Silent Model Updates Are Breaking Your EU AI Act Risk Assessment

July 13, 2026 eu-ai-act compliance model-versioning foundation-models attestation

OpenAI updated GPT-4-turbo's training data cutoff from April 2023 to December 2023 without issuing a new model ID. Anthropic regularly updates Claude models within a named version — claude-sonnet-4-6 today is not bit-for-bit identical to claude-sonnet-4-6 three months ago. Google has retrained Gemini models while keeping the same API endpoint identifiers.

If your production AI system is built on any of these, your EU AI Act Article 9 risk assessment is covering a system that may no longer exist in the form you assessed.

The Article 9 gap

EU AI Act Article 9 requires providers of high-risk AI systems to maintain a risk management system that is "a continuous iterative process run throughout the entire lifecycle." Article 9(2)(c) requires "examination, testing and validation procedures" — procedures that are implicitly tied to the AI system as it exists at test time.

Article 17, which governs Quality Management Systems, requires documentation of "the AI system design, development process and testing methodology." A QMS that references a risk assessment performed against model version X is not a valid QMS if version X is no longer what runs in production.

Neither article defines what constitutes a "model version change" at the foundation model level. Article 83 addresses substantial modifications at the system level, but is silent on what happens when a third-party foundation model component updates beneath your system. Recital 66 suggests that the deployer who builds on a third-party model takes on responsibility for its behavior — including changes introduced by the upstream provider.

The compliance obligation falls on you, not the provider.

What providers expose

Most providers return the following in an API response:

  • The model ID you requested (e.g., gpt-4o, claude-sonnet-4-6)
  • A response ID scoped to the request, not the model version
  • Token usage counts
  • Finish reason

What they rarely expose:

  • The specific training checkpoint that processed your request
  • Any signal that the model has changed since your last call
  • A version fingerprint you can use to verify consistency across calls

Anthropic's dated model IDs come closest to version pinning — they're specific enough that a behavioral change should, in theory, require a new ID. OpenAI's identifiers are less reliable: gpt-4-turbo has resolved to at least three distinct internal versions. Even when you pin to a specific identifier, you're trusting the provider's representation. There's no mechanism for you to independently verify that the model responding today is cryptographically identical to the model you assessed six months ago.

Free tier: 500 proofs/month, no credit card required.

See plans & get free key

The binding problem

Suppose you capture the model ID from every API response. You have a log entry:

2026-07-13T09:14:23Z | model=claude-sonnet-4-6 | request_id=req_01abc | output_length=847

This log entry proves that you believe you called claude-sonnet-4-6. It does not prove:

  1. That the model ID accurately represents the version that processed your request
  2. That the log entry hasn't been modified since it was written
  3. That the stored output matches what was actually returned

Gap 1 is a provider accountability problem you can't fully close. Gaps 2 and 3 are your infrastructure problem — and both are addressable.

What model version attestation requires

Capture at response time, not request time. Record the full model identifier from the provider response, not what you sent in the request. Some providers return different model IDs than were requested due to aliasing, auto-routing, or fallback behavior. The response field is the authoritative record.

Bind output to model identifier. Hash the tuple (model_identifier, sha256(request_body), sha256(response_body)) at the API boundary, before the output enters your application logic. A minimal attestation record looks like:

{
  "ts": "2026-07-13T09:14:23.441Z",
  "model_id": "claude-sonnet-4-6",
  "req_hash": "e3b0c44298fc1c149afb",
  "resp_hash": "a665a45920422f9d417e",
  "attestation": "sha256:req_hash||resp_hash||ts",
  "sig": "..."
}

This fingerprint proves a specific output was produced in the context of a specific model identifier. It doesn't prove the identifier is accurate — but it proves you recorded what the provider told you, and that the output hasn't changed since.

Use tamper-evident storage. A SQL table with an auto-increment primary key is not tamper-evident. An append-only log where each entry includes the hash of the previous entry is. The distinction matters under Article 9(4), which requires that risk management records support retrospective review.

Monitor provider changelogs. Subscribe to provider model update announcements. When a model updates, flag all downstream risk assessments that reference that model for review under your Article 17 QMS. This is a process control, not a technical one, but Article 9's "continuous iterative" requirement makes it mandatory.

Version provider documentation alongside your assessments. Model cards, system cards, and technical reports don't constitute cryptographic proof, but they do support the Article 11 technical documentation requirement. Capture them at assessment time and version them with the assessment itself.

The proxy layer as attestation point

The cleanest implementation sits at the API proxy — a component between your application and the model provider. This layer intercepts every request and response, extracts model version metadata, computes the attestation fingerprint, and writes the record before returning the response to your application.

This is what ArkForge's Trust Layer implements. Rather than requiring each application team to add attestation logic independently, the Trust Layer enforces it structurally: every call through the proxy produces a signed attestation receipt containing the model identifier, input and output hashes, and a timestamp anchored to a tamper-evident log. Coverage is total because it's not opt-in — there's no call path that bypasses the attestation step.

Application-level logging is inconsistently implemented across teams and systems. Proxy-level attestation is structural. For multi-model environments — where a single workflow might call a reasoning model, an embedding model, and a vision model in sequence — the proxy is the only place you can guarantee uniform coverage.

Before December 2027

High-risk AI system obligations under the EU AI Act apply from 2 December 2027 (2 August 2028 for AI embedded in regulated products). If your system falls under Article 6 or Annex III, the following steps are not optional:

  1. Inventory foundation model dependencies. Include indirect usage — models called through RAG pipelines, tool-calling agents, or orchestration layers. Teams often don't know which models their abstractions resolve to.

  2. Anchor existing risk assessments to model versions. If your risk documentation doesn't state which model version was assessed, it has no date relative to what's running. Revise it before the obligation date.

  3. Implement output fingerprinting at the API boundary. The model ID in application logs is necessary but not sufficient. You need the output hash bound to the model identifier at the moment of response.

  4. Build a model update detection process. Treat a provider's model update as a potential QMS-triggering event. Your Article 17 quality management system should have a defined response procedure for upstream model changes.

  5. Review Article 11 technical documentation. If your documentation describes model behavior without a version anchor, it describes a system that may not exist. Revise before a conformity assessment.

The asymmetry

Silent model updates are a provider behavior you don't control. But EU AI Act compliance obligations fall on the deployer — the organization building systems on top of these models. The practical response is not to demand better versioning from providers, though that would help. It is to build an attestation layer at the point where provider behavior enters your system.

A signed attestation record stating "this output came from model identifier X, with this input fingerprint, at this timestamp" doesn't close the provider accountability gap. It does give you a defensible record of what your system actually processed — which is what Article 9's continuous risk management process requires, and what a regulator asking about a decision made eighteen months ago will want to see.


Prove it happened. Cryptographically.

ArkForge generates independent, verifiable proofs for every API call your agents make. Free tier included.

Compare plans → or get free key directly