← Back to News
ANALYSIS

CAISI Pre-Deployment Eval Agreements Now Cover All Five Frontier Labs — What Changes for the Rest of the Industry

The Center for AI Standards and Innovation finalized pre-deployment evaluation agreements with OpenAI, Anthropic, Google DeepMind, Microsoft, and xAI on 2026-05-18. The agreements are voluntary and unfunded, but they establish a public expectation that every frontier model launch runs a third-party eval before release — a norm that will propagate downstream through enterprise procurement faster than any statute could.

By Michael Eakins•• min read
CAISIAI RegulationFrontier ModelsPre-Deployment EvaluationAI SafetyOpenAIAnthropicGoogle DeepMindMicrosoftxAI

The US Commerce Department's Center for AI Standards and Innovation (CAISI) announced on 2026-05-18 that it has finalized pre-deployment evaluation agreements with all five major frontier AI labs — OpenAI, Anthropic, Google DeepMind, Microsoft, and xAI. The agreements cover capability and safety evaluations of frontier models prior to public release. They are voluntary, unfunded by statute, and carry no licensing or enforcement teeth. They are also the most consequential thing that has happened to AI governance in 2026.

The reason is not the formal structure of the agreements. It is the public expectation they codify: that a frontier model launch is preceded by a structured pre-deployment evaluation conducted by an independent third party, with documented methodology and disclosed results. That expectation will not remain at the frontier-lab tier. It will propagate downstream through enterprise procurement, vendor selection, board-level risk reviews, and insurance underwriting — and it will do so faster than any of the formal regulatory tracks the EU and California have been advancing in parallel.

What the agreements cover

CAISI's published summary describes the agreements as covering pre-deployment evaluation across four dimensions: capability benchmarks (what the model can do), dual-use evaluations (biosecurity, cybersecurity, chemical-weapons uplift), agentic-task evaluations (autonomous action and tool-use capability), and a "national-security" residual category that is deliberately vague in the public materials. Each lab agrees to provide CAISI with model access on a pre-release basis, along with internal evaluation documentation, sufficient to allow CAISI's evaluators to run their own methodology and reach independent conclusions.

The agreements do not include a release veto. A lab that disagrees with CAISI's findings can still ship the model. What the agreement does require is that the lab disclose, post-launch, that an evaluation was performed and at what level the model was assessed against each evaluation category. That disclosure is the lever — not the evaluation itself.

The five labs covered are the same five that have dominated US frontier development for the last two years. Notably absent: Meta, which has not released a model at the frontier capability bar that CAISI uses as a coverage threshold; the major Chinese labs (Alibaba, DeepSeek, Zhipu), which are outside CAISI's jurisdiction; and the second-tier US labs (Mistral, Cohere, Inflection's successor entities) whose models do not meet the agreement's capability threshold but who watch the precedent carefully.

The voluntary-agreement frame is doing real work

Pre-deployment evaluation has been the central unresolved question of US AI governance since the Three-Speed AI Governance analysis in early May. The EU AI Act simplification package, the California SB-1047 successor working through the state legislature, and the UK AI Sandbox framework have all converged on the same architectural question: who decides a frontier model is safe enough to launch, and on what evidence?

The CAISI agreements answer that question with a third option that neither EU nor UK frameworks have produced: a voluntary, methodology-disclosed, results- disclosed pre-launch evaluation regime that does not require statutory authority because it operates on reputation incentives. The labs participate because non-participation is a public statement that they would rather not be evaluated.

That dynamic is more politically durable than it looks. As long as the agreements remain voluntary, the political coalition required to maintain them is small — Commerce, the labs themselves, and a handful of national- security stakeholders. A statutory regime requires Congress; the CAISI agreements require an inter-agency MOU that is already in place.

Downstream propagation: the enterprise-procurement vector

The most overlooked feature of the CAISI agreements is that they create a reference standard that downstream buyers can cite. Before today, an enterprise procurement team asking "What evals did you run before shipping this update?" had no public benchmark to reference. After today, the question has a default answer: "Did you run the CAISI capability and safety evaluations? At what level did you score?"

That question will not stop at the frontier-model boundary. It will reach every vendor selling LLM-powered products to enterprise buyers. Within six months I expect the question to appear in major-bank procurement questionnaires; within twelve months I expect it in defense and healthcare RFPs. The vendors that can answer with a structured eval report — even one they ran themselves, not through CAISI — win the contract. The vendors that say "we tested it internally" lose to the ones that say "here is the capability pass rate, the safety pass rate, and the regression delta for the last three releases."

The architecture for that downstream answer is not complicated. It is roughly the same pipeline a frontier lab uses, scaled down to a customer-facing feature — the kind of pipeline I describe in today's tutorial on building a pre-deployment LLM evaluation gate in TypeScript. Engineering teams that have it ready are positioned for the question; teams that do not are scrambling.

What is not in the agreements

Three things the CAISI agreements explicitly do not cover, each of which is substantive enough to deserve attention:

Open-weight models. The agreements cover model releases by the signing labs. They do not cover models the same labs publish as open weights once those models are downloaded — by definition, post-download distribution is outside CAISI's pre-deployment window. The asymmetry creates a curious incentive: if a closed-weight release would not pass CAISI's bar, opening the weights is a clean way to ship the capability without the disclosure requirement.

Fine-tuned variants. A lab evaluates a base model with CAISI; a customer fine-tunes the model for their domain; the fine-tuned model is not re-evaluated. Most production deployments use fine-tunes. The agreements implicitly acknowledge that fine-tuning is the customer's evaluation responsibility — which is correct, but it pushes the eval-pipeline burden downstream to exactly the customers least likely to have one.

Capability creep through deployment context. A model that passes safety evals in a tightly-scoped support-agent context may behave differently when deployed as a general-purpose assistant with tool-use, memory, and a wider prompt surface. The CAISI evaluation occurs against the model, not the deployment. Production safety is a property of the deployed system, not the weights.

Each of these gaps is a legitimate engineering challenge. None of them invalidates the agreements; they define the boundary of what the agreements solve and where the remaining work falls.

What changes for AI labs not in the agreements

The second-tier US labs and the major non-US labs face a sharper question. If CAISI evaluation becomes a procurement reference standard, models from labs that have not been evaluated face a disclosure asymmetry: the buyer can compare evaluated-model A to evaluated-model B, but unevaluated-model C is a risk that the buyer's procurement function does not know how to price.

That asymmetry will push the second-tier labs toward seeking either (a) voluntary CAISI evaluation participation, (b) participation in equivalent schemes the UK AI Sandbox or EU AI Office will offer in the next twelve months, or (c) a credible internal-evaluation methodology with third-party auditing — which is essentially the private eval-harness pattern that has been spreading through engineering organizations independently of the regulatory conversation.

The non-US labs face the additional problem that CAISI is a US Commerce Department entity. Chinese labs cannot participate; European labs face geopolitical friction even if they wanted to. The likely outcome is parallel evaluation regimes — CAISI in the US, an equivalent EU AI Office program by mid-2027, and a UK AI Sandbox sibling — that converge on similar methodology but operate as separate disclosure tracks. Enterprise buyers will demand evaluation against the regime relevant to their jurisdiction.

Reading the next twelve months

Three signals to watch on the CAISI track between now and Q2 2027:

Disclosure depth. Today the labs commit to disclose that an evaluation occurred and the level at which the model was assessed. They have not committed to publishing the full evaluation reports. The pressure on that will increase as the second and third evaluation cycles produce results that diverge — the labs will be tempted to characterize CAISI findings favorably in marketing, and CAISI will be tempted to publish more detailed counter- evidence. The disclosure ceiling is not stable.

Procurement adoption. The leading indicator is whether large enterprise RFPs start asking for CAISI evaluation status as a vendor-eligibility filter. The first such RFP will appear within six months; the question is whether the pattern propagates beyond defense and financial services into general enterprise tech procurement.

Second-tier participation. Whether Meta, Mistral, or any of the major Chinese labs find a path to participate or to publish an equivalent. The absence of Meta is conspicuous given its public model release cadence; a Meta agreement, if and when it appears, will signal that the CAISI track has won the legitimacy contest against the EU and UK alternatives.

Sources