OpenAI Wants to Write the Rules for Its Own Safety Audits
OpenAI's new principles for third-party assessments aim to shape how frontier AI safety is evaluated before regulators finalize their own rules. The framework benefits OpenAI strategically while raising unresolved questions about evaluator independence and access.
- OpenAI published priorities and principles for third-party AI safety assessments of frontier models and safeguards on September 22, 2026.
- The framework arrives as the EU AI Act's systemic-risk rules and NIST evaluation guidance are still being defined, making OpenAI a de facto standard-setter.
- The core tension: independent assessment requires access to model internals, but labs control that access β making true independence structurally difficult.
- Who benefits most is not the safety community but OpenAI itself, which gains a first-mover advantage in defining evaluation norms.
What Did OpenAI Actually Publish?
According to OpenAI News, the document lays out priorities and principles for "rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards." The summary is deliberately sparse β no specific evaluators are named, no access protocols are detailed, and no enforcement mechanism is proposed. That sparseness is itself the story. OpenAI is staking out philosophical territory, not operational commitments. The timing matters. The EU AI Act's obligations for general-purpose AI models with systemic risk entered application in August 2025, and the EU AI Office has been building out its evaluation infrastructure since. NIST's AI Safety Institute has published its own evaluation frameworks. OpenAI's document lands in that gap β after regulation exists on paper but before enforcement mechanisms are fully operational. OpenAI is trying to fill the vacuum with its own definitions.Why Does 'Independent' Assessment Depend on Lab Cooperation?
This is the structural contradiction at the heart of the document. Independent evaluators β organizations like METR (Model Evaluation and Threat Research) and Apollo Research β need access to model weights, inference endpoints, and training data to conduct meaningful assessments. That access is controlled entirely by the labs. An evaluator who depends on OpenAI for API keys and compute credits is not independent in any meaningful sense. OpenAI's principles acknowledge this tension without resolving it. The document calls for security and rigor, which are lab-favorable constraints: they justify restricting access to vetted evaluators, controlling the terms of engagement, and limiting what can be published. Security concerns become a legitimate reason to keep evaluation opaque. The safety community has flagged this dynamic for years, but OpenAI's framework formalizes it.
Who Are the Winners and Losers in This Framework?
OpenAI is the clearest winner. By publishing principles first, it sets the vocabulary that regulators, journalists, and competitors will use. When the EU AI Office drafts its own evaluation standards, OpenAI's document becomes a reference point β whether adopted, adapted, or rejected. That is agenda-setting power. Independent evaluators are ambiguous winners. METR and Apollo Research gain legitimacy from being named in the ecosystem, but their leverage depends on access terms they do not control. If OpenAI's principles become the industry standard, evaluators become contractors rather than auditors. Smaller AI labs are losers. If third-party assessment becomes a de facto requirement, the compliance cost falls hardest on organizations without OpenAI's legal and policy teams. Anthropic, Google DeepMind, and Meta can afford to shape or resist these norms. A twenty-person startup cannot.| Dimension | OpenAI's Framework | EU AI Act Approach | NIST AI Safety Institute |
|---|---|---|---|
| Who defines 'independent' | Lab-influenced principles | Regulatory mandate | Voluntary standards |
| Access to model internals | Lab-controlled, security-gated | Mandated for systemic-risk models | Negotiated, voluntary |
| Enforcement mechanism | None proposed | Fines up to 3% global revenue | None β guidance only |
| Publication of results | Not specified | Required summaries | Evaluator discretion |
| Speed of implementation | Immediate (self-imposed) | Phased through 2027 | Ongoing |
| Verdict | Fastest but least binding | Slowest but enforceable | Advisory only |
What Does This Mean for the Regulatory Landscape?
The EU AI Office has been explicit that it will require independent evaluation for models classified as systemic risk. OpenAI's framework is best understood as a bid to shape how those evaluations are conducted before the Office publishes detailed technical standards, expected in 2027. If OpenAI's definitions of "rigorous" and "secure" are adopted, the compliance burden shifts in its favor. The Financial Times reported in 2025 that AI labs were increasingly hiring former regulators to shape emerging rules β a pattern OpenAI exemplifies. This document is the policy equivalent of that hiring strategy: it is not lobbying, but it achieves similar ends by defining terms.Does the Safety Community Buy It?
Early reactions from independent evaluators have been cautious. METR has publicly emphasized that meaningful assessment requires unmediated access to model capabilities, not lab-curated demonstrations. Apollo Research has made similar arguments. OpenAI's principles do not commit to unmediated access, leaving the core question unresolved. The document's silence on publication rights is equally telling. If evaluators cannot publish their findings, third-party assessment becomes a private audit with no public accountability. That is not independence β it is outsourcing with a confidentiality agreement.Predictions
1. The EU AI Office will publish technical standards for systemic-risk model evaluation by Q3 2027 that require publication of evaluation summaries, diverging from OpenAI's silence on publication rights. 2. At least one major independent evaluator (METR or Apollo Research) will publicly criticize access restrictions within 12 months, citing inability to verify claims without unmediated model access. 3. Anthropic will publish its own third-party assessment framework by mid-2027, adopting a more permissive access stance to differentiate from OpenAI.- August 2025EU AI Act systemic-risk provisions take effect
Obligations for general-purpose AI models with systemic risk enter application, creating demand for evaluation standards.
- September 2026OpenAI publishes third-party assessment principles
OpenAI outlines priorities and principles for rigorous, secure, and independent assessments of frontier models.
- 2027 (expected)EU AI Office technical standards
The EU AI Office is expected to publish detailed technical standards for systemic-risk model evaluation, testing whether OpenAI's framework is adopted or rejected.
Who Controls Frontier AI Safety Evaluation? (estimated influence share)
Article Summary
- OpenAI's September 22, 2026 framework is a standard-setting play, not a binding commitment β it defines terms without accepting enforcement.
- The structural contradiction: independent assessment requires lab-controlled access, making true independence impossible without regulatory mandate.
- OpenAI wins the agenda-setting battle; smaller labs and independent evaluators lose leverage.
- The EU AI Office's 2027 technical standards will be the real test of whether OpenAI's framework holds.
- Watch for evaluator pushback on access and publication rights as the key indicator of whether this framework has teeth.
Source and attribution
OpenAI News
Priorities and principles for effective third party assessments
Discussion
Add a comment