OpenAI Wants to Write the Rules for Its Own Safety Audits

OpenAI Wants to Write the Rules for Its Own Safety Audits

OpenAI's new principles for third-party assessments aim to shape how frontier AI safety is evaluated before regulators finalize their own rules. The framework benefits OpenAI strategically while raising unresolved questions about evaluator independence and access.

OpenAI published a framework on September 22, 2026, outlining priorities and principles for third-party AI safety assessments of frontier models. The move positions the lab as a standard-setter rather than a standard-taker, arriving as the EU AI Act's systemic-risk provisions and NIST's evaluation guidance are still being operationalized. What changed is not the technology β€” it is who gets to define what 'independent assessment' means.
  • OpenAI published priorities and principles for third-party AI safety assessments of frontier models and safeguards on September 22, 2026.
  • The framework arrives as the EU AI Act's systemic-risk rules and NIST evaluation guidance are still being defined, making OpenAI a de facto standard-setter.
  • The core tension: independent assessment requires access to model internals, but labs control that access β€” making true independence structurally difficult.
  • Who benefits most is not the safety community but OpenAI itself, which gains a first-mover advantage in defining evaluation norms.
OpenAI published a document titled "Priorities and principles for effective third party assessments" on September 22, 2026, outlining its vision for how independent evaluators should assess frontier models and safeguards. OpenAI said the framework covers rigor, security, and independence β€” three words that sound neutral but carry enormous strategic weight when the publisher is also the entity being assessed.

What Did OpenAI Actually Publish?

According to OpenAI News, the document lays out priorities and principles for "rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards." The summary is deliberately sparse β€” no specific evaluators are named, no access protocols are detailed, and no enforcement mechanism is proposed. That sparseness is itself the story. OpenAI is staking out philosophical territory, not operational commitments. The timing matters. The EU AI Act's obligations for general-purpose AI models with systemic risk entered application in August 2025, and the EU AI Office has been building out its evaluation infrastructure since. NIST's AI Safety Institute has published its own evaluation frameworks. OpenAI's document lands in that gap β€” after regulation exists on paper but before enforcement mechanisms are fully operational. OpenAI is trying to fill the vacuum with its own definitions.

Why Does 'Independent' Assessment Depend on Lab Cooperation?

This is the structural contradiction at the heart of the document. Independent evaluators β€” organizations like METR (Model Evaluation and Threat Research) and Apollo Research β€” need access to model weights, inference endpoints, and training data to conduct meaningful assessments. That access is controlled entirely by the labs. An evaluator who depends on OpenAI for API keys and compute credits is not independent in any meaningful sense. OpenAI's principles acknowledge this tension without resolving it. The document calls for security and rigor, which are lab-favorable constraints: they justify restricting access to vetted evaluators, controlling the terms of engagement, and limiting what can be published. Security concerns become a legitimate reason to keep evaluation opaque. The safety community has flagged this dynamic for years, but OpenAI's framework formalizes it.
OpenAI Wants to Write the Rules for Its Own Safety Audits

Who Are the Winners and Losers in This Framework?

OpenAI is the clearest winner. By publishing principles first, it sets the vocabulary that regulators, journalists, and competitors will use. When the EU AI Office drafts its own evaluation standards, OpenAI's document becomes a reference point β€” whether adopted, adapted, or rejected. That is agenda-setting power. Independent evaluators are ambiguous winners. METR and Apollo Research gain legitimacy from being named in the ecosystem, but their leverage depends on access terms they do not control. If OpenAI's principles become the industry standard, evaluators become contractors rather than auditors. Smaller AI labs are losers. If third-party assessment becomes a de facto requirement, the compliance cost falls hardest on organizations without OpenAI's legal and policy teams. Anthropic, Google DeepMind, and Meta can afford to shape or resist these norms. A twenty-person startup cannot.
DimensionOpenAI's FrameworkEU AI Act ApproachNIST AI Safety Institute
Who defines 'independent'Lab-influenced principlesRegulatory mandateVoluntary standards
Access to model internalsLab-controlled, security-gatedMandated for systemic-risk modelsNegotiated, voluntary
Enforcement mechanismNone proposedFines up to 3% global revenueNone β€” guidance only
Publication of resultsNot specifiedRequired summariesEvaluator discretion
Speed of implementationImmediate (self-imposed)Phased through 2027Ongoing
VerdictFastest but least bindingSlowest but enforceableAdvisory only

What Does This Mean for the Regulatory Landscape?

The EU AI Office has been explicit that it will require independent evaluation for models classified as systemic risk. OpenAI's framework is best understood as a bid to shape how those evaluations are conducted before the Office publishes detailed technical standards, expected in 2027. If OpenAI's definitions of "rigorous" and "secure" are adopted, the compliance burden shifts in its favor. The Financial Times reported in 2025 that AI labs were increasingly hiring former regulators to shape emerging rules β€” a pattern OpenAI exemplifies. This document is the policy equivalent of that hiring strategy: it is not lobbying, but it achieves similar ends by defining terms.

Does the Safety Community Buy It?

Early reactions from independent evaluators have been cautious. METR has publicly emphasized that meaningful assessment requires unmediated access to model capabilities, not lab-curated demonstrations. Apollo Research has made similar arguments. OpenAI's principles do not commit to unmediated access, leaving the core question unresolved. The document's silence on publication rights is equally telling. If evaluators cannot publish their findings, third-party assessment becomes a private audit with no public accountability. That is not independence β€” it is outsourcing with a confidentiality agreement.
OpenAI's third-party assessment principles are a strategic preemption play dressed in safety language. The company is not wrong to publish them β€” standards need to come from somewhere β€” but readers should be clear-eyed about the incentive structure. OpenAI benefits most when it defines the terms of its own evaluation. In the short term, expect OpenAI to point to this document as evidence of good-faith safety commitment in regulatory conversations. In the long term, the real test is whether the EU AI Office adopts these principles or writes its own. My prediction: the EU AI Office will publish technical standards in 2027 that diverge from OpenAI's framework on access and publication requirements, forcing labs to comply with stricter rules than they would prefer. The losers here are smaller labs and independent evaluators who lack the leverage to negotiate access terms. The winners are OpenAI and, paradoxically, the safety community β€” because even a flawed framework raises the baseline for what counts as assessment.

Predictions

1. The EU AI Office will publish technical standards for systemic-risk model evaluation by Q3 2027 that require publication of evaluation summaries, diverging from OpenAI's silence on publication rights. 2. At least one major independent evaluator (METR or Apollo Research) will publicly criticize access restrictions within 12 months, citing inability to verify claims without unmediated model access. 3. Anthropic will publish its own third-party assessment framework by mid-2027, adopting a more permissive access stance to differentiate from OpenAI.
  1. August 2025
    EU AI Act systemic-risk provisions take effect

    Obligations for general-purpose AI models with systemic risk enter application, creating demand for evaluation standards.

  2. September 2026
    OpenAI publishes third-party assessment principles

    OpenAI outlines priorities and principles for rigorous, secure, and independent assessments of frontier models.

  3. 2027 (expected)
    EU AI Office technical standards

    The EU AI Office is expected to publish detailed technical standards for systemic-risk model evaluation, testing whether OpenAI's framework is adopted or rejected.

Who Controls Frontier AI Safety Evaluation? (estimated influence share)

Article Summary

  • OpenAI's September 22, 2026 framework is a standard-setting play, not a binding commitment β€” it defines terms without accepting enforcement.
  • The structural contradiction: independent assessment requires lab-controlled access, making true independence impossible without regulatory mandate.
  • OpenAI wins the agenda-setting battle; smaller labs and independent evaluators lose leverage.
  • The EU AI Office's 2027 technical standards will be the real test of whether OpenAI's framework holds.
  • Watch for evaluator pushback on access and publication rights as the key indicator of whether this framework has teeth.

Source and attribution

OpenAI News
Priorities and principles for effective third party assessments

Discussion

Add a comment

0/5000
Loading comments...