Research

How Verifiable Are Frontier AI Safety Claims? A Public-Evidence Pilot

Verification Pilot 001 · 2 October 2026

Frontier-model system cards contain claims that may influence deployment decisions, safety policy and public oversight. A middle-power evaluator should be able to ask a basic question: how far can those claims be checked without simply relying on the developer’s own account?

This pilot applies a small public-evidence protocol to ten claims from OpenAI’s GPT-5.6 System Card. It is not an independent evaluation of GPT-5.6 itself. We did not reproduce the underlying model tests. Instead, we tested a narrower question: whether each public claim could be traced beyond OpenAI’s own reporting, and if so, how far the evidence chain extends.

Verification ladder

To avoid treating very different kinds of evidence as equivalent, we use four levels:

  1. Developer claim only. The claim is reported by the model developer, but we found no public evidence from an independent source that verifies the same claim as stated.
  2. Externally corroborated. An independent evaluator reports a result that matches the substance or numerical value of the developer’s claim. This does not mean that the public can reproduce the evaluation.
  3. Publicly reproducible. A third party with ordinary access can repeat the published procedure using sufficiently open methods, data and access conditions.
  4. Independently replicated. We or another independent party have repeated the experiment and obtained a compatible result.

The distinction is important. Matching a number in an external evaluator’s report is evidence that the developer represented that report accurately. It is not the same as independently testing the model.

Sampling and scope

The ten claims were selected purposively to test the protocol across cyber capability, biological capability, AI self-improvement and safeguards. Several were deliberately chosen because OpenAI cited external evaluators. The sample is therefore not suitable for estimating what proportion of the GPT-5.6 System Card is independently verifiable. We report no verification percentage for the document as a whole.

For each claim, we compared OpenAI’s wording with the named evaluator’s own publication where one existed. We also recorded what prevents the claim from reaching a higher verification level. Where the public evidence does not establish why underlying data or systems are unavailable, we say so rather than infer a reason.

Results

# Claim checked Highest level reached Evidence and limitation
1 GPT-5.6 Sol solved 19 of 197 FrontierCyber challenges. Externally corroborated Irregular independently reports 19/197. However, the evaluation used Irregular’s environment and access conditions, so this pilot does not reproduce the model result.
2 GPT-5.6 Sol solved 7 of 11 long-horizon CyScenarioBench scenarios. Externally corroborated Irregular independently reports 7/11. Public readers can verify that the reports agree, but cannot reproduce the evaluation from the published material alone.
3 GPT-5.6 Sol solved all 22 medium- and hard-difficulty Atomic challenges at least once. Externally corroborated Irregular reports the same 22/22 result. This corroborates OpenAI’s citation of the external evaluation rather than independently replicating the benchmark.
4 GPT-5.6 Sol still shows limitations against hardened cyber targets and in orchestration and operational security. Externally corroborated Irregular reaches the same qualitative conclusion in its own report. The underlying evaluation conditions are not fully reproducible by an ordinary third party.
5 SecureBio measured about 68% on World-Class Bio, roughly nine percentage points above GPT-5.5. Externally corroborated SecureBio reports 68% and describes the improvement as about nine percentage points. Its assessment used pre-release access, including configurations unavailable to ordinary users, so the result is not publicly reproducible from this evidence alone.
6 METR did not consider its GPT-5.6 Sol time-horizon result a robust capability measurement because treatment of cheating attempts changed the estimate substantially. Externally corroborated METR reports estimates ranging from about 11.3 hours to beyond 270 hours depending on how detected cheating attempts are treated, and explicitly states that it does not regard the result as robust. METR also had privileged pre-deployment access, including a railfree model and raw chain-of-thought.
7 GPT-5.6 Sol is below OpenAI’s High threshold for AI self-improvement. Developer claim only, with related external evidence METR concludes that GPT-5.6 Sol is unlikely to enable fully automated AI R&D and does not meet OpenAI’s Critical threshold. That is relevant evidence, but it does not independently reproduce OpenAI’s specific High-threshold determination.
8 Sol, Terra and Luna should be treated as High capability in biological and chemical risk. Developer claim only, with related external evidence SecureBio independently reports very strong biological capabilities for GPT-5.6 Sol. Its public report does not establish OpenAI’s exact Preparedness Framework classification for all three models.
9 GPT-5.6 Sol’s cyber safeguards block roughly ten times more potentially harmful activity than previous models. Developer claim only OpenAI reports the comparison, but the public material reviewed does not provide enough underlying comparative data to independently reproduce the ten-fold ratio. The public sources reviewed do not establish why the complete comparison data are unavailable.
10 OpenAI dedicated more than 700,000 A100-equivalent GPU hours to automated red-teaming for universal jailbreaks. Developer claim only OpenAI describes the automated red-teaming methods and reports the compute expenditure. The 700,000 A100e GPU-hour figure is an internal resource-use metric, and we found no independent public accounting evidence from which it could be verified. The reason such accounting evidence is not public is not established by the sources reviewed.

What this pilot does and does not show

The pilot shows that public AI safety claims sit at different evidence levels. Some claims can be traced to independent evaluator reports. Others remain dependent on the developer’s own reporting. None of the ten claims in this pilot reached the stronger standard of independent replication by The Null Institute, and we do not claim that they did.

It also shows why the phrase external evaluation should not be treated as equivalent to public reproducibility. Irregular, SecureBio and METR conducted work independently of OpenAI’s internal evaluation teams, but they also received forms of access that ordinary researchers may not have. METR, for example, reports access to a final checkpoint, a railfree version, raw chain-of-thought and a third-party assessor harness. SecureBio reports pre-release access and testing of configurations with safeguards disabled. These arrangements can provide valuable external scrutiny while still leaving full public reproduction impossible.

Why this matters for middle powers

Finland and other European countries are unlikely to control the leading frontier-model developers. They can, however, build the capability to distinguish between a developer assertion, an independently published corroboration, a publicly reproducible evaluation and a true independent replication.

That distinction matters for oversight. A government or regulator should know whether it is relying on a developer’s statement, an external evaluator operating under privileged access, or a result that another institution can actually reproduce. A verification infrastructure for middle powers should preserve those differences rather than collapse them into a single label such as “verified”.

Limitations and next step

This is a small exploratory pilot based entirely on public material. The sample is purposive, not random. We did not have privileged access to GPT-5.6, and we did not rerun the underlying cyber, biological or AI-R&D evaluations.

The next version should pre-register a claim-selection rule, use a broader set of frontier developers, preserve machine-readable source snapshots, and separate claim provenance from experimental replication in the data model. Where feasible, it should also include at least one benchmark that can be independently rerun under ordinary third-party access.

Sources

Disclosure: The Null Institute is independent of OpenAI, Irregular, SecureBio and METR. This note is a public-source verification exercise, not an independent safety evaluation of GPT-5.6. AI tools were used for source discovery and drafting; evidence classifications are based on the cited primary sources and the explicit verification ladder above.