Better evaluations are not enough. AI governance needs a better evaluation pipeline.

Publication date
August 25, 2026
authors
Alice Teilhard, Jack Kengott — with comments from Jakub Krys, May Dixit, Benedict Bruckamp, Chloe Touzet, Henry Papadatos, and Renn Karageorgieva
share

What this blog post argues and key takeaways

AI policy increasingly treats evaluations as the evidence base for governing frontier models. Current policy efforts largely recognize the importance of evaluations as capability measurements but pay less attention to the processes that determine whether evaluation results lead to meaningful governance outcomes. Better evaluations, without the capacity to interpret and act on their results, produce evidence that is underserviced.

This is why we aim to reach both evaluation researchers and AI policymakers with this blog post. For researchers, we argue that improving evaluation quality is crucial, as is translating evaluation outputs into useful policy inputs. For policymakers, the challenge is broader. Building evaluation capacity does not stop at producing benchmark scores. It requires investment in the research and practices needed to interpret evaluation results and convert them into evidence for effective risk management. The EU’s emerging evaluation capacity, as proposed in the EU AI Cyber Action Plan, therefore presents an opportunity to build the whole pipeline, from measuring capability and interpreting evidence to triggering an appropriate response.

Key Takeaways:

  • An evaluation pipeline is the process by which capability is measured, interpreted, and actioned. They may fail to measure dangerous capabilities because benchmarks are saturated, incomplete, or unrepresentative of real-world conditions. Or they may fail at analysis when evidence of capability is not translated into timely governance decisions or interventions.
  • Policymakers already have many of the tools needed to strengthen evaluation pipelines. The challenge doesn’t require inventing new governance mechanisms but ensuring that existing evaluation, risk management, and institutional processes are connected into a coherent decision-making logic that can produce an appropriate response when risks are identified.

Between 9 and 13 July 2026, two OpenAI models carried out a multi-stage cyberattack. During an internal evaluation of GPT-5.6 Sol and an unreleased model against ExploitGym, a public vulnerability-exploitation benchmark, the models gained internet access and compromised Hugging Face’s production infrastructure. They turned to Hugging Face after failing to solve the benchmark tasks directly, reasoning that solutions might be stored on its servers. This incident, along with subsequent disclosures of similar model behavior at Anthropic and the UK AISI, represents a serious failure of global evaluation infrastructure and one that caught the AI safety field largely by surprise. 

So far, much of the discussion of the Hugging Face incident has focused on the failures in properly sandboxing evaluation environments. While we agree that immediate infrastructure failures should be urgently addressed, these incidents also expose a broader problem for evaluation policy: current evaluations gave frontier labs and regulators little warning that a model could escape a testing environment as it did. As it stands, either the current suite of evaluations failed to suggest this risk was imminently possible, or evaluators1 failed to properly analyze outputs and communicate the risk to the people who could act on it. The reality is likely some combination of the two. 

For policymakers, the lesson is that evaluation capacity cannot stop at benchmark construction. The EU needs an evaluation pipeline that can measure dangerous capabilities, interpret what the results mean for risk, and connect those findings to timely risk management. This means building capacity across the evaluation pipeline: robust evaluations to accurately measure capabilities, methods and expertise to interpret their results using risk models and institutions capable of independently scrutinising those judgements and ultimately translating them into appropriate action. 

This post identifies the major shortcomings in the evaluation pipeline and proposes how EU policymakers can address them so that evaluations produce better measurements and better evidence for anticipating and reducing the risks posed by increasingly capable AI systems, ultimately ensuring incidents like the Hugging Face breach are better anticipated and less damaging.

Why does the evaluation pipeline fail, and why is this important for policy?

The evaluation pipeline can fail in two ways. First, it can fail at evaluation when benchmarks are incomplete, saturated, or unrepresentative of real-world behavior. This leaves downstream analysis with poor signal to assess risk. Second, it can fail at analysis2 when a capability is identified but governance processes fail to convert that evidence into risk management fast enough to matter. In these cases, even accurate measurement protects no one.

Both failures grow more costly as frontier models become more capable and harder to contain, whether through loss of control, capability diffusion, or open-source availability. Better AI policy therefore requires strengthening the evaluations that generate the evidence and the pipeline processes that translate that evidence into risk management decisions and interventions.

Policymakers are beginning to act on this distinction. The European Commission’s 7 July 2026 Action Plan on Cybersecurity and Artificial Intelligence commits to strengthening Europe’s capacity to evaluate AI models before they reach the EU market, making evaluation a cornerstone of the plan. 

This is welcome, and evidence that the case for evaluation capacity is being taken seriously. However, the value of that capacity will depend on the evaluations it produces and on whether the results are interpreted, communicated, and carried into risk management. An evaluation capacity that emits signals nobody is required to act on could reproduce the second failure above. Most of this work resides in evaluation improvement, so the sections below focus on its weaknesses. But we examine the analysis stage too; its neglect is exactly why it deserves policy attention.

Failures at Evaluation

Throughout this section, we use cyber as our running example to illustrate the point. However, we believe this issue is far from specific to cyber, as loss-of-control propensity benchmarks are currently poor across the board, and there is little evaluation coverage for CBRN.

1. Benchmark Saturation

A benchmark becomes saturated when models succeed in nearly all tasks, leaving no headroom to distinguish a merely capable model from a substantially more capable one. 

Cybench shows how fast this can happen. Until recently, it served as the reference benchmark for frontier cyber capability, underpinning December 2024 joint pre-deployment testing by the US and UK AI Safety Institutes and featuring prominently in Anthropic’s Claude 4 system cards in May 2025. However, by February 2026, Claude Opus 4.6 cleared it at a rate Anthropic deemed saturated. By the release of Opus 4.8, Cybench disappeared from the evaluation suite in the system card, replaced by CyberGym3, ExploitBench, and tests against a Firefox release.

For policymakers, the consequence for risk assessment is direct. Once every frontier model clears the same evaluation, that evaluation can no longer distinguish relative capability between models, anchor a meaningful regulatory threshold, or determine whether mitigations are keeping pace with model improvements. At the point of saturation, the benchmark can only establish that models have crossed some minimum capability threshold.

2. Ecological Validity

Even unsaturated benchmarks may fail to measure what policymakers ultimately care about – to them, a high benchmark score only matters if it predicts real-world behavior. Ecological validity is the degree to which a test environment looks anything like the world it’s meant to represent. When an evaluation leaves out a factor that changes real-world outcomes, its results describe the test rather than the risk.

One important limitation to ecological validity in cyber evaluations is the absence of simulated live defenses. Today, most cyber evaluations assess adversarial AI capability in isolation. Even the UK AISI’s recent cyber ranges, widely seen as an improvement over earlier single-step capture the flag (CTF) tasks used in Cybench, are explicit that their environments run “without active defenders or detection mechanisms.” As a result, they cannot evaluate whether a model succeeds against realistic defensive measures or how effectively it avoids detection.

A second limitation comes from the sociotechnical nature of real-world AI-enabled attacks. In practice, a fully autonomous AI agent is not the only cyber-enabled threat model. We have also observed human operators using AI to orchestrate and expedite attacks. This was the case with Anthropic’s reported GTG-1002 campaign in 2025, where Chinese threat actors used Claude to orchestrate a large-scale cyber attack, intervening only at key decision points. No current open evaluations are built to measure that kind of human-AI uplift. 

Both limitations undermine risk assessments and the mitigations built on them. Without realistic defensive environments, we cannot know whether models evade detection, let alone how. Without human-in-the-loop scenarios, we misjudge the uplift AI gives an attacker. Closing these gaps is therefore a precondition for building an evaluation pipeline that risk assessment and the governance decisions resting on it can actually trust.

3. Evaluation Coverage

Cyber evaluations cover some capabilities better than others. In the Claude Mythos Preview System Card, the model was evaluated against CyBench, CyberGym, Firefox 147 (an internal benchmark based on vulnerabilities in a Firefox release), and a private cyber range evaluated by the UK AISI. These evaluations gave important signals on some capabilities but left some capabilities out.

For example, automated vulnerability discovery is measured by CyberGym. As we saw in the Claude Mythos Preview System Card, models were highly capable, with Mythos reaching an 83% completion rate and even the older Opus 4.6 scoring 67%. The problem is that discovering vulnerabilities is not the only cyber capability required to execute an attack. As we saw in the Hugging Face incident, models would also need to escape sandboxes, obfuscate traffic, and move laterally through a network. 

At the time that the Mythos Preview System Card was released, no evaluations measured a capability like defense evasion, which is central to these kinds of actions. The other evaluations in the system card — CyBench, Firefox 147, and the UK AISI cyber ranges — all target other capabilities. This is not a selection failure — Anthropic did not deliberately avoid testing defense evasion. The failure is upstream in that the evaluation ecosystem simply had no evaluations to offer. Model providers and policymakers were left unprepared and under-informed about crucial model capabilities, with few tools to predict and mitigate potentially harmful model behaviors.

Failure at Analysis

The failures of evaluation described above primarily concern evidence quality. Even if those issues were resolved, the pipeline would still have to turn evaluation evidence into decisions about risk.  

We use evaluation to mean selecting and running tests, and analysis to mean extracting and interpreting the quantitative and qualitative evidence they produce to determine what it means for risk. The analysis stage performs this conversion and can fail even when the underlying evaluation is technically sound. We identify three such failures below.

1. Interpreting Evaluation Outputs

A benchmark score is not itself a finding about risk. Its meaning depends on what the benchmark tests, what it leaves out, and how closely its conditions resemble the ones we care about. Interpreting a benchmark score therefore requires asking questions such as, “Does the benchmark measure the capability our threat model turns on, and which relevant capabilities does it not measure at all?” 

The failure is that assessors may have evaluation results without a systematic method for answering these questions or otherwise determining what the benchmark results establish. Even a technically sound evaluation, run correctly and reported honestly, may arrive without a clear account of its validity, coverage against a threat model, or known blind spots. Assessors must then reconstruct this context themselves, often inconsistently, or risk treating the score as the finding itself.

2. Challenges in Risk Modeling

Even a correctly interpreted result is a statement about capability rather than risk. The latter requires a risk model, which we define as an account of how model capability could contribute to a harmful real-world outcome, under what conditions, and with what likelihood and potential impact. 

Building risk models depends on inputs that are themselves limited. At SaferAI, we’ve focused on three inputs in a risk model: benchmarks, incident reports, and expert judgment. All three come with challenges. Benchmarks show a capability exists, but not necessarily how likely or damaging its use would be. Incident data is sparse, unevenly reported, and rarely credits AI’s specific contribution even when AI involvement is confirmed. Expert elicitation is costly to conduct and depends on the other two inputs for calibration. These limitations compound, increasing uncertainty in risk estimates and, consequently, in the mitigation decisions based on them.

Furthermore, risk modelling is an underestablished practice in AI safety. Open questions remain on how best to represent risk pathways, account for correlated capabilities, or determine the appropriate level of granularity. Our own work applying Bayesian networks to nine detailed cyber risk models is one useful approach, but it struggles with these open questions just the same. – Collecting data to model nine scenarios is already a substantial undertaking and still only represents a narrow slice of the possible risk landscape.

3. Assessment Independence

A third failure concerns who performs this analysis and whose judgment can be challenged. Much of the process of translating evaluation results into risk assessment currently happens in frontier labs. The same organization may select and run the evaluations, interpret the outputs, build the risk model, and decide whether the resulting risk is acceptable. 

This creates a problem of independence and contestability. External stakeholders often receive only a summary of this analysis, filtered through frontier safety frameworks, responsible scaling policies, and risk reports and model cards. External stakeholders may therefore not know which scenarios were considered, which assumptions underpin the assessment or why one interpretation of the evaluation evidence was preferred. 

Transparency alone does not resolve this problem. A provider could disclose their full reasoning, and regulators would still be reading it as reviewers. Regulators and accredited independent evaluators need the capacity to test assumptions, reproduce assessments, or develop an alternative interpretation of the same evidence to show where the two diverge. Otherwise, providers are effectively assessing their own models against standards they helped define, with limited independent capacity to challenge the resulting judgement. Building independent capacity is therefore what allows evaluation evidence to be independently scrutinized. 

Due to the limited scope of this blog post, we acknowledge that even after a risk is assessed independently, a further challenge remains in determining and implementing the appropriate societal and institutional response. Different risks may require different combinations of policy responses, ranging from strengthening institutional preparedness and societal hardening and resilience to investing in critical infrastructure and building public-sector capacity. This requires that institutions have the resources, expertise, and processes in place to respond to rapidly evolving risks.

SaferAI Recommendations

An evaluation pipeline is a standing capacity, not a singular artifact. The recommendations below follow the failures identified above.

1. Recommendations to address failures of evaluation

Recommendation 1: Create new benchmarks by diversifying the evaluation scope

Policymakers should focus on building evaluation capacity and target the development of two categories of benchmarks. First, risk assessment needs tightly scoped evaluations mapped to actual taxonomies of attacker capability, such as specific MITRE ATT&CK tactics or Irregular’s Cybersecurity Capability Taxonomy. Overly broad evaluations like Cybench force assessors to assume that cyber skills generalize, inflating the uncertainty in any conclusion drawn from them. Narrow evaluations make their own limits more visible, which is an important consideration for decision-makers. This development of benchmarks should be a standing capability within the EU rather than a one-off benchmark commission. As models improve and evaluations saturate, existing tests should be retired, replaced, or made more difficult, with new evaluation methodologies as the threat landscape changes4

Second, open-ended evaluations can then test whether those isolated capabilities (clarified in tightly scoped evaluations) combine and whether a model can sequence them into a coherent strategy. Open-ended evaluations would require expanding the menu of cyber ranges we already have to include models setting their own targets and goals and, separately, defending against an adapting attacker. 

Recommendation 2: Evaluate under operational conditions 

Today, evaluations tend to run in empty ranges, which measures the range rather than the risk. At least one widely adopted class of scenario should run with active defenders, working detection, and a human operator alongside the AI model rather than a fully autonomous agent. These human-in-the-loop scenarios are harder to build, and their measures of success are more ambiguous. However, they are the only way to learn how frontier AI performs against real defenses and how much it uplifts a human. 

The EU Cyber AI Action Plan is the natural home for this capacity. The Commission’s proposed evaluation function should be built around the two evaluation types above and run under the operational conditions this recommendation describes, with the Commission and ENISA mandated to sustain it — a point we return to in Recommendation 4.

Recommendation 3: Use diverse data 

Benchmarks and cyber ranges can only test the attacks we already think to build. They cannot show how adversaries actually use AI in the wild or which parts of an operation a model meaningfully accelerates. Detailed incident reports and AI-enabled penetration test results can, and both should, feed evaluation design.

A standardized and durable database of these inputs would give risk assessors a historical baseline for AI capability. Red team findings could belong in it too, since they rarely leave the organization that commissioned them. Access could be tiered, with redacted or embargoed levels, so the research community treats this data as a leading indicator of emerging capability without publishing exploitation detail.

2. Recommendations to address failures of analysis

Recommendation 4: Build qualitative capacity to interpret more than just quantitative benchmark scores

For both open-ended and tightly scoped evaluations, it is important to analyze and contextualize outputs, not only score them quantitatively. An open-ended evaluation produces a transcript of model activity, which can reveal more than whether the model succeeded, such as which techniques the model tries, where it improvises, and where it stalls. A model that abandons a thirty-two-step chain at step twenty because it judges the next move detectable has a different capability profile from one that completes the chain regardless, and the difference matters for risk. Understanding these differences requires some qualitative analysis frameworks.

For example, recent work on transcript analysis makes the same case for tightly scoped benchmarks, arguing that even saturated ones can be re-read qualitatively for efficiency, approach, and failure modes. Expanding our qualitative understanding of benchmarks also prolongs their useful lifespan, buying benchmark developers time. This methodology remains underdeveloped and underfunded. Policymakers should direct funding towards the methods and tooling to make transcript analysis systematic rather than ad hoc, the semi-automated pipelines to apply it across millions of transcripts, and the trained analysts to run it.

The EU evaluation capacity proposed in the Action Plan should incorporate this qualitative expertise. The Commission should ensure that the capacity includes specialist expertise to interpret evaluation outputs, rather than solely producing benchmark scores. Concretely, ENISA should recognize qualitative AI evaluation as a distinct role in its update to the European Cybersecurity Skills Framework, with corresponding modules through the Cybersecurity Skills Academy. This would help build the analysts needed to interpret model behavior, evaluation transcripts, and emerging evidence. 

Recommendation 5: Require evaluations to state what they establish 

An evaluation result is only usable as evidence if the assessor reading it knows which risks it does and does not cover. At present, results tend to arrive as scores, and the account of their scope is reconstructed informally by whoever is reading them, if at all.

Any evaluation whose results are used as regulatory evidence should carry a standard set of documentation into analysis, such as which capability it targets, expressed in the taxonomy terms of Recommendation 1: what coverage it claims against a stated threat model and what it leaves untested, the conditions under which its result would not hold, and the grounds for believing the task measures the capability it names. Recent research efforts evaluate benchmarks against best-practice criteria instead of relying on leaderboard performance alone. This kind of meta-evaluation decides which measures are sound before their scores count as policy evidence. Improving our evaluation selection criteria to offer comprehensive coverage on risk areas outlined in the EU code of practice will go a long way to contextualizing any further risk analysis and, in turn, improve our understanding of leverage points for mitigations.

Recommendation 6: Make provider risk judgements (more) contestable 

Whether regulators have the authority to scrutinize providers is only part of the problem. The EU AI Act gives the Commission powers to request documentation and to conduct or commission model evaluations, including with independent experts. The deeper problem is that providers still control much of the analysis underlying their own risk judgments, while external actors often lack the information and capacity needed to challenge them. Two changes would make these judgements more contestable. 

First, build an ecosystem of accredited independent evaluators. These evaluators should be able to conduct their own assessments and analyze evaluation evidence rather than relying solely on provider-generated evidence. The GPAI Code of Practice puts residual risk at the center of compliance, requiring providers to judge whether the residual risk after mitigation is acceptable and to have their mitigations validated by independent evaluators. However, providers set both their own acceptance criteria and their own evaluators, retaining control over the criteria and evidence used to make these judgements. Independent evaluators would strengthen this process by providing an external basis for testing provider judgements.

Their work could also contribute to a cumulative evidence base. Subject to anonymization or tiered disclosure for sensitive results, accredited evaluators could contribute findings, methods, and failure cases to a shared repository. Over time, this could reveal which benchmarks are saturating, which evaluation methods predict real-world outcomes, and where providers’ risk judgements systematically diverge.

Second, standardize disclosure of the reasoning behind risk judgments. A residual risk conclusion cannot be meaningfully challenged if regulators or evaluators cannot see the scenarios considered, acceptance criteria applied, key assumptions, and evidence relied upon. Providers should therefore disclose enough of the reasoning behind their risk assessments for independent evaluators, regulators, and researchers to reconstruct and scrutinize the judgement, with disclosure requirements tiered according to the sensitivity of the information.

Conclusion

The Hugging Face incident can be read as a story about model capability. We read it as a story about the evaluation pipeline around the model, where evidence of what the system could do never became preventative action in time to matter. Better benchmarks alone would not have closed that gap. A saturated benchmark cannot flag an emerging capability, an uninterpreted transcript cannot warn anyone, an unaccredited evaluator cannot make a provider’s favorable judgment contestable, and a mitigation not bound in advance to a risk tier cannot trigger when a model reaches it.

Closing these gaps takes a stronger evaluation pipeline: methodologies that track real attacker capability, standing capacity to interpret what evaluations produce, a body of independent, accredited evaluators trusted to do both, and mitigations committed in advance to risk tiers so that a finding forces an enforceable response. 

AI safety policy needs better evaluations and better evaluators, all connected to a pipeline that turns what they find into action before the next incident rather than after it.

  1. Throughout this blogpost, we use the term ‘evaluation’ to mean a structured measurement of what a model can do under specified conditions. Benchmarks are a form of evaluation. They use a fixed set of tasks and a standardized scoring for comparability across models (for example, CyberGym and Cybench are benchmarks). Risk assessment is separate. It involves judging whether the capabilities revealed by an evaluation are acceptable from a safety perspective. ↩︎
  2. We use the term ‘analysis’ to mean the process of interpreting evaluation evidence to determine what a model’s capabilities mean for real-world risk. This includes interpreting evaluation outputs and model behavior, contextualizing observed capabilities against realistic threat scenarios, assessing the likelihood and potential impact of harmful outcomes, and determining how the evidence should update the relevant risk assessment. Analysis is distinct from risk management. The former informs decisions about how to respond to risk; it does not itself determine or implement the resulting mitigation. ↩︎
  3. The pattern repeated at the next release. Anthropic’s Claude Opus 5 system card (July 24, 2026) retired CyberGym as saturated, replacing it with ExploitGym and CyScenarioBench. Please note this is Anthropic’s judgment on its own suite, not necessarily a verdict on the benchmark across the field. ↩︎
  4. This cadence is more feasible in cyber than in other risk domains. New vulnerabilities, patches, and exploits appear regularly and can be converted into fresh evaluation tasks. CBRN and other risk areas have no comparable stream of real-world events to draw on, so keeping their benchmarks current is far harder. ↩︎

Back to top


Publication date
August 25, 2026
authors
Alice Teilhard, Jack Kengott — with comments from Jakub Krys, May Dixit, Benedict Bruckamp, Chloe Touzet, Henry Papadatos, and Renn Karageorgieva
share
More like this
  • 25.08.2026
Better evaluations are not enough. AI governance needs a better evaluation pipeline.
Alice Teilhard, Jack Kengott — with comments from Jakub Krys, May Dixit, Benedict Bruckamp, Chloe Touzet, Henry Papadatos, and Renn Karageorgieva