Emerging Best Practices for Frontier AI Safety Frameworks

Publication date
July 15, 2026
share

The average AI company scores 22% on our assessment of frontier AI risk management. Adopting practices already in use by its peers would lift that score to 59%. In a recent paper (Stelling et al., 2026), we assessed the Frontier AI Safety Frameworks of the twelve companies that have published one (Amazon, Anthropic, Cohere, G421, Google DeepMind, Magic, Meta, Microsoft, Naver, Nvidia, OpenAI, and xAI) against 65 weighted criteria spanning risk identification, risk analysis and evaluation, risk treatment, and risk governance. For each area, we identified the strongest practices currently on the market. The result is a repository of current best practices, meant to support decision-making and encourage improvements in frontier AI risk management across the industry. It shows a clear and achievable path for improvement, built on measures companies already use. The full best-in-class table, listing the leading practice and the company behind it for each risk management area, can be found here.

Background

There is growing interest in frontier AI risk management. Following the AI Summit in Seoul in 2024, twelve AI companies published voluntary Frontier AI Safety Frameworks, outlining their approaches to managing catastrophic risks from advanced AI systems. These frameworks are now increasingly being used as a basis for external accountability mechanisms and are turning into regulatory infrastructure. In the EU, the AI Act and its voluntary Code of Practice (CoP) prompt signatories to develop a safety and security framework to identify, assess, mitigate, and govern risks when developing AI models with systemic risks. The provisions for general-purpose AI models enter enforcement in August 2026. In the US, the Transparency in Frontier Artificial Intelligence Act (TFAIA) in California, the Responsible AI Safety and Education (RAISE) Act in New York, and the Frontier AI Safety Act in Illinois require large companies developing highly computationally expensive AI models to set up frameworks to manage catastrophic risk. Standardization processes are also under development through NIST and ISO, including standards on evaluation and a preliminary work item on frameworks. However, there remains high uncertainty of what the best practice for these frameworks2 is, limiting clarity, actionability, and practicality across industry and regulators. This work aims to address this gap.

SaferAI has for several years analyzed AI companies’ safety policies and frameworks, using our Frontier AI Risk Management Framework (Campos et al., 2025) as a basis. Building on this work, we recently published a paper summarizing these assessments (Stelling et al., 2026) to identify the companies scoring highest on each dimension and analyzing their distinctive practices. This allows us to highlight current industry best practices for each category of risk management practices, based on the existing practices in the market as of today, which does not necessarily mean that these are best practices overall for each category. 

This work provides a repository to support decision-making and encourage improvements in frontier AI risk management across the industry. We show that a clear and achievable path for improvement exists, built on measures already in use across the industry. Just by adopting existing practices already present among their peers, the average AI company could improve their assessment score from an average of 22% to 59%.

A table with best practices for each risk management area can be found here. The best practices have been distilled from the safety frameworks of 12 companies, including, among others, Anthropic, OpenAI, Meta, and Google DeepMind, and cover 65 individual criteria spanning different aspects of the risk management process.

Areas To Target For Quick Wins

Among the risk management practices, there are some where a small number of companies substantially outperform their peers. These cases reveal clear paths for improvement where one or a few companies demonstrate strong performance, and most others show more emerging practices. Even taking into account the unique circumstances for each company, these should be considered as potential quick wins. These four areas are: Risk Modeling, Key Control Indicators (KCIs), Continuous Monitoring of Key Risk Indicators (KRIs), and Risk Governance. Each of these areas is detailed in the sections below.

Risk Modeling

Risk modeling is a nascent but important component of Frontier AI risk management. It entails constructing detailed, decomposed scenarios that describe how concerning AI capabilities might materialize into real-world harm, considering elements such as threat actors, targets, and bottlenecks in the physical realm. We examine the role of risk modeling in depth in a recent paper (Touzet et al., 2025). The EU AI Act’s Code of Practice Safety and Security Chapter, Commitment 3, Measure 3.3, calls for AI providers to conduct systemic risk modeling. A good example is provided by Meta, who develop risk models for each risk domain, with published threat scenarios. Their Advanced AI Scaling Framework takes a structured, outcomes-led approach: i) identifying a set of catastrophic outcomes; ii) performing threat modeling to identify the potential causal pathways that may be sufficient to realize these outcomes; iii) analyzing the key risk factors, such as model capabilities and propensities, that could lead to realization of each threat scenario; and finally, iv) using evaluations to assess the extent to which a given model could substantially contribute to the threat scenario.

Key Control Indicators for Containment and Deployment

Key Control Indicators (KCIs) are measurable signals representing the effectiveness of mitigations and safeguards. For this, they summarize the effect of combined mitigations into a measurable quantity, making it clearer how much mitigations are actually reducing risk. The EU AI Act’s Code of Practice Safety and Security Chapter, Appendix 3.3, calls for AI providers to assess mitigation effectiveness, and Measure 3.5 calls for post-market monitoring. We analyze best practices for KCIs both for containment (strategies focused on controlling access to the AI system) and deployment (mitigations that allow controlling the potential for misuse).

For containment, G42’s Frontier AI Safety Framework clearly defines Security Mitigation Levels that map to Capability Thresholds. They describe escalating information security measures, allowing for clear links between indicators of risks and the indicators of the effectiveness of mitigations. Google DeepMind similarly provides strong guidance for establishing links between containment/security thresholds and the capability thresholds. Their Frontier Safety Framework defines increasingly stringent containment standards — SL2+, SL3, SL4 — with descriptions of the protection objectives they imply. These levels build on the security levels defined in RAND’s report on securing AI model weights.

For deployment, OpenAI’s Preparedness Framework distinguishes three types of safeguard measures for misuse risks: Robustness (jailbreak resistance), Usage Monitoring (detecting harmful actions), and Trust-based Access (restricting access to vetted users). They provide examples of safeguards that could support these claims, as well as potential efficacy assessments.

Continuous Monitoring of Key Risk Indicators

Key Risk Indicators (KRIs) are measurable signals that act as proxies for how risks are evolving. Current practices provide insight into how the monitoring of KRIs can be improved. The EU AI Act’s Code of Practice Safety and Security Chapter, Measure 3.5 calls for post-market monitoring of capabilities, propensities, affordances, and/or effects. A key aspect in any risk monitoring is proper elicitation. In their Preparedness Framework, OpenAI outlines this well, with multiple elicitation strategies and commitments to match the elicitation efforts of potential threat actors, modeled as high-end adversaries. Its elicitation methods are technically detailed across multiple approaches, treating any one-time capability elicitation as a lower bound rather than a ceiling. This offers a constant assessment of potential emerging risks operated by a diversity of threat actors.

Risk Governance

Finally, risk governance is another practice area where there could be significant improvements by adopting practices already present in the market. The category of risk governance comprises decision-making, advisory, audit, oversight, and culture. Good examples of best practices include the risk owners in xAI’s framework who are responsible for proactively mitigating risks, G42’s risk committee, Anthropic’s Responsible Scaling Officer, OpenAI‘s Safety Advisory Group, and Nvidia’s clear escalation procedures.

Conclusion

Our analysis shows that strong examples already exist across the industry for each dimension of frontier AI risk management. This means that frontier AI developers could already significantly improve their risk management by simply adopting practices similar to those championed by their peers. Doing this will require resources and internal leadership but offers the potential of achievable changes. The case for implementing them now is strong, as regulations increasingly call for state-of-the-art practices across a range of AI risk management domains.

  1. SaferAI contributed to the process of writing G42’s Frontier AI Safety Framework. ↩︎
  2. This goes by different names: Responsible Scaling Policy and Frontier Compliance Framework at Anthropic, Preparedness Framework and Frontier Governance Framework at OpenAI, Frontier Safety Framework at Google DeepMind, Advanced AI Scaling Framework at Meta, Frontier Governance Framework at Microsoft, among others ↩︎

Back to top


Publication date
July 15, 2026
share