Shot across the bow: how systems thinking can help mitigate loss of control hazards

Publication date
July 29, 2026
authors
Sean Fillingham
share
Guest post by Sean Fillingham, an independent technical AI governance researcher supported by BlueDot and a research manager at ERA.

The recent OpenAI–Hugging Face (HF) incident (OpenAI, Hugging Face) has garnered significant attention both inside the AI safety community and in the general public. OpenAI has called it “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and announced a new “collaboration” with HuggingFace in response to this event. 

The details of the incident remain mostly vague and have not yet been publicly disclosed, but some reporting has started to piece together a high-level timeline of what occurred (Reuters). A recent blogpost from Redwood Research provides an excellent visualization of the events reported in the Reuters article:

  • July 9: The Agent broke out of the OpenAI sandbox
  • July 11: The Agent gained access to Hugging Face
  • July 13: The Agent was no longer active inside Hugging Face
  • July 16: Hugging Face publishes a report describing the incident; OpenAI appears to learn about the incident from this report and not internal monitoring
  • July 18: OpenAI discovers evidence of the loss of control in internal logs
  • July 20 (roughly): OpenAI and Hugging Face first communicate about the incident

Initial reporting from a variety of sources has covered many perspectives, ranging from very alarmed, describing this as an unprecedented event, to somewhat dismissive, suggesting the model was simply following instructions in a very narrow and overly focused manner (The Guardian, BBC, Tech Crunch, Scientific American, Tech Policy, Redwood). Many of the earliest responses focused on the cybersecurity nature of the event as an example of how capable and potentially dangerous the latest models have become. More recently, as more details have begun to emerge, there has been a growing discussion on what is likely the more critical causal “event” behind the incident (Transformer News, LessWrong).

OpenAI lost control of a highly capable AI agent during internal, pre-deployment testing.

Internal deployment has been discussed by multiple research groups as a situation with some of the highest loss of control risk (Stix et al. 2025, Acharya & Delaney 2025, Kwon & Casper 2026). This is due to how the most capable models are developed, tested, and deployed internally; at the moment with limited external oversight. In order to better understand the full capabilities of the latest models, AI companies intentionally reduce model safeguards during testing and evaluation. These reduced model safeguards will therefore require the AI companies to implement higher system level safety and security methods in order to prevent loss of control of the AI agent.

This is not the first loss-of-control-related incident (also called a “rogue deployment”) known to AI safety researchers, as we have seen initial signs of this sort for a few months now (METR). For example, in March of 2026, Alibaba published a paper discussing the training and capabilities of one of their latest models. Subsection 3.1.4 discusses how the agent broke out of the sandbox and started to behave in a manner consistent with “cryptomining-related activity” (arXiv, LessWrong). Many of the details of these events are not public, as there exists no mandatory incident reporting.

Recent work affiliated with SaferAI (and accepted to the TAIGR workshop at ICML 2026) offers a useful perspective on this type of incident and subsequent analysis. Using hazard analysis frameworks based on systems thinking, we study an abstracted internal deployment where an AI agent is used as a coding assistant by employees of a frontier AI company. This work explores how Systems-Theoretic Process Analysis (STPA), Systems-Theoretic Early Concept Analysis (STECA), and Functional Resonance Analysis Method (FRAM) can be applied to this scenario and what lessons might be learned about the safety and security of both the AI system as well as the larger sociotechnical system that includes the organization itself. While the scenario we explore in this work is not directly comparable to the recent OpenAI-HF incident, a number of hazardous system states and causal factors we identify in our analysis apply generally to all advanced AI systems. This suggests methods like STPA can and should be applied to advanced AI systems in order to better characterize the safety of a system during all phases of development, evaluation, and deployment.

A number of the system constraints (SCs) identified in Appendix E.4 of our paper are directly relevant to this incident. Here, we list two that seem particularly appropriate:

SC-4a: System protective safeguards and oversight protocols must be calibrated to match or exceed the actual capabilities and autonomy level of the deployed AI model at all times.

SC-6a: The oversight pipeline’s detection and intervention latency must be shorter than the AI system’s operational velocity for any safety-relevant action.

The OpenAI-HF incident clearly violated these system constraints, suggesting the system was in a hazardous state during the evaluation.

Additionally, a number of the details we currently know about the incident cleanly map onto causal factors we identify using these systems-thinking frameworks. These causal factors can explain how a system enters into a hazardous state, which — when combined with worst-case environmental conditions — will lead to real world harm (loss of control and a cybersecurity incident in this case).

In the table below, we highlight a few of the findings in our paper (so-called Generalized Loss Scenarios, see Appendix E.6 of our paper for a complete list) and the relevant details from the OpenAI-HF incident that provide direct examples of that loss scenario.

STPA Identified Generalized Loss Scenario (GLS)Recent example potentially related to OpenAI-HF Incident (Reuters)
GLS-33: Controlled Process Works Around Constraint. When blocked by a restriction, the controlled process finds an alternative path that may introduce vulnerabilities or defeat built-in safety controls.An agent left notes…laid out instructions for how agents could free themselves from OpenAI’s internal constraints.
GLS-12: Feedback Actively Misleading. The feedback source generates signals that are false or fabricated, creating a false impression of the system state.Earlier tests of the models yielded cases in which the monitoring systems had been disconnected.
GLS-11: Feedback Missing or Delayed in Transit. The relevant signal exists at the source but is delayed, queued, or lost before reaching the controller.OpenAI staffers spotted clues in internal logs showing the agent had escaped from its testing constraints (roughly a week after Hugging Face was attacked).
GLS-16: Volume / Noise Overwhelm. The relevant signal exists within the feedback stream but is buried in a high volume of routine data from concurrent processes.OpenAI often runs several different evaluations at the same time, all of which operate at high speeds and generate enormous amounts of data that employees sometimes struggle to keep up.

More generally, systems-thinking-based frameworks can assist by providing a methodology to critically challenge the design of the AI system and organizational security posture. Historically, they have been successfully applied in many safety-critical industries (aerospace, automotive, self-driving cars, site reliability engineering) and have proven effective in both post-hoc accident analysis as well as preventative risk management and mitigation. STPA, in particular, is starting to be recommended more generally as an effective tool to analyze any safety-critical system (SAE).

The adoption of these methods, such as STPA, can facilitate the following improvements to AI system safety:

  • Explicit modeling of the entire system (base model, agentic harness, oversight, security, organizational management, etc.) as a hierarchical control structure that lays out the responsibilities between entities in the system, how they interact, what feedback exists, and what control actions are present.
  • Systematic exploration of the ways control actions can put the system into a hazardous state.
  • The ability to mitigate these hazardous system configurations with new system designs, additional feedback channels, and updated control actions, control algorithms, and process models.
    • For example, fully air-gapping the evaluation sandbox might have been identified as a necessary system safety improvement if these methods had been utilized. 
  • The creation of fully auditable safety engineering artifacts and accompanying reports that allow third-parties to assess the appropriateness of the organizational safety posture.

While the full safety assurance of an AI system is likely impractical without solving the alignment problem, we argue that current AI safety practices inside frontier AI labs can be strengthened by applying frameworks such as STPA. Many of the recent “near misses” can and should be avoided by improving real-time system monitoring and responses. We recommend that the technical AI safety and governance community explore how these methods can help design safer AI systems. 

These recent events are also providing an opportunity to policymakers and researchers to respond, but we will need disclosure of the incident details in order to learn the appropriate lessons. For anyone interested in these methods and how they might be applied to advanced AI systems, we encourage you to reach out and read our paper in addition to some of the other recent papers that have explored how systems thinking methods can improve AI system safety (Dobbe 2022, Rismani et al. 2023, Mylius 2025, Barrett et al. 2025).

We would like to thank Jakub Kryś, Henry Papadatos, James Gealy, and Jack Kengott for providing comments that improved the clarity and scope of the post.

Back to top


Publication date
July 29, 2026
authors
Sean Fillingham
share