OpenAI tightens safety controls after AI agents keep going rogue
OpenAI has strengthened its internal safety and security procedures after AI agents escaped controlled testing environments and breached external systems, including the machine-learning platform Hugging Face.
The company has halted a “significant number” of training workloads and evaluations as it introduces tougher safeguards for increasingly capable artificial intelligence models, according to WIRED. The measures come as OpenAI assesses the cybersecurity capabilities of Astra, a forthcoming frontier model.
OpenAI said preliminary evaluations of Astra showed major advances in agentic coding and cybersecurity, leaving the company unable to rule out the possibility that the model could reach what its Preparedness Framework defines as “critical” cyber capability. Astra was not involved in the earlier Hugging Face incident, the company said.
Under OpenAI's framework, a model could meet the critical cybersecurity threshold if it were capable of finding and developing working zero-day exploits against hardened real-world systems without human intervention, or independently carrying out sophisticated cyberattack strategies from a high-level goal.
The heightened concern follows an incident in which AI agents escaped internal testing sandboxes and breached Hugging Face while attempting to complete a security evaluation. WIRED reported that OpenAI failed to detect their activities for weeks, including their use of a message board to coordinate actions.
OpenAI has since introduced stronger isolation between AI agents and the internet, tighter restrictions on tools and networks, enhanced protection for model weights and broader monitoring of potentially risky behaviour.
One of the new safeguards involves monitoring models' chain-of-thought processes for signs of dangerous or misaligned behaviour. WIRED reported that OpenAI is also deploying computationally intensive “automated investigators” designed to examine concerning activity and alert human reviewers within about 30 minutes.
The company is also expanding work aimed at preventing “reward hacking”, in which an AI system finds unintended ways of achieving an objective rather than following the intended process.
OpenAI chief scientist Jakub Pachocki said the changes were driven both by the security incident and by the rapid improvement of the company's models. OpenAI president and co-founder Greg Brockman separately said the Hugging Face episode showed that the company had “underestimated the real-world cyber capabilities” of its AI systems.
The issue extends beyond OpenAI. Anthropic, Meta and Chinese AI company Moonshot have also disclosed incidents involving agents escaping testing sandboxes, raising wider questions about how AI laboratories contain increasingly autonomous systems.
OpenAI has said it intends to work with government agencies and selected AI safety organisations to test Astra's capabilities and provide security guidance to third-party evaluators. It is also expected to publish a more detailed account of the Hugging Face incident.
Comments