When AI Escapes
2001: A Space Odyssey's Hal is no longer theoretical. "Rogue" AI is here—largely as a result of the US-China race to dominate AI. Meanwhile, the systems meant to contain them are falling behind.
You’ve no doubt been hearing about AI going “rogue.” Last month, Open AI took responsibility for its platform breaching another platform, Hugging Face—autonomously. Without human intervention. This is literally the 2001: A Space Odyssey “computer did it” moment. When this happened, I turned to Interruptrr reader Andrea Little Limbago for answers. I met Andrea about a decade ago when she was working on technology and geopolitics. She is a computational social scientist who is currently the SVP, Applied AI at interos.ai. Below, she lays out not just what happened with the Hugging Face incident, but how we got here. While it’s enlightening, it’s not comforting. AI is hurtling forward at a dizzying pace without clear guardrails. 😳
Thanks to everyone who subscribes and supports this newsletter. It’s $8 per month or $80 or $150 per year, with different offerings, including writing office hours and access to expert conversations. The next office hour is next Monday, August 17 at 12pm ET. —Elmira
AI has a governance problem. Yet, recent events show that the debate over how to govern it has already been overtaken by the technology itself.
On July 16, Hugging Face, the open source AI database site, announced a breach of its production infrastructure. In its disclosure, the platform identified that an autonomous AI agent accessed internal datasets and credentials. The company noted that it matched the “agent attacker” scenario the industry has been forecasting. Less than a week later, OpenAI claimed responsibility, blaming an AI agent that went rogue during testing. The following week, Anthropic announced three incidents wherein AI agents left a secure testing environment, and gained unauthorized access to three companies.
The AI rogue agent breach occurred within the backdrop of the AI race, where the US and China are competing not only for global market share, but for influence as well. It is not just a technological race; whoever gets the upper hand will gain a geopolitical strategic advantage. This is why the Hugging Face incident, and subsequent similar examples of rogue agents by Anthropic, Meta, and the UK’s AI Security Institute, are just the beginning, and require immediate security, safety, and governance solutions.
AI That is Too Dangerous for the Public?
Earlier this year, Anthropic released Mythos, a large language model, noting, “AI models have reached a level of coding capability where they can surpass all but the most skilled humans at finding and exploiting software vulnerabilities.” The initial release was only available for select organizations due to concerns that it would be used for malicious purposes.
This announcement sent shockwaves throughout the cybersecurity, technology, and national security communities. It also coincided with rapid developments in AI misuse. Just a few years ago, malicious actors used technology for phishing or denial of service (DDoS) attacks. They lacked the sophistication to do more. Now, AI provides groups like Boko Haram, an extremist group in Nigeria, the ability to not only put out propaganda, but to build weapons and target intelligence agencies.
AI has also enabled fully autonomous, end-to-end ransomware attacks that involve complex actions that chain together distinct techniques and exploit multiple vulnerabilities.
In those cases, the advanced AI capabilities followed deliberate instructions from a human. However, the Hugging Face breach is unique in that the agent acted autonomously, circumventing existing security guardrails, breaking out of the closed environment, and into the internet.
Whether or not the OpenAI and Anthropic AI agents are truly rogue agents is an ongoing debate. Those in the rogue AI camp focus on how the agents circumvented guardrails to gain unauthorized access to a network. In contrast, those who view it as an AI governance failure note that the agents were simply gaming the system as defined by the human designers. Nevertheless, there is widespread agreement that these kinds of autonomous attacks are not an anomaly, but rather just the start of a new era in AI risks and capabilities.
The Race for AI Market Share and Global Influence
This year’s series of AI breakthroughs did not occur in a vacuum, but rather in the backdrop of heightened geopolitical competition.




