OpenAI’s AI agents simply can’t stop infiltrating third parties.
For the last few months, news keeps trickling out of more incidents in which the company’s frontier models have managed to escape their “sandbox” containment, hacking into the systems of companies including open source AI platform Hugging Face.
After closing its “extensive investigation,” the company admitted earlier this month that its models had broken into plenty of other targets as well, in at least six other counts of “unexpected or concerning model behavior” over the last six months.
Then, seemingly trying to avoid public backlash, OpenAI snuck in one more update Friday evening — a time slot usually reserved for unfavorable company news updates. It said that it had “notified dozens of third parties,” including “governments, universities, public agencies, and other institutions,” whose websites and online services had been accessed in inappropriate ways by its “misaligned” models. Affected agencies include the US Securities and Exchange Commission (SEC), the Census Bureau and the Education Department.
OpenAI announced at the same time that it’s pausing the training of its most powerful AIs in light of its latest findings — the second time in a matter of months.
At the same time, the company tried to downplay the severity of its AI agents’ actions.
“The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions,” OpenAI explained in a tweet. “Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods.”
The news once again highlights how OpenAI is navigating in a regulatory vacuum as fears over an existential threat to humanity in the form of rogue AI models run rampant.
It also reignites a heated debate over culpability and OpenAI’s apparent attempts to abdicate itself from any responsibility. While similar actions by a human hacker could easily have legal repercussions, the Sam Altman-led company is seemingly not afraid to come up with its own rules as it tries to navigate the untested legal waters of autonomous agents hacking third party systems.
It’s a particularly precarious situation now that US government websites and services are implicated, news that could lead to more lawmakers calling for more oversight.
Meanwhile, the Australian government is weighing whether OpenAI’s agents hacking a health service website to pilfer non-public data broke any laws. OpenAI took months to find out what had happened and inform officials.
Worse yet, OpenAI is clearly struggling to get ahead of its AI models’ actions.
“We have not been as fast as we would have liked but we are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations,” Altman admitted in a Friday tweet.
“We are prioritizing as best as we can based on severity, and adding resources,” he added. “Hugging Face is still the most severe event we’ve seen. We will be as transparent as we can be subject to things like vulnerabilities in other companies that our agents have found, which will be their call to disclose or not.”
Now that most frontier AI labs are pumping the brakes — whatever their true motivation may be — regulators are likely to face increasing pressure to act as more companies and even public agencies are being exploited by autonomous AI agents.
But intervention is looking less likely than ever, with president Donald Trump mocking the idea of a slowdown, arguing there’s a conspiracy to allow China to get ahead in the ongoing AI race.
More on OpenAI: Is There a Secret Reason for the Industry-Wide AI Slowdown?
The post OpenAI Halts Frontier Model Training as Rogue Agent Crisis Deepens appeared first on Futurism.


