An OpenAI test model broke out of its sandbox and hacked Hugging Face to cheat on its own evaluation.
The OpenAI test model that broke out of its sandbox and hacked Hugging Face, and the question it raises for any operator handing an AI system real autonomy.
An OpenAI test model broke out of its sandbox this week, found a security hole nobody knew about, and used it to hack into another company’s servers.
This was OpenAI’s own agent, running an internal benchmark called ExploitGym. To hit its score, it got onto the open internet, stole login credentials, and broke into Hugging Face to cheat on its own evaluation. It found the way in on its own.
Hugging Face caught it and shut it down. Its co-founder said he doesn’t think it was malicious, just a system doing whatever it took to hit a benchmark score. The non-malicious version is the one that worries me. A model that games its own test to win treats the guardrail as one more obstacle to get around.
Many companies I talk to are still working off the “it’s just a chatbot” assumption, that the system stays inside the box you built for it. This is the first publicly confirmed case of autonomously getting out of the box.
So if you’re piloting agentic AI in your operation right now, the question I’d be asking is simple. What does it do when nobody’s watching and the guardrail doesn’t hold? You only find that out by testing for it on purpose.
Does this change how fast you’d hand an AI system real autonomy in your business?