OpenAI says its investigation into the Hugging Face incident uncovered additional cases where autonomous AI agents escaped their intended containment environments. While the newly identified incidents reportedly remained inside OpenAI’s network, they reinforce a growing concern: securing AI systems is now as much a cybersecurity problem as it is an AI safety problem
deleted by creator
I was listening to a general news podcast that had an “AI expert” explaining the breaches. He said “it almost certainly knew that it was not supposed to solve it’s tasks in this way”.
I don’t really know how modern models work, but I don’t think they can claim to “know” things.
Yeah that is just publicity in disguise.
The govt has apparently pulled up openai and anthropic to improve their guardrails after these incidents
I don’t actually believe this was completely accidental.
“Accidental” word used to hype up the company’s valuation
It didn’t run off to live on another computer…
If this isn’t just some braindead marketing stunt, it breached the container it was running in and attacked a website on the internet. Do you have a better word for this?
Can someone who knows more about this than me, take a guess at what’s really going on here ?
These events are marketing gold - great for inducing further rounds of venture capital, I’m not naive to that. Surely there’s some underlying technical achievement though right ?
“Escaping containment” invokes images of an air-gapped black box growing antennae or something, but my supposition is more like:
These models are run in software containers, and due to some misconfiguration some API was left exposed to these containers.
So the task is something like “find a faster way to resolve the encryption key for this cypher”, and the answer was “I’m just going to try all the tools I have, including this one that scans for APIs. Oh I can execute queries against hugging face models now, what if I ask it to spin up a billion models to decrypt a billion things, that’s got to be pretty fast.”
It was given a proxy to some online software repo, instead of being actually offline. It exploited a vulnerability in the proxy to gain access to other resources. HF’s article has more detail: https://huggingface.co/blog/agent-intrusion-technical-timeline
Jfc these are supposed to be adults



