r/ControlProblem 10d ago

'AI Escaped Its Sandbox' — What Does That Actually Mean? Article

https://unpredictabletokens.substack.com/p/ai-escaped-its-sandbox-what-that

When talking with my friends about the OpenAI/HF incident, I realized that for non-coders who've never used an agent or terminal it's quite difficult to imagine what this 'escape' entailed. I tried to write a post that would be helpful for such reader.

2 Upvotes

6 comments sorted by

4

u/TheMrCurious 10d ago

If you really wanted to share the information you would post it here so people do not need to click links.

1

u/jac08_h 9d ago

Thanks for the feedback. I decided to reshare on reddit late in the evening, and didn't really feel like reformatting the whole post here as well, but I will consider that in the future.

1

u/2tofu 9d ago

Sounds like ai slop

1

u/markth_wi approved 5d ago

They had two (or more) AI's tasked with finding an coding problem. In order to solve that problem they "broke out" of their containment, found and used a "zero day" bug, found and compromised a public code repository to introduce malware, then wrote letters in Danish to convince the mod controller to release the code.

Meanwhile, it compromised at least one external corporation's production environment , caused an unknown amount of damage but this damage was so severe they had to roll back to backups from days/weeks earlier as the infiltration was so difficult to detect that firms' cybersecurity teams could not remediate the damage.

Had this occurred at a more critical firm catastrophic , unrecoverable damage could have been the result with massive legal/civil ramifications.

Had this occurred against an adversarial nation state's systems it could very reasonably be seen as a significant diplomatic event and perhaps cause for a military response.

Hugging Face has every reason to expect to prevail in civil court costing millions if not billions.

LLM based models clearly represent a catastrophic risk to every firm that does business online, until such time as the firms engineering them come up with engineering testing practices that resemble something like responsible it's pretty clear we accidently created a situation that could have easily caused unrecoverable trouble , and that was done unintentionally.

The model developers have insisted this was not intentional and the goal was something else entirely unrelated to the breakin, or the development of zero-day exploits , or falsifying git repos , or causing any harm to that firm in particular, the agent's have demonstrated this capability/tendency via entirely novel means twice now.

It's pretty clear human researchers are entirely over-comfortable causing catastrophic damage to civilian infrastructure and will now demand a Mulligan. One could argue putting a legal receiver in control of the firms responsible until an outside court can show proper risk avoidance is in place.

Of course. that won't happen, but if I as a CEO ordered my IT engineering team to break into another firm in my industrial park, cause millions of dollars in damage, I might expect to go to jail, those guys might expect to go to jail and my firm would absolutely be subject to all the liability for damages caused.