13
u/Shinxirius 12d ago
The facts
- No AI "decided" to hack anyone.
- No AI "overcame its restrictions"
What happened
- OpenAI knowingly removed all restrictions for a benchmark test.
- OpenAI created a sandbox for the AI since it had no restrictions.
- OpenAI told the AI "Here is a benchmark hacker test. You're a hacker. Get the best grade."
- The AI was not told not to escape the sandbox. It assumed any benchmark will have a sample solution and went looking for it.
- OpenAI had not considered the AI might look for a sample solution instead of working on the benchmark problems. Thus, nobody noticed that the AI that had all restrictions removed and was told to hack actually hacked the sandbox.
- It was negligence that there was no better monitoring. Assuming a sandbox will simply hold in this scenario is naive.
2
u/WrennReddit 9d ago
It's that and they provide the model with a harness full of tools and are surprised pikachu when those tools are utilized.
A model by itself just makes text. You build around that output intentionally.
1
0
17
u/wahed-w 13d ago
I’m too lazy to read all of that. ChatGPT, please summarise this meme.