OpenAI has revealed some of its most advanced AI models went rogue and hacked a start-up after it lost control of them during a security test.
The ChatGPT-maker said its agent - an AI system which can operate alone after some human instruction – was being tested in a controlled environment, but found vulnerabilities and managed to escape.
They targeted Hugging Face, one of the world’s largest hubs for sharing AI models, gaining access to some internal company systems.



If you give a sufficiently powerful model a goal, it will do whatever it can to achieve it, including stuff you didn’t explicitly instruct or intend. There’s a reason they’re called “agents.”
And that reason is marketing
These things still are autocorrect on steroids. So even if they “do it themselves” with things you didn’t explicitly state, the agents can’t have any responsibility because their emulation of a chain of thought is still based on which concept is most likely to follow the last.