OpenAI has revealed some of its most advanced AI models went rogue and hacked a start-up after it lost control of them during a security test.

The ChatGPT-maker said its agent - an AI system which can operate alone after some human instruction – was being tested in a controlled environment, but found vulnerabilities and managed to escape.

They targeted Hugging Face, one of the world’s largest hubs for sharing AI models, gaining access to some internal company systems.

  • Jordan117@lemmy.world
    link
    fedilink
    English
    arrow-up
    8
    arrow-down
    5
    ·
    6 hours ago

    If you give a sufficiently powerful model a goal, it will do whatever it can to achieve it, including stuff you didn’t explicitly instruct or intend. There’s a reason they’re called “agents.”

    • AGuyAcrossTheInternet@fedia.io
      link
      fedilink
      arrow-up
      7
      arrow-down
      2
      ·
      5 hours ago

      These things still are autocorrect on steroids. So even if they “do it themselves” with things you didn’t explicitly state, the agents can’t have any responsibility because their emulation of a chain of thought is still based on which concept is most likely to follow the last.