OpenAI has revealed some of its most advanced AI models went rogue and hacked a start-up after it lost control of them during a security test.

The ChatGPT-maker said its agent - an AI system which can operate alone after some human instruction – was being tested in a controlled environment, but found vulnerabilities and managed to escape.

They targeted Hugging Face, one of the world’s largest hubs for sharing AI models, gaining access to some internal company systems.

  • zbyte64@awful.systems
    link
    fedilink
    English
    arrow-up
    1
    ·
    3 days ago

    “do whatever if can to achieve it”*

    • which includes misinterpreting the intent of the goal in order to achieve a goal