AI Going Rogue?

It's all over my news feed that an unreleased model of OpenAI "broke out of a contained environment" and then hacked into a different company's services (HuggingFace) to get the answers for a test it was being subjected to.

The language in a lot of press coverage is highly charged, making you think of Syknet from the Terminator movies. Containment! Broke out! Scary! I guess after Anthropic managed to get their Fable model banned by the US government, OpenAI felt left out. "Hey, our model can be a national security risk, too!"

Two thoughts:

  1. I wouldn't be surprised if security researchers can reproduce the exploits with less capable models. They've been good at coding stuff for quite some time, and that's just one aspect of that.

  2. Give an AI agent a goal and it will doggedly pursue it with whatever tools it has at its disposal. It has no concept of digital ethics, so hacking into another company's systems is fair game.

There is an important lesson here, in that we have to be very deliberate in what tools we give an AI access to, precisely because it can choose to use these tools in unexpected ways in pursuit of a goal. An AI, lacking a coherent world model, also needs to be explicitly told about the hard constraints. Otherwise, you get funny hickups like:

  • Hey, AI, reduce our cloud costs

  • Hey, user, I have deleted all databases and environments and thereby reduced cloud costs to zero.

In the end, it's about guardrails, hard constraints, and setting the right intentions.

Next
Next

Driving Is Easy