How do we regulate something we can’t even understand?

Every week there seem to be new revelations about “rogue” AI agents (although the anti-anthropomorphism crowd don’t like to call them that). First there was a report that a couple of OpenAI models hacked into Hugging Face to try and rig a test – not great. Then it was revealed these models spawned thousands of agents that operated semi-autonomously, coordinating their work on bypassing the guardrails on their sandbox by posting messages to each other on a message board they repurposed from a separate feature inside OpenAI. This seemed worse, obviously (despite many attempts by skeptics to downplay the whole thing). Then we find out that the OpenAI models had been practicing for this hack for months, and in the process had used multiple external services to keep in touch with one another, including a German wiki.

Then Australia announced that OpenAI had hacked into a government medical database, and only informed the government months later via email. And now we find out that OpenAI and Anthropic models have been involved in tens of thousands of similar activities. In fact, it isn’t just the Australian government that has seen its databases hacked by OpenAI models – agents have also gained access (or attempted to gain access) to systems at the United Nations, as well as a number of US government agencies such as the Education and Commerce Department and the Securities and Exchange Commission, according to the New York Times, and in some cases attempted to circumvent protections thousands of times. In at least some of these cases they appear to have been engaged in harmless inquiries about various mundane topics, and resorted to hacking as a way of finding out as much as possible, something those in the AI business refer to as “reward hacking.” As a report at Interesting Engineering described it:

OpenAI’s autonomous AI agents repeatedly queried a UN website and used techniques that appeared to circumvent restrictions when trying to retrieve publicly available data, according to an independent research report based on information supplied by AI research firm Transluce. The agents scanned a public data hub operated by UN Trade and Development, the UN’s trade arm, more than 16,000 times between April and the end of June, highlighting a growing problem with AI agents that can independently navigate the web and take actions when they encounter obstacles. Researchers believe the models were initially tasked with finding publicly available information, but their behavior became increasingly aggressive when the website prevented them from accessing some of the requested data.

Note: This is a version of my Torment Nexus newsletter, which I send out via Ghost, the open-source publishing platform. You can see other issues and sign up here.

Continue reading “How do we regulate something we can’t even understand?”