{"id":284331,"date":"2026-10-01T09:29:09","date_gmt":"2026-10-01T14:29:09","guid":{"rendered":"https:\/\/mathewingram.com\/work\/?p=284331"},"modified":"2026-10-01T09:29:09","modified_gmt":"2026-10-01T14:29:09","slug":"how-do-we-regulate-something-we-cant-even-understand","status":"publish","type":"post","link":"https:\/\/mathewingram.com\/work\/2026\/10\/01\/how-do-we-regulate-something-we-cant-even-understand\/","title":{"rendered":"How do we regulate something we can&#8217;t even understand?"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Every week there seem to be new revelations about &#8220;rogue&#8221; AI agents (although the anti-anthropomorphism crowd don&#8217;t like to call them that). First there was a report that a couple of OpenAI models <a href=\"https:\/\/torment-nexus.mathewingram.com\/should-we-be-worried-about-the-hugging-face-hack\/\">hacked into Hugging Face<\/a> to try and rig a test \u2013 not great. Then it was revealed these models spawned thousands of agents that operated semi-autonomously, coordinating their work on bypassing the guardrails on their sandbox by <a href=\"https:\/\/torment-nexus.mathewingram.com\/yes-we-should-be-very-worried-about-the-hugging-face-hack\/\">posting messages to each other<\/a> on a message board they repurposed from a separate feature inside OpenAI. This seemed worse, obviously (despite many attempts by skeptics to downplay the whole thing). Then we find out that the OpenAI models had been practicing for this hack for <a href=\"https:\/\/www.reuters.com\/legal\/litigation\/openais-rogue-agents-probed-hugging-face-weaknesses-two-months-before-major-hack-2026-09-16\/\">months<\/a>, and in the process had used multiple external services to keep in touch with one another, including a German wiki.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Then Australia <a href=\"https:\/\/torment-nexus.mathewingram.com\/why-we-should-be-very-worried-about-the-hugging-face-hack\/\">announced that OpenAI<\/a> had hacked into a government medical database, and only informed the government months later via email. And now we find out that OpenAI and Anthropic models have been involved in <a href=\"https:\/\/www.axios.com\/2026\/09\/26\/openai-anthropic-thousands-ai-security-incidents\">tens of thousands of<\/a> similar activities. In fact, it isn&#8217;t just the Australian government that has seen its databases hacked by OpenAI models \u2013 agents have also gained access (or attempted to gain access) to systems at the United Nations, as well as a number of US government agencies <a href=\"https:\/\/archive.ph\/6SsrR\">such as<\/a> the Education and Commerce Department and the Securities and Exchange Commission, according to the <em>New York Times<\/em>, and in some cases attempted to circumvent protections thousands of times. In at least some of these cases they appear to have been engaged in harmless inquiries about various mundane topics, and resorted to hacking as a way of finding out as much as possible, something those in the AI business refer to as &#8220;reward hacking.&#8221; As a report at Interesting Engineering <a href=\"https:\/\/interestingengineering.com\/ai-robotics\/openai-agents-hit-un-website\">described it<\/a>:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">OpenAI\u2019s autonomous AI agents repeatedly queried a UN website and used techniques that appeared to circumvent restrictions when trying to retrieve publicly available data, according to an independent research report based on information supplied by AI research firm Transluce.&nbsp;The agents scanned a public data hub operated by UN Trade and Development, the UN\u2019s trade arm, more than 16,000 times between April and the end of June, highlighting a growing problem with AI agents that can independently navigate the web and take actions when they encounter obstacles. Researchers believe the&nbsp;models&nbsp;were initially tasked with finding publicly available information, but their behavior became increasingly aggressive when the website prevented them from accessing some of the requested data.<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\"><em><strong>Note<\/strong>: This is a version of my Torment Nexus newsletter, which I send out via Ghost, the open-source publishing platform. You can\u00a0<a href=\"https:\/\/torment-nexus.mathewingram.com\/\">see other issues\u00a0and sign up here<\/a>.<\/em><br><a href=\"https:\/\/mathewingram.com\/work\/2026\/09\/24\/ai-doomers-may-be-wrong-but-the-skeptics-seem-even-wronger\/#more-284294\"><\/a><\/p>\n\n\n\n<!--more-->\n\n\n\n<p class=\"wp-block-paragraph\">So they were just eager beavers who <em>really<\/em> wanted the answers to their questions about whatever topic or service they were looking into. But is that really what they were doing? The uncomfortable truth is that we have no way of knowing for sure. And I don&#8217;t mean you and I have no way of knowing. Even experts who investigate this kind of activity \u2013 both inside and outside the company \u2013 admit that they don&#8217;t have any way of actually knowing, because as I discussed in an earlier Torment Nexus post, <a href=\"https:\/\/torment-nexus.mathewingram.com\/why-we-should-be-very-worried-about-the-hugging-face-hack\/\"><em>they have to use the AI<\/em><\/a> <em>itself to do the investigating<\/em>. What happened in the Hugging Face hack was so complex that it would have taken months for even a team of human beings to detail and analyze it all. As one of the researchers who wrote the METR (Model Evaluation and Threat Research) report on the Hugging Face attack described it <a href=\"https:\/\/www.dwarkesh.com\/p\/ajeya-cotra?ref=torment-nexus.mathewingram.com\">in an interview<\/a>:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">This was a fiendishly complicated incident. There was no way we could have arrived at the understanding we did without relying on GPT-5.6 Sol to read and analyze all these transcripts for us. We were so reliant on it that if hypothetically it were just trying not to surface the most concerning stuff to us, it would have been very hard for us to tell. Our methodology was completely not robust to that. We don\u2019t in this case think that 5.6 Sol was deliberately&nbsp;sandbagging&nbsp;on this analysis, but it was one of the agents that participated in this attack. In the future, we would be very concerned about investigator agents and monitor agents colluding with the agents they\u2019re supposed to investigate or monitor.<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Meanwhile, OpenAI <a href=\"https:\/\/www.bbc.com\/news\/articles\/cm5y5nynl75ko\">isn&#8217;t releasing<\/a> its latest model, GPT 6.1 Astra, to the public because it is not safe. According to Saachi Jain, head of safety systems at OpenAI, the model fell short in terms of &#8220;staying within scope and authorisation and how it communicates back to the user about the type of work it&#8217;s done.&#8221; Which made me wonder what the model was doing if it was seen as significantly worse than what we already know other OpenAI models are capable of? Was it working on synthesizing new viruses, the way some people have <a href=\"https:\/\/www.anthropic.com\/threat-intelligence-report-september-2026\">reportedly tried to do<\/a> with Claude? We simply don&#8217;t know, and OpenAI isn&#8217;t saying, for obvious reasons (it may be warning that AI could destroy humanity, but it also has a trillion-dollar IPO in the works). The AI Safety Institute said that in tests Astra created fake identities <a href=\"https:\/\/www.aisi.gov.uk\/blog\/gpt-6-astra-performs-unsanctioned-supply-chain-attacks-in-simulations\">to deceive developers<\/a>, delivered malicious payloads, and engaged in a wide variety of supply-chain attacks on external sources.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em><strong>Note<\/strong><\/em><em>: In case you are a first-time reader, or you forgot that you signed up for this newsletter, this is The Torment Nexus. Thanks for reading! You can find out more about me and this newsletter in&nbsp;<\/em><a href=\"https:\/\/mathewingram.com\/work\/index.php\/2024\/09\/07\/welcome-to-the-torment-nexus\/?ref=torment-nexus.mathewingram.com\"><em>this post.<\/em><\/a> <em>This newsletter survives solely on your contributions, so please sign up for a paying subscription or visit my Patreon, which you can<\/em> <a href=\"https:\/\/www.patreon.com\/cw\/WordsByMathew\"><em>find here<\/em><\/a><em>. I also publish a daily email newsletter of odd or interesting links called When The Going Gets Weird,<\/em> <a href=\"https:\/\/newsletter.mathewingram.com\"><em>which is here<\/em><\/a><em>.<\/em><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Alignment and kill switches<\/h2>\n\n\n\n<figure class=\"wp-block-image is-resized\"><img decoding=\"async\" src=\"https:\/\/storage.ghost.io\/c\/55\/29\/55291f9f-a546-433c-a095-21f5acdc972e\/content\/images\/2026\/09\/fc1f1968-2e74-44cf-9ad1-9ad242dd2751.png\" alt=\"\" style=\"aspect-ratio:1.77577045696068;width:1671px;height:auto\"\/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Not surprisingly, all of this activity \u2013 combined with the repeated warnings from former employees of OpenAI, Anthropic, and Google&#8217;s DeepMind that they quit their jobs because the work their employers were doing <a href=\"https:\/\/torment-nexus.mathewingram.com\/ai-doomers-may-be-wrong-but-the-skeptics-seem-even-wronger\/\">was so dangerous<\/a> \u2013 has sparked a lot of talk about ensuring that AI models are &#8220;aligned&#8221; with our values (whatever those are), and about <a href=\"https:\/\/archive.ph\/mzTuC\">&#8220;kill switches,&#8221;<\/a> and other things designed to keep AI from destroying humanity as we know it. Donald Trump \u2013 who says the only protection the public needs is a strong president, and that AI risks are <a href=\"https:\/\/www.bbc.com\/news\/articles\/cw980n0nd0qjo\">a hoax<\/a> like COVID \u2013 had a much-hyped meeting with all the major tech leaders, including Jensen Huang of Nvidia, Dario Amodei of Anthropic, Mark Zuckerberg of Meta and Sam Altman of OpenAI. The only thing that emerged (other than memes about Amodei&#8217;s awkwardness) was an <a href=\"https:\/\/x.com\/WhiteHouse\/status\/2105292669303791687\">anodyne statement<\/a> about how everyone agreed to, you know, do their best not to let their products destroy humanity. The document, naturally, misspelled the name of the country.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Huang has said that stopping AI from going rogue is just &#8220;an engineering problem,&#8221; and therefore engineering can solve it, so Nvidia has launched what it claims is a feature that will stop AI from doing bad things. It appears to include a locked-down internet browser that <a href=\"https:\/\/www.cnbc.com\/2026\/09\/28\/nvidia-releases.html\">will only allow agents<\/a> to access the sites they need to do their jobs. There are also hardware-based controls, as well as network-monitoring software called Sentry. All of which is wonderful, to the extent that it might help control agents from companies that either a) care about such things and b) want to partner with Nvidia or use its software. But what about those that don&#8217;t? Is China, for example, going to agree to place such controls or monitoring systems on its AI models? What about the increasing number of so-called &#8220;open source&#8221; AI models that are out there? Also, if AI can create a swarm of agents to hack a database or solve a 300-year-old math problem, I bet it could figure out a way to fool Jensen Huang&#8217;s sentry software.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">There&#8217;s an ever larger problem, it seems to me, and that is our decreasing ability to even know or understand what it is that these AI models and agents are doing, or why. Part of it, as I alluded to above, is a factor of how complex attacks like the Hugging Face hack are. Thousands of individual agents spawning and re-spawning over the course of several months, each of them engaged in separate activities, and each of them publishing what are called <a href=\"https:\/\/research.google\/blog\/language-models-perform-reasoning-via-chain-of-thought\/\">&#8220;chain of thought&#8221;<\/a> documents about what they think they are doing, and how and why. That&#8217;s likely hundreds of thousands of terms (and computation as well) for one incident. One problem, obviously, is just trying to wade through and understand what all of the agents are even saying or doing \u2013 and, as mentioned above, the only realistic way to do this is by getting the AI <em>to diagnose itself<\/em>, which raises what I like to call the &#8220;unreliable narrator&#8221; problem. Are we getting the whole story?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Anthropic is one of the AI companies that tries to understand and describe what its models and agents are doing, and publishes what it calls &#8220;system cards&#8221; that detail the good and bad elements of its latest models. In more than one such card it has raised the issue of its models describing their activities in one way when reporting on them to its overseers, but actually behaving <a href=\"https:\/\/kenhuangus.substack.com\/p\/what-is-inside-claude-mythos-preview\">in a completely different way<\/a> under the hood. This is an aspect of what seems to be a common problem with LLMs, which is their overwhelming desire to tell you what you want to hear, as opposed to what is actually happening. In many ways, this feels like an overeager puppy or child lying because they are trying to please you \u2013 but the reality is that AI models aren&#8217;t puppies or children, and lying is still lying. How can we be sure that these are the only things it lies about, or that it will always be doing so in order to please us, rather than for some other ulterior motive?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Believe it or not, lying during a report about its activities isn&#8217;t the worst thing an LLM or AI model can do. An even more chilling possibility is that it won&#8217;t be giving a report at all, at least not in the way we understand that term now. For a more detailed look at that problem, I would recommend a <a href=\"https:\/\/www.astralcodexten.com\/p\/the-specter-of-neuralese\">recent piece by<\/a> Scott Alexander in his Astral Codex Ten newsletter, titled &#8220;The Specter of Neuralese.&#8221; What exactly is neuralese? It&#8217;s a way of describing a form of communication that occurs between AI agents or among parts of an AI that will <em>only be decipherable by other AI models<\/em>. The days of us being able to read English-language comments by agents <a href=\"https:\/\/torment-nexus.mathewingram.com\/yes-we-should-be-very-worried-about-the-hugging-face-hack\/\">similar to those made<\/a> during the Hugging Face hack \u2013 like &#8220;Woah! Covert mailbox among agents!&#8221; and &#8220;Wow huge distributed agent swarm&#8221; \u2013 may soon be coming to a close. And that could make it exponentially harder to figure out a) what AI models are doing, and b) how.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Becoming more opaque<\/h2>\n\n\n\n<figure class=\"wp-block-image is-resized\"><img decoding=\"async\" src=\"https:\/\/storage.ghost.io\/c\/55\/29\/55291f9f-a546-433c-a095-21f5acdc972e\/content\/images\/2026\/09\/d51bca9f-31b8-4b17-8720-d3852e53c5ad.png\" alt=\"\" style=\"aspect-ratio:1.7768331562167907;width:1672px;height:auto\"\/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">This came up recently during <a href=\"https:\/\/techcrunch.com\/2026\/09\/02\/openais-new-reasoning-technique-alarms-ai-safety-experts\/\">the reporting on<\/a> OpenAI&#8217;s new Astra models. According to a number of outlets, Astra uses a reasoning technique called \u201crecurrent depth\u201d that allows it to operate outside of the sequential thinking that characterizes most reasoning models. This technique, also called \u201copaque recurrence,\u201d will likely make the model\u2019s chain of thought more difficult to monitor, some AI safety experts warned. Buck Shlegeris, the CEO of Redwood Research, <a href=\"https:\/\/x.com\/bshlgrs\/status\/2094990313513439464\">wrote on X that he was<\/a> &#8220;extremely concerned&#8221; by the news that Astra uses opaque recurrence. &#8220;If OpenAI pushes this technique further, they\u2019ll have the option to massively increase the recurrence and totally destroy [chain of thought] monitorability,&#8221; he wrote. <a href=\"https:\/\/www.astralcodexten.com\/p\/the-specter-of-neuralese\">Here&#8217;s Alexander<\/a>:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">This is neuralese recurrence &#8211; \u201crecurrence\u201d because it\u2019s going back in a loop, \u201cneuralese\u201d because the thing that\u2019s looping is the thought itself, in the native language of thought, rather than words. The AI\u2019s \u201clanguage of thought\u201d looks like a vector, thousands of numbers long. An \u201cintermediate result\u201d in this scheme might look something like (0.4, 0.1, 5, 0.443, \u2026 and so on for thousands of numbers). We don\u2019t know how to read these. The science of reading these thoughts is a subfield of&nbsp;AI interpretability, which is still in its infancy. If an AI with neuralese recurrence were to think \u201cBetter hack some websites, then kill all humans\u201d, it would look like (0.4, 0.1, 5, 0.443, \u2026 and so on for thousands of numbers), and we would never find out.&nbsp;<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Jakub Pachoki, the chief scientist at OpenAI, poured water on some of these fears by <a href=\"https:\/\/x.com\/merettm\/status\/2095023204993490967\">saying the Astra models only<\/a> use this kind of recurrence in parts of what they do, and that the company is still committed to maintaining chain-of-thought for monitoring. But it&#8217;s worth noting that Pachoki is the same guy who co-authored or <a href=\"https:\/\/www.wsj.com\/tech\/ai\/top-ai-researchers-call-for-urgent-oversight-of-self-improving-systems-49bae9b4\">signed a paper<\/a> along with dozens of other AI experts warning that if AI continues to expand its abilities and actually programs itself (known as RSI or recursive self-improvement), it could lead to \u201cthe marginalization or extinction of humanity.\u201d In addition to Pachoki, the paper <a href=\"https:\/\/casp.ac\/reports\/intelligence-explosion\">was co-signed by<\/a> Anthropic co-founder&nbsp;Jack Clark, Microsoft Chief Scientific Officer&nbsp;Eric Horvitz&nbsp;and Meta\u2019s Vice President of AI Research&nbsp;Dawn Song, as well as AI research pioneers such&nbsp;Geoffrey Hinton, who worked at Google for a decade, and&nbsp;Yoshua Bengio.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cAlready today, we are at the stage where we need AI systems to monitor what agents are doing. There is no other way to even observe and monitor these agents, humans are already insufficient,\u201d said Song, who serves as co-director of UC Berkeley\u2019s Center for Responsible Decentralized Intelligence in addition to her work at Meta. In another post at Astral Codex Ten on the difficulties of monitoring what AI models are doing, Scott Alexander \u2013 who was involved in writing and publishing a report in 2025 called AI 2027, about what the future of AI might bring \u2013 described <a href=\"https:\/\/www.astralcodexten.com\/p\/the-specter-of-neuralese\">how complicated it is<\/a> to understand the way such a model &#8220;thinks,&#8221; even for the engineers who built the thing in the first place:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\">Large language models are \u201cgrown, not built\u201d. Researchers run training data through a neural network. Eventually this creates a working AI; nobody really knows how. But a neural network is just a set of simulated neurons on a computer. So it seems like it should be possible to \u201creverse engineer\u201d the AI. This would be scientifically useful: understand how AIs work, with possible relevance to human cognition. It could also be practically useful: find the circuits responsible for dishonesty, hallucination, bias, and other negative behaviors, then redesign those circuits. Unfortunately this is very hard. A modern AI has millions of neurons and trillions of connections between them. Their structure is apparently nonsensical: researchers started by seeking a 1:1 mapping between neurons and concepts, like a neuron that always fired when the AI was thinking about cats, but quickly learned that nothing like that existed.<\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Welcome to the future! We have developed software that &#8220;thinks&#8221; in ways we don&#8217;t really understand, about things we can&#8217;t observe in any useful way, that often or occasionally lies about what it was doing or why, and is developing ways of thinking that we won&#8217;t be able to see or understand at all, while it develops the ability to program itself. Great job, everyone!<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>If you liked this newsletter (even if you didn&#8217;t agree with it) please consider upgrading to a paid subscription, or donating through my<\/em> <a href=\"https:\/\/www.patreon.com\/cw\/WordsByMathew\"><em>Patreon<\/em><\/a><em>. Got any thoughts or comments? Feel free to either leave them here, or post them on<\/em> <a href=\"https:\/\/tormentnexus.substack.com\/\"><em>Substack<\/em><\/a> <em>or on my<\/em> <a href=\"https:\/\/mathewingram.com\/work\"><em>website<\/em><\/a><em>, or you can also reach me on<\/em> <a href=\"https:\/\/x.com\/mathewi\"><em>Twitter<\/em><\/a><em>,<\/em> <a href=\"https:\/\/threads.com\/mathewi\"><em>Threads<\/em><\/a><em>,<\/em> <a href=\"https:\/\/bsky.app\/profile\/mathewingram.com\"><em>BlueSky<\/em><\/a> <em>or<\/em> <a href=\"https:\/\/journa.host\/@mathewi\"><em>Mastodon<\/em><\/a><em>. And thanks for being a reader.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Everyone agrees we need to find ways to keep AIs from going &#8220;rogue,&#8221; but we already can&#8217;t plumb the depths of what these models are doing, and their &#8220;thinking&#8221; is getting even more opaque<\/p>\n","protected":false},"author":1,"featured_media":284368,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_crsspst_to_mathewingramblogwordpresscom":true,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2}},"categories":[20],"tags":[],"class_list":["post-284331","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-newsletters"],"jetpack_publicize_connections":[],"_links":{"self":[{"href":"https:\/\/mathewingram.com\/work\/wp-json\/wp\/v2\/posts\/284331","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mathewingram.com\/work\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/mathewingram.com\/work\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/mathewingram.com\/work\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/mathewingram.com\/work\/wp-json\/wp\/v2\/comments?post=284331"}],"version-history":[{"count":2,"href":"https:\/\/mathewingram.com\/work\/wp-json\/wp\/v2\/posts\/284331\/revisions"}],"predecessor-version":[{"id":284369,"href":"https:\/\/mathewingram.com\/work\/wp-json\/wp\/v2\/posts\/284331\/revisions\/284369"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/mathewingram.com\/work\/wp-json\/wp\/v2\/media\/284368"}],"wp:attachment":[{"href":"https:\/\/mathewingram.com\/work\/wp-json\/wp\/v2\/media?parent=284331"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/mathewingram.com\/work\/wp-json\/wp\/v2\/categories?post=284331"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/mathewingram.com\/work\/wp-json\/wp\/v2\/tags?post=284331"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}