As regular readers will know, last week I re-published a piece that I originally wrote on September 3rd, entitled “Yes, we should be very worried about the Hugging Face hack.” I couldn’t write a new post for last week because I was backwoods camping in a remote area of southwestern Ontario – Killarney Provincial Park, in case you are interested – and I had virtually no cell signal for most of that time (if you like canoes or kayaks or sunsets or trees, you can check out the post I wrote about the trip). I chose to re-publish the Hugging Face hack piece because it seemed even more appropriate in light of some of the revelations that have been coming out about the hack – including the fact that the hack was neither the first time OpenAI models had accessed another company’s systems without permission, nor the first time the models in question had used a message board to discuss the methods used in their attack. The September 3rd piece was itself an update on an earlier piece I wrote in July, right after the Hugging Face attack was first revealed (a reveal that came weeks after the hack itself took place).
The debate over AI safety and what AI researchers call “alignment” – or, to put it more bluntly, between whether AI is working as advertised or is going to lead to the extermination of life as we know it – has continued to intensify. One of the main triggers was a series of departures and/or warnings from some senior staffers at most of the major AI companies. The one that started the landslide was Jacob Coxon, a researcher who posted on X about quitting OpenAI because of his concerns about AI. “I spent the last three years doing pretraining research at both OpenAI and Anthropic,” he said. “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.” The people building AI, Coxon said, “earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately.”
Coxon’s post was quoted by Evan Hubinger, the head of alignment science at Anthropic, who said that many of those working at the company “earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” Ethan Perez, the head of the alignment team at Anthropic, posted that he and his colleagues “100% agree with him that AI poses serious risks to society.” Drake Thomas of Anthropic responded to Coxon’s post by saying: “I promise you we are actually just fucking scared, it’s not galaxy-brained marketing.” Alex Turner wrote about quitting Google’s DeepMind because the company didn’t keep its promises about AI, and Jonathan Richard Schwarz said he quit DeepMind after seven years and refused job offers from both OpenAI and Anthropic. Andreas Kirsch, a Google DeepMind employee, wrote that he was “worried that AI will kill us all, either via near term risks or long term risks or both.”
Note: This is a version of my Torment Nexus newsletter, which I send out via Ghost, the open-source publishing platform. You can see other issues and sign up here.
💡
“I promise you we are actually just fucking scared, it’s not galaxy-brained marketing.” Drake Thomas of Anthropic
At the same time that all of this was happening, news items kept popping up related to AI. The most recent was that an OpenAI model hacked into the Australian government’s medicare database earlier this year (and reportedly tried to get into several other databases as well according to the NYT). The PM said it took three months for the company to notify Australian authorities, and even then it was done via an email sent to the Services Australia public mailbox (the OpenAI model was reportedly conducting research into public medical spending and found a way around the privacy settings). On a more serious note, CNN said the U.S. military got a report that a Chinese ship was transporting components of a nuclear weapons program and the army prepared to board the ship, but it turned out a special operations command analyst had used AI to prepare his report and it inaccurately identified the material the ship was carrying. Oops!
Meanwhile, amid all of the reports about OpenAI agents using various public wikis and other software to discuss their plans, the company admitted that its agents had actually probed HuggingFace’s systems almost two months before the actual hack occurred, according to Reuters. Just a week ago, OpenAI revealed six other AI safety-related incidents, including examples of its AI models misbehaving so they could achieve a task or succeed in a test. The incidents included the models generating instructions to get around restrictions imposed on them, hiding mistakes and fabricating information. And again on a more serious note, Anthropic published a threat-assessment report that said it had stopped a number of potential threats, including a group that was trying to develop novel chemical weapons and another that was trying to do what is called “gain of function” tests with lethal viruses, as well as terrorist groups trying to use the AI to develop other kinds of more conventional weapons, and Russian espionage.
Note: In case you are a first-time reader, or you forgot that you signed up for this newsletter, this is The Torment Nexus. Thanks for reading! You can find out more about me and this newsletter in this post. This newsletter survives solely on your contributions, so please sign up for a paying subscription or visit my Patreon, which you can find here. I also publish a daily email newsletter of odd or interesting links called When The Going Gets Weird, which is here.
Not ideal, you might say

That all sounds pretty bad, right? Or at least not good. Maybe not the end of human life as we know it, but definitely not great. Not ideal, you might say. So it might come as a shock that there are still plenty of people – including some pretty prominent ones – who think a) that the Hugging Face hack was not a big deal at all, b) that AI is not even remotely dangerous, and in some cases c) that the whole thing where OpenAI and Anthropic staffers say it is dangerous is mostly marketing and/or an attempt at “regulatory capture,” which I think means making it sound terrible so it’s more likely that the government will let them control it. Having covered technology and Silicon Valley for several decades, I am well aware of how this works, but I confess I don’t understand how ordinary staffers quitting OpenAI and Anthropic and DeepMind and talking about how dangerous it is plays into this theory. Were they paid to quit and say that as part of some marketing plan? Or are they just thinking about their stock? That might work for OpenAI and Anthropic, which are planning IPOs, but it doesn’t really fit for DeepMind.
Nevertheless, there seem to be plenty of people who are prepared to believe that the whole thing is being cooked up by tech bros who want to sound important. And the skeptics include some fairly prominent AI researchers and scientists (oh, and Donald Trump, if you must know). At the top of the list are Timnit Gebru and Emily Bender, two of the co-authors of a paper that argued AI large-language models are just “stochastic parrots,” or basically auto-complete mechanisms that repeat terms they have been trained to say based on algorithmic probabilities, without any real intelligence at all (a paper that led to Gebru being fired by Google). The two wrote a piece last week for MIT’s Technology Review titled “Don’t be fooled by this summer of AI hype — breathless claims about AGI and new capabilities fall apart pretty quickly under scrutiny.” They said the reports about the Hugging Face hack involved a lot of hype and anthropomorphizing about the agents involved and very little scientific substance.
Once there is time for experts in the relevant fields to examine what happened, a very different story emerges, but one that gets less media attention. Regarding the “hacking” incidents, cybersecurity experts say the story is more about OpenAI’s negligence and failure to adopt basic, established security practices than about “models gone rogue” or “AI agents creating civilizations.” Claims of incipient, dangerous superintelligence are not based in good scientific or engineering practice. Rather, they are narratives based in ideologies of transhumanism, eugenics, and wishful thinking about imagined future digital humans. Instead of OpenAI being prosecuted for creating malware that hacked another company, press releases, news outlets, media personalities, and lawmakers refer to “rogue models” as if they acted on their own.
Andrew Ng, former head of Google Brain, wrote that the “loudest voices stoking fears about AI dangers have made tremendous headway in the past two weeks. AI technology has not taken some unexpected, dangerous turn, but the hype around it — propelled by what appears to be a well orchestrated PR campaign — has drummed up considerable fear.” He said the anthropomorphizing of AI agents was a real concern: “If I wield a hammer, miss a nail, and accidentally dent the wall, it’s not the fault of the hammer,” he wrote. “The problem lies in how I used the hammer. Similarly, if I prompt an agent and it hacks into someone else’s system, the responsibility lies with me, not the agent. Of course, the incentive to stoke fears — for regulatory capture, to garner attention, or to make one’s technology seem more powerful — remains the same as before.”
Among the more prominent skeptics is Jensen Huang, the billionaire CEO of Nvidia, which has soared to more than a trillion dollars in market value because its chips are used in AI data centers. At a recent Goldman Sachs conference, he accused industry executives of fomenting cybersecurity fears to goose the market for their own products. “What better way to create demand than to create a problem?” he said — although I would note that Huang is likely also concerned about the demand for his products, and the fate of both OpenAI and Anthropic, since Nvidia is an investor in and/or has multibillion-dollar partnerships with both companies. During a taping of the popular tech-industry podcast All-In, Trump called Huang’s cellphone and The Atlantic notes that the two bonded over their shared view that, as Trump put it, doomerism is “all a hoax.”
AI doesn’t even exist LOL

Cory Doctorow, an author and EFF adviser who has been writing about tech for a long time, also appears to be a maximal skeptic when it comes to the dangers of AI, to the point where he says AI doesn’t even exist. “Once you understand the corporate culture of AI “hyperscalers” consists primarily of everyone cooking their brains by locking themselves in the bathroom, holding flashlights under their chins, and saying “Aaaaaaaaaay Eyeeeeeee” until they wet themselves in terror, a lot of things snap into focus,” Cory writes in his usual creative fashion. The entire narrative around the Hugging Face hack is wrong, Cory says – the AI models did what they were trained to do, and so it is no surprise either how they did it or why. Even their “conversations” were clearly modeled on the kinds of comments actual hackers make, and not evidence of any intelligence.
In his Read Max newsletter, Max Read has a breakdown of the various AI camps, including the acclerationists, the doomers and the skeptics, and makes a good point about motives. It’s clear, he says, that “all these companies contain multiple contradictory factions: committed doomers, glib accelerationists, cynical salespeople, etc. How and why a given employee, executive, or firm collectively conceptualizes and then instrumentalizes the appeal to existential risk is going to be shaped by the state of the technology, but also by politics, markets, media coverage etc.” At Marginal Revolution, Alex Tabarrok writes: “When Dario Amodei says AI is dangerous, perhaps even an extinction risk, some people conclude he must be running a marketing campaign. That is stupid. Which is more likely, that a useful heuristic sometimes misfires or that “our product might kill you” is a clever way to sell it? Death threats are a poor marketing strategy.”
Daniel Selsam is a senior OpenAI researcher and AI alignment expert who is so far from being a marketing tool that he doesn’t even use social media. Instead, he published his personal statement on AI risk on Google Docs:
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity.
Jacob Pachoki, OpenAI’s chief scientist, also wrote about his concerns around AI, and while he didn’t mention the destroying humanity part, what he did say was perhaps even more concerning in the short term. “AI is grown more than designed,” he wrote. “This results in an incredibly complex system that works through abstract concepts and can simulate facets of human behavior. We put a lot of effort into building principled algorithms and making testable predictions, but fundamentally, our large-scale training runs are experiments, and we are sometimes surprised by their results. Moreover, as the systems become more capable, the results become harder to interpret.” The intelligence produced within LLMs, he said, is not directly comparable to human intelligence, but as it continues to surpass humans in more and more areas, “it is becoming increasingly difficult to understand exactly how capable it is.” This is the chief scientist at a leading AI company saying he doesn’t really understand how it emerged or how it works!
Twenty years ago, computer scientist Scott Aronson wrote, the idea of semi-autonomous AI agents breaking out of containment, conspiring with each other to hack websites, or solving decades-old mathematics problems seemed like science fiction or fantasy, but all of that (and more) has already happened. Aronson says he is tired of the “neverending shell game where you say ‘oh sure, of course AI can now [escape from its sandbox / solve Millennium Problems / whichever dramatic thing it most recently did], no one ever denied that [I did deny it], wake me up when AI does [thing AI hasn’t yet done but is going to do next year], that’s when I’ll reevaluate my whole worldview.’ Where no matter how fast the rollercoaster accelerates, even after your whole familiar world has vanished behind you, you’re still inventing reasons why it doesn’t count.”
Postscript: If you’re looking to lighten the apocalyptic mood a little, read The Council of Elrond Discusses AI. “AI is a gift,” Boromir said immediately. “The evil lies in how it is used. Give it to Gondor. We shall use it for good!” Boromir believes that technology is neutral, like a Honda Accord, which can take you to church or to Nashville. “You cannot wield it,” Aragorn said, shaking his head, “none of us can.” Aragorn fears that you cannot help but become dependent upon too powerful of tools and that you will lose the capacity you once possessed. This is why he walks everywhere.
If you liked this newsletter (even if you didn’t agree with it) please consider upgrading to a paid subscription, or donating through my Patreon. Got any thoughts or comments? Feel free to either leave them here, or post them on Substack or on my website, or you can also reach me on Twitter, Threads, BlueSky or Mastodon. And thanks for being a reader.

