As regular readers will know, last week I re-published a piece that I originally wrote on September 3rd, entitled “Yes, we should be very worried about the Hugging Face hack.” I couldn’t write a new post for last week because I was backwoods camping in a remote area of southwestern Ontario – Killarney Provincial Park, in case you are interested – and I had virtually no cell signal for most of that time (if you like canoes or kayaks or sunsets or trees, you can check out the post I wrote about the trip). I chose to re-publish the Hugging Face hack piece because it seemed even more appropriate in light of some of the revelations that have been coming out about the hack – including the fact that the hack was neither the first time OpenAI models had accessed another company’s systems without permission, nor the first time the models in question had used a message board to discuss the methods used in their attack. The September 3rd piece was itself an update on an earlier piece I wrote in July, right after the Hugging Face attack was first revealed (a reveal that came weeks after the hack itself took place).
The debate over AI safety and what AI researchers call “alignment” – or, to put it more bluntly, between whether AI is working as advertised or is going to lead to the extermination of life as we know it – has continued to intensify. One of the main triggers was a series of departures and/or warnings from some senior staffers at most of the major AI companies. The one that started the landslide was Jacob Coxon, a researcher who posted on X about quitting OpenAI because of his concerns about AI. “I spent the last three years doing pretraining research at both OpenAI and Anthropic,” he said. “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.” The people building AI, Coxon said, “earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately.”
Coxon’s post was quoted by Evan Hubinger, the head of alignment science at Anthropic, who said that many of those working at the company “earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” Ethan Perez, the head of the alignment team at Anthropic, posted that he and his colleagues “100% agree with him that AI poses serious risks to society.” Drake Thomas of Anthropic responded to Coxon’s post by saying: “I promise you we are actually just fucking scared, it’s not galaxy-brained marketing.” Alex Turner wrote about quitting Google’s DeepMind because the company didn’t keep its promises about AI, and Jonathan Richard Schwarz said he quit DeepMind after seven years and refused job offers from both OpenAI and Anthropic. Andreas Kirsch, a Google DeepMind employee, wrote that he was “worried that AI will kill us all, either via near term risks or long term risks or both.”
Note: This is a version of my Torment Nexus newsletter, which I send out via Ghost, the open-source publishing platform. You can see other issues and sign up here.






















