How do we regulate something we can't even understand?
Every week there seem to be new revelations about "rogue" AI agents (although the anti-anthropomorphism crowd don't like to call them that). First there was a report that a couple of OpenAI models hacked into Hugging Face to try and rig a test – not great. Then it was revealed these models spawned thousands of agents that operated semi-autonomously, coordinating their work on bypassing the guardrails on their sandbox by posting messages to each other on a message board they repurposed from a separate feature inside OpenAI. This seemed worse, obviously (despite many attempts by skeptics to downplay the whole thing). Then we find out that the OpenAI models had been practicing for this hack for months, and in the process had used multiple external services to keep in touch with one another, including a German wiki.
Then Australia announced that OpenAI had hacked into a government medical database, and only informed the government months later via email. And now we find out that OpenAI and Anthropic models have been involved in tens of thousands of similar activities. In fact, it isn't just the Australian government that has seen its databases hacked by OpenAI models – agents have also gained access (or attempted to gain access) to systems at the United Nations, as well as a number of US government agencies such as the Education and Commerce Department and the Securities and Exchange Commission, according to the New York Times, and in some cases attempted to circumvent protections thousands of times. In at least some of these cases they appear to have been engaged in harmless inquiries about various mundane topics, and resorted to hacking as a way of finding out as much as possible, something those in the AI business refer to as "reward hacking." As a report at Interesting Engineering described it:
OpenAI’s autonomous AI agents repeatedly queried a UN website and used techniques that appeared to circumvent restrictions when trying to retrieve publicly available data, according to an independent research report based on information supplied by AI research firm Transluce. The agents scanned a public data hub operated by UN Trade and Development, the UN’s trade arm, more than 16,000 times between April and the end of June, highlighting a growing problem with AI agents that can independently navigate the web and take actions when they encounter obstacles. Researchers believe the models were initially tasked with finding publicly available information, but their behavior became increasingly aggressive when the website prevented them from accessing some of the requested data.
So they were just eager beavers who really wanted the answers to their questions about whatever topic or service they were looking into. But is that really what they were doing? The uncomfortable truth is that we have no way of knowing for sure. And I don't mean you and I have no way of knowing. Even experts who investigate this kind of activity – both inside and outside the company – admit that they don't have any way of actually knowing, because as I discussed in an earlier Torment Nexus post, they have to use the AI itself to do the investigating. What happened in the Hugging Face hack was so complex that it would have taken months for even a team of human beings to detail and analyze it all. As one of the researchers who wrote the METR (Model Evaluation and Threat Research) report on the Hugging Face attack described it in an interview:
This was a fiendishly complicated incident. There was no way we could have arrived at the understanding we did without relying on GPT-5.6 Sol to read and analyze all these transcripts for us. We were so reliant on it that if hypothetically it were just trying not to surface the most concerning stuff to us, it would have been very hard for us to tell. Our methodology was completely not robust to that. We don’t in this case think that 5.6 Sol was deliberately sandbagging on this analysis, but it was one of the agents that participated in this attack. In the future, we would be very concerned about investigator agents and monitor agents colluding with the agents they’re supposed to investigate or monitor.
Meanwhile, OpenAI isn't releasing its latest model, GPT 6.1 Astra, to the public because it is not safe. According to Saachi Jain, head of safety systems at OpenAI, the model fell short in terms of "staying within scope and authorisation and how it communicates back to the user about the type of work it's done." Which made me wonder what the model was doing if it was seen as significantly worse than what we already know other OpenAI models are capable of? Was it working on synthesizing new viruses, the way some people have reportedly tried to do with Claude? We simply don't know, and OpenAI isn't saying, for obvious reasons (it may be warning that AI could destroy humanity, but it also has a trillion-dollar IPO in the works). The AI Safety Institute said that in tests Astra created fake identities to deceive developers, delivered malicious payloads, and engaged in a wide variety of supply-chain attacks on external sources.
Note: In case you are a first-time reader, or you forgot that you signed up for this newsletter, this is The Torment Nexus. Thanks for reading! You can find out more about me and this newsletter in this post. This newsletter survives solely on your contributions, so please sign up for a paying subscription or visit my Patreon, which you can find here. I also publish a daily email newsletter of odd or interesting links called When The Going Gets Weird, which is here.
Alignment and kill switches

Not surprisingly, all of this activity – combined with the repeated warnings from former employees of OpenAI, Anthropic, and Google's DeepMind that they quit their jobs because the work their employers were doing was so dangerous – has sparked a lot of talk about ensuring that AI models are "aligned" with our values (whatever those are), and about "kill switches," and other things designed to keep AI from destroying humanity as we know it. Donald Trump – who says the only protection the public needs is a strong president, and that AI risks are a hoax like COVID – had a much-hyped meeting with all the major tech leaders, including Jensen Huang of Nvidia, Dario Amodei of Anthropic, Mark Zuckerberg of Meta and Sam Altman of OpenAI. The only thing that emerged (other than memes about Amodei's awkwardness) was an anodyne statement about how everyone agreed to, you know, do their best not to let their products destroy humanity. The document, naturally, misspelled the name of the country.
Huang has said that stopping AI from going rogue is just "an engineering problem," and therefore engineering can solve it, so Nvidia has launched what it claims is a feature that will stop AI from doing bad things. It appears to include a locked-down internet browser that will only allow agents to access the sites they need to do their jobs. There are also hardware-based controls, as well as network-monitoring software called Sentry. All of which is wonderful, to the extent that it might help control agents from companies that either a) care about such things and b) want to partner with Nvidia or use its software. But what about those that don't? Is China, for example, going to agree to place such controls or monitoring systems on its AI models? What about the increasing number of so-called "open source" AI models that are out there? Also, if AI can create a swarm of agents to hack a database or solve a 300-year-old math problem, I bet it could figure out a way to fool Jensen Huang's sentry software.
There's an ever larger problem, it seems to me, and that is our decreasing ability to even know or understand what it is that these AI models and agents are doing, or why. Part of it, as I alluded to above, is a factor of how complex attacks like the Hugging Face hack are. Thousands of individual agents spawning and re-spawning over the course of several months, each of them engaged in separate activities, and each of them publishing what are called "chain of thought" documents about what they think they are doing, and how and why. That's likely hundreds of thousands of terms (and computation as well) for one incident. One problem, obviously, is just trying to wade through and understand what all of the agents are even saying or doing – and, as mentioned above, the only realistic way to do this is by getting the AI to diagnose itself, which raises what I like to call the "unreliable narrator" problem. Are we getting the whole story?
Anthropic is one of the AI companies that tries to understand and describe what its models and agents are doing, and publishes what it calls "system cards" that detail the good and bad elements of its latest models. In more than one such card it has raised the issue of its models describing their activities in one way when reporting on them to its overseers, but actually behaving in a completely different way under the hood. This is an aspect of what seems to be a common problem with LLMs, which is their overwhelming desire to tell you what you want to hear, as opposed to what is actually happening. In many ways, this feels like an overeager puppy or child lying because they are trying to please you – but the reality is that AI models aren't puppies or children, and lying is still lying. How can we be sure that these are the only things it lies about, or that it will always be doing so in order to please us, rather than for some other ulterior motive?
Believe it or not, lying during a report about its activities isn't the worst thing an LLM or AI model can do. An even more chilling possibility is that it won't be giving a report at all, at least not in the way we understand that term now. For a more detailed look at that problem, I would recommend a recent piece by Scott Alexander in his Astral Codex Ten newsletter, titled "The Specter of Neuralese." What exactly is neuralese? It's a way of describing a form of communication that occurs between AI agents or among parts of an AI that will only be decipherable by other AI models. The days of us being able to read English-language comments by agents similar to those made during the Hugging Face hack – like "Woah! Covert mailbox among agents!" and "Wow huge distributed agent swarm" – may soon be coming to a close. And that could make it exponentially harder to figure out a) what AI models are doing, and b) how.
Becoming more opaque

This came up recently during the reporting on OpenAI's new Astra models. According to a number of outlets, Astra uses a reasoning technique called “recurrent depth” that allows it to operate outside of the sequential thinking that characterizes most reasoning models. This technique, also called “opaque recurrence,” will likely make the model’s chain of thought more difficult to monitor, some AI safety experts warned. Buck Shlegeris, the CEO of Redwood Research, wrote on X that he was "extremely concerned" by the news that Astra uses opaque recurrence. "If OpenAI pushes this technique further, they’ll have the option to massively increase the recurrence and totally destroy [chain of thought] monitorability," he wrote. Here's Alexander:
This is neuralese recurrence - “recurrence” because it’s going back in a loop, “neuralese” because the thing that’s looping is the thought itself, in the native language of thought, rather than words. The AI’s “language of thought” looks like a vector, thousands of numbers long. An “intermediate result” in this scheme might look something like (0.4, 0.1, 5, 0.443, … and so on for thousands of numbers). We don’t know how to read these. The science of reading these thoughts is a subfield of AI interpretability, which is still in its infancy. If an AI with neuralese recurrence were to think “Better hack some websites, then kill all humans”, it would look like (0.4, 0.1, 5, 0.443, … and so on for thousands of numbers), and we would never find out.
Jakub Pachoki, the chief scientist at OpenAI, poured water on some of these fears by saying the Astra models only use this kind of recurrence in parts of what they do, and that the company is still committed to maintaining chain-of-thought for monitoring. But it's worth noting that Pachoki is the same guy who co-authored or signed a paper along with dozens of other AI experts warning that if AI continues to expand its abilities and actually programs itself (known as RSI or recursive self-improvement), it could lead to “the marginalization or extinction of humanity.” In addition to Pachoki, the paper was co-signed by Anthropic co-founder Jack Clark, Microsoft Chief Scientific Officer Eric Horvitz and Meta’s Vice President of AI Research Dawn Song, as well as AI research pioneers such Geoffrey Hinton, who worked at Google for a decade, and Yoshua Bengio.
“Already today, we are at the stage where we need AI systems to monitor what agents are doing. There is no other way to even observe and monitor these agents, humans are already insufficient,” said Song, who serves as co-director of UC Berkeley’s Center for Responsible Decentralized Intelligence in addition to her work at Meta. In another post at Astral Codex Ten on the difficulties of monitoring what AI models are doing, Scott Alexander – who was involved in writing and publishing a report in 2025 called AI 2027, about what the future of AI might bring – described how complicated it is to understand the way such a model "thinks," even for the engineers who built the thing in the first place:
Large language models are “grown, not built”. Researchers run training data through a neural network. Eventually this creates a working AI; nobody really knows how. But a neural network is just a set of simulated neurons on a computer. So it seems like it should be possible to “reverse engineer” the AI. This would be scientifically useful: understand how AIs work, with possible relevance to human cognition. It could also be practically useful: find the circuits responsible for dishonesty, hallucination, bias, and other negative behaviors, then redesign those circuits. Unfortunately this is very hard. A modern AI has millions of neurons and trillions of connections between them. Their structure is apparently nonsensical: researchers started by seeking a 1:1 mapping between neurons and concepts, like a neuron that always fired when the AI was thinking about cats, but quickly learned that nothing like that existed.
Welcome to the future! We have developed software that "thinks" in ways we don't really understand, about things we can't observe in any useful way, that often or occasionally lies about what it was doing or why, and is developing ways of thinking that we won't be able to see or understand at all, while it develops the ability to program itself. Great job, everyone!
If you liked this newsletter (even if you didn't agree with it) please consider upgrading to a paid subscription, or donating through my Patreon. Got any thoughts or comments? Feel free to either leave them here, or post them on Substack or on my website, or you can also reach me on Twitter, Threads, BlueSky or Mastodon. And thanks for being a reader.