The last few months in AI risk have seen, among other things, a spiraling slew of cybersecurity incidents and a near-miss China-US war. This article draws on recent events to provide an updated accounting of AI risk.
To summarize: Misuse of AI in the military domain may now be elevating nuclear war risk. In cybersecurity, something is clearly happening, but while it could pose a significant risk to society, some evidence suggests that it won’t. Finally, I remain skeptical of near-term AI takeover catastrophe scenarios, but given the stakes, the risk should be taken seriously even if the probability seems low.
I am not addressing AI biosecurity risk in this article only because I haven’t been following it closely enough to comment, but this is also an important domain.
The types of AI discussed in this article are large language models (LLMs) and LLM agents. An LLM agent consists of an LLM plus another program that interacts with the LLM and other tools. Instead of a human chatting with an LLM chatbot for step-by-step guidance on a task, a human can describe the task to the other program and it will chat with the LLM and carry out steps for completing the task. This is how the AI systems can execute tasks like the cybersecurity tasks that have been in the news.
In the AI field’s jargon, that other program is called a “harness”, though this is a bad metaphor. The LLM is supposed to be the horse: “powerful on its own, but useless for farming without a harness”. However, a horse harness just makes the horse easier for the human to ride. An LLM harness does a lot more; it’s more akin to a horse harness plus an assistant rider who the lead farmer can instruct to carry out agricultural tasks. Despite the faulty metaphor, it is a common term and an important concept for understanding current AI risk.
Some of the recent incidents, including the much-discussed OpenAI-Hugging Face incident, involve interacting groups of LLM agents. These agents are typically multiple instances of the same LLM-harness pair, roughly analogous to multiple automated characters in the same video game. Research has found improved performance on some tasks when agents are trained to play different roles and sometimes even sacrifice themselves to advance overall task completion, analogous to role-playing in groups of humans and other animals.
Nuclear War
My immediate concern is not about rogue AI taking over nuclear weapons systems and launching the weapons at us—it’s about humans misusing AI in ways that lead to nuclear war. In the international security jargon, it’s about inadvertent escalation and inadvertent nuclear war. This is a concern in particular because current LLMs are known to be error-prone, including hallucinating false information as if it was true, and because humans are prone to have an excessive faith in AI outputs, a phenomenon known as automation bias. This can cause humans to engage in escalatory military activity that they otherwise wouldn’t have, with potential for escalation all the way up to nuclear war.
Two recent incidents demonstrate this concern. First, the tragic February 28 US attack on an Iranian school was prompted in part by faulty targeting guidance from the Palantir Maven AI system, per reporting from Bloomberg. Maven might not have been the source of the error: Bloomberg’s reporting indicates that the error instead came from the US military feeding bad information into Maven and then being overly trustworthy of Maven’s outputs. Second, a CNN investigation uncovered a near-miss incident from this past spring in which a spurious AI report almost led the US military to intercept a Chinese ship that the report falsely claimed was transporting components for a nuclear weapons program. The false claim was described as an AI hallucination. The incident occurred in the context of the US-Israel war against Iran, which was motivated in part by concerns of Iran building nuclear weapons. One of CNN’s sources stated that the spurious AI report “almost started a war”.
Both incidents are, at their core, organizational failures of the US military. Under its current leadership, the US military has pursued aggressive AI adoption, slashed its civilian protection staff, and called for “maximum lethality, not tepid legality”. This is a recipe for AI-related harm. Over 100 Iranian schoolchildren have already paid the price with their lives. Unless and until the US military takes a more responsible stance, the risk of nuclear war will, in my opinion, be elevated. How elevated? That is difficult to say. But the clear solution is for the US military to act more responsibly with AI, which it should do anyway. Ditto for other countries.
Cybersecurity
In my July 2026 article, I presented evidence suggesting that headlines about AI cyber capabilities may have been due mainly to corporate marketing hype. A few months later, I can say that while yes, the companies probably are guilty of hype, it is also the case that the AI itself has significant cyber capabilities.
For example, Google reports a sharp uptick in the rate of fixing vulnerabilities in the Chrome browser starting around April 2026 and attributes this to an internal AI tool. Even if we question that as marketing hype, there is no reason to doubt the people who maintain open-source software and report facing an acute challenge: AI systems can now scan the online spaces where the maintainers coordinate to fix vulnerabilities and rapidly convert mere rumors of vulnerabilities into functional exploits. And then there are the tens of thousands of recent incidents of LLM agents escaping from their “sandboxes” during testing, such the Hugging Face incident, which an independent cybersecurity expert describes as “really impressive cyberoffense work”.
That said, there is also evidence suggesting that these AI cyber capabilities may not translate into significant harms to society. Two separate sources document substantial growth in the number of identified cyber vulnerabilities but not in cyberattacks exploiting the vulnerabilities, suggesting that LLMs and LLM agents may be more useful for identifying vulnerabilities than for executing attacks. It has also been reported that cyber insurance companies have had zero losses from AI cyberattacks.
Perhaps harmful cyberattacks will accelerate as the latest LLMs and agents diffuse across the cyber crime sector. Or, perhaps the tools are more useful for cyber defense than offense [1]. Indeed, a lot of cyber crime uses relatively low-tech exploits, such as tricking someone into clicking on a bad link, which may not benefit from advanced AI cyber capabilities. Furthermore, LLMs and agents may be too unreliable and insufficiently stealthy to work well for cyber criminals. Indeed, a closer look at the Hugging Face incident shows erratic behavior: the LLM agents kept repeating certain actions even after the corresponding task had already been completed, they pursued clumsy and inefficient strategies, and they hallucinated nonsensical text [2]. This creates a distinctive attack signature; it likewise suggests countermeasures for defenders such as honeypots.
Another suggestion is that LLMs and agents will matter less for cyber crime and more for cyber warfare. Compared to cyber criminals, countries generally have greater resources and different motivations, with more potential for harm. Attacks on critical infrastructure, especially globally critical infrastructure, are of particular concern. It is believed that Iran is behind recent cyber attacks against US water utilities. Those attacks were not highly sophisticated—water utilities are often poorly defended—but other targets would benefit from advanced cyber capabilities. However, the erratic and unstealthy nature of LLMs and agents could again make them undesirable to attackers by threatening mission success and plausible deniability. They may be most attractive to terrorist groups and other rogue actors who just want to see the world burn.
Or, perhaps the harm will come mainly from AI companies continuing to fail to contain their own products during testing. Cybersecurity experts present compelling evidence that AI companies have failed to take basic cybersecurity measures and that the companies likewise display “a pretty profound unfamiliarity” of cybersecurity. Other testimonial describes AI companies as having a “Wild West” culture with “a ton of pressure to move really quickly”. The companies may well be more reckless and harmful than cyber criminals and hostile nations. Government regulators should take note.
Takeover Catastrophe
The erratic behavior of LLM agents in the Hugging Face incident, as described above, is at the heart of why I remain less concerned about AI taking over the world, or for that matter taking over the internet, as has also been recently discussed. The prospects even seem week for more limited scenarios in which humans lose the ability to shut down one or more LLM agents.
In previous articles, I explained the “bag of heuristics” theory, which posits that LLMs mainly consist of a large number of scattered facts about the world instead of a more elegant and coherent representation of how the world works. This theory can explain, among other things, LLMs’ tendency to hallucinate made-up facts about the world, such as the “fact” that that Chinese ship was transporting nuclear weapons components. A well-designed harness can sometimes overcome this limitation of LLMs—for example, if an LLM hallucinates a book title, the harness could look it up in a book database and correct it. However, the erratic behavior in the Hugging Face incident indicates that state-of-the-art LLM agents are still making heavy use of simplistic, error-prone heuristics. This is to be expected: the latest agents are not said to involve a radical change in AI system design. As my previous articles detailed, I am skeptical about the potential for takeover by AI bags of heuristics.
Another line of evidence comes from comparing human and AI performance. In software engineering, AI has substantially increased the amount of code produced by human workers, but the effect on shipping usable software has been much smaller, apparently because the AI lacks capability in other parts of the software engineering process. Similarly, in mathematics, AI can now complete advanced proofs, but are humans still the ones developing the theory and setting research directions. In both cases, AI functions mainly (perhaps entirely) by applying existing, human-developed ideas and techniques. This is seen in OpenAI’s notorious scooping of a rival group’s idea for the Navier-Stokes problem in mathematics and in AI systems generating cybersecurity exploits from rumors of vulnerabilities in open-source software. AI likewise performs poorly in figuring out what to work on in the first place and in managing AI research projects. These examples all come from computing and mathematics, two domains where AI performance has been relatively strong. In the domain of playing simple puzzle video games, the latest LLM agents can now match human performance, but only with $20,000 of computing power, compared to the roughly $100 given to human participants. This evidence further suggests limitations of LLM-based AI systems that make takeover less likely.
In my view, much of what LLM agents do appears to be a brute-force style of work. They process large numbers of heuristics, engage in wide-ranging trial-and-error, and burn through massive amounts of computing power and electricity—hence all the new data centers. The use of groups of agents enhances the process by enabling brute force work in parallel. It turns out, much can be accomplished with this style of work, but with limitations that leave me skeptical of its ability to take over.
The massive computing power needed is a substantial impediment. LLM agents are not like the typical computer virus that can readily self-replicate and spread itself around the internet. In the recent cybersecurity testing incidents, there is an important sense in which the humans of the AI companies did not lose control: the LLM agents remained within the companies’ data centers, which the humans could simply turn off at any time. The incidents are less akin to a prisoner breaking out of jail and more akin to a prisoner surreptitiously gaining internet access from a computer in the jail. There is talk of LLM agents “breaking out of jail” (in the AI jargon, “exfiltrating their weights”), but that is a nontrivial step beyond the status quo.
Furthermore, for agents to be self-sustaining or “self-sovereign”, they may need to generate revenue in the human economy to pay for their computing power, where they would be competing against teams of humans that can have access to their own LLM agents. Given the comparative advantages of humans discussed above, independent LLM agents may struggle to sustain themselves, making it that much easier for humans to shut them down if we so choose.
Concluding Thoughts
One might protest that my skepticism about cybersecurity and takeover risks is overly rooted in past and present AI capabilities, whereas the biggest risks come from more capable future systems. I certainly agree that AI technology continues to change and that future AI systems could be considerably more capable. There is significant uncertainty here. Likewise, given the stakes, I strongly believe that the more extreme scenarios should be taken seriously in both research and governance. This also applies to nuclear war scenarios. Even if the probability seems low, it can still be a large risk.
That said, I do also believe in the importance of rigorous analysis of the risks. This includes a close look at what is already happening because it can be very relevant to what might happen in the future. For example, my initial exploration of the “bag of heuristics” theory was for LLM chatbots, but it remains relevant for LLM agents. With AI risk becoming a high-level policy issue and topic of public debate, it’s important for discussions of AI risk to be well-founded.
***
[1] AI and cybersecurity expert Joshua Saxe posits that it may be “a net positive to release those [AI] models faster” because doing so helps cyber defenders more than attackers.
[2] See p.10 of the report Hugging Face Incident Initial Post-Mortem written by a large group of AI and cybersecurity experts with the Cloud Security Alliance.
Image credit: lemmling




