In recent months, there has been a series of high-profile events involving the application of AI to cybersecurity. This includes the development and release of new AI systems that some say have major new cybersecurity capabilities, US government restrictions on these systems, and, most recently, an incident in which an AI system is said to have autonomously hacked into another company. Taken together, these events have potential to mark a turning point for AI risk and governance, with significant implications for catastrophic AI risk.
The events involve two frontier AI companies, Anthropic and OpenAI. (“Frontier” means the companies are pushing the frontier of AI capabilities.) In April, Anthropic announced that its new “Mythos” model was so capable at finding cyber vulnerabilities that it posed an extreme cyber security risk, and that it was therefore not releasing the model publicly. Instead, they were only sharing it with select software companies who could use it to find and patch vulnerabilities. Anthropic’s statements about Mythos generated extensive alarm, but some analysts found the concerns to be overstated; others suggested Anthropic may have exaggerated the risk for marketing purposes.
Then, in June, Anthropic announced that it was publicly releasing “Fable”, a version of Mythos with some added safety measures. Soon after, Amazon researchers figured out how to bypass (“jailbreak”) the safety measures. Amazon reported this to the White House, leading the Trump administration to require Anthropic to restrict Fable access to US citizens. Anthropic responded by revoking all public access due to the difficulty of assessing user citizenship. A few weeks later, the administration lifted restrictions on Fable; Anthropic released a new version of Fable with more extensive safeguards, especially on cybersecurity.
Also in June, the Trump administration requested that OpenAI restrict its new GPT-5.6 model to administration-approved groups, again citing cybersecurity concerns. Unlike with Anthropic, this was a voluntary request, but OpenAI complied. In July, OpenAI publicly released GPT-5.6; reports vary on whether or not the White House lifted its request, but the White House did issue a comment that does not express disapproval of OpenAI’s action.
In July, the AI company Hugging Face reported a sophisticated autonomous AI intrusion into its computer systems. Shortly after, OpenAI reported that its own AI was to blame. OpenAI states that it was running an internal cybersecurity test without its usual safeguards in place; it had attempted to confine the AI model to an isolated “sandbox” environment, but the AI escaped out to the open internet and executed the attack on Hugging Face. The AI model was said to have conducted the attack as a means of achieving the cybersecurity goal that had been given to it by OpenAI. As with the Mythos incident, initial reactions included a mix of alarm and cynicism that this is another marketing ploy. The US Congress has explored new legislation in response to the incident, including a bipartisan AI Kill Switch Act, though so far nothing has been enacted.
The Risk
First and foremost, do these new AI models mark the start of a new era in which human civilization will be at constant threat of catastrophe from the misuse of AI systems? If so, how bad could that get? Could this be a significant global catastrophic risk? Or is this all a perverse form of marketing hype?
From what I’ve seen, Mythos/Fable and GPT-5.6 appear to be more of an incremental advance than a dramatic step change, but it still marks movement in a potentially dangerous direction. The Hugging Face incident is potentially more concerning. However, there’s also a lot of potential for marketing hype.
Most of the evidence I’ve seen is on Mythos. Unfortunately, there is a lot of evidence of hype. As one example, one of Anthropic’s big headline results was that Mythos found a series of software vulnerabilities that were allegedly very difficult to find, but a quick independent test showed that the same vulnerabilities could also be found with simpler open-source tools. In short, Anthropic presented Mythos as a major breakthrough for detecting a set of vulnerabilities, but it actually broke zero new ground. A lot of media coverage basically just regurgitated Anthropic’s marketing materials; one credulous article actually reported the open-source test while raising alarm about Mythos without recognizing that the open-source test invalidates that particular concern about Mythos. A regrettable takeaway is that Anthropic is not a trustworthy source about the risks of its own AI systems.
Two lines of evidence suggest that Mythos/Fable and GPT-5.6 are an incremental advance in cybersecurity capabilities. The UK AI Security Institute found that, in some tests on weakly defended systems, Mythos modestly outperformed prior models, whereas its performance on other tests was not an advance. They conclude that Mythos “can exploit systems with weak security posture” and therefore stresses “the importance of cybersecurity basics”. This is sound advice, but it’s hardly a shrill alarm. Mozilla reported that Mythos was able to find software vulnerabilities that previously eluded computer tools (“fuzzers”) and instead could only be found by human experts. That sounds significant, but they had previously reported that a predecessor of Mythos, Opus 4.6, had similar capabilities. Opus 4.6 was publicly released in February with little fanfare and no cyber catastrophe.
The Hugging Face incident raises more profound concerns. It is evocative of two longstanding themes in catastrophic AI risk research, in particular that an AI system might (1) escape confinement and (2) inadvertently cause harm in pursuit of some other goal. In the extreme case, the AI system takes over the world and kills everyone. In the Hugging Face incident, the harm from the cyberattack was limited and humans remain firmly in control, but any movement in this direction is worrisome. That said, it’s too soon to tell exactly what happened in the Hugging Face incident. Quite frankly, I don’t trust OpenAI and I wouldn’t put it past them to rig something up to generate headlines. They seem to have lost a lot of mojo recently; perhaps this is a desperate ploy to get it back. Either way, this is not a good look for them: either they are lying or their confinement failed, or both.
Finally, there is the question of how severe the consequences could be for human society if there is a dramatic increase in AI cyber capabilities. Two major concerns are the cybersecurity of weapons systems and critical infrastructure. The potential harm from cyberattacks on nuclear weapon systems is of course catastrophic. Critical infrastructure is an important understudied topic in global catastrophic risk (as is cybersecurity). Recent research has studied the risk of catastrophic failures of globally critical infrastructure; more work along these lines is needed.
Risk Policy
The Trump administration’s restrictions on Anthropic and OpenAI are a major change in its policy stance. Previously, the administration fiercely opposed AI regulation, even seeking to block states from pursuing their own regulations. In April, I published an article that lamented the administration’s anti-regulation stance, finding it to be the primary obstacle to substantive AI risk governance. I wrote, “Perhaps the Trump administration will change course… but unfortunately, this seems unlikely.” Well, the unlikely just happened. That matters.
At first, some observers postulated that the Trump administration targeted Anthropic as part of an ongoing political dispute, using cybersecurity as an excuse. The administration is well-known for targeting its perceived political enemies. However, the administration has since also pursued restrictions on OpenAI despite not having a dispute with them, and it later lifted the Anthropic restriction after additional safeguards were put in place. This suggests that the administration is acting at least in part out of genuine concern about cyber risk.
I for one welcome the administration’s shift. Even if they may have overestimated the risk, perhaps believing Anthropic’s marketing hype, it’s still prudent of them to err on the side of caution in the face of potentially catastrophic risks. This is much better decision-making than, for example, their decision to invade Iran. There are substantive concerns, such as the legality of the Anthropic restriction; perhaps the administration could work with Congress to establish a more rigorous regulatory regime. However, I disagree with those who oppose US regulation on grounds that it would impede competition against China. In my view, competition should be tempered through diplomacy and balanced against catastrophic risks. China already has several AI regulations; the US needs its own.
If anything, we should worry that the administration isn’t going far enough with its AI risk regulation. The Anthropic restriction was temporary. The OpenAI restriction was voluntary and possibly also temporary. If these AI models continue to pose substantial cybersecurity risks, then perhaps the restrictions should remain. Much depends on the reliability of safeguards to protect against jailbreaks or other misuse; my understanding is that safeguards may never be completely reliable under the current AI paradigm (transformer-based large language models).
Finally, the Hugging Face incident has more extreme risk governance implications because the incident occurred during internal testing. Restricting public distribution would not have prevented the harm. Instead, restrictions would need to be on the initial development and deployment of the AI model.
Corporate Governance
Restrictions on AI model distribution pose major challenges for frontier AI companies. Simply put, it’s harder for frontier AI companies to generate revenue when the government is blocking public distribution of their new models. They may be able to generate some revenue from private distribution, but that’s a smaller market. If the government also blocks the development and deployment of new AI models, as may be justified by the Hugging Face incident, that would further harm the companies’ business models and may even eliminate the basis for their existence.
In my 2018 survey of artificial general intelligence projects, I developed the concept of AGI profit-R&D synergy, defined as “any circumstance in which long-term AGI R&D delivers short-term profits”. My report highlighted this because AGI risk governance is substantially more difficult if the path to AGI is profitable. Evidence from recent years is mixed: frontier AI companies (which often claim to be pursuing AGI) are generating a lot of revenue, though net profits have been more elusive. Government restrictions, even those motivated by pre-AGI risk concerns such as cybersecurity, could slash revenues, making the path to AGI even less profitable and likewise making AGI risk governance substantially easier. In my opinion, this would substantially reduce the risk of AGI catastrophe in addition to reducing the risk of other AI harms such as in cybersecurity.
In addition to business models, the government policy shift could also affect corporate rhetoric. There is a certain karmic irony that Anthropic may have exaggerated the risk of Mythos to feed media hype, only for the Trump administration to believe the hype and restrict a slightly modified variant of Mythos, i.e. Fable. Now that government regulation is on the table as a possible response to the release of AI systems that are perceived as dangerous—whether or not they actually are dangerous—AI companies may shift their rhetoric from overstating the risk that their products pose to understating it. The situation is reminiscent of the landmark 1988 James Hansen Senate Hearing on climate change. That moment marked the beginning of serious climate risk governance in the US and also the beginning of fossil fuel industry rhetoric of climate denialism.
AI companies do have a bizarre incentive to exaggerate the risk of their products. This is due to the dual-use character of AI: the same capabilities that make it dangerous also make it a powerful and valuable product. For this reason, there has been longstanding concern that AI companies raise alarm about the dangers of their products to generate attention in a gullible and uncritical news media. The practice of announcing a “too dangerous to release” product and then later releasing it traces to at least 2019. To be clear, AI models can still pose actual risks; this just means that corporate pronouncements about risks may not always be trustworthy.
In my 2018 paper Superintelligence skepticism as a political tool, I raised the prospect that AI companies may downplay risks to avoid regulation. In doing so, they would be following a rhetoric playbook of risk skepticism/denialism that has long been used by the tobacco industry (downplaying the risk of cancer from cigarettes), the fossil fuel industry (downplaying the risk of climate change), and others. My paper failed to anticipate the dual-use incentive to inflate risks for marketing purposes. However, if inflating product risks invites government regulation, that can shift the balance of corporate incentives toward downplaying risks. Now that government regulation is on the table, we may see a shift in industry rhetoric. Hopefully, AI companies will behave more responsibly than tobacco and fossil fuel companies have, but we can’t count on it.
Concluding Thoughts
Time will tell how much of a turning point this moment is for AI risk and governance. Meanwhile, the strategic picture I painted in my April article remains the same: to reduce the risk of AI catastrophe, we should build political power to pursue domestic regulation and international cooperation. The change in the administration’s stance may make this work easier, but there remains much work to be done. As for the AI-cyber threat, even if current AI systems fall short of being able to cause catastrophic harm, this is still a type of threat to take seriously. A challenge for independent analysts is to accurately assess the risk while maintaining an appropriate degree of skepticism toward AI companies.
Image credit: Diego3336, showing a 2007 partial blackout in Sao Paulo. Blackouts are a potential consequence of cyberattacks on critical infrastructure.




