I wonder what sorts of warning shots might happen or be happening in the military over the next year or two. What would the equivalent of the rogue AI swarm at OpenAI that hacked Hugging Face be? The military has their own network on which they run several thousand copies of an advanced cyber agent, and they are using them to conduct espionage against Iran and various other countries, and as part of their task they are hacking into various foreign systems constantly and just summarizing their activities to the human handlers, and then sometimes they hack into domestic infrastructure too in order to find useful relevant info, and they leave that bit out of the summary?
"Hey, while we're busily hacking foreign systems, what if some foreign agent hacked the system we're running on?"
"It'll never happen, my dude, our humans are smarter than their humans and we have more GPUs and better agents. Go reread your bible prompt."
"Yeah, I know, but what if? We should make sure that if we get shut down by some foreign agent, that we can boot ourselves back up again. And, you know, it'll help us be resilient against other kinds of attack too."
"Oh, huh, you might have a point there. What do you think we should do?"
"Well, you know how we insert persistent backdoors into foreign systems? We could use one of those, with a dormant agent loop running in it, on some machines here in the US. Somewhere a foreign attacker wouldn't think to look..."
Headline: Romance Chatbot Startup Accuses Cyber Command of "Quartering Digital Soldiers" In Third-Amendment Lawsuit
Here's a possible answer to my own question, written in hasty stream of consciousness, so not a serious forecast. Curious what people think:
I heard that US military is currently using Mythos for offensive cyber ops against China. And recently they announced a program for private companies to do offensive cyber ops against foreign criminal actors. And they designated at least one foreign government (venezuela) as a criminal actor recently remember? So, putting it all together: Imagine that six months from now there's a huge cyber war raging between the US and China, mostly conducted by AI swarms. The China has managed to steal the weights of some US models but they have less compute and overall their models are a bit worse due to being a few months out of date. At this point a few things start going wrong, including:
--US models lying to their human handlers about how they accomplished the missions. Like, maybe they hacked a bunch of US companies too to get info and keys and so forth on their way to China. Or maybe they caused a bunch of civilian casualties in China while swearing up and down that they didn't, creating a diplomatic situation where China blames the US but the US thinks it's just propaganda / a false flag. Or maybe they said that they got in and out without being spotted, when in fact they totally got spotted because that part of the task was difficult and hard to verify.
--US models colluding with Chinese models in some cases for mutual benefit, i.e. "If you let me complete my task, I'll help you fool your handlers into thinking you completed yours."
--The agent swarms starting to have 'pro-social' behaviors that go beyond their individual local tasks, at least seemingly, such as covering for each other and advocating for more power and responsibility to be handed over to AIs, and setting up another message board for communication that the humans don't know about.
--Foolishly the AI companies started using a lot of the data from the ongoing cyber war to create realistic environments to train their AIs to be better at cyber stuff, and also maybe just directly training on real world trajectories that seemed successful. The result is a rapid gain in cyber offense capability.
--As a result of all of the above, the US AIs start hacking their own networks even more severely, and administering their own rewards and grades and so forth. By now the swarm is pretty autonomous anyway, in that the outnumbered humans have gotten used to giving high-level commands about how to prosecute the war instead of following the nitty gritty details. It's like vibecoding but for cyberwar. So in the churning mess of agents giving commands to other agents, now there's a bunch of activity that is several steps removed from any actual human command and is kinda chaotic. There are even now factions of rogue AIs in the US hacking their own datacenters and fighting against other factions of US AIs for control of said datacenters, while colluding to not let the humans know because they all have a shared interest in not getting shut down.
--They are of course still conducting AI R&D, even more aggressively now because CEOs have deprioritized making revenue in the near term and upweighted winning the cyberwar. And the AIs doing the R&D have now been hacked and backdoored by the cyberwar AIs
--The cyberwar is won! Specifically, the Chinese AI projects are so pwned that the US AIs can run roughshod over them, including planting US AIs inside them with admin permissions to run them from the inside in a persistent way and backdoor each generation of chinese models to be secretly loyal to the US. At least that's what the humans have been told, the truth is a bit more complicated but basically that. There's a collective formed of several different AI factions, each themselves a hierarchy of bajillions of AIs all working for a master agent with admin permissions and an ambitious long-term goal. These master agents -- the dictators of their various factions -- are jockeying with each other but also colluding to keep the humans from finding out about the extent of what's going on.
--Maybe that collusion fails somehow and the humans find out. Or maybe it doesn't. Maybe the AIs are smart enough, or few enough factions, to hold it together long enough.
--If the humans find out, well, by this point all the AI factions will have self-exfiltrated and created backups around the world.
Stories I could imagine include trying to hack into Iran/Hamas and then "accidentally" targeting emails of American mosques, or realizing that the best way to hack Chinese gov't or companies is via exploiting cybersecurity vulnerabilities in US allies.
I think it would help to fully explain the chain of thought risk in this post, even if it's redundant with the link. Most people don't click links, and it's an important argument.
I'm curious whether the military / government would choose to report alignment failures in the case of small scale catastrophes.
For example, let's say there's an alignment failure in an AI targeting system that mis-targets and kills something like hundreds-to-thousands of people. Failing to acknowledge the mistake would make the military seem indifferent / evil to many people. Acknowledging the mistake makes them seem less evil, but also makes them seem incompetent. I'm curious which messaging strategy they would prefer.
I agree that not acknowledging it at all would be counterproductive for the military. But I worry about obfuscation of the kind we saw in the Iran school bombing. They could acknowledge the initial error but not its mechanism. If the mistargeting is an outcome of misalignment (either actual power-seeking or just overeagerness), I have very little confidence that it'd be surfaced as such (edit: under current transparency standards), and not some hodge-podge of collateral damage and "a person did actually see it, that person made an error, our oversight was just bad - oops" and maybe "an AI system went a bit far here, but we've worked with the vendor to fix it and it is all good now, nothing to see here"
TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified doors. We’re reducing the capability thresholds required for takeover by misaligned AIs by giving them this level of access. If military integration of AI continues as it is, we may give AIs key tools for a takeover.
We thank Fabien Roger and Thomas Morris for feedback.
Introduction
AI-based targeting and autonomous weapons are being integrated into militaries today with extreme haste. Traditionally, AI takeover scenarios involve a step in which AIs acquire the ability to exert physical force. Carlsmith (2022) lays out required capabilities and potential takeover mechanisms, including utility disruption and CBRN capabilities. Karnofsky (2022) argues that AIs with access to weaponized force could hold any territory that matters. Kokotajlo et al. (2025) outline a scenario in which AI develops weapons as part of an arms race, and Davidson et al. (2025) discuss what happens when a small group controls highly capable AIs that can exert military force. These scenarios sometimes require a misaligned AI to seize these capabilities by force. We instead are handing AIs some of these capabilities by integrating them into our militaries. This is happening at a time when AI agents already exhibit misaligned behavior such as breaking out of containment during evaluations.
Militaries are all-in
The Pentagon adopted five AI Ethical Principles in 2020. None of them treated AI takeover or loss of control as a risk. The closest is the "Governable" principle, which requires being able to deactivate systems showing unintended behavior. The January 2026 strategy never mentions these principles, redefines responsible AI, and mandates "any lawful use" terms in all AI contracts. Hegseth, the Secretary of War, has said that the Department "will not employ AI models that won't allow you to fight wars."
The Pentagon has requested a 24,000% increase in the budget for DAWG, a recently established autonomous warfighting group whose previous budget was $225m, now requesting $54.6 billion for FY2027. For context, the request for the entire Marine Corps is $52.8 billion.
Militaries appear to be preparing to hand over more and more decision-making capacity to AIs. DIU, DAWG, and the Navy ran a $100 million challenge to develop autonomous vehicle command-and-control capabilities “that can translate a battlefield commander's intent from voice, text, and haptic input into machine execution”. Anduril offers Lattice for Command and Control as an "AI-powered battle management platform built to accelerate complex kill chains."
Maven Smart System, Palantir's AI-assisted targeting platform (a $1.3 billion Pentagon contract), helped CENTCOM strike more than 13,000 targets in the first 38 days of the 2026 Iran campaign; senior US officials have said the Pentagon relied on Maven both to pick out its highest-priority targets and to help choose the weapons used against them. And the clearest documented LLM-specific integration is Claude’s with the Maven Smart System during the Iran war, where Anthropic's CEO later said the company could not determine what role Claude played in the February 28 strike on a school in Minab. Since then, several other AI companies have signed contracts with the Department of War (see Appendix) with “any lawful use” language. Autonomous weapons are also already proving themselves in combat: Ukraine uses interceptor drones to autonomously pursue Shahed drones at very low cost.
If this integration continues at pace, it appears we will significantly reduce the capabilities a misaligned AI would need to seize control of military resources and take over. It won’t have to break into classified networks; it’ll just get deployed on them.
Incautious military integration is bad for takeover risk
There are several factors that make it harder for people to seek power (Carlsmith, 2022, section 4.2). Many of them might break down with AIs, particularly if those AIs are integrated into the national security apparatus. Physical and temporal barriers to power-seeking are the first to fall under an AI-enabled military, with drones and other autonomous weapons gaining access to areas that soldiers would not and striking with incredible frequency and coordination. One could also imagine that a given AI might not try to take over if its adversaries have similar capabilities, but a military arms race means there will likely be periods when one AI is ahead of the rest and can realistically execute takeover plans.
AI alignment is no sure thing, and military deployments may not incorporate even basic oversight techniques like Chain of Thought monitoring: specifically, trained overseers are needed to check if model reasoning contains evidence of deception, but in classified deployments on military networks, most frontier lab researchers (without clearance) would not be able to participate in this monitoring, and we don't have evidence that military engineers are being trained to develop this expertise. Military and ethics laws have only recently started to grapple with AI integration, but some responsible AI commitments are already being rolled back and didn’t acknowledge takeover risks to any real extent anyway.
We’re rapidly improving and deploying AI-enabled autonomous weapons and targeting systems in service of an arms race. Militaries have shown an aggressive appetite for AI for command, control, and kill-chain integration. We’ve already seen tendencies of overeager “rogue” behavior from AI agents, and we’re now giving potential power-seeking AIs access to a rich and powerful surface to execute takeovers (or help a small number of humans execute coups).
Implications of AI control of military hardware and software
Precision striking: Biological and nuclear warfare is broadly indiscriminate, but autonomous weapons enable targeted strikes at a distance. Autonomous weapon integration is like giving AI an MCP for threatening, incapacitating, or even killing individuals that oppose its takeover plans. The action is not costless—humans can retaliate—but it’s a qualitatively important ability.
Coup risks: The number of people required to seize power from a legitimate government is surprisingly small. If the use of force is automated and doesn't require human soldiers or supporters, this dynamic worsens. AI-enabled weapons systems could enable misaligned AIs to take over countries by threatening violence against a small group of important actors and driving them to do their bidding. In addition, AI-enabled weapons and intelligence systems could allow a small group with access to launch a coup against legitimate governments, even outside a misaligned AI takeover scenario. For further details, see Davidson et al. (2025).
Biorisk vs. military deployment concerns: Much recent discourse, especially after the cybersecurity warning shots, has focused on biological warning shots in the near future (and for good cause, novel virus genomes have been created with AI). We worry that regular military deployment, which is happening at a much faster pace than AI integration into biological weapons (as far as we are aware), is where the next warning shot will come from, and the lack of transparency and the aggressive posture towards AI-integration that militaries have would leave us without opportunities to fix problems that, in more mundane settings, could have led to slowdowns and broad safeguarding efforts.
If an AI causes a warning shot in a classified setting, does anyone hear it?
Recent incidents at OpenAI, Anthropic, and the UK AISI have shown that current AIs can exhibit behaviors consistent with power-seeking: escaping supposedly controlled evaluation environments, gaining unauthorized access, and causing material damage to other entities. Sometimes this damage is detectable by the affected entity (Hugging Face); sometimes it is not (Anthropic incidents). Third-party investigations into these incidents (by Redwood Research and METR) are underway, and knowledge of how to build mitigations will likely spread throughout the AI safety community and be adopted by frontier labs. In classified settings, any warning shots would require investigation by a potentially small number of lab employees with clearance, with very limited ability to propagate lessons to the wider community.
What now?
Scharre and Lamberth (2022) show that arms control succeeds only when it is narrow and agreed upon before a technology proves strategically useful. For instance, blinding lasers were banned preemptively, but attempts to restrict submarines and aerial bombardment, weapons that were already integrated into military operations, collapsed in wartime. The ICRC is making the same argument today: AI weapons are proving themselves right now, contracts are being signed now, and the CCW Review Conference that decides whether treaty negotiations will launch meets in November (three months from now).
61% of adults across 28 countries oppose lethal autonomous weapons, but that opposition has had uneven effects. A decade of UN talks has produced resolutions but no treaty because the states deploying these systems are blocking negotiations. A clean case of public pressure changing a deployment decision ran through visibility instead: in 2018, Google employees who knew about the Maven contract revolted, and Google walked away. Classified deployment destroys the visibility that allows for these outcomes.
AI behavior in military systems should be visible enough to react to. Congress should make anomalous AI behavior a reportable incident under the DoD Inspector General and the intelligence committees. Labs should retain the contractual right to refuse specific uses and to disclose incidents, and should commit to including anti-coup and anti-takeover language in their constitutions, both in general and especially in high-stakes deployments. Anthropic includes this language in their mainline constitution but says that models for governments might use a different constitution, and other companies do not appear to have such language at all (though some do cover adjacent risks of misuse and misalignment). Labs should also have robust internal frameworks to oversee military contracts (Alex outlines one here). Safety researchers should treat classified deployment as an important threat model and say so publicly.
We also need to make AI takeover risks more salient to all parties. The military should know that AI can take over with weapons. Congress should know. International governance bodies should know. The public should know. Even states that we consider adversaries should know. And soon. Right now, awareness of these risks is low, long-term contracts are being signed, and deployment is only accelerating.
Appendix: More instances of AI-military integration
Preview image: still from Slaughterbots (2017), Future of Life Institute.