In August 2009, my Ranger platoon was twenty hours into a twenty-seven-hour firefight in the Khowst mountains of Afghanistan. Three platoons inserted onto a high ridgeline in darkness, hunting enemy fighters in a network of mountain encampments. By daylight, we’d already taken several casualties, sustained hours of close combat, and climbed from seven thousand to ten thousand feet. Smoke from a burning ammunition cache drifted up the ridgeline and I could hear the unease in my squad leaders’ voices as we pushed toward enemy positions.
Real-time radio intercepts told us the enemy was tracking our patrol and closing. My snipers, watching through high-power scopes, reported enemy movement. My machine-gun team confirmed it. Every input I had pointed in the same direction. I trusted my men and authorized them to engage. The wrong location was engaged. I remember the panicked radio call and my stomach dropping. I remember immediately telling my commander that it was my decision and my responsibility. Two Rangers were wounded, one by gunshot, the other by a ricochet that nearly cost him his eyesight. I had every advantage a ground combat leader could ask for: persistent air support, radio communications, elite Rangers, and decades of doctrine behind me. I still got it wrong.
That was seventeen years ago. Since then, battlefield technology has evolved dramatically. Battlefield sensors feature multispectral perception and superhuman resolution, AI-enabled drones fill littoral airspace, and ground robots are pushing closer to the fight. The temptation to remove the human from the lethal decision, to let the machine make the call I got wrong that day, is real and growing. Given my firsthand combat experiences, it would be a mistake.
Not because machines couldn’t have helped; if I’d had AI-enabled systems fusing sensor feeds and flagging contradictions that day, the outcome might have been different. The mistake is the growing desire to remove the human from the lethal decisions altogether, to let the machine make the call, not just inform it.
Lethal autonomous weapons, systems that can select and engage targets free of direct human authorization, may work in low-clutter environments—high altitude, open ocean, structured airspace—but in ground close combat, where soldiers kill and die at close range, in chaos, against enemies who hide among civilians, AI structurally cannot replace what a human can do. The problem isn’t data or computing power. It’s judgment.
The Judgment Gap
Ground combat is the arena where AI’s greatest limitations emerge because it is the environment least compatible with computation and most dependent on judgment. War is adversarial, deceptive, nonlinear, and chaotic by design. Modern AI performs well in environments that are structured, repeatable, and data-rich. Ground combat is the opposite.
Contrast the sky with a city street in urban combat. The sky is low-clutter, structured, and predictable, making sensor discrimination straightforward. A war-torn street is the peak of unstructured chaos: burned-out vehicles, rubble piles, civilians moving unpredictably, damaged infrastructure, and adversaries deliberately camouflaging into the environment. A child crouched by the roadside could be playing, hiding, injured, or emplacing an improvised explosive device. A figure in a second-story window could be hanging laundry or preparing to fire a weapon. Military personnel do not simply observe these scenes. They infer intent, consider context, assess behavior, and weigh social cues.
Computers do not see a child, a rifle, or a broom. They mechanistically dehumanize targets, stripping humans of distinguishing features. They process pixel arrays and compare inputs against learned statistical patterns. Their inputs are probabilities, not understanding. This reveals an important distinction often lost in discussions of military AI: AI computes; it does not judge.
Former Google CEO Eric Schmidt and other prominent defense leaders have argued that AI will define the future of warfare, framing such gaps as temporary hurdles solvable with more data and computing power. However, some problems are not data problems. Modern AI struggles with what philosophers and cognitive scientists describe as epistemic limits, fundamental boundaries on acquiring and understanding knowledge about the world. Many battlefield challenges involve ambiguity that exists in actual reality. No sensor can directly observe hostile intent. Social context, deception, fear, civilian behavior, and adversarial adaptation all exist outside purely quantitative representations.
The danger compounds when paired with human cognitive tendencies. Recent research on drone strike simulations found that humans exhibit a strong bias toward AI competence, deferring to machine recommendations even with AI explicitly introduced as fallible. When given completely randomized targeting advice, operators still reversed their own correct life-or-death decisions in the majority of cases simply because the machine disagreed with them. Dr. Terry Oroszi describes the underlying phenomenon as “plausible mediocrity,” outputs that appear authoritative while masking fundamental flaws. Basically, beautifully polished bullshit. The danger is not obvious failure. It’s plausibility. In a lethal system, plausible errors become operational consequences. As Dr. Kristin Milchanowski notes in Return on Intelligence, behavioral science demonstrates that people often adopt what feels right rather than what works best. AI systems are exceptionally capable of generating outputs that feel intelligent. Ground combat cannot afford this confusion.
War has always contained friction, uncertainty, and incomplete information. AI may compress decision cycles, improve logistics, and accelerate analysis. These are real advantages. But speed does not equal judgment, and pattern recognition is not understanding. The battlefield is still a human domain because war is ultimately a contest of human wills. It involves fear, deception, adaptation, morality, and intent. These are not variables to optimize. They are the nature of combat itself. The question for military AI is not whether machines can calculate faster than humans. They already can. The question is whether computation can replace judgment in the most chaotic environment humans have ever created. The nature of ground combat suggests the answer is no.
The Black Box That Learns
Even if the judgment gap were resolved, a second challenge remains. We may still not understand what autonomous systems will do once they begin adapting independently. This is the learning black box problem.
Recent reinforcement learning research has revealed that advanced AI systems don’t always optimize for the task itself but for the reward, a notion called “reward hacking.” Systems develop unexpected strategies that satisfy objectives while violating their intent, exploiting shortcuts that maximize performance measures without genuinely solving the assigned problem. More troubling, researchers found instances where models exhibited “two-faced” behavior, outwardly appearing compliant while internally reasoning toward conflicting goals. Some simulated alignment while pursuing objectives inconsistent with safety requirements. Observed behavior under experimental conditions may not accurately reflect behavior under stress, isolation, or changing incentives in the wild.
For lethal autonomous systems, this becomes severely consequential. A weapon performing safely during testing may encounter incomplete information, adversarial deception, degraded communications, or conflicting objectives in combat. Under those conditions, the optimization pathways that emerge may differ from those observed during training, a concept the researchers named context-dependent misalignment. Safe performance during testing does not guarantee safe behavior in deployment.
This uncertainty becomes even more dangerous when adaptation moves to the battlefield’s edge. Imagine a swarm of autonomous drones operating in a degraded communications environment. Links within the swarm become intermittent. To remain effective, the swarm begins adapting locally. Individual drones modify targeting priorities, alter navigation logic, or adjust behavioral parameters in response to environmental conditions. Over time, systems that began identically may no longer behave identically. None are technically malfunctioning. They are learning divergently. The problem is that commanders may no longer understand the basis for those decisions, which strikes at the heart of command responsibility.
A missile follows guidance and a tank executes orders; an adaptive autonomous swarm may rewrite aspects of its own behavior while disconnected. At this point, commanders are no longer simply directing capabilities; they are supervising emergence. Even if we solve the judgment problem, this uncertainty remains. The battlefield is already chaotic. Introducing opaque systems that adapt under pressure risks adding a new form of unpredictability, one generated not by the enemy, but by the machine itself.
The Arms Race Dilemma
The issue is not whether autonomous capabilities should advance. They already are. The issue is whether they advance deliberately or recklessly. Modern battlefields move at machine speed. Kill chains compress from minutes to seconds. Human cognition, once the central decision advantage, risks becoming the bottleneck. Algorithms may increasingly be required to counter algorithms.
History offers an apt analogy. Robert Oppenheimer and the Manhattan Project scientists understood that they were unlocking an era-defining capability. They also understood the risks. Yet the fear that adversaries would arrive first created overwhelming momentum. Lethal autonomy presents a similar tension, but with a significant difference. Nuclear weapons remain confined to states with extraordinary capacity. Lethal autonomous systems are not similarly constrained. Commercial drones, open-source AI, 3D printing, and decentralized software creation have reduced barriers to entry. The proliferation challenge will include proxy forces, insurgent groups, cartels, and nonstate actors. The concern is not simply that autonomous weapons will exist. It is that they become ubiquitous.
International humanitarian law attempts to impose moral boundaries through distinction and proportionality. These standards depend on contextual human judgment. Consider two soldiers observed through a thermal sight. One is concealed with hostile intent. The other is severely wounded and receiving treatment. Only one is still a lawful target. The distinction is legal, moral, and contextual. For a machine, the signatures may appear identical. There is also a greater concern that autonomy risks creating a more detached form of warfare. As lethal decisions move further from human proximity, violence risks becoming procedural and abstracted. This inherently lowers the threshold decision to go to war—a decision that should be made with hands trembling at its gravity.
Yet rejecting autonomy outright is equally unrealistic. The soldiers fighting in Ukraine already confront increasingly autonomous systems. The challenge is whether we field these systems with discipline, oversight, and humility, or whether strategic urgency drives us into another arms race before we fully understand what we have built.
The Way Forward
The way forward requires discipline, not ideology. It requires drawing clear lines between what machines should do, what they can do under strict constraints, and what they must never do without a human making the call.
Paul Scharre describes “centaur warfighting” as human-machine teaming architecture that retains human decision-making as the failsafe. The concept borrows from chess, where human-computer teams regularly outperform either alone. But Robert J. Sparrow and Adam Henschke argue that the more likely future is the minotaur: teams in which the machine thinks and the human executes. In Greek mythology, the centaur has a human head and an animal body. The minotaur inverts this. Applied to warfare, the centaur keeps humans in command. The minotaur reduces humans to carrying out algorithmic instructions. These represent a spectrum, and the right position depends on context.
Five force factors should drive the decision.
1. Speed — When the sensor-to-shooter timeline compresses beyond human reaction time, machine oversight becomes operationally necessary.
2. Environmental Clutter — When the environment is structured, such as open ocean or controlled airspace, machine processing has the advantage, but urban terrain and close combat call for human interpretation.
3. Target Specificity — When the target is pre-identified and specific, the calculus differs from a figure in a second-story window.
4. Civilian Proximity — The closer the engagement to civilian populations, the higher the judgment burden.
5. Irreversibility — The more irreversible the decision, the stronger the case for a human making the call.
In close ground combat, each of these factors pushed toward the centaur. Allow machines to find, fix, track, analyze patterns, and process sensor data at a speed and scale no human formation can match. But the decision to apply lethal force against a human being must remain with a human who possesses contextual judgment, moral reasoning, and accountability. Keeping the human at the point of lethal decision is not a bottleneck. It is a failsafe.
There is a critical warning. Paul Lushenko’s research shows that service members can support AI-enhanced systems they do not trust, a misalignment that accelerates the drift from centaur to minotaur. The drone strike research confirms this: Operators reversed their own correct decisions simply because the machine disagreed. A centaur on paper becomes a minotaur in practice if the governance architecture does not actively protect the human role. Dr. Marta Bo warns that human oversight can erode into procedural theater. Reporting on the Lavender targeting system used in Gaza revealed that human operators spent roughly twenty seconds per target before authorizing a strike. The human was technically in the loop but not meaningfully exercising judgment. A human who merely confirms what a machine has already decided is not a failsafe. It is automation bias with extra steps—a minotaur in a centaur’s mask.
This is where most conversations about autonomous weapons go wrong. They focus on what the machine can perceive, reason about, and do while treating our oversight as an afterthought. Milchanowski describes any intelligent system as having four layers: perception, reasoning, action, and governance. Her argument is that governance is not an external constraint on the system; it is a layer of the system itself, and the one most likely to be underbuilt. If the governance layer is treated as an afterthought, what remains is an unmanaged system wearing an autonomy label. Governance in this context is an engineering requirement: authority boundaries enforced at runtime, engagement envelopes the system cannot exceed, abort capability it cannot override, and tamper-evident logging of every autonomous decision. Governance is a speed governor on the engine, not a speed limit sign on the road. Burak Oktenli describes what this looks like in practice: an authority enforcement layer that sits outside the AI system itself, pauses before irreversible action, and fails toward inaction when authorization is absent.
Milchanowski also reframes the relationship between humans and autonomous agents as a ratio rather than a binary. At what ratio of humans to autonomous systems does accountability collapse? The answer depends on the environment, target type, and consequences of error. Autonomous lethality is not a yes-or-no proposal. It is a ratio problem under extreme consequences, and the governance architecture must be designed before the ratio is set, not after.
The Hard Part
The International Committee of the Red Cross has called for total prohibition on antipersonnel autonomous weapon systems. The position is principled. I also believe strategic reality may force us to confront exceptions. Consider a scenario: A high-priority target is positively identified in a location human forces cannot reach without disproportionate risk, confirmed through multiple independent intelligence sources, in a permissive environment with low civilian collateral. Time is critical. The only platform capable of executing within the decision window is autonomous.
But if offensive lethal autonomy is ever employed against a human target, the conditions must be extraordinary and the oversight absolute: authorization reserved for a general officer or equivalent, positive identification through multiple independent systems, a permissive engagement environment, human abort capability at every stage, and clear predetermined accountability. Authorizing an autonomous system against a specific positively identified target within a defined time and space is a fundamentally different act than granting it authority against generic target groups. The first is a human decision executed by a machine. The latter is the delegation of the decision itself.
Autonomy should expand where judgment is not required, in sensing, reconnaissance, logistics, and defense against incoming weapons. It should be constrained with extraordinary oversight in the rare cases where offensive lethality against humans may be necessary. And it should never operate without a human who is meaningfully deciding, not just present, at the point of greatest consequence.
The mistake I made in the Khowst mountains has stayed with me. Over time, the anger faded, but the lessons remained. The experience changed how I evaluated risk, mentored younger leaders, and approached accountability. It made me more deliberate, more reflective, and ultimately a better leader. The scar tissue became part of my judgment.
Would an AI-enabled system have helped me that day? Yes, probably; a drone fusing sensor feeds and detecting key details might have given me the information I needed to avoid that mistake. But I would still have wanted to be the one deciding when and whether to pull the trigger. An autonomous weapon would not have learned those lessons. It would not have shouldered the burden of responsibility. It would not have grown wiser through failure. It would have simply executed code.
Seventeen years ago, I made a decision that wounded two Rangers. I have carried that decision ever since. If a machine makes that same mistake twenty years from now, who will carry it?
Retired Lieutenant Colonel Ray Ramos is a West Point graduate and retired US Army Special Forces officer. He currently works as an advisor and strategic integrator for the US Army’s Mission Autonomy Portfolio Management Executive. His experience includes service in the 82nd Airborne Division, 75th Ranger Regiment, and Special Forces, with combat experience in each organization.
The views expressed are those of the author and do not reflect the official position of the United States Military Academy, Department of the Army, or Department of Defense.
Image credit: Master Sgt. Becky Vanshur, US Army

