TECHNOLOGY · ETHICS · PEACE

AI and Warfare

Practical solutions for civilian protection, accountable human control, and de-escalation.

Artificial intelligence changes warfare by changing what institutions can see, how quickly they decide, and how easily decisions become action. Its danger extends beyond autonomous weapons: software can shape a human decision so thoroughly that formal human approval becomes an empty ritual. Its promise also extends beyond combat: carefully governed systems could help protect hospitals, coordinate relief, document harm, and support negotiations.

The governing principle should be that an increase in computational capability must be accompanied by stronger responsibility, demonstrable civilian protection, and effective limits on escalation. The deployment programme developed below is a policy proposal, with accountable owners, staged trials, and stopping rules. It is not a claim that these systems have already been implemented.

1 What AI changes in warfare

Military AI includes analytical tools, administrative and logistical systems, decision support, and autonomous functions. These categories should not be conflated. The ICRC defines autonomous weapon systems by their ability to select and apply force without human intervention; such systems need not use machine learning. Conversely, an AI tool may strongly influence the use of force without controlling a weapon. [1][2]

This distinction changes where governance must begin. Reviewing only the final act overlooks the upstream construction of a decision: which information enters the record, which uncertainties disappear, and which alternatives never reach the responsible official. A system that produces an apparently authoritative assessment can determine the practical choice before anyone presses an approval button.

A further problem is institutional dependence. Once staffing, deadlines, and performance expectations assume machine assistance, an operator may technically have the right to reject a recommendation while lacking a workable way to proceed without it. A credible safeguard therefore requires a usable fallback procedure and protection for personnel who question the system.

2 Why greater accuracy may still produce greater harm

Three separate questions matter: whether a system performs its technical task well, whether its use is lawful, and whether its deployment makes conflict less destructive. Success on one does not establish the others.

Consider an illustrative calculation, not a measurement of an actual conflict. A process handling 1,000 assessments with a 5 percent error rate generates 50 incorrect assessments on average. A new process handling 100,000 assessments at a 1 percent error rate generates 1,000. The error rate falls while the number of errors rises twentyfold. Incorrect assessments are not equivalent to casualties, but the arithmetic shows why throughput and downstream consequences must accompany accuracy metrics.

The political consequences can also differ from the technical results. If leaders perceive force as cheaper, faster, or less risky for their own personnel, they may become more willing to use it. This is a possible incentive effect, not an inevitable outcome. It means evaluation should examine whether AI expands the frequency and scale of harmful activity, rather than treating efficiency itself as public benefit.

Errors can also become correlated. Several tools may repeat the same unreliable report, creating an illusion of independent corroboration. Counting agreeing outputs is therefore weaker than tracing whether their underlying evidence is actually independent.

3 Human control must be a practical capability

The ICRC identifies automation bias, poor contextual data, adversarial manipulation, and accelerated errors among the risks of AI decision support. Its recommendations include testing, reliable information, training, and meaningful opportunities to challenge outputs. [2]

A workable implementation should make human control observable. Reviewers need access to the evidence behind an assessment, the age and limitations of that evidence, competing interpretations, sufficient deliberation time, and authority to stop the process. A plausible explanation generated by a language model should never substitute for an evidentiary record.

Oversight tests should deliberately include misleading recommendations in controlled exercises. The question is whether reviewers detect them, explain their disagreement, and successfully interrupt the workflow. An approval log proves that someone clicked; it does not prove independent judgment.

Responsibilities must remain identifiable across procurement, development, authorization, operation, and investigation. These are institutional assignments, not automatic findings of legal liability. The legal responsibility of any person or organization depends on applicable law and the facts. Nevertheless, no institution should accept a workflow in which every participant can claim that responsibility belonged somewhere else.

4 Law sets constraints that an optimization score cannot replace

International humanitarian law protects civilians and limits the conduct of hostilities. Article 51 of Additional Protocol I prohibits indiscriminate attacks and includes the prohibition on attacks expected to cause incidental civilian harm excessive in relation to the concrete and direct military advantage anticipated. Article 36 requires states parties, when studying, developing, acquiring, or adopting a new weapon, means, or method of warfare, to determine whether its employment would be prohibited in some or all circumstances by applicable international law. These provisions have specific treaty scopes; they should not be described as a universal AI licensing regime. [3][4]

A numerical confidence score cannot by itself establish civilian or combatant status, proportionality, or the legality of an operation. Nor can a mathematical objective settle what an institution ought to value. Encoding a mistaken assumption more precisely merely makes the mistake more systematic.

The ICRC advocates new binding rules prohibiting unpredictable autonomous weapons and autonomous weapons designed or used to target people, alongside restrictions on other autonomous weapons. These are policy recommendations and should not be presented as an already universal treaty prohibition. [1]

This chapter supports those proposed prohibitions. It also proposes excluding AI from authority to initiate nuclear use and preventing analytical systems from automatically triggering escalatory responses. The justification is that catastrophic and irreversible consequences require institutional arrangements that preserve accountable deliberation, independent verification, and opportunities for restraint.

5 Ethics must address the institution as well as the machine

An ethics of consequences asks whose suffering is reduced and whose risks are displaced. Saving the lives of one side’s personnel cannot, on its own, establish that a technology improves the overall human situation. Civilian injury, displacement, damaged public services, and enduring fear belong in the assessment.

An ethics of duties asks whether some acts must remain impermissible even when an aggregate calculation promises advantage. Respect for persons becomes fragile when individuals are represented only as predictive categories. People must not lose protection because a system treats association, proximity, or incomplete data as sufficient evidence.

An ethics of character asks what repeated reliance on automation does to judgment. Institutions should cultivate caution, accountability, and the ability to recognize uncertainty. A culture that rewards unquestioning speed can defeat carefully written safeguards.

These perspectives converge on a practical demand: review the organization’s incentives, authority structures, and treatment of affected people alongside the software. Technical reliability cannot compensate for an institution committed to unlawful or reckless objectives.

6 Deploy a civilian protection review service

A first proposed deployment is an independent service that helps authorized reviewers understand risks to civilian life and essential services. It should document uncertainty about civilian presence and examine how disruption to electricity, water, transport, and healthcare can produce interconnected harm.

Its purpose is to surface reasons for caution and further inquiry. It must not issue an automated certificate that an action is safe or lawful. Missing information must remain visible as missing information.

A civilian protection lead should own the service, supported by legal advisers, security specialists, and independent evaluators. Humanitarian organizations should control whether and how they participate; participation must not be assumed or made a condition of receiving assistance.

Humanitarian data requires particularly strict boundaries. Patient identities, shelter locations, and aid recipient records can become sources of danger if reused for surveillance or military purposes. The preferred design minimizes collection, separates access by role, limits retention, and permits withholding sensitive information. An omission from a database must never be interpreted as proof that civilians or protected services are absent.

7 Deploy humanitarian tools under humanitarian control

A second programme should begin with narrowly bounded assistance to relief providers: forecasting medicine shortages, identifying inconsistencies in inventories, translating verified public guidance, and comparing supply needs against available resources. These applications still require testing; calling a system humanitarian does not make it harmless.

For example, a supply forecasting tool could run alongside an established manual process and flag possible stock shortages for trained staff. It should expose source records and uncertainty, allow local correction, and retain a manual fallback. It should not autonomously determine eligibility for aid or deny treatment.

Evaluation should compare service outcomes against the existing process. Relevant measures include stockout duration, staff time, accessibility across languages, errors affecting underserved groups, and complaints resolved. Increased reporting may indicate improved access to complaint channels rather than worsening performance, so metrics need interpretation.

Any trial should stop when a serious privacy breach, persistent exclusion, or inability to correct decisions makes continued operation unsafe. Restart should require documented remediation and independent review.

8 Deploy escalation safeguards and credible evidence systems

A third programme should strengthen the institutions that prevent misunderstandings from becoming armed confrontation. Verified communication channels, staffed diplomatic contact points, and agreed incident procedures are more fundamental than an AI interface.

AI could assist with translation, document comparison, or organizing incident reports, provided officials check consequential outputs against the original material. It should not infer hostile intent as an established fact or send an escalatory message without accountable human authorization.

Governments should negotiate procedures for questioning suspicious media, checking contradictory reports, and communicating uncertainty during crises. No fixed delay is appropriate in every emergency; the objective is to prevent machine speed from eliminating necessary verification.

A related evidence programme should preserve original records, timestamps, provenance, and a documented chain of custody for civilian harm investigations. AI may help organize material, but investigators must assess authenticity and competing explanations. Tamper evidence can reveal that a record changed; it cannot establish that the original claim was true. Access controls must also protect witnesses and survivors.

9 Make procurement and international cooperation enforceable

Procurement should make oversight a contractual requirement. Buyers should obtain access sufficient for independent evaluation, documented operating limits, incident notification, controlled updates, and procedures for suspending unsafe use. A material change in model behaviour or intended application should trigger renewed assessment.

Testing must examine the complete workflow. A reliable component can become unsafe when combined with poor information, excessive workload, or incentives that discourage dissent. Evaluation teams should therefore include operational users, legal specialists, human factors researchers, and people qualified to assess civilian impacts.

International cooperation should pursue verifiable restraint rather than rely on declarations alone. States could exchange information about review procedures, establish incident reporting arrangements, and develop common evaluation requirements while negotiating binding limits. Verification will remain difficult where software is dual use and access is restricted. That difficulty supports layered measures; it does not justify pretending that a declaration is enforcement.

Civilian communities and less wealthy states need a role in setting these standards. Otherwise, the states and companies with the greatest technical capacity may define acceptable risk for the people most exposed to its consequences.

10 A phased implementation roadmap

The following schedule is an illustrative programme design. Each stage requires a decision based on evidence; a calendar deadline does not authorize progression.

StageAccountable ownerDeliverableCondition for progression
First 30 daysExecutive authority and independent oversight leadInventory of AI uses, assigned responsibilities, prohibited uses, and complaint channelsEvery in scope system has an owner and documented purpose
Days 31 to 90Evaluation team and legal advisersWorkflow assessment, privacy review, fallback procedures, and baseline measuresCritical deficiencies addressed before a trial
Months 4 to 6Humanitarian programme ownerSmall trial of supply forecasting or verified information assistanceManual fallback works and users can correct outputs
Months 7 to 12Independent evaluatorsAssessment of benefits, errors, exclusion, incidents, and operating costsBenefits supported by evidence without unresolved serious harms
ContinuingPublic authority and external oversight bodyControlled updates, incident investigations, periodic review, and public reporting where safeSuspension and remedy remain available in practice

The starting investment should cover staff, evaluation, data protection, training, and maintenance as well as software. A nominally inexpensive model can create a costly system if checking its outputs requires more work than the task it replaces.

Expansion should follow demonstrated benefit. A successful humanitarian logistics pilot does not validate military decision support, and a laboratory result does not establish reliability in a conflict environment. Each new purpose needs its own assessment.

11 Measure protection and preserve the possibility of peace

Performance reports should combine technical and institutional measures: evidence completeness, reviewer detection of planted errors in exercises, successful interruption of unsafe workflows, privacy incidents, time to remedy, and real service outcomes. Approval speed and output volume should not stand in for civilian protection.

Civilian harm trends require especially careful interpretation. Reporting access, the intensity of hostilities, population movement, and other changes can affect observed outcomes. A decline in recorded incidents is not, by itself, evidence that AI caused an improvement. Independent investigation and qualitative testimony remain necessary.

The deepest policy question is what institutions choose to automate. They can devote resources to accelerating coercion, or to improving the knowledge, restraint, public services, and accountability that make violence less likely and less destructive. Technology does not choose that purpose for them.

A defensible approach to AI and warfare therefore joins enforceable limits with practical investment in protection and diplomacy. Its test is whether human beings retain the knowledge, authority, and institutional support to refuse an unjustified action—and whether those harmed can obtain an explanation, an investigation, and an effective remedy.

Sources

[1] International Committee of the Red Cross. ICRC position on autonomous weapon systems. Institutional recommendations on prohibitions and restrictions.

[2] International Committee of the Red Cross. FAQ on artificial intelligence in the military domain. Definitions, decision support risks, and human centred safeguards.

[3] Protocol Additional I to the Geneva Conventions. Article 36 on new weapons. Treaty text and scope of weapons review obligations.

[4] Protocol Additional I to the Geneva Conventions. Article 51 on protection of the civilian population. Treaty rules on civilian protection and prohibited attacks.

The implementation programmes, hypothetical calculation, and evaluation framework are the chapter’s proposals, not reported operational results.