This post forms part of the Opinio Juris symposium on International Law and Artificial Intelligence in Armed Conflict (introduced here) and draws on the authors’ chapter in the forthcoming OUP volume of the same title.
Legal reviews of new weapons, means and methods of warfare feature prominently in the international discussion on the governance of artificial intelligence (AI) in the military domain. Participants in the Group of Governmental Experts on Lethal Autonomous Weapons Systems (GGE on LAWS), the Summits on Responsible AI in the Military Domain (REAIM), and the working groups under the US-led Political Declaration on responsible military use of AI and autonomy have consistently identified legal reviews as an important element of proposed governance frameworks. Stakeholders see legal reviews as a safeguard against the development and use of capabilities — not just autonomous weapon systems but a broader range of applications — that cannot be employed in compliance with international humanitarian law (IHL). We examined this mechanism in relation to AI capabilities at the 2025 International Conference on Cyber Conflict (CyCon) (see earlier post) and have since developed that analysis into a book chapter, forthcoming in the volume International Law and Artificial Intelligence in Armed Conflict: The AI–Cyber Interplay (see the introduction to this symposium).
This post explains our central finding: that legal reviews of military AI capabilities are essential, but that on their own they are insufficient. Legal reviews are essential for two reasons: they are required by international law, and, in the absence of an international verification regime, they operate as a confidence-building measure. They are insufficient for two further reasons: the distinctive features of military AI limit the effectiveness of legal review as a means of assuring compliance, and, more fundamentally, legal reviews are a procedural mechanism rather than a normative one.
Essential as a Legal Requirement
The clearest legal basis for the obligation to conduct legal reviews is Article 36 of Additional Protocol I to the Geneva Conventions (API). It requires each High Contracting Party, in the study, development, acquisition or adoption of a new weapon, means or method of warfare, to determine whether its employment would in some or all circumstances be prohibited by international law. A capability that cannot be used in compliance with IHL cannot lawfully be fielded. The review is the means by which a state addresses that issue before the capability is used.
Article 36 formally binds only the parties to API. Several militarily active states are not among them. Some of those states – including the United States and Israel – conduct legal reviews under national policy without conceding a treaty or customary obligation to do so. The obligation to review does not, in any event, rest on Article 36 alone. As we argue in the chapter, it is also supported by the positive due diligence dimension of Common Article 1 to the Geneva Conventions, under which states undertake to respect and to ensure respect for IHL. A state that fields a military AI capability without having assessed whether it can be used in compliance with IHL has not taken a reasonable step to prevent violations. Importantly, this basis for legal reviews does not depend on the state being bound by API or the contested customary status of Article 36. Relevantly for military AI capabilities, because the language of Common Article 1 is not confined (as Article 36 of API is) to ‘a new weapon, means or method of warfare’ and to the pre-employment phase, it offers a more natural foundation for the review of capabilities that do not fit neatly within such categories, as well as capabilities whose functioning or expected effects may change after they are fielded.
Essential as a Confidence-Building Measure
Legal reviews are essential for a second reason. In the absence of an internationally agreed verification regime for military AI comparable to those that exist for nuclear and chemical weapons, legal reviews conducted at the national level are the primary means by which states can not only ensure but demonstrate that their military AI capabilities can be used in compliance with applicable international law.
That confidence-building function depends on a degree of transparency. It does not necessarily require a state to disclose the outcome of any particular review, still less the technical detail on which the review rests. Indeed, states have legitimate reasons to protect both. But some transparency as to process allows legal reviews to operate as a confidence-building measure at a time when few others are available. This might include sharing information as to how, at what points in a capability’s lifecycle, and against which standards the reviews are conducted. The recent multilateral policy discussions appear to reflect this understanding, identifying legal reviews as central to responsible behaviour in terms that extend beyond the formal scope of Article 36.
The Features of Military AI that Limit the Effectiveness of a Legal Review
Legal reviews are, however, also insufficient in and of themselves. One reason for this arises from the unique features of military AI. These features prevent the legal review methodology developed to assess traditional weapons, and the software-based capabilities that preceded military AI, from transferring straightforwardly to AI-enabled capabilities. Four features of military AI, discussed at greater length in the chapter, contribute to this mismatch.
First, breadth. AI is not a category of weapon but a set of techniques applied across a wide range of military functions, including targeting, command and control, and intelligence, surveillance and reconnaissance (ISR). The range of legal rules potentially engaged is correspondingly broad, and extends beyond the IHL rules on weapons and the conduct of hostilities that have been central to Article 36 practice to date. Depending on how a capability is used, a review may need to consider international human rights law (for example, where the capability is used below the threshold of armed conflict) or the law on the use of force (where it informs decisions to use force in self-defence). The lawfulness of many military AI capabilities also depends less on their inherent characteristics than on the circumstances of their use. Accordingly, a robust review must identify the conditions under which a capability can be used lawfully, rather than reaching a single permitted-or-prohibited conclusion based on design characteristics.
Second, iterative development. A legal review is ordinarily conducted at a defined point. This relies on the assumption that an assessment made before a capability is fielded provides assurance for the period in which it is used. Contemporary development of an AI capability, by contrast, involves iteration, with training and learning happening at different points in the development and use of a military AI capability. Moreover, military AI capabilities are increasingly developed and delivered as services, and are updated, retrained or otherwise modified over their operational periods. This raises a question: when does a modification render a capability sufficiently ‘new’ to require a further review.
Third, reliability. A meaningful legal review requires the reviewer to foresee the effects of a capability in its normal or expected use. Legal reviews of traditional weapons draw on testing against defined performance parameters, such as a weapon’s accuracy or error rate. The established methods of testing, evaluation, verification and validation (TEVV) do not transfer straightforwardly to many AI capabilities. Such capabilities can be opaque (their internal reasoning resistant to inspection), brittle (liable to degrade when operating conditions differ from those represented in the training data), and vulnerable to adversarial tampering (for example, where a model trained on a manipulated dataset misclassifies targets in ways that testing under normal conditions would not reveal). Each of these characteristics makes the effects of a capability more difficult to foresee.
Fourth, the role of industry. Much of the information and expertise required for a meaningful review – including in relation to TEVV — is held by the private actors that develop these capabilities rather than by the reviewing state. The obligation to conduct the review remains with the state, but the means to conduct it effectively may not. Proprietary and trade-secret concerns, potential liability, classification, and a shortage of relevant technical expertise among reviewers may each limit a state’s access to the information it requires.
The chapter sets out ways in which legal reviews might be adapted to address these features. Adaptation can improve their effectiveness, but it does not overcome a more fundamental limitation.
Legal Review as a Procedural Mechanism
However carefully it is designed and conducted, a legal review remains a procedural mechanism. It operates within the existing law and reflects it. Legal reviews facilitate compliance with existing obligations; they do not develop new ones. Nor do they, as a national mechanism, promote common understanding about the interpretation of existing international law.
Many of the questions at the centre of the policy debate on military AI are, however, of a normative kind: what standards should govern these capabilities, how much and what form of human involvement the law should require and what applications and uses of military AI should be off-limits. These international normative questions cannot be resolved by a national review process alone.
Concluding Thoughts
Legal reviews of military AI capabilities are therefore essential but insufficient. They are essential because they are required by international law and because, in the absence of an international verification regime, they are the primary means by which states can demonstrate that their capabilities are able to be used lawfully. They are insufficient because the features of military AI limit their effectiveness, and because, however well conducted, they cannot substitute for the elaboration of the standards that should govern these capabilities.
Our chapter concludes with the connection between the essential and insufficient nature of legal reviews: as practices that generate practical experience and technical understanding at the intersection of military AI and international law, legal reviews are indispensable for any informed policy conversation. Whatever normative choices states ultimately make collectively, legal reviews at the national level will be necessary to give those choices practical effect and to generate the understanding of how existing rules apply to specific technologies.
