E-mail:
Jack Balkin: jackbalkin at yahoo.com
Bruce Ackerman bruce.ackerman at yale.edu
Ian Ayres ian.ayres at yale.edu
Corey Brettschneider corey_brettschneider at brown.edu
Mary Dudziak mary.l.dudziak at emory.edu
Joey Fishkin joey.fishkin at gmail.com
Heather Gerken heather.gerken at yale.edu
Abbe Gluck abbe.gluck at yale.edu
Mark Graber mgraber at law.umaryland.edu
Stephen Griffin sgriffin at tulane.edu
Jonathan Hafetz jonathan.hafetz at shu.edu
Jeremy Kessler jkessler at law.columbia.edu
Andrew Koppelman akoppelman at law.northwestern.edu
Marty Lederman msl46 at law.georgetown.edu
Sanford Levinson slevinson at law.utexas.edu
David Luban david.luban at gmail.com
Gerard Magliocca gmaglioc at iupui.edu
Jason Mazzone mazzonej at illinois.edu
Linda McClain lmcclain at bu.edu
John Mikhail mikhail at law.georgetown.edu
Frank Pasquale pasquale.frank at gmail.com
Nate Persily npersily at gmail.com
Michael Stokes Paulsen michaelstokespaulsen at gmail.com
Deborah Pearlstein dpearlst at yu.edu
Rick Pildes rick.pildes at nyu.edu
David Pozen dpozen at law.columbia.edu
Richard Primus raprimus at umich.edu
K. Sabeel Rahmansabeel.rahman at brooklaw.edu
Alice Ristroph alice.ristroph at shu.edu
Neil Siegel siegel at law.duke.edu
David Super david.super at law.georgetown.edu
Brian Tamanaha btamanaha at wulaw.wustl.edu
Nelson Tebbe nelson.tebbe at brooklaw.edu
Mark Tushnet mtushnet at law.harvard.edu
Adam Winkler winkler at ucla.edu
Public Institutions Can’t Outsource Their Reasoning
Guest Blogger
For the Balkinization Symposium on the Global Political Economy of Artificial Intelligence.
Ignacio Cofone
[This essay distills part
of the argument in Ignacio Cofone, Institutional Accountability and
Legitimate Inference in Algorithmic Adjudication: Beyond Trustworthy AI,
forthcoming in Cambridge Forum on AI Law and Governance (2026).]
AI does not relieve courts
and administrative agencies of the duty to defend the reasoning behind their
decisions but, on occasion, it does make that duty harder to satisfy.
For most of the 2010s, the
Dutch tax authority used a self-learning algorithm to flag potential fraud in
claims for childcare benefits, a means-tested subsidy that helps parents cover
daycare costs. The system assigned higher risk scores to families with certain
characteristics, including dual nationality. When civil servants reviewed
flagged claims, they were given no information about why the system had
assigned the score. More than 26,000 families were wrongly accused, many of
them ordered to repay tens of thousands of euros in full, often with penalties
and no installments. Some lost their homes, their jobs, or custody of their
children. In January 2021, the entire Dutch cabinet resigned over the scandal.
The Dutch Data Protection
Authority called the practice unlawful and discriminatory when it fined the
government under the GDPR. A parliamentary inquiry found a violation of the
rule of law. The Dutch high administrative court had reviewed individual cases
for years without catching any of this. All these failures came back to the
same institutional defect. A self-learning model was producing decisions affecting
thousands of families, and the institution running it could not, on demand,
reconstruct any of those decisions in terms the law could evaluate. The
reasoning path leading to action against a family was opaque to everyone in a
position to challenge it: the family, the civil servants reviewing the flag,
and any court asked to review what the institution had done.
In Toeslagenaffaire,
everything that is supposed to make a system like this accountable was in place.
There was an approved algorithm and a procurement process behind it. Civil
servants reviewed every flagged case. Procedures, escalation paths, and appeal
rights existed on paper. None of those gave the institution the ability to
answer for the outcome. The AI does not take that role on. Trustworthiness in
adjudication is the institution’s work.
Courts and agencies derive
their authority from the procedures they follow, such as rules about what
evidence may be considered, requirements to give reasons, opportunities to
contest, and standards of review. None of these properties belong to a model. The
institution still has to justify that output the way law requires, defend it
under cross-examination, and respond on appeal. When AI participates in a
decision, the institution needs to ask whether it can keep doing those things
well. When it cannot, accuracy alone will not save the decision.
The standard policy
response treats this as a technical problem with a technical fix: explainable
AI, meaning systems designed to give an account of how they reach their outputs.
The European AI Act requires explainability, and multiple US bills propose
disclosure of model logic. The premise is that if a system can describe what it
does, accountability follows. But it does not follow necessarily. Such a
description tells a court or an interested party how the system reaches its
outputs, but not whether a decision that incorporates those outputs rests on
grounds the law permits. That is a question courts and agencies have to answer.
Call the capacity to answer
it traceability. Traceability is the ability to reconstruct the reasoning path
from evidence to decision in terms that can be evaluated against legal
standards. Explainability tells a reviewer which features the model weighted.
Traceability requires that those features be legally permissible
considerations, that the weight assigned to them be defensible, and that the
affected party have had a real opportunity to challenge them. Many
explainability tools produce counterfactual statements that show what feature
mattered most by varying it (e.g., “if the defendant had lived in a different
neighborhood, the score would have been lower”). That statement describes the
model output but does not justify the decision; it does not, for example, tell stakeholders
whether neighborhood is a permissible ground for sentencing. The model alone
cannot answer the legal question. The institution has to.
Three things follow. First,
a model’s lack of transparency does not relieve an institution of its
traceability duty. Opacity makes traceability harder to satisfy, but the duty
runs to the reasoning path the institution constructs around the model and not
to the model’s internals: what the output represented, what weight the
decision-maker gave it, how it was integrated with other evidence and
applicable law, and how the affected party could contest each of those steps.
Second, traceability is
what existing doctrine already demands once AI is involved. Due process
requires that an affected party be able to identify and challenge the basis of
a decision against them. Arbitrary-and-capricious review under administrative
law requires that an agency consider the relevant factors and explain how they
connect to the choice it made. Equality doctrines require that decisions not
rest on impermissible grounds. None of these doctrines is satisfied by an
explanation of how a model works. All of them require that the institution show
why the resulting decision rests on grounds the law permits.
Third, this reframes what
deploying an AI system commits an institution to, whether the system is built
in-house or procured from a vendor. The duty to justify a decision is the
institution’s regardless of what tools it uses to aid in the decision.
Deploying AI does not move that duty to the model, the vendor, or the engineer.
The procurement contract or internal documentation must let the institution
obtain the information it needs to justify its reliance on the system and the
decisions that follow. The Dutch tax authority deployed a model whose internal
logic it could not interrogate, even at the level of its own civil servants.
The moment a penalized family asked why, the institution had no answer.
One reply to all of this is
that better accuracy and explainability will close the gap. They will not
because the gap is not technical. A perfectly accurate model still tells a
court only what the case is statistically, not what the law permits an
institution to do with that information. Counterfactual explanations and
feature attributions describe the model with more precision, but they still
cannot tell a court whether a decision the institution reached on the basis of
the model rests on grounds the law permits. Better model accuracy and technical
description, while desirable, do not answer questions of law.
When an institution cannot
account for its decision in terms the law can evaluate, it has tried to
outsource its authority to the AI system it relied on. That is not authority a
court or an agency has to give. In Toeslagenaffaire, when the institution was
finally asked to defend its decisions in legal terms, it could not. Whether the
model was accurate or sophisticated was beside the point. The same standard
applies wherever AI shapes decisions about rights, from risk scores in bail
hearings to generative outputs in administrative decisions.
Ignacio Cofone is Professor of Law & Regulation of AI, University of Oxford. You can reach him by e-mail at ignacio.cofone@law.ox.ac.uk.