Balkinization  

Saturday, August 29, 2026

Public Institutions Can’t Outsource Their Reasoning

Guest Blogger

For the Balkinization Symposium on the Global Political Economy of Artificial Intelligence.

Ignacio Cofone

[This essay distills part of the argument in Ignacio Cofone, Institutional Accountability and Legitimate Inference in Algorithmic Adjudication: Beyond Trustworthy AI, forthcoming in Cambridge Forum on AI Law and Governance (2026).]

AI does not relieve courts and administrative agencies of the duty to defend the reasoning behind their decisions but, on occasion, it does make that duty harder to satisfy.

For most of the 2010s, the Dutch tax authority used a self-learning algorithm to flag potential fraud in claims for childcare benefits, a means-tested subsidy that helps parents cover daycare costs. The system assigned higher risk scores to families with certain characteristics, including dual nationality. When civil servants reviewed flagged claims, they were given no information about why the system had assigned the score. More than 26,000 families were wrongly accused, many of them ordered to repay tens of thousands of euros in full, often with penalties and no installments. Some lost their homes, their jobs, or custody of their children. In January 2021, the entire Dutch cabinet resigned over the scandal.

The Dutch Data Protection Authority called the practice unlawful and discriminatory when it fined the government under the GDPR. A parliamentary inquiry found a violation of the rule of law. The Dutch high administrative court had reviewed individual cases for years without catching any of this. All these failures came back to the same institutional defect. A self-learning model was producing decisions affecting thousands of families, and the institution running it could not, on demand, reconstruct any of those decisions in terms the law could evaluate. The reasoning path leading to action against a family was opaque to everyone in a position to challenge it: the family, the civil servants reviewing the flag, and any court asked to review what the institution had done.

In Toeslagenaffaire, everything that is supposed to make a system like this accountable was in place. There was an approved algorithm and a procurement process behind it. Civil servants reviewed every flagged case. Procedures, escalation paths, and appeal rights existed on paper. None of those gave the institution the ability to answer for the outcome. The AI does not take that role on. Trustworthiness in adjudication is the institution’s work.

Courts and agencies derive their authority from the procedures they follow, such as rules about what evidence may be considered, requirements to give reasons, opportunities to contest, and standards of review. None of these properties belong to a model. The institution still has to justify that output the way law requires, defend it under cross-examination, and respond on appeal. When AI participates in a decision, the institution needs to ask whether it can keep doing those things well. When it cannot, accuracy alone will not save the decision.

The standard policy response treats this as a technical problem with a technical fix: explainable AI, meaning systems designed to give an account of how they reach their outputs. The European AI Act requires explainability, and multiple US bills propose disclosure of model logic. The premise is that if a system can describe what it does, accountability follows. But it does not follow necessarily. Such a description tells a court or an interested party how the system reaches its outputs, but not whether a decision that incorporates those outputs rests on grounds the law permits. That is a question courts and agencies have to answer.

Call the capacity to answer it traceability. Traceability is the ability to reconstruct the reasoning path from evidence to decision in terms that can be evaluated against legal standards. Explainability tells a reviewer which features the model weighted. Traceability requires that those features be legally permissible considerations, that the weight assigned to them be defensible, and that the affected party have had a real opportunity to challenge them. Many explainability tools produce counterfactual statements that show what feature mattered most by varying it (e.g., “if the defendant had lived in a different neighborhood, the score would have been lower”). That statement describes the model output but does not justify the decision; it does not, for example, tell stakeholders whether neighborhood is a permissible ground for sentencing. The model alone cannot answer the legal question. The institution has to.

Three things follow. First, a model’s lack of transparency does not relieve an institution of its traceability duty. Opacity makes traceability harder to satisfy, but the duty runs to the reasoning path the institution constructs around the model and not to the model’s internals: what the output represented, what weight the decision-maker gave it, how it was integrated with other evidence and applicable law, and how the affected party could contest each of those steps.

Second, traceability is what existing doctrine already demands once AI is involved. Due process requires that an affected party be able to identify and challenge the basis of a decision against them. Arbitrary-and-capricious review under administrative law requires that an agency consider the relevant factors and explain how they connect to the choice it made. Equality doctrines require that decisions not rest on impermissible grounds. None of these doctrines is satisfied by an explanation of how a model works. All of them require that the institution show why the resulting decision rests on grounds the law permits.

Third, this reframes what deploying an AI system commits an institution to, whether the system is built in-house or procured from a vendor. The duty to justify a decision is the institution’s regardless of what tools it uses to aid in the decision. Deploying AI does not move that duty to the model, the vendor, or the engineer. The procurement contract or internal documentation must let the institution obtain the information it needs to justify its reliance on the system and the decisions that follow. The Dutch tax authority deployed a model whose internal logic it could not interrogate, even at the level of its own civil servants. The moment a penalized family asked why, the institution had no answer.

One reply to all of this is that better accuracy and explainability will close the gap. They will not because the gap is not technical. A perfectly accurate model still tells a court only what the case is statistically, not what the law permits an institution to do with that information. Counterfactual explanations and feature attributions describe the model with more precision, but they still cannot tell a court whether a decision the institution reached on the basis of the model rests on grounds the law permits. Better model accuracy and technical description, while desirable, do not answer questions of law.

When an institution cannot account for its decision in terms the law can evaluate, it has tried to outsource its authority to the AI system it relied on. That is not authority a court or an agency has to give. In Toeslagenaffaire, when the institution was finally asked to defend its decisions in legal terms, it could not. Whether the model was accurate or sophisticated was beside the point. The same standard applies wherever AI shapes decisions about rights, from risk scores in bail hearings to generative outputs in administrative decisions.

Ignacio Cofone is Professor of Law & Regulation of AI, University of Oxford. You can reach him by e-mail at ignacio.cofone@law.ox.ac.uk. 



Older Posts

Home