LegalAI Space

AI hallucination cases: what UK firms should learn before it is their turn

The reported decisions are not a story about careless lawyers. They share a structure, and the structure is what a firm can actually design against.

Published
Reading time
7 minutes
Written by
The LegalAI Space team, Cognesio LLP

Five UK decisions are now cited whenever this subject comes up: Harber v HMRC [2023] UKFTT 1007 (TC), Zzaman v HMRC [2025] UKFTT 00539 (TC), Bandla v SRA [2025] EWHC 1167 (Admin), and R (Ayinde) v London Borough of Haringey with Al-Haroun v Qatar National Bank QPSC [2025] EWHC 1383 (Admin), heard together by the Divisional Court with Dame Victoria Sharp P presiding. Read as a set rather than one at a time, they share a structure, and the structure is the thing a firm can design against.

The same shape, four times over

In each, material put before a tribunal or court referred to authorities that did not exist. In each, the person putting it forward appears to have believed the authorities were real. And in each, the discovery happened downstream, in the hands of the tribunal, the court or the opponent, rather than inside the organisation that produced the document.

That last feature is the one worth sitting with. None of these were caught by an internal check. The failure was not that a mistake was made, since mistakes are made constantly in legal work and most are absorbed by the process. The failure was that nothing between the drafting and the filing was designed to catch this particular kind of mistake.

Why ordinary quality control misses it

A supervising solicitor reading a draft is doing several things at once: testing the argument, checking the facts against the file, looking at the tone, thinking about what the other side will say. Existence of the cited authority is not on that list, because for thirty years it did not need to be. Citations came from a search, and a search returns things that exist.

So the check that is missing is not a harder version of a check firms already do. It is a new one, and new checks do not appear because people are told to be careful. They appear because somebody makes them a step, gives them an owner, and leaves a record when they run.

The tracker, and what it is good for

The researcher Damien Charlotin maintains a running public tracker of AI hallucination cases worldwide. It is worth ten minutes of any risk partner's time, not for the total, which will be out of date by the time you read this page, but for the range: different jurisdictions, different courts, different levels of seniority, litigants in person and represented parties alike.

The pattern that emerges is that this is not a junior-lawyer problem or a small-firm problem. It is a problem of any workflow where a text generator sits upstream of a document and no verification step sits between them. Firms that read the tracker tend to stop asking whether their people would do this and start asking where in their process it would be caught.

What a check would have had to do

Three things, in order. Fetch the authority from a source the reader can also open. Find the cited passage in it. Report clearly when either step fails, rather than reporting an inability to check as a pass. Every one of the reported failures would have surfaced at step one or step two, because a case that does not exist cannot be fetched.

That is what the verdicts do here. Every authority in an answer carries verified, needs a check, or not found. Verified means the passage was found, word for word, on one of the 74 approved public sources. Not found is the verdict a fabricated case earns, and it is not a soft warning: it is a statement that the checker looked in the places it is allowed to look and the authority was not in any of them.

Why the gates are code rather than a model

Four deterministic gates run on output: match-set, quote-verbatim, confidence-floor and `closed-world-url`. They are ordinary code comparing strings, hosts and sets. None of them asks a model whether the answer was any good, because a model grading its own work reproduces its own blind spots, confidently.

Match-set is the one that maps most directly onto these cases. It takes every authority appearing in the prose and asks whether it is in the set the run actually retrieved and checked. A citation that appears in the text but never appeared in a search result is exactly the footprint a fabricated case leaves, and it is a comparison a computer does perfectly and a tired human does badly at eight in the evening.

What to put in the supervision file

Whatever tool you use, the record needs four elements: what was checked, against which sources, what the result was, and who read the result. Three of those can be produced by software. The fourth cannot, and it is the one a regulator will ask about first.

The judiciary's own guidance is instructive here. The Artificial Intelligence Guidance for Judicial Office Holders, updated 31 October 2025, reaffirms that judges may ask whether AI was used, and warns about hidden white text prompts placed inside documents. The Bar Council updated its considerations on generative AI in November 2025. A firm that can answer the question when it is asked in a hearing is in a materially different position from one that has to go and find out.

The honest edge

Verification catches authorities that do not exist and quotations that do not match. It does not catch an argument that is legally wrong while resting on real, correctly quoted cases, and that remains the more common way for a memo to be dangerous. The gates raise the floor. They do not touch the ceiling.

The verbatim gate also works on quotations, so a passage that paraphrases a holding gives it less to compare, and the risk moves back to the reader. And no product on the market can promise that a model will not write a plausible citation into prose. What it can do is decide what verdict that citation carries, and whether the work can leave the firm while it is unconfirmed.

See it run on your own matter.

Free plan, two seats, 500 welcome credits, no card.