Posted by Patrick Herald, a senior content producer at Legal Futures Associate Access Legal [1]

Herald: Risk of mistaking form for substance
AI doesn’t hallucinate. It doesn’t know the difference between getting something right and getting something wrong. Do legal professionals fully appreciate that they won’t know the difference either without a clear view of the technology?
‘Hallucination’ implies a system that normally gets things right and occasionally doesn’t – a glitch, essentially. I’ll get into why I find that framing to be misleading.
But for now, it’s enough to say that I’ve been thinking about how the legal profession responds to regulatory pressure – and whether that response tends to produce change or just the appearance of it.
On 23 April 2026, His Honour Judge Grimshaw referred two solicitors to the Solicitors Regulation Authority (SRA) [2] after AI-generated case citations were filed in support of an application for permission to appeal in a small claims matter. On its face, the case was a modest one about the return of goods. But the judgment deserves attention.
The failure it describes was not a single error by one person. What happened: a paralegal used AI tools to generate legal research. A consultant solicitor signed a statement of truth without checking the documents it endorsed.
Then, a firm director filed an appeal bundle containing a skeleton argument she claimed not to have reviewed. The document bore her name not in a header, as she suggested, but at the end, in the manner of a lawyer signing off a pleading.
The miscited authorities – which were real cases, but with wrong citations and propositions they did not support – were eventually identified by opposing counsel. (Not the firm, a compliance system, or the regulator.)
Two days before the judgment, the firm’s director filed a supplemental witness statement, assuring the court that procedures had been tightened and lessons learned. Attached was a certificate from an online AI course, dated two days earlier.
Judge Grimshaw noted it. But the statement of truth at its foot – the signed declaration that a court filing is accurate – was the old, superseded version, which drew a raised eyebrow: it had the appearance of preparation rather than the substance of it.
My colleague Brian Rogers has written a thorough analysis [3] of the regulatory dimensions of this case and its relationship to the Mazur decision on supervision and delegation. I want to examine something adjacent: not why the regulatory framework failed, but why training and professional culture keep failing to prevent these situations arising in the first place.
Responses to the wrong problem
The case described above is not an outlier. Barrister Matthew Lee’s tracker [4] of confirmed or suspected AI hallucination cases in UK courts stands at 81 entries at the time of writing; Damien Charlotin’s worldwide database [5] is at 2,149.
Both figures are almost certainly vast undercounts. Cases appear in trackers only when they reach court, are identified and are recorded.
The SRA published its Risk Outlook warning [6] in November 2023, noting that AI “does not have a concept of truth” and should never be trusted to judge its own accuracy.
The cases have multiplied exponentially since: despite the guidance, the pattern accelerated anyway – and the cases that go unchallenged, settle before judgment, or never reach litigation at all likely remain invisible.
The profession’s instinctive response has been reasonably consistent: verify more carefully, implement an AI policy, train people to spot hallucinations before they reach a court file. These are not unreasonable instincts. But what if they’re responses to the wrong problem?
Gerben Wierda coined [7] the term “failed approximations” for what the industry calls hallucinations, and I find it more accurate.
Large language models (LLMs) don’t build an explicit world model or maintain any conscious sense of whether their answers are correct. They generate text based on learned patterns, optimising for fluency and coherence rather than truth.
In a sense, they’re always ‘hallucinating’ – it’s just that the outputs are often accurate. A citation that exists and one that doesn’t can emerge from the same process, presented with the same confident, well-formed prose. The user reading the output has no reliable signal telling them which is which.
Deeply rooted problem
Some firms may assume the solution is better prompting or newer models. But the difficulty is more deeply rooted.
In a September 2025 paper [8], researchers at OpenAI and Georgia Tech showed that a language model’s tendency to produce plausible falsehoods arises from the basic statistics of how it is trained – such errors emerge even when the training data is error-free.
Hallucinations persist, they argue, because training and evaluation reward confident guessing over admitting uncertainty: a model that produces failed approximations scores better on standard benchmarks than one that says ‘I don’t know’.
The authors propose reforming how models are evaluated, but that fix aims to make models more willing to flag uncertainty, not to eliminate error. A genuinely error-free language model remains a theoretical construct, and the one the authors offer barely resembles a working LLM.
The realistic best case is a model that is wrong less often and better at admitting when it might be.
Such a difference of degree would be a meaningful improvement, but still not something a firm could safely treat as reliable, and not something better prompting would fix. Nor does the problem reliably shrink as tools grow more sophisticated: OpenAI’s own reporting has shown some newer models hallucinating more often, not less, on factual-recall tasks.
There are plenty of AI tools that operate in ways that don’t depend on the model grasping legal reality – workflow automation, or contract review with defined parameters.
The problem is specific: it arises when someone deploys an LLM for legal research and document drafting, treats the output as a starting point requiring only light review, and discovers too late that light review was always insufficient.
What firms are often managing, without fully realising it, is not a research tool that occasionally makes mistakes. It is a pattern-matching system that has learned what legal text looks like and produces more of it, whether or not it corresponds to anything real.
Changing the conditions under which AI is used
Without a firm grasp of this, it becomes easy to treat legal AI misuse cases as outliers – the result of a rogue paralegal, an inattentive solicitor or a firm that simply didn’t know better. I think this could be the same instinct that pushes people toward the word ‘hallucination’: a preference for the comforting idea that the system is basically sound and the failures are exceptional.
The evidence suggests otherwise on both counts.
In another recent case [9], a solicitor admitted to a judge that his error arose from time pressure. The judge responded that such pressure was “never an excuse for filing inaccurate or misleading documents”. But he also gestured toward the conditions that caused the failure, noting that it was “in substance a failure of management” at the firm – not an individual lapse.
I spent over a decade teaching at universities in the US, and I encountered a fair deal of plagiarism. The students who plagiarised rarely misunderstood the rules. They usually understood them perfectly well and were under enough pressure – whether that be a deadline, grades or a fear of failure – that a shortcut felt worth the risk. They knew and hoped not to be caught. I always felt empathy for these students.
The parallel is uncomfortable but I think it’s accurate. Some firms deploying LLMs without adequate supervision genuinely don’t understand the tool. Others may understand it well enough and reach for it anyway – say, because the fee earner is under pressure, the deadline is tomorrow and the alternative could be admitting the work can’t be done in the time available.
Regulation catches the cases that reach courtrooms. But what about the cases where the output goes unchallenged, where the matter settles or where opposing counsel simply doesn’t notice? The tracker can only count what the courts can see.
I don’t think good AI training necessarily teaches fee-earners to catch ‘failed approximations’ line by line. That is neither economically viable nor technically possible. What it should do is change the conditions under which AI is used: the decisions about when to use it, what tasks it is and isn’t suited to, what supervision looks like in practice, and who is accountable when something goes wrong.
These are questions of professional judgement, not just technical literacy. And professional judgement is trainable.
The Mazur framework, as Brian Rogers has set out, places supervisory responsibility squarely on the authorised individual. That responsibility cannot be discharged solely by a policy document, a brief e-learning module, or a certificate obtained four days before a court hearing. None of these alone is sufficient.
It requires genuine understanding of the tool, the risk and the conditions that make shortcuts attractive. Firms that build that understanding can create the kind of culture in which cases like the ones described above are less likely to arise, and more likely to be caught when they do.
That’s harder to achieve than updating a policy, but it may be the only thing that will actually work.
Mistaking form for substance
I wrote recently [10] about a structurally similar problem in anti-money laundering compliance – firms filing suspicious activity reports into a system that returns no feedback, generating activity without generating learning. The same logic runs through both: a compliance culture that mistakes the form for the substance, the certificate for the competence, the policy document for the practice.
The legal profession has developed considerable sophistication in talking about AI risk, but responding to it has proved harder. Guidance has been issued, yet misuse cases have multiplied since the SRA’s Risk Outlook warning, prompting the SRA to issue a formal warning notice [11] in August.
So here are the questions I’d put to anyone in a legal practice who is reading this:
- Does your firm have an AI policy – and does it go beyond listing what tools are permitted to ask why someone under pressure might reach for one anyway?
- Do the people responsible for supervision understand what they are supervising when it comes to AI? Not just that generative AI produces errors, but why those errors can be indistinguishable from accurate outputs until it is too late?
- Is your training building the kind of professional judgement that changes behaviour under pressure, or is it only giving people enough to say they’ve been trained?
Will we see a shift in AI supervision, or will the legal profession’s response to AI risk ironically come to be seen as its own kind of failed approximation?