A buyer’s guide to hallucinations in legal AI
Can you trust the answer? We look at the many elements of trust when using AI for legal work.
Generative AI is fast becoming part of everyday legal work. Yet lawyers still have concerns around accuracy.
News headlines increasingly feature high-profile firms that have relied on inaccurate AI-generated information, underlining how quickly errors can become public.
Buyers need to understand what information a legal AI tool is grounded in, how those answers can be verified and what happens when the technology gets something wrong.
Legal AI tools are now everywhere, but not all AI is equal, and not all AI deserves your trust.
Our latest survey found 83% of legal professionals are concerned about fabricated or inaccurate information, up from 57% when we began measuring it in early 2024. At the same time, 81% say they feel more comfortable using AI when it is grounded in legal sources.
In August 2026, the Solicitors Regulation Authority warned of AI-generated inaccuracies or “hallucinations” in legal research, advice, analysis and court submissions.
The Courts and Tribunals Judiciary has also warned that AI hallucinations can generate incorrect or misleading information, while emphasising the personal responsibility of judicial office holders for material produced in their name.
Here are some useful questions to ask vendors when searching for the right solution.
1. Start with the source material
The information available to an AI system is one of the most important elements of its reliability.
Buyers should establish what sits behind the system. Is it drawing on general web information, authoritative legal content, the organisation’s own documents, or a combination?
Then ask how those sources are selected, updated and maintained.
A fluent answer is not necessarily a reliable answer.
Ask the provider: What information can your AI draw upon when answering legal questions, and how is that information selected and maintained?
“Lawyers need to be able to verify sources and test outputs against underlying material. AI-generated work should not be treated as reliable simply because it looks polished.”
2. Ask whether it retrieves the right source
Grounding an AI system in legal information can help provide relevant context and supporting evidence. But buyers should look beyond whether a provider simply says its AI is “grounded”.
Retrieval-augmented generation, or RAG, typically involves retrieving information from a defined collection before an answer is produced. Its effectiveness therefore depends partly on the quality of that retrieval.
Buyers need to ask:
Does the system contain authoritative legal information?
And:
Can it identify the right authoritative information for the question being asked?
Because those are two very different things.
Ask the provider: How do you test whether your system retrieves the correct authority rather than simply a plausible one?
3. Test the citation, not just the answer
A citation can make an AI-generated answer appear more trustworthy. Buyers should test whether that confidence is justified.
Can the lawyer move easily from an AI-generated proposition to the underlying authority? Does the cited material genuinely support the proposition? Can the lawyer establish whether it is current and appropriate for the jurisdiction?
During procurement, don't simply ask whether citations are available. Follow them.
Check whether the authority exists, whether it supports the proposition made, whether the correct jurisdiction and date have been applied, and whether important qualifications are visible.
Ask the provider: Show me exactly how a lawyer can move from an AI-generated proposition to the legal material supporting it.
4. Find out what happens when the AI does not know
A useful procurement test is not simply how a product performs when there is an obvious answer.
It is how it behaves when there isn't one.
This is what is known as confabulation, where a system can confidently present erroneous or false content, including incorrect logic or citations. This is particularly relevant in areas requiring substantial contextual or domain expertise, such as the law.
Buyers should therefore test what happens when:
- a question contains an incorrect premise
- the relevant authorities conflict
- the answer is uncertain
- the law has changed
- the necessary information is not available to the system.
Does the AI communicate those limitations, or simply attempt an answer?
Ask the provider: What does the system do when it cannot find enough reliable information to answer?
New AI workflows. Built for legal standards.
5. Ask how reliability is tested
“Accurate” is not particularly useful as a procurement claim unless you understand how it has been measured.
Ask providers what they test, who assesses the answers, what constitutes an error and how closely their testing reflects the work your lawyers will actually perform.
Testing should also reflect the fact that legal AI systems are not static. Models, retrieval techniques, content and workflows can change.
Ask the provider:
What types of legal problems are included in your testing?
Are outputs assessed by lawyers or subject-matter experts?
How do you test source retrieval as well as the final generated answer?
What happens when weaknesses are discovered?
How frequently is testing repeated?
6. Make verification part of the workflow
The consequences of an inaccurate answer depend heavily on what happens after it is generated.
This is particularly important as AI moves into substantive legal work. Our survey found 69% of respondents use AI for legal research, 62% for document summarisation and 53% for both drafting and document review.
The SRA's warning similarly stresses appropriate human oversight and governance around AI use.
Buyers should therefore assess the product and the surrounding workflow together.
Does the system help lawyers identify sources? Can higher-risk outputs be reviewed appropriately? Is there a clear point at which professional judgement takes over?
Ask the provider: How does the product make verification easier for higher-risk legal work?
The buyer’s six-point reliability test
Before choosing a legal AI provider, assess six things:
- Source quality: What information can the system rely on and how is it maintained?
- Retrieval quality: How does it identify the right authority for the question?
- Answer quality: How are accuracy, completeness and appropriate use of authority tested?
- Uncertainty: How does the system behave when the evidence does not support a clear answer?
- Verification: How easily can lawyers check propositions against the underlying material?
- Accountability: Can you reconstruct what happened if an output is later challenged?
Trust should be earned, not assumed
To find the right solution, focus less on which AI produces the most convincing answers and more on which makes those answers easiest to verify.
Buyers should look across the entire reliability chain:
- the quality of the information behind the system
- how that information is retrieved, how answers are generated
- how uncertainty is handled
- how easily lawyers can verify the result.
"You need to properly understand the functionality, the data flows, what the tool can access, how outputs are generated, and how people are likely to use those outputs in practice."
2026 research & thought leadership reports
The Mentorship Gap
AI can help new lawyers to search, summarise, draft, and reason. But can it teach?
Integrating AI into legal workflows
Anything-goes AI use won't raise your standards. Legal teams need to restructure their workflows based on suitable AI.
Death of the rainmaker
Love them or loathe them, rainmakers are a law firm's bread and butter. But will AI-powered legal solutions see the practice take precedence over the partner?
In CTO we trust
Chief Technology Officers at leading law firms face a new challenge: not just deploying AI, but making it trustworthy at scale.