Answer · updated

Our AI research assistant cited a case that does not exist. Can that be fixed?

Substantially reduced, yes; eliminated, no. A preregistered 2025 study found commercial legal research tools hallucinating 17–33% of the time. Restrict the corpus to approved sources, resolve every citation against a live authority database and show failures, evaluate on real queries continuously, and design so verification takes one click.

This is not a defect in your deployment alone. It is a measured property of the current generation of these tools, and the fix is architectural rather than a matter of a better prompt.

The published position

A preregistered evaluation published in the Journal of Empirical Legal Studies in 2025 tested commercial retrieval-based legal research tools and found each of them hallucinating between 17% and 33% of the time. These were purpose-built products from established legal publishers, using retrieval over curated case law. They were marketed as avoiding this problem and they did not.

The practical conclusion is not that the technology is unusable. It is that a system which produces a citation you have to check is a different product from one whose citations you can rely on, and it must be designed, governed and sold internally as the former.

Why grounding alone is not enough

Retrieval-augmented generation, established by Lewis and colleagues at NeurIPS 2020, materially improves factual grounding: the model answers from passages retrieved from a real corpus. But retrieval and synthesis are separate steps, and the failure usually sits in the second. The system retrieves three real authorities and then writes a sentence that merges two of them into a proposition neither supports — the citation is real, the attachment is invented.

Long context does not solve it either. Research in the Transactions of the ACL in 2024 found models used information best at the beginning and end of their input, with performance degrading significantly when the relevant passage sat in the middle. Adding more documents can reduce accuracy.

What reduces the rate

Four measures, in order of effect. Restrict the corpus to sources you have approved, so a retrieved authority is always real. Require a citation for every proposition, resolved against a live authority database, and surface any citation that fails to resolve as a visible error rather than dropping it silently. Build an evaluation set from real queries, including ones with no good answer in the corpus, and re-run it on every change to the model, prompt or index. And design the interface so the citation is checked in one click, because the reviewer will check what is cheap to check.

What to tell the people using it

The honest framing is that this is a fast first-pass research tool whose output is verified before it leaves the building, with the verification step named and owned. That framing survives a professional negligence conversation. “The AI found it” does not.

This describes engineering practice. The professional conduct obligations that apply to you are a question for your own counsel.

Related questions

Do purpose-built legal AI tools avoid this?

The measured record says no. A preregistered evaluation in the Journal of Empirical Legal Studies in 2025 found commercial retrieval-based legal research tools each hallucinating between 17% and 33% of the time.

If the system retrieves real cases, how does it still get it wrong?

Retrieval and synthesis are separate steps and the failure is usually in the second: real authorities are merged into a proposition that neither supports. The citation exists; the attachment does not.

Would a larger context window fix it?

Not reliably. Transactions of the ACL research in 2024 found models using information best at the start and end of the input, with significant degradation when the relevant passage sat in the middle.

Related Qylis capability

AI Applications

Build generative and agentic AI applications that integrate with the systems you already run.

Discuss your requirement