Skip to main content
New: Project-aware research now ships with jurisdiction filtering. Learn more
AI in Legal Practice|August 16, 2026|11 min read

How Accurate Is AI Legal Research? What the Studies Actually Show (2026)

The honest, study-by-study answer to how accurate AI legal research really is — Stanford's two hallucination studies, the vendor disputes over the numbers, a newer benchmark that flips the story, and how to read any accuracy claim before you trust it.

AI in Legal PracticeLegal ResearchLegal TechAI Legal Research

How accurate is AI legal research? It depends entirely on how the tool works. General chatbots like ChatGPT invented or mis-stated the law on 58% to 88% of tested legal questions. Purpose-built research tools that retrieve real sources cut that sharply, but still erred on 17% to 34% of queries in the only independent audit to date. No tool is close to perfect, and every AI citation still needs verification before it reaches a court.

That range is the honest answer, and it is wider than either the doom takes or the vendor brochures admit. Below is the evidence, study by study, with the methodology explained plainly — because "accurate" turns out to mean three different things, and most accuracy claims quietly pick the flattering one.

The evidence at a glance

StudyWhat it testedHeadline findingThe caveat
Stanford, "Large Legal Fictions" (2024)General chatbots (GPT-4, GPT-3.5, Llama 2) on verifiable questions about real federal casesHallucinated on 58%–88% of queriesTests 2023 general models with no legal retrieval — a floor, not a verdict on purpose-built tools
Stanford RegLab, "Hallucination-Free?" (2024)Paid tools: Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AI, plus GPT-4Lexis+ >17%, Westlaw >34%, GPT-4 43% hallucinationTested mid-2024; vendors dispute the method and have since revised the products
Vals Legal AI Report (Oct 2025)Newer tools (Alexi, Counsel Stack, Midpage) + ChatGPT vs a lawyer baselineSpecialized tools ~80% accurate, beating lawyers' 71%The three largest platforms declined to participate; "general research" only

Study 1 — general chatbots: wrong most of the time

The starting point is the raw language model with no connection to a legal database. Stanford researchers measured this in Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models, published in the Journal of Legal Analysis. They posed specific, verifiable questions about randomly selected real federal cases and checked the answers.

The models hallucinated between 58% (GPT-4) and 88% (Llama 2) of the time, with GPT-3.5 at 69%. The pattern was revealing: they did best on famous Supreme Court cases and worst on obscure district-court details, because their training data is saturated with high-profile opinions and thin on everything else. Worse, they often failed to push back on a false premise in the question, and could not reliably tell when they were making things up.

This is the number behind the phrase "does AI make up cases." A chatbot does not look a citation up — it predicts what a citation should look like, which is why a fabricated case reads perfectly and cites to nothing. If your research tool is a general chatbot, this is your baseline, and it is a bad one.

Study 2 — the specialized tools: better, not solved

The obvious fix is retrieval: connect the model to a real database so it quotes documents instead of inventing them. That is the architecture behind Lexis+ AI and Westlaw's AI-Assisted Research, and it is where the marketing started promising "hallucination-free" answers.

The same Stanford lab tested that claim in Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, the first preregistered, independent audit of paid legal AI, later peer-reviewed in the Journal of Empirical Legal Studies. Running more than 200 open-ended legal queries through each product, the Stanford HAI summary reported:

  • Lexis+ AI: hallucinated more than 17% of the time (accurate on about 65% of queries)
  • Ask Practical Law AI: more than 17%, and gave incomplete or non-answers on more than 60% of queries
  • Westlaw AI-Assisted Research: more than 34% — roughly one answer in three (accurate on about 42%)
  • Raw GPT-4, for comparison: 43%

Retrieval clearly helps. The best tool cut the error rate to about a third of the general-chatbot baseline. But "more than 17%" is not "hallucination-free," and Westlaw's rate was worse than one in three. The AI legal research hallucination rate for even the leading paid products is a real double-digit number, not a rounding error.

The methodology is the part worth internalizing. Stanford counted an answer as a hallucination if it was either incorrect — it stated the law wrong — or misgrounded: it stated the law correctly but cited a source that did not support the claim. That second category is the quiet killer. A misgrounded citation is a real case attached to the wrong proposition, and it survives a quick click-through because the case exists. It is exactly the error a rushed associate waves past.

The vendors pushed back — and the fight is instructive

Thomson Reuters and LexisNexis both disputed the Stanford numbers, and the dispute itself teaches you how to read accuracy claims.

Thomson Reuters' head of Westlaw product management responded publicly that the company's own testing showed accuracy closer to 90%, and noted that Westlaw makes clear its product can produce inaccuracies. There was also a real access problem: the researchers were initially denied full access to Westlaw's tool, and only after the first release did Thomson Reuters grant it, at which point Stanford augmented the study and reported the 42%-accurate figure. LexisNexis argued its internal data suggested a much lower hallucination rate, though it did not publish that data, and it pointed to a newer RAG version released after testing.

A fair critic also noted that "Ask Practical Law AI" is a practice-guidance product that draws only from in-house practice notes, not a case-law research engine, so grouping it with the research tools arguably stacked the deck. That is a legitimate objection.

Here is the tell, though: LexisNexis's own rebuttal leaned on the study's finding that Lexis+ AI was accurate more than three times as often as Thomson Reuters' product. When a vendor cites the study to beat a competitor while disputing the study to defend itself, the relative rankings are probably sound even if the exact percentages move. The durable takeaway is not a specific number but a direction: retrieval-grounded tools beat chatbots, the leaders differ from each other by a lot, and none of them was hallucination-free.

Study 3 — a newer benchmark that flips the mood

Accuracy is a moving target, and a more recent test tells a sunnier story. The October 2025 Vals Legal AI Report, an independent benchmark graded blind by a consortium of law firms and academics, ran 210 general legal-research questions through several newer tools and a panel of practicing lawyers. The AI tools — Alexi, Counsel Stack, Midpage, and even generic ChatGPT — all scored around 80% accuracy, beating the 71% lawyer baseline, and outperformed the humans on 15 of 21 question types.

That sounds like the debate is over. It is not, and the reasons why are exactly the questions you should ask of any benchmark. The three largest platforms — Thomson Reuters, LexisNexis, and vLex — declined to participate, so the report does not measure the market leaders at all. It covered general legal research only, not the citation formatting where fabrications actually surface. And 80% accuracy still means one answer in five is wrong. Beating a distracted lawyer on untimed general questions is real progress; it is not "safe to file unread."

What "accuracy" actually means: three different tests

The reason accuracy numbers scatter so widely is that they measure different things. Pull any citation from an AI answer and it faces three separate tests, in ascending order of difficulty:

  1. Existence — does the case exist? Retrieval-based tools mostly pass this now. This is the failure that got lawyers sanctioned under Rule 11, and it is the easiest to catch.
  2. Support — does the case actually say what the answer claims? This is the misgrounding problem, and it is where good-looking tools still fail double digits of the time. A citation can be real and still wrong for your point.
  3. Treatment — is the case still good law, or has it been reversed, overruled, or superseded? No AI tool fully automates this. It is the deepest layer of "accuracy" and the one benchmarks rarely test.

A tool can ace existence, do passably on support, and say nothing useful about treatment — and still market itself as "accurate." When you read a percentage, ask which of these three it measured. Most measure the first, imply the second, and ignore the third.

How to read any AI accuracy claim

Before you trust a number in a demo or a landing page, ask three questions:

  • What benchmark? Which queries, how many, how hard, and who wrote them? A high score on easy questions is easy.
  • What denominator? Did misgrounded and incomplete answers count as errors, or only outright fabrications? Definitions move a "17%" to a "43%."
  • Who ran it? Vendor-run tests and independent, preregistered ones tell systematically different stories. When a benchmark excludes the biggest tools, it is not measuring the market.

Those are the same habits that separate a credible statistic from marketing generally, which is why our sourced roundup of law-firm AI statistics — including the growing count of court sanctions for fake citations — links every figure to its named study. A number without a named source and a defined denominator is an advertisement.

The practical takeaway: verify at the architecture, then verify again

Two conclusions survive all three studies. First, retrieval beats generation but never reaches zero — grounding a model in a real database is the difference between 58%-plus and 17%, but 17% is still a filing-ending error rate if you trust it blind. Second, the lawyer is the last line, every time. Courts have been consistent that the attorney, not the tool, owns every citation filed.

That is the standard CaseRead is built to, and it is why we lead with verified citations rather than a raw accuracy percentage. Answers cite sources the system actually retrieved, and when a citation can't be verified against real court and statute data, it is flagged rather than asserted — the existence test handled at the architecture level, so the human hours go to support and treatment. If you want the tool landscape compared on exactly this test, our guide to the best AI legal research tools ranks them on citation integrity first, features second.

Whatever tool drafts your research, close the loop by hand. Run any AI-touched citations through the free Hallucination Shield — it checks each citation in pasted text (up to 25 per run) for existence and support against real court data, no signup — and build a repeatable verification habit for the ones headed into a filing. Accuracy is not a number a vendor gives you. It is a step you take.

Frequently asked questions

How accurate is AI legal research? It depends on the tool. In Stanford's testing, general chatbots like GPT-4 hallucinated on 58% to 88% of legal questions, while purpose-built tools that retrieve real sources still erred on 17% to 34%. A newer vendor-run benchmark put specialized tools near 80% accurate on general research, but excluded the largest platforms. No tool is close to perfect, so every AI citation needs verification before filing.

Does AI make up cases? Yes. General-purpose models like ChatGPT, Claude, and Gemini generate citations from patterns rather than looking them up, so they routinely invent plausible-looking cases that do not exist. Stanford's "Large Legal Fictions" study found hallucination rates of 58% to 88% on verifiable legal questions. Retrieval-based tools that pull from a real database make this far rarer, but even they sometimes cite a real case that does not support the claim.

What did the Stanford legal AI study find? Stanford ran two studies. "Large Legal Fictions" (2024) found general chatbots hallucinated on 58% to 88% of legal queries. "Hallucination-Free?" tested paid tools and found Lexis+ AI wrong more than 17% of the time and Westlaw AI-Assisted Research more than 34%, with raw GPT-4 at 43%. It counted an answer as hallucinated if it was factually wrong or cited a source that did not support it.

What is a misgrounded citation? A misgrounded citation is one where the AI states the law correctly but cites a source that does not actually support the claim — a real case attached to the wrong proposition. Stanford's researchers counted misgrounding as a hallucination alongside outright fabrication, because a confident answer with an authoritative-looking but irrelevant citation is exactly the kind of error a busy lawyer waves through. It is harder to catch than an invented case.

How should I read a vendor's AI accuracy claim? Ask three questions: what benchmark was used, what the denominator is, and who ran it. A "90% accurate" figure means little without knowing which queries were tested, whether misgrounded citations counted as errors, and whether the vendor or an independent party graded the answers. Independent, preregistered tests tend to report far higher error rates than vendor marketing. Treat any accuracy number as a starting question, not an answer.

CaseRead

CaseRead Team

AI-powered legal research built for practicing attorneys.

Ready to try AI-powered legal research?

Free to start. No credit card required.

Start Free