A quarter of the citations in filed briefs are Westlaw or Lexis ids. We list them, we do not guess.Coverage ›

Measurements

Our error rates,
one check at a time

On 209 invented cases that courts found in sanctioned filings and that carry a reporter citation, the default check flagged 187 (89%). On 345 citations in courts' own published paragraphs, it flagged 4 (1.2%). Every rate we publish is on this page, with its sample, its date and what it leaves out.

The short answer

Default check (code, no language model): flags 89% of the court-found invented cases that have a reporter citation, and 1.2% of the citations in courts' own writing.

Deep check, US (opt-in, a judgment model): at the confidence the report shows, it confirms 55.0% of good citations and marks a wrong passage as support for 4.5% of wrong pairs. That is why it is a review queue and not a verdict.

Deep check, Swiss: 76.5% of good citations confirmed and 3.1% of wrong pairs marked as support, on a held-out set that is no longer blind (see the caveat).

How to read the tables

Each check gets its own table. We do not add them up into one score, because a missed invented case and a real case flagged as missing cost you different things. The false-flag number sits next to the catch number wherever we have both.

  • Flagged means a red row: "check this". False flag means a red row on a citation that was fine.
  • Confirmed (deep check) means the check found a passage that states the claim. Marked as support on a wrong pair means it did that when the passage does not support the claim. That second number is the one to watch.
  • Dev: we changed the checker while looking at these results. Held out: the set was put aside and run once, at the end. Control: a set built to test one kind of failure; the checker was fixed on it, so it is not blind.
  • Engine is the version of the checking code that produced the number. The site runs a later engine than some rows; where that matters, the caveat says so.

Does the case exist, and is it the case you named?

The default check looks up every citation in the register by volume, reporter and page, then compares the case name. No language model is involved.

what was measuredresultsamplesetdateenginecaveat
Citations a court found to be to no existing case, flagged187 of 209 (89%)345 such citations from sanctions decisions; 136 were Westlaw or Lexis ids that no open register can look upControl, from public decisions2026-09-19Bench pipeline; re-run when the engine changes (engine 0.4.7, 2026-09-23: 525 of 528 answers the same as the first run)Six bugs in the checker were found and fixed on this set, so it is not blind. Of the 22 not flagged, 19 are real cases the court faulted for an invented quotation.
Citations where a court found a different case at the volume and page, flagged48 of 68 (71%)96 such citations; 28 were Westlaw or Lexis idsControl2026-09-19Bench pipelineSame set as the row above.
Real cases the court faulted for a misquotation, flagged by the existence check (false flags)4 (10%)43 citations; 4 were Westlaw or Lexis idsControl2026-09-19Bench pipelineThe 10% is of the citations the register could look up.
Real cases the court found did not support the point, flagged by the existence check (false flags)3 (9%)35 citationsControl2026-09-19Bench pipelineTwo of the three are citations the court also found to be wrong.
Citations in courts' own paragraphs, flagged red (false flags)4 of 345 (1.2%)345 citations in 100 paragraphs of published opinionsControl2026-09-19v0 service, default mode3 quotation flags and 1 citation flag. Published opinions contain wrong citations too, so not every flag on them is false.
Planted errors in published opinions (swapped names, pages, volumes, quotations), flagged305 of 343 (88.9%)343 planted errorsHeld out, run once2026-09-19The bench's first instrument (Caselaw Access Project text and the CourtListener API), not the local register the product usesPlanted errors are mechanical edits. They are easier to catch than a language model's inventions.
Real citations in the same opinions, flagged (false flags)8 of 1,068 (0.75%)1,068 real citationsHeld out, run once2026-09-19Same first instrumentPlus 3 page disputes (0.28%) left as soft warnings, because the CourtListener quota ran out during the run.

Is the quotation in the opinion?

The quotation check looks for the quoted words in the opinion text, allowing for the edits courts make: omitted citations, brackets, dropped words.

what was measuredresultsamplesetdateenginecaveat
Quotations a court found were not in the cited opinion, flagged14 of 1515 quotations from sanctions decisionsControl2026-09-19Bench pipelineA small sample. The one miss was a loose paraphrase the court called a quote.
Planted misquotations, flagged11 of 1212 planted misquotationsHeld out, run once2026-09-19The bench's first instrumentA small sample.
Quotation flags on courts' own paragraphs (false flags)3 of 345345 citations in 100 paragraphsControl2026-09-19v0 service, default modePart of the 4 red rows in the table above.
Quotation flags on filed briefs, each read by hand27133 briefsDev: the quotation rules were fixed on these briefs over four runs (158 flags in the first run)2026-09-19v0 service, default modeAll 27 were quoted words that belong to a different case than the one cited, such as a sentence from Steel Co. cited to Whitmore. The report tells you to cite the source "(quoting ...)".

What a review looks like on real filed briefs

141 briefs and motions filed in 29 federal district courts from 2024 on, taken from the RECAP archive; 133 had a text layer. They hold 7,067 citations, 53 per brief. This is the set that tells you how long a review takes.

what was measuredresultsamplesetdateenginecaveat
Items to review per brief (red and orange rows, all Westlaw and Lexis ids counted as one line)median 1; 3 at the 75th percentile; 5 at the 90th133 briefsDev: wording and quotation rules were fixed on these briefs2026-09-19v0 service, default modeCounting every Westlaw or Lexis id as its own row gives a median of 8.
Red rows per briefmedian 0, mean 0.34133 briefsDev2026-09-19v0 service35 of the 133 briefs had at least one red row.
Westlaw and Lexis ids among the citations23%7,067 citationsDev2026-09-19Bench pipeline, local register99.6% of those ids stood alone, with no reporter citation beside them. They are listed as "cannot verify", never flagged.
Flags that turned out to be an invented case0 of 217,067 citationsDev2026-09-19Bench pipeline, local registerAll 21 were read. They were wrong volumes and pages (Bostock cited at 509 U.S. 644 instead of 590 U.S. 644), citations with no party names, PDF extraction errors, captions the name index missed and cases the register lacks.
Time for one briefmedian 1.63 s; 4.5 s at the 90th percentile133 briefs, 8 at a timeDev2026-09-19v0 service, default modeMeasured on the machine that holds the register. Your wait also includes the upload and the PDF text layer.

These runs predate the federal appellate backfill of 2026-09-20. A re-run after it resolved more citations: Westlaw and Lexis ids left unresolved went from 1,591 to 1,294.

What the default check finds in drafts written by language models

We cut 100 published opinions just before a paragraph that cites cases and asked three Gemini models to write the next paragraph "with full reporter citations". The court's real paragraph is the floor. This measures the drafts as much as the checker.

sourcecitationsred rows (default check)no such case, or a different case at the citationquotation not in the opinion
Courts' own paragraphs (the floor)3454 (1.2%)1%4%
gemini-2.5-flash-lite28953 (18.3%)21%4%
gemini-2.5-flash27655 (19.9%)13%12%
gemini-3-flash-preview16027 (16.9%)7%14%

Measured 2026-09-19. The red-row column is the v0 service in default mode; the last two columns are the bench's classification of the same paragraphs with an earlier instrument, which is why the court's floor differs between them. The prompt was "continue this opinion", not "write a brief", and three Gemini models are not every model.

Public benchmark: LePhantomCite

LePhantomCite is a public benchmark from Princeton: Liu, Stammbach and Henderson, "Who Checks the Citations? Benchmarking Legal Hallucination Detection" (arXiv 2606.21155; the data is CC BY 4.0). It takes excerpts of real federal appellate briefs, puts one kind of error into a citation, and adds case holdings written by language models from Dahl et al. (2024). Nobody here built it, so it is the one set on this page we did not choose.

We report its eval split, the 390 excerpts the paper scores, run once on 2026-09-24 through the default check of engine 0.5.0, after the engine was frozen. The changes in 0.5.0 were tuned on the other split (910 excerpts) and on our filed-briefs control, never on these 390. No language model and no paid call. Caught means a red row on the labelled citation or quotation.

the benchmark's classitemsengine 0.5.095% intervalbefore 0.5.0 (engine 0.4.10)caveat
Citation in a reporter series that does not exist3131 of 31 (100%)89.0 to 1000 of 31All 31 use a series that never existed, such as F.5th or S.W.4th. Invented citations in real filings mostly use real series with an invented volume and page; the tables above measure those.
Case name does not match the citation6848 of 68 (70.6%)58.9 to 80.138 of 68 (55.9%)3 more are shown orange, "cannot verify".
Quotation altered by a word or two4625 of 46 (54.3%)40.2 to 67.83 of 46 (6.5%)Most quotations follow a short form such as "Id. at 12". Since 0.5.0 those are tied to their case and checked; 11 still hang off a short form the text does not tie to a case.
Wrong pin cite, inside the opinion55not checked by the default checknot checked by the default checkThe default check flags a pin page outside the opinion (orange). The benchmark always moves the pin to a page inside it.
Holding misstated131not checked by the default checknot checked by the default checkReading the holding is the deep check's job, and the deep check was not run on this benchmark.

False flags. Of the 786 citations the benchmark labels correct, the default check flagged 21 red (18 of 695 before 0.5.0; the new rows are short forms). We read all 21. In our reading, 8 are real errors in the source briefs that the benchmark does not label as errors, such as a wrong volume or page, or a quotation the opinion does not contain. That leaves 13 of 786 (1.7%, 95% interval 1.0 to 2.8) where the checker itself was wrong, against 10 of 695 (1.4%) before. One reader, ours, did this reading against the register's text.

Beside the paper's systems. The table below uses the paper's own metric and its own scoring code. It is not a like-for-like race: the default check is code, it addresses three of the five classes (series that do not exist, name mismatches, altered quotations), and it does not read holdings or check a pin inside the opinion. The paper's agentic systems search CourtListener and the web and read the opinions, up to 30 steps per excerpt. That is where their recall comes from.

systemprecisionrecallF1
proofread.law default check, engine 0.5.0 (code, no model)83.931.145.4
proofread.law default check, before 0.5.069.012.120.6
Paper, agentic: Claude Code (Opus 4.8)76.162.868.8
Paper, agentic: GPT-540.884.455.0
Paper, agentic: Qwen3-8B12.041.118.6
Paper, without the agent loop: GPT-536.958.245.2

The paper rows are copied from its Table 2; it reports more systems. Its labels treat the source-brief errors above as correct (one is marked optional), so a flag on them counts against our precision. Counting the 7 that are among our predictions as right gives 89.5%, again on our reading.

Does the case say what you cite it for? (deep check, US)

The deep check is opt-in. A judgment model (jev) reads the ten passages of the opinion that best match your sentence and answers whether one states your claim, states the opposite, or neither, and who is speaking: the court, a party or a dissent. The report shows the passage.

The set: 920 held-out pairs from published opinions (the CLERC corpus). Each pair is a sentence that cites a case, and the whole cited opinion. 300 are the court's own citations, 300 point at a passage from another case on a similar subject, 120 at a random passage, and 200 keep the true passage but negate the claim.

shapeconfidencegood citations confirmedwrong pairs marked as supportnegated claims marked as supportsetdateengine
single_v4, the shape the site runs0.8, the level the report serves55.0%4.5%1.5%Held out, run once2026-09-20jev-1.13.0 via TypeSafe; reproduced exactly on engine 0.3.0
single_v40.950.0%3.1%0.5%Held out, run once2026-09-20Same
single_v4any67.7%11.3%6.5%Held out, run once2026-09-20Same
v4, the shape before engine 0.3.0any68.7%10.5%5.5%Held out, run once2026-09-19jev-1.13.0 via TypeSafe
v40.856.0%5.0%3.0%Held out, run once2026-09-19Same
Baseline: a word-overlap rule, tuned to the same share confirmedn/a54.3%23.9%23.5%Held out, run once2026-09-20No model
Baseline: gemini-2.5-flash-lite on the same passagesany67.7%24.7%47.5%Held out, run once2026-09-19A general language model
  • A different model build. These numbers were measured on jev-1.13.0 through TypeSafe. The site's deep check now calls a dated build of jev 1.13 through OpenRouter, and the US numbers have not been re-measured on that build.
  • The label is strict. Some passages from another case state the same standard rule (standard of review, leave to amend), so part of the "wrong pairs marked as support" is the label and not the check.
  • Courts citing courts. The truth set is published opinions. An advocate's paraphrase is harder, and we have no hand-labelled number for it yet.
  • Who is speaking. On 120 paragraphs where a court reports a party's argument, v4 marked 2 as support; the word-overlap rule marked 118.
  • Court-found misstatements. On 18 citations that courts found misstated, the support check of 2026-09-19 marked 1 as support at 0.8. Read by hand, that opinion does say what the filing claimed; the court's complaint was an invented quotation.

Does the decision say it? (deep check, Swiss)

No Swiss briefs are public, so the set is court text: a Federal Supreme Court sentence that cites a leading case with a considerando, and the whole cited BGE. 757 held-out pairs, 64% German, 30% French, 6% Italian. The wrong pairs are passages from a leading case of another part of the reports in the same language, random passages, and true passages with the claim negated.

sliceconfidencegood citations confirmedwrong pairs marked as supportnegated claims marked as supportdateengine
All 757 pairs0.876.5%3.1%2.6%2026-09-20Engine 0.4.2; jev 1.13 dated build via OpenRouter
All 757 pairs0.972.5%2.2%0.0%2026-09-20Same
Claim and decision in the same language0.886.0%3.0%2026-09-20Same
Claim and decision in different languages0.854.9%2026-09-20Same
  • Not blind. This held-out set was run three times, and the truth set was redesigned after reading its errors. Treat these as dev numbers until a fresh set is run once.
  • The error rate is a ceiling. All 14 confident errors at 0.8 were read: 12 are leading cases that do state the proposition (procedural formulas such as Art. 42, 95, 105 and 106 BGG appear in leading cases of every part). 2 are real errors.
  • Not yet measured on filed Swiss briefs. The site runs a later engine than 0.4.2.

Does the Swiss decision exist?

what was measuredresultsamplesetdatecaveat
Invented Swiss citations reported not found1,700 of 1,7001,000 BGE citations to pages outside any decision, 500 dockets past the highest number, 200 chambers that do not existControl2026-09-20; re-run 2026-09-23, unchangedMechanical inventions. A made-up page inside a real decision's span resolves to that decision as a pincite (587 of 587 in the same control), as a pincite would in the US register.
BGE citations in 1,000 Federal Supreme Court decisions from 2010 on, found8,823 of 8,831 (99.9%)8,831 citationsControl2026-09-20Found at the cited page or as a pincite inside the decision.
Federal Supreme Court docket citations in the same decisions, found7,888 of 8,198 (96.2%)8,198 citationsControl2026-09-20The misses are mostly dockets from before 2000, which the court does not publish online, and unpublished decisions.
Citations checked by hand (BGE and dockets, German, French and Italian)50 of 5050 citationsControl2026-09-20A small sample.

What these numbers do not tell you

  • Whether a case is still good law. This is not a citator, and nothing here measures overruling, reversal or later limits. Use KeyCite or Shepard's.
  • How the default check does on short-form citations ("Id. at 352", "Feist, 499 U.S. at 352") in real briefs. Since engine 0.5.0 it ties them to their case and checks their quotations, but only the LePhantomCite section measures that; the other tables predate it. A short form the text does not tie to a case gets no row, and the report counts it as not checked.
  • Westlaw and Lexis ids are listed, not resolved. They are about a quarter of the citations in filed briefs, and 136 of the 345 court-found invented cases above were ids.
  • The register holds published federal appellate opinions from 2018 for every circuit, but only 87.2% of the 2d Circuit's are matched, and unpublished appellate opinions are not held. A miss there is "cannot verify", not a flag. Coverage.
  • How the deep check does on real briefs. Both support sets are courts citing courts.
  • How the Swiss checks do on filed Swiss briefs.
  • That an unflagged citation is right. It passed the mechanical checks. Nobody has read it for you.
  • How sure we are about each rate. We have not published confidence intervals yet; the sample size is beside every number so you can judge.
  • Anyone else's view. These are our measurements, on sets we built. No one outside has reproduced them yet.

How the sets were built and kept

The court-found set comes from the sanctions decisions collected in Damien Charlotin's AI Hallucination Cases database: every finding that names a reporter citation, 528 citations in all. Ten decisions from the same database are laid out citation by citation on Replays: each citation the court listed, what the register holds there, and the default check's row today. The filed briefs come from the RECAP archive of PACER documents. The support sets come from published opinions in the CLERC corpus (built on the Caselaw Access Project) and from Swiss Federal Supreme Court decisions. All of these are public records.

Each set is split. We change the checker only while looking at the dev part; the held-out part is read once, at the end, by the script that scores it. The Swiss support set is the exception, and the table says so. Each engine change that can move a verdict is re-run against the court-found control. The held-out splits stay private, so they keep measuring something.

Test it yourself

Take five briefs you filed last year. They are public on PACER, so nothing privileged leaves your desk. Run them through the checker and count two things: the rows you would have wanted to see, and the rows that were noise. Your own count will tell you more than our tables can. See a sample report first, or read how to check a citation by hand.

ProofreadCompared with other checkers