A checker that hides its blind spot is the thing it is supposed to catch. Every number below was measured on filed briefs and on machine-drafted paragraphs, with the same instrument the checker runs.
of invented cases caught (court-adjudicated set, 528 citations)
invented quotations caught
false flags on the court's own writing
hallucination rate. The default check uses no language model.
| source | what it holds | as of |
|---|---|---|
| CourtListener bulk data | 10.1 million cases, 18.1 million citations, one opinion text per case | dump 2026-06-30, refreshed daily from the search API |
| Caselaw Access Project | opinion text where CourtListener has none (to 2019) | static |
| govinfo | federal appellate opinions, all 13 circuits, 2018 onward | 2026-09 |
| Supreme Court Database | U.S. Reports citations for recent terms the bulk data lacks | 2025 release |
528 citations from court decisions that found fabricated or misused authorities; 7,067 citations in 141 briefs filed in 29 federal district courts in 2024 and later; 300 paragraphs drafted by three language models against the court's own 100 paragraphs as the floor. Real filed briefs carried zero fabricated cases; the LLM drafts carried 7 to 21%.
The deep check (opt-in) confirms 69% of good citations and certifies a wrong passage 10.5% of the time, 5% at the confidence it shows. That is why it is a review queue with the passage, never a verdict.