Anyone can claim a big number. We claim a chain of custody.
The corpus behind the model fleet spans court and regulatory records at institutional scale — but scale is the least interesting thing about it. What makes training data trustworthy is provenance: where each record came from, how duplicates were killed, and what never got in. That is what this page documents.
The composite key exists because we learned the hard way that raw identifiers collide across source systems — around two percent of the time, which at corpus scale is millions of phantom duplicates. The deduplication that fixed it was verified by an exact gated row count before the production data was swapped, and the pre-deduplication file was archived rather than deleted. That event is in the public changelog.
Zero synthetic records in production training. If real data does not exist for a model, the model stays stub — visibly — until it does.
Data that cannot be traced to a real source record does not train production models. When a review found one family in violation, the family was quarantined and its metrics withdrawn.
Blog posts, news coverage, and analyst opinion about litigation are not outcomes. The training target is what courts did, not what anyone wrote about it.
Sources whose terms of use prohibit this application are excluded, full stop — a data moat built on violations is a liability, not an asset.
Four active common-law jurisdictions, with additional markets in data-ingestion stages. Exact corpus figures are disclosed in diligence with tier context; public surfaces describe scale qualitatively by policy — a number without its methodology attached is how stale statistics get loose.