Criterica Intelligence — production models trained on real court records, not synthetic data
Training Data

Anyone can claim a big number. We claim a chain of custody.

The corpus behind the model fleet spans court and regulatory records at institutional scale — but scale is the least interesting thing about it. What makes training data trustworthy is provenance: where each record came from, how duplicates were killed, and what never got in. That is what this page documents.

The path of a record
FILED RECORD
A real docket, in a real court, with a real outcome
COMPOSITE KEY
Source system + source identifier — because raw IDs collide across sources
DEDUPLICATION
One record, once — verified by gated row-count checks
TRAINING CORPUS
Provenance carried through to every model that learns from it

The composite key exists because we learned the hard way that raw identifiers collide across source systems — around two percent of the time, which at corpus scale is millions of phantom duplicates. The deduplication that fixed it was verified by an exact gated row count before the production data was swapped, and the pre-deduplication file was archived rather than deleted. That event is in the public changelog.

What we deliberately exclude
Synthetic data

Zero synthetic records in production training. If real data does not exist for a model, the model stays stub — visibly — until it does.

Unverifiable provenance

Data that cannot be traced to a real source record does not train production models. When a review found one family in violation, the family was quarantined and its metrics withdrawn.

Commentary and secondary text

Blog posts, news coverage, and analyst opinion about litigation are not outcomes. The training target is what courts did, not what anyone wrote about it.

Sources with prohibitive terms

Sources whose terms of use prohibit this application are excluded, full stop — a data moat built on violations is a liability, not an asset.

Where the data reaches
United States

All 50 states and all federal circuits. Judicial assignment modeled against 16,302 verified judges.

Canada

British Columbia, Alberta, Ontario, Quebec — common-law provinces with feature parity to US courts.

United Kingdom

England & Wales and Scotland.

Australia

State and federal coverage across the mature common-law market.

Four active common-law jurisdictions, with additional markets in data-ingestion stages. Exact corpus figures are disclosed in diligence with tier context; public surfaces describe scale qualitatively by policy — a number without its methodology attached is how stale statistics get loose.