Criterica Intelligence — production models trained on real court records, not synthetic data
The Registry

Every number here has a denominator.

Model counts are easy to inflate and hard to audit — so we publish ours the way a registry actually looks: production models next to the entries that never earned the label, and a changelog that records the corrections, not just the wins.

78.2%
OF REGISTRY ENTRIES HOLD PRODUCTION STATUS
29,211
PRODUCTION MODELS
Cleared AUC ≥ 0.70 temporal-holdout gates, individually
37,351
TOTAL REGISTRY ENTRIES
Every model ever attempted — including the ones that failed
16,302
VERIFIED JUDGES
Judicial assignment as a first-class model feature
50
US STATES COVERED
Plus all federal circuits, across 4 active jurisdictions

One model = one jurisdiction × one case type × one prediction target. The gap between37,351 registry entries and 29,211 production models is not overhead — it is the gate working: entries that remain experimental, stubbed awaiting sufficient data, or failed and recorded as failed. Counts are updated only through the same verification protocol that governs every published statistic.

Methodology changelog

Prediction infrastructure that never corrects itself is not being measured. These are the material governance events on the fleet — including the ones where we found our own numbers wanting and withdrew them.

JUL 2026
GOVERNANCE
Below-gate models demoted from production

A fleet-wide temporal-holdout review identified a small set of production models performing below chance on the most recent evaluation window. All were moved back to experimental status pending label and evaluation audit, and the production count was reduced accordingly.

JUL 2026
CORRECTION
Fleet-wide calibration metric withdrawn for re-derivation

An internal audit found a previously computed fleet-level calibration summary had been evaluated on an unrepresentative sample. The figure was withdrawn from every surface rather than corrected in place, and per-model calibration review continues under the standard gates.

JUL 2026
DATA INTEGRITY
Court-record corpus deduplication verified

The primary corpus was deduplicated on a composite source key after cross-source identifier collisions were found, and the unique count was verified by a gated row-count check before the production data was swapped. The pre-deduplication file was archived, not deleted.

JUL 2026
GOVERNANCE
Models with unverifiable training provenance quarantined

A provenance review found a small model family whose training data could not be traced to verifiable source records. The family was reverted to experimental status and its previously published metrics were withdrawn.

APR 2026
TRAINING
Universal retrain on real data completed

Every production model was retrained exclusively on real filed records, completing the removal of all synthetic training data from the production fleet.

Verify it yourself
Registry inspection

Under NDA: walk the registry itself — model entries, training windows, holdout results, version history, status transitions. The counts on this page reconcile to that file.

Methodology walk-through

A working session with the team on pipeline, gates, and evaluation protocol. Bring your hardest statistician.

Audit on your book

The strongest test: we score your historical matters with known outcomes and you measure the models against what actually happened. Your data, your baseline.

Arrange a verification session →