Every number here has a denominator.
Model counts are easy to inflate and hard to audit — so we publish ours the way a registry actually looks: production models next to the entries that never earned the label, and a changelog that records the corrections, not just the wins.
One model = one jurisdiction × one case type × one prediction target. The gap between37,351 registry entries and 29,211 production models is not overhead — it is the gate working: entries that remain experimental, stubbed awaiting sufficient data, or failed and recorded as failed. Counts are updated only through the same verification protocol that governs every published statistic.
Prediction infrastructure that never corrects itself is not being measured. These are the material governance events on the fleet — including the ones where we found our own numbers wanting and withdrew them.
A fleet-wide temporal-holdout review identified a small set of production models performing below chance on the most recent evaluation window. All were moved back to experimental status pending label and evaluation audit, and the production count was reduced accordingly.
An internal audit found a previously computed fleet-level calibration summary had been evaluated on an unrepresentative sample. The figure was withdrawn from every surface rather than corrected in place, and per-model calibration review continues under the standard gates.
The primary corpus was deduplicated on a composite source key after cross-source identifier collisions were found, and the unique count was verified by a gated row-count check before the production data was swapped. The pre-deduplication file was archived, not deleted.
A provenance review found a small model family whose training data could not be traced to verifiable source records. The family was reverted to experimental status and its previously published metrics were withdrawn.
Every production model was retrained exclusively on real filed records, completing the removal of all synthetic training data from the production fleet.
Under NDA: walk the registry itself — model entries, training windows, holdout results, version history, status transitions. The counts on this page reconcile to that file.
A working session with the team on pipeline, gates, and evaluation protocol. Bring your hardest statistician.
The strongest test: we score your historical matters with known outcomes and you measure the models against what actually happened. Your data, your baseline.