Criterica Intelligence — production models trained on real court records, not synthetic data
The Hard Questions

The questions sophisticated buyers actually ask — answered where everyone can read them.

These are the ten we hear from the sharpest rooms — investment committees, actuaries, heads of underwriting. If a claim on this site would need a caveat under cross-examination, the caveat belongs on the page. Here are ours.

They do — for the cases they have time to look at, with judgment that lives in their heads and leaves when they do. We are not claiming to out-judge your best people on their best day. We are systematic where judgment does not scale: every case scored, the same probability meaning the same thing every time, an audit trail behind every number, and no fatigue on the four-hundredth file. The honest framing is coverage, consistency, and auditability layered under your experts — not a replacement for them.

Because a language model outputs plausible text and an underwriting decision needs a calibrated frequency. Jurisdiction-level outcome data largely is not in LLM training corpora; LLM answers are stochastic and unauditable; and no LLM benchmark measures whether a stated litigation probability is true. We use discriminative statistical models trained and gated on real outcomes. The complete argument, including where LLMs genuinely help, is on the Models vs. Generative AI page.

Temporal holdout is the primary defense: every model is evaluated on real outcomes from a later period it never saw, which blocks the memorize-your-era failure that flatters randomly split evaluations. Promotion requires AUC ≥ 0.70 on that holdout, per model, with calibration reviewed. And the registry keeps the receipts — thousands of entries that failed the gate and were never promoted, plus production models later demoted when re-evaluation caught deterioration.

True, and it matters. The cases that reach any underwriting desk are a biased sample of all litigation. Our models are trained on the full population of filed outcomes, not on a funder's inbound flow, so the base rates are anchored to the court, not to the filter. What the model cannot see — the reasons a case was shopped — remains your underwriters' territory. That is one reason we position the system under human judgment rather than in place of it, and the audit on your own book is where selection effects get measured rather than argued about.

The holdout catches drift after it starts; it cannot predict a statute before it passes. Our defenses are structural: jurisdiction-specific models localize a shift to the venues it actually touches rather than contaminating a national model; scheduled retraining pulls new outcomes in on cadence; and re-evaluation can demote a model whose world changed. When a regime genuinely breaks a model, the correct output is a demotion notice, not a stale probability — and the changelog shows we act on that.

Definitions first: one model = one jurisdiction × one case type × one prediction target. Narrow by design, because narrow is what makes a probability meaningful in a specific courtroom. The count is of models that individually cleared promotion gates — and we publish the denominator: 37,351 registry entries, roughly one in five of which never earned production status. A count you can reconcile against a registry under NDA is the opposite of inflation.

Real filed court records — dockets that exist, in courts that exist, with outcomes that happened — deduplicated on composite source keys and carried with provenance through the pipeline. Zero synthetic data in production training. When a provenance review found one model family whose training data could not be verified, the family was quarantined and its metrics withdrawn; that event is in the public changelog because that is what data governance looks like when it is real.

Correct, and we never claim otherwise. The models tell you the outcome distribution for this configuration of venue, case type, posture, and assignment — they do not tell you a witness is weak or a document is fatal. The system prices the risk; your professionals litigate the case. Institutions do not ask their actuaries to argue motions either.

You should not — which is why we do not lead with accuracy claims. We publish the evaluation protocol and the gates, and then we offer the only evidence that actually binds: scoring your historical matters with known outcomes so you can measure lift against your own baseline. No demo-day cherry-picks, no blind-test theater. If the models do not add value on your book, that audit is where it shows, on your side of the table.

Withdraw them, publicly, and say why. This year that has included withdrawing a fleet-level calibration summary computed on an unrepresentative sample, demoting below-gate models, and quarantining a family with unverifiable provenance — all recorded in the changelog. A metrics program that has never retracted anything is not a metrics program. The bar we hold is simple: nothing on these pages should need a caveat we would only give when challenged.

Have an eleventh? Bring it. The verification session exists for exactly the questions this page cannot answer in general form.

Ask it directly →