How to Read a Calibrated Outcome Distribution
A single number — "68% probability of a favorable outcome" — is the least informative thing a calibrated model produces. The distribution behind it is the underwriting object. Here is what to look at instead.
What "calibrated" actually means
A calibrated model is one where, among every case it assigns a given probability, that fraction of cases actually resolves that way. If a model says 0.70 across a thousand cases with that score, roughly 700 should resolve favorably, checked against real recorded outcomes the model never saw during training. That is a testable property, not a marketing claim. It is different from accuracy, and different from confidence. A model can be highly confident and badly calibrated — stating 0.90 when the true rate is 0.60 — which is the specific failure mode that makes a stated probability worthless without a calibration check behind it.
The practical test for a funder evaluating a provider is simple to state and hard to fake: ask for the reliability curve, not the accuracy number. Plot stated probability against observed frequency on holdout data. A calibrated model tracks the diagonal. A model that runs above the diagonal at high stated probabilities is overconfident exactly where underwriting decisions concentrate — the top of the funding queue — which is the worst place for miscalibration to hide.
A point estimate throws away the information that matters for pricing
A probability of 0.68 for a favorable outcome tells you almost nothing about the shape of the 32% you are not funding for, or the dispersion inside the 68% you are. Two cases can both carry a 0.68 probability and have completely different risk profiles: one drawn from a tight distribution where outcomes cluster near the median, the other drawn from a bimodal distribution where the case either resolves strongly in the plaintiff's favor or dismisses outright, with almost nothing in between. A portfolio built on point estimates alone treats these as identical positions. They are not, and they should not carry the same size or the same reserve.
The distribution also carries the information a point estimate cannot: the settlement-value band conditional on a favorable outcome, the duration distribution conditional on each outcome branch, and the correlation structure with other cases sharing the same jurisdiction, judge, or defendant. Pricing decisions — advance rate, discount rate, portfolio concentration limits — should be built against these, not against the single summary number that gets quoted in a one-line pitch.
Reading the bands: what a labeled distribution should tell you
A properly labeled output separates the outcome-probability distribution from the settlement or award band conditional on a favorable outcome, and tags each input by its source — case-specific evidence the model was given, versus prior-adjusted estimates drawn from the comparable-case cohort. This separation matters because it tells you which parts of the number are grounded in this case's actual facts and which parts are inherited from the base rate for similar cases. A number built almost entirely from prior-adjusted inputs, with thin case-specific evidence, deserves a wider confidence band and a smaller position size than one grounded mostly in case-specific facts.
It is also worth stating plainly what a calibrated distribution is not: it is not a prediction of the realized dollar value this specific case will produce, and it is not a forecast of IRR or MOIC on its own. Those are computed downstream, by applying stated deal terms and negotiated positions to the distribution — a separate calculation with its own assumptions, which should be inspectable and stress-tested independently of the underlying outcome model.
Checking a calibration claim instead of taking it on faith
A provider's calibration claim is falsifiable, which means it can and should be checked before it is relied on. Ask for the specific holdout sample size behind any reliability curve — a curve built from a few dozen cases at the high end of the probability range carries far less evidence than one built from several hundred, even when both are plotted on the same chart at the same visual scale. Ask whether the holdout is temporal (later cases the model never trained on) or a random split of the same time period, since a random split can hide a model that has quietly learned the statistical signature of its own era rather than a durable pattern that will hold going forward.
Watch for a single headline accuracy or AUC figure presented without the reliability curve behind it. A model can have a strong AUC — good at ranking cases from lower to higher risk — while still being poorly calibrated, meaning wrong about the actual probability level, and the two failures have different consequences for pricing. An AUC failure costs you the wrong cases. A calibration failure costs you the wrong price on the right cases, which is the more expensive and harder-to-detect mistake of the two.
A related check worth running is whether the provider retrains and revalidates on a defined cadence, and whether performance is re-measured after each retrain rather than reported once and left standing. A model's calibration can drift as courts, dockets, and case mix shift over time, and a reliability curve published a year ago says nothing about whether the model is still calibrated today. Ask when the fleet was last retrained for the specific jurisdiction and case type you are underwriting, and ask to see the post-retrain calibration check, not just the original one used to launch the model.
What to ask for from an intelligence provider
- 01The reliability curve (stated probability vs. observed frequency) for the specific case type and jurisdiction you are underwriting, not an aggregate figure across the whole fleet.
- 02The full outcome distribution and the settlement or award band conditional on each branch — not a single collapsed number.
- 03Which inputs are case-specific evidence versus prior-adjusted estimates from the comparable cohort, tagged individually.
- 04A confidence interval or dispersion measure alongside every point estimate presented in a memo or dashboard.
Statistics shown reflect historical or illustrative model outputs derived from real case data. They are not predictions or guarantees of any individual outcome. Litigation results depend on facts, jurisdiction, judge, and counsel, and vary case by case. Model accuracy is subject to selection effects and changing legal dynamics.
Run a calibration check on your own book.
Start a Portfolio Audit