Glossary

Model calibration

By ScoringFactoryUpdated First published 28 May 20264 min read
Definition

Model calibration is the property that a model's predicted probabilities match what actually happens: of all the founders, deals or candidates it scores at 70 percent, roughly 70 percent should turn out positive. A well-calibrated model's numbers can be read literally as odds.

How model calibration works

A predictive scoring model often outputs a probability: this candidate has a 0.6 chance of passing the final round, this company a 0.3 chance of raising a Series A within two years. Calibration asks whether those numbers are honest. The standard test, set out in the survey Classifier calibration by Silva Filho and colleagues (Machine Learning, 2023): among cases given a probability close to 0.8, about 80 percent should actually be positive.

To check it:

  1. Take past cases where you know the outcome.
  2. Group them into bins by predicted probability, for example 0 to 0.2, 0.2 to 0.4 and so on.
  3. For each bin, compare the average prediction with the share that actually turned out positive.
  4. Plot one against the other. This plot is a reliability diagram. A calibrated model sits on the diagonal.

Two summary numbers are common. The Brier score, from Glenn Brier's 1950 work on verifying weather forecasts, is the mean of (predicted probability - outcome)^2, where outcome is 1 or 0; lower is better. Expected calibration error is the weighted average gap between predicted and observed rates across bins.

Worked example: an overconfident deal model

A fictional fund tests a model on 600 past seed deals, predicting whether each raised a Series A within three years.

Predicted rangeDealsAverage predictionActually raisedObserved rate
0 to 20%20010%189%
20 to 40%18030%4927%
40 to 60%12050%4235%
60 to 80%7070%2840%
80 to 100%3088%1447%

The ranking is fine: higher bins do better. But the model is overconfident at the top. A partner told "88 percent" would expect almost every deal in that bin to raise; fewer than half did. If the fund sized reserves from those numbers, it would set aside money for follow-ons that never happen. Recalibrating would pull the top predictions down toward 40 to 50 percent without changing the order.

Why model calibration matters for investors and hiring teams

If a score is only used to sort a list, as a fit score usually is, calibration matters less: the order is what counts. Once someone reads the number as a probability, it matters a lot. A PE firm comparing expected value across targets, a fund setting reserves, or a hiring team deciding how many finalists to interview to fill one seat are all doing arithmetic with the probability.

Do not assume a model is calibrated. Guo and colleagues reported in 2017 that modern neural networks were often overconfident, and a one-parameter fix called temperature scaling became the standard repair. A 2021 follow-up, Revisiting the Calibration of Modern Neural Networks (NeurIPS 2021), found the newest image models among the best calibrated and pointed to architecture as a major factor. Calibration depends on the model, so measure it.

Model calibration vs accuracy vs human calibration

Model calibrationDiscrimination (accuracy, AUC)Calibration of reviewers
QuestionDo predicted probabilities match outcome rates?Does the model rank positives above negatives?Do reviewers mean the same thing by a score?
Measured byReliability diagram, Brier score, calibration errorAUC, precision, recallInter-rater reliability statistics
Fixed byRecalibration (sigmoid, isotonic, temperature scaling)Better features or a better modelCalibration sessions and clearer rubrics

A model can rank perfectly and still be badly calibrated, as in the example above. It can also be well calibrated and rank poorly, for instance by predicting the base rate for everyone. Reviewer labels from qualitative scoring can be checked the same way: if a partner's "strong yes" calls succeed no more often than her "maybe" calls, the labels are not calibrated.

Common mistakes

  • Checking calibration on the training data. Use cases the model has not seen.
  • Too few cases per bin. Thirty deals in a bin gives a noisy observed rate; widen the bins or pool more history.
  • Calibrating once. When the market or applicant pool shifts, calibration slips. Recheck it as part of watching for score drift.
  • Calibrating to the wrong base rate. A model trained on a hot funding year will overpredict in a cold one. Benchmark against recent outcomes.
  • Presenting a raw model output as a probability. Many scoring models output a score, not odds. Only call it a probability once it has been checked.

How ScoringFactory approaches model calibration

ScoringFactory presents scores as a ranking against the team's own bar, each one cited to the record behind it, rather than as a promise of odds. That keeps the number honest about what it is, and leaves the judgment about likelihood with the people who make the call. See the flow.

Frequently asked questions

What does it mean for a model to be well calibrated?

It means the probabilities it outputs can be taken at face value. If you collect every case the model scored around 30 percent, about 30 percent of them turn out positive. The same holds across the range. A well-calibrated model lets you plan with its numbers as well as sort by them.

How do you measure calibration?

Bin past predictions by probability, then compare each bin's average prediction with its observed outcome rate. Plot the result as a reliability diagram; a calibrated model follows the diagonal. Summarize with the Brier score, the mean squared gap between predictions and outcomes, or with expected calibration error, the average gap across bins.

How is model calibration different from accuracy?

Accuracy, or discrimination, asks whether the model ranks good cases above bad ones. Calibration asks whether its probabilities are the right size. A model that gives every good deal 0.9 and every bad one 0.8 ranks perfectly but is badly calibrated if far fewer than 80 percent of deals succeed.

How do you fix a poorly calibrated model?

Fit a small correction on held-out data that maps raw outputs to observed rates. Common choices are sigmoid (Platt) scaling, isotonic regression, and temperature scaling for neural networks. These change the numbers without changing the order, so a model's ranking stays the same while its probabilities become honest.

Sources

  1. Silva Filho, Song, Perello-Nieto, Santos-Rodriguez, Kull and Flach (2023), Classifier calibration: a survey on how to assess and improve predicted class probabilities, Machine Learning (Springer)
  2. Minderer, Djolonga, Romijnders, Hubis, Zhai, Houlsby, Tran and Lucic (2021), Revisiting the Calibration of Modern Neural Networks, NeurIPS 2021