Glossary

Calibration

By ScoringFactoryUpdated First published 1 July 20264 min read
Definition

Calibration is the practice of aligning the people, or the model, doing the scoring so that a given score means the same thing every time it is given, whichever partner, interviewer or week it came from, usually by scoring shared cases and resolving the differences.

What calibration means in hiring and investing

Two interviewers both give a candidate a 4. One means "strong hire". The other gives everyone a 4 unless something goes wrong. Their scores look the same and mean different things. Calibration is the work of closing that gap.

In practice it happens in a calibration session: reviewers score the same cases independently, compare, and talk through every case where they differ until they agree on what the evidence was worth. The output is a sharper rubric and reviewers who apply it the same way. It matters most in qualitative scoring, where criteria like conviction or clarity of thinking leave the most room for private scales.

The word has a second, statistical meaning: whether a model's predicted probabilities match real outcome rates. That is covered under model calibration. This page is about people.

Why calibration matters

Professional judgment varies far more than most firms expect. In a 2021 interview with McKinsey Quarterly, Olivier Sibony, a coauthor with Daniel Kahneman of the book Noise, described a noise audit at an investment firm. Analysts who were supposed to apply the same valuation methods were given the same company data, and on average two analysts' valuations differed by 44 percent. The firm's leaders had no idea the variation was that large, and Sibony said that reaction is common.

For a venture fund, uncalibrated partners mean a founder's chance of a second meeting depends on who took the first one, and by the time a deal reaches the investment committee, a 4 in one partner's memo may not mean a 4 in another's. For a hiring team, it means the hiring bar moves with the interview schedule. Calibration is how a team gets one bar instead of several.

How to run a calibration session

  1. Pick five to ten past cases with known outcomes: some clear yeses, some clear nos, and a few that were close.
  2. Everyone scores each case alone, using the current rubric, without seeing anyone else's scores.
  3. Collect the scores and sort cases by spread. Start with the widest.
  4. For each wide case, the highest and lowest scorer explain which evidence drove their number.
  5. Decide what the evidence should have scored, and rewrite the rubric anchor that allowed the disagreement.
  6. Measure agreement before and after with an inter-rater reliability statistic, such as an intraclass correlation (see ten Hove, Jorgensen and van der Ark, 2024, for which form to use), so you know the session worked.

OPM's guidance on structured interviews rests on the same idea: every answer judged against the same rating scale and the same standard for an acceptable response.

Worked example: three partners, four deals

The fictional Northwind Ventures has three partners score four past deals on founder quality, 1 to 5.

DealPartner 1Partner 2Partner 3Spread
Ledgerly4431
Fernhill Robotics5233
Quarry Labs2220
Oakline Health3542

Fernhill has the widest spread. Partner 1 scored the founder's previous exit; partner 2 scored the fact that the exit was in an unrelated market. The fund decides prior exits only count at the top level when they were in the same or an adjacent market, writes that into the anchor, and rescoring brings Fernhill to 3, 3, 3.

CalibrationModel calibrationInter-rater reliability
Applies toPeople (and the rubric they use)A model's predicted probabilitiesAgreement among people
Is itA practiceA propertyA measurement
Question it answersDo we mean the same thing by a 4?Do 70 percent predictions come true 70 percent of the time?How often do we agree, beyond chance?

Calibration is also different from benchmarking, which compares a score with a reference set. A benchmark only means something once reviewers score the same way.

Common calibration mistakes

  • Discussing before scoring. The session then measures how persuasive the first speaker was.
  • Only easy cases. Everyone agrees on the obvious ones. The value is in the close calls.
  • Agreeing without rewriting the rubric. The alignment fades within weeks if it is not written down.
  • Letting the senior person settle every case. If the most senior reviewer is always right, the session teaches everyone to guess what they think rather than read the evidence.
  • Calibrating once. Standards slip as the pool and the team change. See score drift.

How ScoringFactory approaches calibration

ScoringFactory applies one learned bar to every founder, company or candidate, so the score does not depend on who opened the file. Because every score is cited to the record, a calibration session can argue about specific evidence rather than impressions. See the flow.

Frequently asked questions

What is a calibration session in hiring or investing?

A calibration session is a meeting where interviewers or partners score the same set of candidates or deals on their own, compare the results, and resolve the differences. The goal is a shared understanding of what each score level means, written into the rubric so it lasts beyond the meeting.

How do you calibrate interviewers or partners?

Use past cases with known outcomes. Have everyone score them independently, start the discussion with the cases that have the widest spread, and ask the highest and lowest scorers to name the evidence behind their numbers. Rewrite the rubric anchor that caused the gap, then rescore to confirm the gap closed.

How often should a team recalibrate?

At least quarterly for teams scoring every week, and whenever something changes: a new interviewer or partner, a new role or fund strategy, or a visible shift in average scores. Short, regular sessions on a handful of cases work better than a long annual one, because drift builds gradually.

Can a model be calibrated the same way?

Not quite. A model is calibrated statistically, by checking whether its predicted probabilities match observed outcome rates and adjusting if they do not. That is model calibration. But the people reviewing a model's scores still benefit from calibration sessions on how they act on those scores.

Sources

  1. Structured Interviews, U.S. Office of Personnel Management, current text
  2. Sounding the alarm on system noise (interview with Daniel Kahneman and Olivier Sibony), McKinsey Quarterly, 2021
  3. Updated guidelines on selecting an intraclass correlation coefficient for interrater reliability (ten Hove, Jorgensen and van der Ark, 2024), Psychological Methods (APA)