Calibration is the process of tuning a scoring system, human or model, so its outputs match reality and stay consistent no matter who applies it or when. A calibrated 8 out of 10 means the same thing every time it's given, whether the evaluator is a partner reading a pitch or a model reading a resume.
Calibration starts with a reference set: a stack of past cases with known outcomes, funded deals, hired candidates, closed accounts, that a team already agrees on. Every evaluator or model runs against the same set, and the results get checked two ways: do different reviewers land on the same score for the same evidence, and does the score actually predict what it claims to predict.
A calibration pass typically checks:
A rubric is only as good as the agreement behind it. If one partner's 7 is another partner's 4, the score stops being useful for ranking anything. Recruiting benchmark research from 2026 makes the same point about hiring panels: the metric that actually predicts hiring quality isn't a satisfaction survey, it's whether different interviewers reach the same call when shown the same evidence. That's calibration. It's why founder scoring and candidate scoring only hold up over time if the bar behind them is checked and re-tuned on a schedule, not set once and left alone. We wrote more about what breaks when that check is skipped in same bar, every partner: killing score drift.
General calibration applies to an entire scoring process, including human panels. Model calibration is the narrower, statistical version of the same idea: whether a model's predicted probabilities match real-world frequency, so that of every 100 cases it scores at 80%, roughly 80 actually turn out the way it predicted. Model calibration is one instrument for measuring the broader kind. A fund or hiring team needs both: a calibrated model underneath, and a calibrated bar across the humans reading its output.
No. Calibration measures consistency and matched confidence, not correctness. A scoring system can be well calibrated, its 80% calls hit 80% of the time, and still be less accurate overall than a different approach that isn't calibrated at all.
Whenever the underlying bar changes, a new fund thesis or a new hiring bar, or whenever outcome data shows the score is drifting from reality. Most teams review calibration quarterly and re-check it any time a new evaluator joins.
Whoever owns the rubric, typically a platform lead or head of talent. Their job is to run the agreement checks, flag drift, and retrain reviewers or retune the model when scores stop lining up.
Bring a founder you're diligencing. We'll score them against your bar, live, with the receipts behind every number.