Model calibration is the process of tuning a scoring model so its predicted outputs match real-world outcomes: if the model gives a batch of candidates or deals a score of 80, roughly 80% of them should actually turn out to be strong. It is a property of the model itself, distinct from calibration between human reviewers.
A scoring model can be perfectly good at ranking, putting the strongest candidates or founders above the weaker ones, while still being badly calibrated. Ranking answers "who is better than whom." Calibration answers a different question: "does a score of 80 mean what we think it means." A model that assigns 90+ scores to half its inputs is uncalibrated even if its ordering is correct, because the number stops carrying information.
Calibration is usually checked with a reliability curve: bucket every prediction by score, then compare the predicted rate against the observed outcome rate in each bucket. A well-calibrated evaluation rubric or model keeps those two lines close together across the full range of scores, not just at the extremes.
ScoringFactory scores founders and candidates against a defined bar, and that bar is only useful if the numbers behind it are calibrated to what actually happened, not just to what feels right in the moment. A 2026 review of recruiting benchmarks put it directly: the real quality benchmark for a scoring system is whether "different recruiters and interviewers make the same call when shown the same evidence," and teams are advised to track how far interviewer scores drift apart and run regular calibration reviews on edge cases rather than treating disagreement as a personnel problem.
The cost of skipping this work shows up in production. A May 2026 Stanford-led audit of hiring algorithms, reported by Fortune, found that when researchers checked outcomes position by position instead of in aggregate, 10.62% of jobs showed adverse impact on Black candidates, with more than 25% of Black applicants' submissions landing in roles where the model's outputs met federal discrimination standards. Aggregate accuracy looked fine; the model was miscalibrated for specific segments of the population, and nobody caught it until they checked the outcomes directly. That is the exact failure mode model calibration exists to catch, and why we treat every score in the platform as something to be audited against what actually happened, not just shipped and trusted.
Calibration is never a one-time step. Every score a partner or hiring manager acts on gets tied back to an outcome later: did the deal close, did the hire perform, did the reference check hold up. Those outcomes feed back into the model on a fixed cadence, and any segment where the model's confidence and the real outcome diverge gets flagged before it silently drifts the whole system, the same discipline behind our post on tying every score to a line of evidence.
No. A model can rank candidates correctly and still be miscalibrated if its scores overstate or understate confidence. Accuracy asks whether the ordering is right; calibration asks whether the number itself is trustworthy.
On a fixed cadence, not just after something goes wrong. Hiring patterns, deal flow, and market conditions shift, and a model that isn't re-checked against fresh outcomes will quietly drift out of step with reality.
No. It keeps the model's confidence honest so a partner or hiring manager can trust what a score is telling them. The decision still sits with the person; calibration just makes sure the input to that decision means what it claims to mean.
Bring a real deal or hire. We'll show you the score, the calibration behind it, and the evidence that backs every point.