Score drift is what happens when the same evidence stops producing the same score. One partner reads a founder's track record as a 9, another reads the identical record as a 6. A model scores a candidate one way in January and a different way in June with no change in the underlying facts. The bar hasn't been redefined on purpose. It has just quietly moved.
Drift rarely announces itself. It shows up as small, accumulating gaps: a partner who has gotten more skeptical after a bad quarter, a hiring manager who calibrates against whoever they interviewed last, a model whose outputs shift as new training data or new prompts get layered in without anyone re-running the old cases to check. Left unmeasured, none of this looks like a problem. It looks like normal variance, until someone lines up two decisions side by side and finds the same facts got two different verdicts.
A May 2026 Stanford-led study of AI hiring algorithms, covered by Fortune, found exactly this kind of hidden inconsistency at scale: when a vendor's scoring model was evaluated in aggregate across 1,746 positions, it looked fine, but when researchers broke the results out position by position, more than 10 percent of roles showed a discriminatory pattern the pooled view had buried. That is drift hiding inside an average. The system was consistent in the way that mattered least, and inconsistent in the way that mattered most.
A venture fund's entire pipeline runs on one assumption: that a 9 today means the same thing a 9 meant last year. If it doesn't, the fund isn't actually applying a bar, it's applying whatever mood the reviewer was in that week. The same failure mode hits portfolio hiring. A 2026 recruiting benchmarks analysis from Humanly puts it directly: the metric that actually predicts whether an assessment process works is calibration, whether different reviewers reach the same call when they're shown the same evidence. Everything else, including speed and volume, is secondary to that.
This is the problem founder scoring is built to solve. A rubric defined once and applied the same way to every founder removes the option for the bar to silently shift between deals. Every score carries the evidence that produced it, so a partner reviewing a decision six months later can check it against the same standard, not against whatever the reviewer happened to be feeling that day. We wrote about this directly in Same bar, every partner: killing score drift.
Not every change in scores is drift. A fund can and should update its bar on purpose, tightening it in a downturn or loosening a specific dimension because the market changed. The difference is that deliberate recalibration is documented, applied uniformly, and versioned, everyone knows the bar moved and why. Drift is the same shift happening by accident, unevenly, with no record of it. The fix for drift isn't freezing the bar forever. It's making every change to the bar visible, and making the reasons for a score checkable against evidence rather than memory.
They're related but not identical. Bias is a systematic skew in one direction. Drift is any change in consistency over time or across reviewers, which can introduce bias but can also just be noise. Both are caught the same way: by checking whether identical evidence produces identical scores.
Yes, and it can be harder to spot. A model's outputs can shift after a retrain, a prompt change, or new data without anyone changing the stated rubric. That's why evidence-backed scoring keeps every score linked to the specific facts that produced it, so drift shows up as soon as you compare cases instead of months later.
Fix the rubric, attach evidence to every score, and periodically re-score a sample of past decisions to check they'd come out the same today. If they don't, that's drift, and you know exactly where to look.
Bring a handful of past decisions. We'll re-score them against your rubric and show you exactly where the bar moved.