Score drift is when the same evidence stops producing the same score over time, because reviewers' standards loosen or tighten, the pool being scored changes, or the link between the evidence and real outcomes shifts, so this year's 4 no longer means what last year's 4 meant.
Three causes of score drift
- Reviewer drift. People change how they score. Interviewers grow tired of the role and get harsher; partners in a hot market get generous. New reviewers join with their own habits. The rubric stays the same on paper while its use shifts.
- Population drift. The pool changes. A fund that gets written up in the press suddenly sees twice as many inbound decks, mostly from earlier-stage companies. A job posted on a new board pulls a different mix of applicants. Average scores move even though nobody's standard did.
- Concept drift. The link between evidence and outcomes changes. A signal that predicted a strong Series A in one funding climate predicts less in the next. Machine learning calls this concept drift (Bayram and colleagues, 2022, review methods for detecting it), and it is the hardest of the three to see, because the scores can look stable while their meaning moves.
Most teams notice drift only when someone asks why a company passed at 3.4 in January and was rejected at 3.6 in September.
How to detect score drift
| Cause | Warning sign | Check | Fix |
|---|---|---|---|
| Reviewer drift | Average scores move with no change in the pool | Re-score a fixed anchor set of past cases each quarter | Calibration session, rubric rewrite |
| Population drift | Mix of sources, stages, sectors or locations shifts | Compare this period's input mix and score distribution with a reference period | Re-baseline, or score against a fixed reference pool |
| Concept drift | High scorers stop outperforming low scorers | Compare scores with later outcomes, cohort by cohort | Revisit criteria and weights; recheck model calibration |
Among people, also track inter-rater reliability over time. Falling agreement is often the first sign that reviewers are drifting in different directions.
Worked example: catching reviewer drift with an anchor set
The fictional Northwind Ventures keeps an anchor set: ten past seed deals with known outcomes. Each quarter, partners re-score them blind, using the current rubric.
| Quarter | Anchor set average | Live inbound average | Share of inbound sent to partner meeting |
|---|---|---|---|
| Q1 | 3.2 | 2.9 | 8% |
| Q2 | 3.3 | 3.1 | 11% |
| Q3 | 3.7 | 3.4 | 17% |
The anchor deals did not change, yet their average rose by half a point. That rules out a better pool as the explanation: the partners are scoring more generously, and the meeting rate has doubled as a result. The fund holds a calibration session on the three anchor deals that moved most, tightens the anchors, and the Q4 anchor average returns to 3.3.
Why score drift matters
Drift breaks comparisons over time, which is most of the value of scoring. A PE firm tracking a target market for two years needs a 4 in month one to mean the same as a 4 in month twenty. A hiring team filling the same role twice a year needs this year's hires to clear the same benchmark as last year's. Sales teams face the same problem: RevOps scoring reviews account tiers each quarter because tiers built on last year's wins drift away from what wins now.
Drift is also invisible from inside a single review. Daniel Kahneman and Olivier Sibony describe professional judgment as far noisier than the professionals themselves expect (McKinsey interview, 2021), and drift is noise that accumulates across months. NIST's AI Risk Management Framework names data, model and concept drift as reasons automated systems need regular maintenance, and calls for monitoring systems once they are in production (NIST AI RMF 1.0).
Common mistakes
- Treating rising averages as a better pool. Check the anchor set before celebrating.
- Changing the rubric without versioning it. Scores from before and after a change cannot be compared unless you know which rubric produced them. This is an auditability problem as much as a scoring one.
- Only watching the average. The spread matters too. Scores bunching in the middle often means reviewers have stopped committing.
- Retraining a model on drifted labels. If reviewers drifted, a model trained on their recent scores learns the drift.
How ScoringFactory approaches score drift
ScoringFactory applies one bar, learned from a team's past yes and no decisions, to every profile, and keeps each score with the person or company along with the record it was cited to. That makes it possible to see when and why a score changed, and to check whether the bar itself still matches how the team decides. See the flow.
Frequently asked questions
What causes score drift?
Three things. Reviewers change how strictly they score, usually gradually and without noticing. The pool changes, so average scores move even under a fixed standard. Or the relationship between evidence and outcomes changes, so the same score predicts less than it used to. Each needs a different check and a different fix.
How do you detect drift in a scoring model?
Watch three things over time: the distribution of scores compared with a reference period, the mix of inputs coming in, and whether high scores still go with good outcomes. For human reviewers, re-score a fixed set of past cases every quarter. If those scores move, the reviewers moved, not the cases.
How do teams keep scoring consistent across partners?
Use one written rubric with concrete anchors, have partners score independently before discussing, measure agreement regularly, and hold short calibration sessions on the cases where partners disagree most. Keeping a fixed anchor set of past deals gives everyone a shared reference point that does not change from quarter to quarter.
Is score drift the same as concept drift?
Concept drift is one cause of score drift. It refers to the relationship between inputs and outcomes changing, a term from machine learning. Score drift is broader: it also covers people scoring differently over time and the pool itself changing. All three make old and new scores harder to compare.
Sources
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST, 2023
- Sounding the alarm on system noise (interview with Daniel Kahneman and Olivier Sibony), McKinsey & Company, 2021
- Bayram, Ahmed and Kassler (2022), From concept drift to model degradation: an overview on performance-aware drift detectors, Knowledge-Based Systems (Elsevier)