A founder pitches Monday morning and gets a 7.5. The same deck, same numbers, same team, pitched to a different partner on Friday afternoon, gets a 5.5. Nothing about the company changed. That gap is score drift, and it is the quiet reason most diligence scores don't survive contact with a second opinion.
I have watched this happen inside our own product before we fixed it. Two partners, same rubric open in two tabs, same founder profile, and scores four points apart on a ten-point scale. Neither partner was wrong. They were both applying honest judgment to a category ("team quality," say) that the rubric described in one sentence and left everyone to interpret on their own.
What score drift actually is
Score drift is any case where the same subject, a founder in diligence or a candidate in a hiring pipeline, gets a materially different score depending on who is scoring, when they're scoring, or what they scored right before it. It shows up in three flavors we see constantly:
- Rater drift. Partner A runs warm on "market size," Partner B runs cold. Same evidence, different number, every time.
- Order effects. The fourth pitch of the day gets scored against the third pitch, not against the bar. A mediocre founder looks great right after a genuinely weak one.
- Time drift. The bar itself quietly moves. A team that would have scored a 6 in January scores a 7 in June because the pipeline got weaker and everyone's sense of "average" recalibrated without anyone deciding to.
None of this requires bad faith. It's what happens by default when judgment isn't anchored to anything external. Researchers building AI systems to score startups have run into the identical problem from the other direction: a multi-agent system called DIALECTIC, described in a February 2026 paper out of TUM and General Catalyst, had to structure its scoring as a simulated debate between arguments for and against a deal specifically because a single model asked to just "rate this startup" reproduces the same inconsistency a single human partner does. Consistency turned out to be a design problem, not a smarter-model problem.
Why a written rubric doesn't stop it
The obvious fix is a written rubric, and every firm we've talked to already has one. It doesn't work on its own, because a rubric only agrees on the categories. It doesn't agree on what an 8 looks like inside "founder-market fit" versus a 6. Two people can both be following the rubric in good faith and land four points apart.
A shared rubric tells two raters what to look at. It doesn't tell them where the line is. Only calibration does that.
This isn't unique to venture or hiring. A May 2026 study from Cambridge's ALTA Institute had up to 80 trained raters score the same set of essays against a detailed, shared analytic rubric, then applied a statistical model (a two-facet Rasch model) just to correct for the fact that individual raters were, on average, harsher or more lenient than their peers even after training. If 80 professionally trained raters working from an explicit rubric still need after-the-fact calibration to produce a fair score, a VC partnership scoring founders between meetings on instinct and a one-page rubric has no chance of staying consistent without the same kind of correction built in.
How we hold the bar
ScoringFactory treats this as infrastructure, not a training problem you solve once in an offsite. Three pieces do the work:
- The rubric is evidence-anchored, not adjective-anchored. Instead of "strong technical founder = 8," each rubric level is tied to worked examples with the actual evidence attached, so "8" means a specific pattern of commits, references, and outcomes, not a feeling. We wrote about how that evidence trail works in how we tie every score to a line of evidence, and it's the same mechanism that keeps two partners pointed at the same bar.
- Every rubric version is dated and every score is tagged to the version it was scored under. When the bar moves, and sometimes it should move, you can see exactly when and compare a March score to a June score honestly instead of pretending the number means the same thing six months later.
- We run blind recalibration passes. A held-out set of previously scored profiles gets rescored periodically without anyone seeing the original number. When a partner's rescore lands more than a point and a half off their own prior score, or off the panel average, that's flagged and reviewed, the same way the rubric itself gets reviewed in the diligence scorecard we run on every inbound founder.
Drift shows up as small, invisible disagreements until you measure it. A partner who scores 0.8 points harsher than the panel average on "GTM clarity" for six straight quarters isn't being difficult, they're an uncalibrated rater, and now it's visible instead of buried in six different memos.
Why consistency beats being "right"
There is no ground truth for whether a specific founder "really" deserves a 7 versus an 8. Nobody has that number. What a fund actually needs is for a 7 in January and a 7 in June to mean the same thing, and for a 7 from one partner to mean the same thing as a 7 from another. That's what lets you rank a portfolio, compare a Q1 cohort to a Q3 cohort, and trust that the founder who got passed over wasn't just unlucky enough to pitch on a busy Friday.
Getting any single score exactly "right" is a matter of opinion. Getting every score consistent with every other score is a matter of process. We built ScoringFactory to make the second one boring and automatic, so partners can spend their judgment on the calls that actually need it.
