I have run hundreds of reference calls in this seat. A good one feels like real diligence. Most of the time it is theater, and the founder wrote the script before I dialed in.
Ask any partner to describe their reference process and they will tell you it is rigorous. Three calls, maybe four, structured questions, careful notes. What they usually will not say out loud is who chose the three or four people on the other end of the line. The founder did. Every single time. That single fact should discount everything that follows, and in most partnerships it does not.
You picked the panel before I called
A reference list is a curated exhibit, not a random sample. Nobody hands an investor the cofounder they fired, the early engineer who quit over a missed equity promise, or the customer who churned after a bad renewal conversation. They hand you the people who will say yes to the call and say nice things once they are on it. That is not dishonesty on the founder's part. It is just how anyone would build a defense exhibit if they got to pick the witnesses.
The scale of the problem is bigger than most partners assume. GCheck's 2026 Trust in Hiring Report found that 93 percent of job seekers admit to embellishing something during a hiring process, from exaggerated scope of a role to invented answers under interview pressure. Founders raising a round are candidates too, just for a bigger job with worse odds. If nine out of ten will shade the story on their own résumé, there is no reason to assume the story their handpicked references tell is any straighter.
Where the format breaks
Selection bias is the root problem, but it shows up in the room as three specific failure modes.
The halo effect. Once a reference decides they like the founder, every answer gets pulled toward that verdict. Ask about weaknesses and you get a strength dressed up as a flaw: "works too hard," "cares too much." The reference is not lying. They are pattern-matching to their overall impression, and the overall impression is favorable because the founder chose them for exactly that reason.
Small sample, big weight. Three or four calls is not a sample, it is an anecdote with a phone number. A recent review of the reference-checking literature by Not-For-Profit People found that referees rating the same candidate often give substantially different scores, and that a favorable report is just as likely to reflect a lenient referee as a genuinely strong candidate. The article also flags something we see constantly in venture: candidates and founders from well-connected backgrounds can source more polished, senior-sounding referees, which rewards network more than substance. This is one of the clearest places where bias mitigation in a diligence process actually matters, because the format itself is quietly favoring people who already have access.
Nobody wants to be the bad reference. Even a reference with real doubts feels social pressure not to be the one negative data point, especially in an industry as small as venture where the founder, the reference, and the investor will likely cross paths again. The same research notes that references are rarely verified and referees suffer almost no reputational cost for talking up an average candidate. The safe move is always a shrug and a compliment, so that is what most people do.
A reference call tells you what one person is willing to say to your face. The record tells you what happened when nobody was managing the narrative.
What the record catches, the call misses
This is why due diligence that stops at phone calls is incomplete diligence. The public record does not get curated the way a reference list does. A founder's commit history, if they are technical, shows how they actually behaved under pressure long before anyone thought to ask a reference about it, the same way we read a founder's GitHub before the first call. Court filings, past company dissolutions, prior investor updates, and customer reviews left before a fundraise was ever on the horizon all share one property references do not: nobody wrote them for your benefit.
The record is not perfect either. It is incomplete, sometimes outdated, and it needs a human to interpret it correctly, which is exactly the argument for evidence-based scoring rather than gut-feel diligence. Every fact we surface links back to its source, so when a partner disagrees with a score, the disagreement is about the evidence, not about whose gut is more trustworthy.
As diligence has gotten faster with AI tools doing the first pass, the hours investors save are not disappearing, they are getting redirected into deeper backchanneling and more customer calls, according to VC Minute's February 2026 breakdown of AI's impact on due diligence. The bar has gone up, not down. Founders should assume their uncurated customers, not just their chosen references, will get a call.
Bring evidence into the call
None of this means we skip reference calls. We run them on every deal, and a good conversation with someone who worked closely with a founder still tells you things no dataset will. What changes is how we walk in. Instead of opening with "tell me about working with them," we open with something specific the record already surfaced: a resignation that lined up with a product pivot, a pattern of late deliverables in a public repo, a customer review that contradicts the growth story in the deck. A reference who is coasting on vague praise has to respond to a fact instead of a feeling, and that is where the real conversation starts.
That is the whole shift. We are not replacing the call, we are refusing to walk into it blind. The reference still gets to talk. We just stop pretending their word is the only evidence in the room.
