Qualitative scoring is the practice of rating judgment-based criteria, such as communication, conviction, or clarity of thinking, on a fixed scale where each level has a written anchor and each rating carries a stated reason, so another reviewer can check the call.
How qualitative scoring works
Some of the things a team cares about most have no number attached. How clearly a founder explains a hard tradeoff. Whether a candidate owns a past failure or blames the team. How much a management team believes its own plan. Qualitative scoring turns those impressions into ratings that can be compared and challenged.
A workable setup has four parts:
- A named criterion. "Communication" is too broad. "Explains a technical decision to a non-technical listener" can be rated.
- A fixed scale. Usually 1 to 4 or 1 to 5. An even number of levels stops reviewers parking everyone in the middle.
- Anchors for each level. A sentence or two describing what a 2 looks like and what a 4 looks like, written before anyone is rated.
- A reason per rating. One line pointing at what the person said or did. No reason, no score.
The anchors are the part most teams skip, and they matter most. A widely used form is the behaviorally anchored rating scale (BARS), where each point on the scale is tied to an example of real behavior collected from people who know the job. SHRM's guidance on interview rating scales makes the same point for hiring: anchor each level to observable, job-related behavior, and give raters examples of poor, acceptable and excellent answers. The anchors live inside a rubric, and the reasons are what make it evidence-based scoring rather than a gut feel with a number on it.
Why qualitative scoring matters
In venture capital, the soft criteria often carry the decision. In a survey of 885 venture capitalists, Gompers, Gornall, Kaplan and Strebulaev (2020) found the management team was named as an important factor by 95 percent of firms and as the most important factor by 47 percent. Most of what a partner learns about a team in a first meeting is qualitative.
Hiring has the same shape. Structured interviews, which ask every candidate the same questions and grade answers against the same standard, are the method the U.S. Office of Personnel Management describes for measuring job-related competencies. A structured interview is qualitative scoring with the question set fixed in advance.
Unanchored judgment is noisy. In a 2021 McKinsey interview, Olivier Sibony describes an investment firm that gave all its analysts the same company to value. On average, any two analysts differed by 44 percent, and the firm's leaders had no idea the gap was that large. Anchors and written reasons are the cheapest way to shrink that gap.
Worked example: scoring founder clarity
Northwind Ventures, a fictional seed fund, scores "clarity of thinking" on a 1 to 4 scale after every first meeting. Its anchors read:
- 1: Cannot say who the first customer is or why they would pay now.
- 2: Names a customer but the reason to buy is generic ("saves time").
- 3: Names a customer, a specific pain, and one piece of proof, such as a pilot or a quote.
- 4: All of level 3, plus a clear answer on what would prove the plan wrong.
Two partners meet the founder of Ledgerly, a fictional Series A fintech. Partner A gives a 3: "Named mid-size accounting firms, month-end reconciliation pain, two paid pilots." Partner B gives a 4: "Same, and said churn above 5 percent a month in pilots would kill the thesis." The one-point gap is easy to settle because both reasons point at things the founder said. Without the reasons, the partners would argue about whether they liked the founder.
Qualitative vs quantitative scoring
The two are not rivals. Most real scorecards mix them.
| Qualitative scoring | Quantitative scoring | |
|---|---|---|
| Input | A reviewer's judgment of what they saw or heard | A measured value, such as revenue growth or years in role |
| Typical criteria | Communication, conviction, ownership, coachability | Retention, burn multiple, quota attainment |
| Main risk | Reviewers disagree or drift | Measuring what is easy instead of what matters |
| How to check it | Anchors, written reasons, inter-rater reliability | Data quality checks and source verification |
| Where it shows up | Interview notes, partner meeting write-ups | Data rooms, CRM fields, financial models |
Once rated, a qualitative criterion behaves like any other input to a scorecard: it can be weighted, combined, and compared, and it can help decide who makes the shortlist.
Common mistakes
- Scales without anchors. "Rate communication 1 to 5" produces five private definitions of a 4.
- Criteria that overlap. "Presence", "confidence" and "leadership" often measure the same impression three times.
- Rating before writing. Write the reason first, then pick the number. The reverse invites rationalising.
- Sharing scores before everyone has rated. The first number spoken becomes the anchor for the room.
- Never checking agreement. Run a calibration session every quarter: score the same three past cases independently and compare. If the ratings feed a model that outputs probabilities, check model calibration as well: a predicted 70 percent should come true about 70 percent of the time.
How ScoringFactory handles qualitative criteria
ScoringFactory learns what a team's yes and no decisions have rewarded on criteria like clarity or ownership, then applies that bar to each new founder, company, or candidate. Every rating comes with receipts: the part of the record it rests on, cited to the line. Your team reads the reason, agrees or overrides, and makes the call. See how the flow works.
Frequently asked questions
How do you score qualitative criteria consistently?
Write anchors for every level of the scale before anyone is rated, require a one-line reason that points at something the person said or did, and have reviewers score independently before discussing. Then compare ratings on a few shared cases each quarter. Where two reviewers keep landing far apart, the anchor for that criterion needs rewriting.
What are behaviorally anchored rating scales?
Behaviorally anchored rating scales (BARS) tie each point on a rating scale to a described example of real behavior. The examples are gathered from people who know the job, sorted into dimensions, and kept only when experts agree on where they belong. Reviewers then compare what they observed to the examples instead of to their own private standard.
Can qualitative judgments be turned into numbers?
Yes. Once each level of a scale has a written anchor, a reviewer's judgment becomes a number that can be weighted and combined with measured data. The number is only as good as the anchor behind it, so keep the written reason next to the score. The reason is what lets someone else check the rating later.
Is qualitative scoring less reliable than quantitative scoring?
Unanchored qualitative ratings are usually less reliable, because reviewers apply different private standards. Anchored ratings with written reasons can reach good agreement between reviewers. Quantitative data has its own failure: it can be precise and still measure the wrong thing. Most strong scorecards use both kinds of criteria.
Sources
- Structured Interviews, U.S. Office of Personnel Management, current text
- Gompers, Gornall, Kaplan and Strebulaev (2020), How do venture capitalists make decisions?, Journal of Financial Economics
- Sounding the alarm on system noise (interview with Daniel Kahneman and Olivier Sibony), McKinsey & Company, 2021
- Build Consistent Hiring Decisions with Job Interview Rating Scales, SHRM, 2026