Glossary

Scoring model

By ScoringFactoryUpdated 4 min read
Definition

A scoring model is the fixed set of criteria, scales, weights and rules that turns evidence about a founder, company, candidate or deal into one comparable score, so the same evidence always produces the same result no matter who runs it.

How a scoring model works

Every scoring model, from a hiring spreadsheet to a fund's screening tool, has the same five parts:

  1. Criteria. The things you judge, such as founder and market fit, revenue growth, or depth of system design experience.
  2. Scales. How each criterion is scored, usually 1 to 5, with a rubric that says what each level looks like. If criteria come on different scales, such as dollars and ratings, normalization puts them on one before they are combined.
  3. Weights. How much each criterion counts toward the total. See weighting.
  4. Rules. Hard filters and tie breaks, for example "no companies outside North America" or "missing evidence scores 1, not 3".
  5. Output. A total, a rank, or a band such as meet, maybe, pass.

The most common form is a weighted sum: score = w1 x s1 + w2 x s2 + ... + wn x sn, where each s is a criterion score and the weights add up to 1. The field of multiple-criteria decision analysis (Cinelli and colleagues, 2020, review its methods) covers more elaborate methods, but most teams never need them.

A model can be hand-built from a team's stated priorities or fitted to past outcomes. The second kind is usually called predictive scoring.

Why a written model beats a gut call

Two partners reading the same deck will weigh it differently, and the same partner will weigh it differently on Monday and Friday. A model holds the weighting still, so disagreements move to the evidence, which is where they belong.

The research on this is old and consistent. Robyn Dawes showed in 1979 that even crude linear models, including ones with equal or random weights, tend to match or beat expert judges. A 2020 hiring study found the same: across three samples of managers, a model with consistent but random weights reliably beat the experts' own predictions of job performance (Yu and Kuncel, 2020). A 2023 review of the meta-analyses found mechanical combination more valid than human judgment across fields, and about 50 percent more valid for predicting job performance (Neumann et al., 2023).

For a venture fund, that means a seed deal seen in March gets judged on the same terms as one seen in October. For a PE firm, it means a target list of 400 companies is ranked on one standard. For a hiring team, every candidate faces the same bar.

Worked example: scoring two seed deals

Northwind Ventures, a fictional seed fund, uses four criteria scored 1 to 5:

CriterionWeightLedgerly (fintech)Fernhill Robotics
Founder and market fit0.354 (1.40)3 (1.05)
Traction0.253 (0.75)5 (1.25)
Market size0.204 (0.80)3 (0.60)
Fit with the fund's thesis0.205 (1.00)2 (0.40)
Total1.003.953.30

Fernhill has the stronger traction, but it sits outside the fund's thesis and the founders are new to the sector, so Ledgerly ranks first. The useful part is the row-by-row view: a partner who wants to back Fernhill now has to argue that thesis fit should weigh less, or that the founder score is wrong. Both are better conversations than "I like it more".

Rules-based scoring models vs predictive models

Rules-based modelPredictive modelUnstructured judgment
Where weights come fromThe team decides themFitted to past outcomesImplicit, shifts by reviewer
Data neededVery littleMany past decisions with known resultsNone
Easy to explainYesDepends on the methodRarely
Main riskWeights reflect opinion, not resultsLearns old biases; drifts as the market changesNoise and inconsistency

Most teams start rules-based and move toward a fitted model once they have enough outcomes. Either way, check that a predicted 70 means what it says, which is model calibration.

Common mistakes when building a scoring model

  • Too many criteria. Twelve criteria at 8 percent each means none of them moves the score. Four to seven is usually enough.
  • Double counting. "Revenue" and "growth rate" and "ARR multiple" can all measure the same thing.
  • Scoring missing evidence as average. A blank should be marked unknown, or it quietly drags weak profiles toward the middle.
  • Never checking against outcomes. If the companies you scored 4.5 do no better than the ones you scored 3, the model is decoration.
  • Scores without receipts. A 4 on traction should point at the figure it came from. That is the idea behind evidence-based scoring.

How ScoringFactory approaches scoring models

ScoringFactory builds the model from the team rather than from a template: it learns the bar from the founders, companies or candidates a team said yes and no to, applies that bar to the whole pool, and cites every score to the record behind it. Your team still makes the call. See how the flow works.

Frequently asked questions

What is a scoring model?

A scoring model is a fixed recipe for turning evidence into a number: a list of criteria, a scale for each, weights that say how much each counts, and rules for edge cases. Because the recipe is written down, two reviewers or two runs of a tool should reach the same score from the same evidence.

How do you build a weighted scoring model?

Pick four to seven criteria that you believe predict a good outcome. Write a rubric for each so a 2 and a 4 mean something specific. Assign weights that add up to 1, score a few past cases, and check whether the totals rank them the way the outcomes did. Adjust and repeat.

Do simple scoring models beat expert judgment?

Often, yes. Studies going back to Dawes in 1979 found that simple linear models, even with equal weights, match or beat the experts they are compared with, mainly because they apply the same weights every time. The expert still matters: experts choose the criteria and spot evidence the model cannot see.

How often should a scoring model be updated?

Review it when you have a meaningful batch of new outcomes, or at least once a year. Check whether high scores still go with good results and whether reviewers are scoring the same evidence the same way. If either has slipped, look for score drift before changing the weights.

Sources

  1. Cinelli, Kadziński, Gonzalez and Słowiński (2020), How to support the application of multiple criteria decision analysis? Let us start with a comprehensive taxonomy, Omega (Elsevier)
  2. Yu and Kuncel (2020), Pushing the limits for judgmental consistency: comparing random weighting schemes with expert judgments, Personnel Assessment and Decisions
  3. Neumann, Niessen, Hurks and Meijer (2023), Holistic and mechanical combination in psychological assessment: why algorithms are underutilized and what is needed to increase their use, International Journal of Selection and Assessment (Wiley)