Glossary

Evaluation rubric

By ScoringFactoryUpdated First published 8 June 20264 min read
Definition

An evaluation rubric is a rubric built for one specific decision, such as a seed investment or a senior engineering hire, that lists the criteria, how much each counts, any hard requirements, and anchored descriptions of weak, average and strong evidence for each criterion.

What an evaluation rubric includes

A general rubric is a scale for one criterion. An evaluation rubric is the complete instrument for one decision. It should fit on one page and contain six things:

  1. Scope. The decision it covers, for example "Seed, B2B software, first meeting" or "Staff backend engineer, final round".
  2. Hard requirements. Pass or fail checks applied before scoring, such as work authorization or fund geography.
  3. Criteria. Four to seven, each measuring one thing.
  4. Weights. Adding up to 100 percent.
  5. Anchors. A concrete description of weak, average and strong evidence for every criterion. Anchors tied to observed behavior are the idea behind behaviorally anchored rating scales, introduced by Smith and Kendall in 1963. A 2020 study that built such scales for rating teams found acceptable agreement between independent trained raters (Georganta and Brodbeck, 2020).
  6. Decision rule. What total, or what pattern of scores, leads to advance, discuss or decline.

The decision rule is the part most teams leave out. Without it, two people can agree on every score and still disagree about what to do.

Why one rubric per decision type

A fund that writes one rubric for every stage will judge a pre-seed founder on revenue they could not have yet. A company that uses one interview rubric for every role will grade a designer on system design. The criteria that predict success change with the decision, so the rubric has to change too.

For hiring, the written rubric is how a team turns its hiring bar into something interviewers can apply the same way. OPM describes structured interviews as asking every candidate the same questions and judging every answer against the same rating scale and standards (OPM, Structured Interviews). For investing, it is where a fund's thesis fit criteria become explicit instead of living in one partner's head.

Example: an evaluation rubric for a staff engineer

Tidewater Health, a fictional 60-person company, is hiring a staff backend engineer. Hard requirement: has run a production system with on-call responsibility. Then:

Criterion (weight)Weak (1)Average (3)Strong (5)
System design (30%)Describes one design, cannot explain trade-offsExplains trade-offs of a design they builtHas redesigned a system under load and can show what changed
Code quality (25%)Work sample runs but is hard to followClear, tested, conventionalClear and tested, and the reviewer learned something reading it
Technical leadership (25%)No examples beyond own tasksHas mentored one or two engineersHas set a technical direction other teams adopted
Domain context (20%)No exposure to regulated dataHas worked near regulated dataHas built systems handling health or financial records

Decision rule: advance at a weighted score of 3.8 or above with no criterion below 2. Candidate B scores 4, 3, 5, 2: weighted total 3.6, so the panel discusses rather than advances, and the discussion is about domain context, not about whether they "felt senior".

Evaluation rubrics for hiring vs investing

Hiring rubricInvestment rubric
Unit judgedOne candidate for one roleOne company and its founders
EvidenceWork samples, structured interview answers, referencesMetrics, founder history, market data, references
Typical hard requirementsWork authorization, location, licenceStage, sector, geography, check size
Main legal exposureEmployment discrimination lawMostly internal policy
Filled-in outputInterview scorecardDeal scorecard, then the memo

Common mistakes

  • Copying a generic template. A downloaded rubric reflects someone else's bar, not yours.
  • No hard requirements stage. Scoring candidates who fail a must-have wastes panel time and muddies averages.
  • Anchors only at 1 and 5. Most evidence lands in the middle, so write the middle.
  • No agreement check. If two interviewers often score the same answer two points apart, the anchors need work. A 2020 review by Panadero and Jonsson notes that when scores must agree across raters, rubrics with fewer performance levels are often the better choice. Measure it with inter-rater reliability.
  • Letting the rubric go stale. When the role or the fund strategy changes, rewrite it. Benchmarking new scores against past ones shows when a 4 today no longer means what a 4 meant last year.

How ScoringFactory approaches evaluation rubrics

Writing a rubric by hand captures what a team says it values. ScoringFactory starts from what the team actually chose: it learns the bar from past yes and no decisions for that role or fund, applies it to every new candidate or company, and cites each score to the line in the record behind it. See the flow.

Frequently asked questions

How do you create an evaluation rubric?

Define the decision it covers, list any pass or fail requirements, pick four to seven criteria, weight them, and write anchors for weak, average and strong evidence on each. Add a decision rule. Then test it on five to ten past cases with several reviewers and rewrite any row where scores spread widely.

What should an evaluation rubric include?

Scope, hard requirements, criteria, weights, anchored score levels and a decision rule. The anchors should describe observable evidence, such as a metric, an event or a specific answer, rather than adjectives like strong or excellent. Leaving out the decision rule is the most common gap, and it is where panels end up arguing.

How many criteria should a rubric have?

Four to seven works for most hiring and investment decisions. Fewer and the rubric misses things that matter. More and each criterion carries so little weight that it barely moves the total, while reviewers tire and start scoring rows by habit. If two criteria always move together, merge them.

Is an evaluation rubric the same as a scorecard?

No. The evaluation rubric is the blank standard for a decision type. A scorecard is that rubric filled in for one candidate or company, with scores and evidence. You write the rubric once per role or strategy and produce a new scorecard for every person or company you assess.

Sources

  1. Structured Interviews, U.S. Office of Personnel Management, current text
  2. Capturing the four-phase team adaptation process with behaviorally anchored rating scales (Georganta and Brodbeck, 2020), European Journal of Psychological Assessment
  3. A critical review of the arguments against the use of rubrics (Panadero and Jonsson, 2020), Educational Research Review