Inter-rater reliability is the degree to which different reviewers, scoring the same evidence independently, give it the same score. It is usually measured with Cohen's kappa for two raters and categories, Fleiss' kappa for more raters, or the intraclass correlation for numeric scores.
How inter-rater reliability is measured
The simplest measure is percent agreement: how often two reviewers gave the same answer. It overstates agreement, because two people who both reject most candidates will agree often by chance alone. The standard statistics correct for that.
Cohen's kappa, for two raters assigning categories such as advance or reject:
kappa = (po - pe) / (1 - pe)
where po is the observed agreement and pe is the agreement expected by chance, given how often each rater uses each category (Jacob Cohen proposed it in 1960; the formula is set out in de Raadt and colleagues, A Comparison of Reliability Coefficients for Ordinal Rating Scales, Journal of Classification, 2021). Kappa is 1 for perfect agreement and 0 for chance-level agreement.
Fleiss' kappa extends the idea to three or more raters. The intraclass correlation (ICC) is used when scores are numbers, such as 1 to 5 ratings, and asks how much of the variation in scores comes from real differences between candidates rather than differences between raters (see Andrade, A primer on the intraclass correlation coefficient, Indian Journal of Psychiatry, 2026).
Worked example: two interviewers, 20 candidates
Two interviewers at a fictional company each make an advance or reject call on the same 20 candidates.
| Interviewer B: advance | Interviewer B: reject | Total | |
|---|---|---|---|
| Interviewer A: advance | 6 | 3 | 9 |
| Interviewer A: reject | 2 | 9 | 11 |
| Total | 8 | 12 | 20 |
- Observed agreement: (6 + 9) / 20 = 0.75.
- Chance agreement: A advances 45 percent of the time, B 40 percent. Both advance by chance 0.45 x 0.40 = 0.18; both reject by chance 0.55 x 0.60 = 0.33. Total 0.51.
- Kappa: (0.75 - 0.51) / (1 - 0.51) = 0.49.
Seventy-five percent agreement sounds solid. A kappa of 0.49 says that once chance is removed, they agree on about half of what they could. On the widely quoted Landis and Koch scale that is "moderate". The five split decisions are where the rubric needs work. The usual fix is an evaluation rubric with a written example for each rating, so both interviewers mean the same thing by a 3.
What counts as good agreement
There is no universal cut-off, and the published guidelines were proposed by their authors rather than derived from data. Two are widely used:
| Statistic | Guideline | Bands |
|---|---|---|
| Kappa | Landis and Koch (1977), as tabulated by Silveira and Siqueira (2023) | 0.41 to 0.60 moderate, 0.61 to 0.80 substantial, 0.81 to 1 almost perfect |
| ICC | Koo and Li (2016), as summarized by Andrade (2026) | Below 0.50 poor, 0.50 to 0.75 moderate, 0.75 to 0.90 good, above 0.90 excellent |
For hiring and investment decisions, kappa above roughly 0.6 or ICC above 0.75 is a reasonable working target. Below that, the outcome depends noticeably on who did the review.
Why inter-rater reliability matters
A score that two qualified reviewers would not reproduce is not measuring the candidate or the company. It is measuring the reviewer. Low reliability also caps how predictive any scoring process can be: if scores are partly random, they cannot track outcomes closely.
This is one reason structured interviews work better than open conversations: same questions, same scale, so agreement rises. For a venture fund, it is a direct test of whether partners share one bar. For qualitative scoring of things like founder insight or communication, it is the main evidence that the scores mean anything.
Common mistakes
- Reporting percent agreement alone. It hides chance agreement, especially when most decisions are rejects.
- Measuring after discussion. Agreement must be measured on scores given independently.
- Using kappa on 1 to 5 scores. Plain kappa treats a 4 versus 5 the same as 1 versus 5. Use weighted kappa or ICC.
- Measuring once. Agreement erodes as reviewers join and standards shift. Track it over time alongside score drift.
- Measuring without acting. A low number should lead to a calibration session and a rubric rewrite.
How ScoringFactory approaches inter-rater reliability
ScoringFactory applies the same learned bar to every profile, so the same evidence gets the same score on every run. When reviewers disagree with a score, each one is cited to the record behind it, which turns the disagreement into a question about specific evidence. Your team makes the decision. See the flow.
Frequently asked questions
What is inter-rater reliability?
It is how consistently different reviewers score the same thing when working independently. High inter-rater reliability means a candidate or company would get roughly the same score whoever reviewed it. Low reliability means the reviewer's identity is shaping the outcome, which makes the scores hard to trust or compare.
How do you measure agreement between interviewers?
Have them score the same candidates independently, then calculate a chance-corrected statistic. For two interviewers making yes or no calls, use Cohen's kappa. For three or more, use Fleiss' kappa. For numeric ratings such as 1 to 5, use the intraclass correlation or a weighted kappa, which gives partial credit for near misses.
What is a good kappa or ICC value?
Common guidelines call kappa between 0.61 and 0.80 substantial agreement (Landis and Koch) and ICC between 0.75 and 0.90 good (Koo and Li). These bands are conventions, not laws. For consequential decisions like hiring, aim for kappa above about 0.6 and treat anything below 0.4 as a sign the rubric needs rewriting.
How do you improve inter-rater reliability?
Write anchored rubric levels that describe observable evidence, ask every candidate or founder the same core questions, have reviewers score before any discussion, and hold short calibration sessions on the cases with the widest spread. Then measure again to see whether agreement actually rose.
Is inter-rater reliability the same as calibration?
No. Inter-rater reliability is a measurement: how much reviewers agree beyond chance. Calibration is the practice of getting them to agree, through shared cases, discussion and clearer anchors. You measure reliability before and after a calibration session to find out whether the session worked, and to decide when the next one is due.
Sources
- de Raadt, Warrens, Bosker and Kiers (2021), A Comparison of Reliability Coefficients for Ordinal Rating Scales, Journal of Classification (Springer)
- Silveira and Siqueira (2023), Better to be in agreement than in bad company: a critical analysis of many kappa-like tests, Behavior Research Methods (Springer)
- Andrade (2026), A primer on the intraclass correlation coefficient as a measure of reliability in medical research, Indian Journal of Psychiatry