AI candidate scoring is the use of machine learning or language models to rate or rank job applicants against a role's requirements. Because the score can decide who a human ever reviews, laws such as NYC Local Law 144 and the EU AI Act treat these tools as regulated.
How AI candidate scoring works
Many of these tools run inside an applicant tracking system. Most follow the same path. They take in applications, extract structured facts (titles, tenure, skills, written answers), compare those facts to the role, and output a score or a rank. The comparison takes one of two forms.
- Learned from past outcomes. A model is trained on who the company hired, or who performed well, and scores new applicants by resemblance. This copies whatever patterns sit in the history, including unfair ones.
- Rated against stated criteria. A language model reads each application against a written rubric and returns a rating per criterion, with the passages it relied on. This is easier to inspect, but only as good as the rubric. Writing that rubric from what the role actually involves is the job of job fit scoring.
Either way, it is still candidate scoring. What changes is scale and opacity: a tool can score 5,000 applicants in an hour, and a reviewer may not see why applicant 4,812 was ranked last.
Rules that apply to AI candidate scoring
This section summarises two well known rules. It is not legal advice, and obligations depend on where the employer and the candidates are.
New York City. Local Law 144 of 2021 bars employers and employment agencies from using an automated employment decision tool unless it has had a bias audit within one year of use, a summary of the results is public, and candidates or employees have received notice. The city's enforcement agency says notice must be given 10 business days before use. Enforcement began on July 5, 2023.
European Union. Annex III of the EU AI Act lists AI systems used for recruitment or selection, including systems that filter applications and evaluate candidates, as high risk. High-risk systems carry duties on risk management, data quality, documentation, transparency and human oversight. The European Commission's AI Act FAQ says the high-risk rules apply from 2 December 2027, after the Digital Omnibus moved the original date.
U.S. federal anti-discrimination law also applies whether a decision is made by a person or a tool. A selection rate gap between groups can be adverse impact regardless of how the score was produced.
Worked example: checking a screening tool
Northwind Logistics, a fictional employer, uses a tool that advances the top-scoring 25% of warehouse supervisor applicants to a phone screen. Before renewing it, the team checks the last quarter.
| Group | Applicants | Advanced | Selection rate | Impact ratio |
|---|---|---|---|---|
| Group 1 | 600 | 180 | 30% | 1.00 |
| Group 2 | 400 | 84 | 21% | 0.70 |
The impact ratio is each group's selection rate divided by the highest group's rate: 21 / 30 = 0.70. Local Law 144 requires impact ratios to be calculated and published, but it does not set a pass mark. Under the federal four-fifths guideline in the Uniform Guidelines, a ratio below 0.80 is generally treated as evidence of adverse impact, so 0.70 calls for a closer look. The team asks the vendor which inputs drive the gap. It turns out gaps in employment history weigh heavily. They remove that input, rerun the quarter, and add a person to review every applicant scored in the bottom half before rejection.
AI candidate scoring vs human candidate scoring
| AI candidate scoring | Human candidate scoring | |
|---|---|---|
| Speed | Thousands of applicants per hour | A few dozen per reviewer per day |
| Consistency | Same input gives the same output | Varies by reviewer, mood and order |
| Where bias comes from | Training data, proxies, rubric design | Individual judgment and group habits |
| How to explain a score | Only if the tool shows its evidence | Ask the reviewer, if notes were kept |
| Legal treatment | Specific AI and automated decision rules, plus general law | General anti-discrimination law |
Neither is neutral by default. NIST Special Publication 1270 describes bias in AI as coming from systemic, statistical and human sources, which means a tool can inherit bias from the people who built and used it.
Common mistakes with AI candidate scoring
- Training on past hires without checking them. If the past hires were skewed, the model learns the skew.
- Auto-rejecting below a cutoff. Keep a human in the loop for rejections, at least near the line.
- Accepting a score without reasons. If the tool cannot point to the parts of an application that drove the rating, nobody can check it. Prefer tools built for explainable AI.
- Auditing once. Applicant pools and models change. Re-check selection rates each hiring cycle, not only when the law requires.
- Assuming the vendor carries the obligation. Under Local Law 144 the employer or agency using the tool is the one that must not use it without an audit and notice.
How ScoringFactory approaches AI candidate scoring
ScoringFactory scores candidates against a team's own bar, learned from its past hiring decisions, and cites each rating to the line in the candidate's record. It ranks and explains. It never rejects anyone, and the hiring manager makes every decision. See the hiring use case.
Frequently asked questions
Is it legal to use AI to screen candidates?
Generally yes, but rules apply. In New York City, an automated employment decision tool needs a bias audit within the past year, a public summary of results, and notice to candidates. The EU AI Act classes recruitment AI as high risk. Anti-discrimination law applies everywhere. This is not legal advice; check the rules where you hire.
Does AI hiring software need a bias audit?
In New York City, yes, if it is an automated employment decision tool used to substantially assist hiring or promotion decisions. The audit must be done by an independent auditor within one year of use. Elsewhere it may not be required by law, but a regular bias audit is the most direct way to find adverse impact early.
How do you explain an AI candidate score?
Show the criteria the score was measured against, the rating on each, and the specific parts of the application behind each rating. A total with no breakdown cannot be explained or challenged. If a tool only returns a number, ask the vendor for per-criterion reasons before using it on real applicants.
Can AI candidate scoring reduce bias?
It can reduce some kinds, such as one reviewer being harsher after lunch, because the same input always gets the same output. It can also scale bias that sits in training data or in the criteria. Whether it reduces bias overall is an empirical question for each tool, answered by measuring selection rates by group.
Sources
- Automated Employment Decision Tools (AEDT), NYC Department of Consumer and Worker Protection, current text
- Regulation (EU) 2024/1689 (Artificial Intelligence Act), official text, EUR-Lex, Official Journal of the European Union
- 29 CFR 1607.4: Information on impact (four-fifths rule), eCFR, U.S. Government, current text
- Towards a Standard for Identifying and Managing Bias in Artificial Intelligence (SP 1270), NIST, 2022
- Navigating the AI Act (FAQ), European Commission, 2026