Normalization is converting inputs that arrive on different scales or in different formats, such as revenue in dollars, a 1 to 5 interview rating, years of experience, or three spellings of one company name, onto a common scale or form so they can be compared and combined fairly.
Two kinds of normalization in scoring
The word covers two jobs that both happen before a score is calculated.
Scale normalization is the statistical sense. Singh and Singh compare fourteen normalization methods, z-scores among them, in Investigating the impact of data normalization on classification performance (Applied Soft Computing, 2020), and find that no single method wins everywhere. Annual revenue of 4 million dollars and a founder rating of 4 out of 5 cannot be added. Each has to be rescaled first, usually to 0 to 1 or to standard deviations from the mean.
Record normalization is the data sense: making the same thing look the same everywhere. "Ledgerly Inc.", "Ledgerly" and "ledgerly.io" are one company. "Sr. Software Eng." and "Senior Software Engineer" are one title. "NYC" and "New York, NY" are one city. Matching records for the same company is entity resolution, which Christophides and colleagues define in An Overview of End-to-End Entity Resolution for Big Data (ACM Computing Surveys, 2020) as identifying different descriptions that refer to the same real-world entity. Cleaning titles and cities to one standard form is the same idea applied to values, and it usually runs alongside data enrichment.
Get the second wrong and you score one company twice. Get the first wrong and one criterion silently outweighs the rest.
Common normalization formulas
| Method | Formula | Output | Use when |
|---|---|---|---|
| Min-max scaling | (x - min) / (max - min) | 0 to 1 | Bounded inputs without big outliers |
| Z-score | (x - mean) / standard deviation | Centred on 0 | Roughly bell-shaped inputs; comparing to the pool |
| Percentile rank | Share of the pool scoring below x | 0 to 100 | Skewed inputs; easy to explain |
| Log, then scale | log(x), then min-max or z-score | Varies | Revenue, funding, follower counts, anything spanning orders of magnitude |
The z-score is also called the standard score. A z of 1.5 means one and a half standard deviations above the pool average. Because it depends on the pool, a candidate's z-score changes when the shortlist changes, so recompute it after every cut.
Normalization changes the scale, not the meaning. A min-max score of 0.8 is not an 80 percent chance of anything. Turning scores into honest probabilities is a separate step, model calibration.
Worked example: three companies, two inputs
A fictional PE firm screening software companies has revenue and monthly growth for three targets.
| Company | Revenue ($M) | Revenue, min-max | Revenue, z-score | Growth (%/month) | Growth, min-max |
|---|---|---|---|---|---|
| Ledgerly | 1.1 | 0.19 | -0.47 | 9 | 0.50 |
| Fernhill Robotics | 4.0 | 1.00 | 1.39 | 3 | 0.00 |
| Quarry Labs | 0.4 | 0.00 | -0.92 | 15 | 1.00 |
Added raw, revenue and growth would mean Quarry's 15 percent growth counts as 15 "points" and Fernhill's revenue as 4. After min-max scaling, both inputs run 0 to 1 and can be weighted on purpose. Notice the catch: Fernhill's 4.0 is far above the others, so Ledgerly's 1.1 lands at only 0.19. With a pool of hundreds and one very large company, min-max would squash everyone else near zero. A log transform or percentile rank handles that better.
Why normalization matters for investors and hiring teams
Every scoring model that combines more than one kind of input depends on it. A fund comparing founders across countries needs funding amounts in one currency. A recruiter comparing candidates needs job titles mapped to a common level, since "Lead" at a 20-person startup and "Lead" at a bank are different jobs. A PE firm using benchmarking needs margins defined the same way for every company before it can say who is above the median.
Common normalization mistakes
- Normalizing against the wrong pool. A z-score against this week's 30 inbound deals says nothing about the market. State the reference pool.
- Min-max with outliers. One extreme value compresses everyone else.
- Treating missing as zero. A company with no published revenue is unknown, not a company with no revenue.
- Collapsing different things. Merging "Head of Sales" and "Sales Lead" may be right at one company and wrong at another. Normalize titles with the company's size in view.
- Losing the original. Keep the raw signal alongside the normalized value, so anyone reading a score can see what it was built from.
How ScoringFactory approaches normalization
ScoringFactory resolves people and companies to one record, puts inputs on a common footing before scoring, and keeps the original evidence attached, so a score can always be traced back to the line it came from. The team sees the ranked list and the receipts, and decides. See the flow.
Frequently asked questions
What is data normalization in scoring?
It is the step that makes different inputs comparable before they are combined. That includes rescaling numbers, such as converting revenue and ratings to a 0 to 1 range, and standardizing records, such as matching three spellings of one company name. Without it, the input with the biggest numbers dominates the score.
Should you use z-scores or min-max scaling?
Use min-max when inputs have natural bounds and no large outliers, such as a 1 to 5 rating. Use z-scores when you want to know how far something sits from the pool average. For skewed inputs like revenue or funding, take the log first, or use percentile ranks, which outliers cannot distort.
Why do scores need to be normalized before weighting?
Weights only work if every input is on the same scale. If revenue is in millions and ratings run 1 to 5, a 20 percent weight on revenue still swamps an 80 percent weight on ratings. Normalizing first means the weight you chose is the influence the criterion actually has.
Is normalization the same as data enrichment?
No. Data enrichment adds information to a record, such as headcount or funding history. Normalization makes existing information consistent and comparable. They run together in most pipelines: enrich a record, then normalize what was added so it matches every other record.
Does normalizing inputs make a score fair?
Not by itself. Normalization makes inputs comparable, but if an input carries bias, such as a pedigree signal that tracks who could afford a certain school, rescaling it keeps the bias intact. Fairness depends on which inputs you score and how you check the results across groups, which is the subject of bias mitigation.
Sources
- Singh and Singh (2020), Investigating the impact of data normalization on classification performance, Applied Soft Computing (Elsevier)
- Christophides, Efthymiou, Palpanas, Papadakis and Stefanidis (2020), An Overview of End-to-End Entity Resolution for Big Data, ACM Computing Surveys