Annual performance reviews promise measurement and deliver horoscopes: confident, vague verdicts that say more about the person reading the chart than about you. The research behind that claim is brutal, with over 60% of a rating reflecting the rater rather than the ratee, and a third of feedback interventions actively making performance worse.
Astrology takes a real input, the position of planets, and generates confident statements about your personality that are vague enough to feel true and specific enough to feel personal. Annual performance reviews run the same trick on a year of your work: “meets expectations,” “should demonstrate more leadership,” “a 3.7 out of 5,” Mercury is in retrograde and your visibility with senior stakeholders could improve.
What annual performance reviews really measure
Ask what a review measures and the official answer is twelve months of performance, but in practice it’s 3 or 4 weeks of it, plus office politics filtered through one person’s mood.
Recency does the first part of the damage; a human being can’t hold a year of someone’s work in their head, so the review quietly becomes an assessment of the last quarter, and the colleague who shipped their big win in November will outscore the one who carried the team through a miserable spring. Everyone who has sat through review season knows the choreography this produces: the strategic October project, the sudden burst of December visibility, the achievements doc that gets updated in a panic the night before, and none of that counts as performance.
Then there’s the politics; ratings pass through calibration meetings where managers negotiate on behalf of their people, and the outcome tracks who argues well, who has capital to spend, and whose team the budget favours this cycle. I wrote about the quiet-quitting panic and who never had to defend themselves in that framing; reviews run the same trick, holding the worker’s output up to the light while the process around it goes unexamined.
Your 3.7 out of 5 is numerology
Here’s the finding that should have ended rating scales 25 years ago: Scullen, Mount, and Goff decomposed thousands of manager ratings in the Journal of Applied Psychology and found the rater’s own idiosyncratic tendencies accounted for 62% and 53% of the variance across two large datasets, more than the actual performance of the person being rated. Three separate studies put the rater’s share at 71%, 58%, and 55%, and no other factor cracked 20%. Marcus Buckingham translates it more usefully:
All the data show that when I rate you, over 60% of your rating is about me, and not you.
That sentence, from Buckingham’s own breakdown of why ratings are bad data, is the whole scandal in one line. The score on your review mostly describes your manager: how harsh a grader they are, what they personally rate as “leadership,” which of your traits remind them of themselves. Companies then attach decimal points to this and call it measurement. A 3.7 versus a 3.4 presents itself as precision, and it’s the precision of a birth chart, a confident number generated by a process that can’t support the confidence.
I’ve seen the same disease outside HR. It was in my own dashboards when I was tracking 17+ KPIs that couldn’t tell me if anything was working, and it’s why I show my methodology whenever I claim a pipeline number: the existence of a figure does all the persuading, whether or not the figure measures anything.
Deloitte looked at this evidence about its own system, which was consuming around two million hours a year, and tore the whole apparatus down in 2015, and when a Big Four consulting firm concludes its measurement ritual measures nothing, the rest of us can stop defending ours.
The feedback evidence nobody in HR quotes
The deeper assumption underneath the ritual is that feedback itself reliably helps, and the largest study ever run on that assumption says otherwise: Kluger and DeNisi’s meta-analysis in Psychological Bulletin covered 607 effect sizes and 23,663 observations, and found that over a third of feedback interventions actively decreased performance. Roughly a third helped, a third did nothing, and a third made people worse at their jobs. We know that nobody would prescribe a medicine with that profile, and yet HR prescribes it to the entire company every December.
Feedback works when it directs attention at the task: this section buried the answer, this campaign targeted the wrong segment, here’s the specific behaviour and the specific fix. “You’re a 3” says nothing about any task and everything about you as a person, so the recipient spends their energy defending their identity instead of improving their work.
A third of feedback makes performance worse, and the format most likely to backfire is exactly the one the annual review uses: self-focused, high-stakes, and tied to pay.
The feedback that works is boring
Feedback that improves performance is task-focused, specific, and cheap to receive: close to the work in time, aimed at the work rather than the person, and low-stakes enough that nobody’s cortisol spikes when the calendar invite lands.
In practice that means the 15-minute conversation this week about the draft, rather than the verdict in March about the year, it means “the opening buries your argument, move the claim up” instead of “your written communication could be stronger.” When I was managing freelancers across four countries, this was the entire system: clear standards agreed upfront, then frequent, small, task-level corrections against those standards, with no scores anywhere.
However, the same research tradition warns that more feedback is a lazy prescription too, since most feedback effects are small and constant commentary can interfere with people who already know how they’re doing. Frequency was never the magic ingredient, specificity and task-focus are, and a company that swaps the annual ritual for weekly rating-flavoured micro-verdicts has just made the astrology more frequent.
What to salvage from the ritual
Torching the ceremony doesn’t mean torching everything inside it, and this is where I’ll defend the managers stuck running a process they don’t believe in. A few pieces of the annual review are genuinely worth keeping, once you stop pretending they’re measurement.
Once or twice a year, a protected hour on where someone’s career is going, what they want more and less of, and what the next role looks like, has real value that weekly check-ins never replace. Keep the conversation and delete the scorecard, because that’s what turns a career discussion into a verdict-defence session.
Keep the written record too: documented examples of what someone shipped and its results, which protects both sides far better than a number does, and which is precisely the kind of ownership evidence I’ve argued people should be keeping anyway. Then there’s the compensation decision itself, which genuinely has to happen on some cadence; make the pay call openly as a judgment, informed by documented work, without laundering it through a pseudo-scientific rating first.
What dies in this version is only the astrology: the scale, the decimal, the ranking, the annual pretence that one person’s memory of your year, filtered through their own tendencies, constitutes data. Companies find that piece surprisingly hard to give up, because the number feels like control, but a confident wrong number is worse than honest judgment, in reviews as everywhere else, and the organisations that admit it get better performance conversations at a fraction of the cost.
Frequently asked questions
Annual performance reviews fail on two documented fronts. Ratings mostly measure the rater rather than the employee, with studies attributing over half the variance in scores to the manager’s own idiosyncratic tendencies, and the feedback format itself is the kind most likely to backfire, since high-stakes, self-focused feedback made performance worse in over a third of cases in the largest meta-analysis ever conducted. Add recency bias and calibration politics, and the annual score reflects the last quarter and the loudest advocate more than the year’s work.
Annual performance reviews are best replaced by continuous, specific, task-focused feedback given close to the work, in low-stakes settings with no scores attached. The evidence shows feedback improves performance when it directs attention at the task and backfires when it directs attention at the self, so short, frequent, concrete conversations about actual work beat annual verdicts about the person. Keep a separate, protected career conversation once or twice a year and make compensation decisions openly on documented work rather than through ratings.
Yes, decades of it. Scullen, Mount, and Goff found the rater’s idiosyncratic tendencies accounted for 62% and 53% of rating variance across two large datasets, and three separate studies put the rater’s share at 71%, 58%, and 55%, which is why researchers say a rating reveals more about the rater than the ratee. Kluger and DeNisi’s meta-analysis of 607 effect sizes found over a third of feedback interventions decreased performance. Deloitte cited exactly this evidence when it dismantled its own ratings system in 2015.
No, and the research doesn’t support that reading either. Feedback helps on average; the problem is the format, since the gains come from specific, task-level input, while scores and rankings do the damage. Managers should give more of the boring kind, concrete and close to the work, and stop delivering the theatrical version once a year with money riding on it. Constant commentary is its own failure, so aim for useful frequency rather than maximum frequency.
