The Product Manager’s Atlas
Assessment & Coaching

Common Assessment Pitfalls

4 min readΒ·749 words
assessmentbiaspitfallscalibrationgovernance

Every assessment instrument in this domain (the matrix, the self-assessment, the score) is only as good as the human running it. And humans are biased in predictable, well-studied ways. This note is the field guide to the traps I have fallen into and watched others fall into, with the countermeasure for each.

#My point of view: name the bias before you score

You don't beat these by being smart or well-intentioned. Every bias here operates on people who believe they are being fair. You beat them with process: pre-committed criteria, first-hand evidence, and calibration. Goodwill is not a control. Structure is.

#The five that wreck PM assessment

β–²1. Recency bias

You weight the last month and forget the other eleven. The PM who shipped a win in week 50 looks great, and the one who carried Q1 and coasted lately looks weak. Countermeasure: keep a running evidence log all year, and at review time deliberately pull examples from every quarter (Running a Performance Review). If all your examples are recent, you haven't done the work.

β–²2. Halo / horns effect

One vivid trait colors everything. The charismatic presenter gets rated up on data they are actually weak at (halo), and the quiet PM who shipped flawlessly gets rated down across the board because they don't perform confidence (horns). Countermeasure: score each competency independently, against its own behaviors, with its own evidence. Force yourself to justify why a strong communicator is also strong at quality. Usually you can't, and the halo dissolves.

β–²3. Grading to the org's median

The pull to cluster everyone around "On Track" or "meets expectations" because the extremes feel risky. This destroys the instrument's whole purpose. If nobody is ever Needs Focus and nobody is ever Outperform, your scores carry zero information and your promotion signal is gone. Countermeasure: demand evidence for the middle too. "Why is this a solid On Track and not a Needs Focus?" is the question that breaks the herd toward the center. A flat distribution is a smell, not a sign of fairness.

β–²4. Confusing activity with impact

The deadliest one for PMs specifically. You reward the person who shipped the most features, ran the most meetings, and wrote the most tickets, while overlooking the one who shipped less but moved the number that mattered. Countermeasure: score against outcomes, not outputs. The question is never "how busy were they." It is "what changed in the business because they were here." A feature factory looks productive on every activity metric and accomplishes nothing.

β–²5. The self-assessment distortions

They cut both ways. Impostors under-rate, discounting real wins as luck. Overconfident PMs inflate the strategic competencies, the ones with the least objective evidence and the most prestige. Countermeasure: anchor to evidence, and read the diff between self-score and manager-score as the signal, not either number alone (The Competency Self-Assessment).

#A few more worth naming

β„ΉQuieter traps
  • Similarity bias: rating people who work like you higher. This one is lethal for building a diverse team, because the spiky PM who covers your valley is exactly the one you will under-rate.
  • Idiosyncratic rater effect: your ratings say as much about you as about them. It is the single strongest argument for calibration.
  • False precision: treating a summed score of 64 versus 66 as a real difference. The number is coarser than it looks (The Proficiency Scale).
  • Potential-vs-performance conflation: promoting on "they could be great" instead of "they are already operating at the next level." See Calibration and Promotion.
✦The meta-countermeasure

Every trap above is neutralized by the same three habits: first-hand evidence (no rumor), independent per-competency scoring (no halo), and calibration (no private scale). If you only remember one thing, make it this: substantiate every score with a specific example, or don't give it. That single discipline kills most of these on contact.

#Continue Reading