Engineering Performance Review Design
A review cycle that produces decisions people can act on, anchored to the ladder and calibrated across managers so a rating means the same thing in two teams.
- Who
- Led by the founder, hands-on for the duration.
What you're seeing
- A rating means something different depending on which manager gave it.
- Usually means There is no calibration, so each manager is applying their own scale. Ratings are then incomparable across teams, which makes promotion and compensation decisions built on them arbitrary.
- Somebody heard something in a review that they should have heard in March.
- Usually means Feedback is being saved for the cycle rather than given continuously. Anything a person first learns in a review arrived too late for them to have done anything about it.
- Reviews assess effort and attitude.
- Usually means There is no levelling framework underneath, so there is nothing objective to assess against. What remains is likeability and visible busyness, and both correlate poorly with contribution.
- The development conversation and the pay conversation happen in the same meeting.
- Usually means Only one of them is being heard. When compensation is in the room it occupies the whole room, and the development half is received as preamble.
Anchored, or it measures likeability
A review with no levelling framework underneath assesses effort and attitude, because those are the only things left to assess.
That is not a lighter version of a performance process. It is a different one, with worse properties: it rewards visibility over contribution, it cannot be explained to someone who disagrees with the outcome, and it establishes precedents that are difficult to unwind when a framework does arrive.
So the ladder comes first. Ratings are assessed against the criteria of the person’s level — scope, autonomy, influence — which is what makes a rating a statement about work rather than about the working relationship.
Calibration is the mechanism
Two managers with the same rating scale will use it differently, and both will be confident.
Calibration is a session where managers present proposed ratings to each other with evidence, before anything is communicated to anyone. It does two things. It makes ratings comparable across teams, which is a precondition for using them in promotion and compensation. And it surfaces the manager who has been consistently generous or consistently harsh — which nothing else surfaces, and which they usually do not know.
It is uncomfortable, it takes half a day, and skipping it means the cycle produced a set of numbers that cannot be added together.
Separate the two conversations
When a compensation number is in the room, it is the only thing anyone hears.
Development feedback delivered in that meeting is received as preamble to the number, however carefully it is framed, and the most valuable part of the whole cycle is lost for the sake of holding one meeting instead of two. Separating them in time costs a scheduling inconvenience and recovers most of what the process was for.
No quota
Forced distribution makes managers argue about allocation rather than about performance, and it penalises an engineer for the strength of the team around them.
Where the budget genuinely constrains what can be awarded, that is a real constraint and it should be stated as one — a compensation conversation held openly, rather than a quota disguised as an assessment. Calibration already delivers the consistency that forced distribution is usually reached for.
Where it sits
This capability sits under Engineering Hiring, which carries the whole people system — the loop, the levels, the first ninety days and this.
It is downstream of Career Ladders and cannot run without it: ratings need criteria to be assessed against, and without them the cycle assesses something else. It shares its mechanics with Interviewer Calibration — the same problem of two people, one scale and no shared meaning, solved the same way, by reconciling real cases before any decision is communicated.
How the work runs
-
Anchor to the ladder
Ratings assessed against level criteria. A review with no levelling framework underneath assesses effort and likeability.
-
Set the cadence
Twice a year for formal review, with feedback continuous. Annual cycles mean feedback arrives up to eleven months late.
-
Calibrate across managers
Managers presenting proposed ratings to each other with evidence, before anything is communicated. Without it a rating means whatever that manager means by it.
-
Separate feedback from compensation
Held apart in time. In one conversation, nobody hears the development half.
What arrives
- A review cycle anchored to level criteria
- A calibration session format with evidence requirements
- A manager guide covering the difficult cases
- A separation of development and compensation conversations
What it costs your team
A design workshop, then facilitation of the first calibration round.
How we decide
Ratings are anchored to the ladder, or the cycle does not run
Costs It makes performance review dependent on levelling work that may not be finished.
A review with no levelling framework underneath assesses effort and likeability, because those are the only things left to assess. That is not a milder version of the same process — it is a different process with worse properties, and running it establishes precedents that are hard to unwind later.
Managers calibrate against each other before anything is communicated
Costs It requires a room of managers presenting evidence about their own reports, which is uncomfortable and takes half a day.
Without it, a rating means whatever that manager means by it, and the organisation has no comparable signal for promotion or compensation. Calibration is also where a manager who has been consistently generous or consistently harsh discovers it, which is information nothing else surfaces.
No forced distribution
Costs It removes the mechanism that guarantees the ratings sum to a budget.
A quota makes managers argue about allocation rather than about performance, and it punishes people for being on a strong team — which is precisely backwards as an incentive. Calibration achieves consistency without the quota. Where the budget is genuinely the constraint, that is a compensation conversation and it should be held as one.
Feedback and compensation are separated in time
Costs It doubles the number of conversations and extends the cycle.
When a number is in the room, it is the only thing anyone hears. Development feedback delivered alongside a compensation decision is not received, however well it is delivered — which means the most valuable half of the process is wasted for the sake of a scheduling convenience.
Frequently Asked Questions
Sources
- Google re:Work — Structured interviewingrework.withgoogle.com
- Microsoft — Code With Engineering Playbookmicrosoft.github.io
- Team Topologiesteamtopologies.com
Page reviewed