Interviewer Training and Calibration
Training your engineers to run the loop and agree on what they saw. A rubric only means something once the people applying it produce the same score for the same candidate.
- Who
- Led by the founder, hands-on for the duration.
What you're seeing
- Two interviewers gave the same candidate very different scores and both were confident.
- Usually means The rubric is being read differently by each of them, which means it is not doing its job. Until that gap closes, the loop's output depends on who happened to be free.
- One interviewer's recommendations are almost always positive.
- Usually means Scoring drift, and it is invisible without looking at the distribution per interviewer. It is rarely deliberate and it silently sets a different bar for whoever gets scheduled with them.
- Interview notes record conclusions rather than observations.
- Usually means 'Strong engineer' is not evidence and cannot be reconciled with another interviewer's 'not convinced'. Without recorded observations the debrief has nothing to work from except opinion.
- Hiring standards loosened during a busy quarter and nobody decided to loosen them.
- Usually means Standards move quietly under pressure. Without periodic drift review the bar becomes whatever the most urgent role required most recently.
The rubric is not the agreement
A rubric agreed by circulation is agreed on the words, not on what they mean.
The reliable demonstration is simple and slightly brutal: give four interviewers the same anonymised past submission, have them score it independently against the current rubric, and put the scores on a wall. The spread is usually much larger than anyone expected, and the people at either end are both confident and both articulate about why.
That disagreement is the training material. Reconciling it — arguing about what the criterion actually requires, against a real artefact — is what produces convergence, and there is no shortcut to it. A better-written rubric does not remove the need; it just moves where the disagreement surfaces.
Observations, not conclusions
“Strong engineer” cannot be reconciled with “not convinced”. There is nothing in either statement to compare.
“Identified the race condition without prompting, then dismissed the retry question as unimportant” can be. Two interviewers who both wrote something of that shape can find where they diverged, and a third person reading both can form a view. It is also the only version that survives the six-month question of why a candidate was rejected, and the only version that can be turned into useful feedback for them.
Getting interviewers to write this way is most of the mechanical half of the training, and it is resisted mainly by people confident in their own judgement — which is a reasonable position and does not help the person who was not in the room.
Drift is the recurring cost
Standards move quietly, and they move most during the quarter when a role is most urgent.
Nobody decides to lower the bar. What happens is that one borderline candidate is passed because the team is short, then the next borderline candidate is compared to that one, and within two quarters the bar is somewhere nobody chose. It is only visible in the distribution — scores per interviewer over time, pass rates by stage — which is why the quarterly review exists and why it is cheap.
Where it sits
This capability sits under Engineering Hiring, and it is the one that makes the rest of it real: a structured loop that has not been calibrated is structured on paper and unstructured in practice.
It is inseparable from Interview Loop Design — the rubrics being calibrated are that loop’s rubrics, and neither is worth much alone. It also shares its mechanics with Performance Review Design, where the same problem appears internally: two managers, one rating scale, and no shared meaning until they calibrate against real cases.
How the work runs
-
Calibrate on real submissions
Interviewers score anonymised past candidates independently, then reconcile. The disagreements are the training material and they are usually large at first.
-
Train the mechanics
How to ask a follow-up that produces evidence, how to record what was observed rather than what was concluded, and how to run a bounded exercise without rescuing or abandoning a candidate.
-
Cover the specific biases
Similarity bias, the halo from one strong answer, and the effect of interview order. Named concretely against your own scoring history rather than in the abstract.
-
Shadow and reverse-shadow
Every new interviewer observes twice and is observed twice before running a stage alone.
What arrives
- A calibration session with your own anonymised submissions
- An interviewer guide covering mechanics and common failures
- A shadowing path with a clear bar for running a stage alone
- A quarterly scoring-drift review
What it costs your team
A half-day calibration session per group, plus shadowing during live loops.
How we decide
Calibration runs on real anonymised submissions, not on hypotheticals
Costs It requires assembling past candidate material and anonymising it, and the disagreements it surfaces are uncomfortable in the room.
A rubric agreed in the abstract is agreed on the words. Interviewers only converge by scoring the same real submission independently and then reconciling, and the size of the initial disagreement is usually a surprise to everyone. That argument is the training material — it cannot be replaced by a document.
Interviewers record observations, not conclusions
Costs It takes longer to write up and feels pedantic to people who are confident in their judgement.
'Strong engineer' cannot be reconciled with another interviewer's 'not convinced' — there is nothing to compare. 'Identified the race condition without prompting, then dismissed the retry question' can be. It is also what makes a rejection explicable to a candidate and to yourself six months later.
Bias is taught from your own scoring data, not in the abstract
Costs It requires looking at real patterns in real decisions, which can implicate people who are present.
Generic bias training has weak evidence behind it and is mostly experienced as compliance. The same content applied to a visible pattern in your own history — scores by interview order, by interviewer, by similarity to the interviewer's own background — changes behaviour, because it is about something the room can see.
Nobody runs a stage alone until they have shadowed and been shadowed
Costs It costs two extra interviewer-hours per new interviewer and slows down widening the pool.
Interviewing is a skill and it is routinely assumed to be a by-product of being senior. Observing twice and being observed twice is a small cost against a candidate lost to a badly run stage, and it is the mechanism that keeps the standard transferring as the team grows.
Where this has run
Frequently Asked Questions
Sources
- Google re:Work — Structured interviewingrework.withgoogle.com
- Microsoft — Code With Engineering Playbookmicrosoft.github.io
Page reviewed
