What calibration actually tests
Calibration is the meeting where managers defend their proposed ratings to each other. Most people prepare for it as though it were a presentation β assemble the evidence, state the case, hope nobody pushes hard. That is the wrong model. Calibration is closer to a peer review, and the ratings that survive it are the ones where the manager has already found the weakest part of their own argument.
The uncomfortable truth is that a rating you cannot argue against is usually a rating you have not examined. If you walk in able only to say why someone is strong, you will be caught out by the first person who asks what they missed.
Where ratings fall apart
In practice, proposed ratings tend to collapse for a small number of recurring reasons:
- Recency. The last six weeks are vivid and the first six months are not. A strong finish can rescue a weak year, and a bad month can erase a good one, unless you deliberately reconstruct the whole period.
- Effort mistaken for impact. Someone visibly working hard reads as high performance. Someone who removed the need for the work entirely often reads as quiet. The second is usually worth more.
- Evidence that is really an impression. "Great communicator" is a conclusion. "Rewrote the incident postmortem so support could use it, cutting repeat tickets" is evidence. Only the second survives a challenge.
- Comparison to the wrong bar. Rating people against each other rather than against the level definition produces a ranking, not an assessment β and rankings shift depending on who else is in the room.
- The unspoken concern. Nearly every manager has one reservation they have not written down. If you do not raise it, someone else will, and it will land much harder unprepared.
Why stress-testing beats polishing
The demo above does not write your review. It argues against it. You paste anonymized notes and your proposed rating, and it returns the objections a skeptical peer would raise, so you meet them at your desk rather than in the room.
That reframing matters more than the tool. Sometimes the counter-arguments reveal that the rating itself is wrong β that the evidence supports a different call than the one you had already decided on. Finding that out beforehand is the entire point. A calibration meeting is a bad place to change your mind under pressure.
Two limits worth taking seriously
Anonymize before you paste. Not as a formality. Employee names, customer names, compensation figures, employee IDs and internal metrics should not go into a third-party AI tool without checking your organization's policy first. The demo works perfectly well with [PERSON A] and [PERSON B].
This is a drafting aid, not a compliance tool. Employment law varies by jurisdiction, and the consequences of getting a review or a termination wrong are serious and personal for the people involved. Anything consequential should go past your own HR and legal people before you act on it. AI output is also confidently wrong on a regular basis β verify anything factual it asserts about a person's record.
More on how we think about this is on the terms page, and the manager prompt set it comes from is at five free prompts.