Short answer
Calibration aligns how leaders apply the same rating scale and Function expectations before reviews are finalized. A useful session compares evidence, tests edge cases, records decisions, and sends managers back with clear updates for each draft.
Set the purpose before the meeting
Calibration is not a popularity vote, a defense of each manager's team, or an automatic distribution. Its purpose is to make the same rating mean the same thing for comparable evidence.
Name the decision scope in advance: which ratings are in review, who can decide, what must be resolved in the room, and what can return to the manager for more evidence.
Prepare comparable inputs
Use the Function snapshot pinned at cycle kick-off, the shared rating scale, proposed ratings, and short evidence mapped to level expectations. Flag cases where the prose and rating disagree.
HR should check for missing evidence, recency bias, unsupported trait language, and different standards applied to similar roles before the session.
- The same Function version used by the review cycle
- Proposed ratings with evidence from the full period
- Edge cases, disagreements, and missing information
- A decision owner and note taker
Run an evidence-first discussion
Read the relevant expectation when participants disagree. Ask what happened, where the evidence sits, and why it supports this point on the rating scale.
Compare like with like. Do not compare communication style, office visibility, hours, confidence, or a single recent event. Pause a decision when the room lacks enough evidence.
Record decisions and reasons
For each changed or confirmed rating, record the expectation, the evidence considered, the decision, and who updates the review. Keep notes factual and suitable for the review record.
If the Function wording caused the disagreement, log it for the next version. Do not rewrite the standard for one person during an active cycle.
Decision note
The room felt this person was not senior enough.
The evidence did not yet show [next-level scope]. The manager will update the draft to reflect the current-level rating and add [specific development step].
Close the loop after calibration
Managers should update ratings and prose together. A changed rating with unchanged narrative creates a review the employee cannot understand.
Before sign-off, confirm every development note includes a next step. Then review patterns across teams for standards that need clarification in the next Function version or rating scale guidance.