Most calibration meetings don't fail because people mean badly. They fail on logistics. Managers arrive without evidence, the first three employees eat half the hour, and the people who most needed a second look get rushed through at the end.

This guide is the facilitator's playbook for running a calibration meeting well. It covers who to invite, what managers must send before the session, a 90-minute agenda, the questions that keep the discussion on evidence, how to settle disagreements, and what happens after the meeting ends.

If you're still deciding whether to calibrate at all, start with our explainer on what performance review calibration is, and its benefits and drawbacks. This article assumes you've decided to run one.

Key takeaways

A calibration meeting tests proposed performance ratings against a shared standard before they become final.

Most of the work happens beforehand: written rating definitions, a one-page pre-read per employee, and a data pack from HR.

Spend the meeting on the highest, lowest, and borderline ratings. Undisputed middle ratings rarely need airtime.

Record every rating change with its reason, and have the manager, not "the committee," deliver the result.

What Is a Calibration Meeting?

A calibration meeting is a structured session where managers present proposed performance ratings for their teams, test them against shared rating definitions, and adjust any rating the evidence doesn't support. HR usually facilitates. The goal is simple: two employees who do the same quality of work should get the same rating, whichever manager they report to.

Calibration meetings sit near the end of a review cycle. Managers draft their reviews first, then calibrate, and only then share ratings with employees or feed them into pay and promotion decisions. You'll also hear them called calibration sessions or performance calibration meetings. The terms mean the same thing.

The stakes are trust. In Gallup's survey of 18,665 U.S. employees, only 22% strongly agreed that their performance review process is fair and transparent. Calibration won't fix that on its own, but it removes one of the most visible causes: managers grading on different scales.

A performance calibration meeting looks back at the review period. Talent calibration looks forward at potential and readiness for future roles, and it usually runs after performance ratings are settled. Our talent calibration guide covers that process separately.

Who Should Attend a Calibration Meeting?

Invite the managers whose teams do comparable work, a neutral facilitator, and a note-taker. Add one senior leader only if the group needs a tiebreaker. We recommend capping each session at about six to eight presenting managers, so every manager gets real time to present and to question others.

Role What they do in the room What they owe before the meeting
Facilitator (HR business partner or People Partner, ideally not rating anyone present) Runs the agenda, holds the group to the rating definitions, challenges vague evidence, and names patterns The data pack and the list of flagged cases
Presenting managers Present each direct report's proposed rating with evidence, and question other managers' ratings A one-page pre-read for every direct report
Senior leader (optional) Joins as a peer and applies the agreed tiebreak rule when the group deadlocks A read of the pre-reads for contested cases
Note-taker Records every rating change and the reason for it A tracking sheet with every employee's proposed rating

How to Group Sessions

Group managers by job family and level, so the room compares similar work. A sales director and an engineering manager can't judge each other's people with much confidence.

Larger organizations usually calibrate in tiers:

  1. First-line managers calibrate their individual contributors.
  2. Managers of managers calibrate those first-line managers and spot-check the first round.
  3. Executives review the overall distribution and sign off.

How Many Employees One Session Can Cover

As a rule of thumb, budget three to five minutes for each employee you expect to discuss and ten or more for a contested case. That gives a 90-minute session room for roughly 12 to 15 discussions. Because most middle ratings won't need discussion, that is usually enough for a group of 40 to 60 employees, provided managers send pre-reads in advance.

How to Prepare for a Calibration Meeting

Preparation decides whether a calibration meeting works. Three things must exist before anyone joins: written rating definitions, a pre-read for every employee, and a data pack from HR. Without them, the meeting turns into managers describing people from memory.

1. Two to Three Weeks Before: Lock the Rating Definitions

Each rating level needs a description managers can point to, not interpret. Write each level in terms of observable outcomes and behaviors.

For example, a definition of "exceeds expectations" might read: delivered all goals for the role, and regularly took on work expected one level up, with evidence from at least two quarters. Compare that with "goes above and beyond," which every manager will read differently. For more wording ideas, see these performance rating scale examples.

Pick two anonymized benchmark employees per level as well: one solid "meets" and one clear "exceeds." They give the group fixed reference points during the session.

2. One Week Before: Collect a One-Page Pre-Read for Each Employee

Ask every manager to complete the same short template for each direct report. A shared format lets the group compare like with like.

Field What to include
Employee details Name, role, level, and time in role
Proposed rating The manager's rating, plus a note if it's a borderline call
Goals and outcomes Two or three goals, each with its result
Evidence Two or three specific examples from across the review period, each with what happened and its impact
Peer feedback Main themes from peer or 360-degree feedback, where collected
Development area One area to work on next cycle
Promotion nomination Yes or no; if yes, examples measured against next-level criteria

Review the pre-reads as they arrive. Send back any pre-read that relies on adjectives such as "strong" or "reliable" instead of examples.

3. Two to Three Days Before: Build the HR Data Pack

The data pack shows the facilitator where to spend the meeting's time. It should include:

  • Rating distribution by manager and by team, to spot lenient or strict raters
  • A comparison with last cycle's ratings for the same group
  • Ratings broken down by gender, work location, and tenure, for the bias check
  • A flagged list: borderline cases, ratings that moved two levels since last cycle, and promotion nominations

Send the agenda and the flagged list 48 hours before the meeting. Ask managers to mark any employee on another team they want to discuss, so those cases join the flagged list.

How to Run a Calibration Meeting, Step by Step

Run a calibration meeting in seven steps: set the ground rules, re-anchor on the rating definitions, review the distribution, calibrate the extremes, work through borderline cases, review promotions, then check for bias and close. The order matters, because each step sets reference points for the next.

1. Open With the Purpose and Ground Rules

State the purpose in one sentence: the group is checking that ratings follow the same standard, not voting on people. Then set four ground rules:

  • Everything said in the room stays confidential.
  • Evidence beats adjectives. A claim without an example doesn't move a rating.
  • Every manager discusses every employee, not only their own team.
  • Everyone knows the decision rule before the first case, including who settles a deadlock.

2. Re-Anchor on the Rating Definitions

Read the definitions for each level aloud and walk through the benchmark employees. It takes about ten minutes. Skipping it is the fastest way to get ten managers using ten slightly different scales.

3. Review the Distribution Before Any Individual

Show the data pack before discussing anyone by name. If one team has 70% of its people rated "exceeds" and a similar team has 10%, the group should know before the case-by-case discussion begins.

Treat a skewed distribution as a question, not a verdict. Some teams really do outperform. Don't push ratings toward a fixed curve; forced distribution makes the group compare people with each other instead of with the standard.

4. Calibrate the Highest and Lowest Ratings First

Start with the top and bottom of the scale. These ratings carry the biggest consequences, such as larger raises or a performance improvement plan, and they usually have the clearest evidence. Settling them first gives the group anchors for the harder calls.

Use the same format for every employee:

  1. The manager presents the evidence first, then the proposed rating, in about two minutes.
  2. Other managers ask questions and share what they've seen.
  3. The facilitator asks whether the evidence matches the definition for that rating.
  4. The group confirms or adjusts the rating, and the note-taker records the outcome and the reason.

Asking for evidence before the rating is deliberate. Once a number is said out loud, the discussion tends to anchor on it.

5. Work Through Borderline and Flagged Cases

Next, discuss employees on the edge between two levels and anyone another manager flagged. These are the cases where calibration changes outcomes.

Say out loud that undisputed middle ratings won't be discussed one by one. Otherwise managers may assume those employees were overlooked.

6. Review Promotion Nominations Separately

Keep promotions as their own block. A promotion case needs evidence that the person already performs at the next level, not just a strong rating at the current one.

For each nomination, ask two questions. Which next-level expectations has this person already met, with examples? And if the answer is "not yet," what would they need to show by next cycle?

7. Run a Bias Check, Then Close

Before anyone leaves, review the final ratings by gender, work location, and tenure, with names removed. Look for patterns the evidence doesn't explain, such as ratings that moved down more often for one group.

Then close in three moves. Read back every change and its reason. Confirm that each manager will deliver their team's results personally. Set a date by which every review conversation must happen.

Calibration Meeting Agenda: A 90-Minute Template

This 90-minute calibration meeting agenda follows the seven steps above. It fits six to eight managers who sent pre-reads in advance. Shift time between the two discussion blocks to match how many flagged cases you have.

Time Block What happens
0–5 min Purpose and ground rules Confidentiality, evidence over adjectives, and the decision rule
5–15 min Rating definitions Read definitions aloud and review the benchmark employees
15–25 min Distribution review Ratings by manager and team, plus changes since last cycle
25–50 min Highest and lowest ratings Evidence first, then the proposed rating, then the group decision
50–70 min Borderline and flagged cases Edge cases and employees other managers asked to discuss
70–80 min Promotion nominations Evidence against next-level expectations
80–87 min Bias check Final ratings by gender, location, and tenure, with names removed
87–90 min Close Read back changes, confirm owners, and set the conversation deadline

For a team of fewer than about 20 people, a 45-minute version works: keep the definitions and the bias check, and merge the two discussion blocks. If a session regularly runs past two hours, split the group by function or level. Late cases deserve the same attention as early ones.

Pre-Work Email Template

Copy and adapt this email. Send it 48 hours before the meeting.

Subject: Calibration meeting on [date]: pre-work due [date]

Hi all,

Our calibration meeting for [team or function] is on [date] at [time].

Before then, please:

1. Complete a pre-read for each direct report by [date]. The template is here: [link].

2. Re-read the rating definitions: [link].

3. Review the attached flagged list. Reply with any employee on another team you'd like to discuss.

In the meeting, we'll present evidence before ratings, and we'll focus on the highest, lowest, and borderline cases.

Everything discussed stays confidential.

Thanks,
[Name]

Calibration Meeting Questions That Keep the Discussion Fair

Good calibration questions turn opinions into evidence. The facilitator asks most of them, but any manager can. These eight cover most situations:

  1. What specific result or behavior supports this rating?
  2. Which part of the rating definition does that example match?
  3. How does this compare with someone at the same level who received the same rating?
  4. Does the evidence cover the whole review period, or mostly the last two months?
  5. What did this person deliver that others on the team didn't?
  6. Who else saw this work, and does their view match yours?
  7. Would this rating hold if the person worked in a different location or on a different team?
  8. What evidence would change your mind?

The last question is the most useful one in a stalemate. It moves the conversation from defending a position to naming what would settle it.

Vague Phrases to Challenge

Some phrases sound like evidence but aren't. When one comes up, ask for the example behind it.

What a manager says What to ask instead
"She has a great attitude." Which situation showed that, and what was the result?
"He's not proactive." When did the role call for initiative, and what happened?
"She's very visible." What did the visible work deliver?
"He's a strong communicator." Which communication made a measurable difference?
"She needs to step up." Which expectation for her level isn't she meeting?
"He's a culture fit." Which of our stated values did he demonstrate, and how?

How to Handle Disagreements and Changed Ratings

When managers disagree about a rating, the facilitator should summarize both views, name the gap against the rating definition, and ask what evidence would settle it. If the group still can't agree, apply the decision rule set at the start. Don't invent a new one mid-meeting.

1. A Script for the Facilitator

"So far we've heard two views. Priya sees the platform migration as work above level. Daniel sees it as a well-scoped project the role already expects. Our definition of 'exceeds' asks for regular work at the next level across at least two quarters. What evidence would move either of you?"

The script works because it restates each position fairly, points back to the written standard, and asks for evidence rather than agreement.

2. the Decision Rule in Advance

Organizations usually choose one of two rules:

  • The manager owns the final rating but must address the group's input in writing before submitting it.
  • The group decides by consensus, and a named senior leader breaks deadlocks.

Either rule works if everyone knows it before the first case. Also cap any single discussion at about ten minutes. If it's still unresolved, park it, gather the missing evidence, and revisit it at the end or in a short follow-up.

3. When a Rating Changes

If calibration moves a rating, the manager still delivers it and explains it against the criteria. A manager should never say "the committee decided" or "I fought for you, but it was out of my hands." That shifts blame and makes the whole process look arbitrary.

A manager who disagrees with the final outcome should raise it with HR or their own manager, not with the employee.

How Calibration Meetings Can Introduce Bias (and How to Stop It)

Calibration meetings are meant to reduce bias, but the meeting itself can add new bias. In a January 2024 Harvard Business Review article, researchers from the Equality Action Center and the Center for WorkLife Law at UC Law San Francisco argue that calibration meetings can introduce bias in several ways. They also note that small changes, such as teaching participants what bias looks like, can help.

These are the distortions facilitators most often need to watch for:

Bias How it shows up in the room What the facilitator can do
Anchoring The first rating said aloud frames the whole discussion Ask for evidence before the proposed rating
Seniority and loud voices The most senior or most confident person sets the outcome Use a fixed speaking order, and have senior leaders speak last
Advocacy Managers argue for the highest rating for their own people Ask managers to argue for the correct rating, not the highest
Leniency or strictness One manager rates everyone high, another rates everyone low Compare each manager's distribution in the data pack
Recency The last few weeks outweigh the full review period Ask for examples from each quarter (more on recency bias)
Halo and horns One standout success or mistake colors the whole rating Assess each goal separately (halo and horn effect)
Proximity People seen in the office more get rated higher Compare remote and on-site employees at the same level (proximity bias)
Fatigue Cases late in a long session get less attention Take a break every 45 minutes, and rotate which team goes first each cycle

Open each session with a two-minute reminder of these patterns, with a real example. Naming bias at the start makes it easier for anyone to point it out later without sounding accusatory. For a wider look at the topic, see our guide to performance review bias.

Performance Calibration Examples

The three examples below are illustrative, with hypothetical names and numbers. They show the kinds of changes a well-run calibration meeting produces and the reasoning behind each one.

Example 1: A Lenient Manager and a Strict Manager

Two customer success managers each lead teams of six to seven people handling similar accounts. Manager A proposes "exceeds" for five of her seven people. Manager B proposes "exceeds" for none of his six.

The group reviews each manager's evidence against the definition. Two of Manager A's five met their goals but didn't take on next-level work, so they move to "meets." One of Manager B's people ran renewals for a second region during a vacancy, which matches the "exceeds" definition, so she moves up.

The final result is three "exceeds" ratings on Team A and one on Team B. The gap still exists, but now it reflects the work rather than the manager.

Example 2: A Borderline "Meets" or "Exceeds"

A manager proposes "exceeds" for a support team lead, citing a 15% drop in ticket backlog. A peer manager asks which part of the definition that matches. The manager explains that the lead redesigned the triage process and trained two other teams on it.

The facilitator asks whether that work came from the lead's own role or a stretch beyond it. The group agrees that training other teams goes beyond the role, and the rating holds. The note-taker records the reason: cross-team process work above level.

Example 3: A Quiet Contributor Others Missed

An engineer's manager proposes "meets." During the discussion, a manager from another team mentions that the engineer led most of a data migration the two teams shared. The engineer's own manager hadn't seen that work in detail.

The group asks for written confirmation from the second manager, then moves the rating to "exceeds." The note-taker records the source of the new evidence. Without calibration, that contribution would have stayed invisible in the review.

After the Meeting: How to Communicate Calibrated Ratings

After calibration, each manager delivers their own team's ratings, explains them against the rating definitions, and shares the useful feedback that came up in the room. HR's job is to prepare managers for those conversations, not to hold them.

We recommend holding review conversations within a week or two of calibration, before details fade and before rumors fill the gap. Our guide on how to conduct a performance review meeting covers the conversation itself.

What Managers Should and Shouldn't Say

The wording matters most when calibration surfaced a different view of someone's work.

Avoid Say instead
"I thought you were great, but the other managers didn't agree." "Your work on our team was strong. Feedback from other teams showed it wasn't as consistent across projects, so let's work on that."
"The committee decided your rating." "Your rating is 'meets.' Here's how your results compare with our definition for each level."
"You just missed 'exceeds.'" "You were close to 'exceeds.' Here are the two things that would get you there next cycle."

Explain the Process to Employees

Employees should know that calibration is part of the review process, why it exists, and which criteria it uses. They don't need a play-by-play of the discussion. A short explanation in the review guide or a team meeting helps calibration feel like a fairness check rather than a closed-door verdict.

Run a Short Retro With the Managers

Send the participants three questions within a day of the meeting. What helped? What slowed us down? What should change next cycle? Calibration improves fastest when the process gets the same scrutiny as the ratings.

How to Run a Calibration Session Remotely

A remote calibration session follows the same steps, but it needs more structure, because video calls make it easier for a few voices to dominate. These adjustments help:

  • Share one live tracking sheet on screen, so everyone sees the same rating and the recorded reason as each decision is made.
  • Use round-robin input for each contested case, rather than open discussion that favors whoever unmutes first.
  • Split long sessions across time zones instead of holding one meeting at an hour that suits only part of the group.
  • Ask for outcomes, not observations of presence. "Always online" and "very responsive on Slack" describe visibility, not results.
  • Compare remote and on-site ratings for employees at the same level during the bias check.

Common Calibration Meeting Mistakes to Avoid

Most calibration problems come from a handful of repeat mistakes. Each one has a straightforward fix.

Mistake What it causes The fix
Treating the meeting as a group writing session Managers draft ratings live, and the meeting overruns Require finished pre-reads before the invite goes out
Discussing pay and ratings in the same session The conversation drifts toward budgets instead of evidence Settle ratings first, then hold a separate compensation discussion
Running sessions back to back Facilitators and managers tire, and later cases blur together Spread sessions over several days, with no more than two a day per facilitator
No record of why ratings changed The same arguments return next cycle, and decisions are hard to defend Note every change and its reason during the meeting
Calibrating only once a year Standards drift between cycles Add a short mid-year check on the rating definitions
No feedback on the process The same friction repeats every cycle Run a three-question retro after every session

When Spreadsheets Stop Working for Calibration

A spreadsheet can handle one team's calibration. It struggles once you calibrate across several groups. Versions multiply, change reasons end up in someone's notebook, and nobody is sure which ratings are final.

A performance review tool with calibration built in keeps the proposed rating, the calibrated rating, and the reason for each change in one place. In ThriveSparrow Performance, you can:

  • Spot inconsistencies before the meeting. Rating distributions, heatmaps, and bell curves show which teams or managers skew high or low.
  • Set up calibration groups by team, function, or level, each with its own employees and panel.
  • Calibrate only the questions that matter, such as an overall rating, instead of revisiting every score.
  • Record a reason for every change, visible to every authorized calibrator.
  • Lock final scores for one employee, a selection, or the whole group, and reopen a locked score if new evidence appears.
  • Keep the manager in charge of the message. Employees see only their final score, not whether it changed during calibration.

Running calibration in spreadsheets? See how calibration groups, change reasons, and locked scores work inside a real review cycle. Try ThriveSparrow Performance free for 14 days!

Start With the Definitions

If you change only one thing before your next calibration meeting, rewrite your rating definitions so managers can point to them rather than interpret them. Clear definitions make the pre-reads sharper, the discussion shorter, and the final ratings easier to explain.

Then decide where the process will live. One team can calibrate in a spreadsheet. Several teams need one place for proposed ratings, change reasons, and locked results. See how calibration works in ThriveSparrow.

Calibration Meeting FAQs

1. How long should a calibration meeting take?

Most calibration meetings take 60 to 90 minutes for a group of six to eight managers, provided pre-reads arrive in advance. If a session regularly runs past two hours, split it by function or level. That way, the last cases get the same attention as the first.

2. How often should you hold calibration meetings?

Hold a calibration meeting before ratings are finalized in every formal review cycle, whether that's annual or twice a year. Teams that change quickly can add a short mid-year check to confirm managers still read the rating definitions the same way.

3. Should a calibration meeting force ratings onto a bell curve?

No. Forcing ratings into fixed percentages makes managers compare people with each other instead of with the standard. Use the rating distribution as a prompt for questions, such as why one team rates much higher than a similar team, not as a quota to hit.

4. Who makes the final call on a rating?

That depends on the decision rule you set before the meeting. In many organizations, the manager owns the final rating but must address the group's input. In others, the group decides by consensus and a named senior leader breaks deadlocks. What matters is that everyone knows the rule before the first case.

5. Do small companies need calibration meetings?

If two or more managers rate employees and those ratings affect pay or promotions, a short calibration meeting is worth it. A 45-minute session with two or three managers, the founder or HR lead as facilitator, and written rating definitions is enough to catch the biggest inconsistencies.

‍

‍

ABOUT THE AUTHOR
Shankari Gurumoorthi
Growth Marketer
Growth Marketer at ThriveSparrow who has spent years deep in the world of HR technology, people strategy, and the real stories behind how teams work and grow.
LinkedIn