How to Run a Performance Review Cycle, Step by Step
By Bodhih Training · UpdatedThe short answer
To run a performance review cycle, agree what it is for, choose a rhythm and rating scale with clear definitions, keep evidence through the year with check-ins, collect self and peer input, rate against anchors, calibrate ratings across managers without quotas, check outcomes by group for fairness, then hold the review conversation, decide pay separately and follow through on development, appeals and a review of the cycle itself.
- Name the purposes of the cycle in order; development and pay pull against each other
- Define every rating level, and make the middle level a good rating
- Rater effects can outweigh real performance in ratings, so use anchors, evidence and calibration
- Calibrate on evidence, never to a forced curve, and record every change
- Check outcomes by group before ratings are final
- Hold the pay conversation separately from the review conversation
What is a performance review cycle?
A performance review cycle is the repeating sequence an organisation uses to agree goals, gather evidence of performance, reach a fair judgement and turn that judgement into development, pay and other decisions. Most organisations run it once or twice a year, with lighter check-ins in between.
The cycle has a bad reputation, and some of it is earned. In a 2019 article, Gallup's Robert Sutton and Ben Wigert reported that only 14% of employees strongly agree their performance reviews inspire them to improve. Writing in Harvard Business Review in 2015, Marcus Buckingham and Ashley Goodall described Deloitte's finding that it was spending close to 2 million hours a year on performance management. The answer is not to stop reviewing performance. It is to run a cycle with a clear purpose, real evidence and fair decisions.
The table below shows the nine steps of a well-run cycle, what each produces and the mistake it prevents.
| Step | What it produces | Mistake it prevents |
|---|---|---|
| 1. Design | A purpose statement, rhythm, rating scale, weights and dates | A process that tries to do everything |
| 2. Goals and check-ins | Weighted goals with measures and quarterly notes | Reviews written from memory |
| 3. Launch and train | A calendar, clear messages and a manager briefing | Missed deadlines and untrained raters |
| 4. Self and peer input | Evidence from the employee and colleagues | One-sided judgements |
| 5. Manager review | Scores against anchors and a specific summary | Vague, biased or recency-driven ratings |
| 6. Calibration | Ratings agreed across managers with recorded reasons | Lenient and severe raters deciding outcomes |
| 7. Fairness check | Outcomes compared across groups | Hidden patterns against a group |
| 8. Conversation, pay, development | A useful conversation, a separate pay decision, a plan | Development drowned out by the pay number |
| 9. Follow-through | Appeals, support plans and a retrospective | Disagreement with no route and a cycle that never improves |
What should a performance review cycle be for?
Most cycles try to serve three purposes: development (helping people grow), reward (evidence for pay and bonuses) and decisions (promotion, talent planning and support plans). They pull against each other. The CIPD's evidence review on performance management, published in December 2016, found that ratings given for development tend to be stricter and drawn from a wider range of evidence than ratings used for administrative decisions such as pay, and suggested separating developmental conversations from administrative ones.
Write a short purpose statement with leadership before you touch any form: the first purpose, the second, what the cycle is not for, how the purposes are kept apart and how you will know it worked. A common choice is development first, evidence for pay second, with the pay conversation held a week or two after the review conversation.
Should you move to continuous performance management?
From around 2012, well-known employers began moving away from the classic annual appraisal. A Wharton Knowledge article in September 2016 named Adobe, Kelly Services, GE, Deloitte and PwC among firms that had ended annual reviews, often in favour of regular conversations, at minimum quarterly. The approach is usually called continuous performance management: frequent check-ins, goals that can change and feedback close to the work.
The useful lesson is not that annual reviews are bad. It is that the year-end moment should summarise conversations that already happened. If you link pay to performance, you will need some form of rating either way, and a visible, defined rating is easier to explain and challenge than a hidden one. Many organisations land on an annual or half-yearly review with quarterly check-ins.
How do you design a fair rating scale?
A scale with labels alone invites every manager to supply their own meaning. Give each level a definition and a line on what it looks like in evidence. Define the middle level as a good rating, earned by most capable people most years, so that a 3 out of 5 does not feel like a fail.
Decide what counts and how much: results against goals, competencies (how the work is done) or both with weights, such as 60% goals and 40% competencies. Then decide the inputs: a self-assessment, peer input that asks for examples rather than scores, and the manager's evidence from the whole period. Pro-rate targets for part-time work, leave and mid-year changes at goal-setting time, not at review time.
- Five levels with definitions and evidence examples
- Score bounds so the arithmetic is consistent across managers
- Behavioural anchors for each competency
- A cut-off date for new joiners, who get a check-in instead of a rating
Which rating errors should managers watch for?
The research case for structure is strong. A 2000 study in the Journal of Applied Psychology by Steven Scullen, Michael Mount and Maynard Goff found that effects specific to the rater explained 62% and 53% of the variance in ratings in two data sets, while the performance of the people being rated explained 21% and 25%. A rating can tell you as much about the rater as the person rated.
Train managers to recognise the common errors and interrupt them before calibration: recency (only the last few weeks count), halo and horns (one strength or mistake colours everything), leniency and severity (everyone high, or nobody at the top), central tendency (a safe middle rating to avoid a hard conversation), similarity (rating people like yourself higher) and contrast (judging someone against the person rated before). Add attribution: blaming a person for results driven by workload, tools or targets.
Reading helps; measuring tells you what to work on. These AI-graded assessments on AssessAll pair with this topic:
How do you run a calibration meeting without forced ranking?
Calibration is a meeting in which managers agree ratings across teams so that the same performance earns the same rating whoever the manager is. It is not forced ranking: nobody is moved to fit a quota. You can look at the distribution of ratings by team and by manager without imposing one; a flag is a question, not an answer.
When you read the distribution, compare each team's and each manager's average with the organisation's and ask what explains any large difference. A strong team may deserve high ratings; a lenient manager may not. Compare people in the same role and level side by side, and look at whether calibration moved ratings only in one direction, which can signal budget pressure rather than evidence.
Calibration can add bias of its own. A November 2025 SHRM article described confirmation bias, groupthink driven by seniority, clustering in the middle to avoid debate, rushed judgements and lobbying. A neutral facilitator, clear ground rules and timeboxed cases guard against all of these.
- Send a pre-read two working days before: scale definitions, distribution by team and manager, cases to discuss
- Read the ground rules aloud: evidence not adjectives, judge against definitions, no quotas or trading
- Discuss flagged teams first, then top ratings and promotions, then the lowest and borderline cases
- Record every change with its evidence before moving on
- Have junior voices speak before senior ones on each case
How do you check performance ratings for fairness?
Before ratings are final, compare outcomes across groups such as gender, work pattern, location, level and age band, where your policy and local law allow you to hold that data. Look at the average rating, the share of top ratings and promotion readiness for each group against the whole organisation.
Treat a gap as a reason to investigate, not proof of bias, and do not report groups below a minimum size such as five people. Check the process first: were part-time targets pro-rated, was leave counted against people, is one group's work less visible to raters, is one manager driving the pattern? Correct individual ratings where the evidence supports it, fix the process for next cycle and take advice before acting on anything with legal implications.
A simple screen is to divide each group's share of top ratings by the organisation's share. A ratio well below 1, such as under 0.8, is worth a closer look. Treat that number as a prompt for questions rather than a legal test, and remember that level, tenure and age often move together, so check whether one explains another before drawing conclusions.
How should performance reviews link to pay?
A published merit matrix makes pay decisions consistent and explainable. It sets guideline increases by final rating and by position in the pay range, measured as compa-ratio (salary divided by the range midpoint). People low in the range get a larger percentage for the same rating, which moves pay towards the market rate over time. Check the total against budget, check that higher ratings really earn more, and compare increases by group within the same rating.
Transparency rules are tightening. The EU Pay Transparency Directive (2023/970) requires employers to make the criteria used for pay and pay progression easily accessible to workers, and those criteria must be objective and gender-neutral; member states may exempt employers with fewer than 50 workers from the progression part. Wherever you are, hold the pay conversation separately from the review conversation so development is heard.
Where does AI fit in performance reviews?
AI assistants can summarise check-in notes, group anonymised peer input into themes, check drafts for vague or biased language and draft routine messages. They should not decide ratings, pay, promotion or improvement plans. The manager decides the rating from the evidence, calibration agrees it, and a person can explain every decision without referring to the tool.
Use only approved tools for employee data, remove names and sensitive details before pasting, and check every output against the evidence. In some places AI that evaluates workers' performance carries legal duties: under the EU AI Act such systems are classed as high-risk, with obligations now due from December 2027. Managers who want to test how well they write and deliver reviews can use an AssessAll assessment such as the Written Feedback Quality Assessment, then turn the results into a Jobulary individual development plan. The Performance Review Cycle Kit from Bodhih Training includes an AI prompt pack with checks for each step.

Run your next review cycle with every file ready
The Performance Review Cycle Kit gives you the method, the workbook, the forms, the calibration guide, the scripts and the pay simulator for the whole cycle. Open the Start Here map and your first hour is planned.
Sources
- CIPD: What works in performance management? (evidence review, December 2016)
- Gallup: More Harm Than Good: The Truth About Performance Reviews (Sutton and Wigert, 2019)
- Harvard Business Review: Reinventing Performance Management (Buckingham and Goodall, 2015)
- Knowledge at Wharton: The End of Annual Performance Reviews (2016)
- SHRM: How Calibration Meetings Can Add Bias to Performance Reviews (2025)
- EUR-Lex: Directive (EU) 2023/970 on pay transparency
More from the Bodhih family
Questions people ask next
How often should performance reviews happen?
Most organisations hold a formal review once or twice a year with quarterly check-ins in between. The right rhythm depends on how fast goals change, how much time managers have and when pay decisions are made. Whatever you choose, the formal review should summarise conversations that already happened.
Is a 3 out of 5 a bad rating?
It should not be. Define the middle level as fully meeting what the role needs, a good rating that most capable people earn most years, and say so in your FAQs and manager briefing.
What is the difference between calibration and forced ranking?
Calibration agrees ratings across managers using evidence and the scale definitions, and any number of people can earn any rating. Forced ranking requires a fixed share of people at each rating regardless of evidence.
Should peer feedback be anonymous?
A common middle path is to name reviewers to the manager and HR but share only themes with the employee. Ask for examples rather than scores, and include at least one reviewer outside the immediate team.
How do you handle an employee who disagrees with their rating?
Listen to the specific points, look again at any evidence they raise and explain the appeal route. If it does not settle, an independent reviewer should decide within a set time and give written reasons. Raising a concern should never count against anyone.
When should a performance improvement plan be used?
Only after concerns have been raised early and informally, causes explored and real support given. A written plan needs observable standards, support, review dates and HR involvement. Rules differ by country, so check your HR team, company policy and local law first.
How long does a performance review cycle take?
From launch to the last conversation, many organisations allow about ten to twelve weeks: three weeks for self and peer input, two for manager drafts, one for calibration and the fairness check, and three to four for conversations. Add design and training time before launch and a retrospective after.
What should HR measure after a review cycle?
Completion on time, the share of ratings changed in calibration, appeals per hundred people, fairness flags, and survey results on whether reviews felt fair and useful. Use them in a retrospective and change two or three things for the next cycle.