Home/Blog/Customer Service
Customer Service · 10 min read

How to Build a Call Center Quality Assurance Program

By Bodhih Training · Updated

The short answer

To build a call center quality assurance program, define quality from the customer's point of view, turn it into a short scorecard with auto-fail, customer-critical and weighted items, sample interactions at random, evaluate with evidence, calibrate evaluators every month, coach within a week, and use Pareto and root-cause analysis to fix the processes behind repeated misses. Check that QA scores broadly agree with what customers tell you, and keep a person in charge of every judgement about an agent.

Key takeaways
  • Write the purpose down: customer outcomes, compliance, coaching and process improvement, not policing
  • Keep auto-fails to two to four compliance and conduct items and weight what customers feel
  • Four evaluations a month find coaching topics; they cannot rank agents reliably
  • Calibrate monthly and rewrite unclear definitions the same day
  • Trace repeated misses to root causes before coaching anyone
  • Use automated QA to flag and draft; let people decide

What is call center quality assurance?

Call center quality assurance (QA) is the practice of reviewing a sample of customer interactions against an agreed standard and using what you find to protect customers, meet legal and regulatory duties, help agents improve and fix the processes behind poor service. It applies to calls, live chat, email and social replies, and increasingly to conversations that start with a chatbot.

QA is different from customer surveys. Surveys tell you what customers felt. QA tells you what happened in the interaction and why. You need both, and they should broadly agree. When they do not, the scorecard is usually measuring the wrong things.

Two international standards are useful reference points. ISO 18295-1:2017 sets out service requirements for customer contact centres, in-house or outsourced, including performance measures. ISO 10002:2018 gives guidelines for handling complaints, including analysing them to improve products and services. Neither tells you how to design your scorecard, but both reinforce the same idea: measure service from the customer's side and use what you learn.

What is QA for, and what is it not for?

A quality program has four legitimate jobs, in this order: protect the customer's outcome, protect customers and the organization from harm (identity, data, payments, regulation), help agents improve, and improve the system. Catching people out and producing a league table are not on the list. Programs that drift into policing get compliance theatre: agents learn the phrases that score and stop thinking about the customer.

Start by writing a one-page QA policy that says what QA is for, how interactions are selected, how scores work, how agents can dispute a score, and how results will and will not be used. In particular, say whether a single month's score can affect pay or discipline. Using QA in formal performance processes raises fairness and employment law questions that differ by country, so check with your HR team and local law before linking the two.

How do you design a QA scorecard?

Begin with customer outcomes in the customer's own words, drawn from complaints, survey comments and repeat contacts: 'they understood my problem first time', 'they told me the right thing', 'I knew what would happen next'. Then list the observable behaviours that produce each outcome. 'Was empathetic' is not observable. 'Named the impact on the customer before offering the fix' is.

Sort the behaviours into three types. Auto-fail items cover compliance and conduct, such as identity verification, payment card handling and serious misconduct; any breach sets the final score to zero. Customer-critical items, such as accuracy and clear next steps, carry heavy weight and are flagged whenever they score zero. Weighted items shape the experience. Give the weighted items 100 points between them and put most of the weight where customers feel it.

Score each item on a short scale with a written definition for each level. A three-level scale (met, partly met, not met, plus not applicable) works well for most teams because evaluators can apply it consistently. Then pilot the scorecard: two evaluators score the same ten interactions alone, and you rewrite every definition that produced a full-point difference.

Item typeExamplesHow it scores
Auto-failIdentity verification, payment card data, abusive conductAny breach sets the final score to zero; reviewed case by case
Customer-criticalAccurate information; resolution or clear next stepHeavily weighted; a zero triggers coaching
WeightedUnderstanding the need, empathy, clear language, notesPoints earned divided by points available
Not applicableWritten quality on a call; hold handling with no holdLeft blank; never counts against the agent

How do you adapt QA for chat, email and social?

Use one scorecard with channel applicability rather than separate scorecards for each channel, so results stay comparable. Mark which items apply where and add a short channel note to each definition. On a call, understanding the need sounds like a spoken playback of the problem. In an email, it reads as a reply that answers every question the customer asked.

Written channels add their own failure points: copy-pasted templates that ignore what the customer wrote, long response gaps in chat, and correct answers buried in dense paragraphs that customers read on a phone. Public social replies are read by many people, so keep them short and move anything personal to a private channel.

When a chatbot starts the conversation, include the handoff in your review. Gartner's survey of 5,728 customers, published in July 2024, found that 64% would prefer companies did not use AI in customer service, and their top concern was that it would make it harder to reach a person. Check that the customer knew they were dealing with an automated system, that the handoff carried the context, and that the agent read it.

How many interactions should you evaluate per agent?

It depends on what the sample is for. Quality scores vary from contact to contact for reasons unrelated to the agent's skill, and that variation sets a limit on what a small sample can tell you. If an agent's scores typically vary by about 10 points either way, the margin of error on a monthly average from four evaluations is roughly plus or minus 10 points at 95% confidence. To get within plus or minus 5 points you need about 16 evaluations.

So use small agent samples for coaching topics and three-month trends, not rankings. Pool evaluations for team, channel and item rates, where the sample is much larger. If you need an operation-wide rate, such as the share of contacts that pass identity verification, the standard formula for estimating a proportion gives about 1,067 contacts for plus or minus 3 points at 95% confidence when you have no prior estimate, slightly fewer once corrected for your monthly volume.

Select interactions at random across days, times and contact reasons. Targeted reviews of complaints, low survey scores or analytics flags are valuable, but record and report them separately; mixed into the random sample, they make agents look worse than they are. Finally, check capacity: multiply evaluations by minutes per evaluation, including the time to write feedback, and compare with the hours your evaluators have after calibration and admin.

  • Random sample for each agent's monthly evaluations
  • Targeted reviews recorded separately
  • Pooled results for team and item rates
  • Three-month trends for individuals
  • Capacity checked before you promise a sample size
Measure where you are

Reading helps; measuring tells you what to work on. These AI-graded assessments on AssessAll pair with this topic:

How do you run a QA calibration session?

Calibration keeps scores meaningful. Every few weeks, evaluators and team leads score the same two or three interactions alone, before the session. In the session they compare item by item, explain what they heard and which words in the definition they relied on, agree a benchmark score, and rewrite any definition that caused a difference. The most senior person speaks last.

Measure agreement simply: the share of scores within a tolerance of the benchmark (for example plus or minus 5 points) and whether each evaluator made the same auto-fail call. Track each evaluator's average difference from the benchmark to spot leniency or harshness drift. More formal statistics such as Cohen's kappa exist for detailed item-level analysis, but a transparent within-tolerance measure is easier to explain and act on. Never cancel calibration to save time; cut the number of evaluations first.

How do you coach agents from QA results?

Coaching is where QA pays for itself. Give a short feedback note within two working days of every evaluation, with one strength and at most one improvement, each tied to a timestamp or a line and to what the customer experienced. Hold a coaching conversation within a week for anything below the pass mark, any customer-critical miss and any auto-fail.

A 20-minute structure works well: ask for the agent's view of what the customer needed first, listen to or read the key 30 seconds together, agree what a top score would sound like in the agent's own words, practise it once, and set a follow-up date to check it stuck. Add side-by-sides for new starters, a call library of short clips of strong moments shared with the agent's consent, and monthly self-evaluation, where agents score one of their own contacts with the same definitions.

If coaching conversations are where your team leads struggle, measure the skill directly. AssessAll's Coaching Conversation Assessment with a Spoken Round for Managers and Internal Coaches is one way to see where a coach stands before and after a few months of practice.

How do you turn QA scores into improvements?

Look past the average. For each scorecard item, calculate points lost (the item's weight multiplied by the misses) and rank the items. A few items usually account for most of the lost points; this is the Pareto principle that Joseph Juran popularized in quality management. Then read the evaluator notes on those misses and assign a root cause: knowledge gap, process or policy, system or tool, agent behaviour, or customer and third party.

If three or more agents miss the same item for the same reason, treat it as a system problem first. A refund rule with two possible timescales, an out-of-date knowledge article or a missing ticket category will defeat any amount of coaching. Take example contacts to the owner of the process and track the fix.

Then check QA against what customers say. Group evaluated contacts by score band and compare survey results and first-contact resolution for each band. Customer measures to consider include CSAT, Net Promoter Score (introduced by Fred Reichheld of Bain & Company in Harvard Business Review in 2003) and Customer Effort Score, from research by Matthew Dixon, Karen Freeman and Nick Toman, published in Harvard Business Review in 2010, which found that customers mostly want a quick, simple solution rather than delight. Be cautious with the service recovery paradox: a 2007 meta-analysis in the Journal of Service Research found a positive effect of good recovery on satisfaction but no significant effect on repurchase intentions or word of mouth.

Where should AI and automated QA fit?

Speech and text analytics can check objective items on every contact, flag interactions worth a human look and spot trends in contact reasons. Generative AI can summarize a transcript against your scorecard and draft feedback. Gartner predicted in March 2025 that agentic AI will autonomously resolve 80% of common customer service issues without human intervention by 2029, so QA will increasingly need to cover machine-handled contacts too.

Automation is weaker on judgement: whether empathy was genuine, whether next steps were clear enough for this customer, whether a hold was reasonable during an outage. Run any tool in parallel with human evaluators first, calibrate it monthly, and require a person to review every score before it reaches an agent or affects a decision about them.

Check the law before you switch features on. In the European Union, the AI Act has prohibited AI systems that infer employees' emotions from biometric data such as voice in the workplace since 2 February 2025, except for narrow medical or safety reasons, and transparency duties for AI systems that interact with people apply under the Act's timetable from 2 August 2026. Recording consent, employee monitoring, data retention and payment card rules also differ by country and sometimes by state. Check your privacy law, your HR team and a qualified adviser. The Call Quality Monitoring Kit from Bodhih Training keeps these specifics in one dated cheat sheet alongside the scorecard, workbook, calibration and coaching tools.

Automation can help withA person must decide or check
Checking disclosures and verification statements on every contactWhether a flagged breach really happened
Flagging long silences, repeat contacts and complaint wordsWhich flagged contacts to review and why
Summarizing a transcript against the scorecardEvery score before it reaches an agent
Spotting trends in contact reasonsRoot causes and who owns the fix
Drafting feedback notesThe final words and the coaching conversation
The Call Quality Monitoring Kit e-book cover
Bodhih Pro Kit

Run your QA program with every step ready

The Call Quality Monitoring Kit from Bodhih Training gives you an e-book, a quality monitoring workbook, evaluation, calibration and coaching forms, an evaluator handbook, a message library and a 60-day launch calendar, so each step in this guide has a file ready to open.

Common questions

Questions people ask next

What is a good QA score for a call center?

There is no universal figure, because the number depends entirely on how your scorecard is built and weighted. Set a target after you have piloted the scorecard and seen the spread of scores, and pay as much attention to auto-fails, customer-critical misses and the trend as to the average.

Should team leads evaluate their own agents?

They can, and it keeps coaching close, but their scores tend to drift and standards differ between teams. Many operations use dedicated quality analysts for most evaluations, with team leads doing a few each month and joining every calibration session.

How long should a QA evaluation take?

Time your own. A voice evaluation usually takes longer than the call because good evaluators listen once for the customer's outcome and again to score with timestamps, then write a short feedback note. Use the time per evaluation to plan evaluator capacity.

What is an auto-fail in QA?

An auto-fail is a compliance or conduct item, such as skipping identity verification or recording card details, where any breach sets the final score to zero. Keep the list to two to four items and record the behaviour score separately so coaching can address both.

How should agents dispute a QA score?

Give agents a short window to raise a dispute, have someone who did not score the original interaction decide within a few working days, and change the definition for everyone when it caused the problem. A dispute route protects trust and often improves the scorecard.

How can evaluators check their own listening and feedback skills?

Calibration results show how closely you match the benchmark. For the underlying skills, AssessAll's Workplace Listening Assessment and its Written Feedback Quality Assessment are options; repeat them after a few months to see what changed.

How do we help quality analysts and team leads keep developing?

Pick one or two skills from calibration and coaching results, such as writing evidence-based feedback or running agent-led coaching conversations, and build them into a personal development plan with measures. Jobulary explains how a development plan built around specific skills works.