How to Evaluate Training Effectiveness and Measure Training ROI
By Bodhih Training · UpdatedThe short answer
To evaluate training effectiveness, agree the business result with the sponsor before design, then gather evidence at each level: reaction (relevance and readiness), learning (scenario tests before and after), behaviour (checks at 30, 60 and 90 days) and business results compared with a similar untrained group. For ROI, convert the isolated change to money, subtract fully loaded costs, and show how the result changes if benefits are lower.
- Plan the evaluation with the sponsor before the training is designed
- Reaction surveys are an early warning, not proof of impact
- Test decisions in realistic scenarios, not recall of facts
- Behaviour at 30, 60 and 90 days decides whether results follow
- Never claim the whole improvement: use a comparison group, trend or adjusted estimates
- Show ROI on fully loaded costs with a sensitivity line
Why does most training evaluation stop at the smile sheet?
Most organisations evaluate the easy levels and stop. In its 2019 report Effective Evaluation, the Association for Talent Development surveyed 779 talent development professionals: about eight in ten organisations evaluated reaction and learning, 54% evaluated behaviour on the job, 38% business results and 16% return on investment. The barrier named most often, by 41%, was difficulty isolating a programme's effect on results.
The cause is usually timing rather than skill. Evaluation is planned after design, or after delivery, when the baseline data is gone, everyone has been trained and the sponsor has moved on. The rest of this guide puts the steps in the order that makes higher-level evidence possible.
Which training evaluation model should you use?
Four approaches cover almost every need, and they work together rather than competing. Use the table to pick the right one for the question you need to answer.
In practice, most L&D teams use the four levels as the structure of every plan, add the Phillips rules on isolation and costs for their biggest programmes, run success case interviews when they need to understand why some people got results and others did not, and use LTEM as a check that they are not mistaking attendance or satisfaction for learning. Decide the depth for each programme from three things: its fully loaded cost, its strategic importance to a senior leader, and how many people it reaches. A short webinar needs a quick check of usefulness; a programme costing as much as two salaries needs evidence a finance director would accept.
| Approach | Credited to | Best for | Watch out for |
|---|---|---|---|
| Four levels: reaction, learning, behaviour, results | Donald Kirkpatrick (1950s), later Jim and Wendy Kirkpatrick | A shared structure for any evaluation plan; starting from results | Stopping at the first two levels |
| ROI Methodology: adds a fifth level and rules for isolation, money and costs | Jack Phillips and the ROI Institute | Programmes big enough that someone will ask about money | Forcing ROI onto results that cannot be valued credibly |
| Success Case Method | Robert Brinkerhoff | Explaining why results happened or did not; credible stories | Presenting stories as the average result |
| LTEM: eight tiers from attendance to effects of transfer | Will Thalheimer (first published 2018) | Checking that your evidence is strong enough | Treating it as a reason never to use simple measures |
What should you agree with the sponsor before the training is designed?
Book 45 minutes with the sponsor and start at the end: six months after the programme, what is different in the business? Agree the business results to move, their baselines and targets, and where the data comes from. Then ask what else could move those numbers in the same period. That answer decides how you will isolate the training's effect.
Work down from there: three to five observable behaviours that drive the result, the support the work environment will provide, the skills and decisions people need, and the few reaction questions you will act on. Before results arrive, agree with finance the value per unit for each result and the ROI hurdle, so nobody can say you chose numbers to fit the answer. This top-down logic is the same idea behind Cathy Moore's Action Mapping in design.
- Business results with baselines, targets and data sources
- Other factors that could move the results
- Isolation method: comparison group, trend line or estimates
- Critical behaviours and how they will be observed
- Values per unit and the ROI hurdle, agreed with finance
- Report date and audience
How do you write a post-training survey that tells you something?
Reaction data has a weak link to later outcomes. A 2008 meta-analysis by Sitzmann and colleagues in the Journal of Applied Psychology found reactions had a modest relationship with immediate learning and none with delayed knowledge, and a 2010 meta-analysis by Blume and colleagues found that reactions about usefulness had a modest link to transfer while enjoyment had none. Treat the survey as an early warning system, not as evidence of impact.
Replace agree-disagree scales, where everyone ticks agree, with direct questions and descriptive answers, an approach Will Thalheimer calls performance-focused learner surveys. Ask how relevant the programme was to the person's work, how ready they are to apply it, how much practice they had, how clear their plan is and what support they expect from their manager. Add a 0 to 10 recommend question and two or three open questions, and run it before people leave the room. If you want to check your own survey-writing skills, AssessAll's Survey and Questionnaire Design Assessment is one option.
How do you test whether people actually learned?
Test decisions, not definitions. Write a blueprint from the skills in the evaluation plan, then write scenario items: a realistic situation with times, numbers and people, one best answer and three plausible wrong answers drawn from real mistakes. Most items should sit at the apply level or above in the revised Bloom's taxonomy. Build the pre-test and post-test from the same blueprint with different items, and observe practical skills with a checklist.
Report normalised gain alongside raw gain. Popularised by Richard Hake in a 1998 physics education study, it divides the gain by the room for improvement: (post minus pre) divided by (100% minus pre). A learner who moves from 55% to 85% has a normalised gain of 0.67, having closed two thirds of the gap.
Reading helps; measuring tells you what to work on. These AI-graded assessments on AssessAll pair with this topic:
- Training Evaluation and Programme Effectiveness Assessment for L&D and HR Teams (AssessAll)
- Post-Training Application and Workplace Support Check for Workshop Participants at 30, 60 and 90 Days (AssessAll)
- Survey and Questionnaire Design Assessment for Research, HR and Customer Insight Teams (AssessAll)
How do you measure behaviour change after training?
Name three to five critical behaviours an observer could see or hear, and check them at 30, 60 and 90 days. Use more than one source: self reports are cheap and good for spotting barriers, manager observation is stronger, peer input adds the team's view, and system data is strongest where a behaviour leaves a trace. Ask observers to write an example before choosing a rating, and to record 'not observed' rather than guess.
Remember that the work environment matters as much as the programme. Research on transfer of training has long pointed to manager support, the chance to practise and the time available as major influences on whether learning is used. That is why the behaviour check is also where L&D finds the levers it can still pull after the programme ends.
Hold a short 60-day conversation between each participant and their manager about what stuck, what did not and what support comes next. Read the barriers at 30 days and fix them then. Waiting until the 90-day report turns a solvable problem into a disappointing result. AssessAll's Post-Training Application and Workplace Support Check is designed for participants at 30, 60 and 90 days if you want a ready-made pulse.
How do you isolate the effect of training on business results?
Never claim the whole improvement. The Phillips approach lists several isolation methods; three cover most situations. A comparison group of similar people not yet trained is the strongest: the training's effect is the trained group's improvement minus the comparison group's improvement. Programmes that roll out site by site or cohort by cohort create comparison groups naturally. A trend line works when you have at least a year of stable data and nothing else changed. Where neither is possible, ask people close to the work for the share of the improvement they attribute to training and their confidence in that estimate, then multiply the two to stay conservative.
| Method | Calculation | Example |
|---|---|---|
| Comparison group | (Trained after - before) - (Comparison after - before) | Errors 0.62% to 0.41% vs 0.60% to 0.55%: effect 0.16 points |
| Trend line | Actual after - forecast | 96 lines per hour vs a forecast of 95: effect 1 |
| Adjusted estimate | Improvement x share x confidence | 3,100 hours x 40% x 70% = 868 hours |
How do you calculate training ROI honestly?
Convert each isolated result to money with a value finance already accepts, such as a standard cost per error or the historical cost of an early leaver. Use first-year benefits unless the sponsor agrees otherwise. Count fully loaded costs: needs analysis, design spread across cohorts, delivery, participants' and managers' time, venue and travel, overheads and the evaluation itself. Participants' time is the cost most often left out.
Finally, sense-check the result before anyone else does. Look for double counting, such as claiming both an overtime saving and a productivity gain that come from the same hours. Check that every value per unit has a named source and that every estimate is labelled as one. If the ROI looks spectacular, that is usually a sign that a cost is missing or an effect has not been isolated.
Benefit-cost ratio is benefits divided by costs. ROI is benefits minus costs, divided by costs, times 100. In a worked example with 75,423 in first-year benefits and 31,874 in fully loaded costs, the BCR is 2.37 and the ROI 137%. Then show sensitivity: in the same example, if only half the benefits were real the ROI would be 18%. Benefits that cannot be valued credibly, such as a handful of avoided safety incidents, belong in the report as intangibles, with evidence, not in the ROI.
How should you report the results to leaders?
Write for the decision. Put a half-page summary first: what the programme was for, what changed, how confident you are and what you recommend. Then the targets with results, the learning and behaviour picture, business results with the isolation method, the money with caveats, intangibles with one verified success story, and recommendations with owners and dates. Add a one-page scorecard with each level's target, result and status for busy leaders.
Preview disappointing results with the sponsor before they go wide, and include what did not work: an all-green report invites suspicion. Every finding should end in a decision to keep, change, stop or investigate. The Training Evaluation and ROI Kit from Bodhih Training includes a workbook that calculates every level from the sheets you fill in, along with the planning form, survey, follow-up forms, report template and scorecard.

Evaluate your next programme with every step ready
The Training Evaluation and ROI Kit from Bodhih Training gives you an e-book, an evaluation and ROI workbook, a planning form, a performance-focused survey, follow-up forms, a report template and a 90-day calendar, so each step in this guide has a file ready to open.
Sources
- ATD: L&D's Struggle With Learning Evaluation (on the 2019 Effective Evaluation report)
- Kirkpatrick Partners: The Kirkpatrick Model
- ROI Institute: ROI Methodology
- Work-Learning Research: LTEM, the Learning-Transfer Evaluation Model
- QIC-WD: Umbrella summary on trainee reactions (Sitzmann et al. 2008; Blume et al. 2010)
More from the Bodhih family
Questions people ask next
What are the levels of training evaluation?
Kirkpatrick's four levels are reaction, learning, behaviour and results. The Phillips ROI Methodology adds a fifth: return on investment. Plan from results down and collect evidence from reaction up.
What is a good ROI for training?
There is no universal benchmark. Agree a hurdle with your finance team before results come in, and judge the result against it together with the sensitivity line and the strength of the isolation method.
How long after training should you measure behaviour?
Check at 30, 60 and 90 days. Thirty days shows early adoption and barriers, 60 days whether habits are forming, and 90 days whether the behaviours have stuck.
Do I need a control group to evaluate training?
A comparison group is the strongest way to isolate the effect, but not the only one. Trend lines and confidence-adjusted estimates are acceptable when labelled clearly in the report.
Should every training programme be evaluated to ROI?
No. Match the depth to cost, strategic importance and reach. Compliance e-learning may need only knowledge checks and an audit; a large leadership programme may justify a full ROI study.
Can AI help with training evaluation?
Yes, for drafting survey items and test scenarios, grouping open comments and drafting reports. A person should check every theme and every number, and no AI tool should decide who passed or who is applying a behaviour.
How do I build my team's evaluation skills?
Pick one programme and run it through every level with the team, then turn the skills that proved hardest, such as isolating effects or writing scenario items, into individual development goals. Jobulary explains how a development plan built around specific skills works.