Organizations spend a huge amount of money training their employees on both technical and behavioral competencies. Every year, new tools are designed to try to cater to individual learning styles and make training more effective. But underneath all that spending sits a simple, uncomfortable question: how does anyone actually know if it worked?
Donald Kirkpatrick, professor emeritus at the University of Wisconsin, began working on this problem early in his career. His original work was published in 1959 in a journal of the American Society of Training Directors, and it went on to become the most widely used evaluation framework in the training world — simple enough to apply in almost any organization, yet complete enough to catch problems that simpler methods miss.
Kirkpatrick laid out four levels for evaluating any training program:
- Reaction — the thoughts and feelings of participants about the training
- Learning — the increase in knowledge or understanding as a result of the training
- Behavior — the extent of change in behavior, attitude, or capability
- Results — the effect on the company’s bottom line as a result of the training
A fifth level — Return on Investment (ROI) — has since been added by later practitioners to extend the model, though it was not part of Kirkpatrick’s original four.
The Four Levels at a Glance
The core idea behind the model is that each level builds on the one below it. Evaluating behavior change is close to meaningless if you haven’t first confirmed that learning actually happened — and evaluating learning tells you little if participants were too disengaged to absorb anything in the first place. Each level can only really be trusted once the level below it has been checked.
| Level | Core Question | What It Measures |
|---|---|---|
| 1. Reaction | Was the environment suitable? | How favorably participants responded to the training experience itself |
| 2. Learning | Did they learn anything? | The actual change in attitude, skill, or knowledge (ASK) |
| 3. Behavior | Are they using it on the job? | Whether learning actually transferred to on-the-job behavior |
| 4. Results | Was it worth it? | The measurable effect on business outcomes |
Level 1: Reaction
Reaction measures how favorably participants responded to the training — not whether they learned anything, just whether the experience itself landed well. This evaluation is primarily quantitative and serves as direct feedback to the trainer and the training design.
The most common collection tool is a post-session questionnaire (often called a “smile sheet”) analyzing the content, methodology, facilities, and overall course quality. Typical questions ask participants to rate the relevance of the material, the trainer’s clarity, the pacing of the session, and whether the venue and logistics supported learning.
Worth remembering: a room full of participants who enjoyed the session tells you the experience was pleasant — it doesn’t yet tell you whether anyone learned anything. Positive reactions correlate weakly with actual learning, which is exactly why the model doesn’t stop here.
Level 2: Learning
At the level of learning, evaluation shifts to measuring the actual change in the ASK (Attitude, Skill, and Knowledge) of the trainees. This involves observation and analysis of what participants say, how they behave, and what they produce during and after the session.
Common tools include:
- Pre- and post-tests — comparing knowledge before and after the session to isolate what the training itself contributed
- Interviews and surveys — gathering qualitative detail that a test alone won’t capture
- Skills demonstrations — having participants perform the task, not just describe it
For example, a sales training program might give trainees a product-knowledge quiz before the session and the same quiz afterward. A meaningful score improvement is decent evidence that learning happened — though, as Level 3 makes clear, it’s no guarantee that the new knowledge ever makes it back to the sales floor.
Level 3: Behavior
Behavior evaluation analyzes whether learning actually transferred from the training session to the workplace — arguably the level that matters most, and the one most organizations are worst at measuring. The primary tool here is observation, often combined with questionnaires and 360-degree feedback from managers, peers, and direct reports.
This evaluation typically happens weeks or months after the training, once participants have had a real chance to apply — or fail to apply — what they learned. A customer service training program, for instance, might be evaluated by reviewing call recordings or customer satisfaction scores three months after the session, rather than immediately afterward when enthusiasm is naturally at its highest.
Behavior change can be blocked even when learning genuinely happened — a supportive environment, management buy-in, and the opportunity to actually practice the new skill all matter as much as the training content itself.
Level 4: Results
The results stage evaluates the training’s effect on the organization’s bottom line. What counts as a “result” depends entirely on the goal of the specific training program — revenue, error rates, retention, safety incidents, customer satisfaction, and productivity are all common measures depending on what the training was meant to fix.
The most rigorous version of this evaluation uses a control group: comparing outcomes for employees who received the training against a comparable group who didn’t, over a defined period, to isolate the training’s actual effect from other factors that might be moving the numbers at the same time.
The Fifth Level: Return on Investment
Because organizations ultimately care about spending, a fifth level — ROI — has been layered onto the original four to put a dollar figure on the results. This level asks the most direct question of all: did the financial benefit of the training exceed what it cost to run? It’s the level furthest removed from the training room and the hardest to measure cleanly, which is part of why relatively few organizations get this far.
A Worked Example: Evaluating a Customer Service Training Program
Seeing all five levels applied to one program makes the logic easier to follow than reading about them in the abstract.
| Level | What Gets Measured in This Example |
|---|---|
| 1. Reaction | Post-session survey: participants rate the trainer, materials, and pacing immediately after the workshop. |
| 2. Learning | A scenario-based quiz on handling difficult calls, taken before and after the session, shows a 40% score improvement. |
| 3. Behavior | Three months later, call recordings and supervisor observations check whether reps are actually using the new de-escalation techniques on real calls. |
| 4. Results | Customer satisfaction scores and complaint volume for trained reps are compared against a control group of reps who haven’t yet been trained. |
| 5. ROI | The reduction in complaint-driven refunds and escalations is converted into a dollar figure and weighed against the cost of running the program. |
Why Most Evaluations Never Get Past Level 1
In practice, most training evaluation stops at reaction data — the smile sheet gets collected, and that’s the end of it. Fewer organizations collect learning data, still fewer measure behavior change, and very few make it all the way to business results or ROI.
The reasons are practical rather than a lack of interest:
- Reaction data is cheap and immediate; results data takes months and a control group.
- Behavior change is hard to observe without disrupting the very work being observed.
- Isolating a training program’s effect on revenue or retention from every other variable moving at the same time is genuinely difficult.
Organizations that do reach the higher levels — evaluating training against actual business results — tend to have a deliberate measurement process built in from the start, not something added on after the fact.
Strengths and Limitations
The model’s popularity comes from being simple, flexible, and complete enough to apply to almost any kind of training — technical or behavioral, one-off workshops or year-long programs. It also builds in a natural discipline: each level forces you to ask a harder question than the one before it.
Its main limitation is the same thing that makes it approachable: the model describes what to measure at each level, but not how to isolate the training’s specific contribution from everything else happening in the organization at the same time. That’s a genuinely hard measurement problem, and the model doesn’t solve it for you — it just makes clear which level you’re stopping at when you don’t.
FAQs
-
Do you have to evaluate all five levels every time?
No. Most organizations scale the depth of evaluation to the cost and risk of the training itself. A short compliance refresher might only need Level 1 and 2; a major leadership development program justifies going all the way to Results or ROI.
-
Which level is most commonly skipped?
Level 3 (Behavior) and Level 4 (Results) are skipped most often, mainly because they require follow-up weeks or months after the session ends, by which point attention has usually moved to the next training cycle.
-
Can Level 1 reaction scores predict Level 2 learning?
Only weakly. A well-liked session doesn’t guarantee real learning, and a session participants rate poorly can sometimes still produce a meaningful knowledge gain — which is exactly why the model treats the two as separate levels rather than one.
-
Who added the fifth ROI level, and why isn’t it part of the original model?
Kirkpatrick’s original 1959 framework covered only the first four levels. The ROI level was added later by other practitioners in the training field, in response to organizations wanting a direct financial answer to whether a program was worth its cost — a question the original four levels don’t answer on their own.
-
Is Kirkpatrick’s model still relevant given how much corporate training has changed?
Yes — the four questions it asks (Did they like it? Did they learn it? Are they using it? Did it work?) apply just as well to an e-learning module or a virtual workshop as they did to an in-person session in 1959. What has changed is the tooling available to answer each question, not the underlying logic of the levels themselves.


