Roleplay Gym — AI-powered training

How to measure sales training ROI without fooling yourself

How to measure the return on sales and service training: what to baseline, which metrics survive scrutiny, and the attribution traps to avoid.

Updated · 4 min read

Most training ROI calculations are constructed after the fact by someone who needs the number to be positive. They tend to follow the same shape: revenue went up, training happened, therefore training caused revenue.

A finance director will dismantle that in one question. Here is a method that survives the question.

Decide what you are actually claiming

There are three defensible claims, in increasing order of difficulty:

  1. People can now do something they could not do before. Cheapest to prove, and often enough.
  2. They are doing it in real conversations. Harder, but observable.
  3. It changed a business outcome. Hardest, and the only one finance genuinely cares about.

Most programmes attempt claim three with evidence that barely supports claim one. Be explicit about which one you are making.

Baseline before you start, not after

This is the step that is almost always skipped, and its absence is what makes every subsequent number arguable.

Before any training, measure the current state of the specific behaviour you intend to change. Not general performance — the behaviour. If the goal is that reps stop conceding on price before exploring the objection, then measure, across a defined sample, how often they currently do that.

A structured practice assessment is a clean way to do this, because every person faces the same scenario and is scored against the same rubric. So is a coded sample of real calls, if you have the analyst time.

Without a baseline you are comparing an outcome to a memory.

Metrics that survive scrutiny

Ramp time. Days from start date to first independent productive conversation, or to a defined quota attainment. Concrete, already tracked in most organisations, and directly monetisable: a rep who ramps three weeks earlier produces three additional weeks of contribution.

Win rate at a specific stage. Not overall win rate, which moves for a dozen reasons. If the training targeted discovery, look at the conversion from first meeting to qualified opportunity.

Discount depth. If the training targeted price defence, average discount is a direct, hard number that responds quickly.

Repeat contact rate. In service, the proportion of contacts that generate another contact within a window. Sensitive to whether agents actually resolve things.

First contact resolution and average handling time together. Either alone is gameable. Together they are meaningful.

Attrition in the first ninety days. Frequently the largest hidden cost in a contact centre, and frequently responsive to whether new starters felt prepared.

The comparison that makes it credible

The single strongest thing you can do is run a control.

Split the population. Train one group, delay the other by a quarter. Measure both against the same baseline. This costs nothing except sequencing, and it removes almost every objection — seasonality, a pricing change, a new competitor, a strong quarter — because both groups experienced them.

If a genuine control is impossible, a staggered rollout across sites or regions gives you most of the benefit. Compare cohorts by start date rather than comparing before and after in aggregate.

Attribution traps

Regression to the mean. Training is usually commissioned when performance is bad, and bad periods are frequently followed by better ones regardless of intervention.

Selection effects. Volunteers improve more than conscripts, and they were going to.

Survivor bias. Measuring only people still in the role at month six removes exactly the people the programme failed.

Correlated investment. New training rarely arrives alone. If a new pitch, a new comp plan and new training land in the same quarter, no single one of them owns the result.

The enthusiasm window. Almost everything works for six weeks. Measure at three and six months, or you are measuring novelty.

A worked frame

Say a team of forty reps, average ramp of five months. Training and practice tooling reduce that to four.

One month earlier, forty times, at whatever a ramped rep contributes per month. That is the number. It is defensible because ramp time is already tracked, the comparison is against last year’s cohort, and it does not require anyone to believe that a workshop caused a good quarter.

Notice that this claim is narrow. Narrow and provable beats broad and disputed, particularly in a budget conversation.

Why practice data changes the arithmetic

The historical difficulty with training measurement is that the middle claim — are they doing it in real conversations? — has been expensive to evidence. It required someone to listen to calls and code them.

Structured practice produces that evidence as a by-product. Every rehearsal is scored against the same rubric, so the change in a behaviour over time is a measured line rather than an impression. Baseline, progress and the post-programme comparison come from the same instrument.

That does not prove the business outcome by itself. It does mean you can prove claims one and two cheaply, and reserve the expensive analysis for claim three.

If you want to see what a baseline assessment would show for your team, request a demo — it is usually the most useful first conversation.

Start with one conversation

Which conversation do you want to improve first?

We design a first scenario with you and show you the potential using your own use case.

No commitment. A first scenario tailored to your team.