Perspective · Berlin
Completion Is Not Readiness: What to Measure Instead
Sohrab Mostaghim · 27 July 2026 · 9 min read

Key claims
- Activity and outcome leave a missing middle layer: capability.
- Capability measurement must be observed, individual, under pressure, and leading.
- The budget-winning sentence is baseline, gap, intervention, movement, business translation.
Completion is not readiness.
Most ROI models quietly promise something they cannot deliver
If you have ever had to justify an enablement budget, you know the shape of the problem.
You have completion rates. You have satisfaction scores. You have attendance. And somewhere on the other side of the organisation there is a revenue number that moved, or did not move, for about fifteen reasons.
Between those two things is a gap, and every framework for measuring sales training ROI is an attempt to bridge it. Most of them bridge it by pretending.
I want to start with the pretending, because until that is named, nothing else you build on top of it will survive a serious conversation with a CFO. Completion is not readiness.
The attribution problem nobody wants to state
The standard advice is some version of this: establish a baseline, train the team, measure the change in win rate or quota attainment, attribute the delta to the training.
The problem is that in the same period, your territories were redrawn. Two competitors changed pricing. The product shipped a feature that unblocked a segment. Comp plans changed. Three strong reps left and four new ones ramped. The macro environment shifted. And the accounts were not randomly assigned in the first place, so your trained cohort and your control group were never comparable to begin with.
Sales is not a controlled environment and it never will be. Any model that claims to isolate the causal contribution of a training programme to a revenue number is producing a number that feels precise and is not.
This matters practically, not just philosophically. A CFO who has been in the job for fifteen years has seen a lot of confident attribution models. When you present one, you are not just making a claim about training. You are making a claim about your own rigour. If the number does not survive two follow-up questions, you have spent credibility you needed for the next conversation.
So the goal is not a cleaner attribution model. The goal is to measure something you can actually defend.
The layer that is missing
Here is the structural issue, and once you see it the path forward gets much clearer.
Almost every organisation measures at two layers.
Activity. Did people do the thing. Completion rates, attendance, hours logged, modules finished, certifications passed. This is what an LMS produces and it is easy to collect.
Outcome. Did the number move. Win rate, quota attainment, average deal size, cycle length, ramp time to full productivity.
And then, in the QBR, someone tries to draw a line directly from the first to the second. Ninety four percent completion, therefore the pipeline should improve. That line does not exist. It has never existed. Everyone in the room knows it does not exist, which is why the slide gets a polite nod and no follow up questions.
What is missing is the layer in between.
Capability. Can this specific person now do this specific thing, under conditions resembling the real ones.
Activity tells you someone was exposed to content. Outcome tells you what happened to the business. Capability is the only layer that tells you whether the exposure produced an ability, and it is the only layer that is both measurable and legitimately attributable to enablement.
This is not a semantic distinction. It changes what you can claim. You cannot honestly claim your training caused a revenue increase. You can absolutely claim that eleven named reps could not previously hold a discounting conversation without conceding on first push, that they now can, and that you have the evidence for both states.
That claim is smaller. It is also true, which makes it stronger.
What capability measurement has to satisfy
If you are going to build measurement at the capability layer, four properties determine whether it survives scrutiny.
It has to be observed, not self reported. Confidence surveys measure confidence. The correlation between how ready a rep feels and how ready they are is weak in both directions, and the two failure modes are different people. The overconfident rep inflates your dashboard. The quietly competent rep deflates it. Ask a rep to rate their qualification skills and you learn about their self perception, which is a real thing but not the thing you needed.
It has to resolve to the individual and the specific skill. A team capability score of seventy eight percent is operationally inert. Nobody can act on it. It averages a rep who cannot map stakeholders together with a rep who cannot handle pricing pressure, and produces a number that describes neither of them and prescribes nothing. Useful measurement names the person and the gap.
It has to be captured under pressure. Capability demonstrated in a supportive room is weak evidence of capability under a hostile CFO in week nine. The measurement conditions have to approximate the difficulty of the real conditions or you are measuring a different variable and calling it readiness.
It has to be leading, not lagging. This is the one that matters most for the budget conversation. A metric that tells you what happened after the quarter closed is a historical record. A metric that tells you which reps will struggle with committee dynamics before the committee meeting is an operational instrument. Same underlying data, completely different value, determined entirely by when you can see it.
Metrics that actually hold up
With those four properties as constraints, here are the measures worth building. None of them requires you to claim causal attribution to revenue.
Stage specific competency, per rep. Not sales skills as a single score. Capability broken down by where in the deal it applies: qualification rigour, stakeholder mapping, multi threading, pricing pressure, committee navigation, close planning. A rep is rarely uniformly weak. They are usually strong in three of these and consistently weak in one, and the one is what costs them deals.
Variance across the team, not just the average. This is underused and it is often the most revealing number you have. An average of seventy five with low variance is a team that performs consistently. An average of seventy five with high variance is two teams wearing the same badge, and it means your top performers are compensating for a structural gap in everyone else. Averages hide exactly the thing you most need to see.
Decay at thirty, sixty and ninety days. Measure capability immediately after an intervention and then again a month later, and a quarter later. This is where most programmes quietly fail, and almost nobody measures it, so almost nobody can defend against the accusation that their training does not stick. If you can show retention curves, you can also show which reinforcement mechanisms flatten them, which is a genuinely valuable finding.
Application rate in live deals. Does the framework actually appear in real opportunities. Not whether the MEDDPICC fields in the CRM are filled in, which measures compliance and can be gamed in four minutes. Whether the content of those fields reflects real qualification. An economic buyer field containing a name the rep has never spoken to is a filled field and an unqualified deal.
Time to first qualified opportunity, for new hires. The cleanest ramp metric available, because it is early, specific, and much less contaminated by market conditions than revenue based ramp metrics. It also correlates with things leadership already cares about, which makes it easy to introduce.
The sentence structure that wins budget
Everything above exists to let you construct one type of sentence. It is worth writing it out explicitly, because the structure is the point.
Baseline. Specific gap. Targeted intervention. Measured movement. Business translation.
In practice: At the start of the quarter, eleven of our forty reps consistently advanced deals without ever testing whether their champion had real authority. We know which eleven, because we watched them do it. We ran targeted coaching for that specific gap with those specific people. Six weeks later, nine of the eleven now test champion authority before stage three. Those nine carry forty one deals in current pipeline. Historically, deals where champion authority was never tested closed at roughly half the rate of deals where it was.
Notice what that sentence does and does not do.
It does not claim training generated revenue. It claims a specific capability gap existed, was closed for named people, and that the gap has a known historical relationship with outcomes. The final link is stated as a correlation from your own pipeline history, not as a causal promise.
A CFO can act on that. More importantly, a CFO cannot easily dismantle it, because every component is either directly observed or drawn from the company's own historical data. There is no borrowed benchmark from a vendor whitepaper anywhere in it.
Why this was not standard practice already
It is worth asking why, if capability measurement is this obviously better, almost nobody has it.
The answer is not that enablement leaders lacked the idea. Most of the people I talk to have wanted exactly this for years. The answer is that producing it required a human being on the other side of the table.
To observe whether a rep can navigate committee dynamics under pressure, historically someone had to play the committee. A manager running a live scenario. A senior rep playing the difficult CFO. That is the highest quality assessment available and it is also the least scalable thing in the entire commercial organisation, because it consumes the most expensive and most contested hours you have.
So capability assessment happened in bursts. Onboarding week. Kickoff. The top decile who earned executive coaching. And for the other forty eight weeks of the year, across the rest of the team, there was content and there was completion, because completion was the only thing that could be counted at scale.
The metric was never a failure of ambition. It was the honest residue of a real economic constraint. That constraint is the thing that has recently changed, and it is why this conversation is worth having now rather than five years ago.
Where to start
If you are going to build this, do not start by instrumenting everything. Start with one question.
Pick the stage where your deals most often stall. Look at your own closed lost data and find it, rather than assuming. For most enterprise teams it is qualification or stakeholder mapping, but yours may differ and the point is to know rather than guess.
Then measure capability at that one stage, per rep, observed rather than surveyed. You will almost certainly find that the distribution is wider than anyone expected, and that two or three named people account for a disproportionate share of the gap.
That finding, on its own, is more useful than an entire dashboard of completion rates. And it gives you the first sentence in the structure above, which is the one that changes the budget conversation.
See capability evidence from one full deal. No login.
Frequently asked questions
Why not attribute win-rate lifts directly to training?
In the same period territories change, pricing shifts, product ships, headcount moves, and cohorts are rarely comparable. A precise-looking attribution number often fails two follow-up questions from a CFO. Measure something you can defend instead: observed capability change for named people.
What is the capability layer between activity and outcome?
Capability answers whether a specific person can now do a specific thing under conditions resembling the real ones. Activity shows exposure. Outcome shows business results. Capability is the only layer both measurable and fairly attributable to enablement.
Where should a team start measuring first?
Pick the stage where your deals most often stall from your own closed-lost data. Measure capability there per rep, observed rather than surveyed. You will usually find a wide distribution, with two or three named people carrying a disproportionate share of the gap.