The implementation gap, quantified
Take every strategy the cohort approved, price what it promised, and audit what the operation verifiably delivered. The median answer is 34%. This essay is an accounting of the missing two-thirds.
The number needs defending before it needs explaining, because 34% sounds like polemic. It is arithmetic. We took 214 programmes, each carrying a written value case at approval and a frozen month-0 baseline, and audited them at month 24 under a published verification standard. We exclude self-reported outcomes, which are the figures most implementation statistics are built on, because when we audit those, 44% fail basic tests.
The other 66 points do not go where the anecdotes send them, and bad strategy accounts for far less of the loss than the war stories suggest, which is the finding that costs us the most argument in the room. When we decompose the gap, the strategy itself accounts for roughly a fifth of it, meaning the wrong market or the wrong economics. The rest is lost between the document and the operation, in four places we can now name and size.
The four leaks
The first leak is load, meaning too many concurrent commitments, which is the single strongest predictor of failure in the dataset. The second is translation. These are strategies that never became operating decisions, with no owner, no baseline and no date, so the operation politely continued as before. The third is latency, which is value that arrived so late its window had closed. The fourth is decay, where the result was delivered, celebrated and quietly reabsorbed within a year because nothing held it. Decay is the quietest of the four.
In median shares of the missing value: load 31%, translation 27%, latency 22%, decay 20%. We would not defend those shares to the decimal. Assigning a loss to one leak rather than another took judgement in a meaningful minority of cases, and a second reader would move the shares, though we doubt they would move the order. The pattern is what matters. Every leak is an operating-model failure rather than an analytical one. The strategy was usually right, and the organisation asked to carry it was usually never checked for load-bearing capacity.
Closing the gap is boring, which is why it works
The top quartile of the cohort delivers 71%, not 34%, and does nothing exotic to get there. Fewer commitments, drawn in phases. Every commitment translated into owners, baselines and dates before launch. A fixed cadence that reviews outcomes weekly. Independent verification before anything is called done. Each repair is dull. Together they compound into a doubling of delivered strategy.
The uncomfortable conclusion for strategy work — ours included — is that the marginal hour spent perfecting the answer is usually worth less than the marginal hour spent load-testing the organisation that must deliver it. It is an awkward thing for a firm that sells thinking to publish. It is also why we run implementation oversight at all.
Cite as: Markham Institute, “The implementation gap, quantified”, Markham Perspectives, April 2026. Republication permitted with attribution.
The Execution Notes are written by the Markham Institute from engagement evidence, reviewed before publication. Positions are argued, priced, and open to challenge.
Discuss this essay