Part IV — Applying Organizational Flow · Chapter 55 · 6 min read · First Public Draft
Measuring Value
The fifth measure, and how to reach it from where most organizations actually stand.
Jump to a chapter
Narrated with Microsoft's neural voice, not me — a recording.
Every organization I have worked with could tell me what it delivered last quarter. In some cases to two decimal places, with a chart.
Almost none could tell me what any of it changed.
That is not a criticism, or not only. It is what happens when one number is available on a Thursday and the other one isn’t available for a year, and something has to go in the report on Friday. Delivery gets called value because delivery is the number that exists.
The four indicators in the previous chapter all measure movement — who owns it, how often it changes hands, how long it waits, whether anything came back. An organization can improve all four, genuinely and measurably, and produce more of something nobody wanted, faster than before. That is why there is a fifth, and why it sits apart from the others: it is the only one that points outward.
The substitution, and how to spot it
The substitution is rarely a decision. It happens in the wording.
Somebody writes delivered the new onboarding flow under a heading that says Value. Somebody else reports 90% of the migration complete as a benefit. A slide says value delivered this quarter and then lists things that were built. Nobody is lying. Everyone would agree, if asked directly, that shipping a thing is not the same as the thing being useful. The heading just quietly does the work of the argument.
The reliable way to catch it is to read your own reporting and ask, of each line, whether it describes something that happened inside the building or outside it. Almost everything will be inside. That proportion is the honest starting position, and it is worth knowing before trying to change it, because it tends to be more lopsided than people expect.
Four rungs
What follows is not a maturity model, and I would rather it were not drawn as a pyramid. It is just the order I have seen organizations actually move in, and each rung is a real place to stand for a while.
We delivered it. Countable immediately, entirely under your control, and it tells you nothing about effect. Most reporting lives here. The only thing wrong with it is the label — called output, it is useful; called value, it stops anyone looking further.
Somebody used it. Adoption, usage, the thing being opened. This is a genuine step, because it is the first number that requires another person to do something. It is also where a lot of organizations settle permanently, since it is countable from your own systems and feels enough like evidence. It isn’t: people use things they dislike, daily, because it is the only way to get paid.
Somebody outside says something is different. The first measure that leaves the building. It is late, it is fuzzy, it will not attribute cleanly to one team’s effort, and it is the first one that could actually contradict a decision you made.
Something is decided differently because of it. The rung that matters. A measure that cannot change a decision is decoration, however rigorous it looks. If the answer arrives and the roadmap is identical either way, you have built a reporting obligation, not a feedback loop.
Most organizations are trying to jump from the first rung to the fourth, usually by way of a large dashboard, and stall. Moving one rung is enough for a year.
The smallest version that works
If there is one practice in this paper I would keep when everything else was cut, it is this one, and it costs nothing.
Before the work starts, write one sentence: we think this will let [these people] [do this thing] better, and we expect to see it by [when]. Not a business case. One sentence, written where it can be found again, by the people choosing the work rather than by anyone in a planning function.
Then, at the date, go and look.
That is the whole practice. It is a smaller cousin of what Melissa Perri calls the product kata, and I recommend the original for anyone who wants it done properly. What makes even this stripped-down version work is not the measurement, which is usually rough. It is that somebody committed to an expectation in writing, in advance, which is the one condition under which being wrong is informative rather than embarrassing.
The organizations I have watched fail at this mostly failed at the second half. Writing the sentence is easy and mildly enjoyable. Going back seven months later, when everyone has moved on and the answer can only be awkward, is a habit somebody has to protect.
Asking, when you cannot count
Value often cannot be counted honestly. It can nearly always be asked about, and the asking is worth more than a weak number that looks strong.
The same discipline applies as with the four indicators. Ask about one concrete instance rather than the general case — what did you do differently last week because of it produces evidence, has it been valuable produces politeness. Ask the person who experiences the outcome, not the person who commissioned the work, because the commissioner’s answer is partly about their own judgement and everyone knows it. And ask often enough that no single conversation carries too much weight.
Five real conversations can reveal something a survey with four hundred responses and a mean of 3.8 never will. The survey is more defensible in a meeting. It is much less likely to change anyone’s mind, which by the fourth rung is the only thing a measure is for.
Two failures worth naming
The first is the proxy. Creating Value covers that problem in detail. The practical safeguard here is simple: keep a direct channel to the people experiencing the outcome.
The second is attribution. A great deal of energy gets spent trying to attribute an outcome precisely to one team’s work, and much of it is wasted, because the honest answer is that several things happened at once and the effect belongs to all of them. Measure the outcome where it is actually visible — at the level of the service, the customer, the process — and accept shared credit. An organization that insists on attributing value to individual teams will end up measuring only the things small enough to attribute, which are reliably the things that matter least.
What to expect
The first reading is worth very little. There is no benchmark for any of this, nothing has been validated across enough organizations for anyone to say what a good number looks like, and the value of the measure is entirely in the comparison against your own earlier one. The fourth reading is worth a great deal.
Expect the early answers to be disappointing, and expect that to be the useful part. The first time an organization genuinely asks whether last year’s work changed anything, the answer is usually less than we said at the time. That is not a sign the measurement is broken. It is the first accurate thing the reporting has produced in a while, and everything worth doing next follows from it.
Editor's Notes
Outcome over output is old ground and this chapter claims none of it. Melissa Perri's *Escaping the Build Trap* is the most useful modern statement of it, and the habit of writing down what you expect before the work starts and returning to it afterwards is hers — she calls it the product kata, and it is a discipline rather than a template. Anyone who wants the practice properly should go there. Outcomes and Key Results have been arguing the same case for longer, with mixed results, mostly because the mechanism gets adopted and the honesty does not.
The four friction indicators in Observing Organizations are diagnostic prompts rather than validated metrics, and this fifth one is weaker still — value resists definition in a way that waiting time does not. What this chapter contributes is the arrangement: four measures pointing inward at movement, one pointing outward at effect, and the observation that an organization can be excellent at the first four and learn nothing, because nothing in them can tell you whether the thing that moved was worth moving.