The Problem With Measuring Productivity: What the Ratio Leaves Out

Every productivity number you have ever seen is a fraction. Output over input. Parts per shift, bushels per acre, billable hours per day, story points per sprint, revenue per employee. Measuring productivity is mostly the business of deciding what goes above that line, what goes below it, and what gets left out entirely — and the deciding is the part that gets forgotten. The finished number shows up looking like an observation. It isn’t one. It’s an engineered artifact, assembled from choices about numerators, denominators, and weights, and it deserves the same scrutiny you would give any other designed system.

The neighboring concepts all circle the same fragility: labor productivity (output per hour), total factor productivity (the residual left over once measured inputs are subtracted), Goodhart’s law (what happens when a measure becomes a target), the Solow residual, the time-and-motion study, throughput accounting. Each is an attempt to make the fraction behave. Each half-succeeds. The half that fails is the interesting half. If you design systems for a living — production lines, software teams, schedules, budgets — the fraction is one of your primary materials, whether or not anyone told you so.

A small team working together at laptops in an office — the raw scene that productivity metrics try to compress into a single ratio.
Most productivity metrics begin as an attempt to compress a scene like this into one number. The compression is where the trouble starts.

Productivity Is a Fraction Before It Is a Fact

Start with the arithmetic, because everything downstream follows from it. Productivity is a ratio: a quantity of output divided by a quantity of input. Someone chose the numerator. Someone chose the denominator. Someone decided what counts as output, what counts as cost, and what never gets counted at all. Those choices are load-bearing. Change the numerator and the same shop floor can look brilliant or stagnant between one afternoon and the next. Change the denominator and a “30 percent improvement” can dissolve into a rounding artifact.

I spent a morning once beside a machinist running a short production job. Cutting fluid in the air, aluminum chips curling off the tool, a parts counter ticking upward on the control panel. The counter counted good parts. The scrap bin appeared in no column. The twelve minutes he took mid-shift to re-dress a grinding wheel appeared in no column either, though it was the reason every part after lunch held tolerance. End of day: 412 units, eight hours, 51.5 units per hour. Accurate. Also mostly silent about how the 412 came to exist.

The numerator is where the first trouble lives. Outputs are heterogeneous. You cannot add an appendectomy, a consultation, and twenty minutes spent reassuring a frightened patient, so statisticians convert everything to money — value added — before summing. But prices are not a neutral scale; they carry bargaining power, subsidy, and convention inside them. When an economy-wide productivity figure rises, part of the rise is real output, part is a shift in relative prices, and part is an accounting decision about which industries entered the sample. The published number hands you the sum. It does not sort the parts.

The denominator is where the second trouble lives. Hours are countable, which is exactly their appeal and exactly their weakness. An hour is a unit of time, not a unit of attention, skill, or effort. The first hour of a shift and the eleventh go into the same column. A rested surgeon and a surgeon six hours past fatigue each contribute one doctor-hour. The denominator asserts an equivalence that the work itself does not have.

What You Add Is What You Believe

Any productivity figure that covers more than one kind of work has an aggregation problem, and aggregation is where the quiet distortion gets in. Dissimilar outputs have to be weighted before they can be summed, and the weights are beliefs. Index-number theory — the machinery underneath every official series — is a long, mostly civil argument about which year’s prices to use as weights, with named positions (Laspeyres, Paasche) that produce different answers by construction. So a single-point claim like “productivity up 2.1 percent” should be heard as “up 2.1 percent, given one coherent set of choices.” The point estimate is the visible tip of a stack of decisions.

The deepest version of the argument is old. During the Cambridge capital controversy, economists in England and Massachusetts spent two decades disputing whether “capital” can be measured independently of the rate of return it is supposed to explain. The answer that survived: not cleanly. Aggregating capital requires prices; pricing capital goods requires, among other things, a rate of return. The measure leans on the very thing it measures. A firm’s internal accounting inherits a smaller version of that loop every time it prices a machine, a license, or a person’s hour.

Total factor productivity was built partly to manage this. It is the residual: the growth in output left over after the growth of measured inputs — labor hours, capital services — has been subtracted. Moses Abramovitz, who did as much as anyone to establish the method in the 1950s, called the residual “a measure of our ignorance.” Still the most honest sentence in the field. When productivity is said to have risen 1.8 percent, a meaningful share of that 1.8 is whatever the accounting could not see: quality change, organizational learning, mismeasured inputs, plain luck.

Robert Solow’s 1987 remark that “you can see the computer age everywhere but in the productivity statistics” became famous because it named the gap from the side everyone could feel. The gap did eventually narrow — partly because the statistics improved, partly because businesses reorganized around the machines with a lag measured in years — and that history is its own lesson: a measurement system can be structurally late to exactly the change it was asked to register. The statistical agencies, for their part, are candid about the construction. The Bureau of Labor Statistics publishes its methods precisely because the ratios are built rather than found: deflators for comparing dollars across years, adjustments for quality change in everything from machine tools to laptops, revisions that arrive years later when better data replaces worse.

What the Denominator Hides

Hours measure presence, not contribution, and the difference accumulates in the dark. The fraction rewards what it can see, and it sees very little. It does not see deferred maintenance, which is work moved into a future that has no column yet. It does not see the senior operator who spends forty minutes showing a new hire why the fixture goes on that way, because those forty minutes produce zero units while they happen. And it does not see the hour of thinking that prevents a month of rework. Every organization runs on a shadow ledger of exactly this kind, and every output-per-hour metric quietly declares that ledger to hold a value of zero.

The stopwatch arrived with Frederick Winslow Taylor; Frank and Lillian Gilbreth followed with therbligs, seventeen named elementary motions. Both were sincere attempts to make the denominator see. Both failed in an instructive way: the more precisely the hour was subdivided, the more of the work fell between the subdivisions. What the time study recorded, it recorded well. What it could not name, it erased — and the erasure was then mistaken for efficiency.

There is also a small piece of queueing theory that deserves to be painted on the wall of every planning meeting: L = λW. Average work in process equals throughput multiplied by cycle time. The uncomfortable corollary concerns utilization. As a busy system approaches its capacity ceiling, waiting time does not rise politely. It climbs steeply, then goes effectively vertical. A plant running every station at 95 percent utilization has not maximized output. It has maximized inventory and queueing — the pallets stacking quietly beside the station with the green efficiency score, each pallet a bet that the downstream constraint will someday catch up.

When the Measure Becomes the Target

Goodhart’s law is the one everyone half-remembers, and the mechanics deserve a precise statement. Charles Goodhart’s original point concerned monetary aggregates; Marilyn Strathern’s compressed version is the one that travels: when a measure becomes a target, it ceases to be a good measure. The mechanism fits in a sentence. Under pressure, optimization flows into whatever dimension of the proxy is cheapest to move — not into the dimension the proxy was meant to represent.

Colleagues reviewing figures on a screen during a meeting — a metric displayed in public changes the behavior it was meant to record.
The moment a number goes up on a wall, the wall starts changing the number.

The software world keeps relearning this in public. Lines of code measured quality once, briefly, until the pressure arrived and the pressure won. Story points were invented as a coarse sizing aid; a point is a promise about the future, and promises inflate under pressure, so velocity — meant as a private capacity signal — becomes a negotiated figure within a few quarters of being reported upward. Billable time in six-minute increments teaches a subtler curriculum: the increment trains people to find work that fits the increment. The proxy does not break loudly. It detaches quietly and keeps ticking.

Manufacturing offers the most physical version. Eliyahu Goldratt’s central claim, delivered in the Theory of Constraints, is arithmetic wearing a novel’s clothes: the throughput of a whole system is set by its constraint, and efficiency everywhere else mostly manufactures inventory. A station with a perfect utilization score, running ahead of a bottleneck, is not diligent. It is converting tomorrow’s capacity into today’s pallets, which someone will count, at a cost someone will pay, at a moment the metric will not record.

How to Measure Without Lying to Yourself

You will not escape measurement, and you probably should not want to. Measurement is how a system talks to itself. The workable discipline is to treat each metric as a designed object with a maintenance schedule, not as a fact with a dashboard. In practice, that looks like this:

  • Measure the constraint, not everything. One well-instrumented bottleneck tells you more about system output than forty station-level efficiency scores, because only the constraint’s rate is the system’s rate.
  • Pair every rate with a counter-metric. Throughput pairs with cycle time and rework; utilization pairs with queue length; billable hours pair with write-offs and client retention. A rate that improves while its counter-metric worsens is not improving.
  • Audit the numerator once a year. Ask what the ratio currently refuses to see — maintenance, mentoring, quality, debt of any kind — and write the answer down. The omissions drift as the work drifts.
  • Apply denominator skepticism to percentage claims. “Thirty percent more productive” means thirty percent of what, divided by which hours, weighted by whose prices? Most impressive claims soften under those three questions. A few collapse.
  • Keep a deliberate margin of unmeasured time — or measure it honestly. If maintenance and teaching are real work, put them in the books as stocks instead of pretending the shadow ledger is empty. What gets measured gets managed; what gets hidden gets spent.
Two professionals reviewing documents together at a desk — the setting where measured work and unmeasured work meet.
The hour spent explaining, checking, or preventing rarely appears in the same column as the hour that produces.

FAQ: The Problem With Measuring Productivity

Why is productivity so hard to measure in knowledge work?

Because knowledge work has no countable unit of output. A machinist produces parts; a designer produces decisions, and decisions cannot be summed without weights, and the weights are judgments. So organizations reach for proxies — hours, tickets, story points, messages sent — and proxies fail in a specific way: once one comes under pressure, effort flows into the cheapest dimension of the proxy rather than into the underlying work. The problem is not that knowledge workers resist measurement. It is that nobody has yet found a numerator that behaves under pressure.

What is Goodhart’s law, and how does it apply to productivity metrics?

In Marilyn Strathern’s compact formulation, Goodhart’s law states that when a measure becomes a target, it ceases to be a good measure. Applied to productivity: any metric attached to rewards or penalties will be optimized, and it will be optimized along its easiest dimension — inflated story points, hours that grow to fill their increments, stations running ahead of the constraint. The metric keeps moving. It stops meaning what it meant.

What is the difference between labor productivity and total factor productivity?

Labor productivity divides output by hours worked (or headcount) alone; it is blunt, easy to compute, and blends in everything — capital, technique, organization — that helps an hour produce more. Total factor productivity subtracts the growth of all measured inputs from output growth and reports what remains: the residual. Moses Abramovitz called that residual “a measure of our ignorance,” which is the honest reading — its size often tells you how much you failed to measure.

Can a productivity metric ever be trusted?

Conditionally. A metric earns provisional trust when its numerator is stable and countable, its denominator is honest about what an hour contains, it is paired with counter-metrics able to veto it, and it is used as a thermometer rather than a thermostat — read for information, not wired to rewards. Remove any one of those conditions and the trust should be withdrawn. Trust in a metric is situational, never intrinsic.

Where This Leaves Us

The tension does not resolve, and I am not going to pretend otherwise. An organization that refuses to measure will drift. An organization that measures badly will drift faster, with better paperwork. The way through is not fewer numbers but numbers that know what they are — fractions with their construction visible, paired against the things they omit, kept away from the rewards that corrupt them. James C. Scott called the underlying move legibility: what an institution can see of the work it governs is always thinner than the work itself, and the thinning is where both efficiency and quiet loss live.

This piece opens a standing column on designed measures — metrics treated as engineered objects and taken apart down to the numerator and the denominator. Next in the series: throughput accounting, and what it means to price the constraint instead of the worker. After that, the stopwatch itself — how time study arrived, what it solved, and what it quietly decided not to see. If there is one number your organization treats as scripture, send it to me. Taking apart readers’ sacred metrics will be a recurring feature here, and the first submissions get disassembled in full.