The Number Did Not Start Lying. You Kept Asking It the Old Question.
The Number Did Not Start Lying. You Kept Asking It the Old Question.
At one point in my career I sat in a calibration meeting where the policy required someone in every group to land below expectations. The ranking curve needed a bottom, and somebody had to be assigned to it. My group had just closed out double-digit growth for the period. Other groups in the division had lost money that period on market conditions rather than effort, and the position in the room was that fairness meant someone from my group should carry the same low rating anyway.
I refused it at the group level, with real numbers tying the team's work directly to business outcomes. The person running the room was not moved by that, and pushed the request down a level: name the individuals inside the group who did less than the others. I sat with what refusing would cost before I knew which way it would go. If I lost, I was not going to carry that message back to people I had spent the year telling they were doing well, which meant leaving, and leaving meant not being there to protect them afterward.
Here is what was actually happening in that room, and it took me a while to name it. A rating system is built to answer one question: did this person do the job well. A forced distribution answers a different question entirely, which is how to keep company-wide compensation inside a budget line. Run the second question through the first system's vocabulary without naming the substitution, and a budget decision arrives at an individual looking like a verdict on their work. The instrument was not broken and it was not lying; it was asked something it was never built to answer, and it answered in the only language it had.
Your engineering dashboard is doing the same thing right now, and nobody involved is being cruel about it. Everyone already agrees that metrics can mislead, and that agreement is the problem, because it is where the conversation stops. Every technical leader reading this has sat through a review where the proposed fix was a better metric: move to DORA, add SPACE, instrument the thing you actually care about. The debate is always which number to adopt. Nobody checks whether a number that was genuinely valid has quietly stopped being valid.
Feature branch throughput rose across the industry this year, and that number is honest: not gamed, not vanity. It accurately reports more of something that is genuinely happening, because generation became cheap. What broke is the inference sitting on top of it: the number earned its place on the dashboard when producing code was the expensive step, so producing more of it correlated with delivering more. That correlation was real. It is not real now, and the number reports exactly as it always did.
Nothing announces this. Dashboards carry thresholds for values and no alarm for a metric whose meaning expired, so the figure keeps arriving in the same slot, in the same green, with all the authority it earned under conditions that no longer hold. Meanwhile the constraint moved downstream into integration, review, and production behavior. Main branch throughput fell over the same period. That number is equally honest, and it reaches nobody, because it was never the headline.
Then every downstream decision gets priced off the number that does reach executives. Headcount, dates, and external commitments are set against a figure measuring the step that stopped being expensive, so the organization commits to volume it can generate and cannot integrate. Months later the dates are missed. The miss reads as an execution failure, and in most rooms I have been in it lands on the team, who were handed a commitment priced two quarters earlier off an inference nobody had checked. In the LeadershipOS™ Stack I built, that is a Cadence failure arriving as a Coaching problem, which is why fixing it where it lands never works.
The leaders who catch this are not better at metrics than you, and they are not running cleaner dashboards. They stopped reviewing the numbers and started reviewing the sentences underneath them.
Every proxy encodes an assumption about what is scarce, and that assumption is almost never written down, because on the day the metric was chosen it was too obvious to state. So write it down: this number matters because producing code is the expensive step. Then put that sentence on the review schedule instead of the number. Numbers get reviewed constantly, and the assumptions holding them up almost never do.
Two things will pull you off this. The first is the reflex to add a second metric alongside the first, which is the conventional fix returning in a new costume, because two numbers with no stated assumptions produce the same failure twice. The second is timing. This argument has to be made while the dashboard is green, with no incident to point at, against a number that is currently making everyone in the room look good. Wait for a missed date and you are no longer making an argument, you are explaining a failure that has already landed on your team.
Take your last missed date. Find the number that was used to commit to it, write the assumption that number rests on in one sentence beginning with this matters because, and ask whether the assumption was still true on the day the commitment was made. Then run the same sentence against whatever number is rising fastest right now, while nothing is on fire.
The LeadershipOS™ Scorecard finds which layer is actually carrying a failure like this, which is almost never the layer taking the blame for it: https://theleadershiposbook.com/scorecard
I write about structural leadership for technical leaders in high-stakes operating environments. If you want to see where your system is load-bearing on you personally, the LeadershipOS™ Scorecard maps it: https://theleadershiposbook.com/scorecard
