It Took Half an Hour. My Pride Was Not the Proof It Worked.
It Took Half an Hour. My Pride Was Not the Proof It Worked.
Every month, my LeadershipOS™ Inner Circle page shows bullets pulled straight from that month's issue. When the month rolls over before I have written the next one, the page needs to fall back to something else automatically, not show stale bullets, not break. Check that page today and you will see exactly that: generic bullets, because I have not finished November's issue yet. In a few weeks, once I have, they will update to the specific ones automatically, the same way they always do.
I asked Claude Code to build that whole capability: the extraction, the fallback, the edge cases. It wrote its own tests, month rollover and DST handling included. It took under half an hour.
In years past, I would have spent days on this myself. I am a software developer at heart, and I know exactly how tempting that rabbit hole is: open the editor, start writing the logic myself, tell myself I am the one who should verify it by building it. I did not fall into it this time. What I felt instead was pride, real pride, at what had actually been built.
That pride was not the evidence the capability worked. I reviewed the code. I checked that the edge-case tests Claude wrote actually covered the case I cared about, a month rolling over with no new issue written yet, and that they passed. The pride came after the check, not instead of it. Without that automation, I would have shipped static bullets, and static bullets would not have given potential readers a true view of what my writing is actually worth to them.
The industry's answer to "is AI making us more productive" is a satisfaction survey dressed up as a productivity metric: ask developers how the tool feels, average the yeses, publish the percentage, and call it proof. Almost every technical leader who has ever reported an AI adoption number upward has done some version of that, because it is the fastest data available and it almost always looks good.
Asking whether a tool felt like it worked produces a real answer. It is an answer about feeling, not about whether the output actually held up. My own pride was trustworthy because something concrete backed it: the code I reviewed, the tests that passed against the exact edge case I cared about. A satisfaction survey skips that step entirely. It asks for the feeling and stops, with no requirement that anyone checked the work underneath it.
The sample is rigged before the first question is even asked. It only reaches people still using the tool. Anyone who tried it, hit a real edge case, and quietly went back to building it by hand, the exact rabbit hole I almost fell into, is no longer in the room to answer. The most common AI productivity metric, a developer survey, systematically excludes the one group whose answer would matter most: the people who tried it and stopped.
The number that comes back is not "does this tool work." It is "how do the people who never had a reason to stop feel about it," and leadership, boards, and budget committees treat that number as if it answered the first question instead of the second. A budget gets set. Headcount plans get made. All of it inherits a number that was never actually measuring what it claims to measure.
The leaders who get this right are not the ones who trust AI less. They stopped accepting a feeling as evidence, and started asking for the specific thing that was checked before they signed off on the number.
Fixing this takes two separate moves, because there are two separate problems baked into the number. Find the people who tried the tool and stopped, not just the people still using it; that is a harder, more deliberate search, because someone who quietly went back to the old way rarely raises a hand. Require an artifact before trusting any adoption claim, too: not "did this feel fast," but the specific test that passed, the edge case that held, the output someone actually reviewed.
Pull up the last AI adoption number your organization reported upward. Can you point to one specific thing that was checked to produce it, a test, a review, an output someone verified, or is the number just an average of how people say they feel? If you cannot name the artifact, you do not have a productivity number. You have a satisfaction score wearing one.
You do not have to take my word for any of this, either. Visit the page yourself: right now, before November's issue ships, it is showing the generic fallback, exactly as designed. That is the artifact. Not my pride about it.
The LeadershipOS™ Scorecard diagnoses whether your system actually holds up, not how good it feels to the people still using it: https://theleadershiposbook.com/scorecard
