It Worked. Nobody Had Decided What Happens When It Can't Keep Up.
It Worked. Nobody Had Decided What Happens When It Can't Keep Up.
At a company I worked for, every patient's incoming data ran through one stored procedure. Not a simple one: a sprawling tangle of sub-procedures, each one updating the same core tables that nurses read from on the other side to pull up a patient's information. It worked. It had always worked. I was stuck, because I could see exactly how complex it was and had no clear way to safely change it.
Each call in was wrapped in a transaction, so any single instance was safe on its own. The real risk sat one level up, in the code receiving the incoming data and calling into that procedure. Under enough volume, that code could hit a breaking point, and at that point a call might simply not fire. Nothing was built to catch it if it didn't: no queue, no retry, no alert. It was the linchpin of the whole system, and the gap was known: real tech debt, tracked and deferred rather than ignored, with no plan yet in place for the moment it could no longer keep up.
"It worked" was true, and it was also the wrong thing to be reassured by. It described the load the system was carrying that day. It said nothing about what would actually happen to the data still arriving once that breaking point was hit. The risk was known and deferred; what still had not been decided was the specific behavior at the moment the deferral ran out. The system had decided that part by default: silently, whenever volume outran what the code could handle.
Conventional wisdom says build the simplest thing that works, and do not spend engineering effort on queuing, retries, or backpressure until you actually hit the ceiling. Premature architecture is waste. The proc worked, it did its job, so touching the pipeline around it before it broke would have looked like solving a problem nobody had yet. Every technical leader who has kept shipping against a system everyone privately calls fragile has made some version of that same trade.
Building the simplest thing that works does not eliminate the decision of what happens when volume outruns capacity. It defers it. The decision does not disappear when nobody makes it; it gets decided implicitly by whatever the code already does when it hits that ceiling. At that company, nothing had been built to catch a call that did not fire, so the decision that was made by default was that the data would simply go missing, unnoticed, exactly when the system was under the most pressure to be reliable.
The fix is not free, and pretending otherwise would be dishonest. Queuing and retrying calls once the breaking point hits guarantees the data eventually lands, and it costs freshness: a nurse's read might reflect data still sitting in the queue, not yet applied, for however long the backlog takes to clear. That is what I call a Structural Bearing™: a tension between two legitimate needs, durability and freshness, that does not resolve, it only gets managed.
American Express runs into a version of the same tension in its payment infrastructure, and resolves it differently: a transaction that cannot be confirmed with confidence gets rejected outright, on purpose, rather than risking a consistency it cannot establish. We landed on a different answer to the same underlying question, queue and retry rather than reject. Amex made that decision in advance, as a designed part of the architecture. Ours came later, once someone finally treated the breaking point as a decision to make rather than a risk to quietly hope never landed.
Most leaders read "it's working" as evidence the shape is fine. That reading is the actual decision being made, quietly, every day the system holds. "It's working" is a fact about today's load. It says nothing about what happens to the data still arriving once that load changes, and treating it as reassurance instead of an open question is how the decision gets skipped instead of made.
The leaders who get this right are not the ones who avoided the trade. They stopped reading "it hasn't broken yet" as an answer, and started treating it as the one question still sitting open. The system does not need "the correct architecture" as a checklist item. It needs someone who names the actual trade, durability against freshness, or whatever the two real costs are, as a permanent tension rather than a problem to solve, and then decides on purpose which side to weight, before the wall arrives. Staying undecided is also a decision, the worst one, because it is the only version nobody ever actually considered.
Name the system in your stack you would describe as working fine. Now ask: if it hit its breaking point tomorrow, would the failure be visible, designed, and contained, or would it just quietly not happen, with nothing built to notice? If you do not know the answer, nobody has decided it yet. The system has.
The Edge Case walks through how to name a Structural Bearing™ precisely enough to decide which side to weight, instead of pretending the tension will resolve on its own: http://TheEdgeCaseBook.com
