Skip to main content
Not Yet
Main Series · The Negative Space · TAM_105

Not Yet

Why the Same Tool Gets Two Answers

In a hurry? Read the executive summary.

Why the Same Tool Gets Two Answers
#

Renata approved the rollout in eleven weeks at the retailer and has spent fourteen months blocking it at the hospital, and it is the same tool.

She is not confused about what it does. At the retailer she ran the operations side of marketing, and when the model arrived she tested it against six months of campaign copy, found it comfortably better than the agency they were paying, and had it in production before the contract renewal. She keeps a printed copy of that first test in her desk, not for any reason she can articulate, and takes it out occasionally the way people take out a photograph.

At the hospital system she runs clinical documentation improvement, which is the function that decides whether what a physician wrote supports what the hospital billed. The same vendor came in with the same architecture and a demonstration that was, by any measure she could apply, more accurate than her staff. She has said no four times. She will probably say no a fifth time in October.

Nobody at the retailer thinks she was reckless. Nobody at the hospital thinks she is a Luddite. She is the same person, applying the same judgment, and the judgment produces opposite answers.

The Rewiring That Is Not There
#

The standard account of why good technologies take decades to land is about the cost of the change around them, and it is a real account with a real case behind it.

Electric motors were available to American factories from the 1890s. The productivity gain did not show up in the statistics until the 1920s. The reason was not that anyone doubted the motors. It was that the factory was built around a steam engine and a central driveshaft, with machines placed according to how much torque they could get from their position on the line, and vertically stacked so the shaft could serve several floors. Getting the benefit of electricity meant tearing that building down and putting up a single-story one where machines sat according to the order of the work. That is not an adoption decision. That is a capital program, and the factories that made it were mostly new factories, because the existing ones were waiting for their buildings to wear out.

Almost none of that transfers.

The model arrives as a subscription. There is no driveshaft, no stacked floors, no building to outlive. Renata’s hospital could have it running against live documentation inside a month, and the cost of finding out whether it works is small enough that it does not require anyone’s approval above hers. The thing that explains the forty-year lag is absent, and the lag is here anyway.

So something else is doing the waiting, and it is not the cost of the rewiring and it is not, in her case, ignorance about what the tool can do.

The Bar Is Set by the Worst Thing That Can Happen
#

Adoption looks like it should be a function of performance. The tool is better than what we have, therefore we use it. Every vendor deck is built on that sentence and it is not how the decision gets made.

What Renata is actually doing, in both jobs, is comparing a distribution against a consequence. Not the average output. The bad tail of the output, multiplied by what a bad output costs, against the same product for the process she already has.

At the retailer the tail was cheap. The worst plausible failure was a campaign that read strangely, went out to a segment, underperformed, and got pulled in a day. She could name the cost, it was four figures, and the process it replaced produced failures of roughly the same size at a rate she had measured. The mean was better and the tail was survivable, so the answer was yes in eleven weeks.

At the hospital the mean is also better. She has the numbers and does not dispute them. The tail is a systematic coding pattern that runs undetected for two quarters across forty thousand encounters and surfaces as a federal audit, at which point the question is not whether the model was more accurate than her staff on average. The question is whether the hospital can produce a person who reviewed the pattern, and the answer determines whether this is a repayment or a fraud allegation.

The bar is set by the tail, and capability improvements move the mean.

That sentence is the whole difficulty. Every release makes the system better at the ordinary case, which is the part that was never blocking anything. The organizations sitting on their hands are not waiting for a better average. They are waiting for a bound on the worst case, and a bound is a different kind of object from an improvement.

This is also why the demonstration never lands. A vendor’s evaluation is built to show the mean, because the mean is what a benchmark measures and what a pilot surfaces in six weeks. A pilot large enough to characterize the tail would have to run long enough for the rare failure to occur, and if it occurs during the pilot the pilot has failed, so no vendor designs one. The evidence Renata needs is structurally the evidence nobody has an incentive to produce, and both parties leave the room believing the other one is being unreasonable.

Two organizations can look at identical demonstrated performance, run identical arithmetic, and land on opposite sides, and both be right. This is not a difference in sophistication or courage. It is a difference in what happens at the end of the distribution, and the end of the distribution belongs to the organization rather than to the tool.

The Asymmetry That Nobody Prices
#

The arithmetic Renata is running has a large error in it, and she knows.

The hospital’s current process also fails. Documentation goes under-specified, encounters get coded to the wrong severity, and the consequences fall on patients in ways that are real and continuous: a readmission risk score that does not reflect what the physician actually saw, a discharge that goes out against a record that missed something. Those failures have been happening every week for years. Nobody has been fired for them. No auditor has arrived because of them. They are the ambient failure rate of the existing arrangement, and the ambient failure rate is invisible in exactly the way a new failure would not be.

An error made by a system somebody chose to install is attributable, novel, and has a name on it. An error made by the arrangement that was already there is diffuse, precedented, and belongs to no one. The two are priced completely differently by every institution that has ever existed, and the difference is not about harm.

So Renata’s arithmetic is correct about the tail and wrong about the baseline, and correcting it would not change her answer, because she is not the party that bears the diffuse cost and she is very much the party that would bear the attributable one. The incentive is doing what incentives do. Naming it does not dissolve it.

It Is Not Just Regulation
#

The obvious reduction is that this is about regulated industries, and that the pattern is really a story about compliance departments.

It does not hold up in either direction. Inside Renata’s hospital, the marketing function adopted the same class of tool eight months ago with no friction at all, because a bad hospital advertisement is a bad advertisement. The same building, the same legal department, the same risk committee, and the answer flips by function. Consequence is doing the work, not the industry.

And it runs the other way too. A payments processor is not a regulated profession in the sense a hospital is, and it behaves like one, because a systematic error in settlement logic is a headline. A structural engineering practice with four employees moves as slowly as a health system, on nothing but the recognition of what a wrong number does. Meanwhile financial firms, which are as regulated as anything in the economy, adopted these systems with startling speed in every function where the worst case is an embarrassing email.

The gradient is consequence, and it cuts across every other category anyone has proposed for it.

What Actually Moves It
#

If the bar is the tail rather than the mean, then three things move it and none of them is a better model.

The first is a way to bound the worst case rather than reduce its probability. These are different in kind. A system that fails once in ten thousand rather than once in a thousand is a better system with the same worst case. Renata’s problem is not the rate; it is that she cannot say in advance what the largest possible failure looks like, and every honest technical answer she has received has been about the rate.

The second is attribution. Not blame, exactly, but a mechanism that makes an error traceable to a decision made by an identifiable party who could have decided otherwise. Her staff provide this by existing. It is most of what she is buying from them, and the vendor’s product does not include it and cannot, because the vendor’s liability position depends on not including it.

The third is precedent, which is the one that arrives on its own. Somebody else’s audit, resolved. A regulator who has published a view. Another system in the state that did this and is still standing. Precedent is how consequence gets bounded in institutions that cannot bound it themselves, and it is why adoption in high-consequence settings looks flat for years and then goes vertical, in a shape that has nothing to do with the capability curve underneath it.

Which means the first adopter in any high-consequence sector is doing something categorically different from the fiftieth. They are not buying a tool. They are manufacturing the precedent that the other forty-nine are waiting for, at their own cost, with no mechanism to charge for it. Sectors differ enormously in whether anyone is positioned to play that role: a large academic center with tolerance for being wrong in public can, a rural hospital with one compliance officer cannot, and the sequence in which a sector adopts is mostly a fact about who could afford to go first rather than about who understood the technology soonest.

There is a fourth thing, and it is the ugly one. Consequence can be reassigned. If the vendor indemnifies, or the regulator issues a safe harbor, or the insurer writes the policy, the tail moves off the organization’s balance sheet and the arithmetic changes without a single technical fact changing. Most of what will look, in retrospect, like a wave of adoption driven by improving models will have been a wave of contracts.

I wonder whether the language of readiness has been backwards this whole time, and organizations described as not ready are more often organizations that have correctly identified something the people selling to them have no way to supply.

None of which tells anyone how long. The hospital will adopt this. Renata expects to be the person who approves it, and expects to do so on materially the same technical evidence she has been declining for fourteen months, because what will have changed is the environment around the evidence rather than the evidence.

The printed copy of the retailer test is still in her desk drawer, at the bottom, under a stapler she does not use. She has been asked twice why she keeps it and both times said she was not sure. It is the only piece of paper she owns that records her being right about something before anyone else in the building agreed.


This essay follows the argument in the first entry of The Dissents (The Half-Arrival), which established that adoption gates on the threshold of usable value rather than on capability, and takes up the question that exchange left open: what sets the threshold in a particular place. It connects to Part 95 (The Many Clocks), whose deployment clock is plural for the reason described here, and to the Arbitrage sequence, where the same gap between capability and deployment appears as a position somebody holds.


References
#

The Lag Between Capability and Productivity:

David, Paul A. “The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox.” The American Economic Review, vol. 80, no. 2, 1990, pp. 355-361.

Devine, Warren D. “From Shafts to Wires: Historical Perspective on Electrification.” The Journal of Economic History, vol. 43, no. 2, 1983, pp. 347-372.

Brynjolfsson, Erik, Daniel Rock, and Chad Syverson. “The Productivity J-Curve: How Intangibles Complement General Purpose Technologies.” American Economic Journal: Macroeconomics, vol. 13, no. 1, 2021, pp. 333-372.

Risk, Tails, and Institutional Decision-Making:

Perrow, Charles. Normal Accidents: Living with High-Risk Technologies. Princeton University Press, 1999.

Vaughan, Diane. The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA. University of Chicago Press, 1996.

Kahneman, Daniel, and Amos Tversky. “Prospect Theory: An Analysis of Decision under Risk.” Econometrica, vol. 47, no. 2, 1979, pp. 263-291.

Attribution, Accountability, and Automation:

Elish, M.C. “Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction.” Engaging Science, Technology, and Society, vol. 5, 2019, pp. 40-60.

Parasuraman, Raja, and Victor Riley. “Humans and Automation: Use, Misuse, Disuse, Abuse.” Human Factors, vol. 39, no. 2, 1997, pp. 230-253.

Diffusion:

Rogers, Everett M. Diffusion of Innovations. 5th ed., Free Press, 2003.

How this essay connects to others across The Approximate Mind.

The Half-Arrival established that adoption gates on a threshold of usable value rather than on capability, and left open what sets the threshold in a given place. Not Yet answers with consequence: the bar is set by the tail of the distribution while every release moves the mean.
The Many Clocks records the deployment clock as plural by jurisdiction without saying why. Not Yet supplies the reason, which is that the clock is set by what a failure costs the organization rather than by anything about the technology or the regulator.
The Mapped Territory finds a relational function that was free because the logistics required hands. Not Yet finds the same accounting error running the other way: an ambient failure rate that belongs to nobody is priced at zero against an attributable one that has a name on it.
The Alibicompanion
Not Yet handles the case where a refusal is correct arithmetic. The Alibi handles the case where the arithmetic is a costume, and neither essay can be read safely without the other, because the two are indistinguishable in the room.
The Valid Ticket supplies the imported half of this essay's two answers: the deployment model that is true on the floor and false at the desks, carried across the border by people who learned it honestly, which is why the same tool keeps getting two verdicts from the same honest reader.
The Lag Between Capability and Productivity
  1. David, Paul A. “The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox.” The American Economic Review, vol. 80, no. 2, 1990, pp. 355-361.
  2. Devine, Warren D. “From Shafts to Wires: Historical Perspective on Electrification.” The Journal of Economic History, vol. 43, no. 2, 1983, pp. 347-372.
  3. Brynjolfsson, Erik, Daniel Rock, and Chad Syverson. “The Productivity J-Curve: How Intangibles Complement General Purpose Technologies.” American Economic Journal: Macroeconomics, vol. 13, no. 1, 2021, pp. 333-372.
Risk, Tails, and Institutional Decision-Making
  1. Perrow, Charles. Normal Accidents: Living with High-Risk Technologies. Princeton University Press, 1999.
  2. Vaughan, Diane. The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA. University of Chicago Press, 1996.
  3. Kahneman, Daniel, and Amos Tversky. “Prospect Theory: An Analysis of Decision under Risk.” Econometrica, vol. 47, no. 2, 1979, pp. 263-291.
Attribution, Accountability, and Automation
  1. Elish, M.C. “Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction.” Engaging Science, Technology, and Society, vol. 5, 2019, pp. 40-60.
  2. Parasuraman, Raja, and Victor Riley. “Humans and Automation: Use, Misuse, Disuse, Abuse.” Human Factors, vol. 39, no. 2, 1997, pp. 230-253.
Diffusion
  1. Rogers, Everett M. Diffusion of Innovations. 5th ed., Free Press, 2003.