Chapter 16
What Can Be Caught Wrong
Aa
The grand object of the modern manufacturer is, through the union of capital and science, to reduce the task of his work-people to the exercise of vigilance and dexterity.
Institutions do not automate whatever a machine can do. They automate what they can afford to let it do.
Capability is only the opening condition. A useful deployment must be integrated into a workflow, supplied with authority, monitored for failure, and connected to a remedy when failure matters. These costs vary independently. A system may perform brilliantly on a benchmark and remain unusable because the institution cannot recognize a bad answer before it acts.
That is the central provocation of this chapter. Adoption turns not only on what a machine can produce, but on what someone can catch it doing wrong.
From Reinstatement to Exposure
David Autor described how automation displaces particular tasks while creating new ones, allowing labor demand to be reinstated elsewhere.1 Acemoglu and Restrepo formalized the contest between displacement and new task creation.2 History supplies no guarantee that the two effects will continue to balance. General computational systems can enter newly created cognitive tasks more quickly than a loom could enter bookkeeping or an engine could enter design.
The old question—what can machines do?—therefore divides into several. Can a system perform the task under the actual distribution? Can the organization connect it to data and tools? Can error be detected in time? Who may authorize the act? Who bears the loss? Can an injured party obtain a correction?
Automation advances through that conjunction. No single ratio determines it.
A Candidate Ordering Variable
V/C compares the value at stake in successful completion with the cost of checking an output to acceptable confidence:
The expression is a heuristic, not a law. Its numerator is difficult: willingness to pay, social value, and avoided loss need not coincide. Its denominator also gathers unlike forms of checking—tests before deployment, review of each output, delayed observation, audit, and regulatory approval. A number computed without declaring those choices has false precision.
Even when specified carefully, V/C is only one term in a larger deployment account. Capability, execution cost, integration, expected error loss, complementary labor, regulation, demand, organizational capacity, and the authority to act can alter or dominate the sequence. Verification deserves isolation because it is often hidden, not because it is sovereign.
Oliver Williamson placed measurement and monitoring inside the larger economy of transactions and governance.3 Yoram Barzel showed why measurement costs can shape rights and institutions in their own way.4 V/C extends that intuition to computational output. It asks whether the cost of knowing that a task was done well changes the form and pace of substitution after capability is held constant.
Three Different Clocks
Code makes the hypothesis vivid. Syntax can be rejected instantly, tests can examine declared behavior, and version control preserves a path back. None proves that software serves the user's purpose, but the verification apparatus is unusually rich and correction can be cheap. Models entered coding early partly because production already possessed machinery for catching some kinds of error.
Insurance reveals a slower clock. Structured data can support automated underwriting, yet claims may require inspection, causation, medical evidence, fraud inquiry, and legal judgment. The activities live inside the same firms and regulatory systems but expose different facts to verification. That difference is evidence for the hypothesis, not a clean experiment: frequency, stakes, data, and organizational incentives also vary.
Medicine exposes the longest clock. A model can classify an image or recommend a treatment. Clinical success may emerge months later, and the counterfactual—what would have happened under another decision—may never become observable. Human review does not solve that epistemic problem. It supplies an answerable office, a professional standard, and a route for contest when uncertainty becomes harm.
These examples do not “confirm” V/C. They show what the variable is trying to see. Software, insurance, and medicine differ not simply in intelligence required, but in when error becomes legible and what can still be done after it appears.
Verification Is Not Remedy
The distinction matters because the language of verification can become too clean. A compiler verifies grammar. A test suite checks specified behavior on selected cases. A cryptographic proof establishes a formal relation. None decides that the specification was legitimate or that an injured person has been made whole.
Verification can itself be automated where ground truth is machine-readable and the checking rule is stable. Liability cannot simply be “assigned algorithmically” by the same move. A smart contract can transfer collateral when an encoded predicate fires. It cannot, without further institutions, determine whether the predicate expressed valid authority, whether an oracle was corrupted, or whether justice requires a different result.
This is why falling verification cost can accelerate deployment while increasing constitutional risk. An institution may become very good at detecting conformity to a rule and very poor at answering whether the rule should have governed the case.
An Autonomy Map, Not a Ladder
Assistive, supervised, delegated, and autonomous arrangements are useful descriptions, but they do not form an inevitable ladder. A system may act autonomously on a narrow high-frequency task and remain merely advisory on a neighboring one. A mature institution may deliberately restore human approval after an automated system becomes more capable because stakes, law, or public legitimacy changed.
The form of deployment should therefore be mapped across dimensions: scope of action, frequency of review, reversibility, loss allocation, and power of remedy. “Human in the loop” says almost nothing unless the human has time, information, authority, and a real capacity to refuse.
The Test
The hypothesis can still lose. Across comparable tasks, V/C should explain some variation in time to durable deployment after capability, integration, expected loss, regulation, and demand are taken into account. If a broader transaction-cost or institutional model explains the same variation without a distinct verification term, V/C becomes useful vocabulary rather than a separate analytical contribution. If organizations automate opaque, costly-to-check work as quickly as transparent work, its ordering claim weakens further.
Early enterprise surveys show how little raw adoption figures settle. Many projects remain pilots and many are abandoned.5 Those outcomes could reflect verification, weak capability, integration failure, absent demand, organizational politics, or all of them. The missing evidence is task-level comparison using declared denominators.
What remains is not a universal sequence but a durable question. Computation lowers the cost of producing candidate judgments. Institutions then discover that judgment was never the only expensive thing. Proof, authorization, correction, and consequence keep their own clocks.