Res Agentica
Reading

No saved reading position.

Reading

No saved reading position.

Value Needs Work

The Selection Gradient

13 min read
Aa
Text size

The machines being of themselves unable to struggle, have got man to do their struggling for them: as long as he fulfils this function duly, all goes well with him.

— Samuel Butler, Erewhon (1872)


Armen Alchian offered a reframing of the firm that still feels like a corrective. Writing in 1950, he did not require firms to be rational, or even particularly coherent. He required only that markets be competitive enough to prune. A clumsy firm that blunders into a superior method will outlast a careful firm that plans an inferior one. Outcomes survive. Intentions do not. Richard Nelson and Sidney Winter extended the insight into an evolutionary account of economic change: firms as carriers of routines, habitual patterns that once worked well enough to escape pruning, modified through search containing luck, error, imitation, drift. Order emerges without omniscience, through the differential survival of routines that happen to fit the world as it is.

The previous chapter separated the work spent on a result from its usefulness and return. Between them lie decisions made by different parties for different purposes. Researchers evaluate a model against an objective. Users decide whether it helps. Organizations decide what to support, develop, or abandon. Their judgments can conflict without any of them being merely decorative.

Training, deployment, and capital allocation provide three distinct filters. Lower measured loss can guide parameter adjustment while leaving deployment problems unresolved. A code model can be useful in the frameworks its users actually work with despite an imperfect benchmark result. A translation model can perform well on an evaluation and attract little demand. Investors, boards, and budget committees then make further choices about what to finance, sometimes before deployment supplies much evidence. Their priorities shape which experiments become possible and which useful results receive no further support.

A result can fail at one of these stages after succeeding at another. In a stipulated throughput account, the share clearing declared filters can be represented together; that bookkeeping is not a growth law or a universal sequence through which all value must pass. A failed commercial investment may leave a useful technique behind. What is selected for continued production depends on who is choosing, what they can examine, and what they are prepared to fund.


The V/C Ratio

Consider two hypothetical cases. In the revenue-cycle office of a regional hospital, a prior-authorization request arrives for a proposed knee replacement. An analyst, or increasingly a system, checks coverage against policy terms, verifies eligibility, compares the procedure code with the coverage table, and issues a determination. Many inputs are structured: diagnostic codes, procedure codes, formulary lists, medical-necessity guidelines. Interpretive questions remain at the margins, but a large class of determinations can be checked against the same documents that governed the decision.

Down the corridor in the claims-resolution office, a different analyst sits before a surgical complication. The insurer must ask whether it was foreseeable, whether the team's response met the applicable standard of care, and how the costs divide between coverage and liability. Operative notes, nursing records, expert judgment, legal standards, and an unobservable counterfactual all enter the file. The comparison no longer reduces to a table lookup.

The contrast is not a ranking of intelligence. Both tasks involve trained work and ambiguity. They differ in the cost of checking an answer against an agreed standard.

Some parts of the authorization decision can be checked against the governing policy and supplied records. Others require medical judgment or evidence those records do not contain. A dispute over a surgical complication can add chart review, expert consultation, and a contested counterfactual. The comparison concerns the work needed to establish an answer, not a guarantee that one department’s decisions take minutes and the other’s cannot profitably be assisted.

The V/C ratio is a sequencing heuristic for this difference. Within a defined domain, let V be the expected surplus from accepting a correct completion rather than using the relevant baseline, and let C be the marginal cost of validating that completion to the confidence the domain requires. High V/C makes substitution easier to justify; low V/C makes checking consume the claimed surplus. The ratio is not a universal law. Adoption also depends on liability, regulation, integration cost, data rights, capital already sunk in the old process, and whether errors can be detected before they compound. Its narrower claim is useful: when those conditions are held roughly constant, checkability often predicts automation order better than an informal ranking of cognitive difficulty.

Oliver Williamson's transaction-cost economics identified incomplete contracting, enforcement, and asset specificity among the forces shaping firm boundaries. V/C adds checkability to that account. Custom software may be easy to describe at a high level and expensive to validate against edge cases, scale, and adversaries. Commodity inference can be easier to compare when accuracy, latency, and uptime are measured against a stable benchmark. Firm boundaries can move where performance is inspectable and remain firmer where verification is expensive. Checkability supplements the older account; it does not replace the other reasons organizations internalize work.

Chess illustrates the extreme. Grandmasters spend decades developing intuitions about position, sacrifice, and timing, but the game has a mechanically checkable win condition. Code generation often follows the same logic when a test suite captures the behavior that matters. Running tests is cheap; deciding whether the suite covers security, performance, and unanticipated use may not be. Automation is safest inside the part of the specification the tests actually witness.

Cheap verification gives a purchaser grounds for comparing offers and refusing defective work. Those grounds become competitive pressure when another provider can supply an acceptable result and the purchaser can actually move. A benchmark does not carry the data, retrain the staff, or terminate the existing contract. Expensive verification presents a different difficulty: the recipient must decide which additional checking the use requires and whether the assistance still warrants its cost. A lawyer may have to read the cases cited in a generated memorandum. That work can expose a fluent invention; it can also leave the lawyer with useful drafting or analysis that would otherwise have taken time to produce. The economic comparison turns on the work the lawyer can responsibly save, including the work needed to discover what cannot be accepted.

And in the middle range, where verification is possible but burdensome, the most insidious pathology emerges.


The Competence Trap

Consider a hypothetical radiologist who sits down each morning and opens the queue of images the system has flagged. In the first weeks she works as she has always worked, reading with care, comparing what she sees against what the model suggests, disagreeing when the model is wrong, catching what it missed. She is alert. She is practicing the skill that twenty years of training built. The system is usually right.

As agreement with the model becomes routine, disagreement becomes the rare event. A flag appears. She glances at the image, finds the flag consistent with what she sees, and confirms. The department records faster throughput.

Something has changed that does not show up on any dashboard. She is no longer performing independent diagnosis. She is performing confirmation of the model's diagnosis, and confirmation is a different cognitive act, requiring less sustained attention, less of the slow muscular pattern recognition that distinguishes a competent reader from a great one. In this scenario the skill atrophies: not in a dramatic failure but in the quiet erosion that follows from disuse, the way a surgeon's hands lose their precision during a long sabbatical, the way a musician's intonation drifts when she stops practicing scales. She understands the image less thoroughly than the system that flagged it. But she is the one who signs the report. She is the one who will be deposed if the patient's family sues.

Her signature has become a ceremony, a ritual performance of oversight whose epistemic weight has thinned to nothing without anyone noticing, least of all herself. Delegation has hollowed out the competence required to oversee the delegation. The loop has emptied the human of the skill the loop was supposed to preserve. And the pathology is invisible from within: the hospital sees faster throughput, the insurer sees lower costs, the radiologist herself sees a manageable workload. The only evidence that anything has gone wrong will be the error she would once have caught, arriving one morning in a study she approves in three seconds and a keystroke, producing a consequence that unfolds over months of misdiagnosis while her signature assures the world that a qualified human was watching.

This is the Competence Trap: delegation without maintained understanding. One route into it lies where the V/C ratio sits in a middle range: high enough that automation is economically justified, low enough that a human signature is still demanded. Oversight becomes the justification for delegation and then becomes incapable of catching delegation's errors, because the oversight function has been converted into assent.

An institution that demands a human signature on automated output inherits an obligation it rarely acknowledges: the obligation to keep the signer competent to evaluate what is being signed. Periodic independent readings where the system's output is withheld and the radiologist must decide on her own. Audits that measure diagnostic accuracy, not throughput. Requalification that tests the skill the automation is displacing rather than the speed with which the human can approve its conclusions. Without these investments, the signature becomes an institutional fiction: an assurance to the patient, the insurer, and the court that a qualified human reviewed the image, when in truth a qualified human glanced at a screen for three seconds and pressed Confirm.

Attorneys sign AI-generated memoranda without independently verifying the cited cases. Financial auditors accept automated analyses without reperforming the calculations. In aviation, the automation paradox has been documented since the 1990s and has contributed to accidents in which the crew's first manual intervention in months was the one that occurred during an emergency. The pathology appears wherever verification is expensive enough to tempt the human into deference but not so expensive that the institution gives up on human oversight entirely. The middle range. The range where institutions pretend that someone is watching, and no one, in any meaningful sense, is. The trap reaches individuals and the pipeline that produces them. When agents perform the junior work through which future verifiers are trained — the associate's first brief, the resident's first unassisted read, the apprentice auditor's first independent check — institutions can lose occasions through which verification capacity renews itself. Assistance can also open work to new entrants; whether it develops future judgment depends on the work people continue to perform and the responsibility they are allowed to assume.


What Verification Cost Predicts

Arithmetic automated first because a correct sum is immediately valuable and almost costless to verify: mechanical calculators made the point, electronic computers scaled it, spreadsheets domesticated it. Pattern recognition followed: optical character recognition, speech-to-text, image classification. These tasks are cognitively nontrivial, but their outputs can be compared cheaply against labeled ground truth that centuries of human practice left behind as an unintended subsidy to both training and verification. Filing clerks who typed index cards for library catalogs, court reporters who transcribed testimony, medical illustrators who drew anatomical plates: all produced labeled datasets as a byproduct of their work. The labels were never intended as training data. They became training data: examples of correct performance in standardized formats, mechanically comparable.

Medical imaging has attracted far more authorized AI devices than some adjacent diagnostic domains. Verification cost may explain part of that difference, alongside data availability, workflow maturity, reimbursement, regulation, and the cost of digitizing physical specimens. The comparison supports the heuristic only after those alternatives are considered.

Coordination is beginning to automate, and the V/C ratio identifies one line within the domain. Bilateral coordination with explicit terms can be mechanically checkable: did the delivery arrive by the specified date, was the quantity correct, did the payment settle at the agreed price? Multilateral coordination among parties with conflicting objectives may add a contested standard, such as whether a reallocation was equitable. Bilateral cases are often easier to automate. The distinction is checkability, not the number of parties by itself.

Beyond coordination lies the setting of objectives. The Specification Horizon names the point at which the cost of making a purpose operational exceeds what the principal can state, observe, or afford to check. Better specifications can move that horizon. They cannot make every future circumstance available in advance, and no finite contract settles every dispute about what fidelity to a purpose required.

The V/C ratio illuminates why. Verifying whether an agent followed an instruction requires comparing the action with the instruction. When the instruction says “deliver the cargo by Tuesday,” part of the check may be cheap. When it says “act in the client's best interest,” verification requires reconstructing intent and arguing about a standard never fully articulated. The cost can exceed the value of the delegated act. Beyond that practical boundary, the agent may obey the encoded objective while departing from what the principal meant.

The structural parallel to the participation horizon is precise. One marks where governance cannot reach the governed: deliberation too slow for the decisions it must constrain. The other marks where delegation cannot reach the intent: specification too narrow for the purposes it must serve. Both are boundaries of the same constitutional problem: coordination outrunning the human capacities on which its legitimacy depends.

The hospital administrator from Act I operates at this horizon. The objective function specified aggregate patient outcomes but did not encode the weight she assigns to the third-floor ward serving single parents who cannot travel across the city. A more elaborate objective might represent some of that knowledge. The constitutional point is that the institution must still name who may revise the objective when lived consequences expose what its designers omitted.

The Specification Horizon marks the V/C heuristic's limit. Where the standard itself remains politically or morally contested, a cheaper check cannot decide which standard should govern. The book assigns that decision to an answerable human institution—not because future machines are metaphysically barred from proposing purposes, but because coercive purposes require an author who can be questioned, overruled, and held responsible.

A corporation that sets the wrong objective for its agent fleet may not discover the error until consequences have accumulated. The quality of an objective often reveals itself through those consequences, on a timescale no verification procedure can abolish. Machines can propose purposes. Their adoption must remain reversible because the evidence needed to judge them may arrive only later.

Judgment persists because evidence for evaluating ends arrives through time, consequence, and the lived experience of what a purpose produces at scale. Computation can shorten analysis; it cannot make future evidence present. The constitutional role remains humanly attributable so that the institution choosing the end can answer while that evidence accumulates.

Energy structured through computation and disciplined by selection is the emerging factor of production. Within a defined domain, holding the named counterforces roughly constant, verification cost should predict automation order better than cognitive difficulty alone. Persistent inversions under those conditions, with lower-V/C tasks automating before higher-V/C tasks, would count against the hypothesis.

Search the book

Use ↑ ↓ to move through results; Escape to close.

Search every published chapter, section and reference.

    In this chapter