Res Agentica
Reading

No saved reading position.

Reading

No saved reading position.

Coherence Cost Model

Coherence costs money

21 min read
Aa
Text size
A21Written accountDeclared checking costs, reusable results and choices under a verification budget.

At Tana you should furnish yourself with a dragoman. And you must not try to save money in the matter of dragomen by taking a bad one instead of a good one. For the additional wages of the good one will not cost you so much as you will save by having him.

— Francesco Balducci Pegolotti, La Pratica della Mercatura, chapter II; Henry Yule translation, Cathay and the Way Thither, vol. III, pp. 151–155

A21 identifies expenditures under a declared checking regime. Its categories make search, maintenance and organizational decisions available for comparison without assigning a universal price to coherence. The examples of changing schemas and event representations test what those comparisons require. Vol I, Chapters 7–8 develop the related institutional question: what must another provider supply to replace the work it proposes to make cheaper?

Coherence Cost

Inventing a predicate requires search, checking and decisions about its use. Maintaining it can require further work as vocabularies, populations and commitments change. The coherence fee names the expenditure on that work. A certificate may be reusable; a new context may require another investigation. What must be paid again depends on the claim being maintained and the checking regime. A provider's bill does not disclose that distinction by itself.

Three sources of expenditure need to be distinguished:

  • The overlap surface: comparisons required between views
  • The migration surface: dependencies affected by a change
  • The organizational surface: decisions required when stakeholders disagree

Schemas can move work into migrations and approval procedures, making later queries cheaper. Learned representations can make useful associations available without first settling the corresponding agreement of meaning. Neither architecture fixes the checking regime: both can incorporate further validation, and both can leave a recipient with work still to do.

Declaring the work and its unit of account makes an estimate examinable. The estimate can guide a budget; whether the work was performed remains another question.

The Cost Model

A21
A21: Coherence Cost Model

The coherence cost of an operation is an accounting decomposition for a declared scope, authority and checking regime. Its components can be summed only after specifying a common unit and accounting period; otherwise retain them as a vector:

CoherenceCost(op, U, auth) = C_compile(op) + C_runtime(op, U) + C_org(op, auth)

Compile-Time Costs (C_compile): One-time costs incurred at invention, promotion, or migration.

C_compile = C_search + C_certify + C_witness + C_migrate

Run-Time Costs (C_runtime): Per-query costs incurred at evaluation and coherence checking.

C_runtime = C_local + C_overlap + C_verify + C_drift

Organizational Costs (C_org): Per-decision costs incurred in human coordination.

C_org = C_approval + C_audit + C_dispute

A coherence budget is a policy object that declares the maximum cost the system will pay:

CoherenceBudget(U, auth) = {
  max_compile, max_runtime, max_org, max_expected_loss,
  staleness_tolerance, scope_boundary
}

The budget declares the checking work the institution proposes to fund, its scope and its schedule. A separate report must establish which of that work was completed.

Remark(Relative vs. absolute costs)

The cost model separates expenditures and makes budget choices explicit. It supplies no general ordering of costs under cover refinement. Relative comparisons require a declared checking regime, a common unit, and estimates for the work each alternative actually retains or replaces. It does not provide an absolute cost function — the actual values of C_search, C_verify, C_approval depend on domain-specific factors (vocabulary size, verification regime, organizational structure) that the framework does not calibrate. Worked examples in this chapter use illustrative numbers to show cost structure; practitioners must calibrate to their own systems. The claim is structural (these cost categories exist and interact this way), not predictive (this operation will cost this many units).

Coherence vs Consistency: Coherence (Third Mode) is agreement of meaning on overlaps under declared scopes, witnesses, and absence policies. Consistency (distributed systems) is agreement of values under a storage protocol. Consistency asks: do replicas agree on the data? Coherence asks: do views agree on what the data means? The cost model concerns the declared checking work. Its required storage guarantees must be specified; agreement of replicas alone does not establish agreement of meaning.

Compile-Time Costs

Compile-time costs are one-time. You pay them when a predicate is invented, promoted, or migrated.

C_search: The cost of predicate search from Chapter 18. Proposal, expansion, pruning, scoring. Scales with grammar complexity, derivation depth, and parameter space size.

C_certify: The cost of checking A17 obligations. Local grounding requires evaluating the predicate on witness sets. Overlap agreement requires checking pairwise reconciliation. Invariant preservation requires constraint validation.

C_witness: The cost of producing witnesses. Derivations, exemplars, provenance records, calibration curves. These artifacts are not free; they require computation and storage.

C_migrate: The cost of producing migration witnesses when a breaking change occurs. Chapter 16 established that breaking changes require explicit artifacts. Producing those artifacts has a cost that scales with the number of dependent predicates.

Compile-time costs are often underestimated because they are invisible in steady state. A system with stable vocabulary pays them rarely. A system with rapidly evolving vocabulary pays them constantly.

Consider a fashion platform adding 50 new style predicates per quarter. Each predicate requires search, certification, and witnessing. If C_compile averages 10 CPU-hours per predicate, the platform pays 500 CPU-hours per quarter just for predicate invention. That is the hidden tax on vocabulary growth.

Run-Time Costs

Run-time costs are per-query. You pay them every time you evaluate a predicate or verify coherence.

C_local: The cost of evaluating a predicate in a single view. For measurement predicates, this is computation. For rule-based predicates, this is rule evaluation. For learned predicates, this is model inference.

C_overlap: The cost of verifying agreement on overlaps. This is where the cost model becomes interesting.

Overlap Cost Structure

Define an overlap graph G_U = (V, E) where vertices V are views and edges E connect views with non-empty overlap. Let d̄ be the average degree.

We measure C_overlap per unit time (e.g., per day):

C_overlap(p, U) = f(p) · |E| · k̄(p)
                ≈ f(p) · |V| · d̄ / 2 · k̄(p)

where:

  • f(p) is the check frequency (checks per day)
  • |E| is the number of overlap edges
  • k̄(p) is the average per-check cost (e.g., in seconds or dollars)

Scaling regimes:

  • Dense (d̄ ≈ |V|): O(|V|²) checks per day. Every view overlaps every other.
  • Sparse (d̄ << |V|): O(|V| · d̄) checks per day. Most views are disjoint.
  • Hub-and-spoke: One central view overlaps all others. Sparse globally, but hot edges on the hub.

A sparse graph reduces the number of pairwise checks in this scheme. Concentrated activity at a hub can still dominate expenditure. The actual graph and schedule must be measured rather than inferred from the architecture’s name.

What Refinement Does Not Order

A finer cover need not require more overlap checks. Consider the discrete space U={a,b,c}U=\{a,b,c\} and two covers:

C={{a,b},{b,c}},C′={{a},{b},{c}}.\mathcal C=\{\{a,b\},\{b,c\}\},\qquad \mathcal C'=\{\{a\},\{b\},\{c\}\}.

Each member of C′\mathcal C' lies in a member of C\mathcal C, so this is a refinement under the definition used here. The coarse cover has one nonempty overlap between distinct members; the fine cover has none. With one unit charged per such check, the overlap cost falls from one to zero.

The former Coherence Budget Monotonicity theorem is therefore withdrawn. Its proof assumed that refinement preserves every old overlap as a distinct check among the new members. Factorization through a coarse member does not impose that condition. Neither the cost inequality nor its claimed equality case follows.

The missing check matters. Agreement between two independently supplied descriptions of bb is a different obligation from using a single description on the singleton {b}\{b\}. A cheaper representation may discard that comparison; another implementation may preserve it elsewhere. Counting the new cover alone does not decide whether the old evidentiary task has been performed. The budget must name the checks it funds and the obligations that continue to bind when the representation changes.

For a declared overlap graph and checking schedule, the expression above still accounts for repeated work. Dense pairwise checking has a quadratic edge count; sparse checking does not. Frequencies, per-check costs, retained certificates and organizational work determine the expenditure. These are grounds for comparing specified alternatives, not an ordinal law of refinement. A21 supplies the categories and the budget in which those choices can be made; it establishes neither a universal thermodynamic fee nor the part of a provider's charge that constitutes economic rent.

Staleness tolerance reduces f(p). If you re-check coherence every query, f(p) equals your query rate (potentially millions per day). If you re-check once per day, f(p) = 1. The tradeoff is freshness versus cost. A schedule can combine checks at each use with periodic checks, provided those intervals meet the receiving contract.

C_verify: The cost of checking witness validity at query time. This depends on the verification regime:

  • Proof-carrying artifacts: check the proof and its hypotheses; a signature or hash check alone establishes a different claim
  • Attestations: validate the relevant chain and evidence; cost depends on chain length, algorithms and revocation checks
  • Re-certification: repeat the applicable A17 obligations when the cached grounds no longer suffice

No generic constant-time or logarithmic bound follows from these labels. A cached result needs its own binding to the current inputs and scope.

C_drift: The cost of detecting and handling semantic drift over time. Measured by:

  • Disagreement rate: percentage of overlap checks that fail
  • Calibration decay: drift in calibration witnesses
  • Exemplar churn: percentage of witness set that becomes stale

Drift is the silent coherence killer. A system that was coherent at deployment becomes incoherent as the world changes and the predicates do not.

Consider a "trending" predicate defined as "items with sales velocity in the top 10%." The predicate was coherent when defined: all views agreed on what "top 10%" meant. Six months later, one view has switched to a rolling 30-day window while another still uses a 7-day window. The overlap checks pass (both return items), but the items are different. This is semantic drift: the syntax is stable, but the meaning has changed.

Why do naive overlap checks miss this? If the overlap contract only checks well-typedness or non-emptiness (both views return Item sets), drift slips through. Catching it requires either stronger overlap predicates (agreement on specific items, not just non-empty intersection) or periodic re-certification.

Version-change records, renewed checking and monitored disagreements can each reveal a changed dependency. Their costs and coverage differ; the budget must identify which work it funds.

Organizational Costs

Organizational costs are per-decision. You pay them when humans must coordinate.

C_approval: The cost of obtaining authority approval for promotion. User to org requires one approval. Org to global may require multiple stakeholders, legal review, compliance signoff.

C_audit: The cost of maintaining audit trails and provenance. Regulated industries require extensive documentation. Every predicate, every witness, every migration must be traceable.

C_dispute: The cost of resolving disagreements when overlaps fail. Two merchants define "sustainable" differently. The overlap check fails. Someone must decide what to do.

Organizational Cost Estimators
C_approval ≈ num_approvals × avg_active_labor_hours × loaded_labor_rate
C_audit    ≈ audit_frequency × audit_hours × auditor_rate
C_dispute  ≈ dispute_rate × avg_resolution_hours × stakeholder_count × labor_rate

These are illustrative estimators, not calibrated measurements or established bounds. They are useful even with rough inputs because they make the cost visible.

Organizational costs are often the dominant component. A predicate that is cheap to compute and cheap to verify may still be expensive to approve. A schema change that is trivial technically may be blocked for months by stakeholder coordination.

Model deployment and schema migration can both require extensive approval or permit local experimentation. The organizational burden follows the institution’s actual allocation of authority, not a necessary contrast between the two technologies.

The Third Mode makes C_org visible and lets you choose. Some predicates (user-level, local scope) can have C_org ≈ 0. Other predicates (system-level, global scope) may require extensive organizational investment. The budget declares what you will pay for each category.

The Cost of Incoherence

The cost model so far describes the cost of achieving coherence. But coherence is not an end in itself. The question is: what do you lose if you do not pay?

Expected Loss from Incoherence
expected_loss(p, U) = P(conflict) × impact(conflict)

where:
  P(conflict):     Probability of overlap disagreement going undetected
  impact(conflict): Business cost of acting on inconsistent data

Without overlap checks, this particular source of evidence about conflict is absent; the probability still depends on the sources and other controls. If conflicting data leads to wrong decisions, impact(conflict) is high. The expected loss is the product.

The coherence budget becomes an optimization problem:

optimal_budget: minimize [ CoherenceCost + expected_loss ]

You pay coherence cost to reduce expected loss. At an interior differentiable optimum, marginal cost equals marginal expected-loss reduction. Constraints, discrete choices and uncertain estimates can place the optimum elsewhere.

This supplies one risk-management calculation. It does not authorize an institution to trade away duties to people outside its budget, or to count only losses it expects to bear itself. Required checks and limits on use remain constraints on the optimization.

Coherence Budgets in Practice

A coherence budget is a policy object with parameters:

CoherenceBudget = {
  max_compile:        1000 CPU-hours per predicate
  max_runtime:        100 overlap checks per query
  max_org:            2 approval gates, 1 week cycle time
  max_expected_loss:  $10K/month in conflict-induced errors
  staleness_tolerance: 24 hours for non-critical predicates
  scope_boundary:     primary suppliers only; secondary = best-effort
}

Budget-driven behavior:

  • Operations exceeding max_compile are rejected or deferred
  • Queries exceeding max_runtime use cached results or approximations only where their contract permits; otherwise they defer or return an incomplete result
  • Promotions exceeding max_org are blocked until authority is obtained
  • Predicates exceeding max_expected_loss trigger mandatory review

Scope boundaries determine the coherence horizon. Beyond that boundary, the budget establishes no guarantee. Which operations may proceed remains a matter for their evidence and authority requirements.

Money can fund more investigation without making incompatible requirements jointly satisfiable or an undecidable problem decidable. A budget states the work the institution will support. Where that work cannot establish the needed grounds, the available action may have to change.

Example(50-Supplier Catalog Coherence)

A fashion platform integrates catalogs from 50 suppliers. Each supplier has their own vocabulary for "material," "style," "fit."

Without cost model: Attempt global coherence. Every attribute reconciled across all 50 suppliers. This requires 50 × 49 / 2 = 1,225 pairwise overlap checks per attribute. With 100 attributes, that is 122,500 checks per reconciliation cycle. Whether that load is affordable depends on the cost and schedule of each check.

With cost model: Declare a coherence budget.

CoherenceBudget = {
  max_runtime: 1000 overlap checks per cycle
  scope_boundary: { primary: top 10 suppliers, secondary: best-effort }
  staleness_tolerance: { primary: 24 hours, secondary: 24 hours }
}

Cost-aware behavior:

  • Primary suppliers (top 10 by volume): 10 × 9 / 2 = 45 pairwise checks. The declared pairwise checks, on the schedule below.
  • Secondary suppliers: checked daily, flagged but not enforced.
  • Tertiary suppliers: advisory only, no coherence guarantee.

Worked calculation for one predicate:

This schedule counts the declared primary-primary and secondary-secondary edges. It supplies no primary-secondary comparison or other omitted check.

Primary suppliers: 10 views, d̄ ≈ 9 (near-complete overlap)
  → |E| = 10 × 9 / 2 = 45 edges
  → f(p) = 1 check/day (daily batch)
  → k̄(p) = 20ms per check
  → C_overlap(primary) = 45 × 20ms = 900ms/day (~1 second/day)

Secondary suppliers: 40 views, d̄ ≈ 3 (sparse, some clusters)
  → |E| ≈ 40 × 3 / 2 = 60 edges
  → f(p) = 1 check/day (daily batch)
  → k̄(p) = 20ms per check
  → C_overlap(secondary) = 60 × 20ms = 1.2 seconds/day

Total overlap cost for this predicate: 2.1 seconds/day of compute

If secondary suppliers were instead checked at each of 10,000 queries per day:

C_overlap(secondary) = 60 edges × 20ms × 10,000 queries = 12,000 seconds/day = 3.3 hours/day

The budget makes the difference between 2 seconds and 3 hours.

Result: This one predicate requires 105 pairwise checks per cycle, within the 1,000-check budget. Applying the same schedule independently to all 100 attributes would require 10,500 checks and exceed it. The budget therefore leaves a real choice about coverage, reuse or additional resources. The arithmetic has not established the adequacy of the regime, costs outside this component, or which suppliers’ mistakes would matter most.

T9: Schema Evolution as Coherence Cost

T9 (Schema Evolution) is the problem of vocabulary change over time. "Employee" and "contractor" merge into "worker." The schema changes; dependent queries must adapt.

The failure mode is breaking changes without migration semantics. The schema changes; queries silently break or return wrong results.

The cost model view: migration is a coherence cost.

Migration cost components:

  • C_migrate: producing compatibility witnesses (A17b artifacts)
  • C_runtime: evaluating compatibility shims at query time
  • C_org: coordinating stakeholders on the change
Example(The 'Worker' Evolution)

Three teams define "worker" differently:

  • HR: employee ∪ contractor
  • Payroll: anyone with a pay stub
  • Legal: anyone with a signed agreement

Legal wants to add interns with unsigned agreements to their definition.

Cost analysis:

  • C_migrate: produce witness showing old queries still work (or explicitly documenting breaking changes)
  • C_runtime: compatibility shim evaluates both old and new definitions during transition
  • C_org: Legal, HR, Payroll must agree on timeline

Budget-constrained decision: If C_org exceeds budget (three teams cannot coordinate within cycle time), the system may:

  • Restrict the change to Legal's scope (no promotion)
  • Defer until coordination budget increases
  • Fork: maintain worker_legal_v2 as a separate predicate with explicit compatibility witness

Schema evolution is not a bug to be fixed but a cost to be budgeted. The Third Mode makes that cost visible.

The key insight: migration cost is not a one-time payment. The compatibility shim runs at every query during the transition period. If the transition takes 6 months and the predicate is queried 1000 times per day, that is 180,000 shim evaluations. The run-time cost of migration can exceed the compile-time cost of the change itself.

T10: Representational Complexity as Coherence Cost

T10 (Higher-Arity Events) is the problem of n-ary relations. "Alice introduced Bob to Carol at the conference" involves four entities in a single event.

The cost model asks which checks a representation requires. One possible scheme for an n-ary event materializes every pairwise projection and checks their compatibility. That scheme has the following number of pairs:

pairwise projections = n × (n-1) / 2

A 4-ary event (Alice, Bob, Carol, conference) has 6 pairwise projections. Six is the count for that scheme, not a lower bound on every faithful encoding or its verification cost. An event identifier can preserve the common event in a relational representation; separate relations that discard that identifier may lose it, as Chapter 29 examines. The accounting question is which information and checks survive the chosen encoding. Arity alone does not answer it.

Practical Budgeting

The unit economics of vocabulary growth:

total_coherence_cost = Σ_p [ C_compile(p) + T_p · qps(p) · C_runtime(p) + C_org(p) ]

where T_p is the expected lifetime in seconds and qps(p) is queries per second. Here C_runtime is cost per query; periodic batch work must first be converted to this accounting basis or retained as a separate term. The components must use the common unit and period required by A21.

Rules of thumb:

Dependencies identify work to examine. A change may affect many uses while a preserved interface or reusable check serves several of them. Count the required migrations and comparisons rather than price the change by connectivity alone.

Scope and authority identify further obligations. A wider requested use may need new evidence or approval; the new work depends on what is already adequate and who can supply it. The relevant question is whether a narrower use would serve the purpose while leaving other claims unmade.

Cost reduction strategies:

  • Restrict scope (fewer overlaps to check)
  • Reduce connectivity (fewer dependencies to migrate)
  • Increase staleness tolerance (less frequent checking)
  • Use approximations (probabilistic coherence instead of exact)

Each strategy proposes a change in the checking task. Whether it saves cost, and what it gives up, must be assessed under that changed regime.

What you trade:

  • Restricting scope trades reach for cost. Your predicate applies to fewer contexts.
  • Reducing connectivity trades reusability for cost. Other predicates cannot depend on yours.
  • Increasing staleness trades freshness for cost. Your coherence checks may be outdated.
  • Using approximations trades exactness for cost. Your coherence is probabilistic, not certain.

The calculation makes the proposed trades visible. Their permissibility depends on what the operation owes its users and other affected people. A smaller bill for checking is not evidence that the obligation itself has shrunk.

When the Budget Is Blown

The cost model assumes systems operate within their declared budgets. What happens when they do not?

A system can accumulate coherence debt. Predicates multiply without retirement. Overlaps grow dense. Staleness tolerance creeps upward. Eventually the system reaches a state where the cost of maintaining coherence exceeds the budget, and the gap keeps widening. This is coherence bankruptcy: the system cannot pay the tax it owes.

Bankruptcy has symptoms. Overlap checks time out. Migrations stall because the compatibility surface is too large. Organizational approvals backlog because stakeholders cannot coordinate on the accumulated changes. The system still runs, but coherence is a fiction maintained by hope rather than verification.

Recovery can involve additional resources, a better procedure or an authorized reduction in the promised work. The following possibilities concern the last course.

Scope shedding: Narrow the coherence boundary. Predicates that were global become org-level. Predicates that were org-level become local. The system explicitly abandons coherence guarantees for the shed scope. Downstream consumers receive notice: this predicate is no longer maintained at your scope level.

Predicate retirement: Remove predicates that no longer earn their cost. A predicate with few queries and many overlaps is a candidate. Retirement is not deletion; it is deprecation with a migration path. Queries against retired predicates receive a warning and eventually an error. The predicate's package remains in the audit log.

Invariant relaxation: Convert hard invariants to soft constraints. The system no longer blocks on violations; it flags and continues. This trades correctness for operational continuity. The tradeoff is explicit: relaxed invariants appear in the coherence report with their relaxation date and authorizing decision.

Staleness amnesty: Accept that some predicates will not be fresh. Declare a staleness floor: predicates with staleness tolerance below the floor are raised to it. Queries receive a warning that results may be stale. The system stops pretending it can maintain freshness it cannot afford.

None of these recoveries is costless. Scope shedding breaks downstream assumptions. Predicate retirement forces migrations. Invariant relaxation admits errors. Staleness amnesty reduces trust. The point is not to make recovery painless but to make it structured. A system that recovers through explicit shedding can explain what it gave up. A system that recovers through silent degradation cannot.

A budget can accompany a recovery plan, but declaring max_expected_loss does not authorize the system to shed another person’s protection or abandon a contract. Each recovery action requires its own authority and must preserve or explicitly resolve the obligations affected. The plan should identify who can decide, what they can change and what must happen to claims the reduced checking no longer supports.

Consequence

A specified checking regime can exceed the resources available to perform it. Declaring that limit makes the proposed service examinable. It does not establish that the checks were completed, and it cannot make the omitted work somebody else’s evidence.

The budget and checking contract state:

  • How far coherence extends (scope boundary)
  • How fresh coherence is (staleness tolerance)
  • How expensive operations can be (cost limits)
  • How much risk the system accepts (expected loss)

An intermediary may charge for useful checking and also control whether anyone else is permitted to supply it. The cost categories help investigate the work; they do not by themselves identify rent, prove a minimum charge or justify the intermediary's position. Someone who proposes to replace it must still say which obligations they will meet. Someone who insists that it remain must explain what its exclusive position contributes beyond the work another provider could perform.

Part V turns these declared obligations into a proposed interface. Its implementation must make the checking regime and the consequences of an exhausted budget inspectable.

Search the book

Use ↑ ↓ to move through results; Escape to close.

Search every published chapter, section and reference.

    In this chapter