Res Agentica
Reading

No saved reading position.

Reading

No saved reading position.

Predicate Acceptance

From vibe to citizen

12 min read
Aa
Text size
A29Written accountThe lifecycle of a predicate from informal proposal to formal acceptance.

Most striking at first is this appearance of sudden illumination, a manifest sign of long, unconscious prior work. The role of this unconscious work in mathematical invention appears to me incontestable…

— Henri Poincaré, Mathematical Creation; G. B. Halsted translation, The Foundations of Science (1913)

The Dress That Demanded a Dimension

In this hypothetical catalog, a buyer types "puffy dresses" into the search bar. The records, dates, counts and timings below are illustrative. The system returns results ranked by embedding similarity and filtered by available attributes: color, size, price, brand. None of these captures what the buyer wants.

"Puffy" is not a color. It is not a size. It is a property of silhouette: volume from fabric structure, the architectural drama of tulle and organza, the way a ball gown occupies space. The request asks for a distinction the stated catalog fields cannot evaluate.

A17 permits the proposal, and A24 specifies what its package must contain. Admission now asks whether that package supplies adequate grounds for the proposed filter.

The answer is not "it works on some examples." The answer is not "the merchandiser approved it." A29 defines the proposed acceptance checks. Their scope matters: exemplars test specified cases, while a universal invariant needs a proof or an appropriate bounded verification regime.

The Acceptance Gate

Once admitted as a filter, the proposed distinction can determine which items another query excludes. The gate therefore has to establish more than the availability of an evaluator. A29 states six checks for the declared receiving context. The examples below show why their results cannot be substituted for one another: acceptable latency does not repair poor discrimination, and matching exemplars do not prove a universal invariant.

Exploratory proposals can remain available while admission is unfinished. An established failure, an incomplete check and separately permitted restricted use retain the distinct outcomes stated below.

Anchor A29: Predicate Acceptance

A29
A29: Predicate Acceptance

Two gates:

  • AdmitLocal(q, U): Predicate q enters context U
  • Promote(q, U→V): Predicate q's scope widens from U to V

Six acceptance checks (AdmitLocal):

  1. DISCRIMINATION: Exemplars are correctly classified by evaluator

    • Bool predicates: all positives true, all negatives false
    • Score predicates: positives > τ_high, negatives < τ_low, boundary in [τ_low, τ_high]
  2. ABSTAIN_COVERAGE: Boundary handling is explicit

    • If abstain_policy ≠ "no_abstain_zone": boundary exemplars required, must fall in gray zone
    • If no_abstain_zone claimed: justification required
  3. CONFOUNDER_REJECTION: Hard negatives correctly rejected

    • All confounders score below τ_low
    • Minimal coverage: ≥K confounders (default 3) spanning ≥M subclasses (default 2)
  4. INVARIANT_SATISFACTION: Declared invariants hold

    • Logical: Establish any claimed old-language conservativity by its stated argument. EntailmentRegressionSuite checks the included derivations; it does not prove absence of all new old-language consequences.
    • Operational: Check the declared stable-query contract. A tolerance ε, where permitted, states approximate compatibility rather than exact preservation; changed results require a ChangeReceipt.
  5. OPERATIONALITY: Predicate is deployable within budget

    • Latency within budget, cost within budget (A21)
    • Dependencies pinned, caching semantics declared, revalidation triggers specified
  6. SCOPE_AUTHORITY: Requester has permission for requested scope

Acceptance result:

Admitted(q, AdmissionReceipt)
| Rejected(reason, failing_checks, alternative_standing?)
| Provisional(q, permitted_restricted_use, conditions_for_full_admission)
| Inconclusive(q, unchecked_obligations, reason)

Provisional use requires its own justified, restricted contract; it is not certification of the unfinished ordinary check. Claim registration, admission of a predicate and authorization of an action remain separate decisions. If unfinished checking keeps a consequential matter pending, the receiving process must also meet Appendix I, I.1.6. The six checks do not decide how long an institution may prolong that use without the required review.

The implementation can order checks to avoid needless expenditure after a decisive failure. Their cost depends on the evaluator and evidence: an exemplar may require an expensive service, and an authority check may require investigation. The proposed budgets describe this admission regime; they are not a theorem about the ordering of costs.

What the Receipt Contains

Example(Admission Receipt)

This illustrative receipt assumes the required proofs and operational checks have been supplied. The printed PASS fields report those assumed results; they are not demonstrations of them, and the exemplar counts do not establish conservativity.

AdmissionReceipt(puffy) = {
  predicate_id: pred_puffy_v1,
  package_hash: 0x8f3a...,
  admitted_to: U_fashion_catalog,
  admitted_at: 2026-01-15T11:00:00Z,
  admitted_by: substrate_admission_service,
  
  admission_profile: { confounder_minimum_K: 2, subclass_minimum_M: 2 },
  test_results: {
    discrimination: { positive_correct: 3, negative_correct: 2, boundary_in_zone: 1 },
    abstain_coverage: { boundary_count: 1, in_zone: true },
    confounders_rejected: 2,
    invariants: [
      ("range [0,1]", PASS),
      ("monotonicity", PASS),
      ("logical_conservativity", PASS),
      ("operational_conservativity", PASS)
    ],
    operationality: { latency_p99_ms: 47, cost: $0.002/eval }
  },
  
  standing: "org",
  promotion_eligibility: true,
  expiration: "indefinite",
  revalidation_triggers: ["3d_model_service version change", "calibration drift > 0.1"]
}

Queries can cite the admission record, and auditors can inspect the checks it names. The receipt establishes what was admitted under which procedure. The evidence and verified obligations behind it establish how far the filter may be relied on.

The Puffy Lifecycle

Stage 1: Proposal

In the hypothetical catalog, searches and marked examples suggest that the current filter fails to distinguish what some users are seeking. That is a reason to investigate a proposed predicate, not proof that the schema cannot represent it.

A merchandiser proposes "puffy" as a candidate predicate with initial intension: "dress with voluminous silhouette from fabric structure."

Stage 2: Grounding

The merchandising team gathers exemplars:

Positive (must be puffy):

  • Ball gown with multiple tulle layers
  • Bubble hem dress with organza
  • Tiered skirt with voluminous sleeves

Negative (must not be puffy):

  • Fitted sheath dress
  • A-line dress with no volume

Boundary (uncertain, explicit):

  • Dress with dramatic sleeves but fitted bodice (partial puffiness)

Confounders (look like puffy but aren't):

  • Quilted puffer-jacket dress (volume from insulation, not fabric structure)
  • Ruffled dress (texture, not silhouette volume)

The confounders are the precision gate. The quilted dress has volume. The ruffled dress has visual complexity. Neither is "puffy" in the intended sense. If the evaluator accepts them, the predicate fails.

Stage 3: Packaging

The team assembles A24's package: the Dress-to-Score signature, the volume-based intension, the silhouette_volume_v1 evaluator and its 3d_model_service@v2.1 dependency, with the examples and scope just described. The full seven-part record remains necessary, including provenance and authority.

This receiving test selects thresholds of 0.3 and 0.7, with boundary values flagged for review. It allows 100ms latency and $0.01 per evaluation, requires pinned dependencies and declared caching conditions, and uses the explicitly selected confounder profile K=2, M=2. Its invariant evidence must establish the promised range, monotonicity and preservation conditions independently of the handful of scores printed below. Scope is the fashion catalog; wider use in the search index remains a separate promotion.

Stage 4: Acceptance

The substrate runs the acceptance test suite:

Check 1 (Discrimination): Positive items score [0.87, 0.91, 0.78], all above τ_high (0.7). Negative items score [0.12, 0.08], all below τ_low (0.3). Boundary item scores 0.52, in the gray zone. PASS.

Check 2 (Abstain Coverage): One boundary exemplar declared, correctly in [0.3, 0.7]. Abstain policy is "flag_for_review." PASS.

Check 3 (Confounder Rejection): Quilted dress scores 0.22, ruffled dress scores 0.18. Both below τ_low. Two confounders from two subclasses (insulated, textured), meeting this example's explicitly selected K=2, M=2 profile. Under the stated default K=3 this check would fail.

Check 4 (Invariant Satisfaction): The illustrative record assumes the required invariant proofs and standing-query checks have passed. Running the displayed exemplars alone would not establish logical conservativity or absence of query effects; a new predicate can alter rules that use it.

Check 5 (Operationality): Latency p99 is 47ms (budget: 100ms). Cost is $0.002/eval (budget: $0.01). Dependencies pinned. PASS.

Check 6 (Scope Authority): Merchandising team has org-level authority for U_fashion_catalog. PASS.

Result: Admitted. AdmissionReceipt issued.

Stage 5: Deployment

The predicate enters production. Queries can now filter by puffy > 0.7:

Query: "puffy dresses under $200"
Constraints: [puffy(d) > 0.7, price(d) < 200, category(d) = "dress"]
Predicate receipts: [AdmissionReceipt(puffy) ✓, ...]

Stage 6: Monitoring

After deployment, the system monitors for drift:

DriftMonitor(puffy) = {
  calibration_check: weekly,
  last_check: 2026-01-22,
  exemplar_scores: [within margin],
  result: STABLE,
  
  revalidation_triggers: [
    "3d_model_service version change" → reassess affected behavior, invariants and operationality,
    "calibration drift > 0.1" → recheck discrimination
  ]
}

Stage 7: Revision

Suppose later catalog use reveals ruffled dresses among the results. The team proposes a more explicit intension: volume from tulle, organza or layering, excluding ruffles and quilting. Whether that revision preserves the existing filter is a further question.

The revised predicate is puffy_v2. Compatibility check:

CompatibilityTest(puffy_v1, puffy_v2) = {
  v1_positives_under_v2: all still positive ✓
  v1_negatives_under_v2: all still negative ✓
  confounders_under_v2: rejected more strongly ✓
  
  compatibility_class: "refinement"
  witness: ∀x. puffy_v2(x) > 0.7 ⇒ puffy_v1(x) > 0.5
}

The displayed implication, if verified, takes a v2 score above 0.7 only to a v1 score above 0.5. It does not preserve the deployed v1 filter above 0.7. Agreement on the listed exemplars also does not prove the universal implication. The migration must therefore distinguish the claimed relation from the stronger filter-preservation obligation.

VersionWitness = {
  from: puffy_v1, to: puffy_v2,
  direction: "refinement",
  transport: {
    v2_above_0_7_to_v1_above_0_5: "requires verification of displayed implication",
    v2_above_0_7_to_v1_above_0_7: "not established",
    existing_v1_queries: "remain pinned pending compatibility check or migration"
  }
}

This version record makes the unresolved compatibility visible. Keeping the old evaluator and its required dependencies available allows the old query to continue under its existing contract while the new predicate is assessed; labeling a version “refinement” cannot do that work.

What Rejection Looks Like

Not every proposal passes. Consider "stylish":

ProposedPredicate: "stylish"
Exemplars: { positive: [SKU_X, SKU_Y, SKU_Z], negative: [SKU_W], boundary: [], confounders: [] }

Check 1 (Discrimination): Positive scores [0.72, 0.68, 0.81]. Negative score 0.54. With τ_low=0.4, τ_high=0.6, the negative item is in the gray zone, not clearly rejected. FAIL.

Check 2 (Abstain Coverage): no_abstain_zone claimed without justification. No boundary exemplars. FAIL.

Check 3 (Confounder Rejection): No confounders provided. FAIL.

Example(Type-Directed Rejection)
RejectionReceipt(stylish) = {
  rejected_as: "predicate",
  failing_checks: [
    { check: "discrimination", reason: "negative exemplar in gray zone" },
    { check: "abstain_coverage", reason: "no_abstain claimed without justification" },
    { check: "confounders", reason: "none provided" }
  ],
  
  alternative_standing: "preference_signal",
  alternative_contract: {
    type: SoftPreference,
    usage: "ranking boost, not hard filter",
    requirements: ["user attribution", "no global claims", "display as suggestion"]
  },
  
  remediation_if_predicate_desired: [
    "Add boundary exemplars for 'somewhat stylish' items",
    "Provide confounders from at least 2 subclasses",
    "Narrow scope: 'stylish_for_cocktail' may be tractable"
  ]
}

"Stylish" is not rejected as meaningless. It is routed to a different standing. Preference signals have different obligations than predicates: they can influence ranking but cannot be hard filters, they carry user attribution, and they make no global claims. The system offers a path forward, not just a gate.

Partial Backfill and Undefined Semantics

Admission to U does not authorize application outside its declared scope. In this example the query layer represents such an application as undefined, even if the underlying code could return a score:

q is admitted to U.
Item r exists in V where V ∩ U = ∅.

q(r) = undefined (not false)

A fashion filter supplies no classification of the furniture in this example. Separately, an item inside the admitted scope may still lack an evaluation because backfill is unfinished. Neither absence establishes a negative result; the record should distinguish outside-scope use from an outstanding evaluation.

How undefined behaves:

Query Modeundefined Behavior
FilterExcluded unless query sets INCLUDE_UNKNOWN
Rank boostNeutral (0 boost) unless query sets EXPLORE_UNKNOWN
AggregationExcluded from counts; reported separately as unknown_count

Queries must handle undefined explicitly. The result must disclose its treatment of unknown items, including exclusions.

Promotion

Admission is local. Promotion widens scope. The checks are different.

Promotion Criteria

P1. OVERLAP_AGREEMENT: If U ∩ V ≠ ∅, q's evaluations must agree on shared items.

If disagreement: PromotionObstruction with options (reconcile, fork, arbitrate).

P2. CONFLICT_RATE: Disagreement with V's existing predicates below threshold.

If exceeded: ConflictReport with specific predicate pairs.

P3. AUTHORITY: The receiving context must authorize the wider use under its applicable rules. A numbered higher tier is one possible organizational policy, not a mathematical condition on scope.

P4. INVARIANT_EXTENSION: q satisfies V's external invariants (V may be stricter).

Promotion also requires the receiving context's applicable evidence, abstention, confounder, operational and compatibility checks. Agreement on a sampled overlap is not a universal agreement proof. A predicate that works in one context may conflict with another. The system produces obstruction artifacts when promotion fails, with the same remediation spirit as A28's sense obstructions.

Consequence

A29 makes admission a decision with inspectable grounds. Its six checks do different work, and a later use must respect what each has actually established.

Under the example’s assumed checks, the catalog gains a filter it previously lacked. The admission record connects that use to its exemplars, confounders, invariants and scope. A later revision has to answer for the relation it claims to preserve: the displayed implication from v2 to v1 does not keep the existing threshold intact. The old query therefore remains pinned while that stronger obligation is investigated.

The unsuccessful “stylish” proposal has a different future. Its failed discrimination check does not make preference meaningless; the proposed ranking contract gives it a restricted use with different grounds. Admission can enlarge what the catalog offers without pretending that every useful distinction has earned the same reliance. These are manuscript requirements and assumed demonstration results, not a machine-verified deployment of the complete gate.

Chapter 28 takes up T7: Contextual Equivalence. When are two things the same? The answer is: sometimes, in some contexts, for some purposes. A30 formalizes scoped equivalence.

Search the book

Use ↑ ↓ to move through results; Escape to close.

Search every published chapter, section and reference.

    In this chapter