Proposal and Certification
The two operations of the Third Mode
Aa
Sameness occurs in substance, while likeness occurs in quality.
A proposal can save a search without settling the question for which the result will be used. A19 specifies the candidate-producing operation; A19b specifies the checks that can establish a scoped conclusion about its output. The difference becomes practical when the recipient proposes to merge records, carry a property between them or use a recommendation without making either stronger claim. Volume I, Chapter 7 (The Witness Protocol) develops the related institutional question.
The Embedding Empire's Claim
A useful retrieval can place records beside one another that their keeper would not have thought to compare. The saving is real even when the records remain different. Two descriptions may be close enough to make a recommendation interesting and insufficiently supported to combine their inventory.
Combining the inventory asks a further question: what establishes that the records concern the same stock, in the same units and at the relevant time? A ranking score can contribute evidence to an identification procedure. It cannot supply all those conditions by being high. The recipient must identify the operation it proposes to perform before deciding what the resemblance warrants.
T2: Morning Star, Evening Star
The classical problem comes from Frege1. "Morning Star" and "Evening Star" both refer to Venus. Are they the same?
A retrieval may place the names near one another. If it helps someone find the astronomical connection, that is useful. No particular model, score or retrieval ranking is reported here. The following records illustrate a proposed identification and the uses that identification could support.
But the similarity does not tell us:
- In what sense are they the same? Identity of reference? Identity of meaning? Identity of use?
- In what scope? Colloquial usage? Scientific taxonomy? Navigation system? Poetry anthology?
- What operations does this sameness license? Can I substitute one for the other in a star chart? In a database of astronomical objects? In a poem about longing?
A witness answers these questions.
EquivalenceWitness(morning_star, evening_star) = {
relation_kind: ReferentialIdentity,
verification_regime: attested, // stipulated orbital-identification report
evidence: [IAU_designation(Venus), ephemeris_orbital_parameters],
scope: astronomical_objects,
transport_rights: {
orbital_position: Full,
apparent_magnitude: Full,
cultural_association: None,
poetic_meaning: None
}
}
The record must distinguish its relation from its verification regime. ReferentialIdentity names the claim about the object; decidable, probabilistic and attested describe how supporting evidence is checked. Position and apparent magnitude also require the relevant time and observing conditions. Identifying Venus does not make observations at different times interchangeable.
Per A16, an equivalence witness is not a binary flag; it is a license specifying which properties may be transported across the equivalence and which may not. The common referent can support a calculation with aligned time and observing conditions. A question about which name appears in a poem still requires the name that was used.
What Similarity Actually Measures
A similarity score depends on the representation, training objective and comparison rule. It may reflect linguistic co-occurrence, visual features or other learned structure. Its usefulness must be evaluated for the intended task; closeness in the representation is not a general identity test.
A calibrated identification procedure can use such a score. Its result needs the procedure, evaluated domain and error interpretation. This lets the recipient assess an actual identification claim instead of treating either a statistical method or an exact-looking label as sufficient assurance.
Consider two product listings:
- Listing A: "Vintage silk evening gown, emerald green, size 6"
- Listing B: "Green silk formal dress, retro style, size S"
Embedding similarity: High. Both are green silk formal dresses with vintage/retro styling.
Are they equivalent?
- Same product? Unknown. Could be different items.
- Same size? Unknown. Size 6 and Size S might not match.
- Same condition? Unknown. "Vintage" might mean used; "retro style" might mean new.
- Can I merge their inventory counts? Absolutely not.
The similarity is real. The equivalence is not established. Merging them without establishing the stock relation risks adding distinct inventories or counting the same stock twice. The similarity has supplied candidates for that investigation, not its result.
The Proposal Operator
Embeddings are excellent at generating candidates. The Third Mode uses them for exactly that purpose.
A proposal operator P is a function:
P : PromptOrObject → List[(Hypothesis, Score)]
where:
- PromptOrObject is a query (structured or natural language) or an object to find neighbors of
- Hypothesis is a CandidateItem or CandidateClaim (including equivalence hypotheses)
- Score is a real number indicating confidence or relevance
Implementation modes:
- Embedding similarity:
score(q, x) = cos(embed(q), embed(x)) - Retrieval: BM25, TF-IDF, learned retrievers
- Statistical pattern matching: n-gram overlap, fuzzy matching
- Learned rankers: neural rerankers, cross-encoders
Key property: P produces hypotheses, not certified truths.
A proposal procedure can save checking work by reducing the candidates examined. Suppose, illustratively, it reduces a million records to a hundred. The saving depends on the cost of proposing and checking, and on which acceptable candidates the filter leaves out. Some checks are cheap; some candidate generators are expensive. The roles remain distinct without a universal speed ordering.
Clicks or a successful recommendation can give reasons to retain a retrieval procedure. They do not establish that its output is safe to merge into another record. A service using the result for that purpose must examine the identification and the properties it intends to carry across. Better retrieval can reduce that work without completing it.
The Certification Contract
Certification takes the proposed claim and the receiving context together. The question is whether the required grounds have been established for this use.
A certification operator C is a function:
C : (CandidateClaim, Context) → CertificationResult
where CandidateClaim = Claim with a declared relation_kind (identity, equivalence,
recommendation adjacency, etc.) and the objects it relates
where CertificationResult =
| Success {
relation_kind: CandidateClaim.relation_kind,
witness: π,
witness_class: WitnessClass,
scope: Scope,
transport_rights: PropertyMap
}
| NotCertified {
reason: NonCertificationReason,
evidence: DiagnosticEvidence,
remediation: Option[Guidance]
}
where NonCertificationReason =
| MissingEvidence { what_is_missing: EvidenceSpec }
| InvariantViolation { violated: Invariant, obstruction: ObstructionWitness }
| ScopeMismatch { claimed_scope: Scope, valid_scope: Scope }
| TransportDenied { property: Property, reason: String }
| Inconclusive { unfinished_obligation: String, resource_or_evidence_limit: String }
Key property: An unsuccessful certification is not necessarily a false claim. The result distinguishes a proved violation, missing evidence and an incomplete check. None acquires certified standing merely because the procedure has returned.
Scope of certification: The procedure can establish a claimed identity, a specified recommendation relation or another property under the receiving contract. The relation kind is separate from the A2c witness class. A recommendation may be certified on adequate statistical grounds without becoming an equivalence or authorizing an inventory merge.
The nested result representation above maps to Appendix I as follows: Success records completed certification; InvariantViolation, a demonstrated ScopeMismatch or an established TransportDenied records the failed requirement. MissingEvidence and nested Inconclusive preserve Appendix I's incomplete outcome. NotCertified retains these distinct reasons without turning an unestablished obligation into a proved violation.
Success returns:
- The relation kind certified (what relationship was established?)
- A witness typed per A2c (the claim has standing)
- A witness class (how was the supporting evidence checked?)
- A scope (where is this valid?)
- Transport rights (what operations does this license?)
NotCertified returns:
- A structured reason distinguishing a demonstrated failure from an unestablished obligation
- Evidence sufficient to diagnose the problem
- Optionally, guidance on what would make certification succeed
An unsuccessful attempt must retain what was found and what remains unexamined. Evidence of a violated invariant can justify rejection; an unavailable record leaves a different task for the next investigator. Neither result should disappear into an undifferentiated “no.”
Candidate from P: (morning_star, evening_star, similarity=0.87)
Certification attempt in astronomical_database context:
C(morning_star ≃ evening_star in astronomical_database) =
Success {
witness: [IAU_designation, ephemeris_data, orbital_parameters],
relation_kind: ReferentialIdentity,
witness_class: attested, // stipulated report; its grounds remain examinable
scope: astronomical_objects,
transport_rights: {
orbital_position: Full,
apparent_magnitude: Full,
observation_time: None,
cultural_name: None
}
}
Certification attempt in poetry_corpus context:
C(morning_star ≃ evening_star in poetry_corpus) =
NotCertified {
reason: ScopeMismatch {
claimed_scope: poetic_substitution,
valid_scope: astronomical_reference_only
},
evidence: "Different cultural and poetic associations;
substitution changes meaning of verse",
remediation: "Certify as astronomical_identity only,
or provide poetic equivalence witness"
}
Same candidate, different contexts, different certification results. The proposal operator found the connection; the certification operator determined what operations it licenses.
Propose, Certify, Glue
The Third Mode's operational architecture separates concerns.
1. PROPOSE
P(query) → [(candidate_1, score_1), ..., (candidate_n, score_n)]
Retain the candidate domain, ranking rule and any search limits.
2. CERTIFY
C(candidate_claim, context) → CertificationResult
Successful certification records the relation, scope and grounds.
Demonstrated violations and incomplete checks remain distinguishable.
Evidence may be exact, statistical or attested under the contract.
3. GLUE, WHEN THIS IS THE REQUIRED COMPOSITION
Glue(cover, local_sections) → GlobalClaim | ObstructionWitness
| RejectionWitness | Inconclusive
Establish the sheaf, cover, maps and effective construction.
Exact matching yields the amalgamation under those hypotheses.
Invalid inputs, unequal restrictions and unfinished work differ.
Certification does not automatically make its results a matching family. The next operation must establish that the results inhabit the proposed construction and that its comparisons have been performed. A recommendation relation may be useful on its own; it need not acquire a sheaf interpretation before it can serve the declared recommendation task.
A better embedding can improve evidence used in checking. A proof or a finite exact comparison can discharge a different obligation. The division is among the claims these operations establish, so the same technology may contribute at more than one stage. What must not cross unnoticed is a change in the assurance being claimed.
A user searches for "dresses like item X."
1. Propose:
P(similar_to(X)) → [
(dress_A, 0.94), (dress_B, 0.91), (dress_C, 0.89),
(dress_D, 0.87), (dress_E, 0.85), ...
]
From 50,000 dresses, P returns top 50 candidates. Time: 100ms.
2. Certify: For each candidate above threshold 0.80:
C(dress_A ~style X in catalog_view) =
Success {
witness: shared_attributes(silhouette, fabric_type, occasion),
relation_kind: RecommendationAdjacency, // not equivalence
witness_class: decidable, // stipulated finite attribute checks
scope: catalog_view,
transport_rights: { recommendation: Full, inventory_merge: None }
}
C(dress_B ~style X in catalog_view) =
NotCertified {
reason: InvariantViolation {
violated: occasion_compatibility,
obstruction: "X is formal, dress_B is casual"
}
}
Time: 50ms per candidate, 2.5s total for top 50.
3. Glue: The displayed checks support recommendation relations. They have not supplied a cover, restriction maps or a sheaf of those results. This branch therefore returns the supported recommendations and their receipts. If a subsequent use requires exact composition across views, it must supply and check the missing construction; until then, exact gluing remains unestablished.
The timings in this constructed pipeline are illustrative, not measured performance. Its recommendations carry only the scope and assurance actually established by the specified checks; the sketch does not report an implemented end-to-end certification system.
What the Score Leaves to Establish
An embedding’s accuracy matters to its use. The score reports a relation in a learned representation; its relevance to a claim depends on the model, evaluation and receiving task. Neither co-occurrence in training nor usefulness for a particular query follows from the number alone.
The recipient needs the proposed relation—identity, approximation or something else—alongside its scope, property footprint and evidence. The witness class records how that evidence is checked: decidable, probabilistic or attested. Those are different questions. Asking a bare score for permission to substitute is like asking a photograph for permission to enter the building it depicts.
The omitted operation can be consequential before a merge occurs. A poor proposal may hide a useful candidate; treating its score as an identity certificate may corrupt an inventory. A larger procedure can use embeddings as evidence while checking further grounds and preserving the distinction between a demonstrated violation and work not yet done.
Where the Operations Meet
The catalog need not choose between similarity search and constraints. Existing relational, retrieval and hybrid systems can combine them. The inquiry is whether a candidate’s subsequent use inherits the checks it needs. A valid recommendation relation is already useful, even if it supplies no warrant for merging stock or transferring a price.
Consequence
The narrowed candidate set saves work. What remains is to establish the particular use: whether the item can enter a recommendation, share an inventory count or carry a claim into another view. Success at one of those tasks need not settle the others. The record of certification should let the next recipient see both the achieved result and its remaining alternatives.
Chapter 18 asks where a proposed predicate comes from. A search can assemble a new expression from existing primitives; it must also retain what its ranking and limits kept it from examining. The saving then concerns not only the candidates found, but the constructions that another search can reuse.