Res Agentica
Reading

No saved reading position.

Reading

No saved reading position.

Scoped Equivalence

NYC is not always NYC

15 min read
Aa
Text size
A30Written accountWhen two things are the same in one context and different in another.

The reference of "evening star" would be the same as that of "morning star," but not the sense.

— Gottlob Frege, On Sense and Reference (1892); Max Black translation (1952)

This chapter formalizes scoped equivalence as a typed, witnessed coercion with proof obligations, defining Anchor A30 (Scoped Equivalence). The central construction treats contexts as a poset under refinement and scopes as downward-closed regions within it; an equivalence between entities is valid only when the current context lies within the declared scope. A30 closes the loop opened by A10 (Witnessed Equivalence) and A16 (Transport Certificates), adding enforcement discipline via composition rules, scope violation artifacts, and an equivalence admission procedure. The narrative motivation appears in Volume I, Chapters 6 and 8, where context-free substitution is diagnosed as a systematic failure mode of both the string and schema paradigms.

The Substitution That Seemed Safe

Consider a hypothetical real estate search system merging two data sources. Its source conventions, identifiers and counts below are stipulated for the example; they are not observations of a postal service, listing provider or geocoder. One contains Manhattan apartment listings; the other contains five-borough property records. Both use location fields. The integration team writes:

if location.city == "NYC" or location.city == "New York City":
    # treat as same

The test ignores the convention under which each source uses the name.

In the stipulated address source, both terms identify the five-borough city. In the stipulated listing source, “NYC” is used for Manhattan. That narrower convention is not asserted as a general real-estate usage. The terms look identical. Their denotations diverge.

An unconditional substitution can turn a Manhattan-only request into a five-borough search, or narrow the wider request to Manhattan. Which error occurs depends on the direction of the mapping. The declared coverage makes that difference inspectable.

This is T7, the seventh touchstone. We named it in Chapter 5: "Equivalence requires manual normalization." The pathology is not missing data. It is missing structure. This integration has omitted the source convention that would distinguish the two uses. A schema or a manual mapping can retain it.

A30 makes that structure explicit.

Terms and Entities

The first mistake is conflating terms with entities.

A term is a string: "NYC," "New York City," "Big Apple." Terms appear in documents, listings, query inputs, and user interfaces. They are syntax.

An entity is a referent: the municipality, the island, the colloquial region. Entities are what terms denote. They are semantics.

The relationship between terms and entities is mediated by context. In one context, "NYC" denotes the five-borough city. In another, it denotes Manhattan. The term is the same. The denotation differs. The two-step model makes this precise:

Two-Step Model

Step 1: Denotation Map

Given term tt and context CC, the denotation map returns an entity:

⟦t⟧C→Entity\llbracket t \rrbracket_C \to \text{Entity}

Step 2: Entity Equivalence

Equivalence is defined between entities, not terms:

e1∼Se2e_1 \sim_S e_2

where SS is the scope in which the equivalence holds.

The NYC problem becomes precise:

Context⟦"NYC"⟧⟦"New York City"⟧Same referent?
Postalnyc@address_sourcenew_york_city@geocoderYES (different IDs, same referent)
Real estatemanhattan@listing_sourcecity_of_new_york@geocoderNO (different referents)
Legalnyc@registrynew_york_city@registryYES (different IDs, same referent)

Under the proposed A30 contract, the postal and legal contexts require a correspondence: the terms denote different entity IDs from different sources for the same underlying referent, so scoped equivalences bridge them. Real estate fails earlier: the terms denote genuinely different referents (manhattan vs city_of_new_york), and the stated identity substitution has no support.

The postal case is not trivial. "NYC" from address_source data and "New York City" from the example geocoder are different entity IDs. Substitution requires a scoped equivalence:

nyc@address_source∼postalnew_york_city@geocoder\text{nyc@address\_source} \sim_{\text{postal}} \text{new\_york\_city@geocoder}

with the witness and property-specific transport certificate required by A30.

In real estate context, the terms denote genuinely different entities (manhattan vs city_of_new_york). No exact identity witness between the stipulated borough and city is supported. A task might use a separately defined approximation, but that is a different relationship. The system that ignores this distinction will conflate manhattan with city_of_new_york. The error changes which properties the search covers.

Scope as Proof Obligation

A10 introduced witnessed equivalence with a scope field. A16 established that transport along equivalences requires certificates. A30 requires the receiving operation to check that scope before using the witness.

A30
A30: Scoped Equivalence

Context Structure: Contexts form a poset (C,≤)(C, \leq) under refinement. C1≤C2C_1 \leq C_2 means C1C_1 is more specific than C2C_2. For a fixed property footprint and refinement maps proved to preserve its evidence, a scope SS can be represented as a downward-closed region: if C∈SC \in S and C′≤CC' \leq C, then C′∈SC' \in S. Arbitrary added constraints or new properties do not inherit the witness through this definition.

Entity Equivalence:

e1∼Se2:=witnessed equivalence between e1 and e2, valid in scope Se_1 \sim_S e_2 := \text{witnessed equivalence between } e_1 \text{ and } e_2 \text{, valid in scope } S

Components:

ScopedEquivalence = {
  lhs: Entity,
  rhs: Entity,
  scope: Scope,                      // downward-closed region in context poset
  kind: identity | isomorphism | equivalence | approximation,
  witness: EquivalenceWitness,       // evidence
  transport_certificate: {
    properties_preserved: [Property...],
    properties_not_preserved: [Property...]
  }
}

Coercion Formalization:

Entity tokens are values, not types. A witness ww for e1∼Se2e_1 \sim_S e_2 supports a specified transport of data or propositions about them. Schematically:

transportw,S:Data⁡F(e1)→Data⁡F′(e2)\text{transport}_{w,S} : \operatorname{Data}_F(e_1) \to \operatorname{Data}_{F\prime}(e_2)

where F and F′ identify the aligned property footprints and the map’s correctness has been established

with proof obligation: current_context∈S\text{current\_context} \in S.

Established use outside the scope is a ScopeViolation. An unfinished membership check leaves the transport unsupported without establishing that violation. Neither result permits successful transport under this contract.

Transport Rule:

For any substitution e1→e2e_1 \to e_2 in context CC:

  1. Retrieve witness ww for e1∼Se2e_1 \sim_S e_2
  2. CHECK: C∈SC \in S (scope membership)
  3. CHECK: required_properties ⊆\subseteq transport_certificate.properties_preserved
  4. Check the actual evidence, maps, validity conditions and required authority; if established, perform the permitted transport and record it
  5. Distinguish ScopeViolation from missing or inconclusive evidence, expired grounds and a property-map failure

Naming convention: A context (e.g., postal_context) is a specific point in the poset where a query executes. A scope (e.g., postal_scope) is a downward-closed region containing one or more contexts. The check postal_context ∈ postal_scope asks whether the query's context lies within the equivalence's valid region.

A type discipline can encode these obligations. An arbitrary entity ID is not itself a type, and an untyped implementation can also enforce an explicit transport contract. Using the coercion requires discharging a proof obligation. If the current context is not in scope, the obligation fails. The system refuses the substitution.

What Violations Look Like

Example(Scope Violation Artifact)
Operation: Substitute "New York City" for "NYC"
Context: real_estate_context

ScopeViolation = {
  attempted_substitution: ("NYC", "New York City"),
  in_context: real_estate_context,
  denotations_in_context: {
    lhs: entity:manhattan@listing_source,
    rhs: entity:city_of_new_york@geocoder
  },
  reason: "Terms denote different entities in this context",
  
  candidate_witnesses: [
    { 
      witness_id: "postal-nyc-equiv",
      equivalence: "nyc@address_source ~_{postal} new_york_city@geocoder",
      applicable: false,
      reasons: [
        "real_estate_context ∉ postal_scope",
        "denotations differ: manhattan@listing_source ≠ nyc@address_source"
      ]
    }
  ],
  
  remediation_options: [
    "Narrow query to postal context (where A30 equivalence applies)",
    "Query without substitution, using 'NYC' as written",
    "Define and justify a different relationship if the intended operation permits one"
  ]
}

The artifact is auditable. It explains why the substitution failed. It shows which witnesses exist and why they do not apply. It offers paths forward. The user or downstream system can make an informed decision.

Example(Valid Substitution via A30)
Operation: Substitute "New York City" for "NYC"
Context: postal_context

System:
  1. Compute denotation: ⟦"NYC"⟧_{postal} = entity:nyc@address_source
  2. Compute denotation: ⟦"New York City"⟧_{postal} = entity:new_york_city@geocoder
  3. Different entity IDs; search for equivalence
  4. Found: nyc@address_source ~_{postal_scope} new_york_city@geocoder
  5. Check proof obligation: postal_context ∈ postal_scope? YES
  6. Check the witness and required property map before transport

TransportReceipt = {
  substitution: ("NYC", "New York City"),
  in_context: postal_context,
  denotations: (entity:nyc@address_source, entity:new_york_city@geocoder),
  equivalence_used: "nyc@address_source ~_{postal} new_york_city@geocoder",
  witness: postal_authority_ruleset,
  properties_preserved: [delivery_zone, mailing_address],
  properties_not_preserved: [source_id]
}

The postal substitution is not trivial. The terms denote different entity IDs from different sources. A30 bridges them with a scoped equivalence, produces a receipt, and documents which properties survive transport. The system earns the substitution; it does not assume it.

An unconditional string rule loses the source distinction. A correspondence table can preserve it if its keys, restrictions or surrounding code express the relevant scope. Where no adequate correspondence is available, the recipient must either investigate or leave the proposed substitution unfinished. The proposed contract makes that obligation inspectable across implementations.

A30 makes scope an explicit condition of the substitution. Its validity still depends on the source interpretation and evidence.

Composition Rules

Identity graphs compose. If e1∼Se2e_1 \sim_S e_2 and e2∼Te3e_2 \sim_T e_3, what is the relationship between e1e_1 and e3e_3?

Without explicit rules, transitivity leaks scope. The system might conclude e1∼e3e_1 \sim e_3 with unbounded scope, reintroducing the very conflation A30 was designed to prevent.

Composition Rules

Transitive Closure:

If e1∼Se2e_1 \sim_S e_2 and e2∼Te3e_2 \sim_T e_3, then:

e1∼S∩Te3e_1 \sim_{S \cap T} e_3

The composed equivalence holds only where both original equivalences hold. Scope is intersection, not union.

Conditions:

  • Transport operators and their property maps must compose; intersecting property names alone does not prove this
  • Witnesses must justify the same intermediate interpretation and the claimed composite; approximate or statistical evidence requires its own error and dependence rule

If conditions fail: no composed equivalence. The system stores the path but does not claim the composition.

Kind Meet:

When composing equivalences of different kinds, the result is the weaker kind:

Kind 1Kind 2Composed Kind
identityidentityidentity
identityisomorphismisomorphism
identityapproximationapproximation
isomorphismapproximationapproximation

The table is a non-escalating label policy. A metric approximation may accumulate error, and an arbitrary similarity relation need not compose at all. No bound follows merely from keeping the weaker label.

Multiple Equivalences:

If e1∼Se2e_1 \sim_S e_2 and e1∼Te2e_1 \sim_T e_2 with S≠TS \neq T (same pair, different scopes), the system selects the tightest scope containing the current context. Scopes are ordered by set inclusion; "tightest" means minimal-by-inclusion among scopes that contain CC.

  • If C∈SC \in S and C∉TC \notin T: use S-equivalence
  • If C∈SC \in S and C∈TC \in T and S⊂TS \subset T: use S-equivalence (tighter)
  • If C∉SC \notin S and C∉TC \notin T: ScopeViolation with both as candidate_witnesses
  • If both contain C but neither contains the other, there may be multiple minimal scopes. Compare their evidence and maps or report the unresolved choice; do not claim a unique tightest witness

These rules prevent the "just compose everything" failure. Scope flows correctly through chains. Enforcing them requires checking the actual witness and comparison, not just carrying its scope label through the path.

Admitting Equivalences

Scoped equivalences are not declared by fiat. They are admitted with evidence, following the same discipline as predicate admission (A29).

Equivalence Admission

EquivalenceAdmission(e1∼Se2)(e_1 \sim_S e_2) requires:

  1. Positive Evidence: Witness supporting the equivalence (authority ruling, ontology, measurement, declaration)

  2. Counterexamples: Record known contexts where the proposed relationship does not hold. Their absence does not prove the scope valid; a universally supported identity need not have a counterexample.

    CounterexampleWitness = {
      context: C,
      reason: "different denotations" | "different properties" | "authority conflict",
      evidence: Reference
    }
    
  3. Scope Declaration: Explicit scope SS, defined as region in context poset. Scope must exclude counterexample contexts.

  4. Transport Certificate: Properties preserved and not preserved under substitution.

  5. Authority Tier: Who can admit this equivalence (local, organizational, global).

Result:

  • Admitted(e1∼Se2(e_1 \sim_S e_2, EquivalenceReceipt))
  • Rejected(reason, conflicting_evidence)
  • Provisional(permitted_restricted_use, conditions)
  • Inconclusive(unchecked_obligations, reason)

Scope widening follows the same pattern as predicate promotion. To extend e1∼Se2e_1 \sim_S e_2 to e1∼S′e2e_1 \sim_{S'} e_2 where S⊂S′S \subset S':

  • New evidence required for S′∖SS' \setminus S
  • If counterexample exists in S′∖SS' \setminus S: widening rejected
Example(Scope Widening Rejection)
ScopeWideningRequest = {
  equivalence: "nyc@address_source ~_{postal_scope} new_york_city@geocoder",
  current_scope: postal_scope,
  requested_scope: all_contexts,
  proposed_use: "Substitute the two location terms after their context-specific denotation"
}

Result: REJECTED

CounterexampleWitness = {
  context: real_estate_context,
  reason: "In real estate, 'NYC' denotes manhattan@listing_source; 
           'New York City' denotes city_of_new_york@geocoder. 
           Different entities.",
  evidence: [real_estate_listing_analysis_ref]
}

Remediation: "Widening to all_contexts impossible while counterexample exists.
              You may admit a separate witness for legal_scope, giving two 
              equivalences whose combined coverage is (postal ∪ legal)."

Here the rejected request would carry the term substitution into contexts where the terms denote different places. That does not make the original address-source identifiers cease to co-refer. Two witnesses may instead cover a union of scopes; they can be combined piecewise if their interpretations and maps agree wherever the scopes overlap and the representation supports the combination. The union of scope sets alone does not establish those conditions.

Three Failure Modes

T7 named the pathology: context-free equivalence handling. The contract addresses three specific failure modes:

Silent Conflation: The system treats x=yx = y unconditionally. In some contexts, this is wrong. The real estate integration conflates manhattan and city_of_new_york because it lacks scope. Users see wrong results. No artifact records the error.

A30 prevents conflation by requiring scope checks. The coercion fails if the context is not in scope. The failure is explicit.

Silent Separation: The system treats x≠yx \neq y unconditionally. In some contexts, this is wrong. A postal search fails to find "New York City" records when querying "NYC" because no equivalence is recorded. Users miss valid results. No artifact records the gap.

Recording valid scoped equivalences can recover matches known to the index. A30 does not guarantee discovery of every missing correspondence. Where the equivalence holds (postal context), the system expands queries and merges results. The expansion is witnessed.

Scope Leak: The system uses x∼Syx \sim_S y transport in context TT where T∉ST \notin S. The equivalence is recorded but applied too broadly. Scope leak is reusing a valid equivalence in the wrong scope. This is the integration team's error: they found a postal equivalence and applied it in real estate context. The system does not catch the mismatch.

A30 prevents scope leak by making scope a proof obligation. The coercion e1→e2e_1 \to e_2 requires C∈SC \in S. If the check fails, the system refuses. No silent leak.

A related search policy may expand “blue” to include navy results. Inclusion in that query does not establish symmetric identity between blue and navy, nor an inventory-merging right. "Dr. Smith" in a hospital and "Dr. Smith" in a university share a string; the name alone establishes neither one person nor two. Denotation and co-reference require the relevant evidence. A28 handles sense disambiguation; A30 handles scope enforcement. Together, they specify distinct obligations for interpretation and use; they do not prove that every identification has been found or correctly resolved.

What the User Sees

Example(Cross-Context Query)
Query: "Properties in NYC"
Candidate contexts: [real_estate_context, postal_context]  // inferred from data sources
Status: Ambiguous; user must select declared context

System response:

ContextDisambiguation = {
  term: "NYC",
  denotation_by_context: {
    real_estate_context: entity:manhattan@listing_source,
    postal_context: entity:nyc@address_source
  },
  message: "Query spans multiple contexts where 'NYC' denotes different entities.",
  options: [
    { context: real_estate_context, meaning: "Manhattan", result_preview: "847 listings" },
    { context: postal_context, meaning: "All five boroughs", result_preview: "4,231 listings" }
  ],
  default_action: "none (context declaration required)"
}

This proposed interface asks the user to choose between the two coverages before it runs the search. A disclosed default or an exploratory display could serve a different contract. In either case, the output must preserve which coverage governed the result.

Consequence

A30’s entity equivalence supports a declared transport under its hypotheses. Terms can also have equivalence relations; this example needs their denotations kept distinct from the entities they identify.

A system that treats "NYC" and "New York City" as unconditionally equivalent will conflate where it should separate. A system that treats them as unconditionally different will separate where it should conflate. Tracking scope can prevent reuse of a witness outside its established domain. An incorrect identification inside the declared scope remains incorrect; a missing witness remains missing. The records make those questions available for examination.

Manual normalization can already be scoped and well supported. A30 proposes a common record for those grounds and permitted uses. The recipient gains a substitution it need not establish again, while retaining the means to see where it stops. A pricing question cannot inherit a name-normalization witness merely because the same two records occur in both operations. It needs grounds for the property it proposes to carry across.

Chapter 29 asks the final touchstone question: what happens when named entities participate in events? "Alice introduced Bob to Carol" is not three strings plus a verb. It is a structured relation with roles. Identifying Bob does not establish whether he introduced someone or was introduced. The event must preserve who occupied each role.

Search the book

Use ↑ ↓ to move through results; Escape to close.

Search every published chapter, section and reference.

    In this chapter