Identity Maintenance
Tracking entities across representations
Aa
This chapter formalizes identity maintenance as Anchor A23: a discipline for building context-indexed equivalence over the substrate defined in A22. A23 specifies a scope algebra (meet, leq, compatible, overlap_on), a kind order over equivalence types, scope-conditional transitive closure, and conflict handling that distinguishes scoped disagreement from genuine conflict. The formalization draws on the witnessed equivalence framework of A10 and applies it to keyless joins, widening, and kind degradation. For the narrative treatment of identity as a declared, context-dependent relation rather than a global assumption, see Vol I, Chapter 7 (The Witness Protocol).
The Key Fallacy
Primary keys are useful identifiers. The fallacy is treating them as eternal metaphysics.
A primary key identifies a row within its declared relation. Trouble begins when a recipient treats that identifier as a permanent, global answer to who or what the row describes. Two records with different keys may concern the same person; the same local key may be reused by another source. Revision and entity resolution can address these problems, but the resulting match still needs a scope when another operation relies on it.
Keyless joins are joins that do not assume a shared primary key. They rely on equivalence witnesses: explicit, auditable evidence that two tokens should be treated as the same for specific purposes in specific contexts.
Identity as Declaration
We do not assume a global referent. We maintain licenses to treat tokens as the same for specific purposes in specific contexts. Canonicalization (a "golden record" or master entity) is a downstream view built atop those licenses, not the definition of identity itself.
Entity-resolution methods already investigate uncertain matches and different matching purposes1. This proposal concentrates on keeping each declared relationship, its evidence and its permitted uses available after a match has been made. It tracks what has been declared equivalent, by whom, with what evidence, in what scope. If two contexts disagree about identity, both assertions are recorded. The disagreement is data, not error.
An entity does not "exist" in the substrate until claims mention it. Identity between entities does not "exist" until an equivalence is declared with a witness. These are conditions for representation and standing in the proposed substrate, not claims that entities begin to exist when recorded.
Anchor A23: Identity Maintenance
Identity maintenance builds context-indexed equivalence over the substrate (A22).
Scope Algebra (requiring an effective representation of the chosen site and its maps):
- : the specified greatest lower bound, or an empty-scope result. If common refinements exist but no greatest one does, this interface must report that limitation rather than label them disjoint
- : U refines V (U is finer)
- : consults entity-incidence index to check if x has incident claims in meet(U,V); returns unknown if co-reference hasn't stabilized
Kind Order (poset):
Declaration: with witness creates Equivalence(e) with scope U and kind K.
Transitive Closure (scope-conditional): and where , only if the evidence restricts to W and the relation has a justified composition law there.
- A non-escalating kind label records the composite; it does not prove its validity. Approximate relations require an error and dependence rule
- Closure scope = unless explicit widening witness
Conflict Handling: and :
- If : ScopedDisagreement (not a conflict)
- A nonempty meet and incident claims identify a candidate comparison. Return ConflictWitness only after aligning the pair, relation/property, versions and restrictions and establishing opposed claims there; missing grounds leave the comparison incomplete.
Obligation: Equivalence assertions must be witnessed. Unwitnessed equivalence is not asserted. Inequivalence also needs grounds for the particular relation. Different jurisdictions or disjoint scopes can limit comparison without establishing that two tokens name different entities.
Scope Compatibility
If and , does follow?
For an equivalence relation, transitivity is part of the definition. For these records, the scopes must be compatible and the evidence must justify composition of the same relation on their common domain.
The scope algebra gives us the tools to decide. Two contexts U and V are compatible if , meaning they share a common refinement in the site structure. When they do, exact equivalence witnesses can compose if their intermediate interpretations and restrictions agree. Approximate similarity is not automatically transitive, even in one scope.
Same context: For an established transitive relation with composable evidence, and imply . Sharing U alone does not make an approximate relation transitive.
Compatible contexts: and where . The derived equivalence holds in the common refinement, not in either original scope.
Incompatible contexts: and where . No automatic closure. The system does not infer in any scope.
This prevents a common failure mode: chaining equivalences across incompatible scopes to conclude that everything is equivalent to everything. The scope-conditional closure is the discipline that keeps identity maintenance honest.
Kind Degradation
The specification orders kind labels to prevent escalation of permitted operations. Composing the actual relationships requires more: for a metric bound, two errors may add; for statistical bounds, dependence matters. A minimum label records no such calculation.
If asserts with kind "identity" and asserts with kind "approximation," the derived equivalence has kind "approximation." You cannot chain approximations into identities. The kind label does not measure information-theoretic loss.
This matters for transport. An identity-kind equivalence can license broad transport, but it still respects the transport certificate and footprint (A16) and context invariants. An approximation-kind equivalence licenses transport of some properties with acknowledged loss. The kind tells downstream consumers what standing the equivalence has.
Conflict and Disagreement
What happens when but ?
Case 1: Disjoint scopes. If , this is not a conflict. Different contexts, different truths. The system records a ScopedDisagreement artifact:
ScopedDisagreement(
x: token_a,
y: token_b,
equivalent_in: U,
inequivalent_in: V,
meet: ⊥
)
This is informative, not erroneous. It tells you that identity depends on context.
Case 2: Overlapping scopes. If , check whether the overlap is relevant to x and y. The predicate returns true if the meet contains claims incident to x.
If overlap_on returns false under a completed check, that overlap supplies no conflict for these entities. Unknown leaves the comparison unsettled; it does not establish that the entities fail to overlap.
If overlap_on returns true, compare the claims. A conflict requires opposed claims about the same pair, relation and property under aligned versions and restrictions. Incidence merely identifies a comparison worth making. Once that opposition is established, the system produces a ConflictWitness:
ConflictWitness(
entities: [x, y],
equivalence_in_U: e1,
inequivalence_in_V: e2,
overlap: meet(U, V),
incident_claims: [...],
compared_relation_and_property: ...,
aligned_versions_and_restrictions: ...,
opposition_evidence: ...,
resolution_options: [
"restrict e1 scope to exclude overlap",
"restrict e2 scope to exclude overlap",
"produce obstruction and escalate"
]
)
The conflict is explicit, computed, and stored. It does not disappear because someone ran a query.
T2: Morning Star and Evening Star
The canonical identity problem2. "Morning Star" and "Evening Star" are names for the same celestial body (Venus), but the discovery that they refer to the same object was informative. The names had different senses, different contexts of use, and different associated properties.
In a schematic record, suppose the observation comparison has been checked for the declared astronomical properties and times. Two entity tokens are linked as follows:
declare_equivalence(
x: morning_star,
y: evening_star,
context: U_astronomy,
witness: EphemerisAlignment(
observation_set: <stipulated observation record>,
method: "orbital_parameter_match"
),
kind: identity
)
The result is an Equivalence(e_venus) with scope U_astronomy and kind identity.
Transport: The certificate on e_venus specifies which properties can be transported. Orbital period and mass may be reused under that checked comparison. Position additionally requires the observation time and coordinate convention: identifying the object does not identify every recorded position.
Poetic associations: no. The certificate does not cover cultural properties. A query asking "what poems mention evening_star?" cannot use e_venus to return poems about morning_star. Different scope, different properties, different standing.
Scope restriction: In this example, no equivalence supporting name substitution in U_poetry has been established. A query asking "are these the same?" in U_poetry returns: Unknown. No witness exists in that scope. The system does not guess.
T7: NYC Strings
"NYC," "New York City," "New York, NY," "Manhattan," "10001." In different contexts, these may or may not refer to the same entity.
In a constructed routing example, suppose a documented address rule accepts two strings for the same destination after the other routing fields have been checked. Its scope is that rule’s domain. A tax or historical query requires the relevant jurisdiction and date; it cannot acquire its answer by reusing the routing result. No particular USPS, tax-authority or tourism convention is asserted here.
The strings may denote the same city in all three uses. The error would be to claim that one verified operation had established every interpretation. Nor does calling Manhattan an approximation to New York City establish a metric or a safe composition rule. The intended task must supply that relation.
Exact equivalences can compose in a common scope when their maps and meanings fit. Otherwise the index retains the original claims without inventing a further link.
Keyless Joins
A traditional join assumes shared keys:
SELECT * FROM orders JOIN customers ON orders.customer_id = customers.id
A keyless join uses equivalence witnesses:
SELECT * FROM orders JOIN customers
ON equivalent(orders.customer_ref, customers.entity, U_billing)
The equivalent predicate:
- Looks up the equivalence index for the token pair in context U_billing
- Checks scope compatibility with the query context
- Returns true only if the witness is applicable and adequately supports this join under the receiving contract
- Returns false if inequivalence is witnessed
- Returns unknown if no witness exists either way
Unknown is not false. This is a join policy decision:
Strict policy: Join only where the required grounds establish the declared relation. An exact or statistical comparison must meet the applicable contract; its evidence class stays visible.
Best-effort policy: Include additional candidate matches under a declared evaluation and threshold, retaining their uncertainty and the uses permitted for them. Recording a score does not establish its calibration.
Exploratory policy: Return proposed joins and their scores without certified standing. Further exploration can use them; a consequential reliance requiring stronger grounds must obtain those grounds.
The policy is declared at query time. The specification permits these different declared uses; an implementation must supply their checks and enforce their boundaries.
The Equivalence Index
The substrate maintains an equivalence index for efficient queries:
EquivalenceIndex:
by_token: Map<EntityToken, Set<Equivalence>>
by_context: Map<Context, Set<Equivalence>>
by_pair: Map<(EntityToken, EntityToken), Set<Equivalence>>
closure_cache: Map<(EntityToken, EntityToken, Context), ClosureResult>
Operations:
lookup(x, y, U): Are x and y equivalent in context U? Returns Equivalence | NotEquivalent | Unknownclosure(x, U): All entities equivalent to x in context U (transitive closure)conflicts(x): All conflicts involving x across contexts
Cost:
- Declaration: an indexed write plus the required evidence and invariant checks
- Lookup: indexed retrieval plus applicability checks; the data structure determines the lookup bound
- Closure: for an explicit finite graph, ordinary reachability takes O(vertices + edges), before evidence and scope checks; the relevant cache inputs must be pinned
- Conflict detection: depends on the candidate comparisons, maps and evidence checks; the context count alone supplies no general bound
The cost model (A21) applies. Large closures are expensive. The coherence budget constrains how much closure you can afford.
Widening
Equivalence scope can be extended via explicit widening. If has scope U, widening to scope V (where ) requires a widening witness .
widen(e, U → V, π_widen) → Equivalence(e') with scope V
Widening is never inferred from closure but is an explicit governance act. The widening witness records who authorized the extension, under what authority, and what evidence supports it.
This prevents scope creep. An equivalence that was valid for shipping purposes does not automatically become valid for tax purposes. A scope extension requires adequate evidence and the applicable authorization. A proved preservation law may supply part of that evidence; a new declaration cannot substitute for it.
Consequence
Identity maintenance is specified. Equivalences are declared with witnesses, propagated with scope-conditional closure, and constrained by conflict handling. The substrate now has rules for populating and querying Equivalence nodes.
A retained match lets the next recipient examine which relation was established. A widened use needs additional grounds where the earlier check does not reach, while a disputed link can remain visible without supplying the requested join.
Chapter 22 asks what a predicate must carry to be accepted into this substrate. Chapter 23 asks what a query promises when it runs against equivalence-aware data. The identity discipline is established. The predicate discipline follows.