The Coherence Topos and Vocabulary Evolution

Appendix L

51 min read

This appendix addresses a specific problem: how autonomous computational agents might invent new concepts, certify them against existing commitments, transport them across institutional boundaries, and account for the cost — within a mathematically rigorous structure.

Existing approaches address fragments of this problem. Retrieval-augmented generation retrieves without coherence guarantees. Multi-agent frameworks compose outputs without gluing conditions. Knowledge graphs structure without vocabulary invention. Schema systems enforce without evolution. The composition of these fragments under formal guarantees remains open.

Parts I–VI of The Proofs assembled the components: commitment sets (A1), witnessed equivalence (A10), context sites (A12b), the sheaf condition (A13), fibrations (A14), transport discipline (A16), predicate invention (A17), conservative extension (A17b), and the coherence cost model (A21). This appendix states the structural consequence that these components jointly entail, develops several results that are (to our knowledge) novel, positions the work explicitly against the existing landscape, and identifies concrete research programs for the mathematical community.

The topos theorem (L.2) is a consequence, not a contribution — it follows from Giraud's theorem applied to the context site. We state it because its corollaries are the contribution: the internal logic subsumes the logic-selection machinery of A15, the subobject classifier provides the multi-valued truth that A4 reached for, and the monadic structure of predicate invention (L.6) gives a formal theory of vocabulary evolution that has not, to our knowledge, been developed elsewhere.

Remark

Status and scope. Three boundaries govern how the results below should be read.

This is a specification, not the deployed kernel. The object here is the mathematics of vocabulary evolution in context sites, not Bulla's operational seam-fee diagnostic. That diagnostic uses a cheaper invariant — a rank difference, computed as a component count — where cheaper suffices, and the program has retired the practice of describing it in cohomological language. That retirement concerns the tool-composition seam complex, whose real corpora are triangle-generated, whose relevant obstruction is a component count rather than a higher class, and whose public framing had drifted ahead of what was measured. It does not touch the cohomology used in this appendix, which lives on a different object — context sites for identity resolution and vocabulary evolution, where triple and higher overlaps arise by construction in the federated case (L.4.2) and higher cohomology is genuinely engaged rather than capped at triangles. The two are the same machinery applied to different structures, and their empirical standing differs accordingly.

The theorems are proved; the operational readings are hypotheses. The cohomology computations, the acyclicity theorem, and the monad characterization are proved within standard sheaf theory and model theory. The operational corollaries drawn from them — that hierarchical organizations integrate cheaply, that four-way federations face a coordination obstruction that three-way federations do not — are conjectures supported by the theorems about idealized sites and by synthetic witnesses, not yet by measurement on real federated-identity data. Their falsification route is stated where they appear.

Everything here is the exact-agreement fragment. The presheaves are set-valued and the sheaf condition demands equality on overlaps. The probabilistic and attested witnesses the program also admits, and the tolerance-based agreement its cost model assumes, require an enriched theory that this appendix does not supply. That gap is Problem 8, and it is the most consequential limitation of the framework, not a footnote to it.

L.1 What This Appendix Claims

We distinguish three levels of novelty:

Standard results applied to a new domain (L.2, L.3): The coherence topos theorem and its internal-logic corollary are instances of known mathematics (Giraud, Mac Lane–Moerdijk). We claim only that the instantiation is well-formed and that the corollaries are operationally significant for distributed systems.

Novel results (L.4–L.6): The Obstruction Cohomology computation, the acyclicity theorem for hierarchical sites, the H2H^2 meta-obstruction for federated sites, the time-indexed instability of overlap agreement, and the characterization of the predicate invention monad as a quotient of a free monad are (to our knowledge) new. They are provable within standard sheaf theory and model theory but have not appeared in the literature because the combination — sheaf-theoretic coherence applied to vocabulary evolution under conservative extension constraints — has not been studied.

Landscape comparison (L.7): We position this work explicitly against Spivak's functorial data migration, Goguen's sheaf semantics, Abramsky's sheaf-theoretic contextuality, and Caramello's bridge program. The comparison identifies what is shared, what is new, and where the framework extends existing work.

Open problems (L.8): Eight precisely stated problems for the mathematical community, including three new problems motivated by the results of this appendix (the Eilenberg-Moore category of I\mathcal{I}, persistent cohomology of evolving sites, and enriched/graded coherence).

Notation

Symbols from Parts I–VI (commitment sets, anchors, etc.) follow the conventions in Appendix H. The following notation is specific to this appendix or used here with specialized meaning. Standard category-theoretic and sheaf-theoretic notation follows Mac Lane & Moerdijk(Mac Lane 1992)Saunders Mac Lane, Sheaves in Geometry and Logic: A First Introduction to Topos Theory (New York: Springer-Verlag, 1992).View in bibliography.

SymbolMeaningIntroduced
(Ctx,J)(\mathbf{Ctx}, J)Context site: category Ctx\mathbf{Ctx} with Grothendieck topology JJA12b
PSh(Ctx)\mathbf{PSh}(\mathbf{Ctx})Presheaf category [Ctxop,Set][\mathbf{Ctx}^{\mathrm{op}}, \mathbf{Set}]L.2
Sh(Ctx,J)\mathbf{Sh}(\mathbf{Ctx}, J)Sheaf category (the coherence topos)L.2
Ω\OmegaSubobject classifier; Ω(U)\Omega(U) = JJ-closed sieves on UUL.2
aaSheafification: left exact left adjoint to inclusion i:ShPShi : \mathbf{Sh} \hookrightarrow \mathbf{PSh}L.2
WP\mathcal{W}^{\mathcal{P}}Exponential sheaf: certification space from proposals to witnessesL.3
Cˇn(U,F)\check{C}^n(\mathcal{U}, F)Čech nn-cochains of presheaf FF with respect to cover U\mathcal{U}L.4
Hˇn(U,F)\check{H}^n(\mathcal{U}, F)Čech nn-th cohomology groupL.4
δn\delta^nČech coboundary map CˇnCˇn+1\check{C}^n \to \check{C}^{n+1}L.4
Ci×UCjC_i \times_U C_jOverlap (fiber product) of CiC_i and CjC_j over UUL.4.1
E2p,qE_2^{p,q}Second page of the Čech-to-derived-functor spectral sequenceL.4.1
Hq(F)\underline{H}^q(F)Presheaf of local cohomology groupsL.4.1
Sig\mathbf{Sig}Category of signatures with inclusion morphismsL.5
Σ,Σ\Sigma, \Sigma'Signatures (finite sets of typed predicate/function symbols)A17
PPProposal endofunctor: P(Σ)P(\Sigma) = single-predicate extension proposalsL.6.1
PP^*Free monad on PP: finite sequences of proposalsL.6.1
I\mathcal{I}Predicate invention monad: admissible extensions of Σ\SigmaL.6.2
π:PI\pi : P^* \twoheadrightarrow \mathcal{I}Quotient monad morphism (surjective)L.6.2
η,μ\eta, \muMonad unit and multiplicationL.6.2
SigI\mathbf{Sig}_{\mathcal{I}}Kleisli category: vocabulary evolution pathsL.6.2
kerπ\ker \piKernel of the quotient: inadmissible proposal combinationsL.6.3

L.2 The Coherence Topos

Throughout this appendix, (Ctx,J)(\mathbf{Ctx}, J) is the context site from A12b and Ctx\mathbf{Ctx} is assumed essentially small.

Theorem(Coherence Topos)

Sh(Ctx,J)\mathbf{Sh}(\mathbf{Ctx}, J) is a Grothendieck topos(Verdier 1972--1973)Michael Artin and Alexander Grothendieck and Jean-Louis Verdier, Théorie des Topos et Cohomologie Étale des Schémas (SGA 4) (Berlin: Springer-Verlag, 1972--1973).View in bibliography(Mac Lane 1992, ch. III, §4)Saunders Mac Lane, Sheaves in Geometry and Logic: A First Introduction to Topos Theory (New York: Springer-Verlag, 1992), ch. III, §4.View in bibliography. It has all finite limits, all small colimits, exponentials, a subobject classifier Ω\Omega, and the inclusion i:Sh(Ctx,J)PSh(Ctx)i : \mathbf{Sh}(\mathbf{Ctx}, J) \hookrightarrow \mathbf{PSh}(\mathbf{Ctx}) has a left exact left adjoint aa (sheafification).

Proof

By Giraud's theorem(Verdier 1972--1973)Michael Artin and Alexander Grothendieck and Jean-Louis Verdier, Théorie des Topos et Cohomologie Étale des Schémas (SGA 4) (Berlin: Springer-Verlag, 1972--1973).View in bibliography(Mac Lane 1992, ch. III, Theorem 1)Saunders Mac Lane, Sheaves in Geometry and Logic: A First Introduction to Topos Theory (New York: Springer-Verlag, 1992), ch. III, Theorem 1.View in bibliography. The topology JJ determines a Lawvere-Tierney operator j:ΩPShΩPShj : \Omega_{\mathrm{PSh}} \to \Omega_{\mathrm{PSh}} via j(S)={f:VUf(S)J(V)}j(S) = \{f : V \to U \mid f^*(S) \in J(V)\}, which is idempotent, preserves top, and preserves meets. The jj-sheaves are the JJ-sheaves, and the category of jj-sheaves in a topos is a topos (Mac Lane & Moerdijk, Ch. V, Theorem 1).

Corollary: The Subobject Classifier and Epistemic Status

Truth Values in the Coherence Topos

The subobject classifier Ω(U)={SS is a J-closed sieve on U}\Omega(U) = \{S \mid S \text{ is a } J\text{-closed sieve on } U\}. Truth values are not {0,1}\{0, 1\} but JJ-closed sieves: families of contexts in which a claim holds, closed under the covering relation.

A4 Epistemic StatusTopos Interpretation
True in UUThe maximal sieve  ⁣U\uparrow\! U (all refinements)
False in UUThe empty sieve \emptyset
Undetermined in UUA proper non-empty JJ-closed sieve
Conflict at UUBoth χφ\chi_\varphi and χ¬φ\chi_{\neg\varphi} are non-empty, proper sieves

The internal logic is intuitionistic. In the posetal context sites used throughout this appendix, excluded middle holds at UU iff the restricted topology admits only the trivial local truth values — every JJ-closed sieve below UU is maximal or empty — which is the closed-world assumption. (In full generality Booleanness is a property of the sheaf topos rather than a syntactic flag on the site; the posetal case is where the two coincide.) Non-discrete topologies yield open-world reasoning natively — no adapter required.

Theorem(Logic Selection as Topos Relativization)

The indexed logic selection of A15 is a special case of relativizing to sub-topologies. Specifically: Uφ¬φU \Vdash \varphi \lor \neg\varphi for all φ\varphi iff JJ restricted to the sieve below UU is the discrete topology. The CWA/OWA distinction is not an engineering parameter but a structural property of the topology over each context.

Proof

A Grothendieck topos is Boolean iff ¬¬=idΩ\neg\neg = \mathrm{id}_\Omega, which holds iff every JJ-closed sieve is maximal or empty — the discrete topology(Mac Lane 1992, ch. VI, §6)Saunders Mac Lane, Sheaves in Geometry and Logic: A First Introduction to Topos Theory (New York: Springer-Verlag, 1992), ch. VI, §6.View in bibliography. Restricting to the slice Sh/U\mathbf{Sh}/U yields a sub-topos whose Booleanness depends on the induced topology on the under-category U/CtxU/\mathbf{Ctx}.

L.3 Exponentials and Certification

The topos has exponentials. For sheaves P\mathcal{P} (proposals, per A19) and W\mathcal{W} (witnesses, per A2c):

WP(U)HomSh/U(PU,WU)\mathcal{W}^{\mathcal{P}}(U) \cong \mathrm{Hom}_{\mathbf{Sh}/U}(\mathcal{P}|_U, \mathcal{W}|_U)

A certification contract (A19b) is a global section cΓ(WP)c \in \Gamma(\mathcal{W}^{\mathcal{P}}): a natural transformation PW\mathcal{P} \Rightarrow \mathcal{W} that commutes with all restriction maps. The topos guarantees the space of certifications is a well-defined sheaf. Coherence of certification across contexts is naturality. Whether a particular certification exists is the engineering problem; the topos provides the space in which to search.

L.4 Obstruction Cohomology: A Worked Computation

This section contains what we believe to be novel: an explicit computation of the first sheaf cohomology group H1H^1 for a concrete context site arising in data integration, and its interpretation as classifying ambiguous identity resolution. The companion paper Predicate Invention Under Sheaf Constraints (SCPI) proves that the same H1H^1 classifies obstructions to predicate invention across heterogeneous agent contexts, formalizing the descent problem that A17's three obligations address. The SHEAF Protocol extends this diagnostic to a distributed setting with mechanism-design enforcement.

The Setup: Three-Merchant Catalog

Let Ctx\mathbf{Ctx} be the poset category with objects {U,A,B,C,AB,AC,BC}\{U, A, B, C, A \wedge B, A \wedge C, B \wedge C\} where A,B,CA, B, C are merchant contexts covering the catalog context UU, and the \wedge-objects are pairwise overlaps. Morphisms are inclusions (each overlap refines both parents).

The topology JJ declares {AU,BU,CU}\{A \to U, B \to U, C \to U\} as a cover.

Let FF be the presheaf of product identifiers:

  • F(A)={a1,a2,a3}F(A) = \{a_1, a_2, a_3\} (merchant A's products)
  • F(B)={b1,b2,b3}F(B) = \{b_1, b_2, b_3\} (merchant B's products)
  • F(C)={c1,c2}F(C) = \{c_1, c_2\} (merchant C's products)

On overlaps, restriction identifies shared products:

  • F(AB)F(A \wedge B): product a2a_2 and b1b_1 are "the same item" — but the identification is ambiguous (two possible matchings exist)
  • F(AC)F(A \wedge C): product a3a_3 and c1c_1 are unambiguously identified
  • BCB \wedge C is initial (merchants BB and CC share no sub-context), so it contributes no comparison

Two facts about this site are worth separating, because they are the two independent sources of obstruction that L.4.3 will classify. First, the nerve of the cover {A,B,C}\{A, B, C\} — the simplicial complex recording which merchants overlap — is the path BACB - A - C, a tree, hence contractible: it carries no topology of its own. Any obstruction that appears here therefore comes not from the shape of the cover but from the coefficient system FF, specifically from the ambiguous matching on ABA \wedge B. This is a descent obstruction. It is the opposite mechanism from the one in L.4.2, where the coefficients are constant and the obstruction lives entirely in the nerve.

The Čech Complex

The Čech cohomology of FF with respect to the cover U={A,B,C}\mathcal{U} = \{A, B, C\} is computed from the cochain complex:

Cˇ0(U,F)δ0Cˇ1(U,F)δ1Cˇ2(U,F)\check{C}^0(\mathcal{U}, F) \xrightarrow{\delta^0} \check{C}^1(\mathcal{U}, F) \xrightarrow{\delta^1} \check{C}^2(\mathcal{U}, F)

where:

  • Cˇ0=F(A)×F(B)×F(C)\check{C}^0 = F(A) \times F(B) \times F(C) — local sections (one per merchant)
  • Cˇ1=F(AB)×F(AC)\check{C}^1 = F(A \wedge B) \times F(A \wedge C) — comparison on the two non-trivial overlaps
  • higher terms vanish: BCB \wedge C and the triple overlap are initial

Because FF is set-valued (its values are sets of product identifiers, not abelian groups), the differential is not a subtraction. The correct object is the non-abelian Čech complex, in which δ0\delta^0 records, for each overlap, the pair of restrictions to be identified, and a global section is a choice of local products together with a witnessed identification on every overlap that is compatible where overlaps meet. Writing subtraction here would presuppose a group structure the identifiers do not carry; the honest structure is descent, and we treat it as such.

Global sections and the descent obstruction

H0(U,F)H^0(\mathcal{U}, F) is the set of global sections: assignments of local products that admit a consistent identification across all overlaps — the coherent global catalog. When every overlap identification is forced, H0H^0 is a single glued catalog.

The obstruction is the failure of that section to be unique. For set-valued FF the classifying object is not an abelian cohomology group but the non-abelian first Čech cohomology Hˇ1(U,F)\check{H}^1(\mathcal{U}, F) — the pointed set of 1-cocycles modulo coboundary, equivalently π0\pi_0 of the groupoid of global matchings. Its elements are the distinct global catalogs assemblable from the same local data. The abelian H1H^1 appears only after one replaces FF by its free abelianization ZF\mathbb{Z}F; the resulting Hˇ1(U,ZF)\check{H}^1(\mathcal{U}, \mathbb{Z}F) is the linear shadow of the descent groupoid, a computational convenience, not the primary object.

Theorem(Descent Ambiguity for Identity Resolution)

For the three-merchant site above, whose nerve is contractible, the classifying set Hˇ1(U,F)\check{H}^1(\mathcal{U}, F) has more than one element whenever the overlap ABA \wedge B admits multiple consistent identifications of shared products. Concretely: if a2a_2 could match either b1b_1 or b2b_2 (both consistent with the restriction maps), then Hˇ1\check{H}^1 has at least two elements, and they correspond bijectively to the distinct global catalogs assemblable from the same local data. The obstruction is coefficient-driven: it survives even though the nerve carries no topology.

Proof

A 1-cocycle assigns to each overlap an identification satisfying the compatibility condition where overlaps meet (here vacuous: ABA \wedge B and ACA \wedge C meet only at AA, and no triple overlap exists). Two cocycles are cohomologous when a relabeling of the local products F(A)F(A) or F(B)F(B) carries one to the other.

The two matchings m1:a2b1m_1 : a_2 \leftrightarrow b_1 and m2:a2b2m_2 : a_2 \leftrightarrow b_2 define distinct cocycles. They are cohomologous iff some relabeling of F(A)F(A) or F(B)F(B) transforms one into the other. If b1b2b_1 \neq b_2 and neither lies in the image of any other identification, no such relabeling exists, and the two cocycles represent distinct classes in Hˇ1\check{H}^1. Each class is a distinct global catalog: the same local data assembled into different global pictures according to which identification is chosen.

Remark

This is the formal version of a problem every data-integration practitioner knows: two sources share some entities, the matching is ambiguous, and different matchings produce different downstream results. A non-trivial Hˇ1(U,F)\check{H}^1(\mathcal{U}, F) is the mathematical name for that ambiguity, and the size of the set counts the distinct resolutions. It is not metaphor — it is computable for finite context sites — but it is a set of descent classes, not an abelian group, and the distinction matters: the ambiguity here comes from the coefficients, not from any hole in the cover.

For the agentic substrate specifically: when two agents in different contexts propose identity claims about shared entities, Hˇ1\check{H}^1 counts the irreducible ways of reconciling those claims. No amount of embedding similarity collapses the set; only an explicit choice of representative — a witnessed identification — does.

L.4.1 Acyclicity of Hierarchical Sites

The three-merchant example has non-trivial H1H^1 because the overlap structure admits ambiguity. A natural question: for which site structures does ambiguity vanish? The answer connects organizational topology to coherence cost.

Hierarchical Context Site

A context site (Ctx,J)(\mathbf{Ctx}, J) is hierarchical if:

  1. Ctx\mathbf{Ctx} is a finite rooted tree (poset where every element except the root has exactly one immediate predecessor)
  2. The topology JJ is generated by parent-children families: for each non-leaf node UU with children {C1,,Ck}\{C_1, \ldots, C_k\}, the family {CiU}i=1k\{C_i \to U\}_{i=1}^k is a cover
  3. For distinct siblings Ci,CjC_i, C_j (children of the same parent), the overlap Ci×UCjC_i \times_U C_j is the initial object \emptyset (no shared sub-context between different branches)
Theorem(Acyclicity of Hierarchical Sites)

Let (Ctx,J)(\mathbf{Ctx}, J) be a hierarchical context site. For any abelian presheaf FF on Ctx\mathbf{Ctx} and any cover U\mathcal{U} in JJ:

Hˇn(U,F)=0for all n1\check{H}^n(\mathcal{U}, F) = 0 \quad \text{for all } n \geq 1

In particular, H1=0H^1 = 0: there is no ambiguity in identity resolution for hierarchical organizations.

Proof

We prove this by analyzing the Čech complex directly.

Step 1: Structure of overlaps in a tree.

Let UU be a node with children {C1,,Ck}\{C_1, \ldots, C_k\} forming a cover. For iji \neq j, the overlap Ci×UCj=C_i \times_U C_j = \emptyset by the tree condition (distinct branches share no sub-context). Therefore for any presheaf FF:

F(Ci×UCj)=F()={}(terminal, for an abelian presheaf: the zero object)F(C_i \times_U C_j) = F(\emptyset) = \{*\} \quad \text{(terminal, for an abelian presheaf: the zero object)}

Step 2: Collapse of the Čech complex.

The Čech complex for cover U={C1,,Ck}\mathcal{U} = \{C_1, \ldots, C_k\} of UU is:

Cˇ0=iF(Ci)δ0Cˇ1=i<jF(Ci×UCj)δ1Cˇ2=i<j<lF(Ci×UCj×UCl)\check{C}^0 = \prod_{i} F(C_i) \xrightarrow{\delta^0} \check{C}^1 = \prod_{i < j} F(C_i \times_U C_j) \xrightarrow{\delta^1} \check{C}^2 = \prod_{i < j < l} F(C_i \times_U C_j \times_U C_l) \to \cdots

Since Ci×UCj=C_i \times_U C_j = \emptyset for all iji \neq j, every term Cˇn=0\check{C}^n = 0 for n1n \geq 1. The complex is:

iF(Ci)00\prod_i F(C_i) \to 0 \to 0 \to \cdots

Therefore Hˇn=0\check{H}^n = 0 for all n1n \geq 1.

Step 3: Extension to composite covers.

For a cover of a non-root node, the same argument applies locally: each non-leaf is covered by its children, which are pairwise disjoint. By the Čech-to-derived-functor spectral sequence (or directly by Leray's theorem applied to the refinement of any cover by the canonical parent-children covers), the vanishing extends to all covers in JJ, not just the generating ones.

Step 4: Recursive argument for depth >1> 1.

For a tree of depth dd, consider the cover of the root by its children, then each child by its children, etc. The Čech-to-sheaf cohomology spectral sequence for this iterated cover has:

E2p,q=Hˇp(U,Hq(F))E_2^{p,q} = \check{H}^p(\mathcal{U}, \underline{H}^q(F))

where Hq\underline{H}^q is the presheaf of local cohomology. By induction on depth: Hq=0\underline{H}^q = 0 for q1q \geq 1 (each sub-tree is acyclic by the inductive hypothesis), so E2p,q=0E_2^{p,q} = 0 for q1q \geq 1. And E2p,0=Hˇp(U,F)=0E_2^{p,0} = \check{H}^p(\mathcal{U}, F) = 0 for p1p \geq 1 by Step 2. Therefore the spectral sequence degenerates and Hn(Ctx,F)=0H^n(\mathbf{Ctx}, F) = 0 for all n1n \geq 1.

Remark

This theorem has a precise operational meaning, and precise limits. In the idealized hierarchical site, where sibling branches are stipulated to share no sub-context, there is no room for multiple consistent identifications, because there is nothing on which distinct branches can disagree. The hierarchy resolves identity by construction.

The operational reading follows only to the extent that a real structure meets the idealization. It suggests why hierarchies are comparatively easy to integrate — a corporate merger of divisions with genuinely disjoint operations, a taxonomy with strict inclusion, a file tree with no links — and, read the other way, it locates exactly where real hierarchies stop being easy. Ambiguity re-enters a real tree through every violation of the sibling-disjointness assumption: aliases and symlinks, shared services, cross-cutting policies, duplicated entities, informal equivalences maintained outside the tree. The theorem is therefore a diagnostic as much as a guarantee: where a nominal hierarchy exhibits identity ambiguity, the cohomology says the tree is not the real site, and the shared sub-contexts hiding the obstruction can be named.

The price of acyclicity is rigidity. A tree cannot express "A and B share some context but neither subsumes the other." Peer-to-peer and federated structures can, and they pay for it with non-trivial cohomology.

L.4.2 Higher Obstructions in Federated Sites

Federated structures are the opposite extreme from hierarchies: multiple overlapping authorities, no single root, non-trivial shared contexts. We show that federated sites can have non-trivial H2H^2, which classifies meta-conflicts — disagreements not about identity itself but about how to resolve identity disagreements.

Federated Context Site

A context site (Ctx,J)(\mathbf{Ctx}, J) is federated if:

  1. Ctx\mathbf{Ctx} contains a set of federation nodes {F1,,Fm}\{F_1, \ldots, F_m\} and member nodes {M1,,Mn}\{M_1, \ldots, M_n\}
  2. Each member belongs to at least one federation: for each MjM_j, there exists FiF_i with a morphism MjFiM_j \to F_i (membership)
  3. The topology JJ includes the cover {MjFiMjFi}\{M_j \to F_i \mid M_j \in F_i\} for each federation FiF_i
  4. Members of distinct federations may share non-trivial overlaps: Mj×FiMkM_j \times_{F_i} M_k need not be initial
  5. There exists a global context GG covered by {F1,,Fm}\{F_1, \ldots, F_m\}
Theorem(Non-Trivial H^2 in Federated Sites)

There exists a federated context site (Ctx,J)(\mathbf{Ctx}, J) and presheaf FF such that H2(U,F)0H^2(\mathcal{U}, F) \neq 0 for a cover U\mathcal{U} of the global context. Elements of H2H^2 classify meta-obstructions: situations where pairwise identity resolutions exist but no globally consistent resolution strategy exists.

Proof

Construction. Let GG be covered by four federation nodes F1,F2,F3,F4F_1, F_2, F_3, F_4, with a shared member MijM_{ij} for every pair i<ji < j (so Fi×GFj=MijF_i \times_G F_j = M_{ij}), a shared member MijkM_{ijk} for every triple i<j<ki < j < k, and no member common to all four (F1×GF2×GF3×GF4=F_1 \times_G F_2 \times_G F_3 \times_G F_4 = \emptyset). The nerve of this cover is the boundary of a tetrahedron — four vertices, six edges, four triangles, no filled interior — the simplicial sphere Δ3S2\partial\Delta^3 \cong S^2.

Let FF be the constant presheaf of identification conventions, Z/2Z\mathbb{Z}/2\mathbb{Z}-valued (two conventions per node, "match by name" vs "match by code"), with identity restriction maps. The Čech complex is

(Z/2)4δ0(Z/2)6δ1(Z/2)4δ20,(\mathbb{Z}/2)^4 \xrightarrow{\delta^0} (\mathbb{Z}/2)^6 \xrightarrow{\delta^1} (\mathbb{Z}/2)^4 \xrightarrow{\delta^2} 0,

with Cˇ3=F()=0\check{C}^3 = F(\emptyset) = 0 because no member is common to all four federations.

Computation. For a constant abelian coefficient system, the Čech cohomology of the cover is the simplicial cohomology of its nerve — here H(S2;Z/2)H^\bullet(S^2; \mathbb{Z}/2). The dimension count confirms the complex: the alternating sum 46+4=2=χ(S2)4 - 6 + 4 = 2 = \chi(S^2). Since H0(S2;Z/2)=Z/2H^0(S^2; \mathbb{Z}/2) = \mathbb{Z}/2 and H1(S2;Z/2)=0H^1(S^2; \mathbb{Z}/2) = 0, the ranks are forced: dimkerδ0=1\dim\ker\delta^0 = 1, so dimimδ0=3\dim\operatorname{im}\delta^0 = 3; H1=0H^1 = 0 forces dimkerδ1=dimimδ0=3\dim\ker\delta^1 = \dim\operatorname{im}\delta^0 = 3, so dimimδ1=63=3\dim\operatorname{im}\delta^1 = 6 - 3 = 3; hence

Hˇ2=Cˇ2/imδ1=(Z/2)4/(Z/2)3=Z/20.\check{H}^2 = \check{C}^2 / \operatorname{im}\delta^1 = (\mathbb{Z}/2)^4 \big/ (\mathbb{Z}/2)^3 = \mathbb{Z}/2 \neq 0.

A non-trivial 2-cocycle assigns a convention to each triple overlap so that the compatibility condition holds, yet cannot be written as a coboundary of pairwise data. This is the meta-obstruction: every pair of federations can resolve its identity disagreements, and every triple can find a consistent resolution, but no single global resolution strategy is compatible with all four at once.

Remark

Why three federations are not enough. The natural first attempt uses three federations, and it fails in a way that proves the point. With F1,F2,F3F_1, F_2, F_3, pairwise members MijM_{ij}, and one common member M123M_{123}, the nerve is the filled triangle Δ2\Delta^2, which is contractible; the complex (Z/2)3(Z/2)3Z/2(\mathbb{Z}/2)^3 \to (\mathbb{Z}/2)^3 \to \mathbb{Z}/2 computes to Hˇ1=Hˇ2=0\check{H}^1 = \check{H}^2 = 0: no obstruction at any level. This is the theorem's content. A higher obstruction is manufactured not by adding parties but by leaving a hole in the nerve: three federations sharing a common member fill their triangle; four federations sharing no common member leave the tetrahedron hollow, and the hollow is the S2S^2 whose H2H^2 carries the meta-conflict. The obstruction is in the topology of the cover, not in the number of participants.

Remark

The H2H^2 meta-obstruction has a vivid operational interpretation. Consider four regulatory bodies (F1,,F4F_1, \ldots, F_4) each overseeing a set of financial institutions. Any two regulators can agree on how to identify shared entities. Any three can find a consistent protocol. But when all four try to federate, a global obstruction emerges: the pairwise agreements, though locally consistent in triples, cannot be simultaneously satisfied. This is a higher-order coordination failure — not a conflict about data but a conflict about conflict-resolution strategies.

For the agentic substrate: H20H^2 \neq 0 means that even if every pair of AI agents can resolve their identity disputes, and every triple can coordinate, the system as a whole may still lack a globally consistent identity protocol. The obstruction is structural, residing in the topology of the federation, not in any particular data disagreement.

The nerve of the cover is the key invariant: when it has non-trivial higher homotopy, higher cohomology obstructions emerge. This connects the formal theory to classical algebraic topology in a precise and computable way.

L.4.3 The Cohomological Hierarchy: A Classification

The results of L.4, L.4.1, and L.4.2 fit into a single classification — provided one keeps separate the two independent sources of obstruction they exhibit. An obstruction can come from the coefficients (an ambiguous identification on an overlap, which produces a non-trivial descent class even when the nerve is contractible — this is L.4) or from the nerve (a hole in the cover, which produces a non-trivial class even when the coefficients are constant — this is L.4.2). The earlier literature on this material tended to collapse the two; they are genuinely different, and the flagship example of each lives on a nerve where the other source is switched off.

Site structureNerveCoefficientsObstruction sourceClassifying objectOperational meaning
Hierarchical (tree)ContractibleanynonetrivialHierarchy resolves all identity
Ambiguous overlap (L.4)Contractiblenon-constant (ambiguous match)descent / coefficientsHˇ1(U,F)\check{H}^1(\mathcal{U}, F), a pointed set with >1> 1 elementFinitely many distinct global catalogs
Flat peer-to-peerS1\bigvee S^1constanttopological / nerveH1(nerve)H^1(\text{nerve})Identity ambiguity from holes in the cover
Federated, four-way (L.4.2)S2S^2 or higherconstanttopological / nerveH2(nerve)0H^2(\text{nerve}) \neq 0Meta-obstruction: no global strategy
Fully connectedContractible (Δn1\Delta^{n-1})anynonetrivialTotal overlap; everyone sees everything
Remark

The two acyclic rows are trivial for opposite reasons: in a tree, siblings share nothing; in a complete graph, everyone shares everything. Between them lie the realistic cases — partial overlap, partial authority, partial sharing — which are exactly the structures that arise in multi-agent systems, federated databases, and inter-organizational data sharing.

Reading the table by column rather than by row is the point. The descent obstruction and the topological obstruction can each be present or absent independently: a contractible nerve with ambiguous coefficients carries the first and not the second; a hollow nerve with constant coefficients carries the second and not the first; a real federation typically carries both, and the total coherence cost is not one invariant but a pair. An architect choosing an organizational topology is choosing a point in a two-axis space, and a diagnostic that reports a single number has already lost the distinction that tells it which repair to attempt — disambiguate a matching, or close a hole in the cover.

L.5 Vocabulary Evolution: Composability and Its Limits

Signature Category

Sig\mathbf{Sig} is the category of signatures (finite sets of typed predicate/function symbols) with morphisms the signature inclusions ΣΣ\Sigma \hookrightarrow \Sigma'.

Theorem(Composability of Conservative Extensions)

If ΣΣ\Sigma \hookrightarrow \Sigma' and ΣΣ\Sigma' \hookrightarrow \Sigma'' are both conservative extensions (A17b), then ΣΣ\Sigma \hookrightarrow \Sigma'' is conservative.

Proof

Let φ\varphi be a Σ\Sigma-sentence with (Σ,I,L)φ(\Sigma'', I'', L) \vdash \varphi. Since φ\varphi is also a Σ\Sigma'-sentence, conservativity of ΣΣ\Sigma' \hookrightarrow \Sigma'' yields (Σ,I,L)φ(\Sigma', I', L) \vdash \varphi. Conservativity of ΣΣ\Sigma \hookrightarrow \Sigma' then yields (Σ,I,L)φ(\Sigma, I, L) \vdash \varphi. The converse is monotonicity.

This composability is what makes incremental vocabulary evolution safe. A chain of conservative extensions is conservative. You verify each step; the chain is automatic.

But overlap agreement, checked extensionally, is not stable under the growth of the population it is checked against. This is the central tension in the theory of vocabulary evolution, and it is worth stating precisely, because the naive version — that agreement fails to compose at a fixed moment — is false. At a fixed population, agreement composes: if two predicates each agree pointwise on the overlap, so does every Boolean combination of them, and no composite can fail where its parts passed. The instability is temporal, not combinatorial, and that is the sharper and more consequential fact.

Theorem(Instability of Extensional Overlap Agreement)

There is a predicate with context-dependent definitions that satisfies Obligation 2 against the overlap population present at certification time t0t_0, yet violates it at a later time t1>t0t_1 > t_0 once the overlap acquires a single item on which its two definitions disagree — with no change to any definition. A certificate of overlap agreement is therefore a statement about the population sampled at issue time, and carries no guarantee for any later population.

Proof

Let Ctx\mathbf{Ctx} have objects UU, VV, UVU \wedge V, and let Σ\Sigma contain a sort DD (dresses) with base predicates material_quality:D[0,1]\text{material\_quality} : D \to [0,1] and certified_sustainable:D{,}\text{certified\_sustainable} : D \to \{\top, \bot\}.

Define q:D{,}q : D \to \{\top, \bot\} by two context-local definitions:

  • in UU: q(d)=[material_quality(d)>0.7]q(d) = [\,\text{material\_quality}(d) > 0.7\,]
  • in VV: q(d)=[certified_sustainable(d)]q(d) = [\,\text{certified\_sustainable}(d)\,]

Obligation 2 requires the two definitions to agree on the overlap UVU \wedge V. At time t0t_0, let every item present in the overlap satisfy material_quality(d)>0.7certified_sustainable(d)\text{material\_quality}(d) > 0.7 \Leftrightarrow \text{certified\_sustainable}(d) — each is either high-quality-and-certified or neither. On this population the two definitions coincide pointwise, so qq passes Obligation 2 and ΣΣ{q}\Sigma \hookrightarrow \Sigma \cup \{q\} is admissible.

At time t1t_1 a new item dd^\ast enters the overlap with material_quality(d)=0.6\text{material\_quality}(d^\ast) = 0.6 and certified_sustainable(d)=\text{certified\_sustainable}(d^\ast) = \top. The UU-definition now returns q(d)=q(d^\ast) = \bot and the VV-definition returns q(d)=q(d^\ast) = \top: the definitions disagree on UVU \wedge V, Obligation 2 fails, and the previously certified extension is no longer admissible. No definition changed — only the population did.

Crucially, dd^\ast could not have belonged to the t0t_0 population: it fails the very condition every certified item satisfies, since material_quality(d)0.7\text{material\_quality}(d^\ast) \le 0.7 while certified_sustainable(d)=\text{certified\_sustainable}(d^\ast) = \top gives [material_quality(d)>0.7]certified_sustainable(d)[\,\text{material\_quality}(d^\ast) > 0.7\,] \neq \text{certified\_sustainable}(d^\ast). The disagreeing item is therefore genuinely new at t1t_1, not one present-but-overlooked at t0t_0; the certificate at t0t_0 ranged soundly over a population that excluded it. The instability is a property of population growth, and the exact validity domain of the certificate is the agreement region A={d:[material_quality(d)>0.7]=certified_sustainable(d)}A = \{\, d : [\,\text{material\_quality}(d) > 0.7\,] = \text{certified\_sustainable}(d)\,\} — the certificate holds for populations contained in AA and is voided by, and only by, the arrival of an item outside it.

Remark

Two consequences follow, and they are why the coherence cost model (A21) charges for re-verification rather than certifying once.

First, admissibility is time-indexed. An extensional overlap check is a measurement against the sample present at issue time, and — like every measurement in the program's empirical layer — it expires. The monad multiplication of L.6 therefore re-verifies Obligation 2 against the current population at every step rather than trusting a prior certificate; this is the only sound reading of what the certificate claims, not defensive engineering.

Second, the cost this imposes is bilinear — O(noverlaps)O(n \cdot |\text{overlaps}|) per round, quadratic only when overlaps|\text{overlaps}| scales with nn — in a specific and honest way. Composition at a fixed population holds no surprises (Boolean combinations of pointwise-agreeing predicates agree pointwise), so the burden is not an explosion of composites but the re-verification itself: each of the nn predicates must be re-checked on every overlap whenever the population it governs changes, giving the per-round O(noverlaps)O(n \cdot |\text{overlaps}|) scaling recorded in A21. The limit is temporal. Agents cannot bank an overlap certificate and compose against it later, because the overlap it certified is not the overlap they will compose over.

For the agentic substrate: autonomous agents cannot invent vocabulary against a snapshot and assume it holds. Invention must re-synchronize at the overlap boundary on every material change to the shared population — the exact-agreement counterpart of the persistence question raised in Problem 7, where the lifetime of such an obstruction becomes the invariant of interest.

L.6 The Predicate Invention Monad

Despite the time-indexed instability of Obligation 2 (L.5), predicate invention has a well-defined algebraic structure when the full A17 pipeline (including re-verification) is included. We develop this structure in three stages: the free monad of unconstrained proposals, the quotient that enforces admissibility, and the resulting algebraic characterization.

L.6.1 The Proposal Endofunctor

Proposal Endofunctor

Define the proposal endofunctor P:SigSigP : \mathbf{Sig} \to \mathbf{Sig} by:

P(Σ)={(Σ{q},δq)qΣ,  δq is a grounding definition for q}P(\Sigma) = \{(\Sigma \cup \{q\}, \delta_q) \mid q \notin \Sigma,\; \delta_q \text{ is a grounding definition for } q\}

where δq\delta_q specifies the sort, arity, and local definition of qq in each context. PP sends a signature to the set of all single-predicate extension proposals (without checking admissibility). On morphisms: an inclusion ΣΣ\Sigma \hookrightarrow \Sigma' maps a Σ\Sigma-proposal (Σ{q},δq)(\Sigma \cup \{q\}, \delta_q) to the Σ\Sigma'-proposal (Σ{q},δq)(\Sigma' \cup \{q\}, \delta_q) when qΣq \notin \Sigma', and discards it otherwise (the proposed predicate already exists).

Typing. Because P(Σ)P(\Sigma) is a set of extension proposals rather than a single signature, PP is an endofunctor not on Sig\mathbf{Sig} itself but on its free coproduct completion Sig^\widehat{\mathbf{Sig}}, the category whose objects are sets of signatures under disjoint union, with Σ\Sigma identified with the singleton {Σ}\{\Sigma\} and P({Σ})P(\{\Sigma\}) its set of single-predicate extensions. The free monad P(Σ)=n0Pn(Σ)P^*(\Sigma) = \coprod_{n \geq 0} P^n(\Sigma), its unit and multiplication, and the Kleisli category SigI\mathbf{Sig}_{\mathcal{I}} are all formed in Sig^\widehat{\mathbf{Sig}}, where the coproduct exists; an element of P(Σ)P^*(\Sigma) is then the finite proposal sequence described below. The inadmissible combinations collected in kerπ\ker\pi (L.6.3) are sent to a formally adjoined bottom element \bot, so that the admissibility quotient I\mathcal{I} is a genuine quotient monad rather than a partial operation.

Free Monad on Proposals

The free monad PP^* on the endofunctor PP is defined by:

P(Σ)=n0Pn(Σ)=Σ+P(Σ)+P(P(Σ))+P^*(\Sigma) = \coprod_{n \geq 0} P^n(\Sigma) = \Sigma + P(\Sigma) + P(P(\Sigma)) + \cdots

An element of P(Σ)P^*(\Sigma) is a finite sequence of extension proposals (q1,δ1),,(qn,δn)(q_1, \delta_1), \ldots, (q_n, \delta_n) applied to Σ\Sigma. The monadic structure:

  • Unit η:IdP\eta : \mathrm{Id} \to P^* embeds Σ\Sigma as the empty sequence of proposals.
  • Multiplication μ:PP\mu : P^{**} \to P^* flattens a sequence-of-sequences into a single sequence by concatenation.

PP^* is the free monad on PP in the sense of the universal property: for any monad TT and natural transformation α:PT\alpha : P \Rightarrow T, there exists a unique monad morphism αˉ:PT\bar{\alpha} : P^* \to T extending α\alpha.

L.6.2 The Admissibility Quotient

The free monad PP^* allows any sequence of proposals. The predicate invention monad I\mathcal{I} is the quotient that enforces the three obligations of A17.

Predicate Invention Monad

Define the admissibility relation \sim on P(Σ)P^*(\Sigma): two proposal sequences are equivalent if they yield the same final signature and both pass (or both fail) the A17 admissibility check. Define:

I(Σ)={ΣΣΣΣ passes A17}\mathcal{I}(\Sigma) = \{\Sigma' \supseteq \Sigma \mid \Sigma \hookrightarrow \Sigma' \text{ passes A17}\}

ordered by inclusion. There is a surjective monad morphism π:PI\pi : P^* \twoheadrightarrow \mathcal{I} that sends each proposal sequence to its composite extension (if admissible) or discards it (if not). The monadic structure:

  • Unit ηΣ:ΣI(Σ)\eta_\Sigma : \Sigma \hookrightarrow \mathcal{I}(\Sigma) — the identity extension (always admissible).
  • Multiplication μΣ:I(I(Σ))I(Σ)\mu_\Sigma : \mathcal{I}(\mathcal{I}(\Sigma)) \to \mathcal{I}(\Sigma) — compose extensions and re-verify Obligation 2 for the composite. μ\mu is well-defined because conservative extension composes (L.5) and Obligations 1 and 3 are monotone in signature; only Obligation 2 requires re-checking.

The Kleisli category SigI\mathbf{Sig}_{\mathcal{I}} has:

  • Objects: signatures
  • Morphisms ΣΣ\Sigma \to \Sigma': admissible extensions
  • Composition: extension-then-re-verify

This is the category of vocabulary evolution paths. A morphism in SigI\mathbf{Sig}_{\mathcal{I}} is a certified route from one vocabulary to another.

Theorem(Predicate Invention as Quotient of Free Monad)

I\mathcal{I} is a quotient monad of PP^*. Specifically, there is a surjective monad morphism π:PI\pi : P^* \twoheadrightarrow \mathcal{I} whose kernel is the congruence generated by two relations:

  1. Path independence: (q1,δ1),(q2,δ2)(q2,δ2),(q1,δ1)(q_1, \delta_1), (q_2, \delta_2) \sim (q_2, \delta_2), (q_1, \delta_1) when both orderings yield the same composite extension
  2. Admissibility filtering: (q1,δ1),,(qn,δn)(q_1, \delta_1), \ldots, (q_n, \delta_n) \sim \bot when the composite Σ{q1,,qn}\Sigma \cup \{q_1, \ldots, q_n\} fails any obligation of A17

Consequently, the category of I\mathcal{I}-algebras is a reflective subcategory of PP^*-algebras, consisting of those PP^*-algebras where the Obligation 2 equations hold.

Proof

That π\pi is a monad morphism: We must show π\pi commutes with unit and multiplication. For the unit: π(ηP(Σ))=π(Σ,empty sequence)=Σ=ηI(Σ)\pi(\eta_{P^*}(\Sigma)) = \pi(\Sigma, \text{empty sequence}) = \Sigma = \eta_{\mathcal{I}}(\Sigma). For multiplication: let s=((q1,δ1),)s = ((q_1, \delta_1), \ldots) be a sequence in P(P(Σ))P^*(P^*(\Sigma)), consisting of a sequence of sequences of proposals. Then π(μP(s))\pi(\mu_{P^*}(s)) = the composite of the flattened sequence, and μI(π(π(s)))\mu_{\mathcal{I}}(\pi(\pi(s))) = the composite of the composites. Since extension composition is associative (signature union is associative), these agree when both are admissible. When either is inadmissible, both map to \bot.

Surjectivity: Every admissible extension ΣΣ\Sigma \hookrightarrow \Sigma' with Σ=Σ{q1,,qn}\Sigma' = \Sigma \cup \{q_1, \ldots, q_n\} is the image of the proposal sequence (q1,δ1),,(qn,δn)(q_1, \delta_1), \ldots, (q_n, \delta_n) under π\pi.

Kernel characterization: Two proposal sequences have the same image under π\pi iff they yield the same composite signature (path independence) or both are inadmissible (admissibility filtering). These generate a congruence on PP^* because both relations are compatible with the monad multiplication (re-verification depends only on the composite, not the path).

Reflective subcategory: An I\mathcal{I}-algebra is a signature Σ\Sigma equipped with an action α:I(Σ)Σ\alpha : \mathcal{I}(\Sigma) \to \Sigma — a way to "absorb" admissible extensions. This is a PP^*-algebra that additionally satisfies the admissibility relations: whenever a composite extension admissible against one overlap population becomes inadmissible against a later, larger one (L.5), the algebra's action must reject the stale composite rather than carry it forward. The reflector is the functor that takes a PP^*-algebra and quotients by these Obligation 2 relations.

Remark

The monad laws hold:

  • Left unit: μηI=id\mu \circ \eta_{\mathcal{I}} = \mathrm{id} (extending by nothing, then composing, is identity).
  • Right unit: μI(η)=id\mu \circ \mathcal{I}(\eta) = \mathrm{id} (composing with the identity extension is identity).
  • Associativity: μμI=μI(μ)\mu \circ \mu_{\mathcal{I}} = \mu \circ \mathcal{I}(\mu) — this holds because re-verification of Obligation 2 for the composite is independent of the order in which we compose three extensions. The overlap structure depends only on the final signature, not on the path taken to reach it.

The last point is significant: the cost of re-verification may depend on the path (some orderings may allow caching), but the result does not. The monad captures what is invariant (the admissibility condition); the cost model (A21) captures what varies (the verification effort).

L.6.3 The Algebraic Content of the Quotient

The quotient structure π:PI\pi : P^* \twoheadrightarrow \mathcal{I} makes precise what kind of algebraic object vocabulary evolution is — and, as importantly, what it is not.

Theorem(Non-Freeness of the Invention Monad)

I\mathcal{I} is not the free monad on the proposal endofunctor PP: the quotient π:PI\pi : P^* \twoheadrightarrow \mathcal{I} has a non-trivial kernel.

Proof

The free monad PP^* has as its elements the sequences of proposals; two sequences are equal in P(Σ)P^*(\Sigma) only if they are literally identical. But π\pi identifies any two sequences that reach the same admissible composite signature — in particular the two orderings of a pair of independent proposals, which yield the same union Σ{q1,q2}\Sigma \cup \{q_1, q_2\} by the associativity and commutativity of signature union. As soon as a signature admits two independent extensions, π\pi collapses distinct elements of P(Σ)P^*(\Sigma); hence π\pi is not injective, kerπ\ker\pi is non-trivial, and IP/kerπ\mathcal{I} \cong P^*/\ker\pi is a proper quotient. It is therefore not the free monad on PP.

The kernel carries more than reordering. By the instability theorem (L.5), an extension admissible against one overlap population can become inadmissible against a larger one; the admissibility-filtering relations of the quotient are exactly these time-indexed exclusions, and they are not reorderings. This is what separates I\mathcal{I} from a mere symmetrization of PP^*.

Whether I\mathcal{I} is free on some other endofunctor is a strictly stronger question, which we do not settle here. Freeness on any QQ would require the equational theory of I\mathcal{I}-algebras to be trivial; the Obligation-2 relations make that implausible, but establishing it is the content of Problem 6 (the Eilenberg–Moore category of I\mathcal{I}). We claim only what path-independence already forces: non-freeness on the generating endofunctor PP.

Remark

This resolves, in the correct direction, why predicates cannot simply be invented in parallel and merged. It is not that composition fails at a fixed moment — it does not (L.5) — but that I\mathcal{I} is a genuine quotient of the free proposal monad, carrying relations the free monad lacks: the reordering relations, which are harmless, and the time-indexed admissibility relations, which are not. A free monad would permit unrestricted parallel composition against a frozen world; the quotient forces re-synchronization whenever the world moves.

This also locates the cost model correctly. The coherence budget (A21) does not compute the cardinality of kerπ\ker\pi (that would be a static count of a set) but prices the verification surface the kernel induces: the re-checks the admissibility relations demand as predicates and populations grow. A richer kernel means more re-verification per predicate added, and it is the growth of this surface with signature size, not any literal kernel size, that the O(noverlaps)O(n \cdot |\text{overlaps}|) re-verification cost of Obligation-2 checking records.

L.7 Relation to Existing Frameworks

The coherence topos framework occupies a specific position in the landscape of categorical approaches to data integration and distributed systems. We make the comparisons explicit to identify precisely what is shared, what is new, and what remains open.

L.7.1 Spivak's Functorial Data Migration

Spivak's program(Spivak 2012)David I. Spivak, "Functorial Data Migration," Information and Computation 217 (2012): 31–51.View in bibliography models databases as functors I:CSetI : \mathbf{C} \to \mathbf{Set} from a schema category C\mathbf{C} (encoding tables, columns, and foreign keys) to Set\mathbf{Set} (the actual data). Data migration between schemas C\mathbf{C} and D\mathbf{D} is a functor F:CDF : \mathbf{C} \to \mathbf{D} inducing three adjoint operations:

ΣFΔFΠF\Sigma_F \dashv \Delta_F \dashv \Pi_F

where ΔF\Delta_F is pullback (direct image), ΣF\Sigma_F is left Kan extension (existential migration), and ΠF\Pi_F is right Kan extension (universal migration).

What the coherence topos shares with Spivak: Both use category theory to formalize data integration. Both treat schemas as categories and data as functors. The restriction maps of our presheaves correspond to Spivak's pullback functors ΔF\Delta_F.

What the coherence topos adds that Spivak does not:

  1. Vocabulary invention. Spivak's framework migrates data between fixed schemas. The functor F:CDF : \mathbf{C} \to \mathbf{D} exists before migration begins. In our framework, the signature Σ\Sigma itself evolves: agents invent new predicates, and the admissibility of the invention is the central question. Spivak has no analog of Obligation 2 (overlap agreement for invented predicates) because his schemas do not grow during operation.

  2. Scoped truth and non-Boolean logic. Spivak's instances are Set\mathbf{Set}-valued functors: a row either exists or does not. Our sheaves carry epistemic status (A4): claims can be true, false, undetermined, or in conflict, with the logic varying by context (A15). The subobject classifier Ω\Omega of the coherence topos (L.2) subsumes this; Spivak's Set\mathbf{Set}-valued model does not.

  3. Cohomological obstruction theory. Spivak does not develop obstruction theory for migration. When ΔF\Delta_F fails (the pullback does not exist or is trivial), the failure is unstructured. Our H1H^1 computation (L.4) provides a classification of the distinct ways migration can fail: a pointed set of descent classes, refined to higher classes in L.4.2. The acyclicity theorem (L.4.1) and the H2H^2 meta-obstruction (L.4.2) have no analogs in Spivak's work.

  4. Cost accounting. Spivak's adjunctions are "free" — there is no cost model for migration. Our coherence budget (A21) makes the cost of maintaining sheaf conditions explicit, and the O(noverlaps)O(n \cdot |\text{overlaps}|) re-verification cost that L.5's instability makes unavoidable — priced by the A21 framework, not derived as a bound — quantifies the engineering tradeoff.

Remark

Spivak's framework is the right foundation for structural data migration: moving data between known schemas with known relationships. The coherence topos is designed for the harder problem: semantic data integration where the schemas themselves are evolving, the relationships are being discovered (not given), and the correctness of the discovery must be certified against formal obligations.

A precise connection: the Kleisli category SigI\mathbf{Sig}_{\mathcal{I}} of the predicate invention monad (L.6) can be viewed as a category of schemas with certified evolution paths. Spivak's functors F:CDF : \mathbf{C} \to \mathbf{D} correspond to morphisms in SigI\mathbf{Sig}_{\mathcal{I}} where the evolution is a single-step conservative extension. The framework developed here extends Spivak's to the setting where schemas evolve under formal governance.

L.7.2 Goguen's Sheaf Semantics

Goguen(Goguensheaf 1992)Citation not found: goguensheaf1992View in bibliography proposed sheaves as a semantics for concurrent interacting objects, where each object has a local state and objects interact by sharing state on overlaps. This is the closest ancestor to our use of sheaves.

What we share with Goguen: The core insight — sheaves formalize when local information composes into global information — is Goguen's. Our site structure (Ctx,J)(\mathbf{Ctx}, J) is a descendant of his interaction sites.

What we add: Goguen's sheaves are on fixed interaction structures. He does not develop: predicate invention (the site's presheaf growing during operation), the instability of overlap agreement under population growth (L.5), obstruction cohomology as a classification of integration failures (L.4), or the monad structure of vocabulary evolution (L.6). Goguen also does not develop the connection to model-theoretic conservativity (A17b), which is essential for the safety guarantees of predicate invention.

L.7.3 Abramsky's Sheaf-Theoretic Contextuality

Abramsky and Brandenburger(Abramsky 2011)Citation not found: abramsky2011View in bibliography use sheaf theory to formalize contextuality in quantum mechanics: a family of local measurements is contextual if it has no global section — a presheaf that fails the sheaf condition. Their Čech cohomology detects contextuality, with H10H^1 \neq 0 implying strong contextuality.

What we share with Abramsky: The Čech cohomology machinery and the interpretation of H1H^1 as measuring obstruction to global consistency. Our H1H^1 computation (L.4) follows the same pattern.

What differs: Abramsky's presheaves are empirical models — probability distributions on measurement outcomes. Ours are data claims — assertions by computational agents about shared entities. The obstruction in Abramsky is physical (no hidden-variable model exists); ours is semantic (no consistent global identity assignment exists). The mathematics is the same; the domain and operational consequences are different. Critically, we develop the higher cohomology (H2H^2, L.4.2) and the structural classification (L.4.3), which Abramsky does not pursue in the same setting.

L.7.4 Caramello's Toposes as Bridges

Caramello's program(Caramello 2018)Olivia Caramello, Theories, Sites, Toposes: Relating and Studying Mathematical Theories through Topos-Theoretic `Bridges' (Oxford: Oxford University Press, 2018).View in bibliography uses Morita equivalence of toposes as a tool for transferring results between mathematical theories. Two theories are "Morita equivalent" if they classify the same topos, and the topos serves as a "bridge" for transferring invariants.

Connection to our work: Problem 1 in L.8 asks for the geometric theory classified by the coherence topos. If this theory can be identified, Caramello's bridge technique would immediately transfer invariants from other Morita-equivalent theories, potentially connecting coherent vocabulary evolution to problems in algebraic geometry, logic, or topology that have been studied independently.

What we add: Caramello's program is a meta-mathematical tool — it relates theories via their classifying toposes. We provide a specific instantiation: the coherence topos, with its specific site, specific presheaves, and specific theorems (acyclicity, overlap-agreement instability, the monad characterization). Our work provides a concrete object for Caramello's program to analyze.

L.7.5 Institution Theory

Goguen and Burstall's theory of institutions(Burstall 1992)Joseph A. Goguen and Rod M. Burstall, "Institutions: Abstract Model Theory for Specification and Programming," Journal of the ACM 39, no. 1 (1992): 95–146.View in bibliography is the closest prior art to the vocabulary-evolution half of this appendix, and the framework a formal-methods reader will reach for first. An institution abstracts a logical system into a category of signatures, a functor assigning sentences to each signature, a functor assigning models, and a satisfaction relation required to be invariant under signature change — the satisfaction condition, that truth is preserved along signature morphisms. This is exactly a discipline for composing and translating vocabularies through morphisms between them, and much of what A17b calls conservative extension is an institution-theoretic statement.

What we share: signatures as objects of a category, signature morphisms as the vehicle of vocabulary change, and a preservation condition along those morphisms. The Kleisli category SigI\mathbf{Sig}_{\mathcal{I}} (L.6.2) is, in institution-theoretic terms, a category of signatures equipped with certified translations.

What we add is exactly what institutions omit by design. Institutions are static and cost-free: signatures relate by morphisms that are given, and the framework says nothing about inventing a new symbol, about paying to certify it, or about the overlap obligations certification must discharge. There is no institution-theoretic analog of Obligation 2, no admissibility that can lapse as a population grows (L.5), and no cost model (A21). Institution theory tells you when a translation between vocabularies preserves truth; this appendix asks how an agent may earn such a translation, at what price, and under what conditions the earning expires. The two are complementary — the institution is the ambient logical setting, the invention monad the governed process that generates new signatures within it.

L.7.6 Capabilities Comparison

The landscape can be summarized in a table. The final column carries status markers rather than an unbroken row of affirmatives — proved in this appendix, specified elsewhere in The Proofs and instantiated here, ? open — so it reports what is established, not what is hoped for:

CapabilitySpivakGoguenAbramskyCaramelloInstitutionsThis appendix
Sheaf-theoretic coherenceimplicityesyesmeta-levelno
Vocabulary inventionnonononono△ (A17, L.6)
Obstruction cohomologynonoH1H^1 onlynonoH0H^0H2H^2 (L.4)
Overlap-agreement instabilitynonononono✓ (L.5)
Non-freeness of the invention monadnonononono✓ on PP (L.6.3)
Cost accountingnonononono△ (A21)
Scoped non-Boolean logicnonoimplicityesno✓ (A15, L.2)
Hierarchical acyclicityn/anononono✓ (L.4.1)
Enriched / graded agreementnonononono? (Problem 8)
Multi-agent operational semanticsnopartialnonono△ (L.9)

The gap is not that sheaf theory is unapplied to data integration — Goguen applied it in 1992. The gap is that no existing framework addresses the full lifecycle of vocabulary in a distributed system: invention, certification, transport, versioning, cost, and the algebraic structure of the evolution process. Each prior framework addresses a fragment. This work addresses the composition of these fragments under formal guarantees.

L.8 Open Problems for the Mathematical Community

The following problems are precisely stated and, we believe, tractable for researchers in topos theory, HoTT, and categorical logic. They are not speculative — each connects to concrete phenomena in distributed data systems and multi-agent AI.

Problem 1: Classify the geometric theory of the coherence topos. Every Grothendieck topos classifies a geometric theory T\mathbb{T} such that models of T\mathbb{T} in any topos E\mathcal{E} correspond to geometric morphisms ESh(Ctx,J)\mathcal{E} \to \mathbf{Sh}(\mathbf{Ctx}, J). What is T\mathbb{T} for the coherence topos? This theory would axiomatize exactly those structures admitting coherent vocabulary evolution. Connection: Caramello's "bridge" program(Caramello 2018)Olivia Caramello, Theories, Sites, Toposes: Relating and Studying Mathematical Theories through Topos-Theoretic `Bridges' (Oxford: Oxford University Press, 2018).View in bibliography.

Problem 2 (Partially resolved): Cohomology of structured context sites. Section L.4.1 proved Hn=0H^n = 0 for hierarchical sites, confirming the tree-acyclicity conjecture. Section L.4.2 constructed a federated site with H20H^2 \neq 0, confirming the meta-obstruction conjecture. Remaining open: (a) Compute HnH^n for random context sites (Erdős–Rényi overlap graphs) and determine the threshold for vanishing. (b) For sites arising from real organizational structures, characterize the relationship between the Betti numbers of the nerve and the operational cost of coherence maintenance. (c) Determine whether the Čech cohomology equals the derived-functor cohomology for context sites with non-Hausdorff nerve (this holds for paracompact nerves by Leray's theorem but may fail in general).

Problem 3: Extend to (,1)(\infty,1)-toposes. Witnesses (A10) carry structure: kinds, composition, coherence conditions. The correct categorical home may be an (,1)(\infty,1)-topos where witnesses are 1-morphisms and witness-equivalences are 2-morphisms. Does the coherence topos extend to an \infty-topos? Does the resulting type theory validate a scoped univalence axiom? Connection: Lurie(Lurie 2009)Citation not found: lurie2009View in bibliography, Shulman.

Problem 4: Morita equivalence of context sites. When do two context sites (Ctx1,J1)(\mathbf{Ctx}_1, J_1) and (Ctx2,J2)(\mathbf{Ctx}_2, J_2) produce equivalent sheaf categories? This would formalize when two institutional arrangements — different organizations, different view decompositions — provide the same coherence guarantees. A Morita equivalence theorem for context sites would be a formal version of "organizational isomorphism from the coherence perspective."

Problem 5: Decidability frontier for the predicate invention monad. For which fragments of the ambient logic LL is the admissibility check for predicate invention (A17) decidable? The conservativity check is decidable for propositional and equality fragments, semi-decidable for first-order, undecidable for higher-order (see Appendix K, §K.1.1). What is the precise decidability frontier when overlap agreement (Obligation 2) is included? This connects to classical questions in mathematical logic but in a new setting where the signature itself is evolving.

Problem 6 (New): Eilenberg-Moore category of the predicate invention monad. Characterize the category of I\mathcal{I}-algebras (L.6.2). An I\mathcal{I}-algebra is a signature equipped with a "vocabulary absorption" operation satisfying the monad laws. What are the free I\mathcal{I}-algebras? Can the category of I\mathcal{I}-algebras be described as a variety of algebras (in the sense of universal algebra) with explicit equational axioms? The non-freeness theorem (L.6.3) implies the equational theory is non-trivial; its explicit description would connect to Birkhoff's HSP theorem and the theory of algebraic theories(Mac Lane 1971, ch. VI)Saunders Mac Lane, Categories for the Working Mathematician (New York: Springer-Verlag, 1971), ch. VI.View in bibliography.

Problem 7 (New): Persistent cohomology of evolving context sites. As vocabulary evolves (new predicates are added, overlaps change), the context site (Ctx,J)(\mathbf{Ctx}, J) changes and its cohomology groups evolve. Does the sequence {Htn}t0\{H^n_t\}_{t \geq 0} of cohomology groups over time form a persistence module in the sense of topological data analysis? If so, the persistence diagram would classify the lifetime of obstructions: some ambiguities are transient (resolved by adding a predicate that disambiguates), others are persistent (structural, arising from the federation topology). The barcode of this persistence module would be a novel invariant of vocabulary evolution paths.

Problem 8 (New): Enriched coherence and graded agreement. This is the most consequential gap in the present theory. The framework uses set-valued presheaves, for which the sheaf condition (A13) demands exact agreement on overlaps. Yet the witness typology (A2c) admits probabilistic and attested witnesses, and the cost model (A21) operates in a regime of approximation, where local sections agree to within a tolerance rather than exactly. The two are not yet reconciled. Develop sheaves valued in a quantale Q\mathcal{Q} (or in a metric or measurable space), define a graded restriction under which "agreement on an overlap" is a distance bounded by ε\varepsilon rather than an equality, and determine whether graded local sections glue to a graded global section. The target is a graded first cohomology HQ1H^1_{\mathcal{Q}} that measures how far a family of contexts sits from coherence rather than only whether it is coherent, recovering the Boolean H1H^1 of L.4 in the case Q={0,1}\mathcal{Q} = \{0,1\}. The decisive question is whether the gluing axiom survives enrichment with the obstruction-localization guarantee (Theorem K.7) intact. Until it is answered, the probabilistic-witness layer rests on a theory proven only for the exact-agreement fragment. Connection: enriched category theory (Lawvere(Lawvere 1973)F. W. Lawvere, "Metric spaces, generalized logic, and closed categories," Rendiconti del Seminario Matematico e Fisico di Milano 43 (1973): 135–166.View in bibliography, Kelly(Maxkelly 1982)Citation not found: maxkelly1982View in bibliography).

L.9 Connection to Agentic Systems

This section connects the mathematical framework to a concrete open problem in multi-agent AI: coherent vocabulary evolution in distributed computational agents.

Current multi-agent AI systems compose outputs (text, actions, tool calls) without composing meaning. Agent A proposes "this product is sustainable." Agent B proposes "this product is eco-friendly." Are these the same predicate? Are they consistent? If Agent C needs to act on both claims, what guarantees does it have?

The coherence topos provides the mathematical infrastructure for answering these questions:

  • Predicate invention (A17, L.6): An agent can propose a new concept. The proposal carries obligations. The monad structure (L.6.1–L.6.3) ensures that sequential inventions compose safely (conservative extension), while the instability theorem (L.5) and the non-freeness theorem (L.6.3) identify exactly where synchronization is required. The algebraic content is precise: the kernel of π:PI\pi : P^* \twoheadrightarrow \mathcal{I} is the set of inadmissible combinations, and its growth rate determines the cost of coordination.

  • Obstruction cohomology (L.4): When agents in different contexts make identity claims about shared entities, H1H^1 measures the irreducible ambiguity. The acyclicity theorem (L.4.1) says hierarchical agent organizations are free of this ambiguity. The H2H^2 meta-obstruction (L.4.2) says federated agent systems face a qualitatively harder problem: not just conflicts, but conflicts about conflict-resolution strategies. The cohomological hierarchy (L.4.3) gives a system architect a precise menu of tradeoffs.

  • Scoped transport (A16, Theorem K.3): An agent's claim is valid within its scope. Transporting that claim to another agent's scope requires a certificate. The topos provides the space of possible certifications (L.3); the engineering problem is constructing one.

  • Cost accounting (A21, Theorem K.4): Coherence has a price. The cost model makes the price explicit. The scope boundary in the coherence budget is the system's declaration of how far it will pay for meaning to compose.

  • Coordination without consensus. The framework does not require agents to share objectives, adopt common logics, or trust one another. It requires only overlap discipline: where two agents' domains intersect, their assertions on the intersection must agree (A13). This is a weaker assumption than shared values, and it is computationally verifiable — the only kind of constraint agents can enforce on each other. Conservative extension (A17b) serves each agent's self-interest: it protects prior commitments. The coherence budget (A21) prices coordination without moralizing it. The sheaf condition is a structural consequence of wanting local outputs to compose globally, not a norm imposed from outside.

    The enforcement layer lies outside The Proofs. The conservative extension condition (A17b) is a mathematical specification; Factor Prime (Vol II, Ch 17) provides an enforcement mechanism — collateralized bonds whose thermodynamic cost makes defection expensive without requiring trust; The Sovereign Syntax (Vol III, Epilogue) provides the verification artifact — the receipt that gives affected parties standing to contest.

  • Landscape position (L.7): This is not a reimagining of Spivak's functorial data migration or Goguen's sheaf semantics. It is their extension to the setting where schemas evolve, agents invent vocabulary, and the cost of coherence is a first-class citizen. The comparison table (L.7.6) makes the precise contribution explicit.

The target is a substrate where computational agents can invent, certify, transport, and version predicates under formal guarantees — where inter-agent coherence is a checkable property with a computable cost. The mathematics developed here provides a specification. The engineering required to realize it at scale remains substantial.

← Back to AppendicesBack to The Proofs →