Provenance and Witness Typing
Witnessed assertions and witness classes
Aa
An unapparent connection is stronger than an apparent one.
This chapter formalizes provenance judgements (A2), witnessed assertions (A2b), and a witness-class composition algebra (A2c) that types evidence by verification regime -- decidable, probabilistic, or attested. The chapter examines retrieval pipelines that lose the source and scope information needed for downstream composition. Its witness structures make those dependencies explicit. For the institutional questions of provenance, received evidence and independent assessment, see Vol I, Chapter 6 (Evidence Without Custody).
The Confident Synthesis
The legal research system had access to everything it needed. In its corpus of appellate opinions (a toy jurisdiction used here for illustration), two holdings from the same court, both correctly retrieved, both authoritative within the corpus. The user asked a simple question: "Are liquidated damages clauses enforceable?"
The system synthesized a confident answer. The prose was measured, the citations real. It stated that liquidated damages clauses are enforceable when the stipulated amount represents a "reasonable forecast" of anticipated harm. This was the standard from a 2019 holding. The answer did not mention that the same court, in a 2023 opinion, had held such clauses unenforceable when the underlying contract was "procedurally unconscionable." The later holding did not contradict the earlier one; it refined it, carving out an exception that might or might not apply to the user's situation. The answer omitted the condition that could change its application.
This is not hallucination. The system did not fabricate a case. Both holdings are real within the corpus; both are authoritative; both were retrieved correctly. The failure is subtler and more structural: the system committed to a synthesis without surfacing the scope relationship between its sources. It had two attestations and no mechanism to notice that one qualified the other. The problem was not inconsistency of sources; it was loss of scope—a refinement was flattened into a slogan.
The retrieved holdings were adequate to expose the qualification. Their incorporation into an answer lost it. Preserving their source relationship would let another reader examine the condition before relying on the general rule.
The Pattern Repeats
The failure is not limited to law. Wherever systems aggregate from multiple authoritative sources, the same structure appears.
Consider the fashion catalog aggregating inventory from three suppliers. Supplier A's feed lists a dress as "silk blend (70% silk, 30% polyester)." Supplier B's feed lists the same SKU (same product, same barcode) as "100% polyester." Both suppliers are in good standing; both feeds are canonical as far as the system knows. A third supplier lists the fabric as "satin," which is a weave rather than a fiber composition, orthogonal to the question entirely.
A user asks: "What is the fabric composition of this dress?"
The system might confidently state one value, choosing a supplier without disclosure. It might report “silk-polyester satin blend,” combining one composition claim with the weave while concealing the contrary composition report. It might refuse to answer without explaining why. None of these is correct. The correct response acknowledges the conflict: Supplier A claims silk blend; Supplier B claims polyester; resolution may require further source investigation or testing, while an interim decision needs its own governing rule.
The two cases ask different things of provenance. The legal example needs a condition on the use of an earlier holding; the catalog needs a comparison of opposed composition claims and a separate weave claim. Flattening either into a single answer loses the work a recipient needs to perform, but a refinement of scope is not itself a contradiction.
Retrieval as Operating System
Retrieval was once a read operation. You queried a search engine; the engine returned a ranked list; the list was the answer. The user's job was to evaluate the results. The system's job was to find them.
Retrieval-augmented generation changed the verb1. Now retrieval is a write operation. The system searches, retrieves, and injects the results into a context window. The language model then generates a response that treats the injected text as part of its available knowledge. The boundary between "what the system knows" and "what the system just retrieved" dissolves. The retrieved text becomes raw material for the next token prediction.
Tool use generalized the pattern. Retrieval became one action among many: search the web, query a database, call an API, execute code, check a calendar. Each action returns a result. Each result enters the context. Each context shapes the generation that follows. From the model's perspective, a tool call is a syscall: dispatch a request, receive a response, incorporate the response into the ongoing computation.
Agent architectures orchestrated these syscalls into plans. The system reasons about which tools to use, in what order, with what parameters. It dispatches multiple calls, aggregates results, handles failures, and synthesizes responses. The agent is a scheduler. Scheduling determines ordering: which evidence enters first, which commitments get made, which conflicts become visible.
Each transition expanded the commitment surface area.
| Paradigm | Operation | Systems Analogy | Commitment Surface |
|---|---|---|---|
| Search | Read | I/O | Results are the answer |
| RAG | Write | Memory injection | Claims enter context |
| Tool use | Action | Syscall | Results become premises |
| Agents | Orchestration | Scheduler | Multiple sources multiply conflicts |
The more sources consulted, the more potential for disagreement. The more actions taken, the more assertions implied. The more complex the orchestration, the harder it becomes to trace which source contributed which claim.
Within the composition discipline developed here, retrieval results must remain addressable when they enter generation. A stable handle lets a later stage identify the claim and recover its governing information. One representation is a tuple of the form:
Addressability is what lets downstream reasoning distinguish a claim from a statute and a claim from a blog post, a measurement from a calibrated instrument and a guess from an uncalibrated one.
A context window does not by itself enforce preservation of these addresses. If a pipeline injects claims as bare text and drops their identifiers and provenance links, a downstream stage may no longer distinguish which supplier, item or holding each claim concerns. The loss follows from what the pipeline discards, not from the use of text as its representation.
Claims Without Addresses
The claim "Paris is in France" looks the same whether it comes from an encyclopedia, a user's personal note, or a hallucinated interpolation. The claim text alone does not establish its evidentiary status. Their identical proposition does not carry identical evidence. A claim from a verified database carries different obligations than a claim from an unverified source. A claim attested by a laboratory carries different weight than a claim inferred by pattern matching.
This is the provenance gap: the space between what a claim asserts and what would license believing it2. A pipeline that retains only the claim text can discard the information needed to examine that gap. Identifiers and provenance relations can instead be encoded in text or carried alongside it. The failure considered here occurs when the downstream stage receives neither those relations nor a means of recovering them.
The consequence is structural:
Lemma (Address Loss): If a downstream stage receives claims as bare strings without stable identifiers for (referent, attribute, source, time, logic-regime), then it cannot, in general, decide whether two claims are in conflict—because conflict detection reduces to a join problem over those identifiers, and dropping provenance drops the keys.
A module that receives the string "fabric: silk blend" cannot know whether another module received "fabric: 100% polyester" for the same item. A module that receives "liquidated damages clauses are enforceable" cannot know whether another module received a holding that carves out exceptions. The fabric conflict and the legal qualification remain in their sources; this processing stage has lost the information needed to examine them.
Citations and audit trails already perform substantial parts of this work. A case name and pinpoint reference can lead a reader to the qualification an answer omitted. A financial audit can follow a recorded amount into its supporting documents. Those practices make an inquiry possible; they do not guarantee that every source is recoverable or that its contents establish the proposed claim.
But a citation is not the same as a witness.
A citation locates: it tells you where a claim came from, providing enough information to retrieve the source. A witness licenses: it tells you how to apply the claim by specifying verification regime, scope, and consumption contract.
When a system appends "Source: Supplier A" to a claim, it has added a citation. When a system passes the pair , where includes the attestation type, the verification method, and the conditions under which the claim may be used, it has provided a witness.
A citation that merely names a source is insufficient for this composition discipline. Claims must remain linked to the evidence, scope and verification conditions on which downstream use depends. The witnessed assertion makes that dependency explicit. Strings with footnotes can support retrieval and independent assessment; when they omit the identifiers and scope needed for a proposed composition, that composition remains unwarranted.
The Provenance Judgement
Given the failure (systems that retrieve correct information and still produce incoherent outputs), what is the minimal object that would prevent it?
The answer is a typing judgement that keeps claims tethered to their evidence.
Define a typing judgement for provenance:
where:
- is a context: background assumptions, available sources, trust anchors
- is a proposition: the claim being made
- is the witness type (read: "the type of witnesses for "): the type of evidence that would support
- is a witness: a term of type , constituting the evidence itself
The judgement asserts: under context , is evidence that holds.
Consumption rule: Downstream steps must receive , the claim paired with its witness, not bare .
The judgement records the dependency the opening synthesis omitted. A procedure must still preserve and examine it. In the court-record case, the pairs would distinguish:
- : "Liquidated damages clauses are enforceable if reasonable forecast" witnessed by 2019 holding
- : "Such clauses are unenforceable if procedurally unconscionable" witnessed by 2023 holding
With both pairs in hand, downstream reasoning can see the relationship. The 2023 holding is not a contradiction of the 2019 holding; it is a refinement that carves out an exception—a scope constraint on the earlier rule. Without retained information about the sources and their scopes, that relationship is unavailable to the stage deciding how to combine them. The opening synthesis failed by collapsing a rule lattice into a single sentence.
The judgement tells us that a witness exists. It does not yet tell us what kind of witness: what verification regime governs it, what operations are valid on it, what level of trust downstream modules may assume.
A2b — Witnessed Assertion:
A witnessed assertion is a pair where:
- is a proposition
- is a witness
- is a verification function
The verification function is typed according to the witness class.
A2c — Witness Classes:
Witnesses are typed by verification regime:
| Class | Verification | Output | Examples |
|---|---|---|---|
| Decidable | Total, deterministic | Type check, arithmetic, schema validation | |
| Probabilistic | Terminates with confidence | ML classifier, embedding similarity, statistical test | |
| Attested | Checks provenance chain | Supplier contract, certificate signature, human attestation |
The class identifies a checking regime. A decidable check establishes the proposition it actually checks. Statistical evidence needs a specified procedure and interpretation of its bounds; a model score alone supplies neither. An attestation identifies a source whose competence, grounds and authority must be assessed for the proposed use.
A metadata convention. One can order the labels and combine labels by minimum. That operation is associative, with as its unit. These elementary facts concern labels under a chosen order. They neither rank the reliability of all evidence nor establish that the underlying witnesses compose. A flawless check of a signature may establish less about a fact than a well-supported human report.
Operational Consequences of Witness Class
The operation must be justified by the evidence and the relationship being claimed, not by its class label alone.
| Regime | What must be established for further use |
|---|---|
| Decidable | The checked proposition and its hypotheses; compatible maps or proofs for composition; the exact sheaf hypotheses for unique gluing |
| Probabilistic | The target quantity, population and procedure; dependence and error propagation under composition; the declared comparison rule |
| Attested | The source's grounds, scope and authority; the receiving use's requirements; how conflicting reports are investigated |
Two completed arithmetic checks can disagree because they have different inputs. A valid signature can cover a false assertion. An attestation from one authority can combine with independent evidence from another without both belonging to a single chain of command. These are ordinary differences in evidentiary work, which the witness record must retain.
Exact gluing applies where the local values and restrictions satisfy A13. A tolerance policy or statistical reconciliation performs a different operation unless further hypotheses connect it to that theorem. Widening an interval to encompass two results may hide incompatible models; it does not automatically preserve coverage.
The fabric composition claim from Supplier A has an attested witness: the supplier's feed, governed by a contractual relationship, checkable by "do we trust Supplier A for fabric data?" The holding from the 2019 case has an attested witness: the court reporter, checkable by "is this a valid citation to the official record?"
An embedding model can supply a similarity score. Probabilistic interpretation requires a calibration or statistical model that relates it to a defined target; cosine similarity does not carry confidence bounds by itself. A type checker's claim that a program is well-typed has a decidable witness: the derivation tree, checkable by replaying the algorithm.
The three classes are not exhaustive (finer distinctions exist), but they capture the major verification regimes that matter for system design. Keeping the regimes distinct lets a recipient ask the right question of each. A type label alone cannot answer it.
Taking the minimum of labels is associative. Composition of actual witnesses remains a separate, partial operation: their meanings, maps, property footprints and scope conditions must fit. Empty common scope need not prove a contradiction; it can mean that there is no shared use to certify. Problem 8 in Appendix L records the unestablished extension from exact agreement to general probabilistic or attested reconciliation. The label algebra supplies no proof of that extension.
A verification policy can end a certificate chain at a root it accepts. That decision establishes where the policy stops asking for another authentication link. The attester’s competence, the grounds for its report and its authority for the receiving purpose remain different questions. An accepted root does not make the report materially true or give every act it describes legitimate authorization.
The Fashion Catalog Revisited
Return to the catalog with the new machinery in hand.
For SKU DRESS-2847, the system now maintains:
- : "fabric = silk blend (70/30)" witnessed by Supplier A feed, class: attested, authority: Supplier A
- : "fabric = 100% polyester" witnessed by Supplier B feed, class: attested, authority: Supplier B
- : "fabric = satin" witnessed by Supplier C feed, class: attested, authority: Supplier C
When queried, the system sees that and are in conflict: they assert different values for the same predicate (fiber composition) on the same entity. It sees that is orthogonal: "satin" is a weave, not a fiber, so Supplier C is making a claim about a different predicate that the string pipeline had merged into "fabric." The conflict is not only value-level; it is predicate-level. The system can now respond correctly:
"Conflicting provenance for fabric composition. Supplier A (attested) claims silk blend. Supplier B (attested) claims 100% polyester. Supplier C's claim (satin) addresses weave, not fiber. Further testing or source investigation may resolve the composition question; a decision made before that resolution needs its own governing rule."
This is not a refusal to answer. It is an answer that respects the evidential situation. The user learns that the data is contested and can choose how to proceed. The system has not committed to a claim it cannot support.
Same Referent, Different Sources
A harder case: the system retrieves a record from Catalog A that calls an item a "cocktail dress." It retrieves a record from Catalog B that calls what appears to be the same item an "evening dress." Are these the same dress? If so, which label is correct?
Embedding similarity might be high; the images look alike, the descriptions overlap. But similarity is not identity. The items could be the same dress sold under different names. They could be different dresses that happen to look similar. They could be the same physical garment at different price points.
The commitment-discipline question is: can the system assert "these are the same item" without a witness? The answer is no. Similarity can propose identity; only a witness can certify it.
A shared key gives a directly checkable equality of identifiers. Using it to identify a product also depends on the source’s assignment convention, namespace and relevant version. Matching measurements or a merchandiser’s examination can supply other grounds. The checking class does not rank those grounds by strength: a perfect comparison of the wrong keys can support less than an informed report about the garments.
Without witnesses, the system cannot safely merge the records3. Adequate evidence can justify a merge; merely attaching a witness record does not make an identification correct. The provenance travels with the claim.
This is the reference problem, what philosophers call the "morning star / evening star" puzzle4. Two names can refer to the same object (Venus) without the system knowing it. The solution is not to assume identity; it is to require witnessed equivalence before asserting it.
Negation Depends on Source
One more complication. Supplier A's catalog operates under the closed-world assumption5: if an attribute is not listed, the item does not have that attribute. For Supplier A, if sustainable = null, the item is not sustainable.
Supplier C's catalog operates under the open-world assumption: if an attribute is not listed, the item's status is unknown. For Supplier C, if sustainable = null, the item's sustainability is undetermined.
A user asks: "Show me sustainable dresses."
If the merge discards these source conventions, the recipient cannot recover their different implications from the merged value alone. Treating Supplier C’s missing values as negative would add an unsupported completeness inference. Keeping them unknown still leaves a query-policy choice: a confirmed-positive filter can exclude them, while an exploratory result can display them separately. OWA supplies no instruction to include every unknown item.
The solution is to include logic in the provenance. The witness for Supplier A's sustainability claims carries a CWA tag: absence means negation. The witness for Supplier C's claims carries an OWA tag: absence means unknown.
When the query arrives, the system can now respond appropriately:
"Results from Supplier A reflect closed-world semantics (missing = not sustainable). Results from Supplier C reflect open-world semantics (missing = unknown). Do you want to see: (a) only confirmed sustainable items, (b) confirmed + unknown, or (c) all items with sustainability status displayed?"
This is not pedantry. It is the difference between a system that silently makes epistemic commitments and a system that surfaces them for user decision. Logic is part of provenance.
The Typed Pipeline
The full picture is now clear. Retrieval produces candidate claims. Each claim should be wrapped with a witness before passing downstream. The witness specifies the source, the verification regime, and the logic under which the claim was made. Downstream reasoning consumes the pair , not the bare string .
Retrieval Candidate Claims Provenance Reasoning
Sources --> (raw strings) --> Wrapper --> (p, π) with
verify(p,π) obligations
The failure case is when the wrapper and equivalent recoverable provenance information are dropped. Claims become "free"—untethered from their evidence, indistinguishable from each other, impossible to reconcile when they conflict. Downstream modules see strings; they cannot see sources. Conflicts that exist in the retrieval layer become invisible in the reasoning layer.
The join needs the identifying information its comparison consumes. Text can encode that information, and an independent inquiry can sometimes recover it. The failure occurs where neither the information nor adequate means of recovery is available.
A downstream stage deprived of both the relevant identifiers and a means of recovering them lacks the information this join requires. Supplying that information repairs the loss; renaming the unexamined claim does not.
Consequence
Retrieval gives a recipient material to examine beyond a generator’s trained parameters. In the legal example, the second holding changes the reach of the first. In the catalog, a weave and a fiber composition answer different questions despite arriving in one field. Preserving a source address makes those inquiries possible; classifying the source does not perform them.
A recipient may discover enough to settle a question, obtain new evidence, or find that the proposed use remains unsupported. The wrapper’s work is to keep the grounds and their scope available through that inquiry. Bare repetition of an assertion cannot count as another investigation.
The next chapter follows a different part of the difficulty. A catalog may preserve every source and still lack a maintained expression for what its user wants to ask. Adding that expression changes more than the space available in a row.
Litmus Cases
This chapter advanced two tests from the touchstone battery.
| Case | Name | Chapter 1 Status | Chapter 2 Status |
|---|---|---|---|
| T2 | Reference | Foreshadowed as "identity without witnesses" | Advanced: same referent, different sources; provenance distinguishes |
| T5 | Negation | Foreshadowed as "negation without witnesses" | Advanced: logic is part of provenance; CWA/OWA as source metadata |
T2 (Reference): Two expressions may refer to the same entity ("cocktail dress" in Catalog A, "evening dress" in Catalog B), but the system cannot assert identity without a witness. Embedding similarity proposes; witnessed equivalence certifies. Resolution requires the machinery of Part II.
T5 (Negation/Absence): The meaning of a missing attribute depends on the source's logical regime. Closed-world sources treat absence as negation; open-world sources treat absence as unknown. Merging sources without logic annotations produces silent commitments. Full resolution requires the epistemic-status machinery of Chapter 4.
The touchstones are not resolved here. They are sharpened: the failures now have names, and the names point toward the objects we must build.