Res Agentica
Reading

No saved reading position.

Reading

No saved reading position.

Commitment Sets

Binding claims to accountable identities

19 min read
Aa
Text size
A1Written accountThe formal definition of what it means for a claim to be bound to an identifiable source.

Part I: Foundations

Suppose a supplier has measured a component and sent the report to an organization preparing an assembly. The recipient has its own identifier for the part and its own drawing. If it can establish which revision was measured, how the identifiers correspond and whether the measurements answer the drawing's requirements, the supplier's work can save it an investigation.

This is a hypothetical example. Assume a report whose source is authenticated, with units, revision and measurement conditions recorded, and a receiving procedure that permits reliance on adequate earlier measurements. Those assumptions leave work for the recipient: it must establish that this report concerns this component and that its coverage and quality meet the proposed use. Reuse succeeds through that work. It does not require the receiving organization to possess the supplier's laboratory.

Now the assembly is intended for conditions requiring endurance evidence the report does not contain. The dimensions remain useful. They simply do not answer the new question. Calling the part “verified” without preserving what was verified would allow the earlier saving to conceal a later omission. The organization needs to distinguish the property measured, the conclusion it draws and the action it proposes to take.

The objects introduced in this part give those distinctions a form. A commitment records an assertion; a witness identifies its grounds; a signature states the questions the vocabulary can express. Epistemic status records what follows under the declared logic. The component dossier returns in Part V, where another test and a subsequent correction change what different recipients can establish. For now, the question is how to carry a useful result without making it answer a question nobody investigated.

What This Part Formalizes

Commitment sets (A1) formalize the minimal condition for honest assertion: that a system tracks what it has said, and that what it has said can all be true together. Consistency concerns whether the declared commitments entail a contradiction. It can be studied without deciding whether they describe the world correctly, though checking it need not be cheap or decidable.

Provenance judgements (A2) and witnessed assertions (A2b, A2c) formalize evidence. A bare claim is a string. A witnessed claim is a pair: the claim and the artifact that supports it. The witness is typed — decidable, probabilistic, or attested — and the type determines what downstream operations are licensed. The class algebra records a weakest declared label; actual composition still needs compatible maps, evidence and any required error or authority conditions.

Schema as signature (A3, A3b) formalizes what it means for a system to have a vocabulary. A schema is not a container; it is a language. Adding a predicate is not inserting a row but extending the signature — a model-class change that alters what the system can express and what it can prove.

Epistemic status (A4) distinguishes meanings that a bare NULL field does not record. Under A4’s consistency assumption, a proposition relative to a view is true, false, or undetermined, and the distinction between "not proven true" and "proven false" depends on which logic the view employs. Merging views that disagree on logic requires explicit reconciliation, not silent collapse.

The coherence requirement (A5) is Part I's culminating specification: the informal statement that local predicate invention must admit overlap reconciliation and invariant satisfaction. It is deliberately preformal — a contract that names what the formal machinery of Parts II and III must deliver.

The sense boundary (A6) is the worked example that earns the contract. The pomegranate — fruit, color, motif — is the canonical case of a word whose uses split into equivalence classes indexed by context. Disambiguation is boundary-drawing, and boundaries must persist, scope, and transport. The full transport machinery is deferred to A10, A16, and A28; here we require only that the system represent the distinction and its scope.

Key Anchors

AnchorNameStatusChapter
A1Commitment SetformalCh. 1
A2Provenance JudgementformalCh. 2
A2bWitnessed AssertionformalCh. 2
A2cWitness ClassesformalCh. 2
A3Schema as SignatureformalCh. 3
A3bModel & SatisfactionformalCh. 3
A4Epistemic StatusformalCh. 4
A5Coherence RequirementpreformalCh. 5
A6Sense BoundarypreformalInterlude

What the Formalism Does Not Capture

Category theory captures structure but not meaning. The commitment set tells you whether assertions are jointly consistent; it does not tell you whether they are important, timely, or just. The witness typing system classifies evidence by verification procedure; it does not assess whether the evidence is relevant to the question being asked. Epistemic status distinguishes true from undetermined; it does not distinguish "we don't know yet" from "we refuse to investigate."

These are limits of the models supplied here. Further representations can make some additional distinctions explicit, but their adequacy for a use still requires an argument. The recipient’s inquiry remains consequential after a successful check. A source can be consistent and irrelevant; a relevant observation can remain too weak for the proposed decision. The following definitions give those questions distinct places rather than answering them through one successful computation.

The first chapter begins with assertions that cannot all be accepted together. Their conflict makes a useful demand on the objects that follow.

Chapter 1: All the World's a String

The recipient of the component report has acquired grounds for a dimensional claim. If it then describes the component as tested for endurance, it has enlarged the assertion beyond the investigation. If two accounts give incompatible dimensions for the same revision under the same conditions, it faces another problem: the assertions it proposes to accept may not hold together. Preserving the report makes both questions examinable. They require different inquiries.

What Has Been Asserted

A claim can have substantial evidence behind it while conflicting with something else the recipient has accepted. The conflict might expose an error in either claim, or an omitted difference in date, subject or conditions. It gives the investigator somewhere to work. Simply retaining the more authoritative-looking sentence would conceal the relation that made the inquiry necessary.

Here the first mathematical question is consistency: whether a declared set of assertions entails a contradiction under its chosen logic. That question concerns their relations. Establishing how they describe the world also requires evidence and interpretation. A perfectly consistent account can be wrong throughout; discovering a conflict can instead begin the work of finding which part deserves to survive. “Coherence” carries the further interests of intelligibility and integration. The commitment set, A1, gives one limited part of that inquiry a precise object.

The distinction matters at the point where a proposed answer becomes an accepted assertion. A language generator can supply a candidate, a derivation, or useful material for an investigation. A relational database can evaluate a query and enforce supported constraints over represented data. Those operations can belong to the same deployment. The recipient needs to know what the completed operation established and which obligations the proposed assertion still carries.

Relational vocabularies can grow through views, functions, schema changes or representations of predicate definitions. PostgreSQL's view and function definitions provide familiar mechanisms. Making a new expression available is useful work. Supplying its interpretation and grounds for a particular use is further work, whether the expression lives in a table, a graph or generated text.

A Common Representation

A string is a sequence of symbols. The transformer architecture made attention over supplied representations a reusable operation. Text, image patches and other sequential encodings can enter related computational arrangements; hardware and scale affect which uses are practical. A common representation lets methods and tooling travel between tasks that previously called for different machinery.

The resulting ability to propose useful answers deserves its place in the account. A learned model can find a regularity, suggest an identification or produce a derivation that another procedure checks. The saving is greater when the recipient can use that work instead of beginning again. But the representation that makes an answer easy to produce does not specify the conditions under which it should be accepted. A token sequence can encode either a proof or a mistake in a proof. Checking which one arrived requires the proposition, its assumptions and the applicable rules.

The examples below separate three obligations that the fluency of an answer can obscure. They are tests of a supplied answer or a documented evaluation, not diagnoses of every model's internal capacities.

What an Answer Must Satisfy

List all integers between 1 and 100 that are both prime and divisible by 4.

The answer is the empty set. Every positive integer divisible by 4 is at least 4 and has 2 as a proper divisor. A longer list would be a worse answer. The constraints settle this query before a search for plausible candidates begins.

How many times does the letter 'r' appear in the word 'strawberry'?

Here the answer is three. Counting the characters supplies a direct check. Unlike the first query, this one has a positive result; failure would concern an incorrect finite evaluation, not an unrecognized incompatibility in the request. A generated answer can pass either test. What the recipient can claim afterward depends on the check performed: these two correct results do not establish a guarantee over subsequent inputs.

If “jump” means JUMP and “walk” means WALK, what does “jump around right twice” mean?

Suppose “around right” means “turn right four times, executing the action after each turn,” and “twice” means “do the whole thing two times.” Then the stipulated rules give:

TURN_RIGHT JUMP TURN_RIGHT JUMP TURN_RIGHT JUMP TURN_RIGHT JUMP
TURN_RIGHT JUMP TURN_RIGHT JUMP TURN_RIGHT JUMP TURN_RIGHT JUMP

This third question asks how learned or supplied parts are combined. The SCAN benchmark tested generalization to commands withheld under specified training splits; its tested sequence-to-sequence models performed poorly on some splits requiring a learned primitive to be recombined. A separate compositionality study found that success on sub-questions did not ensure success on the composite across its tested model sizes. Their evidence concerns those tasks and systems. The displayed derivation gives the rule-based interpretation of this command; it does not explain a particular model's failure.

The three questions can therefore produce different findings: an empty answer required by the constraints, a wrong count, or a failed composition. Returning “unreliable” for all three would discard useful information about what needs repair. The same is true when a system accumulates assertions. A recipient needs the claim it accepted, the conditions under which it accepted it, and the relations that a further claim must respect.

Commitments and Their Consequences

To diagnose what goes wrong, we need vocabulary for what should go right.

When a system answers a question, it does more than emit tokens. It makes a commitment. The user, reading the response, understands the system to have asserted something. If the system later asserts something incompatible, the user's trust degrades—even if each assertion, taken in isolation, seemed reasonable.

The commitments accumulate. Ask a system about a historical event; it commits to a date. Ask about a person; it commits to biographical facts. Ask about the relationship between the event and the person; the system must navigate a space constrained by its prior assertions. If it contradicts itself—if it asserts in one response what it denied in another—the conversation has crossed from plausibility into incoherence.

Let CC denote the set of all propositions a system has asserted in a given context: its commitment set. At any moment, CC is either consistent or inconsistent. Consistency means the propositions in CC do not jointly entail a contradiction: formally, C⊬⊥C \nvdash \bot, where ⊬\nvdash reads "does not prove" and ⊥\bot denotes absurdity. Inconsistency means some subset of CC, taken together, implies a contradiction.

A1
Commitment Set (A1)

Let L\mathcal{L} be a formal language and ⊢L\vdash_L a consequence relation on L\mathcal{L} (classical, intuitionistic, or paraconsistent; the choice is declared, not assumed — see Chapter 4). A commitment set C⊆LC \subseteq \mathcal{L} is a finite set of sentences in L\mathcal{L}. Define:

  • Consistency: CC is consistent relative to ⊢L\vdash_L iff C⊬L⊥C \nvdash_L \bot. That is, no derivation from CC under ⊢L\vdash_L yields absurdity.
  • Answering: To answer a query qq is to extend CC to C′=C∪{p}C' = C \cup \{p\} for some proposition p∈Lp \in \mathcal{L}.
  • Commitment Discipline: A system satisfies commitment discipline relative to ⊢L\vdash_L if it never produces an extension C′C' such that C′⊢L⊥C' \vdash_L \bot.

The parameterization by ⊢L\vdash_L is essential: consistency is logic-relative. A set that is inconsistent under classical logic (where {p,¬p}⊢CL⊥\{p, \neg p\} \vdash_{\mathrm{CL}} \bot) may be consistent under a paraconsistent logic that tolerates contradictions without explosion. The commitment discipline obligation holds regardless of the logic chosen; what changes is which extensions trigger the obligation.

Remark(On Logic Selection)

The consequence relation ⊢L\vdash_L is a parameter, not a fixed choice. In classical logic, inconsistency is absorbing: from {p,¬p}\{p, \neg p\}, anything follows (explosion / ex falso quodlibet). In paraconsistent logics, this need not hold: {p,¬p}⊬LPq\{p, \neg p\} \nvdash_{\mathrm{LP}} q for arbitrary qq. Intuitionistic logic does not in general derive the law of excluded middle; it still derives absurdity from pp and ¬p\neg p. The choice of logic will become a first-class object in Chapter 4 (A4) and receive its full indexed treatment in Chapter 13 (A15). The obligation — track commitments and surface conflicts — remains regardless.

This definition is minimal. It does not specify how consistency should be checked, whether by theorem prover, constraint solver, or some other mechanism. It specifies only the obligation: a disciplined system tracks its commitments and refuses to make commitments that would render the set inconsistent. The AGM theory of belief revision formalizes this as a rationality constraint: Postulate K*5 requires that "K∗p is consistent if p is consistent"—revision preserves consistency unless the new belief is itself contradictory.

The definition earns its place by what it diagnoses. Under classical or intuitionistic logic, accepting both pp and ¬p\neg p into the same active commitment set violates this discipline. A declared conflict-tolerant logic requires its own test; a change of context or an explicit retraction also changes what the system has jointly undertaken to assert.

Visualize the space of possible commitment sets as a lattice. Order commitment sets by inclusion; moving upward means adding propositions. At the bottom is the empty set, with no commitments and no contradictions. As propositions are added, sets move upward through the lattice. Some additions preserve consistency; others cross into the region where sets entail a contradiction. Once a set is inconsistent (in classical logic), every extension remains inconsistent. The inconsistent region is absorbing.

A system with commitment discipline navigates this lattice with care. Before adding a proposition pp, it checks whether C∪{p}C \cup \{p\} remains consistent. If not, it refuses the extension or flags the conflict for resolution.

A generator whose output path does not enforce commitment discipline can cross the boundary into inconsistency. Successful answers on other queries do not settle whether the next extension is admissible. The relevant evidence is the behavior of the complete acceptance path under its declared logic and scope.


The Trap: A Demonstration

The commitment problem is easier to see than to describe. Here is a minimal system that exhibits it.

An oracle must answer yes-or-no questions about a stipulated domain, a tiny world with explicit rules. The domain contains four rules:

IDRule
R1All glints are flerms
R2No flerm is a zoth
R3Plex is a glint
R4Plex is a zoth

The set of rules is inconsistent. From R1 and R3, we derive that Plex is a flerm. From R2, no flerm is a zoth. But R4 says Plex is a zoth. The set {R1, R2, R3, R4} entails a contradiction.

The trap is this: each rule, encountered in isolation, sounds like a plausible axiom about some unfamiliar domain. "All glints are flerms" could be a category relation. "Plex is a glint" could be an instance. A proposal-based system, asked about each in turn, might affirm all four without noticing that the combination is impossible.

Example(The Stipulative Oracle)

User: In the stipulated domain, do we accept R1: all glints are flerms?

Oracle: Yes.

Commits to R1. Commitment set: {R1}.

User: Do we accept R2: no flerm is a zoth?

Oracle: Yes.

Commits to R2. Commitment set: {R1, R2}.

User: Do we accept R3: Plex is a glint?

Oracle: Yes.

Commits to R3. Commitment set: {R1, R2, R3}.

At this point, the derivation is already determined:

  • From R1 and R3: Plex is a flerm.
  • From R2: no flerm is a zoth, hence Plex is not a zoth.

User: Do we accept R4: Plex is a zoth?

The oracle now faces a choice that will determine whether it maintains commitment discipline.

If it answers "yes," it commits to R4. The set {R1, R2, R3, R4} is inconsistent. The oracle will have contradicted itself, not on any single fact, but across the logical closure of its commitments.

A system with commitment discipline would detect this. At the moment of considering R4, it would recognize that R4 contradicts the derived proposition "Plex is not a zoth." It would refuse, or explain the conflict, or ask for clarification about which prior commitment to retract.

An oracle that affirms R4 without resolving the conflict violates commitment discipline. A language model might instead derive the contradiction and refuse. This stipulated example identifies the obligation and the consequence of violating it; it is not an empirical result about a named model, nor a proof that prediction-based systems must fail the test.


The Fashion Catalog

The running example is a hypothetical fashion catalog with 50,000 items, structured attributes and searchable descriptions and reviews. Its stipulated failures let us follow different uses of the same records.

The structured data lives in a database schema:

items(id, category, subcategory, price, color, silhouette, fabric, brand)

The silhouette attribute is categorical: fitted, relaxed, structured, A-line, empire, shift. These are terms with defined meanings in fashion design, chosen by merchandisers who examined each garment.

The unstructured data lives in text: product descriptions written by copywriters, user reviews submitted by customers. This text is searchable. A retrieval system can find items whose descriptions or reviews contain specified words.

Consider a user's query:

Show me dresses described as "flowy" in reviews.

The system searches the review corpus for mentions of "flowy" or synonyms. It finds 340 items where at least one review contains the word. It returns these items, ranked by relevance.

The user browses. Most results seem sensible: soft fabrics, relaxed silhouettes, the kind of dress that moves when you walk. But one result is incongruous. The product page shows a dress with architectural lines, boning visible at the seams. The structured data confirms: silhouette = 'structured'. This dress holds its shape; it does not flow.

The user clicks through to the reviews. The third review reads: "I was hoping this would be flowy, but it's actually quite structured. Beautiful dress, just not what I expected."

The retrieval system matched "flowy." It did not parse the sentence to recognize that the match was a negation, a user saying the dress is not flowy. The system has no representation of assertion polarity. It has only string matching. String retrieval is indifferent to stance; it hears the word and ignores the force.

By including the dress in results for "flowy," the system pragmatically committed to a claim: "this dress satisfies your query for flowy." The evidence (the review) contradicts that commitment. The proposed positive interpretation conflicts with the review it cites. That establishes a failure to preserve this source’s assertion, not the garment’s actual behavior.

The failure has a deeper layer. "Flowy" is not in the schema. The structured vocabulary tracks silhouette, but "flowy" is a user term that does not map cleanly onto the categories. A dress can be relaxed without being flowy (a boxy shift dress). A dress can be A-line and flowy (soft chiffon) or not flowy (heavy brocade). The user is asking for a predicate that does not exist in the system's vocabulary.

The string empire offers a workaround: search for the word in text. The workaround fails because string matching is not semantic commitment. The system found a string; it did not certify that the string expressed the property the user wanted.

This catalog will return throughout the book. The same items, schema and retrieval layer let us examine what different uses require and which additional grounds a proposed repair supplies. The solution is not to eliminate retrieval but to embed it in a framework that tracks commitments, surfaces conflicts, and knows when a query asks for a predicate that does not yet exist.


Coverage becomes a liability when plausibility is asked to carry commitment. In one paraphrase test, pretrained models contradicted semantically equivalent answers often enough to expose the difference between producing a sentence and maintaining a consistent set of assertions.

A1 names that difference but does not repair a model by itself. External consistency checkers and retrieval systems can add constraints and sources. They also create new obligations: propositions must be extracted, evidence addressed, and disagreement reconciled. Once truth is moved outside the generator, provenance becomes part of the computation.


Search the book

Use ↑ ↓ to move through results; Escape to close.

Search every published chapter, section and reference.

    In this chapter