Res Agentica
Reading

No saved reading position.

Reading

No saved reading position.

Epistemic Status

What a system knows about what it knows

18 min read
Aa
Text size
A4Written accountDerivable, refutable and undetermined claims within a declared consistent theory.

What we cannot speak about we must pass over in silence.

— Ludwig Wittgenstein, Tractatus §7; D. F. Pears and B. F. McGuinness translation (1961)

This chapter formalizes epistemic status (A4): the truth value of a proposition relative to a view is true, false, or undetermined, and the distinction depends on which logic the view employs. One unqualified NULL field can conceal several reasons for missingness; a merge must preserve the distinctions its inferences require. The reader who wants the historical argument for why absence has always been contested territory should read Vol I, Chapter 5 (The Empire of Tables).

The Lie of NULL

In the hypothetical catalog, a merge combines feeds from three suppliers. Supplier A's feed includes a sustainable field: items are marked TRUE, FALSE, or NULL. Supplier C's feed has the same field, but with different conventions. For Supplier A, a NULL in sustainable means "this item is not sustainable": the supplier operates under a completeness declaration for this predicate, where absence of a positive claim implies the negative. For Supplier C, a NULL means "we don't know": the supplier acknowledges incomplete information.

Supplier B's feed has no sustainable field at all.

A user asks: "Show me sustainable dresses."

What should the system return?

If only the merged field survives, the table cannot recover the omitted source conventions. Supplier A's NULLs, Supplier C's NULLs, and the synthetic NULLs created for Supplier B's items all look identical. The completeness profiles that govern their interpretation—what absence licenses you to conclude—were never stored. The database sees the same symbol; the semantics are different.

The merge has made an inference harder to inspect without changing the printed symbol. Sometimes the predicate is absent from a source’s vocabulary; sometimes the predicate exists and its value is missing. The recipient needs that distinction before deciding what the blank permits it to conclude.

A missing value does not state the conditions under which absence becomes evidence.


Four Interpretations of Missingness

NULL is not one thing. It is a symbol that collapses at least four distinct situations into a single representation.

Three storage states describe what the data actually records:

Storage StateExampleWhat It Means
Unknown"We haven't tested whether this dress is machine-washable"The predicate applies, but the value is not known; it might be TRUE or FALSE
Not applicable"Machine-washable is not a meaningful attribute for jewelry"The predicate does not apply to this entity; the question is ill-formed
Absent-from-source"The supplier's feed doesn't include this field"No information was transmitted; the absence is about the data source, not the entity

Each storage state implies different behavior. Unknown values might become TRUE or FALSE with more information. Not-applicable is a typing failure: the predicate's domain does not include this entity (or includes it only via a partial function)—asking whether a necklace is machine-washable is a category error, and such values should not satisfy either positive or negative queries about the predicate unless the query explicitly ranges over applicability. Absent-from-source values should be handled according to the source's conventions, which may differ from source to source.

One derived state is not stored but inferred:

Derived StateInferenceWhat It Means
Inferred false"If sustainable isn't listed, it's not sustainable"Under a completeness assumption, absence of a positive claim licenses the negative

This fourth interpretation is different in kind. It is not a fact about what the data contains; it is a conclusion drawn from what the data omits, given an assumption about completeness. The inference requires a declared closure rule for the predicate in question. Using its conclusion as evidence about the world also requires adequate grounds for that completeness claim.

A NULL field alone does not distinguish these situations. SQL can store an absence reason, source and policy in additional fields or related tables. If the representation omits them, sustainable IS NULL selects missing values without identifying why they are missing or what may be inferred from them.

The conflation propagates. A CASE statement that maps NULL to a default may implement one policy correctly and misrepresent another. COUNT(sustainable) counts non-NULL entries, including both TRUE and FALSE; it is not a count of sustainable items. Choosing a denominator or counting potential positives requires the question and absence convention to be represented. The query cannot recover a distinction the supplied records omit.

The missingness policy supplies an inference that the marker alone does not contain.


Completeness and Inference

The interpretation of absence depends on a completeness declaration—a claim about what the source takes itself to have covered.

The Closed-World Assumption (CWA) is the reasoning of completeness. If P(x)P(x) is not derivable from the source, infer that P(x)P(x) is false. The source claims to know everything relevant; what it cannot derive, it denies.

A complete list of the entries in a directory can establish that a number is not listed there. It does not establish that a person has no phone. A catalog may declare a particular inventory predicate complete; that declaration does not close every attribute, and it needs adequate support for the use made of it.

The Open-World Assumption (OWA) leaves absence open. If neither P(x)P(x) nor its negation is derivable, the question remains undetermined. An explicit or derived negative can still establish a negative conclusion; an unreported positive alone cannot.

A biography’s omission of a birthdate need not assert that its subject had none. Nor does a study’s silence about a side effect, by itself, establish that the effect was absent. The evidentiary question is what the inquiry covered and what its reporting convention permits the reader to infer.

The inferred result depends on the policy; the unchanged SQL query does not. SELECT * FROM items WHERE NOT sustainable excludes NULL rows under ordinary SQL three-valued evaluation. To treat a missing value as false, the application must implement that policy—for example with an explicit conversion—and justify the completeness assumption it applies. Declaring CWA in prose does not change the query engine’s behavior.

Example(Birds and Flying)

Schema: birds(id, species, can_fly)

idspeciescan_fly
1sparrowTRUE
2penguinFALSE
3ostrichNULL
4kiwiNULL

Query: "Which birds can fly?"

Both regimes agree: {sparrow}. The NULL values are not TRUE.

Query: "Which birds cannot fly?"

Under an explicit policy that converts these missing values to false: {penguin, ostrich, kiwi}. This is a changed evaluation, such as WHERE NOT COALESCE(can_fly, FALSE), whose inference needs its own justification.

Under the ordinary query WHERE NOT can_fly: {penguin}. The NULL comparisons remain UNKNOWN, and neither positive nor negative filters return those rows.

The twist: What if the NULL for ostrich came from a source that simply omits can_fly for ratites (absent-from-source), while the NULL for kiwi came from a source that includes the field but hasn't assessed the value (unknown)?

The single can_fly field does not express this. SQL can store a reason code, source identifier and policy version alongside it; the schema and queries must preserve and use those fields.

The bird table is the catalog's sustainable column writ small: same storage, same NULLs, different meaning depending on regime.

A uniform convention can make a single source easy to query. Once different conventions meet, an integration that omits them gives the receiver fewer grounds than either source possessed.


The Merge Problem

When Supplier A (CWA: "NULL means not sustainable") is merged with Supplier C (OWA: "NULL means unknown"), the system needs to preserve which convention governed each entry.

Supplier A's NULL items should be treated as not sustainable—under A's completeness profile, the absence of a positive claim is evidence of the negative. Supplier C's NULL items should be treated as sustainability unknown—under C's open-world reasoning, the absence of a claim is just absence.

If the system applies CWA globally, it misrepresents Supplier C's data. Items that Supplier C genuinely doesn't know about get labeled "not sustainable"—an unsupported negative classification, which may or may not be factually false.

If the system applies OWA globally, it misrepresents Supplier A's data. Items that Supplier A deliberately didn't mark sustainable get labeled "unknown"—obscuring a negative claim that the supplier intended to make.

Neither is correct. The correct answer is: it depends on the source.

But the merged table has lost the source. The NULL values are indistinguishable. The completeness profile was never stored. The merge performs semantic erasure: meaning is destroyed when provenance and completeness are dropped.

This is why Chapter 2's provenance typing matters. A witness should carry not just the source but the inference regime under which the claim was made. Without regime annotations, merging is a silent corruption of meaning.

Remark(On Vocabulary Absence vs Value Absence)

Supplier B presents a different problem: the sustainable field doesn't exist in B's feed at all. This is not a NULL; it's a schema mismatch—a vocabulary problem from Chapter 3.

When the system creates a merged table with a sustainable column, it must synthesize NULLs for Supplier B's items. But these synthetic NULLs mean something different again: "the source doesn't traffic in this concept." They are neither unknown (B didn't fail to learn the value) nor false (B didn't claim the items aren't sustainable). They are outside B's vocabulary.

Conflating predicate-absence with value-absence is the first compounding error. Treating all NULLs as equivalent, regardless of whether they came from an explicit NULL, a schema gap, or an inference, is the second.


Epistemic Status

The solution is to make the inference regime explicit. The status defined here is derivability relative to a view and its completeness profile. That status does not establish the factual reliability of the source’s premises.

A4
Epistemic Status (A4)

The epistemic status of proposition pp relative to view (U,RU)(U, R_U) is:

  • true: U,RU⊢pU, R_U \vdash p — the view, under its inference regime, proves pp
  • false: U,RU⊢¬pU, R_U \vdash \neg p — the view, under its inference regime, proves ¬p\neg p
  • undetermined: neither — the view, under its inference regime, is silent on pp

Assume (U,RU)(U, R_U) is consistent on pp: it does not derive both pp and ¬p\neg p. (A different treatment of conflict needs its own declared logic and representation; Part III develops exact gluing under its hypotheses.)

CWA and OWA are inference rules within the regime RUR_U, scoped to predicates:

  • CWA (predicate-scoped): For a predicate PP under a justified completeness rule in view UU, established non-derivability can license ¬P(x)\neg P(x). A timed-out or incomplete proof search does not establish non-derivability
  • OWA: Where neither P(x)P(x) nor its negation is derived, leave its status undetermined; incomplete search remains a separate operational status

The local theory T(U)=(Σ,I,RU)T(U) = (\Sigma, I, R_U) packages signature, constraints, and inference regime together. Two sources with the same signature but different completeness profiles are different theories.

The definition is predicate-scoped because completeness is rarely uniform. A supplier might claim completeness for price (every item has a price; absence would be an error) while acknowledging incompleteness for sustainable (some items haven't been assessed). The inference regime must specify which predicates are closed.

With A4 in hand, the merge problem becomes tractable. When Supplier A's data enters the system, it carries its completeness profile: sustainable is closed. When Supplier C's data enters, it carries a different profile: sustainable is open. The merged representation preserves both:

  • (pA,πA)(p_A, \pi_A): sustainable = NULL for item X, witnessed by Supplier A, regime: CWA for sustainable
  • (pC,πC)(p_C, \pi_C): sustainable = NULL for item Y, witnessed by Supplier C, regime: OWA for sustainable

A downstream query can now distinguish: A’s regime derives a negative for X; C’s regime leaves Y undetermined. Adopting A’s negative still requires the receiver to assess the completeness claim and its applicability. The system can partition results by epistemic status and let the user choose which partition to include.

Remark(On Logic Selection)

The choice between CWA and OWA is sometimes framed as a "logic choice"—different logics with different inference rules. This framing is correct but can be misleading. In knowledge representation, CWA is often modeled as a completeness axiom added to an otherwise open-world theory, or as a closure operation on the minimal model. The practical effect is the same: what you infer from absence depends on what you assume about completeness.

We use "inference regime" rather than "logic" to emphasize that the choice is about what the source claims to know, not about the fundamental rules of reasoning. A source can be complete for some predicates and incomplete for others. The regime is a profile, not a global setting.


SQL's Incomplete Remedy

SQL attempted to handle uncertainty with three-valued logic. Instead of TRUE and FALSE, SQL uses TRUE, FALSE, and UNKNOWN. NULL values propagate as UNKNOWN through most operations. Comparisons involving NULL yield UNKNOWN. Boolean operations follow Kleene's three-valued truth tables.

Three-valued evaluation can leave a comparison undecided instead of turning its missing operand into a negative answer. That behavior is useful even when the reason for missingness must be recorded separately.

But three values are not enough. SQL's UNKNOWN conflates "unknown whether true or false" with "not applicable" with "absent from source." The conflation produces counterintuitive behavior.

Gotcha 1: NOT IN with NULL

SELECT * FROM items WHERE category NOT IN (SELECT category FROM banned)

If the banned table contains any NULL value in the category column, this query can return zero rows—even for items whose category is clearly not among the non-null banned categories. The NULL comparison returns UNKNOWN; NOT IN requires all comparisons to be FALSE; a single UNKNOWN poisons the entire result.

This behavior surprises even experienced SQL developers. NOT EXISTS or a filter excluding NULLs expresses a different, often intended question. The choice must account for the treatment of missing values on both sides; changing the operator does not settle the source’s meaning for us.

Gotcha 2: NULL = NULL yields UNKNOWN

SELECT * FROM items WHERE color = color

This query returns only rows where color IS NOT NULL. The comparison NULL = NULL yields UNKNOWN, not TRUE. Every row with a NULL color fails the WHERE clause.

SQL’s = is not reflexive on NULL markers under this evaluation. A null-safe comparison such as IS NOT DISTINCT FROM asks a different, defined question about the stored values. Neither operation tells us why a field is missing.

No completeness declaration. SQL provides three-valued logic but leaves completeness assumptions to schema design and application convention. Different constructs—NOT IN, NOT EXISTS, outer joins, IS NULL—have specified behavior that a query author chooses. The query engine doesn't know whether a column is closed or open; it applies syntactic rules that may or may not match the intended semantics.

Per-source regimes need representation. A relational table can retain source identifiers, completeness profiles and versions, with queries joining the appropriate policy to each row. The constructed merge lost those fields. That is a defect of its representation and use, not a theorem about the limits of SQL.

Work on incomplete information has long distinguished different missingness semantics. The requirement here is to carry the source and inference regime through the transformation, using the database’s resources to preserve the difference rather than expecting UNKNOWN to encode it by itself.


The Fashion Catalog with Epistemic Status

Return to the catalog with the full machinery in hand.

For the sustainable attribute across three suppliers, the system maintains:

Supplier A (CWA for sustainable):

  • Items with sustainable = TRUE: reported sustainable
  • Items with sustainable = FALSE: reported not sustainable
  • Items with sustainable = NULL: not sustainable (closed-world inference)

Supplier B (no sustainable field):

  • All items: sustainable outside vocabulary—not unknown, not false, but undefined in this source

Supplier C (OWA for sustainable):

  • Items with sustainable = TRUE: reported sustainable
  • Items with sustainable = FALSE: reported not sustainable
  • Items with sustainable = NULL: sustainability unknown

When a user queries "Show me sustainable dresses," the system can now respond:

"Results are partitioned by epistemic status (counts illustrative):

Sustainable (reported): ~1,200 items—sustainable = TRUE from any source.

Not sustainable (reported): ~900 items—sustainable = FALSE from any source.

Not sustainable (CWA inference): ~3,400 items—sustainable = NULL from Supplier A, which claims completeness.

Sustainability unknown: ~8,900 items—sustainable = NULL from Supplier C, which does not claim completeness.

Sustainability not in vocabulary: ~12,800 items—from Supplier B, whose feed doesn't include this attribute.

Which partition(s) would you like to include?"

The distinction matters operationally: a system that surfaces epistemic commitments lets users choose, while a system that hides them chooses for them. A user who selects only positive reports gets that partition. Whether the reports establish sustainability under the user’s intended standard remains an evidentiary question. A user who is willing to include unknowns can make that choice explicitly. A user who wants to exclude Supplier A's CWA inferences (perhaps suspecting the supplier over-claims completeness) can do so.

The system has not hidden the complexity; it has structured it.


Touchstones Advanced

T5 (Negation/Absence): This touchstone asked whether a system can correctly handle negation—whether it can distinguish "known to be false" from "not known to be true."

Chapter 1 foreshadowed T5 as "negation without witnesses." Chapter 2 advanced it to "logic is part of provenance." Chapter 4 specifies this distinction within a view: epistemic status makes its inference regime explicit. CWA and OWA are declared inference rules scoped to predicates. A system that tracks completeness profiles can distinguish negative facts from mere absences.

The operational question remains: how do you actually track completeness profiles in deployed systems? That's Part V's engineering concern. Its adequacy depends on the represented source conventions and the checks the receiving use requires.

T8 (Uncertainty/Value): "Best restaurant in Berlin" is not a factual predicate. Its truth depends on preference context: best for whom? By what criteria? Under what constraints?

T8 requires the preference rule to be part of the question. A fixed rule can select one answer or tie several; changing that rule can change the result. The disagreement need not arise from missing evidence or an inconsistent theory.

A4 handles T8 partially. The machinery of views and regimes can express that "best restaurant" is undetermined in one view and TRUE for a specific restaurant in another (the view that encodes a particular preference function). Chapter 12 develops fibered predicates as one account of systematic context dependence. Explicit parameterized functions can also represent these choices; neither representation establishes that a selected preference rule is appropriate.

For now, T8 is advanced: the framework acknowledges context-dependence. Resolution is deferred.


Consequence

Chapter 3 distinguished changing a signature from establishing that its new predicates and constraints serve the intended use. Chapter 4 shows that even within a fixed vocabulary, the meaning of absence varies across sources. Two sources can share a signature and still disagree on what an empty cell means. The disagreement is not about values; it is about completeness—about what the sources claim to know.

A merge that carries these regimes lets the recipient distinguish the source’s negative conclusion from its silence. The recipient can then examine whether that conclusion is warranted for the intended use.

Generation, retrieval and structured storage can each carry more evidence than the failures examined here allowed. Their capacities do not remove the receiving institution’s task: determine which distinctions survived the transformation and which inference the available grounds support. The missing convention can be represented. Someone must ensure that it remains operative when the record is used.

Chapter 5 brings these requirements together at the point where local definitions are used in a shared result. It asks what the receiving operation must preserve when its sources can each support something useful but cannot simply be treated as one account.


Litmus Cases

CaseNameChapter 4 Status
T5Negation/AbsenceSpecified distinction: epistemic status explicit; CWA/OWA predicate-scoped
T8Uncertainty/ValueAdvanced: preference context represented; Chapter 12 develops a construction

T5 Progression:

  • Chapter 1: Foreshadowed as "negation without witnesses"
  • Chapter 2: Advanced to "logic is part of provenance; CWA/OWA as source metadata"
  • Chapter 4: Specified—epistemic status is formal; inference regime is part of local theory

T8 (Uncertainty/Value): "Best restaurant in Berlin" depends on preference context. A4 provides the framework for context-indexed truth; fibered predicates (Chapter 12) supply one construction for context dependence.

Search the book

Use ↑ ↓ to move through results; Escape to close.

Search every published chapter, section and reference.

    In this chapter