Epistemic Status
What a system knows about what it knows
Aa
What we cannot speak about we must pass over in silence.
This chapter formalizes epistemic status (A4): the truth value of a proposition relative to a view is true, false, or undetermined, and the distinction depends on which logic the view employs. One unqualified NULL field can conceal several reasons for missingness; a merge must preserve the distinctions its inferences require. The reader who wants the historical argument for why absence has always been contested territory should read Vol I, Chapter 5 (The Empire of Tables).
The Lie of NULL
In the hypothetical catalog, a merge combines feeds from three suppliers. Supplier A's feed includes a sustainable field: items are marked TRUE, FALSE, or NULL. Supplier C's feed has the same field, but with different conventions. For Supplier A, a NULL in sustainable means "this item is not sustainable": the supplier operates under a completeness declaration for this predicate, where absence of a positive claim implies the negative. For Supplier C, a NULL means "we don't know": the supplier acknowledges incomplete information.
Supplier B's feed has no sustainable field at all.
A user asks: "Show me sustainable dresses."
What should the system return?
If only the merged field survives, the table cannot recover the omitted source conventions. Supplier A's NULLs, Supplier C's NULLs, and the synthetic NULLs created for Supplier B's items all look identical. The completeness profiles that govern their interpretation—what absence licenses you to conclude—were never stored. The database sees the same symbol; the semantics are different.
The merge has made an inference harder to inspect without changing the printed symbol. Sometimes the predicate is absent from a source’s vocabulary; sometimes the predicate exists and its value is missing. The recipient needs that distinction before deciding what the blank permits it to conclude.
A missing value does not state the conditions under which absence becomes evidence.
Four Interpretations of Missingness
NULL is not one thing. It is a symbol that collapses at least four distinct situations into a single representation.
Three storage states describe what the data actually records:
| Storage State | Example | What It Means |
|---|---|---|
| Unknown | "We haven't tested whether this dress is machine-washable" | The predicate applies, but the value is not known; it might be TRUE or FALSE |
| Not applicable | "Machine-washable is not a meaningful attribute for jewelry" | The predicate does not apply to this entity; the question is ill-formed |
| Absent-from-source | "The supplier's feed doesn't include this field" | No information was transmitted; the absence is about the data source, not the entity |
Each storage state implies different behavior. Unknown values might become TRUE or FALSE with more information. Not-applicable is a typing failure: the predicate's domain does not include this entity (or includes it only via a partial function)—asking whether a necklace is machine-washable is a category error, and such values should not satisfy either positive or negative queries about the predicate unless the query explicitly ranges over applicability. Absent-from-source values should be handled according to the source's conventions, which may differ from source to source.
One derived state is not stored but inferred:
| Derived State | Inference | What It Means |
|---|---|---|
| Inferred false | "If sustainable isn't listed, it's not sustainable" | Under a completeness assumption, absence of a positive claim licenses the negative |
This fourth interpretation is different in kind. It is not a fact about what the data contains; it is a conclusion drawn from what the data omits, given an assumption about completeness. The inference requires a declared closure rule for the predicate in question. Using its conclusion as evidence about the world also requires adequate grounds for that completeness claim.
A NULL field alone does not distinguish these situations1. SQL can store an absence reason, source and policy in additional fields or related tables. If the representation omits them, sustainable IS NULL selects missing values without identifying why they are missing or what may be inferred from them.
The conflation propagates. A CASE statement that maps NULL to a default may implement one policy correctly and misrepresent another. COUNT(sustainable) counts non-NULL entries, including both TRUE and FALSE; it is not a count of sustainable items. Choosing a denominator or counting potential positives requires the question and absence convention to be represented. The query cannot recover a distinction the supplied records omit.
The missingness policy supplies an inference that the marker alone does not contain.
Completeness and Inference
The interpretation of absence depends on a completeness declaration—a claim about what the source takes itself to have covered.
The Closed-World Assumption (CWA) is the reasoning of completeness2. If is not derivable from the source, infer that is false. The source claims to know everything relevant; what it cannot derive, it denies.
A complete list of the entries in a directory can establish that a number is not listed there. It does not establish that a person has no phone. A catalog may declare a particular inventory predicate complete; that declaration does not close every attribute, and it needs adequate support for the use made of it.
The Open-World Assumption (OWA) leaves absence open. If neither nor its negation is derivable, the question remains undetermined. An explicit or derived negative can still establish a negative conclusion; an unreported positive alone cannot.
A biography’s omission of a birthdate need not assert that its subject had none. Nor does a study’s silence about a side effect, by itself, establish that the effect was absent. The evidentiary question is what the inquiry covered and what its reporting convention permits the reader to infer.
The inferred result depends on the policy; the unchanged SQL query does not. SELECT * FROM items WHERE NOT sustainable excludes NULL rows under ordinary SQL three-valued evaluation. To treat a missing value as false, the application must implement that policy—for example with an explicit conversion—and justify the completeness assumption it applies. Declaring CWA in prose does not change the query engine’s behavior.
Schema: birds(id, species, can_fly)
| id | species | can_fly |
|---|---|---|
| 1 | sparrow | TRUE |
| 2 | penguin | FALSE |
| 3 | ostrich | NULL |
| 4 | kiwi | NULL |
Query: "Which birds can fly?"
Both regimes agree: {sparrow}. The NULL values are not TRUE.
Query: "Which birds cannot fly?"
Under an explicit policy that converts these missing values to false: {penguin, ostrich, kiwi}. This is a changed evaluation, such as WHERE NOT COALESCE(can_fly, FALSE), whose inference needs its own justification.
Under the ordinary query WHERE NOT can_fly: {penguin}. The NULL comparisons remain UNKNOWN, and neither positive nor negative filters return those rows.
The twist: What if the NULL for ostrich came from a source that simply omits can_fly for ratites (absent-from-source), while the NULL for kiwi came from a source that includes the field but hasn't assessed the value (unknown)?
The single can_fly field does not express this. SQL can store a reason code, source identifier and policy version alongside it; the schema and queries must preserve and use those fields.
The bird table is the catalog's sustainable column writ small: same storage, same NULLs, different meaning depending on regime.
A uniform convention can make a single source easy to query. Once different conventions meet, an integration that omits them gives the receiver fewer grounds than either source possessed.
The Merge Problem
When Supplier A (CWA: "NULL means not sustainable") is merged with Supplier C (OWA: "NULL means unknown"), the system needs to preserve which convention governed each entry.
Supplier A's NULL items should be treated as not sustainable—under A's completeness profile, the absence of a positive claim is evidence of the negative. Supplier C's NULL items should be treated as sustainability unknown—under C's open-world reasoning, the absence of a claim is just absence.
If the system applies CWA globally, it misrepresents Supplier C's data. Items that Supplier C genuinely doesn't know about get labeled "not sustainable"—an unsupported negative classification, which may or may not be factually false.
If the system applies OWA globally, it misrepresents Supplier A's data. Items that Supplier A deliberately didn't mark sustainable get labeled "unknown"—obscuring a negative claim that the supplier intended to make.
Neither is correct. The correct answer is: it depends on the source.
But the merged table has lost the source. The NULL values are indistinguishable. The completeness profile was never stored. The merge performs semantic erasure: meaning is destroyed when provenance and completeness are dropped.
This is why Chapter 2's provenance typing matters. A witness should carry not just the source but the inference regime under which the claim was made. Without regime annotations, merging is a silent corruption of meaning.
Supplier B presents a different problem: the sustainable field doesn't exist in B's feed at all. This is not a NULL; it's a schema mismatch—a vocabulary problem from Chapter 3.
When the system creates a merged table with a sustainable column, it must synthesize NULLs for Supplier B's items. But these synthetic NULLs mean something different again: "the source doesn't traffic in this concept." They are neither unknown (B didn't fail to learn the value) nor false (B didn't claim the items aren't sustainable). They are outside B's vocabulary.
Conflating predicate-absence with value-absence is the first compounding error. Treating all NULLs as equivalent, regardless of whether they came from an explicit NULL, a schema gap, or an inference, is the second.
Epistemic Status
The solution is to make the inference regime explicit. The status defined here is derivability relative to a view and its completeness profile. That status does not establish the factual reliability of the source’s premises.
The epistemic status of proposition relative to view is:
- true: — the view, under its inference regime, proves
- false: — the view, under its inference regime, proves
- undetermined: neither — the view, under its inference regime, is silent on
Assume is consistent on : it does not derive both and . (A different treatment of conflict needs its own declared logic and representation; Part III develops exact gluing under its hypotheses.)
CWA and OWA are inference rules within the regime , scoped to predicates:
- CWA (predicate-scoped): For a predicate under a justified completeness rule in view , established non-derivability can license . A timed-out or incomplete proof search does not establish non-derivability
- OWA: Where neither nor its negation is derived, leave its status undetermined; incomplete search remains a separate operational status
The local theory packages signature, constraints, and inference regime together. Two sources with the same signature but different completeness profiles are different theories.
The definition is predicate-scoped because completeness is rarely uniform. A supplier might claim completeness for price (every item has a price; absence would be an error) while acknowledging incompleteness for sustainable (some items haven't been assessed). The inference regime must specify which predicates are closed.
With A4 in hand, the merge problem becomes tractable. When Supplier A's data enters the system, it carries its completeness profile: sustainable is closed. When Supplier C's data enters, it carries a different profile: sustainable is open. The merged representation preserves both:
- :
sustainable = NULLfor item X, witnessed by Supplier A, regime: CWA forsustainable - :
sustainable = NULLfor item Y, witnessed by Supplier C, regime: OWA forsustainable
A downstream query can now distinguish: A’s regime derives a negative for X; C’s regime leaves Y undetermined. Adopting A’s negative still requires the receiver to assess the completeness claim and its applicability. The system can partition results by epistemic status and let the user choose which partition to include.
The choice between CWA and OWA is sometimes framed as a "logic choice"—different logics with different inference rules. This framing is correct but can be misleading. In knowledge representation, CWA is often modeled as a completeness axiom added to an otherwise open-world theory, or as a closure operation on the minimal model. The practical effect is the same: what you infer from absence depends on what you assume about completeness.
We use "inference regime" rather than "logic" to emphasize that the choice is about what the source claims to know, not about the fundamental rules of reasoning. A source can be complete for some predicates and incomplete for others. The regime is a profile, not a global setting.
SQL's Incomplete Remedy
SQL attempted to handle uncertainty with three-valued logic. Instead of TRUE and FALSE, SQL uses TRUE, FALSE, and UNKNOWN. NULL values propagate as UNKNOWN through most operations. Comparisons involving NULL yield UNKNOWN. Boolean operations follow Kleene's three-valued truth tables3.
Three-valued evaluation can leave a comparison undecided instead of turning its missing operand into a negative answer. That behavior is useful even when the reason for missingness must be recorded separately.
But three values are not enough. SQL's UNKNOWN conflates "unknown whether true or false" with "not applicable" with "absent from source." The conflation produces counterintuitive behavior.
Gotcha 1: NOT IN with NULL
SELECT * FROM items WHERE category NOT IN (SELECT category FROM banned)
If the banned table contains any NULL value in the category column, this query can return zero rows—even for items whose category is clearly not among the non-null banned categories. The NULL comparison returns UNKNOWN; NOT IN requires all comparisons to be FALSE; a single UNKNOWN poisons the entire result.
This behavior surprises even experienced SQL developers. NOT EXISTS or a filter excluding NULLs expresses a different, often intended question. The choice must account for the treatment of missing values on both sides; changing the operator does not settle the source’s meaning for us.
Gotcha 2: NULL = NULL yields UNKNOWN
SELECT * FROM items WHERE color = color
This query returns only rows where color IS NOT NULL. The comparison NULL = NULL yields UNKNOWN, not TRUE. Every row with a NULL color fails the WHERE clause.
SQL’s = is not reflexive on NULL markers under this evaluation. A null-safe comparison such as IS NOT DISTINCT FROM asks a different, defined question about the stored values. Neither operation tells us why a field is missing.
No completeness declaration. SQL provides three-valued logic but leaves completeness assumptions to schema design and application convention. Different constructs—NOT IN, NOT EXISTS, outer joins, IS NULL—have specified behavior that a query author chooses. The query engine doesn't know whether a column is closed or open; it applies syntactic rules that may or may not match the intended semantics.
Per-source regimes need representation. A relational table can retain source identifiers, completeness profiles and versions, with queries joining the appropriate policy to each row. The constructed merge lost those fields. That is a defect of its representation and use, not a theorem about the limits of SQL.
Work on incomplete information has long distinguished different missingness semantics4. The requirement here is to carry the source and inference regime through the transformation, using the database’s resources to preserve the difference rather than expecting UNKNOWN to encode it by itself.
The Fashion Catalog with Epistemic Status
Return to the catalog with the full machinery in hand.
For the sustainable attribute across three suppliers, the system maintains:
Supplier A (CWA for sustainable):
- Items with
sustainable = TRUE: reported sustainable - Items with
sustainable = FALSE: reported not sustainable - Items with
sustainable = NULL: not sustainable (closed-world inference)
Supplier B (no sustainable field):
- All items:
sustainableoutside vocabulary—not unknown, not false, but undefined in this source
Supplier C (OWA for sustainable):
- Items with
sustainable = TRUE: reported sustainable - Items with
sustainable = FALSE: reported not sustainable - Items with
sustainable = NULL: sustainability unknown
When a user queries "Show me sustainable dresses," the system can now respond:
"Results are partitioned by epistemic status (counts illustrative):
Sustainable (reported): ~1,200 items—
sustainable = TRUEfrom any source.Not sustainable (reported): ~900 items—
sustainable = FALSEfrom any source.Not sustainable (CWA inference): ~3,400 items—
sustainable = NULLfrom Supplier A, which claims completeness.Sustainability unknown: ~8,900 items—
sustainable = NULLfrom Supplier C, which does not claim completeness.Sustainability not in vocabulary: ~12,800 items—from Supplier B, whose feed doesn't include this attribute.
Which partition(s) would you like to include?"
The distinction matters operationally: a system that surfaces epistemic commitments lets users choose, while a system that hides them chooses for them. A user who selects only positive reports gets that partition. Whether the reports establish sustainability under the user’s intended standard remains an evidentiary question. A user who is willing to include unknowns can make that choice explicitly. A user who wants to exclude Supplier A's CWA inferences (perhaps suspecting the supplier over-claims completeness) can do so.
The system has not hidden the complexity; it has structured it.
Touchstones Advanced
T5 (Negation/Absence): This touchstone asked whether a system can correctly handle negation—whether it can distinguish "known to be false" from "not known to be true."
Chapter 1 foreshadowed T5 as "negation without witnesses." Chapter 2 advanced it to "logic is part of provenance." Chapter 4 specifies this distinction within a view: epistemic status makes its inference regime explicit. CWA and OWA are declared inference rules scoped to predicates. A system that tracks completeness profiles can distinguish negative facts from mere absences.
The operational question remains: how do you actually track completeness profiles in deployed systems? That's Part V's engineering concern. Its adequacy depends on the represented source conventions and the checks the receiving use requires.
T8 (Uncertainty/Value): "Best restaurant in Berlin" is not a factual predicate. Its truth depends on preference context: best for whom? By what criteria? Under what constraints?
T8 requires the preference rule to be part of the question. A fixed rule can select one answer or tie several; changing that rule can change the result. The disagreement need not arise from missing evidence or an inconsistent theory.
A4 handles T8 partially. The machinery of views and regimes can express that "best restaurant" is undetermined in one view and TRUE for a specific restaurant in another (the view that encodes a particular preference function). Chapter 12 develops fibered predicates as one account of systematic context dependence. Explicit parameterized functions can also represent these choices; neither representation establishes that a selected preference rule is appropriate.
For now, T8 is advanced: the framework acknowledges context-dependence. Resolution is deferred.
Consequence
Chapter 3 distinguished changing a signature from establishing that its new predicates and constraints serve the intended use. Chapter 4 shows that even within a fixed vocabulary, the meaning of absence varies across sources. Two sources can share a signature and still disagree on what an empty cell means. The disagreement is not about values; it is about completeness—about what the sources claim to know.
A merge that carries these regimes lets the recipient distinguish the source’s negative conclusion from its silence. The recipient can then examine whether that conclusion is warranted for the intended use.
Generation, retrieval and structured storage can each carry more evidence than the failures examined here allowed. Their capacities do not remove the receiving institution’s task: determine which distinctions survived the transformation and which inference the available grounds support. The missing convention can be represented. Someone must ensure that it remains operative when the record is used.
Chapter 5 brings these requirements together at the point where local definitions are used in a shared result. It asks what the receiving operation must preserve when its sources can each support something useful but cannot simply be treated as one account.
Litmus Cases
| Case | Name | Chapter 4 Status |
|---|---|---|
| T5 | Negation/Absence | Specified distinction: epistemic status explicit; CWA/OWA predicate-scoped |
| T8 | Uncertainty/Value | Advanced: preference context represented; Chapter 12 develops a construction |
T5 Progression:
- Chapter 1: Foreshadowed as "negation without witnesses"
- Chapter 2: Advanced to "logic is part of provenance; CWA/OWA as source metadata"
- Chapter 4: Specified—epistemic status is formal; inference regime is part of local theory
T8 (Uncertainty/Value): "Best restaurant in Berlin" depends on preference context. A4 provides the framework for context-indexed truth; fibered predicates (Chapter 12) supply one construction for context dependence.