Chapter 5
The Empire of Tables
Aa
Future users of large data banks must be protected from having to know how the data is organized in the machine (the internal representation).
— E. F. Codd, “A Relational Model of Data for Large Shared Data Banks” (1970)
On the American census form in 2000, the instruction for the race question was to mark one or more races. A household respondent could mark several boxes for the same person without having failed to follow the form. The census had changed what it was prepared to hear.1
The alteration was small enough to fit above the boxes. Its administrative life was considerably larger. In 1997 the Office of Management and Budget had revised the federal standards, allowing multiple racial designations and separating the former Asian or Pacific Islander category into two. The decision followed years of consultation and research. Its stated concerns included respect for individual dignity, the usefulness of the data and continuity with earlier statistics. The categories themselves were described as social and political constructions, not biological divisions. There was no official claim that the old form had discovered the natural number of kinds of people.2
Nor was every advocate of continuity defending the old form for its own sake. Comparable records helped make discrimination visible. An institution accused of treating a population differently could be examined across places and years because the population had a sufficiently stable description in the records. OMB’s account located the earlier standards partly in the work of enforcing civil-rights laws. This is one of the difficulties of a successful classification: people acquire reasons to depend upon it, including reasons for challenging the institutions that use it. The accumulated record can protect them. Its vocabulary can also cease to describe them adequately.
The revised standard did not resolve that conflict by putting everyone who selected several races into a single new box. It allowed combinations. The census form retained more detailed choices, and an opportunity to write in “Some other race,” which were then organized into broader groups for tabulation. It remained a form, with choices and instructions, but it admitted answers that the preceding decennial census had not accepted in the same way. A person could now appear in the statistics through a combination that had previously been unavailable.1
There was an awkward consequence. A fuller answer on one form did not supply a fuller answer on every other form already in use. Birth and death records, collected through state systems, continued to use older categories during the transition. For health statisticians, the new population counts therefore arrived with a problem in the denominator.3
It is easy to pass over that word. A death rate requires both the deaths counted and the population to which they are related. Changing the description of the population beneath the line can change the rate even when nothing has changed in the deaths above it. The resulting difference might then be read as an improvement or deterioration in health. Before asking what had happened to the people, the statistician had to ask what had happened to the comparison.
The previous counts could not simply be kept in service indefinitely. Estimates based on the 1990 census had been extended to provide denominators for 2000 and 2001, but they needed the correction supplied by a new enumeration. Keeping the older vocabulary by keeping an increasingly inadequate estimate would have preserved one continuity at the expense of another. The population had changed, the permissible answers had changed, and the inherited series still had work to do. None of these facts was an error one could remove from a row.3
Statisticians and demographers from the National Center for Health Statistics (NCHS), the Census Bureau and OMB considered excluding the multiple-race population from the single-race rate calculations, while retaining those people in overall statistics. That would avoid assigning them to a single category. It could also bias the rates for the remaining categories, with different effects according to the size of the groups and their underlying rates. Leaving the new answers alone was not equivalent to leaving the old calculation undisturbed. The calculation had always depended on how people entered it.
The alternative was to make the classifications comparable through an adjustment. Either the single-race counts of births and deaths could be expanded toward the newer categories, or the multiple-race population estimates could be brought back toward the older ones. The researchers chose the latter. It required modeling a smaller share of the population and offered fewer possible destinations. Someone counted as Black and White presented two single-race possibilities; someone previously recorded as White might have selected any of several combinations had the newer question been asked. The direction of translation was selected partly because it demanded less invention.3
I find the decision more instructive than a story in which a rigid database rejects a complicated person. The statisticians were trying to keep a useful inquiry possible. They could not retrospectively ask everyone the same question, and they could not make the answers comparable by preserving their spelling. Their method had to introduce something that the newer census responses did not contain: an estimate of how those responses would have been distributed under the older convention.
“Best represents”
For this they had another source. The National Health Interview Survey had long allowed respondents to indicate more than one race and then asked for a primary race. In the surveys used for the bridge, the follow-up was: “Which of these groups would you say best represents your race?” The researchers pooled four years of data, from 1997 through 2000, and modeled the answers using demographic information and the context of the respondent’s county. Those relationships supplied allocation proportions for the population counts.4
That further question had been asked of the survey respondents, not newly put to every person counted by the census. The report expressly assumed that its answer could serve as a guide to the older single-race question. Of the 4,898 multiple-race respondents in the relevant survey groups, 3,956 supplied a primary race and entered the analysis. The others did not provide one. We do not know their reasons. The model learned from those who answered and extended those relationships beyond them; the missing answers leave us unable to treat that extension as something everyone had confirmed. They do not, by themselves, tell us how well or badly it worked.
Even among those who supplied the answer, there was not enough information to fit every group separately. Smaller groups required a composite model. The report explained the approximation and the scarcity of observations behind it; the available information, as well as the desired classification, helped determine the form of the method. A population can be large enough to name and too sparsely represented in a survey to estimate on its own. Adding a category is not the same achievement as acquiring the evidence needed to say much about it.4
The adjustment also began further along the documentary chain than the census form suggests. The population estimates program had prepared a Modified Race Data Summary File. In that derivative, a response combining “Some other race” with a specified category lost the former component; a response consisting only of “Some other race” was assigned a category through imputation under stated rules. For the subsequent bridge, Asian and Native Hawaiian or Other Pacific Islander were recombined to match the older standard. A new estimate could thus contain a distinction recently introduced, omit it for a particular calculation, and remain part of the same statistical undertaking.5
These operations did not amount to a single act of correcting what people had said. The modification aligned one set of categories with another; the bridge then estimated a distribution needed for older classifications. Nor does the derivative establish that the original answers were erased. The Census Bureau published its account of the multiple-race population, and the NCHS report described the changes through which its different population estimates were obtained. The forms of the answer coexisted.
There is a further wrinkle in the Census Bureau’s own account. Its first footnote explains that “reported” includes responses assigned during editing and imputation, as well as answers supplied by respondents. Even before the bridge, the published table was a processed account of the returns. This does not make a census a fiction. It makes its documentation part of knowing what a census count is. A reader looking for untouched testimony cannot find it merely by choosing the earlier table.1
In the NCHS report’s illustrative calculation, probabilities of 0.6 and 0.4 would distribute a Black-and-White population group sixty percent to Black and forty percent to White. These were example proportions, not a reported national result. The operation allocated population counts. It did not discover which particular people would have given which answer. An estimate could be adequate for calculating a rate and contain no warrant for attributing a single-race response to a named person.3
That difference is easy to preserve in a methodological report and easy to lose in a succession of displays. A later table may have an ordinary column heading, an ordinary integer and no room for the sequence we have just followed. Its simplicity can be useful. A health researcher need not repeat the bridging analysis before every calculation. What must remain available is enough of the method to know which calculation the number can enter, what it estimates and what it has never established. The point of the bridge is to spare the next researcher some work without silently changing the question that work answered.
What the user need not know
This is also an ambition of database design. Codd’s 1970 paper identified dependencies on ordering, indexing and access paths that tied programs to particular representations of their data. It acknowledged earlier progress toward independence and proposed a relational model that could extend it. The implementation was left for other work. His opening promise was that users should be protected from needing to know the machine’s internal arrangement; it was not a promise that they could remain ignorant of what their information meant.6
That division of knowledge is a considerable achievement. The person asking about a set of records can attend to the relationships that matter to the inquiry while the system attends to how those records are found. Shared constraints can reject some errors before they spread through separate applications. A declared unique key can prevent duplicate identifiers; an enforced foreign-key constraint can prevent a reference to an absent row. These are dependable operations within their declared conditions. The error would be to require them to establish that two rows describe the same person, or that a recorded relationship is one the world actually contains.7
We have no reason to resent the ignorance a good abstraction permits. It is part of what makes a complicated undertaking available to someone who did not build it. The useful question is which knowledge the user has been relieved of needing. Rearranging storage while preserving a relation is one matter. Reclassifying responses to estimate a different population is another. The latter may be entirely reasonable, but knowing the resulting field’s name does not tell us which of those operations has occurred. A system can conceal the first as a service to its user. Concealing the second can change what the user believes the service has done.
The pomegranate and the motif
The Metropolitan Museum holds a velvet chasuble, dated to the fifteenth or sixteenth century, whose catalogue title includes a pomegranate design. Under classification, the catalogue places it among ecclesiastical costumes. The pomegranate belongs to the description of an ornament; velvet describes its material; the chasuble names a garment with a particular use. A competent catalogue has no difficulty allowing these descriptions to meet in one object.8
It would be an impoverished account that made the pomegranate defeat the table. A fruit, a motif and a color name need not be confused merely because a word moves among them. Nor must a catalogue choose one sense for every appearance of the word. A collection could include both the garment and a painting of fruit, and distinguish them while making their relation searchable. The interesting question is what someone hopes to find through that relation. A search for a textile pattern and an inquiry into religious ornament may begin with the same object and need quite different neighbors.
The catalogue’s restraint matters as much as its detail. It does not have to decide everything the object could mean before allowing it to be found. The object has survived into purposes that its makers could not have specified for a museum database. We need not supply their intentions to notice how many inquiries can now take it up. A classification can make an object available without claiming to have exhausted it; a second description can enlarge that availability without making the first false. Not every difference needs reconciliation into one preferred account.
This is familiar work in the study and design of classifications. Bowker and Star treated their institutional histories and consequences as central to Sorting Things Out; their 1999 introduction already discussed the federal multiple-race revision. Their attention to the inherited infrastructure helps explain why a revised form encounters more than a vocabulary problem. Much of its past is still being used by other people.9
An engineer can give that coexistence an explicit representation. Separate vocabularies can retain their own identifiers and be related by mappings. They can be extended, versioned and used for different purposes. The W3C’s SKOS vocabulary, for example, distinguishes a close match from an exact match, giving a recipient a way to preserve the degree of agreement that has actually been declared.10 Such distinctions are real technical resources. They need no impossible master vocabulary in which every future use of an object has already been anticipated.
A demographic bridge makes a more particular undertaking. It does not declare two descriptions interchangeable for all purposes. It uses information about the relation between them to estimate counts for a specified comparison. Keeping the newer categories alongside that estimate is not a failure to finish the migration. It preserves a question the older series cannot answer. There is no general reason that the calculation most convenient for an inherited institution should become the only account the institution permits to survive.
What NULL refuses to say
Missing information brings the distinction down to a single field. Suppose a small bird database records whether a species can fly. The sparrow is TRUE, the penguin FALSE, and the ostrich and kiwi have NULL in that position. A query for can_fly = FALSE returns the penguin. It does not return the other two flightless birds, because their entries do not say FALSE. Comparing NULL with FALSE yields an unknown result, which does not satisfy the query’s requirement for a true condition.11
Our knowledge of the birds makes the result seem perversely incomplete. The engine has not been asked to apply that knowledge. It has been asked to select rows under a specified condition, and has declined to make an unrecorded assertion on the programmer’s behalf. A separate query can retrieve the NULL entries. Someone can investigate or impute the missing values, but that is further work, with grounds of its own. The missing value does not tell us whether the observer was uncertain, the field was never collected, or information available elsewhere failed to reach this table.
All of these distinctions can be represented. A design can retain a reason for missingness, a source, an assessment status or the rule under which absence permits an inference. There is no incapacity in the relational model that forces unknown, inapplicable and never received to share an institutional meaning. What matters is whether the chosen representation preserves the differences the inquiry will need. Adding them later may require a return to sources that are no longer within reach. Adding a field does not supply the history to put in it.
No record for that combination
In the NCHS release, some omissions are part of the specification. Appendix I describes the file containing its bridged estimates for April 1, 2000, broken down by county, race category, sex, Hispanic origin and age. When the population count for one of those combinations is zero, the file contains no record for it.12
Within that documented file, the omission has a meaning: zero in the estimate. It is not an invitation to speculate about why a respondent left a box empty. It is not a suppressed answer, a missing transfer or an unresolved assessment. The producing institution has specified a convention that a recipient can apply without inventing one. A complete copy of the release and its documentation permits a definite operation precisely where an isolated absence would not.
That certainty has a boundary worth keeping intact. The zero belongs to a population estimate under the stated classifications and method. It does not prove that no person in the county would have described themselves with a particular combination on another form. The method may have grouped that combination differently, and the estimate is not a collection of newly obtained individual answers. Reading the omission correctly requires both pieces of knowledge: what the file’s convention establishes, and what sort of result the file contains.
The public report is part of what makes the file usable. It identifies the source populations, explains transformations, supplies model details and describes the fields through which the result travels. Its reader can distinguish an estimated count from an enumerated category and an absent row from an absent explanation. None of that requires opening identifiable household records to every downstream user. The relevant account can travel at the level of the statistic and the operation.
There is a political choice in making that account available. If a classification can be understood only by asking its present custodian what the numbers mean, those who inherit its results remain dependent on the custodian’s willingness to explain them. Documentation can reduce that dependence. It can also leave a later investigator able to ask a question the original producer did not favor. Preserving the newer population descriptions alongside the older comparison matters for the same reason: the convenience of one inquiry should not quietly extinguish the material of another.
The NCHS bridge was used to calculate birth and death rates and revise previously published rates. It made a comparison possible without requiring the country to go back and fill in the older form. The saving is available to later researchers using the estimates for comparisons they support. If an office instead treats an estimated allocation as evidence of what particular people said, it owes an inquiry that the allocation has not performed.
The table’s precision is valuable because the questions differ. How many people can be counted under this convention, what they reported, what a model inferred and what an institution may do next are not interchangeable merely because each can be given a field. A well-made representation lets one result support the work for which it is fit and leaves the other questions open to their own evidence. The burden of supporting that attribution belongs to the institution making it. It cannot make the people in its records answer a question they were never asked by renaming the column.
Source notes
Footnotes
-
Nicholas A. Jones and Amy Symens Smith, The Two or More Races Population: 2000, Census 2000 Brief C2KBR/01-6 (November 2001), pp. 1–2, especially figure 1 and p. 1 n. 1. The reproduced question concerns the person being reported on; household respondents could answer for others. The report expressly includes editing and imputation in “reported.” Fifteen response choices and write-ins are distinguished from the six broad tabulation categories. The figure, wording and footnote were inspected together. ↩ ↩2 ↩3
-
Office of Management and Budget, “Revisions to the Standards for the Classification of Federal Data on Race and Ethnicity”, Federal Register 62, no. 210 (October 30, 1997), pp. 58782–90; background and governing principles at pp. 58782–83, multiple-selection decision at p. 58786, tabulation and confidentiality at pp. 58788–89. This is an account of the 1997 revision and its historical application, not current federal standards. The proposed consequence concerning continuity and dependence is the chapter’s interpretation. ↩
-
Deborah D. Ingram et al., United States Census 2000 Population With Bridged Race Categories, Vital and Health Statistics 2, no. 135 (September 2003), pp. 1–4. Pages 3–4 distinguish the excluded-from-race-specific-rates option from bridging, explain the chosen direction and identify actual rate calculations and revisions. The 60/40 allocation is the report’s hypothetical example on p. 4, not an empirical proportion asserted by this chapter. The method is a population-estimation procedure; no inference about an identified individual is drawn. CDC archival record. ↩ ↩2 ↩3 ↩4
-
Ingram et al., pp. 6–8 and tables 5–8, pp. 18–19. The primary-race question, the assumption relating its response to the earlier standard and the 4,898/3,956 counts appear on p. 7. Nonresponse differed among groups. The composite model assumes relationships for small groups can be approximated using those found in larger groups. The differing nonresponse rates do not establish either the reasons for missing answers or the magnitude of selection bias. Models incorporated the survey’s design and weights; details and covariates are in the report. ↩ ↩2
-
Ingram et al., pp. 4–7. The Modified Race Data Summary File’s removal or imputation of “Some other race” is a distinct preparation step; the subsequent merging of Asian and Native Hawaiian/Other Pacific Islander categories reflects the older target standard and limited survey observations. The transformations concern specified derivatives and do not establish destruction of the original census returns or published multiple-race tabulations. The chapter does not compare totals from incompatible processing stages as though they had the same definition. ↩
-
E. F. Codd, “A Relational Model of Data for Large Shared Data Banks”, Communications of the ACM 13, no. 6 (June 1970), pp. 377–87, §§1.1–1.2, 2.4. The epigraph retains the original parenthetical. Later SQL and optimizer implementations are not attributed to this paper. ↩
-
The examples concern declared and enforced constraints, not a guarantee that every database installation supplies them. See SQLite, CREATE TABLE, “UNIQUE constraints,” and Foreign Key Support, §§1–2. Foreign-key enforcement must be enabled in the relevant SQLite configuration. A unique row identifier is not proof of a person’s identity; a reference to an existing row is not proof that the represented event occurred. ↩
-
Metropolitan Museum of Art, Chasuble with Pomegranate Design, object 16.32.320; British, fifteenth–sixteenth century, velvet; catalogue classification “Textiles-Costumes-Ecclesiastical.” The discussion uses the documented object and its descriptions, not an invented encounter with it or an attribution of the makers’ intentions. ↩
-
Geoffrey C. Bowker and Susan Leigh Star, Sorting Things Out: Classification and Its Consequences (MIT Press, 1999), introduction pp. 3–5, especially the 1997 revision on p. 4; inspected introduction and author-hosted text, opening of chapter 1. The later NCHS application extends an established inquiry. ↩
-
Alistair Miles and Sean Bechhofer, eds., SKOS Simple Knowledge Organization System Reference, W3C Recommendation (August 18, 2009), §10, especially §10.6.3. A chain of close matches does not entail another close match. The example concerns an existing vocabulary’s declared semantics; it is not evidence that NCHS used SKOS or that vocabulary mappings substitute for statistical models. ↩
-
SQLite, SQL Language Expressions, §§2, 8 and 14. The explicit illustrative bird rows and
NOT INqueries were executed on SQLite 3.51.0. Testing forIS NULL, excluding NULLs from the subquery and usingNOT EXISTShave distinct effects recorded in the accompanying research file. They do not determine what a missing source value means. No claim about a NULL-related defect in the Census or NCHS implementation is made. The additionalNOT INexample is retained here: if the comparison list includes NULL, an otherwise unmatched value can yield UNKNOWN rather than TRUE for the negated condition, leaving no selected rows in the illustrated query. Filtering NULLs or usingNOT EXISTScan change that result; either operation must fit the intended question, rather than merely produce a convenient list. ↩ -
Ingram et al., Appendix I, p. 52, file layout for
br040100.txt, released January 7, 2003. The documented omission rule applies to zero counts for the specified county/race/sex/Hispanic-origin/age combinations in this complete release. The zero is an estimated population count, not proof about every person’s unobserved answer. This chapter examines the documented convention; it does not claim to have audited the population file row by row. The preceding account of NCHS documentation follows pp. 4–9 and Appendix I; the consequences for later institutional use are the author’s argument. ↩