Res Agentica
Reading

No saved reading position.

Reading

No saved reading position.

Chapter 10

Time Crystallized in Weights

17 min read
Aa
Text size

We can think of capital, indeed, as frozen knowledge or knowledge imposed on the material world in the form of improbable arrangements.

— Kenneth Boulding, 'The Economics of the Coming Spaceship Earth' (1966)

At 10:35 on the evening of 3 December 2021, the logbook for Meta's OPT language model acquired an instruction in red capitals: nobody was to reboot or replace the machines until further notice. Recent checkpoints were sitting on their local disks. The process meant to copy them to cloud storage had stopped working, and replacing a machine could destroy the state from which training was supposed to resume. One checkpoint had been copied manually to another location. A new storage destination was proposed; the following morning's entry records the upload of the scattered files and the replacement of a faulty node.1

The distinction between a calculation having happened and its result being available was, for the moment, a matter of where the files were. A large computation had been performed. That did not make its continuation secure. The work depended on a rather ordinary accommodation between people maintaining machines and people who needed those machines to retain something. Replacing defective equipment, usually a contribution to progress, had become a way of losing it.

OPT would eventually be released as a family of trained language models. A recipient could obtain parameters produced by an undertaking they had never joined, load them with the appropriate software and use the resulting model. They would still need equipment, but they would not need to recreate that December evening. The value of the copy begins in this dispensation: a result can be made available without requiring another person to undergo its formation.

There is more than one result to preserve. A weight file contains the learned numerical parameters used in the model's computation. A checkpoint for resuming training can also contain the optimizer's state, settings and information about the run's progress. The distinction appears in the OPT record itself. On 19 December, the team made a version of a checkpoint without the optimizer state and placed it in a location for evaluation. What had been needed to continue the learning procedure was not all needed to examine the model it had so far produced.2

This separation is convenient until someone mistakes one kind of continuity for another. Earlier that month the team had tried changing the rule by which the parameters were updated. On 2 December their debrief identified a problem: restoring the saved optimizer state had also restored settings they had meant to change. Comparisons based on the supposed change could no longer be trusted. The logbook marks them as invalid. The attempt using a different optimizer made little progress and had its own implementation problem; the next morning's entry returns to AdamW, the earlier optimizer, with a lower learning rate.3

The record is unusually helpful because it has not been tidied into a succession of successful decisions. The saved state was doing something it had been designed to do—carrying an earlier arrangement forward—while frustrating an attempt to alter that arrangement. The team needed to know more than whether the file could be loaded. It needed to know which experiment had actually taken place. Nothing in the final model's usefulness would, by itself, disclose this mistake to someone studying the abandoned comparison.

There had already been a larger change of course. In November, with the time available for training running short, the team abandoned a troublesome series of runs and adopted more of the settings reported for other large models. Their update of 17 November explains the decision as a choice under a deadline. Continuing the earlier experiments might have worked, but they did not know whether it would work soon enough. Even the account of the more promising configuration ends with expectations about difficulties still to come.4

That is the history which an image of crystallization can too easily polish away. The weights emerged from a course of work, but the course was neither a single uninterrupted calculation nor a demonstrated minimum. It included efforts whose results were discarded, interventions that changed what continued, and decisions made without knowing the ending. The completed artifact preserves consequences of that history. The logbook preserves reasons to distrust a few of its apparent lessons. Both are worth receiving.

The logbook goes with it

The completion notice was posted on 7 January 2022. Training had finished the previous day. Its authors translated the reported computation into roughly thirty-three days on a large, continuously operating cluster, explicitly assuming away hardware failures and numerical instability. The same notice describes an accidental deletion of the cluster in December and the subsequent development of automatic recovery. Between Christmas and New Year, the team reported eight hardware failures from which training had recovered automatically.5

The uninterrupted run belonged to the calculation of resources. The interrupted one produced the model. By the end, some of the work of recovering from failure had itself become something the machinery could do. This was another useful result of development, though it was not a language capability encoded in the trained parameters. It belonged to the means of keeping the undertaking in operation.

In May, Meta made weights available at several smaller scales and offered researchers access to the largest model upon application. The release also supplied code, model and data documentation, and the development chronicles. Its permissions had limits: the weight license authorized noncommercial research and excluded commercial or production use. The code's licensing was different. Receiving the files was therefore a substantial enlargement of what a researcher could do, without being a grant to use them for any purpose whatever.6

The openness of the history matters as much to this inquiry as the existence of the weights. There is no need to imagine the original laboratory as a treasury of secrets in order to acknowledge its work. Here it published an account of things going wrong. A later researcher could inspect the warning on the invalid optimizer comparisons without conducting those comparisons again. The account does not do the reader's work, and it does not establish that the same remedy will apply to a different run. It gives the reader a more informed place to begin. Some dependence on the people who originally discovered a difficulty can be reduced by their willingness to explain it.

We should allow this generosity its full consequence. An inheritance is not defective because the heir did not acquire it in the manner of its maker. Someone using a trained model need not first become capable of designing and conducting its entire training process. If that were required, much of the point of releasing it would disappear. A working capability can become an instrument of research for people whose question was absent from the project that produced it. They may be better placed to ask that question precisely because their undertaking is different.

Yet a model is not a complete workshop packed into a file. The parameters require an architecture and code that interpret them, the appropriate treatment of inputs, and an environment capable of executing the computation. Testing a result requires further choices about what counts as success. Repeating the original training requires the relevant data and procedure; describing a corpus does not place the corpus, in its prepared form, on another person's machines. A recipient may supply these things through another organization, simplify the task until fewer are needed, or develop a use for which reproducing the original training would be wasteful. None of these responses is adequately described as failing to inherit the whole.

Nelson and Winter placed organizational knowledge partly in routines: people knowing how to respond to one another, supported by records and equipment whose arrangement helped the work continue. A blueprint was not the functioning organization. Their account of replication makes room for the time and difficulty of establishing that organization elsewhere.7 The OPT files do not repeal this distinction. They allow a particular result to cross it. The recipient need not reconstruct the organization before making any use of what the organization learned to produce.

Expertise in training a large model and expertise in testing one use of a released model are different requirements. A transfer that removes the first can enlarge participation even when the second remains demanding. To answer merely that expertise remains necessary is to stop at a truth too general to locate the change. The released model brings the recipient past some work while leaving other work ahead, and someone whose earlier project was impossible may now have a difficulty they are equipped to solve.

Learning from ten

Nor must the capability always travel in the same parameters. In 2015 Geoffrey Hinton, Oriol Vinyals and Jeff Dean described transferring useful behavior from a cumbersome model into another, more convenient one. They acknowledged an earlier compression method developed by Cristian Bucilă, Rich Caruana and Alexandru Niculescu-Mizil. The idea was to train a new model from the older system's responses, rather than reproduce the whole process that had produced the older system.8

Their speech-recognition experiment used ten acoustic models built on an earlier version of the architecture used by Android voice search. A single distilled model retained most of the ensemble's improvement in classifying acoustic frames and matched its reported word-error rate. It still required training; it did not become the ten models in every respect. But the tested improvement could be obtained without running all ten whenever a prediction was needed.9

That result changes the meaning of the original expense. The ensemble's work remained part of the explanation of the student's performance. It no longer specified the equipment that had to be kept in use to obtain that performance. Production had yielded something from which a second, differently organized production could begin. The recipient of such a result is doing real work too: choosing what to transfer, constructing a learner and finding out what survived. A shorter route opened by someone else's achievement is still a route.

There are losses and costs along it. Generating the teaching responses consumes computation; the new model has to be trained and tested; and the behavior selected for transfer need not include everything the original system could do. These are conditions on a particular saving, not reasons to deny that it is a saving. Requiring the student to repeat the teacher's formation would preserve an expense after the reason for incurring it had changed.

This is also why the title of this chapter cannot be turned into a physical valuation. Time crystallized in weights is an image of work leaving a usable result. It does not measure the work another route must require. The formal notions of logical and thermodynamic depth ask more exact questions under assumptions that a training bill does not establish.10 Nor does a bit-for-bit copy have to repeat the search that first produced the bits. Where the claim concerns comparable performance in another model, it is the relevant performance that must be compared. Historical expense cannot settle the comparison in advance.

What the saving is worth

Boulding's frozen knowledge belonged to an argument about the difference between a stock and the flow that maintains it. He wanted economists to attend to what remained available, not only to the activity passing through an economy. In the same essay he wondered whether being well fed mattered more than eating, then hesitated over a world that would maintain our bodies without the activity of taking a meal. Something might be lost in achieving the condition too efficiently.11

One can share the hesitation without extending it to a failed checkpoint upload. There is no general virtue in requiring the same trouble to be taken twice. What deserves preservation is not settled by the fact that time was spent on it. A usable model can save work for which its recipient has no reason to feel nostalgic, while its documentation preserves a difficulty that someone else may need to understand. The expenditure is past; the decisions about what to keep using are still ahead.

For the producer, this can be an uncomfortable division. A discovery may be useful enough to alter other people's work and insufficiently exclusive to repay the undertaking that financed it. For a recipient, the same division can make an otherwise unaffordable project possible. The social gain and the producer's return need not move together. That is an old problem of invention, made particularly visible when a costly computation yields parameters that can be copied. The ease of copying does not render the initial achievement trivial. The initial achievement does not by itself make every subsequent saving the producer's property.

Teece's account of profiting from innovation locates part of the answer in the arrangements around an invention: protection against imitation, and the manufacturing, distribution or other complementary capacities needed to turn it into a sale. Those capacities need not belong to the innovator.12 Where a model's release permits commercial use, the corresponding inquiry concerns who can actually operate it, adapt it to a useful task and reach the people prepared to pay. A capable recipient may already have what the original producer lacks. Another may obtain the weights and remain unable to afford their use.

There is nothing in that second possibility which establishes a permanent advantage for the original laboratory. Dependence may move elsewhere. If suitable computing resources are available from competing suppliers, the recipient can buy a service without submitting the purpose of the project to the model's maker. If only a few suppliers can support the intended use, those suppliers acquire a position which the diffusion of the model does not itself dissolve. The difference has to be investigated in the market concerned; it cannot be read from the number of parameters. Making the file available and making its productive use widely attainable are related achievements, with different remaining costs.

The recipient also has choices which an inventory of missing resources can obscure. A smaller task, a different implementation or a distilled model may make the inheritance useful on other terms. Adaptation need not be an attempt to catch up with the originator. It can be a decision not to pursue what the originator was pursuing. This is one reason to resist treating the producing pipeline as the only consequential asset. The pipeline can remain valuable while its output becomes valuable to undertakings which neither possess nor want it.

Once that happens, the relation between the maker and the user has changed. The maker's contribution remains in the causal history, but the user can continue without purchasing its repetition. A license may reserve particular rights; further discoveries may give the producer further advantages; operation may remain costly. None of those facts restores historical expenditure as a measure of what must now be paid. The economic question concerns the alternatives still available to each party, including the alternative of doing something useful that the producer never intended.

Crystallization, then, deserves its place as an image only if we allow the crystal to leave. Accumulated work can equip another undertaking without transferring the institution that performed it, and without requiring that institution to remain indispensable. What travels is neither the whole past nor merely a receipt for its expense. It is a means of proceeding. The past can explain why that means exists without deciding who must be paid each time it is used.

Source notes

Footnotes

  1. Meta, OPT-175B logbook, PDF pp. 39–40, entries for 3 December 2021, 10:35 p.m. ET, and 4 December, 5:35 a.m. ET. The record distinguishes local checkpoints, a manual copy and the subsequent upload to a new destination. The page order is reverse chronological. The chapter follows the dates, not the order of extraction. ↩

  2. Logbook, PDF p. 16, 19 December 2021 entry. The copy without optimizer state was prepared for evaluation during development; it is not identified here as the eventual public release file. Susan Zhang and colleagues, OPT: Open Pre-trained Transformer Language Models, arXiv:2205.01068v4 (21 June 2022), §§2.2–2.5, describe the optimizer, numerical representation and restart practice. Resuming a run additionally depends on compatible software, configuration and data; weights alone do not restore all training state. ↩

  3. Logbook, PDF pp. 40–43 and 48–50, especially the 2 December debrief at 17:16 ET and the 3 December return to Adam at 7:20 a.m. ET. The restored settings are the team's documented diagnosis. The subsequent SGD attempt also had a weight-decay implementation issue. No controlled ranking of optimization algorithms is inferred. ↩

  4. Zhang, Roller, Goyal, Shleifer and Ott, 17 November 2021 development update, sections on the 11.xx and 12.xx lineages. This is a contemporaneous progress report, including provisional expectations; the later paper's §2.5 summarizes the completed process. ↩

  5. Zhang, Roller, Goyal and Shleifer, 7 January 2022 completion notice, opening and “Cluster Deletion.” The thirty-three-day equivalent assumes uninterrupted training on 1,024 80GB A100s; OPT paper §2.4 reports training on 992. The notice's estimated compute duration is not elapsed calendar time, a minimum necessary cost or a cost of all development. Recovery counts are the team's report. ↩

  6. OPT release README at 3 May 2022 and model license, §§1–2. The snapshot lists downloadable weights from 125M through 30B, 66B as forthcoming, and an application for 175B. The June paper describes release through 66B and identifies eligible research constituencies. The repository README distinguishes the code's MIT license and component exceptions. The chapter describes these historical terms, not current permissions or a legal opinion about enforceability. ↩

  7. Richard Nelson and Sidney Winter, An Evolutionary Theory of Economic Change (1982), chapter 5, pp. 99–106 and 117–120: organizational memory and replication. Their argument supplies a serious antecedent. The claim that a trained artifact can spare a recipient from reconstructing particular producing activities is the present application, not a refutation of their account. ↩

  8. Cristian Bucilă, Rich Caruana and Alexandru Niculescu-Mizil, “Model Compression”, KDD 2006, pp. 535–541, §§1–2. Their method trains a compact model using outputs assigned to additional data by an ensemble; data generation and the learner's capacity constrain the result. Hinton and colleagues explicitly acknowledge this predecessor. ↩

  9. Geoffrey Hinton, Oriol Vinyals and Jeff Dean, “Distilling the Knowledge in a Neural Network”, arXiv:1503.02531v1 (9 March 2015), §§1–4, especially Table 1. Word-error rates were 10.9% for the baseline and 10.7% for ensemble and student on the reported evaluation. The experiment does not establish deployment of the distilled model, universal equivalence or total production cost. The body distinguishes retained tested behavior from identity of systems. ↩

  10. Charles Bennett, “Logical Depth and Physical Complexity” (1988), §§1 and 3. The familiar deterministic characterization uses the least runtime of a program within a specified number of bits of the shortest description, for a fixed universal machine. Bennett's discussion also develops probability-sensitive alternatives and a definition using an s-incompressible program; the near-shortest formulation is his tentative Definition 0.2, not the sole formulation in the paper. Seth Lloyd and Heinz Pagels, “Complexity as Thermodynamic Depth”, Annals of Physics 188 (1988), 186–213, defines depth over processes leading to a state; under its Hamiltonian treatment it is the difference between coarse- and fine-grained entropy. Applying that construction requires a history ensemble and coarse-graining. Neither measure has been established here for OPT, nor does either convert observed expenditure into a proof of unavoidable search or economic value. ↩

  11. Kenneth Boulding, “The Economics of the Coming Spaceship Earth”, in Henry Jarrett, ed., Environmental Quality in a Growing Economy (1966), pp. 3–14. Epigraph: p. 6; knowledge stocks: pp. 6–7; stocks, throughput and the hesitation over eating: pp. 9–11. The application to the retention of computational results is the chapter's interpretation. ↩

  12. David Teece, “Profiting from Technological Innovation”, Research Policy 15 (1986), 285–305, especially pp. 285–290, §§3.1 and 3.3. The distribution of gains depends partly on appropriability and on whether complementary assets are generic, specialized or cospecialized. The subsequent discussion of computing suppliers is a conditional economic argument, not an observed profit distribution from OPT. ↩

Search the book

Use ↑ ↓ to move through results; Escape to close.

Search every published chapter, section and reference.

    In this chapter