Menu

Part Three

Chapter Nine The New Riddle

7 min

An object is grue if it is examined before time t and green, or not examined before t and blue.

— after Nelson Goodman, Fact, Fiction, and Forecast (1955)

I am going to run a confidence game on you, and I am telling you so in advance, and it will work anyway.

In 1946, at the University of Pennsylvania, the logician Nelson Goodman built a monster out of five letters, though the world would not feel its full force until he set it out in a book nine years later. Goodman had come up through symbolic logic in Russell’s long shadow, and his gift was not for solutions but for cracks: for showing that the obvious was not obvious at all. He was working on induction, on why we project some patterns forward and not others, and the previous chapter’s victory makes his question sound settled. Emeralds, every one ever examined, have been green; the update rule takes the accumulated green and raises, lawfully, your credence that the next one is green too. Forced, we said. Unique, we said. Now watch the game.

Define a new word. An object is grue if it is examined before some future time t and green, or not examined before t and blue. Read it twice; the definition is exact and the trap is not in fine print. Now check the evidence. Every emerald ever examined was examined before t, and every one was green, so every emerald ever examined has been, by the definition, grue. Perfectly grue. Not one exception.

So here stand two hypotheses before the same tribunal of evidence. All emeralds are green. All emeralds are grue. Every stone in the record confirms both, completely, identically, and the two hypotheses disagree about every emerald examined after t, when grue requires blue. Your forced, unique update rule was supposed to take the evidence and tell you what to believe about the next emerald. Which pattern does it project? The evidence, Goodman showed, cannot say, because the evidence fits both. Something else has been doing the work all along, some silent principle choosing green over grue before the counting ever starts, and until you can name it and justify it, the proud machinery points in every direction at once. Hume showed induction could not be justified from outside. Goodman seemed to show it could not even be specified from inside. That is the new riddle, and it resisted solution for seventy years, and I will not insult it with a quick answer, because the strait path here passes within an inch of a loss the framework would not have survived.

The first answer looks clean, and you can probably build it yourself from Part Two. MU says spread your credence and let constraints gather it; and hypotheses are not all the same size. “All emeralds are green” posits one stable property. “All emeralds are grue” posits a property before t, a different property after t, and the special time t itself: three moving parts where green has one. Spread prior probability plainly across a space of hypotheses and the many-parted ones receive it spread thinner, since their probability must cover more ways of being wrong; this is Occam’s Razor not as taste but as theorem, the friar’s heuristic finally given its engine, and it seems to end the riddle in a paragraph. Grue is the complicated hypothesis. Complicated hypotheses start behind. The evidence never distinguishes the two, but the starting line does, and green wins by inheritance.

Enjoy that paragraph for a moment, because Goodman is about to take it away from you, and this is the part of his argument that most retellings omit, and omitting it is how a book cheats. Count again, Goodman says, but count in my language. Define a second word, bleen: examined before t and blue, or not examined before t and green.

Now speak the dialect in which grue and bleen are the primitive colours, and describe your two hypotheses again. “All emeralds are grue”: one word, one stable property, simplicity itself. “All emeralds are green”: ah, in this language green must be defined, as grue before t and bleen after, a property that switches at a special time, three moving parts. The complexity you counted was not in the hypotheses. It was in the dictionary you counted with. Simplicity is relative to a language, and languages are symmetric, and the razor cuts whichever way the vocabulary tilts it, and the vocabulary was the thing to be justified. This is the real riddle, the adult version, and I want to be plain about what just happened: the clean answer is circular, and the trial is genuinely going badly, and if this is where the argument stood, it would be over.

Here is what survives, and what it costs. The rescue is not a cleverer count. It is noticing what the count was always relative to, and admitting it into the constraints where it belonged. MU never operated on free-floating hypotheses in no language at all; there is no such place. It operates on the constraints of an actual reasoner, and among any actual reasoner’s constraints is its apparatus: the sensors it measures with, the concepts its channels natively carve, the code its hypotheses are actually written in. Your eye is an instrument that responds to reflectance and does not consult the calendar; a photometer is a device whose readings mean the same thing on both sides of any t; and relative to that apparatus, the apparatus you in fact have, the counting is not symmetric and never was. Green is what your instruments report directly; grue is a construction that must be assembled from a reading plus a date, and the asymmetry is now a physical fact about the machinery of your evidence, not a prejudice of your dictionary. Given an apparatus, MU’s ordering is forced, the razor cuts true, and green wins, lawfully. That is the rescue.

Now the cost, stated as plainly as I know how, because this is the promised wound, kept for good. What has been shown is narrower than what the confident paragraph claimed. MU does not deliver induction from nowhere, valid for every describer in every language; it delivers induction for a reasoner with an apparatus, relative to the channels and code that reasoner actually has. Goodman wins a permanent point: there is no language-neutral, apparatus-free simplicity, and any account that claims one is smuggling its dictionary. The riddle is not dissolved the way the ghost was. It is scoped. And if you ask the natural next question, whether our apparatus itself is arbitrary, whether evolution and engineering could as easily have handed us grue-eyed instruments, the reply is a best explanation rather than a theorem. Sensors built by a world of stable causes get shaped to track the stabilities, which is why eyes track reflectance and not calendars. That story is good. It is abductive, and I mark it so. The scar stays visible. It is the most instructive mark on the framework, because it shows what the framework is: not a view from nowhere, but the forced discipline of a situated reasoner, which is the only kind of reasoner there has ever been.

And the scar itself earns its keep on the same night you read this, because the same confidence game is running right now at planetary scale. Every learned system that has ever shipped was trained on data examined before its own time t and deployed on a world after it, and the gap between the pattern that fit the past and the pattern that governs the future has a name in the machine-learning laboratories, several names: distribution shift, shortcut learning, the model that aced every benchmark and failed in the clinic because it had learned the hospital’s scanner artifacts rather than the disease. That is grue, industrialised. Somewhere tonight a model is generalising on grue, projecting its training’s dialect into your morning, and the engineers fighting it have rediscovered, in code, every inch of this trial: that the evidence alone cannot choose the projection, that the choice lives in the apparatus and the representation, and that the right response is not a guarantee but a discipline. Goodman never lived to see his riddle get a deployment pipeline. It did.

The court will move faster now, and here is why. Three trials in, a pattern has emerged, and naming it converts the pattern into a toolkit, because every dissolution in this act, past and coming, is built from four moves, alone or in combination. Underdetermination is the exclusion engine: if two reasoners could hold identical constraints and land on different outputs, then whatever separated them was never determined by the constraints at all, and out it goes; this single move drives most of the framework’s uniqueness results, and you just watched its negative image, two hypotheses one evidence-set could not separate. Retention is the conservation engine. When new constraints arrive, keep every piece of what you already had that can consistently survive them, change only what joint survival forbids, and where even that leaves options open, hold the whole set of survivors. The update rule itself, in both its classical forms, is this principle’s offspring, and so is this act’s habit of keeping each plaintiff’s true discovery at full strength. Refinement is the symmetry engine: adding a perfectly symmetric copy of what you already have carries no new content, so no consistent answer is permitted to change under the addition, and from that quiet demand the uniform starting points of Part Two follow.

Diagnosis is the engine of this whole act: aim a demand at constitutive structure, at the very thing that makes demanding possible, and the demand instantiates what it questions, leaving only a hypothetical residue the framework can answer, which is what happened to Hume’s request for a bridge and will happen to every plaintiff still waiting. Four forms. You have now seen each one work. The remaining plaintiffs meet them faster, and you are equipped to check every move, which is the point: the method was never meant to stay mine.

The next trial begins in a courtroom inside the courtroom, with a kind of inference so common you ran it before breakfast, and so treacherous it has hanged innocent people: the leap to the best explanation.

The assay

Nothing in this chapter is proved elsewhere. The panel says what holds it up instead.

Open the assay for Chapter Nine. 2 graded sentences, 1 where the book narrows, 4 in the margin, 1 term.

The marks used here

argued Some are the best explanation I can offer for what the evidence shows, and I will say that too.

The paper beneath

  • argued The book's own argument. Nothing lies beneath it.

    And if you ask the natural next question, whether our apparatus itself is arbitrary, whether evolution and engineering could as easily have handed us grue-eyed instruments, the reply is a best explanation rather than a theorem.

    No numbered result stands under this sentence. It is marked argued and nothing more.

  • argued The book's own argument. Nothing lies beneath it.

    It is abductive, and I mark it so.

    No numbered result stands under this sentence. It is marked argued and nothing more.

Where the book narrows

  • Now the cost, stated as plainly as I know how, because this is the promised wound, kept for good.

    The scar: induction for a reasoner with an apparatus, not from nowhere.

In the margin

Terms

Update rule The book

The unique consistent rule within the stated prior, constraint, and consistency-condition scope; the one way to move credence when constraints arrive. In the glossary

In the paper

  • §21–§22 A right answer can still be bad inference Read the section

    Grue survives the simplicity count until the apparatus enters the constraints, and the scope is narrowed for good.

  • §6 An empty class is a finished answer Read the section

    Simplicity is relative to a language, and the razor cuts whichever way the vocabulary tilts it.

Where to swing

A best explanation breaks against a better explanation.

The whole book