Menu

Interlude The Story of Reasoning

Probability theory is nothing but common sense reduced to calculation.

— Pierre-Simon Laplace

You have watched one principle dissolve the puzzles that broke philosophy after philosophy. Now set the logic aside and ask a different question, not what reasoning requires but where it came from, because we are living through the most important part of that story and cannot see it clearly without the earlier ones.

For almost all of time there was no reasoning that left a trace. None. Four billion years of chemistry, three billion of life, cells dividing and species rising and falling, and nothing anywhere drawing a conclusion from evidence. Then nervous systems, and with them a kind of inference that was not yet thought: the mouse learning the cat’s hours, the crow remembering the face that threw stones, prediction encoded in neurons and bounded by a single lifespan. For hundreds of millions of years every insight died with the animal that had it, and knowledge accumulated only in genes, blindly and without intention, one slow correction per generation.

Then, somewhere in the last hundred thousand years, humans began to talk, and thought could leave one skull and enter another for the first time. This was reasoning’s first escape from the body, and it changed the unit of knowing from the individual to the group: a hunter could say where the game had gone, an elder could describe the road to water, a mother could warn her children about the snake that had killed their uncle, and wisdom began to accumulate in culture rather than in blood. The transmission leaked, memory faded, details drifted in the retelling, but imperfect inheritance is infinitely more than none, and oral cultures built astronomy and agriculture and law out of it. The ceiling was memory. You could keep only what a living mind could hold, and complex arguments could not be checked, and when the elder died some of the world died too.

About five thousand years ago came the second escape: writing, marks that outlived the hand that made them, invented in Mesopotamia and Egypt near the start and again, independently and some two thousand years later, in China. Now memory was external, and a thought could be set down, left, and returned to years later, and arguments too long to hold in a head could be laid out and checked step by step. Euclid’s geometry outruns the unaided mind; there are too many dependencies; but it can be done on papyrus, proof stacked on verified proof, and so mathematics became possible in a way it had never been. Alexandria was the dream of it, four hundred thousand scrolls in one place, until the dream showed its flaw, by fire and neglect and slow dispersal rather than in one blaze, and plays that survived in single copies vanished, and we no longer even know the full list of what we lost. Written memory could accumulate, and written memory could also burn.

For a thousand years the bottleneck was copying, every text reproduced by one hand at a time, until the 1450s and Gutenberg’s press, and the cost of a book fell off a cliff. A workshop now made more copies in a day than a scribe managed in a year, and ideas that had stayed local went across a continent in weeks, and the deeper change was not speed but discipline: print meant standardisation, a thousand identical copies where before two Aristotles differed in a hundred places; it meant verification, a result published and read and tested by strangers far away; it meant that each generation could reliably start where the last had stopped. Reasoning had become social, then permanent, then fast.

And by the nineteenth century a stranger question surfaced. Not how should we reason, the question so far, but what is reasoning, structurally, such that it could be written down as a process rather than only its conclusions. Could thought itself be formalised? George Boole believed the laws of thought could be made algebra, and made a start. Gottlob Frege built the first system rich enough to carry real mathematical proof, and we watched, in the self-grounding chapter, what happened when Russell’s letter reached him. Hilbert dreamed of reducing all of mathematics to mechanical rule, and Gödel, whom we have already met, proved the dream impossible in principle, and then, in proving it, something unexpected fell out of the wreckage. To state exactly what a formal system could and could not do, Alan Turing had to define, precisely, what it means to carry out a mechanical procedure at all, and his definition, an imagined machine reading and writing symbols on a tape by fixed rules, was the blueprint for every computer that now exists. The attempt to formalise reasoning had accidentally specified the machine that would come to perform it. Nobody planned that.

And here the history stops being background and becomes the room you are sitting in, because the thread of this history and the argument are about to touch, and the contact is the point of the whole detour. The people who formalised inference were not building philosophy. Cox, asking in the 1940s what rules any consistent measure of belief must obey, was doing mathematics. Shannon, asking how much information a channel can carry, was solving a problem for the telephone company. Jaynes, insisting that the least-committal distribution consistent with your constraints is the only honest one, thought he was cleaning up statistical physics. None of them knew they were writing the operating manual for a kind of mind that did not yet exist.

But when engineers finally built machines that learn, the mathematics those machines turned out to run on was not something new invented for the occasion. It was this. The same mathematics, exactly. A modern learning system adjusts itself to reduce a quantity its designers call a loss. The particular loss that governs the language engines, and a great share of machine classification besides, is the cross-entropy score against which a model’s every next-token belief is corrected. Written in different notation, it is a measure of surprise: the same one Shannon defined and Jaynes deployed and the argument has been circling since the rain first fell. Be exact about the sense of “same”: the mathematical form is identical, the objects and proofs and objectives are not, and a shared equation is not a shared purpose. The machines are not doing something adjacent to the mathematics of consistent inference. They are doing that mathematics, at scale, by the trillion operations a second, and the equations were there first, waiting, and both the philosophers and the engineers walked into the same room from different doors.

Which is why everything that remains is about them. Everything up to now has traced a single mind at a window, weighing a friend’s word about the rain. But reasoning left the single skull a hundred thousand years ago, and it has never stopped leaving, into speech and script and print and formal rule, and now into engines that perform the very inference we have put on trial, faster than any human and soon, in domain after domain, better. The framework held under every classical paradox. The harder question is what it says once the reasoner is no longer only human, once the channel delivering the sentence about the rain was trained rather than raised, once the mind on the far side of the argument is made of the same mathematics as the argument itself. That is the implication that matters now, and the story has finally caught up to the present in order to ask it.

The history is told. The debt is paid, and the machines are already in the room, running the mathematics the dead left unfinished.

The story, in 16 stops

They stand in the order the Interlude names them. Each one carries the sentence that names it.

Drag the rail. The arrow keys step between stops. The rail breaks where the Interlude’s paragraphs break.

01 / 16

Four billion years

Four billion years of chemistry, three billion of life, cells dividing and species rising and falling, and nothing anywhere drawing a conclusion from evidence.

The Story of Reasoning paragraph 2

02 / 16

Nervous systems

For hundreds of millions of years

Then nervous systems, and with them a kind of inference that was not yet thought: the mouse learning the cat’s hours, the crow remembering the face that threw stones, prediction encoded in neurons and bounded by a single lifespan.

The Story of Reasoning paragraph 2

03 / 16 First escape

Speech

Then, somewhere in the last hundred thousand years, humans began to talk, and thought could leave one skull and enter another for the first time.

The Story of Reasoning paragraph 3

Beneath this stop The explainer, channels The paper, page 36

04 / 16 Second escape

Writing

About five thousand years ago came the second escape: writing, marks that outlived the hand that made them, invented in Mesopotamia and Egypt near the start and again, independently and some two thousand years later, in China.

The Story of Reasoning paragraph 4

Beneath this stop The explainer, channels The paper, page 36

05 / 16

Euclid

Euclid’s geometry outruns the unaided mind; there are too many dependencies; but it can be done on papyrus, proof stacked on verified proof, and so mathematics became possible in a way it had never been.

The Story of Reasoning paragraph 4

06 / 16

Alexandria

Alexandria was the dream of it, four hundred thousand scrolls in one place, until the dream showed its flaw, by fire and neglect and slow dispersal rather than in one blaze, and plays that survived in single copies vanished, and we no longer even know the full list of what we lost.

The Story of Reasoning paragraph 4

07 / 16

Gutenberg

For a thousand years the bottleneck was copying, every text reproduced by one hand at a time, until the 1450s and Gutenberg’s press, and the cost of a book fell off a cliff.

The Story of Reasoning paragraph 5

08 / 16

What is reasoning

And by the nineteenth century a stranger question surfaced. Not how should we reason, the question so far, but what is reasoning, structurally, such that it could be written down as a process rather than only its conclusions.

The Story of Reasoning paragraph 6

10 / 16

Gottlob Frege

Gottlob Frege built the first system rich enough to carry real mathematical proof, and we watched, in the self-grounding chapter, what happened when Russell’s letter reached him.

The Story of Reasoning paragraph 6

11 / 16

Hilbert and Gödel

Hilbert dreamed of reducing all of mathematics to mechanical rule, and Gödel, whom we have already met, proved the dream impossible in principle, and then, in proving it, something unexpected fell out of the wreckage.

The Story of Reasoning paragraph 6

Beneath this stop The explainer, ceilings The paper, page 11

12 / 16

Alan Turing

To state exactly what a formal system could and could not do, Alan Turing had to define, precisely, what it means to carry out a mechanical procedure at all, and his definition, an imagined machine reading and writing symbols on a tape by fixed rules, was the blueprint for every computer that now exists.

The Story of Reasoning paragraph 6

Beneath this stop The explainer, ceilings The paper, page 11

15 / 16

Jaynes

Jaynes, insisting that the least-committal distribution consistent with your constraints is the only honest one, thought he was cleaning up statistical physics.

The Story of Reasoning paragraph 7

Beneath this stop The explainer, projection The paper, page 24

16 / 16

The machines

But when engineers finally built machines that learn, the mathematics those machines turned out to run on was not something new invented for the occasion.

The Story of Reasoning paragraph 8

Beneath this stop The explainer, science-and-machines The paper, page 64

The assay

Part of this chapter is proved elsewhere. The panel says which part, and where.

Open the assay for Interlude. 1 graded sentence, 1 with a result beneath, 1 where the book narrows, 1 in the margin.

The marks used here

proved Some claims are proved, and I will say so.

The paper beneath

  • proved The warrant is the open literature, named in Sources and Notes.

    Written in different notation, it is a measure of surprise: the same one Shannon defined and Jaynes deployed and the argument has been circling since the rain first fell.

    Sources and Notes 7: the loss-form identity rests on standard machine-learning theory rather than on the paper, and is graded where it is made. §10 is the paper's own account of the same quantity.

    §10 The explainer, Count every change exactly once The paper, page 20

Where the book narrows

  • Be exact about the sense of “same”: the mathematical form is identical, the objects and proofs and objectives are not, and a shared equation is not a shared purpose.

    The loss-form identity is a shared equation, and the sentence refuses to let it be read as a shared purpose.

In the margin

In the paper

  • §10 Count every change exactly once Read the section

    The loss the language engines reduce is a measure of surprise, and the equations were there first.

  • §24–§29 Where inquiry gets done Read the section

    Reasoning leaves the skull into speech, script, print, formal rule, and finally engines.

Where to swing

The Interlude's graded sentence is marked proved. Its standing as a whole is a best explanation.

A best explanation breaks against a better explanation.

The whole book