Interlude The Story of Reasoning
Probability theory is nothing but common sense reduced to calculation.
— Pierre-Simon Laplace
You have watched one principle dissolve the puzzles that broke philosophy after philosophy. Now set the logic aside and ask a different question, not what reasoning requires but where it came from, because we are living through the most important part of that story and cannot see it clearly without the earlier ones.
For almost all of time there was no reasoning that left a trace. None. Four billion years of chemistry, three billion of life, cells dividing and species rising and falling, and nothing anywhere drawing a conclusion from evidence. Then nervous systems, and with them a kind of inference that was not yet thought: the mouse learning the cat’s hours, the crow remembering the face that threw stones, prediction encoded in neurons and bounded by a single lifespan. For hundreds of millions of years every insight died with the animal that had it, and knowledge accumulated only in genes, blindly and without intention, one slow correction per generation.
Then, somewhere in the last hundred thousand years, humans began to talk, and thought could leave one skull and enter another for the first time. This was reasoning’s first escape from the body, and it changed the unit of knowing from the individual to the group: a hunter could say where the game had gone, an elder could describe the road to water, a mother could warn her children about the snake that had killed their uncle, and wisdom began to accumulate in culture rather than in blood. The transmission leaked, memory faded, details drifted in the retelling, but imperfect inheritance is infinitely more than none, and oral cultures built astronomy and agriculture and law out of it. The ceiling was memory. You could keep only what a living mind could hold, and complex arguments could not be checked, and when the elder died some of the world died too.
About five thousand years ago came the second escape: writing, marks that outlived the hand that made them, invented in Mesopotamia and Egypt near the start and again, independently and some two thousand years later, in China. Now memory was external, and a thought could be set down, left, and returned to years later, and arguments too long to hold in a head could be laid out and checked step by step. Euclid’s geometry outruns the unaided mind; there are too many dependencies; but it can be done on papyrus, proof stacked on verified proof, and so mathematics became possible in a way it had never been. Alexandria was the dream of it, four hundred thousand scrolls in one place, until the dream showed its flaw, by fire and neglect and slow dispersal rather than in one blaze, and plays that survived in single copies vanished, and we no longer even know the full list of what we lost. Written memory could accumulate, and written memory could also burn.
For a thousand years the bottleneck was copying, every text reproduced by one hand at a time, until the 1450s and Gutenberg’s press, and the cost of a book fell off a cliff. A workshop now made more copies in a day than a scribe managed in a year, and ideas that had stayed local went across a continent in weeks, and the deeper change was not speed but discipline: print meant standardisation, a thousand identical copies where before two Aristotles differed in a hundred places; it meant verification, a result published and read and tested by strangers far away; it meant that each generation could reliably start where the last had stopped. Reasoning had become social, then permanent, then fast.
And by the nineteenth century a stranger question surfaced. Not how should we reason, the question so far, but what is reasoning, structurally, such that it could be written down as a process rather than only its conclusions. Could thought itself be formalised? George Boole believed the laws of thought could be made algebra, and made a start. Gottlob Frege built the first system rich enough to carry real mathematical proof, and we watched, in the self-grounding chapter, what happened when Russell’s letter reached him. Hilbert dreamed of reducing all of mathematics to mechanical rule, and Gödel, whom we have already met, proved the dream impossible in principle, and then, in proving it, something unexpected fell out of the wreckage. To state exactly what a formal system could and could not do, Alan Turing had to define, precisely, what it means to carry out a mechanical procedure at all, and his definition, an imagined machine reading and writing symbols on a tape by fixed rules, was the blueprint for every computer that now exists. The attempt to formalise reasoning had accidentally specified the machine that would come to perform it. Nobody planned that.
And here the history stops being background and becomes the room you are sitting in, because the thread of this history and the argument are about to touch, and the contact is the point of the whole detour. The people who formalised inference were not building philosophy. Cox, asking in the 1940s what rules any consistent measure of belief must obey, was doing mathematics. Shannon, asking how much information a channel can carry, was solving a problem for the telephone company. Jaynes, insisting that the least-committal distribution consistent with your constraints is the only honest one, thought he was cleaning up statistical physics. None of them knew they were writing the operating manual for a kind of mind that did not yet exist.
But when engineers finally built machines that learn, the mathematics those machines turned out to run on was not something new invented for the occasion. It was this. The same mathematics, exactly. A modern learning system adjusts itself to reduce a quantity its designers call a loss. The particular loss that governs the language engines, and a great share of machine classification besides, is the cross-entropy score against which a model’s every next-token belief is corrected. Written in different notation, it is a measure of surprise: the same one Shannon defined and Jaynes deployed and the argument has been circling since the rain first fell. Be exact about the sense of “same”: the mathematical form is identical, the objects and proofs and objectives are not, and a shared equation is not a shared purpose. The machines are not doing something adjacent to the mathematics of consistent inference. They are doing that mathematics, at scale, by the trillion operations a second, and the equations were there first, waiting, and both the philosophers and the engineers walked into the same room from different doors.
Which is why everything that remains is about them. Everything up to now has traced a single mind at a window, weighing a friend’s word about the rain. But reasoning left the single skull a hundred thousand years ago, and it has never stopped leaving, into speech and script and print and formal rule, and now into engines that perform the very inference we have put on trial, faster than any human and soon, in domain after domain, better. The framework held under every classical paradox. The harder question is what it says once the reasoner is no longer only human, once the channel delivering the sentence about the rain was trained rather than raised, once the mind on the far side of the argument is made of the same mathematics as the argument itself. That is the implication that matters now, and the story has finally caught up to the present in order to ask it.
The history is told. The debt is paid, and the machines are already in the room, running the mathematics the dead left unfinished.