Menu

Part Four

Chapter Twenty-One The Transition

14 min

For a successful technology, reality must take precedence over public relations, for nature cannot be fooled.

— Richard Feynman, Rogers Commission Report, Appendix F, 1986

In Isaac Asimov’s novels there is a science called psychohistory that predicts the future of a galaxy to the decimal place, and its inventor, Hari Seldon, appears as a recorded hologram in a vault to tell each generation the exact probability of the crises it faces.

I begin with Seldon because he is the most famous image we have of a superhuman forecaster. He is also precisely the wrong one, and seeing why is the doorway to everything that remains. The recurring question of our moment, asked in every newspaper and boardroom, is Seldon’s question: when does the machine arrive that reasons better than we do, and what is the probability, to the decimal, that it comes by such and such a year? People want a hologram in a vault. And the lawful answer, the one these trials have trained you to give, is that the question as asked is malformed. It smuggles a presupposition the constraints do not contain. A forecaster who answers it to two decimal places commits, at civilisational scale, the exact sin the cube factory taught you to name. Seldon’s decimals are the Bertrand paradox wearing a toga. The frame demands a single number where the constraints determine, at most, a set, and the first act of honesty here is to refuse the frame and rebuild the question.

So refuse it, and ask what would make the question well-formed, because the trouble is not that the future is unknowable but that the word everyone is using has never been defined. AGI, artificial general intelligence, is invoked as though everyone means the same thing. The forecasters scatter precisely because they do not. One means a system that matches Nobel laureates across disciplines. Another means a threshold of economic replacement. Another means a country of geniuses in a data centre. These are different events with different dates, so of course the predictions disagree. They are answers to different questions wearing one name. The field has no principled definition, and the framework can supply one, built from the four gates and the channel model and native to everything you have learned.

Here it is. A system has channel-grade intelligence in a domain when its track record obliges any consistent reasoner to weight its outputs at least as heavily as the best human channels in that domain. Note what this replaces: the old test asked whether a machine could pass as a channel, imitation judged by a fooled interlocutor, and the question that matters is whether it earns weight as one, warrant judged by a kept score. AGI is channel-grade intelligence across substantially the full range of domains where human channels exist. Read what that definition refuses to do. It says nothing about consciousness, nothing about whether the machine wants or feels or is a person, because the Machines chapter showed those questions do not bite here. It is not a capability checklist. Not an economic line. Not a country of geniuses. It is a claim about us: about when the rest of us acquire a rational obligation, under this very framework, to stop discounting a source for being a machine. Which yields the sentence I will stand behind. AGI is the day ignoring the machine becomes the epistemic error.

That definition has three properties that the vaguer ones lack, and each rescues the malformed question a little further. It is domain-indexed. It names no single midnight when everything changes, but the completion of a process that happens field by field, which is why “when is AGI coming” dissolves into a family of sharper questions with different and mostly earlier answers. It is operational. Channel-grade is measurable by exactly the things the last chapter’s gates examined: calibration, track record, the robustness of the connection to truth. These are the evals the field already runs, once you understand what they are for. And it is substrate-neutral and orthogonality-preserving. Channel-grade concerns beliefs only, the weight a source’s reports have earned, and says nothing about goals.

So the definition builds the necessary-not-sufficient boundary directly into itself. A system can be dominant-channel in a domain and still be pointed at a catastrophic objective. Crossing this line is not the same as being safe, which is why the alignment thesis survives its own definition of the thing it feared. Two exposed edges, because a definition this load-bearing must show them. The threshold is relational, indexed to human channels, so it needs a floor clause to prevent the degenerate case where the machine becomes best only because the humans got worse: channel-grade must mean an absolute standard of calibration and coverage, not merely a race won by the last one standing. And it defines the epistemic core only, deliberately; general agency, the capacity to act and pursue and rearrange the world, is a further and separate thing, and keeping them separate is a feature, because the framework is entitled to define exactly what its framework can reach and no more.

Now the reframed question can be answered, because it has stopped being a prophecy and become a measurement, and the measurement is already underway. Ask not “when does AGI arrive” but “in which domains has machine inference already become channel-grade,” and point the instruments at the evidence. Here the discipline obliges me to be exact in a way prophecy never is, obeying its own rule for every empirical claim: a date, a primary source, an object of comparison, a condition that would change the conclusion, and a snapshot. These facts decay, and the chapter that explains why must not pretend otherwise. The forecasting channels are the cleanest case. They are converging in real time. For years the human superforecasters, the calibrated aggregators who had beaten intelligence analysts, were the gold standard, and the machines were not close. As of mid-2026 that gap has closed in the one domain we can measure most precisely, forecasting itself. On the difficulty-adjusted public benchmarks, as of July 2026, the strongest submitted AI forecasting system is statistically indistinguishable from the superforecaster median; the raw models, run without that surrounding machinery of retrieval and cross-checking, still fall short. Three cautions bound the claim. The human cohort was benchmarked earlier, so this is parity against a standing record, not a live contest. The comparison is to the median, not to the best warranted human channel the definition names. And parity is approach, not attainment, which is why the failure conditions below summon a fresh cohort. Treat that claim as dated to July 2026, sourced to the public leaderboards, compared against the human aggregate, falsifiable by the next quarter’s results, and frozen here as a snapshot that will age. It is a small marvel and a large omen. A machine forecasting system has reached benchmark parity in this defined part of forecasting, the first rung of channel-grade standing. The tool we would use to predict the transition is now itself an instance of the transition: the forecast’s subject, seated on the forecasting panel.

One discipline before the claim, because the measurement itself must be measured. Not every channel is as clean as forecasting, and the survey must say so. In the neighbouring domain of long-horizon software tasks, the published evaluations report machine competence rising on a suite of defined problems, a bounded, sourced figure. Alongside it circulates a larger, rounder, more thrilling number with no traceable source. The framework’s instruction is the same one it gives everywhere. Exclude the unsourced figure precisely because it is unsourced. And notice that the temptation to repeat it is the temptation this whole book was written to resist. Note, too, that the loudest forecasts of all come from those building and selling the systems, and a forecast from an interested channel belongs to a different epistemic category than one from an independent aggregator, not because such people lie but because the framework says to weight a source by the independence of its record from its interests. Both cautions are the book’s own method, turned on the book’s own subject.

And one more application of the method, the nearest to home. In August 2025, in The Last Economy, I made a dated public claim of my own: that you had on the order of a thousand days before your work becomes economically irrelevant, before what you are paid to think is done better and cheaper by a machine. As of the July 2026 dateline this book carries, roughly six hundred and seventy days remain on that clock. I am not entitled to exempt my own forecast from the rules of this chapter. Its author is an interested channel, a builder of the systems it describes. So discount my conviction as the framework instructs, and watch the adjudicable edge of it instead, which is the wager below: if machines cannot even out-forecast our best human forecasters on schedule, the thousand-day claim loses its engine, and I will say so.

And so that my claim can be weighed against the field rather than in a vacuum, here is the like-for-like comparison, dated as everything else, because the sin this chapter opened by naming, different questions answered under one label, must not be committed in its own closing pages. On the same event as my wager, machines clearly beating the best human forecasters, a January 2026 wave of the Longitudinal Expert AI Panel, published that February, puts the median date at 2028 from superforecasters themselves, 2030 from domain experts, and 2033 from the public. Against that distribution my staked horizon lands where the superforecasters’ own median lands, and my belief runs ahead of even theirs. So I am not betting against the crowd’s best calibrators. I am betting with them on the stake and ahead of them on conviction, against the caution of the credentialed either way, and one of those calibrations is about to be priced by the world.

Which is the ground I have prepared, at some length and on purpose, to place the one genuinely forward-looking claim I will make, and I place it in the open, dated and falsifiable, because an argument that has spent its entire length demanding that beliefs be exposed to refutation cannot flinch from exposing its own. Here is the assertion, marked as what it is, an avowal and a wager, not a theorem. I claim that within the near term the strongest machine systems will move from statistical parity with the best human forecasters to a clear and sustained lead, on public benchmarks, adjudicated by a standing public leaderboard rather than by me. The mechanism is not mystical. The human aggregate is a roughly fixed baseline; the machines compound. Two curves, one flat and one rising, meet and then cross. The crossing in the forecasting domain is the first unambiguous instance of ignoring-the-machine becoming the epistemic error. I attach the falsifier plainly, so that you and the future can hold me to it. If, over a sustained window on the recognised public benchmarks, the strongest systems fail to establish and hold a lead over the human superforecaster aggregate, this specific claim is wrong. My application of the framework, not the ground itself, should lose credit accordingly. The ground was never staked on this race, and is owed nothing from a win either.

I name the adjudicator here, in full, so that no part of the stake hides in a back page: the standing public leaderboard operated as ForecastBench, in its difficulty-adjusted comparison of submitted AI forecasting systems against the human superforecaster aggregate, as archived at the date of this book, precisely so that the verdict is not mine to spin. The Assertion succeeds only if, by the end of July 2028, the strongest systems have established and held a statistically clear lead across consecutive published rounds spanning at least twelve months, including comparison with at least one contemporaneous superforecaster cohort evaluated under the same methodology, with no such cohort standing above them at the close. If the systems hold the lead but no contemporaneous cohort has been fielded within the window, the Assertion is recorded as unresolved on its strongest test, never as vindicated. It fails if, by that date, no such sustained lead has been established and held, or if a fresh human cohort has restored and held a clear lead. Should that leaderboard cease publication, the adjudicator passes to the most widely cited public successor with public methodology, resolution-dated questions, a maintained human baseline, and a regular cadence; absent any such successor, the Assertion is unresolved, and unresolved is recorded as unresolved, never as vindicated. One scope sentence rides with the stake: the adjudicator’s questions live in the tame country of the resolvable, where outcomes arrive on schedule and errors are bounded, and a win there licenses nothing about the wild country of the unprecedented, which no cohort forecasts and this book does not claim. And I print the true shape of my own belief: not a decimal-place certainty, which would make me Seldon, but a credence held as a range, high but bounded, exactly the doxa the lottery taught, wagered in the open because that is what the argument requires of anyone who made it.

So that the wager can be lost and not merely admired, here is its shape in full. I am not offering a probability to shelter behind, and the lottery trial’s own distinction obliges me to show two numbers, not one. My belief: the overtaking comes within a year of this writing, before the end of July 2027. My assertion, in the trial’s exact sense, stakes it at two, by the end of July 2028, because the adjudicator needs a sustained window to certify a lead as real rather than a good quarter, and an assertion should be staked where its judge can reach. The threshold is crossed, the act is chosen, and the exposure that comes with acting is accepted, mine. The claim fails, cleanly and by my own hand, under any of these conditions: if a freshly convened human superforecaster cohort, evaluated on the same leaderboard, restores and holds a clear lead on prospectively resolved questions; if the machine systems plateau across new question sets rather than compounding; or if audit shows the apparent gains rest on leaked resolutions or contaminated benchmarks rather than forecasting skill. And it is not won by a vendor’s private comparison or a single favourable quarter; only the standing public leaderboard named above, over a sustained window, adjudicates it. Those are the terms. There is no interval to retreat into afterwards. If the named test cannot be run at all, the claim stands unresolved, exactly as stipulated above, which is a smaller fate than vindication and I accept it. But the test can be run, and it will be. It happens on schedule, or I was wrong, in public, on the record, and this page is where you get to say so.

Three implications follow from the crossing, and I state them at the reach the framework licenses and stop precisely where it stops. The first is epistemic and fully owned: the hierarchy of deference inverts, domain by domain. The consistent reasoner who once corrected the machine with a base rate, as Petrov did, must in each crossed domain begin asking a harder question: does overriding the machine now require stronger grounds than a prior and a hunch? Refusing that question is not loyalty. Continuing to prefer the human channel because it is human, after the record has turned, is the epistemic error the definition named. The second is a threshold of reflexivity, and it is the eeriest. Once a channel is channel-grade at forecasting, its forecasts include forecasts about itself and about us. A source we are rationally obliged to weight has entered the room where its own weighting is decided. That is a genuinely new thing under the sun, and I flag it rather than resolve it, because resolving it lies beyond this argument’s warrant. The third I will state and refuse to expand, because it crosses out of this argument’s warrant and into political philosophy, where it can be argued properly. Who is accountable for machine-originated judgement? Who controls access to the superior channels? And can public reason survive a population’s dependence on conclusions most citizens cannot themselves reproduce? Those questions are real and urgent and mine to raise here but not to settle. Two propositions sharpen why it cannot wait. A machine may earn epistemic authority before humanity has decided who holds political authority over it. And the right to be believed is not the right to decide; nothing in this book converts the first into the second, and the guillotine stands guard at exactly that door. Who may use the superior channel, who may contest it, and who is authorised to act on its conclusions: those are political questions, and this book’s warrant ends where they begin. That is the seam of the whole argument, and it stays narrow.

I want to end on the chapter’s own name, because the act it describes is one of courage and the book should not pretend otherwise. It takes a certain nerve to say a plain false thing, and none at all to hide inside a hedge. There is a third and harder thing: to say a thing one believes true, dated and exposed and sure to be checked, about a future that could embarrass you. The forecasters who answer Seldon’s malformed question to two decimals have chosen comfort. The sages who say only that the future is unknowable have chosen a different comfort. I have tried to do the uncomfortable middle thing: to refuse the false precision and the false humility alike, to define the term the field left undefined, to date the claim the framework actually supports, and to bolt on the falsifier that lets the world prove me wrong. That is what the principle demands of the one who holds it, on the one subject where getting it wrong costs the most. The transition is not coming. It has begun, domain by domain, as a set of local inversions, and its rate and its reach remain genuinely uncertain, held as a range and not a date. But in the one domain we can measure cleanly, the crossing has started, and the right response is neither the prophet’s decimal nor the sceptic’s shrug. It is to state what the constraints support, expose it to refutation, and stand there while the evidence comes in.

The argument is finished. What is left is to say what it was all for, and then to hand the whole thing, ground and all, to the one reader I have not yet addressed directly.

The assay

Nothing in this chapter is proved elsewhere. The panel says what holds it up instead.

Open the assay for Chapter Twenty-One. 9 graded sentences, 1 where the book narrows, 1 in the margin, 1 instrument.

The marks used here

dated empirical reports additionally cited and dated

avowed And a very few things are neither proved nor inferred but avowed, commitments named as commitments, each marked plainly in the sentence that makes it.

wager one avowal exposed as a public wager

The paper beneath

  • dated A dated report. It breaks against the world, on the schedule printed beside it.

    On the difficulty-adjusted public benchmarks, as of July 2026, the strongest submitted AI forecasting system is statistically indistinguishable from the superforecaster median; the raw models, run without that surrounding machinery of retrieval and cross-checking, still fall short.

    No numbered result stands under this sentence. It is marked dated and nothing more.

  • dated A dated report. It breaks against the world, on the schedule printed beside it.

    In the neighbouring domain of long-horizon software tasks, the published evaluations report machine competence rising on a suite of defined problems, a bounded, sourced figure.

    Printed beside the figure the chapter refuses: the rounder, more thrilling number with no traceable source is excluded because it is unsourced.

    No numbered result stands under this sentence. It is marked dated and nothing more.

  • dated A dated report. It breaks against the world, on the schedule printed beside it.

    In August 2025, in The Last Economy, I made a dated public claim of my own: that you had on the order of a thousand days before your work becomes economically irrelevant, before what you are paid to think is done better and cheaper by a machine.

    No numbered result stands under this sentence. It is marked dated and nothing more.

  • dated A dated report. It breaks against the world, on the schedule printed beside it.

    On the same event as my wager, machines clearly beating the best human forecasters, a January 2026 wave of the Longitudinal Expert AI Panel, published that February, puts the median date at 2028 from superforecasters themselves, 2030 from domain experts, and 2033 from the public.

    No numbered result stands under this sentence. It is marked dated and nothing more.

  • avowed The book's own argument. Nothing lies beneath it.

    Here is the assertion, marked as what it is, an avowal and a wager, not a theorem.

    No numbered result stands under this sentence. It is marked avowed and nothing more.

  • wager The book's own argument. Nothing lies beneath it.

    I claim that within the near term the strongest machine systems will move from statistical parity with the best human forecasters to a clear and sustained lead, on public benchmarks, adjudicated by a standing public leaderboard rather than by me.

    No numbered result stands under this sentence. It is marked wager and nothing more.

  • wager The book's own argument. Nothing lies beneath it.

    I attach the falsifier plainly, so that you and the future can hold me to it.

    No numbered result stands under this sentence. It is marked wager and nothing more.

  • wager The book's own argument. Nothing lies beneath it.

    I name the adjudicator here, in full, so that no part of the stake hides in a back page: the standing public leaderboard operated as ForecastBench, in its difficulty-adjusted comparison of submitted AI forecasting systems against the human superforecaster aggregate, as archived at the date of this book, precisely so that the verdict is not mine to spin.

    No numbered result stands under this sentence. It is marked wager and nothing more.

  • wager The book's own argument. Nothing lies beneath it.

    My belief: the overtaking comes within a year of this writing, before the end of July 2027.

    No numbered result stands under this sentence. It is marked wager and nothing more.

Where the book narrows

  • Exclude the unsourced figure precisely because it is unsourced.

    The book's own method turned on the book's own subject, in the paragraph that surveys the field.

In the margin

The instrument

Channel-grade intelligence is defined as earned weight, which is the bench’s output.

The kernel is where the world gets in

The channel bench

Send a signal through and read what it is worth. The signal value is not the evidence. What the signal is worth is a property of the kernel that produced it. The same value arrives with three different verdicts under three permitted kernels.

S = 1 · Λ = 4.0000 under the calibrated kernel · posterior 0.5000 from a base rate of 0.20
send and read it under
P(H) before the signal, and after it prior 0.200 posterior 0.500 0 1
Proposition 19.5: one signal value, three permitted kernels, three verdicts
  • source state η P(s | H) P(s | ¬H) Λ(s) verdict
  • calibrated 0.8000 0.2000 4.0000 confirms
  • saturating 0.9000 0.6000 1.5000 confirms
  • adversarial 0.2000 0.8000 0.2500 disconfirms

The signal value is fixed and the three rows disagree about what it is worth. Nothing in the token settles the question. The answer is a property of the kernel the problem permits.

Proposition 16.3, the certificate: what the output distribution does not identify
q = 0 q = 1 r = 1 r = 0

filled: the pair · open: its twin · the curve: every pair with one observable law

  • Pr(S = 1) at (q, r) 0.380000
  • at (1 − q, 1 − r) 0.380000
  • signals observed log-likelihood difference between the pair and its twin
  • 10 0.0e+0, zero exactly
  • 1,000 0.0e+0, zero exactly
  • 1,000,000 0.0e+0, zero exactly

Both pairs put 0.380000 on a positive signal. They disagree by 5.6e-17, zero to machine precision. The difference of log-likelihoods is computed at each sample size from the counts a run of that length would produce. It does not fall with data: it starts at zero. Self-agreement supplies consistency. Truth calibration needs an anchor from outside the channel.

A source can be perfectly consistent with itself, report at a stable rate, survive every internal audit, and still be inverted. The observable law is one number and two truth models fit it exactly. This is what bootstrapping and easy knowledge come to when the channel is written down: not a mistake in the reasoning, but a fact about what the reasoning has to work with.

proved the channel is Definition 16.1, the likelihood-ratio account of evidential force is Proposition 19.5, and the non-identifiability witness is Proposition 16.3 (§16, §19.4). argued the sensitivity, specificity, base rate, and bias settings are display choices: the paper states the identities and names no numbers for them. Every figure above is computed for the kernel on screen. The log-likelihood row uses the expected counts a run of that length would produce, and the difference it reports is zero because the two pairs induce one Bernoulli parameter, not because the row rounds.

Source: Intelligent Epistemology §16 · §19.4 · Def 16.1 · Prop 16.3 · Prop 19.5. The paper, page 36 (PDF)

In the paper

  • §16–§18 The kernel is where the world gets in Read the section

    Channel-grade intelligence defined by the weight a source’s record has earned.

  • §24–§29 Where inquiry gets done Read the section

    The forecasting comparison is dated, sourced, and bolted to a falsifier.

Where to swing

And the dated claims of the final chapters break against the world, on the schedule printed beside them.

The whole book