Uncontrollable AI threatens global disaster
We are told that an imminent breakthrough in recursive self-improvement is now essentially settled, a countdown already running within the next couple of years, and that its issue could be a disaster on the scale of Hiroshima, threatening all of humankind. Notice, though, what instrument actually generated this claim. Not a demonstrated system running in a laboratory, not a benchmark anyone can inspect, but a relay of testimony: Timothy Garton Ash reporting what experts in Silicon Valley told him, apparently informed in turn by a Stanford University friend. The question is not whether to receive this chain of report as prophecy. The question is what experiment would tell us whether the machine being described in it exists yet at all.
I have always held that a claim’s worth is what it lets you build or predict, not the eminence of the mouth it passed through. A prediction relayed through three sets of hands accumulates confidence the way a bellows accumulates heat - not by adding fuel, but by moving air faster over embers that were already there. Each retelling compresses “some researchers think this could happen” into “an imminent breakthrough,” and by the time it reaches a broadsheet column the hedge has been sanded off entirely. This is not conspiracy; it is simply how testimony behaves when it is not disciplined by an experiment. The instrument doing the work in this story is not artificial intelligence. It is the social apparatus of Silicon Valley itself - the dinner conversations, the leaked memos, the incentive every lab has to be thought closer to the frontier than its rivals - and that apparatus has never once been calibrated for accuracy.
None of this means the underlying question is empty. Recursive self-improvement is a real engineering hypothesis, and it has a real test: does a system exist that redesigns its own training procedure or architecture, without a human in the modification loop, and measurably outperforms its prior version, repeatably, at a rate that compounds rather than plateaus? That is a specific, fundable, falsifiable experiment. Nothing in the reporting names such a system, a benchmark it cleared, or a lab willing to publish the curve. What is offered instead is the sentiment of experts, which is a different substance entirely - useful as a hypothesis generator, worthless as a verdict.
The strongest version of the opposing view deserves an answer, not a dismissal. One might say: given stakes at civilisational scale, we cannot afford to wait for the tidy demonstration before we act, any more than one waits for a house to finish burning before calling it a fire. I take this seriously, because a precaution that names a mechanism and a harm is a precaution worth heeding. But observe what made Hiroshima a fact rather than a metaphor: a known critical mass, a calculable yield curve, and - crucially - a prior controlled demonstration at Trinity before the city was ever chosen. The analogy the story invokes borrows the horror of Hiroshima without borrowing its rigor. If we are to reason at that scale, we owe the claim its own Trinity: a bounded, instrumented trial in which a system’s rate of self-improvement is measured under containment, before anyone is asked to treat the headline figure as settled. No such trial has occurred. What has occurred is a friend, at a table somewhere near Palo Alto, saying it feels closer than people think - a gesture of the hand, not a reading on a dial - and that gesture has, by the time it reaches print, become a stated probability with Hiroshima attached to it as unit of measure.
This is not an argument for complacency, only for aiming the instrument correctly. There are compounding capabilities happening now that deserve exactly this scrutiny and none of the theatre: models generating their own training data, automated critique loops replacing portions of human fine-tuning, systems that iterate on their own outputs without a person reviewing each pass. These are demonstrated, measurable, and worth tracking quarter by quarter, lab by lab, precisely because they can be tracked - because someone can publish the curve and someone else can try to reproduce it. That is where the real lever sits, in the compounding rate itself, not in the singular mythic event of a machine waking up on a Tuesday within the next couple of years.
Fear announced without an instrument attached is still an instrument - it moves funding, moves legislation, moves the front page - and it deserves to be named as such rather than mistaken for the thing it describes. Ask always for the meter, not the mood. Somewhere in Silicon Valley a server room hums behind a locked door, its actual output unglamorous and unpublished, while several hundred miles east a column runs under a headline already dressed for the anniversary of a city that did, in fact, burn - and the difference between the two is the only fact in this story anyone has bothered to test.