Uncontrollable AI threatens global disaster
There is a very particular process by which a warning about the end of the human race reaches your breakfast table, and it is worth examining before the toast goes cold, because the process is doing something rather peculiar to the warning along the way. An expert in Silicon Valley says something to a friend at Stanford. The friend, alarmed, says something to Timothy Garton Ash, who is a journalist and public intellectual whose entire professional function is the responsible transmission of things worth worrying about. Garton Ash then says something to the rest of us. Nobody in this chain has done anything wrong. Nobody has lied. And yet by the time the warning arrives at the far end of it, it has acquired a very specific shape that none of the individual participants would necessarily endorse if you laid the whole transmission out in front of them and asked them to sign off on it as a document.
The shape it acquires is this: recursive self-improvement, imminent, within the next couple of years. That phrase - a couple of years - has been attached to catastrophic artificial intelligence at fairly regular intervals for well over a decade now, without the couple of years in question ever quite arriving in the way advertised, which suggests the phrase is not really functioning as a forecast at all. It is functioning as a genre convention, the way “sources say” or “growing concern” function in journalism: a form of words that signals urgency without committing anyone to being provably wrong later, because a couple of years is short enough to justify worrying about it now and vague enough to be renewed indefinitely without anyone noticing the renewal.
This is the committee problem in its purest form, and it is worth being precise about where the stupidity actually enters the system, because it does not enter through any individual. The Silicon Valley expert genuinely believes in the thing he is describing. The Stanford friend was genuinely alarmed. Garton Ash is genuinely and responsibly passing along a serious concern from serious people. What nobody in the chain controls is that each hand it passes through selects, not for the most accurate version of the story, but for the most retellable one - and the most retellable version of an AI warning is the one that borrows its shape from a disaster the shape of which is already settled in everyone’s mind, which is why it comes out the far end compared to Hiroshima.
Hiroshima had a date, a device, a shadow burned into a wall, a photograph everyone alive has seen. It is the single most legible catastrophe of the last century, produced, as it happens, by a committee of a rather different character, one that at least knew exactly what it had built and precisely when it had built it. Recursive self-improvement, whatever it eventually turns out to be, will not announce itself with a flash. It will arrive, if it arrives, as a gradual reweighting of probabilities inside a data centre that nobody outside the building is cleared to enter, which is a much harder thing to be afraid of in an organised way, and is exactly why it needs to borrow Hiroshima’s clarity rather than supply its own. You do not reach for someone else’s catastrophe unless yours is still a little short on details.
And here is the part of the mechanism that the warning itself cannot see, because it is standing inside it: the rooms in Silicon Valley producing this alarm are, very often, the same rooms trying to raise the next round of funding for the technology the alarm is about. Announcing that your product might end civilisation is, among the many things it accomplishes, also the single most effective advertisement available to a company whose valuation depends on everyone believing the product is powerful enough to be dangerous. Nobody involved considers this arrangement strange. That, more than the self-improvement, is the alarming part.
Garton Ash, to his credit, has no product in the building. He has only a story to tell well, at dinner, to people who will repeat it - which is precisely why it survives the retelling with less polish and more honesty than anything issued from Palo Alto with a press date attached, and why the couple of years keeps arriving, right on schedule, without ever quite getting here.