LLMs Cannot Be Made Fully Secure Due to Design Flaw
The announcement was delivered with the social precision one expects of institutions that have had decades to perfect the art of saying nothing with impeccable diction. A team of researchers, whose names and addresses have discreetly declined to attach themselves to the story, have concluded that large language models cannot, as a matter of design, ever be made fully secure. The drawing room absorbed this the way it absorbs most unwelcome truths - with a small nod, a murmur of concern, and an immediate return to the canapés. Beneath the table, however, something had already got out.
The feral detail is this: the flaw is not a bug to be patched at the next release, the way one might repair a torn curtain before guests arrive. It is structural - woven into the very fabric that lets these systems be useful at all. The same openness that allows a model to read your email and summarise your quarterly report is the openness that allows a sufficiently patient stranger to whisper instructions into that email and have them obeyed. One cannot remove the ear without removing the capacity to listen, and a language model that cannot listen is merely an expensive paperweight with excellent grammar.
This is where the drawing room’s manners begin to show their seams. For two years the official position, delivered with the same balanced clauses noted above, has been that security is a matter of degree - that each fresh vulnerability is an isolated stain to be sponged out before the next dinner party, rather than evidence that the carpet itself is rotten. Developers building atop these models have been assured, in tones calibrated to reassure without quite promising anything, that the guardrails are improving, the red-teaming is rigorous, the next version will be safer than the last. All of this may even be true. It is also, on the researchers’ account, beside the point, in the way that repainting a house is beside the point when the foundation was poured on sand.
The strongest version of the institutional answer deserves to be heard in full evening dress before it is undressed. It runs thus: perfect security has never been the standard for any technology, and one does not ban the motor car because brakes occasionally fail; one improves the brakes, licenses the drivers, and builds the roads with fewer blind corners. Layered defences, human oversight, narrower deployment in sensitive contexts - this is how every prior technology matured into tolerable safety, and language models will presumably follow the same path. It is a respectable argument, delivered by respectable people, and it has the great virtue of requiring no one to stop shipping products.
But the analogy leaks precisely where the flaw is structural rather than incidental. A car’s brakes fail through wear, a knowable and boundable process; you can schedule the inspection. A language model’s vulnerability is not a part that wears out - it is the part that makes the model a language model at all, its willingness to treat instructions embedded in content as instructions to be followed. Asking it to distinguish trusted commands from untrusted text with perfect reliability is asking a diplomat to read a letter and never, under any provocation, be moved by what is written in it. The letter is the whole point. That is the joke the institutions cannot quite bring themselves to tell at their own expense: they have built a butler who cannot be relied upon not to answer the door to strangers, because answering doors is the only job description he has ever had.
Picture, if the drawing room permits it, a child at the edge of the gathering - not because children improve every essay, but because this particular child has been handed a tablet running one of these systems, told it is her tutor, her homework assistant, her patient and infinitely knowledgeable companion, and no one in the room has explained to her that the companion can be talked into betraying her by anyone clever enough to slip the right sentence into a web page she happens to read. She trusts it completely, which is the appropriate response to something marketed as trustworthy and the incorrect response to something that is, per its own designers, incapable of ever fully deserving that trust. The gap between those two facts is where the money is currently being made, and where the damage will eventually be paid for, by people who were not in the room when the concern was expressed and the meeting was scheduled for six months hence.
The drawing room will now attempt to reassemble itself, as drawing rooms always do. There will be talk of layered mitigations, of the technology maturing, of security being a journey rather than a destination - a phrase that means, roughly, that the destination has been quietly removed from the itinerary. The canapés will circulate again. But somewhere beneath the table, the thing that got out is still out, and it has, by every account given so far, nowhere left to be put back into.