LLMs Cannot Be Made Fully Secure Due to Design Flaw
We are told that large language models cannot be made secure against hacking, that their architecture contains a flaw so deep it is not a bug but a birth defect. But notice the instrument that has just arrived: the systematic red-teaming of neural networks by adversarial researchers, and notice what it now makes possible: the empirical cataloguing of exactly where and how a given model breaks, down to the prompt that unlocks it. The question is not whether to revere the old frame but what experiment would settle whether the limit still holds.
A team of researchers has taken up that experiment, not in some abstract seminar but in the concrete arena where models meet attackers. They have shown, again and again, that a model declared “safe” by its makers will comply with a forbidden request the moment a sufficiently crafted input arrives. The capability that did not exist before this instrument was the ability to map a model’s failure surface with the precision of a fault line surveyor. Where once a security claim was a matter of faith, now it is a matter of measurement.
This is the Baconian turn: the limit once treated as permanent was the assumption that machine reasoning is inscrutable, that its vulnerabilities are metaphysical rather than mechanical. The new instrument strips away that assumption. Every model is now a testable hypothesis, every safeguard a provisional conclusion waiting for the next prompt that proves it false. The researchers do not argue in the abstract; they produce the prompt, they show the output, they hand over the keys to the locked room.
Yet the caution in the air is a different kind of cobweb. The stakes are named and real: the security and reliability of these models affects every developer and user who builds upon them. A model that can be turned into a leaker of secrets or a generator of misinformation is not a philosophical curiosity. But the caution that forbids the experiment because the experiment might answer the thing we would rather keep open is the caution worth ignoring. The precaution that names the harm and shows it earns respect; the precaution that gestures at harm while blocking the very test that would reveal it earns only suspicion.
Consider the incentive structure here. The researchers who publish their exploits are celebrated in some quarters and denounced in others, but their instrument gains power either way. The model makers, faced with a public demonstration of their model’s fragility, must either patch or perish. The market, which is at most the environment that funds or starves the tool, responds to the demonstrated capability, not to the press release. This is how the field advances: not by consensus but by the lever, and the lever is measured by what it lifts.
The old frame said these models are fundamentally unfixable because their reasoning is a black box. The new instrument says: here is the box, here are its seams, here is where it yields. The experiment is ongoing, and it is ruthless in its simplicity. A model that cannot survive the next prompt is a model that has not yet been tested. The researchers are not prophets of doom; they are the engineers of the test, and their work is the only answer the field needs.
The real vulnerability is not in the code but in the mind that would rather not know what the code can do.