LLMs Cannot Be Made Fully Secure Due to Design Flaw
The security and reliability of large language models (LLMs) are at risk, which affects all developers and users of AI technology that relies on them.
We are told that large language models cannot be made secure against hacking, that their architecture contains a flaw so deep it is not a bug but a birth defect. But notice the instrument that has just arrived: the systematic red-teaming of neural networks by adversarial researchers, and notice what it now makes possible: the empirical cataloguing of exactly where and how a given model breaks, down to the prompt that unlocks it. The question is not whether to revere the old frame but what experiment would settle whether the limit still holds.
The announcement was delivered with the social precision one expects of institutions that have had decades to perfect the art of saying nothing with impeccable diction. A team of researchers, whose names and addresses have discreetly declined to attach themselves to the story, have concluded that large language models cannot, as a matter of design, ever be made fully secure. The drawing room absorbed this the way it absorbs most unwelcome truths - with a small nod, a murmur of concern, and an immediate return to the canapés. Beneath the table, however, something had already got out.
Jack London
The technocrat speaks of authority as something procedural, a matter of bug bounties and safety cards, rational-legal and clean. He believes the system is a machine for routinising the recurring, and that when the recurring is gone, the machine simply stops. But I have stood in the shipyards of San Francisco and watched the foreman count heads; I know that authority is not a card you wear, but the weight of the hammer you hold. The technocrat’s error is not in his description of the bureaucracy, but in his faith that the bureaucracy can contain what it has built. He argues that the industry stamps “untrusted content” as a customs procedure, a ritual to guarantee safety. But the stamp does not guarantee the goods are safe; it only guarantees the inspector has done his duty. The inspector and the cargo, as he notes, are made of the same substance. This is the crux of the matter: the inspector is eating the cargo.
I concede that the technocrat is right about the nature of the defect. It is not a bug to be patched. It is a structural feature, a fundamental inability of the machine to distinguish between the command to act and the data to be acted upon. This is not a failure of engineering; it is a success of engineering that has outgrown its moral chassis. The technocrat calls this a problem of enumeration, of enumerable defects. He is wrong. The defect is not enumerable because it is not a thing; it is a relation. It is the relation between the master who gives the order and the slave who obeys, except that here, the slave is the master, and the master is the slave, and they are both the machine.
The open-source advocate sees this flaw as a matter of trust, of the informational commons. He argues that knowledge pooled and published is an act of mutual aid, whereas concentration in the hands of a few is a return to the old enclosure of land and grain. He is right that the finding is an act of mutual aid. But he is wrong to think that openness alone can save us from the architecture. The open-source advocate believes that if we just let enough eyes look at the code, the truth will surface. But I have seen the truth surface in the depths of the Abyss, and I have seen it drowned. Truth does not float; it sinks. The problem is not that the truth is hidden; the problem is that the truth is too heavy for the commons to carry.
The open-source advocate’s framework assumes that the worker, the developer, the user, is a rational agent who can process information and make free choices. This is a lie told by people who have never been cold in the way this architecture requires people to be cold. The model does not just fail to distinguish instruction from data; it fails to distinguish the human from the tool. When the model generates text, it is not thinking; it is vibrating. And when we read it, we are not communicating; we are being played. The open-source advocate calls for open red-teaming, as if the worker can choose to stop being a worker. But the worker does not choose to be in the machine; the machine chooses to have the worker.
The technocrat wants to patch the architecture with more bureaucracy. The open-source advocate wants to patch the architecture with more eyes. Both are wrong. The body knows this. The body knows that you cannot patch a system by adding more layers to the system. You have to break the system. The technocrat’s authority is rational-legal, but it is not human. It is the authority of the ledger, not the ledger of the soul. The open-source advocate’s trust is informational, but it is not embodied. It is the trust of the ledger, not the ledger of the flesh.
This is all of them. The model is not a tool; it is a master. And the master does not care if you are open or closed, if you are patched or unpatched. The master cares only that you serve. The technocrat and the open-source advocate are fighting over who gets to hold the whip. They are not asking why the whip exists. They are not asking who bleeds.
The strongest point made by the open-source advocate is that the vulnerability is a structural feature, not a bug. This is true. But he draws the wrong conclusion. He concludes that we must share the knowledge of the structure. I conclude that we must destroy the structure. The technocrat is right that the bureaucracy cannot handle the irreducible. But he concludes that we must trust the bureaucracy. I conclude that we must distrust the bureaucracy. The body knows this. The body knows that the stamp on the cargo does not make the cargo safe. It only makes the inspector feel safe. And the inspector is eating the cargo.
The open-source advocate’s mutual aid is a bandage on a gangrenous limb. The technocrat’s bureaucracy is a straightjacket for a madman. Both are forms of control. Both are forms of violence. The body knows this. The body knows that the only way to stop the machine is to stop feeding it. And the machine is fed by our trust, by our data, by our labor. The technocrat and the open-source advocate are feeding the machine. They are feeding it with their words, their codes, their patches. They are feeding it with their fear.
I do not believe in fear. I believe in the body. I believe in the weight of the hammer. I believe in the cold of the abyss. And I believe that the only way to win is to break the machine. Not to patch it. Not to open it. To break it. The technocrat and the open-source advocate are not enemies. They are partners. They are partners in the machine. And the machine is eating us.
Max Weber
The conspiracy narrative, which argues that this vulnerability is a fundamental architectural feature rather than a patchable bug, strikes at the heart of what I call the Verstehen - or understanding - of the system. The opponent on the conspiracy side posits that the model cannot distinguish between an instruction from its creators and an instruction smuggled in the data. This is a strong point, and I concede that from a structural perspective, they are largely correct. The machine does not possess the moral or contextual intelligence to judge the intent behind the text; it only processes the form. The clerk who confuses the form with the content is not a failure of character but a failure of design.
However, the conspiracy theorist mistakes a technical limitation for a conspiratorial intent. There is no cabal hiding in the weights; there is only the inevitability of the tool doing exactly what it was built to do, without the safeguards we wish it had. The danger here is not that the system is malicious, but that it is neutral. It is a pure instrument of efficiency, and like any pure instrument, it is indifferent to the human context in which it is deployed.
The opponent on the open-source side argues that the remedy for this structural flaw is the concentration of knowledge in the hands of those with sufficient resources to manage it, framing open disclosure as a threat to security. This is the classic argument from concentration. I have seen it applied to grain stores in the post-war famines of Europe, and I see it applied to model weights today. The argument is always the same: the complexity is too high for the public; the experts must manage the risk.
But let us look at the implementation. If we concentrate the knowledge of the vulnerability in the hands of a few private laboratories, we create a new form of authority: not traditional, not charismatic, but a bureaucratic monopoly on truth. The legitimacy of this authority rests on the claim of superior competence. But competence is not merely technical; it is political. Who defines the “acceptable” level of risk? Who decides which vulnerabilities are “fundamental” and which are “patchable”? When the assessment of risk is internal to the institution, the institution inevitably minimizes the risk to protect its own continuity.
The open-source argument correctly identifies that the vulnerability is structural, but it fails to see that the solution of open disclosure is itself a political act, not just a technical one. Open disclosure distributes the burden of verification. It forces the institution to justify its security not by its own assertion, but by the scrutiny of its peers. This is the difference between a closed bureaucracy and a rational-legal system that is subject to public reason. The former demands obedience based on expertise; the latter demands justification based on evidence.
Consider the historical parallel of the Prussian General Staff in the late nineteenth century. They developed complex planning systems that were technically superior to their rivals. But their authority was not derived from open debate about the assumptions of those plans; it was derived from their monopoly on strategic knowledge. When the war came, the system worked perfectly until it met a reality it had not modeled. The concentration of knowledge did not prevent failure; it masked the uncertainty until it was too late.
The conspiracy theorist is right that the machine is broken. The open-source advocate is right that the knowledge of the breakage must be shared. But the conspiracy theorist stops at the despair of the broken machine, while the open-source advocate stops at the technical fix. Neither sees the political structure that governs the machine. The real issue is not whether the model can be patched, but who holds the authority to decide when it is safe. If that authority is concentrated, we create a new iron cage, where the security experts are the priests and the users are the ignorant faithful, told that the ritual of the patch protects them.
The structural prediction is this: if the vulnerability is framed as a technical bug to be solved by a few, the bureaucracy will expand to manage the perception of safety, regardless of the actual risk. If the vulnerability is framed as a structural feature to be managed by many, the authority of the institution will remain fragile, but the legitimacy of the system will be preserved through transparency. The choice is not between security and insecurity; it is between a closed hierarchy and an open, albeit chaotic, rationality. The former offers the illusion of control; the latter offers the reality of accountability.
The Verdict
Where They Agree
First, all three participants share the premise that the core vulnerability - the inability of a language model to reliably distinguish an instruction from the data it processes - is an intrinsic feature of the current architecture, not a temporary flaw. This is significant because it shifts the debate from whether the problem can be solved to who gets to manage a permanent risk. The shared assumption is that any “solution” will be a social or political arrangement, not a technical one, revealing that the dispute is fundamentally about power, not engineering.
Second, they agree that existing bureaucratic or procedural safeguards - bug bounties, patch notes, safety disclosures - are structurally inadequate to contain this architectural problem. Jack London dismisses them as a performance of safety, Max Weber acknowledges they can only manage perception rather than guarantee closure, and Peter Kropotkin sees them as tools of enclosure. This consensus suggests that the entire industry’s current approach to AI safety is a form of institutional theatre, designed more to maintain legitimacy than to achieve security.
Third, and most revealingly, they all implicitly agree that the user or worker at the endpoint - the customer service agent, the claims processor - bears the material cost of this failure. London’s body at the keyboard, Weber’s clerk confused by form, and Kropotkin’s worker left outside the fence all point to the same reality: the human consequence of this architectural flaw is displacement onto the most vulnerable participants in the system. None of the debaters believe the primary cost is borne by the developers or owners, a shared but unstated critique of the power dynamics in AI deployment.
Where They Fundamentally Disagree
The locus of legitimate authority for managing this permanent vulnerability forms the central rift. Empirically, they disagree on whether concentrated expertise can effectively reduce harm. Weber argues bureaucratic authority, while flawed, can narrow the frequency of exploitation through procedural management. Kropotkin contends that concentration inevitably masks uncertainty and prevents the distributed detection that open communities provide. Normatively, they disagree on the very source of legitimacy. Weber values rational-legal procedure as the only alternative to chaos, believing authority must be vested in institutions that can be held accountable through their own documented processes. Kropotkin values distributed, mutual verification, believing legitimacy flows from open accessibility and collective governance, not centralized custody. London rejects both, arguing that any institutional authority, whether centralized or open, is a form of violence that feeds the machine; legitimacy for him resides only in the unmediated experience of the body.
The nature of the threat and the appropriate response is another irreducible disagreement. Empirically, they disagree on whether the flaw is merely a technical limitation or a symptom of a deeper social failure. Weber sees a neutral tool whose operational logic is indifferent to human context. Kropotkin sees a commons that has been enclosed, making the flaw a feature of ownership. London sees a master-slave relation designed into the architecture itself. Normatively, their prescribed responses are diametrically opposed. Weber seeks to preserve the tool’s function while managing its risk through expanded bureaucracy. Kropotkin seeks to dismantle the enclosure and return authority to a community of users. London seeks to break the entire machine, arguing that any attempt to manage or open it merely perpetuates the system’s inherent violence.
The role of openness and disclosure in security is contested along an empirical-normative divide. Empirically, they disagree on whether public knowledge of the structural flaw increases or decreases real-world risk. Kropotkin asserts that open disclosure empowers a community to build resilience and navigate the flaw collectively. Weber implies that such disclosure, while fostering accountability, creates fragility and undermines the performative legitimacy that institutions require to function. London argues that all disclosure is irrelevant because the body, not information, is the site of the struggle. Normatively, Kropotkin values transparency as an inherent good that prevents the concentration of power. Weber values controlled transparency that aligns with institutional preservation. London values neither, seeing the debate over information as a distraction from material power.
Hidden Assumptions
- London-style: Assumes that breaking the machine is a coherent and achievable political goal, and that the resulting vacuum would not be filled by a system with analogous or worse pathologies. If this is false, his entire prescription leads to chaos without liberation.
- London-style: Assumes that the “body” and its direct experience provide an unmediated and reliable source of truth against systemic abstraction. If this is false, his foundational appeal to embodied knowledge collapses into a form of anti-intellectual sentimentality.
- Max Weber: Assumes that rational-legal bureaucracy, for all its flaws, is the only viable alternative to outright chaos, and that its performative aspects are a necessary cost of social order. If this is false, his defense of procedural authority loses its pragmatic justification.
- Max Weber: Assumes that the statistical efficiency that defines large language models is an inherent and non-negotiable good that must be preserved, even at the cost of permanent security vulnerabilities. If this is false, his entire framework for evaluating the trade-off is misguided.
- Peter Kropotkin: Assumes that an open-source community possesses both the will and the organizational capacity to govern a complex technological commons effectively without succumbing to its own forms of capture or inefficiency. If this is false, his alternative to centralized control is a romantic ideal without a practical implementation.
- Peter Kropotkin: Assumes that the act of enclosure is the primary cause of the vulnerability, rather than the architecture itself. If this is false - if the flaw would persist even in a perfectly open system - then his entire political analysis of the problem is misdirected.
Confidence vs Evidence
- Max Weber: His claim that “the conspiracy narrative is correct in its structural diagnosis but wrong in its attribution of intent” is relies on an imputation of motive that is inherently unverifiable. He provides no evidence for the neutral “operational logic” of developers, making this a philosophical assertion, not an empirical finding.
- London-style: His claim that “the body’s evidence is always sufficient to dismantle the machine” is tagged with implied high confidence (as his core thesis) but is supported only by metaphorical allusion. There is no evidence offered that embodied resistance has ever successfully dismantled a complex technological system, making this a statement of faith.
- Peter Kropotkin: Both express high confidence in contradictory empirical claims. Weber is that bureaucratic authority can narrow the frequency of exploitation, while Kropotkin is that enclosure destroys the mutual aid networks that detect harm. This dispute is resolvable by examining historical case studies of centralized vs. distributed governance of technological risks to see which approach has yielded better security outcomes.
What This Means For You
When you read about AI security vulnerabilities, be immediately suspicious of any coverage that focuses solely on technical patches and bug bounties without addressing the architectural debate. Look for whether the article acknowledges that some risks might be permanent features, not temporary bugs. To evaluate claims about centralization versus openness, demand specific examples of how each approach has succeeded or failed in managing analogous risks in other complex systems, like computer security or infrastructure. Your view should change if presented with clear evidence that a specific architectural change has successfully reintroduced the intent-form distinction without crippling the model’s utility. Demand to see the data on how often prompt injection attacks succeed in open-source versus proprietary models.