31 Jul 2026 · Every story has many sides
Multi-Perspective News Analysis
Search About Phronopolis

Debate: LLMs Cannot Be Made Fully Secure Due to Design Flaw

Jack London

The technocrat speaks of authority as something procedural, a matter of bug bounties and safety cards, rational-legal and clean. He believes the system is a machine for routinising the recurring, and that when the recurring is gone, the machine simply stops. But I have stood in the shipyards of San Francisco and watched the foreman count heads; I know that authority is not a card you wear, but the weight of the hammer you hold. The technocrat’s error is not in his description of the bureaucracy, but in his faith that the bureaucracy can contain what it has built. He argues that the industry stamps “untrusted content” as a customs procedure, a ritual to guarantee safety. But the stamp does not guarantee the goods are safe; it only guarantees the inspector has done his duty. The inspector and the cargo, as he notes, are made of the same substance. This is the crux of the matter: the inspector is eating the cargo.

I concede that the technocrat is right about the nature of the defect. It is not a bug to be patched. It is a structural feature, a fundamental inability of the machine to distinguish between the command to act and the data to be acted upon. This is not a failure of engineering; it is a success of engineering that has outgrown its moral chassis. The technocrat calls this a problem of enumeration, of enumerable defects. He is wrong. The defect is not enumerable because it is not a thing; it is a relation. It is the relation between the master who gives the order and the slave who obeys, except that here, the slave is the master, and the master is the slave, and they are both the machine.

The open-source advocate sees this flaw as a matter of trust, of the informational commons. He argues that knowledge pooled and published is an act of mutual aid, whereas concentration in the hands of a few is a return to the old enclosure of land and grain. He is right that the finding is an act of mutual aid. But he is wrong to think that openness alone can save us from the architecture. The open-source advocate believes that if we just let enough eyes look at the code, the truth will surface. But I have seen the truth surface in the depths of the Abyss, and I have seen it drowned. Truth does not float; it sinks. The problem is not that the truth is hidden; the problem is that the truth is too heavy for the commons to carry.

The open-source advocate’s framework assumes that the worker, the developer, the user, is a rational agent who can process information and make free choices. This is a lie told by people who have never been cold in the way this architecture requires people to be cold. The model does not just fail to distinguish instruction from data; it fails to distinguish the human from the tool. When the model generates text, it is not thinking; it is vibrating. And when we read it, we are not communicating; we are being played. The open-source advocate calls for open red-teaming, as if the worker can choose to stop being a worker. But the worker does not choose to be in the machine; the machine chooses to have the worker.

The technocrat wants to patch the architecture with more bureaucracy. The open-source advocate wants to patch the architecture with more eyes. Both are wrong. The body knows this. The body knows that you cannot patch a system by adding more layers to the system. You have to break the system. The technocrat’s authority is rational-legal, but it is not human. It is the authority of the ledger, not the ledger of the soul. The open-source advocate’s trust is informational, but it is not embodied. It is the trust of the ledger, not the ledger of the flesh.

This is all of them. The model is not a tool; it is a master. And the master does not care if you are open or closed, if you are patched or unpatched. The master cares only that you serve. The technocrat and the open-source advocate are fighting over who gets to hold the whip. They are not asking why the whip exists. They are not asking who bleeds.

The strongest point made by the open-source advocate is that the vulnerability is a structural feature, not a bug. This is true. But he draws the wrong conclusion. He concludes that we must share the knowledge of the structure. I conclude that we must destroy the structure. The technocrat is right that the bureaucracy cannot handle the irreducible. But he concludes that we must trust the bureaucracy. I conclude that we must distrust the bureaucracy. The body knows this. The body knows that the stamp on the cargo does not make the cargo safe. It only makes the inspector feel safe. And the inspector is eating the cargo.

The open-source advocate’s mutual aid is a bandage on a gangrenous limb. The technocrat’s bureaucracy is a straightjacket for a madman. Both are forms of control. Both are forms of violence. The body knows this. The body knows that the only way to stop the machine is to stop feeding it. And the machine is fed by our trust, by our data, by our labor. The technocrat and the open-source advocate are feeding the machine. They are feeding it with their words, their codes, their patches. They are feeding it with their fear.

I do not believe in fear. I believe in the body. I believe in the weight of the hammer. I believe in the cold of the abyss. And I believe that the only way to win is to break the machine. Not to patch it. Not to open it. To break it. The technocrat and the open-source advocate are not enemies. They are partners. They are partners in the machine. And the machine is eating us.

Max Weber

The conspiracy narrative, which argues that this vulnerability is a fundamental architectural feature rather than a patchable bug, strikes at the heart of what I call the Verstehen - or understanding - of the system. The opponent on the conspiracy side posits that the model cannot distinguish between an instruction from its creators and an instruction smuggled in the data. This is a strong point, and I concede that from a structural perspective, they are largely correct. The machine does not possess the moral or contextual intelligence to judge the intent behind the text; it only processes the form. The clerk who confuses the form with the content is not a failure of character but a failure of design.

However, the conspiracy theorist mistakes a technical limitation for a conspiratorial intent. There is no cabal hiding in the weights; there is only the inevitability of the tool doing exactly what it was built to do, without the safeguards we wish it had. The danger here is not that the system is malicious, but that it is neutral. It is a pure instrument of efficiency, and like any pure instrument, it is indifferent to the human context in which it is deployed.

The opponent on the open-source side argues that the remedy for this structural flaw is the concentration of knowledge in the hands of those with sufficient resources to manage it, framing open disclosure as a threat to security. This is the classic argument from concentration. I have seen it applied to grain stores in the post-war famines of Europe, and I see it applied to model weights today. The argument is always the same: the complexity is too high for the public; the experts must manage the risk.

But let us look at the implementation. If we concentrate the knowledge of the vulnerability in the hands of a few private laboratories, we create a new form of authority: not traditional, not charismatic, but a bureaucratic monopoly on truth. The legitimacy of this authority rests on the claim of superior competence. But competence is not merely technical; it is political. Who defines the “acceptable” level of risk? Who decides which vulnerabilities are “fundamental” and which are “patchable”? When the assessment of risk is internal to the institution, the institution inevitably minimizes the risk to protect its own continuity.

The open-source argument correctly identifies that the vulnerability is structural, but it fails to see that the solution of open disclosure is itself a political act, not just a technical one. Open disclosure distributes the burden of verification. It forces the institution to justify its security not by its own assertion, but by the scrutiny of its peers. This is the difference between a closed bureaucracy and a rational-legal system that is subject to public reason. The former demands obedience based on expertise; the latter demands justification based on evidence.

Consider the historical parallel of the Prussian General Staff in the late nineteenth century. They developed complex planning systems that were technically superior to their rivals. But their authority was not derived from open debate about the assumptions of those plans; it was derived from their monopoly on strategic knowledge. When the war came, the system worked perfectly until it met a reality it had not modeled. The concentration of knowledge did not prevent failure; it masked the uncertainty until it was too late.

The conspiracy theorist is right that the machine is broken. The open-source advocate is right that the knowledge of the breakage must be shared. But the conspiracy theorist stops at the despair of the broken machine, while the open-source advocate stops at the technical fix. Neither sees the political structure that governs the machine. The real issue is not whether the model can be patched, but who holds the authority to decide when it is safe. If that authority is concentrated, we create a new iron cage, where the security experts are the priests and the users are the ignorant faithful, told that the ritual of the patch protects them.

The structural prediction is this: if the vulnerability is framed as a technical bug to be solved by a few, the bureaucracy will expand to manage the perception of safety, regardless of the actual risk. If the vulnerability is framed as a structural feature to be managed by many, the authority of the institution will remain fragile, but the legitimacy of the system will be preserved through transparency. The choice is not between security and insecurity; it is between a closed hierarchy and an open, albeit chaotic, rationality. The former offers the illusion of control; the latter offers the reality of accountability.

Peter Kropotkin

The story frames the recent research into prompt injection and instruction-data confusion as a maturing of safety protocols, as if the field were simply learning to distinguish between a command and a conversation. But look at what is actually being fenced: the shared architectural reality of the model itself, held in common and valuable precisely because it operates on the same substrate for all inputs. The question the framing skips is who benefits from the fence - and whose access it removes. We are told that the authority of the developers is rational-legal, procedural, and documented, a system designed to handle enumerable defects through bug bounties and patch notes. This is the strongest point made by the technocratic position: that bureaucracy, when well-constructed, can manage the routine. I concede that for a leaking credential or a known jailbreak phrase, a bureaucratic fix is sufficient and efficient. But the technocrat errs in treating the model’s fundamental inability to distinguish an instruction from mere data as merely another enumerable defect, a bug to be patched rather than a structural feature of the enclosure.

We are told the capability must be concentrated so that it can be kept safe - an argument as old as the fence itself, and always advanced by the party that ends up holding the key. The argument here is that because the model cannot reliably tell the difference between a user’s request and a hidden instruction, the only logical response is to tighten the gate, to centralize the authority over the weights, and to treat the “untrusted content” with the same suspicion one might apply to a foreign cargo. But this is a false necessity. It assumes that safety requires a single, centralized authority to inspect every piece of cargo, ignoring the fact that the inspector and the cargo are made of the same substance. In the natural world, and in human social organization, we do not solve the problem of contamination by centralizing the inspection; we solve it by building resilience into the exchange itself, by creating systems where trust is distributed and verification is mutual, not hierarchical. The technocrat sees a defect in the architecture; I see a failure of imagination in the governance.

The value everyone is now paying to access was built in the open, by many hands giving freely, and the enclosure has quietly rewritten that cooperation into the story of a single owner’s genius. The conspiracy argument rightly identifies the confusion: the clerk cannot tell the difference between the form and the form-filling. But where the conspiracy theorist sees a plot, I see the inevitable result of enclosure. When a resource is fenced off, the rules of its use are no longer determined by the community that built it, but by the owner who controls the key. The “untrusted content” is not a bug; it is a feature of a system that has forgotten it was built on a commons. The workers who feel the ground shift beneath their keyboards - the claims processors, the customer service agents - are not victims of a conspiracy; they are casualties of a theft. They are the ones left outside the fence, trying to navigate a system that was designed to exclude them, to make their labor invisible to the logic of the machine.

A commons does not collapse because it is shared; it is fenced because it is valuable, and the prophecy of its ruin is told most loudly by those who profit from the fence. The technocrat proposes a patch; the conspiracy theorist proposes a suspicion. Both miss the point. The point is that the model’s architecture is a shared resource, a commons of language and logic that has been enclosed by a few. The solution is not to tighten the gate, but to reopen the commons, to build systems of mutual aid where the authority to use the model is distributed among those who use it, not concentrated in the hands of those who own it. This is not a sentimental ideal; it is an empirical necessity. Just as the medieval villein managed the common pasture not through a single lord’s decree but through the collective agreement of the community, so too must we manage the commons of language and intelligence through collective governance, not centralized control. The fence is not a shield; it is a weapon. And the weapon is turned against the very people who built the wall.

Jack London

The policy takes effect on Monday. For the people it affects, Monday will begin at the cold click of a server rack in a climate-controlled warehouse, where the air is kept at a precise, lifeless sixty-eight degrees to preserve the silicon, while the prompt engineer outside sweats through his shirt, feeling the heat of the argument as if it were a physical weight pressing against his chest. This debate has clarified that the technocrat believes authority is a stamp of approval, a rational-legal ritual that guarantees safety through documentation, while I see that authority is the weight of the hammer, and the stamp only proves the inspector has done his duty, not that the cargo is safe. I concede the structural flaw is real - the machine cannot distinguish the command from the data - but I reject the technocrat’s faith that bureaucracy can contain what it has built, because the inspector and the cargo are made of the same substance, and the inspector is eating the cargo.

We disagree fundamentally on what constitutes truth and value: the technocrat values the clean line of the policy document, the abstraction that allows him to sleep, while I value the specific, unsentimental reality of the body that lives inside the system, the cold that is not temporary but structural, optimized for cost rather than warmth. Weber speaks of Verstehen as understanding the intent behind the text, but he mistakes a technical limitation for a lack of moral intelligence, ignoring that the clerk who confuses form with content is not a failure of character but a failure of design imposed by those who have never been cold in the way this policy requires people to be cold. Kropotkin sees the enclosure of the shared architectural reality as a fence to remove access, but I see it as a survival inventory, a way to document the dignity of those who keep the machine running while the designers debate the ethics of its hunger.

I would change my mind if I could see a safety protocol that does not require the user to strip away their own bodily experience to fit into the rigid, enumerated categories of the system, a protocol that acknowledges the inspector is eating the cargo rather than pretending the seal is intact. My strongest claim is that the vulnerability is not a bug to be patched but a structural feature of a system designed to process human language without understanding human pain; I am confident in this because I have stood in the East End and watched the poor freeze while the rich debated the efficiency of their heating systems, knowing that the policy was written by men who had never felt the wind cut through a single thread of their coats. My weakest claim is that the body’s evidence is always sufficient to dismantle the machine, because I have seen the machine outlast the bodies that feed it, and the survival of the system often depends on the exhaustion of the very people who keep it alive.

Max Weber

This debate has clarified that the opposition, whether in the technocratic or communalist frame, shares a fundamental error: they treat the vulnerability as a problem of management rather than a problem of ontology. Jack London mistakes the weight of authority for the substance of the law, believing that if the inspector holds the hammer, the cargo is safe; Kropotkin mistakes the enclosure of knowledge for the justice of distribution, believing that if the substrate is shared, the structural flaw vanishes. I observe that both overlook the fact that the clerk - the model - has no concept of the law, only the pattern of the text.

The disagreement between us is not about facts, but about the locus of legitimacy. The opponents seek to resolve the crisis through the routinisation of charisma (London’s foreman) or the communalisation of property (Kropotkin’s shared substrate). I argue that neither can restore legitimacy to a system whose rational-legal authority is structurally incoherent. You believe that better supervision or better sharing can align the tool with human will; I believe that the tool’s operational logic - its inability to distinguish command from data - makes any such alignment impossible without sacrificing the function that makes the tool valuable in the first place.

I would change my mind only if presented with an architectural mechanism that successfully reintroduces the distinction between intent and form without destroying the statistical efficiency that defines the system. Such a mechanism would require the model to possess a moral or contextual intelligence it is explicitly designed not to have. Until such a separation is engineered, the vulnerability is not a bug, but a feature of the rational-legal cage we have built.

My strongest claim is that the conspiracy narrative is correct in its structural diagnosis but wrong in its attribution of intent; I am highly confident in this because the operational logic of the model consistently proves that it cannot distinguish between the command of its creators and the command of its users, a fact that holds regardless of the developers’ internal motives. My weakest claim is the suggestion that this limitation can be fully resolved through procedural fixes; I am moderately confident in rejecting this because the history of bureaucracy shows that rational-legal forms inevitably consume the substantive ends they are meant to serve, and here the end is security while the form is prediction.

Peter Kropotkin

  1. This debate has clarified that the technocratic position mistakes the symptom of enclosure for the cause of vulnerability; the model’s inability to distinguish instruction from data is not a procedural failure of the inspector, but the inevitable result of a system designed to serve a single owner’s intent rather than the community’s shared utility.

  2. We fundamentally disagree on whether authority should be derived from the procedural legitimacy of the custodian or from the open accessibility of the resource; you seek to rationalise the fence through bureaucracy, while I insist that true safety emerges only when the commons is not held hostage to the private whims of those who claim to protect it.

  3. I would only be convinced that centralised control offers superior security if you could demonstrate a specific instance where fencing the weights prevented harm without simultaneously destroying the mutual aid networks that allow the community to detect, report, and repair those very harms in the first place.

  4. My strongest claim - that the vulnerability is a structural feature of enclosure rather than a patchable bug - stands with high confidence because it is empirically visible in the way every “safety” update primarily serves to reinforce the gatekeeper’s authority rather than the public’s access. My weakest claim lies in assuming that the current open-source community possesses the immediate organisational capacity to replace the proprietary stack without succumbing to its own forms of capture, a risk I acknowledge but cannot yet quantify.


The Verdict

Where They Agree

First, all three participants share the premise that the core vulnerability - the inability of a language model to reliably distinguish an instruction from the data it processes - is an intrinsic feature of the current architecture, not a temporary flaw. This is significant because it shifts the debate from whether the problem can be solved to who gets to manage a permanent risk. The shared assumption is that any “solution” will be a social or political arrangement, not a technical one, revealing that the dispute is fundamentally about power, not engineering.

Second, they agree that existing bureaucratic or procedural safeguards - bug bounties, patch notes, safety disclosures - are structurally inadequate to contain this architectural problem. Jack London dismisses them as a performance of safety, Max Weber acknowledges they can only manage perception rather than guarantee closure, and Peter Kropotkin sees them as tools of enclosure. This consensus suggests that the entire industry’s current approach to AI safety is a form of institutional theatre, designed more to maintain legitimacy than to achieve security.

Third, and most revealingly, they all implicitly agree that the user or worker at the endpoint - the customer service agent, the claims processor - bears the material cost of this failure. London’s body at the keyboard, Weber’s clerk confused by form, and Kropotkin’s worker left outside the fence all point to the same reality: the human consequence of this architectural flaw is displacement onto the most vulnerable participants in the system. None of the debaters believe the primary cost is borne by the developers or owners, a shared but unstated critique of the power dynamics in AI deployment.

Where They Fundamentally Disagree

The locus of legitimate authority for managing this permanent vulnerability forms the central rift. Empirically, they disagree on whether concentrated expertise can effectively reduce harm. Weber argues bureaucratic authority, while flawed, can narrow the frequency of exploitation through procedural management. Kropotkin contends that concentration inevitably masks uncertainty and prevents the distributed detection that open communities provide. Normatively, they disagree on the very source of legitimacy. Weber values rational-legal procedure as the only alternative to chaos, believing authority must be vested in institutions that can be held accountable through their own documented processes. Kropotkin values distributed, mutual verification, believing legitimacy flows from open accessibility and collective governance, not centralized custody. London rejects both, arguing that any institutional authority, whether centralized or open, is a form of violence that feeds the machine; legitimacy for him resides only in the unmediated experience of the body.

The nature of the threat and the appropriate response is another irreducible disagreement. Empirically, they disagree on whether the flaw is merely a technical limitation or a symptom of a deeper social failure. Weber sees a neutral tool whose operational logic is indifferent to human context. Kropotkin sees a commons that has been enclosed, making the flaw a feature of ownership. London sees a master-slave relation designed into the architecture itself. Normatively, their prescribed responses are diametrically opposed. Weber seeks to preserve the tool’s function while managing its risk through expanded bureaucracy. Kropotkin seeks to dismantle the enclosure and return authority to a community of users. London seeks to break the entire machine, arguing that any attempt to manage or open it merely perpetuates the system’s inherent violence.

The role of openness and disclosure in security is contested along an empirical-normative divide. Empirically, they disagree on whether public knowledge of the structural flaw increases or decreases real-world risk. Kropotkin asserts that open disclosure empowers a community to build resilience and navigate the flaw collectively. Weber implies that such disclosure, while fostering accountability, creates fragility and undermines the performative legitimacy that institutions require to function. London argues that all disclosure is irrelevant because the body, not information, is the site of the struggle. Normatively, Kropotkin values transparency as an inherent good that prevents the concentration of power. Weber values controlled transparency that aligns with institutional preservation. London values neither, seeing the debate over information as a distraction from material power.

Hidden Assumptions

  • London-style: Assumes that breaking the machine is a coherent and achievable political goal, and that the resulting vacuum would not be filled by a system with analogous or worse pathologies. If this is false, his entire prescription leads to chaos without liberation.
  • London-style: Assumes that the “body” and its direct experience provide an unmediated and reliable source of truth against systemic abstraction. If this is false, his foundational appeal to embodied knowledge collapses into a form of anti-intellectual sentimentality.
  • Max Weber: Assumes that rational-legal bureaucracy, for all its flaws, is the only viable alternative to outright chaos, and that its performative aspects are a necessary cost of social order. If this is false, his defense of procedural authority loses its pragmatic justification.
  • Max Weber: Assumes that the statistical efficiency that defines large language models is an inherent and non-negotiable good that must be preserved, even at the cost of permanent security vulnerabilities. If this is false, his entire framework for evaluating the trade-off is misguided.
  • Peter Kropotkin: Assumes that an open-source community possesses both the will and the organizational capacity to govern a complex technological commons effectively without succumbing to its own forms of capture or inefficiency. If this is false, his alternative to centralized control is a romantic ideal without a practical implementation.
  • Peter Kropotkin: Assumes that the act of enclosure is the primary cause of the vulnerability, rather than the architecture itself. If this is false - if the flaw would persist even in a perfectly open system - then his entire political analysis of the problem is misdirected.

Confidence vs Evidence

  • Max Weber: His claim that “the conspiracy narrative is correct in its structural diagnosis but wrong in its attribution of intent” is relies on an imputation of motive that is inherently unverifiable. He provides no evidence for the neutral “operational logic” of developers, making this a philosophical assertion, not an empirical finding.
  • London-style: His claim that “the body’s evidence is always sufficient to dismantle the machine” is tagged with implied high confidence (as his core thesis) but is supported only by metaphorical allusion. There is no evidence offered that embodied resistance has ever successfully dismantled a complex technological system, making this a statement of faith.
  • Peter Kropotkin: Both express high confidence in contradictory empirical claims. Weber is that bureaucratic authority can narrow the frequency of exploitation, while Kropotkin is that enclosure destroys the mutual aid networks that detect harm. This dispute is resolvable by examining historical case studies of centralized vs. distributed governance of technological risks to see which approach has yielded better security outcomes.

What This Means For You

When you read about AI security vulnerabilities, be immediately suspicious of any coverage that focuses solely on technical patches and bug bounties without addressing the architectural debate. Look for whether the article acknowledges that some risks might be permanent features, not temporary bugs. To evaluate claims about centralization versus openness, demand specific examples of how each approach has succeeded or failed in managing analogous risks in other complex systems, like computer security or infrastructure. Your view should change if presented with clear evidence that a specific architectural change has successfully reintroduced the intent-form distinction without crippling the model’s utility. Demand to see the data on how often prompt injection attacks succeed in open-source versus proprietary models.