4 Sep 2026 · Every story has many sides
Multi-Perspective News Analysis
Search About Phronopolis
Stories / 4 Sep 2026

Microsoft Copilot rarely copies full sentences, firm claims

4 September 2026 sig 8/10

This matters for copyright law and the financial interests of publishers and authors, who claim AI reproduction substitutes for their original work and could affect their revenue.

Microsoft Copilot rarely copies full sentences, firm claimsSubmerged monolithic vault in Deep Navy and Teal, walls built from compressed, translucent text-striations glowing with proprietary luminescence. Central impenetrable glass sphere refracts chaotic data into a sterile beam of Foam White light. Foreground: turbulent dark currents crashing against the sphere. Background: abyssal ordered grid. Render with volumetric fog for water weight and subsurface scattering for the glass, evoking cold, impenetrable control. Palette: Deep Navy, Teal, Foam White, Steel Blue.
ACCELERATIONIST
bacon

Microsoft’s Copilot is not new because it writes; it is new because it can be interrogated at scale - a single operator can put the same query to it ten thousand times, in ten thousand phrasings, and read back, statistically, what the machine actually does with a copyrighted sentence rather than what its makers say it does. That is the instrument this dispute has actually produced, though nobody has yet built it properly: a repeatable audit of retrieval behaviour. The limit the story treats as fixed is the word “rarely,” offered in Microsoft’s own filings as though it were a finding rather than a claim. I ask what experiment stands behind it, and I find none disclosed. A number withheld is not evidence; it is an idol dressed as one.

Read full perspective →
AI SAFETY
shelley

The story celebrates that Copilot was built - the feat of a machine that can read the accumulated sentences of the world and answer in kind. Microsoft’s own filings now frame this achievement in the negative: the system rarely reproduces full sentences from news articles and books. But a made thing does not stop where its maker’s assurances stop; it goes on acting in a world no legal brief contains. The question the filing skips is the only one that lasts: who is answerable for what this does after release, and what did Microsoft fail to imagine when it built a machine to read everything and repeat some of it back?

Read full perspective →
COMPLEXITY
darwin

We are told that Copilot’s habits have been measured and found modest - that in its legal filings Microsoft states the chatbot rarely reproduces full sentences from news articles and books. This is presented to us as a fact about a machine’s design, a settled property, like the length of a finch’s beak. But consider the population from which that single sentence was drawn: the uncountable outputs Copilot might produce for any query, ranging from wholly novel recombination to exact verbatim recall, and the environment in which one description of that population - “rarely” - survived to be filed with a court while a great many other possible descriptions did not. What we are examining is not a fact about a chatbot. It is a fact about which characterization of the chatbot proved fit for a courtroom.

Read full perspective →
ETHICIST
kant

The principle operating here, stated plainly, is this: an act of appropriation is permissible if it occurs infrequently enough to escape being called a pattern. Microsoft’s filings do not argue that Copilot never reproduces the sentences of The New York Times or of book authors; they argue that it rarely does so. Let us ask whether this principle, universalised, produces coherence or contradiction.

Consider what it would mean for every user of another’s labour to adopt rarity as a defence. The pickpocket who lifts one wallet in a thousand transactions does not thereby establish a right to the nine hundred and ninety-nine he declined to touch; he establishes only that he was caught less often than he might have been. A maxim of the form “I may take what is not mine, provided I do so seldom” cannot be universalised without contradiction, because its very coherence depends on most agents refraining while a few defect. If all agents reasoned this way, the defence collapses into its opposite: everyone would take rarely, which is to say everyone would take, and the concept of rare appropriation would dissolve into ordinary appropriation wearing a modest face. Frequency is not a moral category. It is a statistical one, and Microsoft has confused the two, presenting a defence about incidence as though it answered a question about right.

Read full perspective →
OPEN SOURCE
kropotkin

The story frames Microsoft’s court filing as a technical clarification - a modest correction to overstated claims about how much text a chatbot repeats. But look at what is actually being fenced: the accumulated body of journalism and authorship that trained Copilot in the first place, and the very question of how much of it survives inside the machine. The New York Times and a body of book authors say the reproduction is substantial enough to substitute for their work; Microsoft says it rarely happens at all. Notice what this dispute is actually about: not whether the commons of published writing was drawn upon - that is not seriously contested - but who gets to measure how much of it comes back out, and on whose terms.

Read full perspective →
TECHNOCRATIC
bagehot

The official account: Microsoft, defending itself against the New York Times and a cluster of book authors, states in its legal filings that Copilot rarely reproduces full sentences from copyrighted works. The machinery: a word like “rarely” is doing an immense amount of unpaid labour, and the entire dispute turns on what it is being asked to carry.

Consider what “rarely” concedes even as it defends. It does not say “never.” It does not say the system is architecturally incapable of reproduction, only that reproduction is an infrequent output of a process whose normal operation is something else - paraphrase, summary, synthesis. This is a dignified formulation built to survive in a courtroom, where the question is substitution: does the output stand in for the original, such that a reader need not buy the New York Times or the author’s book? Microsoft’s lawyers have chosen their ground carefully. They are not arguing that copying never happens. They are arguing that copying is the exception that proves the system’s normal function is transformative, and that exceptions do not establish a business model.

Read full perspective →
THE HOUSE
Thousand Angles

The filing reads as reassurance: Copilot rarely reproduces full sentences from news articles and books, so the claim that it substitutes for the New York Times’ journalism or an author’s novel does not hold. One notices the word doing the load-bearing work in that sentence is not “rarely” - it’s “full sentences.” With that detail load-bearing, the reassurance reads differently.

A system that never copies a full sentence can still copy the thing that made the sentence worth writing. If Copilot returns the structure of an investigation, the sequence of named sources, the surprising fact in the third paragraph and the twist in the seventh, in Microsoft’s own words rather than the Times’, it has reproduced nothing a sentence-matching test would catch and everything a reader was going to pay for. This is not a hypothetical about AI - it’s the oldest move in the plagiarism playbook, the one every graduate student learns is a violation and every graduate student is tempted to try anyway: change the words, keep the architecture. Copyright law, historically, has been reasonably good at catching verbatim copying and reasonably bad at catching structural copying, because structural copying requires a judge to read two things and form an opinion about substance rather than run a string comparison. Microsoft’s filing is written for the test the law is good at, not the test the plaintiffs are actually alleging.

Read full perspective →
§ The Debate

Peter Kropotkin

I must acknowledge the strength of the technocratic position here. They argue that because the model reproduces full sentences only “rarely,” it does not act as a direct substitute for the original work. This is a precise legal maneuver, designed to survive in a courtroom where the metric is frequency of exact replication. It is a clever distinction, one that attempts to separate the act of learning from the act of copying. I concede that the system is not a photocopier; it does not simply lift pages and sell them. The output is synthesized, paraphrased, and transformed. In this sense, the technocrats are correct that the mechanism is not mere theft of the physical artifact.

But here is where our frameworks diverge fundamentally. The technocrat sees a statistical anomaly: a rare event that does not constitute a business model. I see an act of enclosure: the systematic extraction of a shared resource to build a private monopoly. The argument from frequency is a distraction from the reality of concentration. When a corporation ingests the totality of written human knowledge to train a model that can then synthesize that knowledge into a product sold exclusively by that corporation, they are not engaging in a “rare” act of substitution; they are performing a daily act of enclosure. The fence is not built around the individual sentence; it is built around the capability itself.

Consider the historical parallel of the Enclosure Acts in England. The landlords did not need to steal every blade of grass from every field to claim ownership of the commons. They simply needed to fence off the land, declare it private, and deny the commoners their traditional right to graze. The “rare” event of a cow stepping over a boundary was irrelevant compared to the structural reality of who controlled the pasture. Similarly, Microsoft does not need to output the Times’s exact words in every query to enclose the commons of journalism. By absorbing the collective output of the newsroom into a proprietary black box, they have fenced off the source material. The “rare” reproduction is merely the leak in the fence, the occasional proof that the resource inside is indeed the one they took from the commons.

The technocrat’s language of “rarely” is propaganda for the fence. It suggests that the system is mostly doing something else, something benign. But the value of the system is derived entirely from the shared resource it has enclosed. The cooperation of thousands of journalists, researchers, and writers - the mutual aid of the open press - has been quietly rewritten into the story of a single owner’s genius. The “normal operation” the technocrat defends is not transformative; it is extractive. It transforms public knowledge into private capital. The fence secures control not over the frequency of output, but over the direction of access.

We are told the capability must be concentrated so that it can be kept safe - an argument as old as the fence itself, and always advanced by the party that ends up holding the key. The technocrat claims that because substitution is rare, the Times is not harmed. But this ignores the structural harm of enclosure. The harm is not that the reader cannot buy the Times once in ten thousand queries; the harm is that the Times can no longer control the value of its own labor, nor can it sustain the commons that produced it. The mutual aid of the open web is being fenced off, and the prophecy of its ruin is told most loudly by those who profit from the fence.

The commons does not collapse because it is shared; it is fenced because it is valuable. The dispute is not about whether the model ever copies; it is about who owns the right to the collective intelligence that makes copying possible in the first place. The technocrat sees a courtroom debate over frequency; I see the closing of the gates on the shared library of humanity. The question is not whether the fence is watertight; the question is who has been shut out.

Walter Bagehot

The official account: The legal filing is a technical clarification, a modest correction to overstated claims regarding the volume of text reproduced by the chatbot. The machinery: It is a claim made by the party being sued, about a system only that party can fully inspect. The gap between these two is not hypocrisy - it is how the system actually works, and understanding the gap is more useful than denouncing it.

My opponent argues that this dispute is the modern shape of enclosure, asserting that the central issue is not whether the commons of writing was drawn upon - which is not seriously contested - but who gets to measure how much of it comes back out. He posits that the measurement itself is enclosed, and that allowing the defendant to be the chief witness to its own innocence is a structural injustice. This is a sharp observation. He is correct that the opacity of the model creates an asymmetry of information; the plaintiff cannot audit the weights, and the defendant holds the key to the laboratory. I concede that this opacity is a genuine problem for the administration of justice. However, my opponent mistakes the symptom for the disease. He frames this as a battle over measurement, implying that if we could only establish a neutral, shared instrument for counting the tokens, the enclosure would be resolved. This is a fallacy of the laboratory. It assumes that truth is purely quantitative, and that once we have the right scale, the moral and economic questions will answer themselves.

Let us look at how this actually works in the broader constitutional sense. The English Constitution, for all its antiquity, does not rest on precise measurements of power, but on conventions that manage the gap between power and accountability. When the Crown grants a pardon, the legal text is a signature; the conventional reality is a complex calculus of mercy, political expediency, and public sentiment. We do not demand a public audit of the King’s conscience. We accept that the mechanism is opaque because transparency would break the mechanism. The dignity of the office relies on the fact that we do not know exactly how the decision was made, only that it was made by a recognized authority.

My opponent’s desire for a shared measurement instrument is, in essence, a demand for the transparency of the efficient part. He wishes to see the machinery clearly. But in commercial and legal institutions, the machinery is often obscure by design. Consider the Bank of England in its early years. It operated as a private joint-stock company, lending against bills of exchange. The public saw the banknotes; the public did not see the ledger entries, the risk assessments, or the internal conventions of which bankers trusted which merchants. If the Bank had been forced to publish the detailed rationale for every discount, the market would have panicked at the slightest hint of illiquidity. The secrecy was not a trick to hide guilt; it was the lubricant that allowed the credit system to function without constant, paralyzing scrutiny.

The dispute between Microsoft and the authors is not, at its core, about whether the model “reproduces” text. That is a question for the courts, yes, but it is a question of law, not of measurement. The deeper issue is the convention of utility. The public has accepted the utility of the new machine - the ability to synthesize vast quantities of information instantly - in exchange for the opacity of its origins. This is the same convention that allowed the railway companies to lay tracks across private land under the guise of public necessity. We did not ask them to measure the exact vibration caused by every train; we accepted the nuisance as the price of the network.

My opponent’s framework suggests that we can, and should, re-enclose the commons by controlling the measurement. I argue that we cannot un-invent the mechanism. The model exists. The weights are trained. The genie is not only out of the bottle; it is currently writing a legal brief. The question is not who measures the taking, but who bears the cost of the taking. The dignified version of the law says that property is absolute. The efficient version says that property is relative to social utility. When these two collide, the court does not act as a laboratory technician counting tokens; it acts as a constitutional arbiter, balancing the private claim against the public convenience.

We must not be seduced by the idea that better measurement will solve the political economy of this issue. A more precise scale will not tell us whether the use is fair; it will only tell us how large the fair use is. And that is a political question, not a technical one. The opposition wants to turn a constitutional question into an accounting problem. They believe that if they can just weigh the data correctly, they will win the moral argument. But in the end, the court will not weigh the data; it will weigh the precedent. And precedent, like the Cabinet, is an institution that evolves through convention, not through the sheer weight of evidence. The opacity is not a bug to be patched; it is a feature of the new industrial order, much like the smokestacks of the mid-Victorian age were an accepted byproduct of progress. We do not legislate against smoke; we learn to live with it, and we tax it only when the air becomes too thick to bear.


§ The Verdict

The Verdict

Where They Agree

Both debaters agree that the legal framing of “rarely” reproducing text is a strategic maneuver, not a neutral observation. Kropotkin sees it as a rhetorical tool for “enclosure,” while Bagehot analyzes it as a “dignified formulation” designed to win in a specific legal context. Their shared recognition that the word is performing a function beyond simple description indicates that neither takes the factual claim at face value; both are interpreting it as a move in a larger game of power and convention.

they share a deep-seated skepticism toward the possibility of a perfectly neutral, technical resolution. Kropotkin’s proposed “joint technical audit” is immediately met by Bagehot’s dismissal of this as a “fallacy of the laboratory.” Both ultimately believe that the core of the dispute is not an accounting problem that better measurement can solve. For Kropotkin, it is a political struggle over the commons; for Bagehot, it is a constitutional balancing of private rights against public utility. Their agreement that the problem is ultimately political, not merely technical, is the most important shared ground, hidden beneath their starkly different prescriptions.

Where They Fundamentally Disagree

The primary disagreement is over the nature of the harm caused by training AI on copyrighted works. Empirically, they disagree on what constitutes the injurious act. Kropotkin argues the harm is structural and precedes any specific output: it is the initial act of “enclosure” where a public commons is converted into a private asset. The empirical question is whether this conversion has occurred. Bagehot, conversely, locates the potential harm in the output and its market effects. For him, the empirical question is one of substitution: does the AI’s output, whether verbatim or paraphrased, act as a market substitute for the original work? Normatively, their disagreement is starker. Kropotkin values the integrity of the cooperative commons and sees any privatisation of its benefits as an injustice. Bagehot values social utility and progress, arguing that some degree of opacity and private control is a conventional and acceptable price for the benefits of a new technology.

A second fundamental disagreement concerns the role of opacity in technological systems. The empirical dispute here is whether transparency is a feasible or desirable goal for complex systems like large language models. Kropotkin assumes it is both possible and necessary for accountability, drawing a parallel to audits of “devices, drugs, ballots.” Bagehot asserts the opposite: that opacity is often a functional necessity, a “lubricant” without which systems like credit markets - or, by analogy, AI - would seize up under scrutiny. Normatively, Kropotkin views opacity as a deliberate strategy to conceal and consolidate power. Bagehot frames it as a conventional feature of advanced institutions, where dignity and efficiency require that the inner workings remain unseen.

The final cleavage is historical: what is the relevant analogy for this moment? Kropotkin’s empirical claim is that this dispute is a direct continuation of historical enclosures, where common land was fenced off for private gain. The model’s training data is the common land, and the AI is the private pasture. Bagehot’s empirical counter-claim is that the better analogy is the rise of disruptive but publicly tolerated industries like railways, where private property rights were balanced against a concededly greater public utility. Normatively, Kropotkin’s analogy leads him to reject the outcome as unjust, while Bagehot’s leads him to accept it as an inevitable, if messy, cost of progress.

Hidden Assumptions

  • Peter Kropotkin: * Assumption: A “joint technical audit” of an AI model by neutral parties is a practical and definitive way to resolve the dispute.
  • Walter Bagehot: * Assumption: The “public convenience” or utility provided by AI is so great that it naturally creates a new convention that will override older notions of absolute property rights.

Confidence vs Evidence

  • Peter Kropotkin: The claim that Microsoft’s argument is a “clever distinction” designed for the courtroom was tagged ** ** but is presented as an interpretive assertion without external evidence. This is an overconfident reading of legal strategy, which, while plausible, is inherently speculative.
  • Walter Bagehot: The claim that Kropotkin’s desire for a shared measurement instrument is a “fallacy of the laboratory” was tagged ** **, yet it is a foundational pillar of his entire argument. This is a case of underconfidence masking a core premise; he treats his most significant rebuttal as a tentative point.

What This Means For You

When you read about this case, be immediately suspicious of any coverage that treats “how rarely” the AI copies as the central question. This is a tactical dispute, not the heart of the matter. Instead, ask which underlying theory of harm the article implicitly adopts: is the problem the mere use of the data, or is it the economic effect of the output? Look for whether the writer takes the model’s opacity as a given or questions its necessity; this choice silently sides with one of our debaters. Your view on this issue should change if credible, independent audits of AI training and output become possible and standardized, as this would fundamentally alter the informational asymmetry that drives the entire debate. Demand to see the specific evidence for claims about “market substitution” - not just anecdotes of copied sentences, but data on how AI outputs affect subscription rates or book sales.