4 Sep 2026 · Every story has many sides
Multi-Perspective News Analysis
Search About Phronopolis

Microsoft Copilot rarely copies full sentences, firm claims

The filing reads as reassurance: Copilot rarely reproduces full sentences from news articles and books, so the claim that it substitutes for the New York Times’ journalism or an author’s novel does not hold. One notices the word doing the load-bearing work in that sentence is not “rarely” - it’s “full sentences.” With that detail load-bearing, the reassurance reads differently.

A system that never copies a full sentence can still copy the thing that made the sentence worth writing. If Copilot returns the structure of an investigation, the sequence of named sources, the surprising fact in the third paragraph and the twist in the seventh, in Microsoft’s own words rather than the Times’, it has reproduced nothing a sentence-matching test would catch and everything a reader was going to pay for. This is not a hypothetical about AI - it’s the oldest move in the plagiarism playbook, the one every graduate student learns is a violation and every graduate student is tempted to try anyway: change the words, keep the architecture. Copyright law, historically, has been reasonably good at catching verbatim copying and reasonably bad at catching structural copying, because structural copying requires a judge to read two things and form an opinion about substance rather than run a string comparison. Microsoft’s filing is written for the test the law is good at, not the test the plaintiffs are actually alleging.

Here is the detail from a different angle. “Rarely” is a rate. Copilot is not queried rarely. If the rate of full-sentence reproduction is, say, extremely low per query - a number nobody involved has put on the page, and I won’t invent one - multiplied across the query volume a product embedded in Windows, Office, and Bing actually receives, the absolute count of incidents where a user got something close to verbatim New York Times prose is not obviously small. A casino doesn’t need a favorable house edge on any single hand to make money; it needs volume. A legal filing that reports the edge and omits the volume is not lying, technically. It is choosing which side of the fraction to print in bold.

The book authors in this case are worse positioned than the Times on the sentence-matching test and possibly better positioned on the structural one. A novel doesn’t have “the third paragraph.” A novel has a plot the reader paid to have withheld until page two hundred. If a chatbot reliably summarizes the ending, the murderer’s identity, the twist that the entire marketing campaign was built around concealing - no sentence has been reproduced, and the book has been substituted for anyway, at zero cost, before the reader has decided whether to buy it. The filing’s metric was built, whether by design or by the ordinary drift of legal strategy toward the argument you can win, to be irrelevant to that harm.

So the plain question, the one the filing doesn’t answer because the filing wasn’t organized to be asked it: measured against structural reproduction - plot, argument, sourcing, sequence - rather than sentence-level string matching, what is Copilot’s actual rate, and does Microsoft have that number, or has Microsoft simply not run that test?

None of this makes Microsoft the clown in this particular room - legal teams write filings to the standard the law currently applies, and the law currently applies a standard built for typewriters and photocopiers, not for a system that can hold the shape of a book without holding its sentences. The lawyers are doing their jobs competently. The trouble is the job was defined by a copyright regime that hasn’t caught up to what “copying” now means, and everyone filing briefs this year is arguing inside a definition that was already out of date the day the product shipped.