4 Sep 2026 · Every story has many sides
Multi-Perspective News Analysis
Search About Phronopolis

Microsoft Copilot rarely copies full sentences, firm claims

Microsoft’s Copilot is not new because it writes; it is new because it can be interrogated at scale - a single operator can put the same query to it ten thousand times, in ten thousand phrasings, and read back, statistically, what the machine actually does with a copyrighted sentence rather than what its makers say it does. That is the instrument this dispute has actually produced, though nobody has yet built it properly: a repeatable audit of retrieval behaviour. The limit the story treats as fixed is the word “rarely,” offered in Microsoft’s own filings as though it were a finding rather than a claim. I ask what experiment stands behind it, and I find none disclosed. A number withheld is not evidence; it is an idol dressed as one.

Consider what is actually being contested. The New York Times and a set of book authors allege that Copilot reproduces their work closely enough to serve as a substitute for it - the reader gets the sentence, the plot, the analysis, without the purchase. Microsoft answers that full-sentence reproduction is rare. Notice that both sides have skipped the instrument and gone straight to the verdict. “Rarely” is a word calibrated for a courtroom, not a laboratory; it tells us nothing about what counts as rare, over what sample of prompts, against what definition of “full sentence,” on what date the corpus was queried.

The stronger version of the authors’ complaint does not even need full-sentence reproduction to hold. A chatbot that reliably delivers the substance of a Times investigation, or the twist of a novel, in paraphrase, has arguably done the economic harm the lawsuit is about, whether or not it strings together the author’s exact clauses. This is the point where I must be fair to the plaintiffs: their claim is not really about verbatim copying at all, it is about substitution, and substitution is a market fact, testable independently of sentence-level fidelity. So Microsoft’s statistic, even if true, may be answering a question nobody who understands the harm was asking.

Which is why the honest move - the Baconian move - is to build the test neither party has offered. Take a controlled set of Times articles and copyrighted books not in Copilot’s training cutoff conversation but demonstrably within its corpus; query the system across a spread of prompts designed to elicit summary, paraphrase, and quotation; measure, across the whole sample, what fraction of queries return content close enough that a reader would not need to buy the original. That number, run by a neutral party and disclosed with its methodology, would settle more of this case than either “rarely” or “systematically” ever will. I do not say the answer favours Microsoft. I say we do not know, and a company confident in its defense should welcome the instrument that could prove it right, just as a plaintiff confident in the harm should welcome the instrument that could prove the substitution real and repeated. The reluctance of either party to propose such a protocol is itself informative.

Picture an author - not an abstraction, but someone at a desk, a specific paragraph of theirs open on one screen and Copilot open on another, typing in a question about their own chapter to see what comes back. That is the experiment already available to any single person with a library card and an afternoon; it has simply not been run at the scale a court requires, and multiplied into a verdict. The technology that makes such an audit trivially repeatable - millions of queries, logged, compared, scored - is the very technology on trial. Copilot is accused of a crime it is also, uniquely, equipped to help prosecute or acquit itself of, if anyone will point the instrument at it rather than at a press release.

The old dispute over reproduction was fought with testimony and a handful of exhibits. This one can be fought with a sampling frame and a spreadsheet. Whoever refuses that instrument first is telling you what they expect it to find.