If you've seen headlines quoting a Microsoft executive calling AI training "the largest theft of labor in human history," here's what you need to know before repeating that line: it's an allegation from unsealed court filings in a news-publishers' lawsuit, not a court finding. Newly unredacted documents in the consolidated AI copyright litigation against OpenAI and Microsoft have surfaced internal communications that plaintiffs say undercut the companies' legal defense — but the underlying cases are far from decided.
Two related but distinct sets of plaintiffs are suing OpenAI and Microsoft in the same consolidated litigation. Book authors, represented by the Authors Guild, filed a motion for summary judgment on September 4, arguing the companies used copies of copyrighted books, including material obtained from LibGen, to train AI models. In a separate strand, news publishers, including The New York Times and other outlets, allege OpenAI and Microsoft scraped their content and bypassed paywalls; newly unredacted filings in that news-publisher case include a quote from Microsoft's Director of Applied Science, Brent Hecht, describing the practice as potentially "the largest theft of labor in human history." Separately, the US Department of Justice filed a Statement of Interest supporting arguments favorable to OpenAI and Microsoft's fair-use defense in the news-publisher strand. Reuters describes the broader litigation as a major test of whether copyrighted material can legally be used to train AI systems under fair-use law.
For anyone using AI tools built on large language models, the practical stakes are significant: if courts eventually rule that training on copyrighted books and news articles without permission is not fair use, it could force AI companies to change how they source training data, potentially affecting model availability, licensing costs, or both. None of that has happened yet — the cases remain unresolved, and the unsealed documents are evidence being argued over, not a verdict.
What the Unsealed Documents Establish, and What They Don't
In the book-authors' case, a February 2026 court order, as reported by The Bookseller, established that an OpenAI employee downloaded books from LibGen and that the resulting datasets, known as Books1 and Books2, were used to train GPT-3 and GPT-3.5. That existence of the datasets is not itself in dispute; what remains contested is whether using them was copyright infringement or protected fair use. Internal messages in the book-authors' filings reportedly describe the practice as "sketchy," which plaintiffs argue undercuts the companies' fair-use defense — that is the plaintiffs' characterization of internal messages, not an admission by OpenAI or Microsoft, and no court has adopted it as fact. Fair use is the legal doctrine at the center of this dispute — it allows limited use of copyrighted material without permission under certain conditions, such as commentary, criticism or transformative reuse. Whether training an AI model on copyrighted text counts as sufficiently "transformative" to qualify is exactly the question the courts have not yet answered.
Why the DOJ's Position Complicates the Narrative
The Department of Justice's Statement of Interest, filed in the news-publisher strand of the litigation, supports arguments favorable to the AI companies' fair-use position specifically in that case — it does not address the separate factual allegations in the book-authors' filing. This signals that at least one federal body sees a legitimate legal argument on OpenAI and Microsoft's side in the news-scraping dispute, even as publishers push their own claims using the Hecht quote and other unsealed internal messages. For US, UK and Australian readers who rely on AI writing and research tools, this litigation will likely shape how AI companies license or restrict training data going forward, regardless of which side's characterization of the unsealed documents ultimately prevails in court.
What Happens Next in the Case
The book-authors' summary-judgment motion filed on September 4 is still pending, and a ruling would represent a significant step toward resolving whether AI training on copyrighted books falls under fair use, though the factual record in the news-publishers' case is separate and would need its own resolution. Until judges rule on each strand, both the plaintiffs' framing and the companies' fair-use defense remain competing legal arguments rather than settled conclusions.
Frequently Asked Questions
Did OpenAI train ChatGPT on copyrighted books?
A February 2026 court order confirmed an OpenAI employee downloaded books from LibGen and that the resulting datasets were used to train GPT-3 and GPT-3.5. The contested legal question is whether that use was copyright infringement or protected fair use, not whether the datasets existed.
Is AI training on copyrighted material legal? That is precisely what the pending litigation is meant to determine. The Department of Justice has filed a statement favorable to AI companies' fair-use argument in the news-publisher case, while book-author plaintiffs' summary-judgment motion argues the opposite in a separate strand; court rulings have not yet been issued in either.
The unsealed filings give both sets of plaintiffs new internal messages to support their claims, but the cases remain legally unresolved, with the DOJ backing arguments favorable to OpenAI and Microsoft's fair-use position in the news-publisher strand specifically. A ruling on the pending summary-judgment motion in the book-authors' case would mark the next major turning point. Check back for updates once the courts issue decisions.