Newly unsealed court documents in the federal copyright battle against Microsoft and OpenAI reveal striking internal concerns about how artificial intelligence companies acquired and used copyrighted material to build their models. Microsoft Director of Applied Science Brent Hecht warned that large-scale AI scraping could be regarded as “an astonishing theft of unprecedented proportions” and potentially the “largest theft of labor in human history.” Other internal communications cited by the plaintiffs acknowledged that AI products could substitute for publishers, reduce traffic to original sources, and threaten the economic model supporting journalism. The disclosures could complicate the companies’ defense that using copyrighted works for AI training qualifies as transformative fair use, although Microsoft and OpenAI continue to dispute the plaintiffs’ characterization of their conduct and the legal significance of individual employees’ comments.
Key Takeaways
- Newly unsealed documents show that Microsoft personnel internally recognized serious concerns about AI models absorbing enormous quantities of human-created work without traditional licensing or compensation.
- The filings indicate that executives and employees understood AI products could substitute for original news websites, potentially depriving publishers of traffic, subscriptions, advertising revenue, and ultimately the resources needed to produce new journalism.
- The legal dispute now centers not merely on how AI systems were trained, but whether large technology companies can invoke fair use while commercially benefiting from copyrighted material whose owners argue was obtained and exploited without authorization.
In-Depth
Newly unsealed filings in a federal copyright case have exposed internal discussions at Microsoft and OpenAI about the economic and legal implications of training artificial intelligence systems on copyrighted material. Microsoft applied-science director Brent Hecht warned that large models absorbing creators’ work could be viewed as “an astonishing theft of unprecedented proportions” and potentially the “largest theft of labor in human history.” Plaintiffs argue those statements undercut the companies’ contention that training AI on millions of articles constitutes lawful, transformative fair use.
The documents also sharpen a question: whether AI companies can build commercial products from copyrighted reporting without compensating the organizations that financed and produced it. Internal discussions cited in the filings reportedly acknowledged that AI products can substitute for publisher websites, reducing referral traffic and potentially creating a “doom loop” in which weakened publishers produce less original material for future models.
The companies dispute the plaintiffs’ interpretation. Their legal position remains that model training is transformative and protected by fair use, while Microsoft argues isolated employee comments do not determine the legality of the technology. No final ruling has established that the challenged training practices constitute infringement.
Still, the disclosures matter because they move the debate beyond outside criticism. They show concerns over uncompensated appropriation, market substitution and damage to content producers existed inside the companies developing these systems. The court must decide the legal questions, but the record raises a broader issue: innovation does not automatically erase property rights merely because copying occurs at unprecedented technological scale.
Sources
- https://www.reuters.com/business/media-telecom/openai-microsoft-executives-quotes-ai-training-threaten-copyright-defense-news-2026-09-17/
- https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/
- https://arstechnica.com/tech-policy/2026/09/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history/
- https://authorsguild.org/news/plaintiffs-file-motion-for-summary-judgment-v-openai-and-microsoft/

