ModelsArchitectures & capability
Unsealed NYT filings intensify scrutiny of Microsoft and OpenAI training-data practices
Newly unredacted litigation material alleges Microsoft and OpenAI internally recognized publisher harm and used disputed web-data pipelines, while the companies maintain that AI training and related products are lawful fair use. The underlying exhibits remain only partly public.

The newly public material is significant less for a single internal comment than for its allegations about data provenance, publisher substitution, and internal risk assessments around AI scraping. [1] [8] [9]
Reported allegations include Common Crawl-derived data, Bing-related data sharing, and projects described as Project Taxi and Project Mango, but these remain plaintiffs’ characterizations rather than court findings. [1] [8]
Evidence in the reviewed sources shows a copyright case moving beyond abstract fair-use arguments into concrete questions about crawler design, dataset lineage, paywall access, copyright-management information, and output substitution.
Read the full assessment
The implications for practitioners are broader than this lawsuit: AI teams may need auditable data manifests, clearer separation of training and retrieval systems, and stronger records of permissions and filtering. For business leaders, vendor indemnity, licensing exposure, and traffic substitution are becoming operational risks, not only legal issues.
Executive brief
On September 17, 2026, TechCrunch reported newly unredacted material from the publishers’ summary-judgment filing in *The New York Times Co. v. Bloomberg Law reported that Project Taxi allegedly supplied OpenAI with a Bing-derived compilation of “billions” of webpages, while Project Mango allegedly involved OpenAI paying Microsoft to develop and operate a crawler. OpenAI, Microsoft Leaders Admitted AI Threatens News, Books (2) OpenAI’s own case page says the dispute centers on fair use and argues that training models can deliver broad public benefits; its September 4, 2026 summary-judgment memorandum says OpenAI collected broad, diverse internet text for generalist LLMs and did not target news as a category.
Read the full section
On September 17, 2026, TechCrunch reported newly unredacted material from the publishers’ summary-judgment filing in The New York Times Co. v. Microsoft Corp. et al. and the consolidated OpenAI copyright multidistrict litigation. The headline allegation is that Microsoft Director of Applied Science Brent Hecht internally characterized large-scale AI scraping as “an astonishing theft of unprecedented proportions” and possibly “the largest theft of labor in human history.” TechCrunch and Ars Technica both emphasize an important caveat: much of the public record still consists of the publishers’ brief quoting sealed exhibits, not the underlying exhibits in full context. Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal | TechCrunch
For AI practitioners, the material matters less as a single explosive quote and more as a window into data provenance, market-substitution evidence, and internal risk assessment. The newly public allegations describe: use of Common Crawl-derived data; alleged exchanges of training data between OpenAI and Microsoft; internal projects named Project Taxi and Project Mango; alleged paywall-circumvention discussions; and alleged stripping of copyright-management information before training. Bloomberg Law reported that Project Taxi allegedly supplied OpenAI with a Bing-derived compilation of “billions” of webpages, while Project Mango allegedly involved OpenAI paying Microsoft to develop and operate a crawler. OpenAI, Microsoft Leaders Admitted AI Threatens News, Books (2)
The defendants’ public position remains that AI training is fair use and transformative. OpenAI’s own case page says the dispute centers on fair use and argues that training models can deliver broad public benefits; its September 4, 2026 summary-judgment memorandum says OpenAI collected broad, diverse internet text for generalist LLMs and did not target news as a category. Reporting the facts about the New York Times’ lawsuit | OpenAI Microsoft has said the quoted internal comments do not represent its legal position and that its court filings explain why the uses are lawful and why Copilot is not a substitute for publisher journalism. Inside Microsoft and OpenAI, Worry About Damaging the Publishing Industry - The New York Times
No court has yet adopted the publishers’ new factual characterizations as findings. The case is now in a high-stakes summary-judgment posture where the court must decide what can be resolved as a matter of law and what factual disputes, if any, go to trial.
What changed and event timeline
Original NYT complaint
The Times sued OpenAI and Microsoft in the Southern District of New York, alleging that the companies used Times works without permission to build AI products that compete with it.
More detail
The complaint alleged that Bing/GenAI products could return verbatim excerpts or detailed summaries longer than traditional search snippets and that The Times had not licensed GenAI training use.
Partial narrowing
Judge Sidney H. Stein dismissed some DMCA-related claims concerning outputs that were excerpts rather than full “copies” under the court’s reading of 17 U.S.C. §1202(b)(3), while other litigation issues continued.
Government and defendants’ summary-judgment push
The Trump administration filed a statement backing a broad fair-use theory for AI training, arguing that training generally transforms copyrighted material and that AI leadership implicates national interests. [](https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1496.0.pdf)
More detail
Reuters described it as the first U.S. government intervention in the broader wave of AI-training copyright suits. OpenAI and Microsoft then filed summary-judgment motions; OpenAI’s memorandum argues that LLM pretraining and Browse are lawful uses of facts and language, not protected expression.
Newly unredacted publishers’ material
The publishers’ partially unsealed motion surfaced internal statements and alleged dataset counts. TechCrunch, Ars, Bloomberg Law, The Verge-linked coverage, Reuters, and Nieman Lab all reported on the material, but the most important constraint is that many underlying exhibits remain sealed or only selectively quoted.
Capabilities and access
The public filings and reporting refer to ChatGPT, OpenAI LLMs, Browse, Microsoft Copilot/Bing Chat, and historical training assets such as the GPT-3 training dataset. [](https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1496.0.pdf) Where exact versions matter, the public record is incomplete.
Read the full section
The public filings and reporting refer to ChatGPT, OpenAI LLMs, Browse, Microsoft Copilot/Bing Chat, and historical training assets such as the GPT-3 training dataset. The exact model checkpoints, full dataset manifests, filtering code, and post-training recipes are not fully public in the unsealed reporting.
OpenAI’s September 4 memorandum describes the models as “generalist” LLMs trained from a broad and diverse internet corpus, including a filtered version of Common Crawl plus other datasets. It also discusses Browse as an automated web-browsing service that reads web sources in response to user requests and returns synthesized answers with links. [](https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1496.0.pdf)
Where exact versions matter, the public record is incomplete. The original Times complaint included examples involving GPT-family systems and Microsoft’s Bing/Copilot products, but the newly unredacted material does not, by itself, provide reproducible model-version documentation for every challenged output.
Technical analysis for researchers and developers
The record does not newly disclose model architecture internals such as layer counts, attention variants, optimizer details, or full training-token mixtures. Bloomberg Law reports that Project Taxi allegedly delivered a Bing-gathered web compilation to OpenAI, while Project Mango allegedly involved a Microsoft-operated crawler funded by OpenAI.
Read the full section
Architecture and data pipeline issues
The record does not newly disclose model architecture internals such as layer counts, attention variants, optimizer details, or full training-token mixtures. The technical significance is instead in data acquisition and governance.
The publishers allege multiple content channels: Common Crawl-derived datasets, OpenAI-built WebText-style datasets, Bing Index-related sharing, Project Taxi, Project Mango, mid-training datasets, and data passed between OpenAI and Microsoft. Bloomberg Law reports that Project Taxi allegedly delivered a Bing-gathered web compilation to OpenAI, while Project Mango allegedly involved a Microsoft-operated crawler funded by OpenAI. OpenAI, Microsoft Leaders Admitted AI Threatens News, Books (2)
For builders, the implementation implication is clear: “training data” is not one thing. Risk varies by stage:
- Pretraining: broad web-scale ingestion; hard to audit after the fact without manifests and hash-level lineage.
- Mid-training / post-training: curated “high-quality” or freshness-oriented data may create higher concentration of publisher works.
- Retrieval / Browse / grounding: live access to paywalled or licensed sources raises separate contract, robots.txt, paywall, and CMI issues.
- Output controls: memorization filters and refusal policies mitigate regurgitation but do not answer acquisition-lawfulness questions.
Evaluation methodology
The parties contest how often models reproduce protected expression. OpenAI’s September 4 memorandum says its expert reviewed a sample of 20 million ChatGPT conversation logs and found 24 alleged verbatim regurgitations, while plaintiffs’ experts found similarly low rates in some analyses; OpenAI also says adversarial or “forced” prompting experiments produced limited extraction under conditions unavailable to ordinary ChatGPT users. These are defense-reported expert results, not independent peer review. [](https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1496.0.pdf)
Independent research supports a narrower point: LLMs can memorize and regurgitate training text, but the rate and extractability vary by model, data duplication, prompt method, and target work. The paper The Files are in the Computer defines memorization as the ability to reconstruct a near-exact, substantial portion of training data from a model, while other work on copyrighted books finds that most models do not memorize most books but some books and some models are notable exceptions. The Files are in the Computer: Copyright, Memorization, and Generative AI Earlier ACL and security literature likewise treats verbatim memorization and training-data extraction as measurable but nonuniform phenomena. Copyright Violations and Large Language Models - ACL Anthology
Reproducibility
Independent reproduction is currently limited because the key artifacts are unavailable: sealed exhibits, nonpublic model checkpoints, internal datasets, filtering pipelines, user-query samples, and expert code. Practitioners should treat both sides’ numerical claims as litigation evidence, not benchmark results.
Claims and evidence
- Microsoft executive Brent Hecht used “theft” language about AI scraping.
- OpenAI/Microsoft products were internally described as substitutive for publisher sites.
- Project Taxi and Project Mango involved Microsoft/OpenAI data-sharing or crawling arrangements.
Read the full section
| Material claim | Evidence status |
| Microsoft executive Brent Hecht used “theft” language about AI scraping. | Reported by TechCrunch and Ars from newly unredacted plaintiffs’ filing; underlying full exhibit context still limited. Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal | TechCrunch |
| OpenAI/Microsoft products were internally described as substitutive for publisher sites. | Reported by Bloomberg Law, Reuters, and TheWrap from the unsealed motion; Microsoft disputes that these observations decide the legal issue. OpenAI, Microsoft Leaders Admitted AI Threatens News Outlets (1) |
| Project Taxi and Project Mango involved Microsoft/OpenAI data-sharing or crawling arrangements. | Bloomberg Law reporting on the motion; still plaintiffs’ characterization pending judicial findings. OpenAI, Microsoft Leaders Admitted AI Threatens News, Books (2) |
| OpenAI argues training and Browse are fair use. | Vendor/defense position in OpenAI’s public case page and September 4 memorandum. Reporting the facts about the New York Times’ lawsuit | OpenAI |
| Models rarely regurgitate publisher works in ordinary use. | Defense-reported expert analysis; contested and not independently reproducible from public artifacts. [](https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1496.0.pdf) |
| LLM memorization is technically real but variable. | Independent academic literature; does not by itself prove infringement in this case. The Files are in the Computer: Copyright, Memorization, and Generative AI |
Context and prior work
The core legal question is not simply “did the model read copyrighted text?” [](https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1496.0.pdf) The publishers argue the products substitute for their journalism and harm licensing and referral markets. The original Times complaint explicitly alleged that The Times permits search indexing for traditional search purposes but never authorized GenAI use.
Read the full section
The core legal question is not simply “did the model read copyrighted text?” It is whether copying for training, post-training, grounding, and/or output generation is protected by fair use, and whether any outputs or data exchanges separately infringe copyright or DMCA provisions.
The defendants lean on a tradition of cases involving transformative use, search, indexing, and facts. OpenAI’s memorandum emphasizes that facts are not copyrightable and that LLMs answer in their own words after training on broad corpora. [](https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1496.0.pdf) The publishers argue the products substitute for their journalism and harm licensing and referral markets. The original Times complaint explicitly alleged that The Times permits search indexing for traditional search purposes but never authorized GenAI use. The New York Times Company v. MICROSOFT CORPORATION et al, 1:23-cv-11195, No. 1 (S.D.N.Y. Dec. 27, 2023)
The policy context has shifted toward AI-industry support: the U.S. government’s September 2026 statement backs a broad fair-use theory for AI training, while creator groups and news plaintiffs continue to frame unlicensed training as market-damaging extraction. US government backs OpenAI in New York Times copyright case By Reuters
Limitations, safety, and contested findings
- Selective quotation risk: The strongest quoted language comes through the plaintiffs’ brief. Without full exhibits, tone, audience, and surrounding caveats are uncertain.
- No final liability ruling: Unsealed allegations are not judicial findings.
- Dataset counts ≠ output infringement: A corpus may contain works that never become extractable outputs.
Read the full section
- Selective quotation risk: The strongest quoted language comes through the plaintiffs’ brief. Without full exhibits, tone, audience, and surrounding caveats are uncertain.
- No final liability ruling: Unsealed allegations are not judicial findings.
- Dataset counts ≠ output infringement: A corpus may contain works that never become extractable outputs; conversely, low observed regurgitation does not resolve whether acquisition was lawful.
- Traffic harm is disputed: Plaintiffs point to answer-engine click-through reductions; OpenAI says expert analysis found no negative impact on publisher traffic or finances. [](https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1496.0.pdf)
- Safety and privacy overlap: The dispute also touches user-log discovery. OpenAI says it has resisted demands affecting private ChatGPT conversations. Reporting the facts about the New York Times’ lawsuit | OpenAI
Business and practitioner implications
- Data provenance is now board-level risk: Maintain dataset manifests, license terms, crawl permissions, CMI handling records, and deletion/retraining playbooks.
- Separate training, retrieval, and output policy: A fair-use view on pretraining does not automatically bless paywall bypass, live retrieval, or verbatim output.
- Contract for indemnity carefully: Enterprise AI buyers should ask vendors what training-data warranties, copyright indemnities, and retrieval-source controls actually cover.
Read the full section
- Data provenance is now board-level risk: Maintain dataset manifests, license terms, crawl permissions, CMI handling records, and deletion/retraining playbooks.
- Separate training, retrieval, and output policy: A fair-use view on pretraining does not automatically bless paywall bypass, live retrieval, or verbatim output.
- Contract for indemnity carefully: Enterprise AI buyers should ask vendors what training-data warranties, copyright indemnities, and retrieval-source controls actually cover.
- Measure substitution, not just hallucination: For publishers and AI search providers, click-through, citation placement, snippet length, and answer completeness are economic design choices.
- Prepare for licensing markets: Even if some training is held fair use, curated mid-training, grounding, and premium-news access may move toward licensing or collective-management models.
Sources
Key sources used: TechCrunch’s original report; Ars Technica, Bloomberg Law, Reuters, Nieman Lab, OpenAI’s public litigation page, OpenAI’s September 4 memorandum, the original NYT complaint, the April 2025 court opinion, the Authors Guild update, and academic work on LLM memorization and extraction.
The source trail.
Sources (12)
Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal | TechCrunch
techcrunch.comOpenAI and Microsoft knew they were starting a ‘doom loop’ for the web
Related coverage; assess separately
theverge.com