In July 2026, a federal district court in San Francisco approved a $1.5 billion settlement in Bartz v. Anthropic, establishing the largest copyright settlement in US legal history concerning the ingestion of unauthorized book archives for model training (Reuters).
Concurrently, ongoing litigation between The New York Times and OpenAI/Microsoft in the Southern District of New York has generated over $28 million in plaintiff legal expenses without reaching an initial liability verdict (Variety).
While aggregate legal trackers catalog more than 200 active AI disputes representing billions in theoretical damages claims (AI Lawsuit Tracker), judicial outcomes remain fragmented across distinct factual contexts—including shadow library scraping, structured database extraction, direct output memorization, and commercial licensing.
Landmark US district and appellate precedents
1. Bartz v. Anthropic (N.D. Cal.)
- Legal Issue: The download and curation of central training corpora from unauthorized repositories (LibGen and Pirate Library Mirror).
- Substantive Finding: Judge William Alsup ruled that while algorithmic training on lawfully acquired text constitutes transformative fair use, the unauthorized bulk acquisition and retention of pirated digital libraries constituted actionable copyright infringement (Justia Docket Analysis).
- Settlement Resolution: Anthropic agreed to a $1.5 billion settlement fund covering past conduct across an eligible class of approximately 482,000 works (~$3,000 per claimed title), approved by Judge Araceli Martínez-Olguín on 20 July 2026. The settlement resolves historical training liabilities through August 2025 without granting prospective training licenses.
2. Kadrey v. Meta Platforms (N.D. Cal.)
- Legal Issue: Training of Llama foundation models on dataset collections containing works by authors including Richard Kadrey.
- Substantive Finding: Judge Vince Chhabria granted summary judgment in favor of Meta in June 2025, holding that the plaintiffs had failed to demonstrate quantifiable market dilution under the fourth fair use factor, as the models did not reproduce the plaintiffs' expressive content in user outputs (Justia Docket).
3. Thomson Reuters v. Ross Intelligence (D. Del. / 3rd Cir.)
- Legal Issue: Direct competitive substitution through the unauthorized reproduction of 2,243 proprietary Westlaw headnotes to train a legal research tool.
- Substantive Finding: Judge Stephanos Bibas held in February 2025 that Ross's commercial use of structured editorial headnotes was non-transformative and caused direct market harm to Thomson Reuters' core product. Oral arguments for the appeal were heard before the Third Circuit in June 2026 (Baker Botts Analysis).
4. The New York Times Co. v. Microsoft and OpenAI (S.D.N.Y.)
- Legal Issue: Synthetic reproduction of paywalled journalistic articles, synthetic search summaries, and alleged removal of Copyright Management Information (CMI).
- Procedural Status: Following Judge Sidney Stein's March 2025 ruling upholding the core copyright claims, litigation expanded into extensive discovery disputes regarding model training logs and synthetic output benchmarks across 17 publisher plaintiffs (CourtListener Docket).
European and UK jurisdictional rulings
1. Getty Images v. Stability AI (UK High Court)
- Legal Issue: Whether importing pre-trained neural network weights (
.safetensors) into the UK constitutes the importation of an infringing article under the Copyright, Designs and Patents Act 1988 (CDPA). - Substantive Finding: Mrs Justice Joanna Smith ruled in November 2025 that model parameters and latent representations do not store static reproductions of training images ([2025] EWHC 2863 (Ch)), dismissing the primary copyright infringement claims while preserving narrow trademark claims concerning synthetic watermarks.
2. GEMA v. OpenAI (Regional Court of Munich I)
- Legal Issue: The memorization and verbatim emission of copyrighted musical lyrics within GPT-4 outputs.
- Substantive Finding: In November 2025, the court held that verbatim output generation constitutes an unauthorized reproduction not shielded by text and data mining (TDM) exemptions under Sections 44b and 60d of the German Copyright Act (UrhG).
Summary of core litigation tracks
| Case | Jurisdiction | Central Legal Question | Judicial Determination / Status |
|---|---|---|---|
| Bartz v. Anthropic | US (N.D. Cal.) | Scraping and storing pirated book repositories | $1.5B class settlement approved July 2026 |
| Kadrey v. Meta | US (N.D. Cal.) | Fair use defense for pretraining on internet text | Summary judgment for Meta; appeal denied |
| Thomson Reuters v. Ross | US (3rd Cir.) | Training competitive products on proprietary headnotes | Partial summary judgment for TR; on appeal |
| NYT v. OpenAI | US (S.D.N.Y.) | Ingestion of paywalled news & conversational RAG | Active discovery & sanctions phase |
| Getty v. Stability | UK (High Court) | Model weights as infringing articles under CDPA | Infringement dismissed; weights do not store copies |
| GEMA v. OpenAI | Germany (Munich I) | Verbatim lyric reproduction in model responses | Injunction granted against verbatim output; on appeal |
Commercial bilateral licensing market
Alongside active litigation, major AI model developers have established substantial commercial licensing agreements with content holders:
- News Corp & OpenAI: Disclosed multi-year partnership covering major international news titles, reported at over $250 million over five years (Wall Street Journal).
- Axel Springer & OpenAI: Multi-year licensing agreement for real-time news retrieval and training access (OpenAI Announcement).
- Reddit & Google / OpenAI: Multi-million-dollar annual API licensing agreements providing real-time data ingestion for search synthesis and model training.
These commercial agreements create established market pricing for authenticated, real-time data access feeds, operating alongside ongoing court determinations of fair use boundaries.