The DOJ just told a federal judge that training AI on your work is fair use
In re OpenAI, Inc. Copyright Infringement Litigation: the United States filed a 20-page Statement of Interest on September 1, 2026 backing OpenAI on training. Here is what it argues, where it breaks from the Copyright Office, and what it means for musicians and writers.

On September 1, 2026, the United States walked into the Southern District of New York and took a side.
The filing is a Statement of Interest in In re: OpenAI, Inc. Copyright Infringement Litigation, No. 25-md-3143 (SHS) (OTW) — the consolidated multidistrict case that includes The New York Times Co. v. Microsoft Corp. and the book-author and publisher suits riding alongside it. It runs 20 pages, it is signed by Senior Counsel to the Associate Attorney General Michael Weisbuch, and its argument is short enough to fit in one line: copying copyrighted text in order to train a large language model is fair use.
You can read the whole thing yourself. I have posted it here:
Download the filing (PDF, 20 pages)
If you write songs, write words, or make images for a living, this is the single most consequential document filed in the AI copyright fight so far — not because a court has agreed with it, but because the government of the United States has now put its institutional weight behind one reading of section 107 while the judge is still deciding.
What the filing actually argues
The government splits the AI pipeline into stages and insists the court do the same: acquisition (collecting the data), training (copying works and feeding them through the model), and output (what the model says back to a user). It then argues only about the middle one.
The United States focuses on the question whether the use of copyrighted works at the training stage — by copying works in order to feed data into the model as learning material — constitutes fair use.
That framing is the whole ballgame, and everything else follows from it.
Factor one: training is "exceedingly transformative"
The first fair-use factor asks whether the new use supersedes the original or adds something with a further purpose or different character. The government says training is not close:
- The purpose of a Times article is to use language to inform or entertain a reading audience.
- The purpose of copying that article into a training run is to learn statistical patterns — vocabulary, syntax, relationships between concepts.
Different in kind, the filing says, quoting Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007 (N.D. Cal. 2025), which called LLM training "transformative — spectacularly so," and Kadrey v. Meta Platforms, 788 F. Supp. 3d 1026 (N.D. Cal. 2025), which called it "highly transformative."
The anchor case is Google LLC v. Oracle America, 593 U.S. 1 (2021). Google copied 11,500 lines of Java declaring code precisely and the Supreme Court still found fair use, because the copying let programmers carry existing skills into a new environment. If verbatim code copying clears the bar, the government argues, so does copying text as learning material.
Commercial motive gets waved off in two sentences, both from precedent: nearly all the uses listed in the preamble to § 107 are conducted for profit, and "the more transformative the work, the less will be the significance of other factors, like commercialism."
Factors two and three: barely a speed bump
The filing spends one footnote on them. Nature of the work follows the transformative purpose. Amount used — yes, training copies entire works, but under Authors Guild v. Google, 804 F.3d 202 (2d Cir. 2015), what matters is not how much was copied but how much was made accessible to the public. Training makes nothing accessible. Footnote 16, and on it goes.
Factor four: the fight
This is where the document gets sharp, and where it picks a target.
The fourth factor asks about the effect on the potential market for the work. The government's position: only substitutive harm counts. Not lost sales generally. Not competition generally. The use has to act as a substitute — it has to reveal enough protected expression that a reader takes it instead of the original.
Training reveals nothing to anybody. So, the argument goes, it cannot substitute for anything.
Then it goes after Kadrey. Judge Chhabria's opinion floated a theory of market dilution: even if no output is substantially similar to any book, LLMs can flood the market with competing works and gut the incentive to write. The government calls that "deeply flawed," and lists why:
| The Kadrey theory | The government's answer |
|---|---|
| Training and output are one continuous use | They are distinct uses; § 107 requires a use-by-use analysis (Warhol v. Goldsmith, 598 U.S. 508) |
| Similarity between works is irrelevant here | Substantial similarity is the most basic element of infringement |
| Genre-level competition is market harm | A genre is an uncopyrightable idea; competing "in the same category" is not harm copyright recognizes |
| The court cited no case for it | The opinion concedes the theory has "never made a difference in a case before" |
And then the analogy the filing clearly wrote for the press: a teenage Joan Didion typing out Hemingway stories to learn how the sentences worked. On the Kadrey logic, the filing says, Didion would owe Hemingway money every time she published, because the learning and the writing were one use and her books competed with his.
The parts that are not about copyright at all
Roughly a quarter of the filing is policy, and it is worth being honest about that.
National security. The brief opens with GAO warnings about AI in intelligence analysis, weapons systems and battlefield targeting, and the January 2025 and June 2026 executive orders on American AI dominance. Rules that make US model training harder, it argues, hand an advantage to adversaries who ignore US intellectual property anyway.
Competition. This is the argument most likely to land with independent creators, and it is aimed squarely at them. If training requires licensing, only companies with the deepest pockets can pay — and the money flows disproportionately to legacy publishers with the largest archives. In the filing's words, licensing fees would "function primarily as large subsidies for old mainstream media companies."
The independent creator. The government name-checks writers using LLMs to generate an image that would otherwise need a photographer, and an independent physics teacher who used an LLM to publish a contrarian critique of a Times story about data-center water use. It also notes that Times journalists themselves use LLMs to conceptualize and edit.
Footnote 13 is the honest one: the United States takes no position on whether a licensing regime would even be feasible, and concedes publishers can and do sign voluntary deals for real-time, paywalled and proprietary access. Nothing in the brief stops anyone from licensing. It only says the law should not require it.
What it does not say
Read the limits carefully, because the headline version of this story will flatten them.
- Outputs are still live. The brief repeatedly reserves the question. If a model reconstructs and disseminates a protected work, that is a different use and may well not be fair. The Times allegations about regurgitation are not answered here — they are set aside.
- Acquisition is untouched. How the data got collected — torrented libraries, scraped paywalls — is a separate question. Bartz itself was fair use on training and still faced a pirated-library problem.
- Music is not in this case. The MDL is text: news, books. Sound recordings, compositions and voice have their own statutes, their own § 114 and § 115 machinery, and their own suits.
- Remedy is capped anyway. Footnote 15 argues any relief must be limited to what plaintiffs need, citing Trump v. CASA, 606 U.S. 831 (2025) — a "tiny sliver of anomalous reconstructive outputs" cannot justify liability that shuts down training.
- It is not law. It is one page of a docket. Judge Stein rules.
Why musicians should care about a text case
Because the reasoning travels, and the reasoning is what gets cited.
If a court adopts the use-by-use frame, then every argument you have heard about AI music splits in two. Ingesting a catalogue to learn what a snare sounds like in a mix becomes one legal question. A model producing a track that sounds substantially like your record becomes an entirely different one — and on the government's own logic, that second question stays wide open, decided output by output on substantial similarity.
That is not the worst outcome for artists, and it is worth saying plainly. The dilution theory is emotionally satisfying — the machines are flooding my market — but as written in Kadrey it was untested, unbriefed and, by the court's own admission, unprecedented. A doctrine built on it would be fragile. A doctrine built on substantial similarity at the output stage is boring, old, and enforceable, and it is the one that has protected sampling claims for forty years.
The real risk in the filing is not the fair-use holding. It is the competition argument being used in both directions. The government says licensing entrenches the majors. It is right. But "no licensing at all" also means no payment at all, and a market where the only people who get paid are the ones with the leverage to sign a private deal — which is to say, again, the majors. Independent artists lose either way unless something else changes.
What to do this week
- Read the filing. It is 20 pages, mostly plain English, and it will be cited at you for the next two years. The download is above.
- Register your work. Statutory damages and fees require timely registration, and every one of these theories collapses without a registration to sue on. Our guide to copyrighting your music in the US walks the form.
- Get your metadata right. Output-stage claims turn on proving what you made and when. Clean metadata is evidence.
- Read your AI clauses. Distribution, sync and sample-pack agreements are quietly acquiring training-rights language. That is a contract question, not a copyright question, and contracts you sign beat any court ruling.
The government's closing line is a quotation from Google: LLM training is "consistent with that creative progress that is the basic constitutional objective of copyright itself." Whether Judge Stein agrees is the story of the next year.