AI Training, Piracy, and the Publisher Rights-Provenance Standard
The Anthropic litigation separates AI training questions from pirated acquisition. Build the rights-provenance record authors and publishers need now.
Direct answer: AI training and the acquisition of training material are separate copyright questions. The Bartz v. Anthropic litigation shows why publishers need a title-level rights file that connects the work, identifiers, copyright registration, contracts, licensed uses, source provenance, and the current owner of each right.
Publishers have spent years debating whether training artificial-intelligence systems on copyrighted books can qualify as fair use. The more durable lesson from the Anthropic litigation is narrower and more practical: where a copy came from, who owns the rights, and what records exist can determine whether a rightsholder can identify a work and assert a claim.
In 2025, the district court treated training on lawfully acquired books as fair use in the circumstances before it, while distinguishing Anthropic’s acquisition of books from pirate libraries. The later settlement process made the paper trail impossible to ignore.
By September 2026, claimants were reviewing title-level allocations, rights reversions, publisher claims, and registration issues. The Authors Guild reported further allocation guidance on September 14 and an extension of the co-claimant dispute window on September 17. On September 18, it publicly pressed publishers over long-out-of-print works and works they had failed to register.
The publishing standard that follows is not “all AI training is fair use,” because the law is not that simple. The standard is: every publication needs a defensible rights-provenance record.
Fair use and acquisition are different questions
Copyright disputes often collapse several issues into one headline. That is dangerous. A court can analyze the use of a work for training separately from the way the defendant obtained the copy. A ruling in one case also does not decide every future model, dataset, output, plaintiff, contract, or jurisdiction.
For publishers, that means broad slogans are less useful than disciplined records. If a dispute arises, the publisher may need to show not only that a work exists, but which edition is at issue, who owned the relevant rights at the relevant time, whether those rights reverted, when the copyright was registered, and what permissions were granted.
An ISBN identifies an edition. It does not prove copyright ownership.
ISBNs are essential publishing identifiers, but they solve a different problem from copyright registration and contracts. An ISBN can identify a particular edition or format. It does not, by itself, establish who owns copyright or who is entitled to infringement proceeds.
Copyright registration creates a public registration record and can affect the remedies available under U.S. law. A publishing contract allocates rights between parties. A rights-reversion notice can change that allocation later. Permissions may cover only a particular territory, format, term, language, or use.
| Record | What it answers |
|---|---|
| ISBN / ASIN | Which edition or product is this? |
| Copyright registration | Who is identified as claimant, when was registration effective, and what record exists? |
| Publishing contract | Which exclusive rights were granted, retained, sublicensed, or reserved? |
| Reversion / termination record | Did ownership or control change later? |
| Permissions file | What third-party material may be used, where, and for how long? |
| AI-use terms | Is training, retrieval, translation, narration, adaptation, or data mining permitted? |
The settlement process exposes weak rights housekeeping
Settlement administration is not glamorous, but it reveals where publishing records fail. Title matching, registration status, publisher claims, author claims, reversions, and allocation disputes all require evidence. If those records live in separate inboxes, old contracts, spreadsheets, and forgotten folders, even a legitimate rightsholder can spend enormous effort reconstructing the chain.
That is why the rights file should be assembled when the book is published, not after a dispute begins.
The Rights-Provenance Minimum Record
Every title should have one record, physical or digital, linking the following information:
- Work title and internal title ID.
- Author and contributor legal names and credited names.
- ISBN or ASIN for every format.
- First-publication date by format and territory.
- Copyright claimant and registration status.
- Registration number and effective date, when registered.
- Exclusive rights granted, retained, sublicensed, or reverted.
- Third-party text, image, audio, font, data, and AI-tool permissions.
- AI-use position: prohibited, reserved, licensed, or undecided.
- Contract language governing infringement recoveries.
- Custodian and last-audit date.
Add provenance for the actual source files
Rights management is stronger when it connects legal records to production evidence. Keep the final manuscript, dated source files, cover and interior source files, contributor releases, commissioned-art agreements, image licenses, font licenses, and key correspondence. Where practical, retain checksums, version history, or repository history that can show which file was final at a given date.
For AI-assisted workflows, record the tool, function, date, and whether confidential or licensed source material was supplied to the system. Do not upload third-party content to a tool merely because the interface allows it. The permission to possess material is not automatically the permission to submit it for every machine-processing purpose.
Audit rights at lifecycle events
A title should be re-audited when a contract is amended, rights revert, a new format is created, a translation is licensed, an audiobook is produced, an edition is substantially revised, or the publisher adopts a new AI-use policy. The record should show not only the current position but the history of how it changed.
This matters for discoverability too. The AI search participation standard governs distribution choices, but a distribution decision does not replace the underlying rights analysis.
What self-published authors should do
Self-publishing can simplify the chain of title because the author may retain most rights, but it does not eliminate recordkeeping. Independent authors still need edition identifiers, registration decisions, contributor agreements, cover-image and font licenses, and proof of rights for quotations, illustrations, maps, photographs, and other third-party material.
When an author later licenses audio, translation, foreign, film, serialization, or other rights, that license belongs in the same title-level record.
What publishers should add to contracts now
Contracts should be explicit about AI-related uses rather than forcing both sides to infer intent from older clauses. Depending on the relationship, terms may address machine learning, text and data mining, retrieval systems, synthetic narration, translation, adaptation, promotional summarization, and the handling of infringement recoveries.
Clear drafting reduces later disputes. It also makes the publisher’s AI-use and disclosure record easier to maintain because the editorial team knows which uses are actually authorized.
FAQ: AI training, copyright, and publishing records
Is AI training on copyrighted books always fair use?
No universal rule can be drawn from a single case. Fair use is fact-specific, and litigation, statutes, contracts, and international rules can produce different results.
Why did copyright registration matter in the Anthropic settlement?
The settlement process used work and registration information to determine eligible works and claims. Registration also matters independently under U.S. copyright law for litigation procedure and potential remedies.
Does an ISBN prove ownership?
No. An ISBN identifies an edition or product. Ownership and control of rights come from copyright law, contracts, assignments, licenses, and reversions.
What is the minimum defensible record?
Connect the work, identifiers, registration status, chain of title, contracts, permissions, source files, AI-use position, and current custodian in one auditable record.
The paper trail is part of the product
Publishers often think of rights files as administrative overhead. In an AI-shaped market, they are part of the publishing asset itself. A clean provenance record makes licensing faster, rights reversions clearer, enforcement more defensible, and future claims easier to evaluate.
No title should be considered publication-complete until its identity, ownership, registration decision, permissions, and provenance are linked in one place.
Legal information only. This article describes publishing practices and public developments and is not legal advice. Rights questions depend on the work, contract, jurisdiction, and facts.
Primary Sources and Further Reading
- Authors Guild: Bartz v. Anthropic Settlement, What Authors Need to Know
- Authors Guild: September 17 allocation-dispute update
- U.S. Copyright Office: Copyright and Artificial Intelligence
The Four Records Every Modern Publisher Needs
This article is part of the September 19 Publishing Standards mini-series. The four records work together: discovery, rights, AI-use transparency, and accessibility.