COPYRIGHT NEWS: USA TODAY and…


By Saurabh Kashyap, B.A., M.A., LL.B., LL.M.

The publications alleged that OpenAI copied their journalism for AI training, generated outputs that substitute for their reporting, and removed identifying copyright information.

USA TODAY Co. and 18 affiliated national and local news publications have sued OpenAI, alleging that it copied hundreds of thousands of protected articles and other journalistic works without permission to develop, train, and operate its generative AI products. The publications seek more than $250 million in damages, injunctive relief, and the destruction of GPT models and training datasets that incorporate their content (USA Today Co., Inc. v. OpenAI Foundation, No. 1:26-cv-08892 (S.D.N.Y. Oct. 8, 2026)).

The action concerns content published by USA TODAY and affiliated publications owned by the plaintiffs: The Tennessean, Indy Star (Indianapolis Star), The Bergen Record, The Enquirer, Asbury Park Press, Democrat & Chronicle, The Knoxville News-Sentinel, Naples Daily News, The Oklahoman, Milwaukee Journal Sentinel, The Columbus Dispatch, The Arizona Republic, The Courier-Journal, The Des Moines Register, Detroit Free Press, The Detroit News, The Palm Beach Post, and StarNews. The publications asserted direct and vicarious copyright infringement claims, as well as a claim under 17 U.S.C. § 1202 for removal or alteration of copyright-management information.

The complaint is the latest in the copyright litigation confronting generative AI companies over the material used to train large language models. OpenAI had not yet responded to the allegations when the action was filed.

Copied content allegedly used for model training. The publications alleged that OpenAI copied their reporting while creating and using datasets to train its GPT models, including the WebText, WebText2 and Common Crawl-derived datasets.

According to the complaint, the WebText dataset contained more than 160,000 entries from the publications’ websites, including 83,266 entries from usatoday.com. The publications also alleged that their domains represented more than 122 million tokens in the C4 dataset, a filtered English-language subset of a 2019 Common Crawl snapshot.

The complaint maintained that OpenAI treated high-quality news content as particularly valuable in training. It cited OpenAI materials describing datasets regarded as higher quality as being sampled more frequently, and alleged that the company curated datasets of current news material to improve the freshness of its models’ answers.

The publications alleged that OpenAI copied their works during the acquisition, storage, and processing of the datasets, and again during pre-training, training, and fine-tuning its models. They said OpenAI neither sought licenses nor offered compensation for using their content. The complaint also alleged that OpenAI continues to obtain new copyrighted journalism, citing its models’ changing knowledge-cutoff dates and the demand for current information.

The action identifies a broad range of GPT models and products, including ChatGPT, ChatGPT Plus, enterprise offerings, and OpenAI’s API platform. It alleges the models were trained on material obtained directly from the publications’ websites or through third-party datasets containing that material.

Outputs allegedly reproduce or supplant reporting. The publications also alleged that OpenAI’s products reproduce protected works, create derivative works, or generate detailed paraphrases and summaries that replace the need to access the original reporting.

The complaint asserted that the GPT models memorized portions of their training data and could generate near-verbatim text when prompted. It alleged that, before certain safeguards were implemented, OpenAI products could return verbatim or near-verbatim copies of copyrighted articles. The publications cited internal communications they say showed OpenAI’s awareness that models could regenerate protected works.

Beyond alleged memorization, the complaint challenges the use of retrieval-augmented generation, or RAG, in OpenAI’s search-related products. It alleged that these tools retrieve source material from search indexes or the internet in real time, provide that material as context to a model, and produce a narrative answer that can perform the same informational function as the underlying article.

The publications said these responses go beyond ordinary search-result snippets. Although a response may include links to source material, it allegedly provides enough expressive content that readers have less reason to click through to the publisher’s website, subscribe, or view advertising there.

That conduct, the complaint alleged, harms the publications’ subscription, advertising, and content-licensing businesses. It also contended that unauthorized outputs threaten the publications’ reputations because readers may receive incomplete, inaccurate, or unattributed accounts of their reporting.

Alleged circumvention of publisher restrictions. The publications asserted that they use copyright notices, paywalls, terms of service and robots.txt files to protect their journalism. Their terms of service allegedly prohibit copying or exploiting content to develop or improve AI systems, including through model training, fine-tuning, grounding, and retrieval-augmented generation.

The complaint further alleged that the publications blocked OpenAI web crawlers through robots.txt files and directed detected AI crawlers to a landing page stating that crawling and scraping were not permitted. It accused OpenAI of failing to review website terms of use when collecting web content and of failing to identify or remove paywalled content from training datasets.

The publications also cited OpenAI’s licensing arrangements with other news and media organizations, alleging that those arrangements show OpenAI recognizes journalism’s commercial value and the need to obtain permission to use it.

Copyright-management information claim. A separate claim alleges that OpenAI removed copyright-management information from the publications’ works in violation of 17 U.S.C. § 1202. The asserted information includes copyright notices, bylines, article titles, publisher names, terms of use, and other identifying material.

The complaint alleged that OpenAI used text-extraction tools called Dragnet, Newspaper, and Gutentag to construct training datasets. These tools allegedly isolate article text while removing what they treat as extraneous material, including headers, footers, copyright notices and, in some circumstances, author and title information.

The publications alleged that OpenAI knew these tools operated in that manner and deliberately created datasets containing their articles without the associated copyright-management information. They further alleged that OpenAI distributed outputs containing copies or derivatives of journalism without restoring that information.

According to the complaint, removing this information both conceals alleged infringement and enables further infringement by end users, who may not know an output reproduces protected material from a publication.

Requested relief. The publications seek statutory damages, actual damages, restitution, disgorgement of profits, costs, and attorneys’ fees. They alleged that the infringement was willful. The complaint states that statutory damages for willful copyright infringement may reach $150,000 per work and that the DMCA provides additional statutory damages for violations involving copyright-management information.

They also seek a declaration that OpenAI’s past, present, and future use of their content is unlawful; a permanent injunction against the alleged conduct; and an order under 17 U.S.C. § 503(b) requiring destruction of GPT or other large language models, as well as training sets, that incorporate their content.

The Case is No. 1:26-cv-08892.

Attorneys: Jennifer Maisel (Rothwell Figg Ernst & Manbeck, P.C.) for USA Today Co., Inc.

Companies: USA Today Co., Inc.; OpenAIFoundation

MainStory: TopStory AINews NewYorkNews TechnologyInternet Copyright GCNNews



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *