A coalition of more than a dozen US news publishers filed a formal sanctions motion against OpenAI in a Manhattan federal court on Thursday, accusing the AI company of concealing its ability to search its own systems for proof that it used their journalism without permission to train ChatGPT. The group, led by the New York Times, also alleges that OpenAI deleted billions of relevant ChatGPT conversation logs or rendered them unsearchable during the ongoing litigation.
The case, which began when the Times sued OpenAI and Microsoft in December 2023, has grown to include the New York Daily News, the Center for Investigative Reporting, the Intercept, and digital publisher Ziff Davis. Reuters reports that the publishers told the court OpenAI falsely claimed it lacked the tools to search its large language models for their copyrighted material, while hiding the fact that it had already conducted such searches before the first plaintiff filed suit.
A deposition changes the picture
The turning point came during a court-ordered deposition of Vinnie Monaco, OpenAI's head of privacy engineering. According to TechCrunch, Monaco's testimony allegedly revealed that OpenAI had already built an internal database of around 78 million de-identified ChatGPT conversations it was using to assess how much it was reproducing others' content. Separately, OpenAI allegedly developed a tool called a "Bloom" filter as part of an internal initiative known as "Project Giraffe", which tracked and recorded when its outputs reproduced text from its training data — launched shortly after the original lawsuit was filed.
“"If OpenAI genuinely believed that copying our clients' journalism was fair and legal, it wouldn't have hid the truth about having done it," — Ian B. Crosby, lead counsel for the plaintiffs”
The publishers are asking the court to bar OpenAI from using a sample of 20 million ChatGPT conversations it has offered as evidence, arguing it is unreliable given the alleged deletions. They also want the court to treat it as established fact that the full logs would have shown substantial and systematic reproduction of their copyrighted articles, and to require OpenAI to cover their legal costs in pursuing the withheld material.
OpenAI pushes back on privacy grounds
OpenAI has consistently argued that handing over ChatGPT conversation logs would violate the privacy of users who have nothing to do with the case, and that searching its systems at the scale requested would be technically prohibitive. Spokesperson Drew Pusateri rejected the publishers' motion as "blatantly false allegations," saying the company would continue defending "the long-established principles of fair use." Fair use is a legal doctrine in US copyright law that allows limited use of protected material under certain conditions, including for purposes such as commentary, research, or transformation.
“"This motion asks the court to punish OpenAI for hiding and destroying evidence showing how ChatGPT was trained on stolen journalism" — Steven Lieberman, attorney for the New York Daily News”
The broader stakes extend well beyond the parties involved. The dispute sits at the intersection of two fragile industries: a news sector facing structural decline in advertising revenue and digital traffic, and an AI industry whose legal right to train on publicly available text remains unsettled. OpenAI and other AI developers argue that ingesting online text is shielded by fair use, a theory being contested simultaneously in dozens of lawsuits brought by authors, visual artists, and music labels. The largest settlement to date came when Anthropic agreed to pay book authors 1.5 billion dollars over the use of their works to train its Claude models, though that figure is a fraction of Anthropic's reported valuation. The Times has spent more than 28 million dollars on its AI-related litigation to date, with 4.2 million spent in the first quarter of 2026 alone.
Meanwhile, the media industry's relationship with AI companies remains deeply divided. Many publishers have chosen licensing agreements with OpenAI, Google, and Meta, accepting fees in exchange for allowing their archives to be used in AI training. The Associated Press was among the first to sign such a deal with OpenAI, in 2023. The Times has taken a different path, pursuing litigation while also striking a separate content licensing agreement with Amazon for AI-related uses, illustrating the commercial and legal tightrope many news organisations are walking as AI chatbots increasingly compete with them as a first stop for information.
This article is free to read. It always will be — no paywall, no account, no tracking.




