OpenAI accidentally deleted potential evidence in NY Times copyright lawsuit

Date:

Share post:


Lawyers for The New York Times and Daily News, which are suing OpenAI for allegedly scraping their works to train its AI models without permission, say OpenAI engineers accidentally deleted data potentially relevant to the case.

Earlier this fall, OpenAI agreed to provide two virtual machines so counsel for The Times and Daily News could perform searches for copyrighted content in its training data sets. In a letter, attorneys for the publishers say that they and experts have spent over 150 hours since November 1 searching OpenAI’s training data.

But on November 14, OpenAI engineers erased all the publishers’ search data stored on one of the virtual machines, according to the aforementioned letter, which was filed in the U.S. District Court for the Southern District of New York late Thursday.

OpenAI tried recover the data — and was somewhat successful. However, because the folder structure and file names were “irretrievably” lost, the recovered data “cannot be used to determine where the news plaintiffs’ copied articles were used to build [OpenAI’s] models,” per the letter.

“News plaintiffs have been forced to recreate their work from scratch using significant person-hours and computer processing time,” counsel for The Times and Daily News wrote. “The news plaintiffs learned only yesterday that the recovered data is unusable and that an entire week’s worth of its experts’ and lawyers’ work must be re-done, which is why this supplemental letter is being filed today.”

The plaintiffs’ counsel makes clear that they have no reason to believe the deletion was intentional. But they do say the incident underscores that OpenAI “is in the best position to search its own datasets” for potentially infringing content using its own tools.

We’ve reached out to OpenAI for comment and will update this piece if we hear back.

In this case and others, OpenAI has maintained that training models using publicly available data — including articles from The Times and Daily News — is fair use. In other words, in creating models like GPT-4o, which “learn” from billions of examples of ebooks, essays, and more to generate human-sounding text, OpenAI believes that it isn’t required to license or otherwise pay for the examples — even if it makes money from those models.

That being said, OpenAI has inked licensing deals with a growing number of new publishers, including The Associated Press, Business Insider owner Axel Springer, Financial Times, People parent company Dotdash Meredith, and News Corp. OpenAI has declined to make the terms of these deals public, but one content partner, Dotdash, is reportedly being paid at least $16 million per year.



Source link

Lisa Holden
Lisa Holden
Lisa Holden is a news writer for LinkDaddy News. She writes health, sport, tech, and more. Some of her favorite topics include the latest trends in fitness and wellness, the best ways to use technology to improve your life, and the latest developments in medical research.

Recent posts

Related articles

DOJ: Google must sell Chrome to end monopoly

The United States Department of Justice argued Wednesday that Google should divest its Chrome browser as part...

WhatsApp will finally let you unsubscribe from business marketing spam

WhatsApp Business has grown to over 200 million monthly users over the past few years. That means there...

OneCell Diagnostics bags $16M to help limit cancer reoccurrence using AI

Cancer, one of the most life-threatening diseases, is projected to affect over 35 million people worldwide in...

India’s Arzooo, once valued at $310M, sells in distressed deal

Arzooo, an Indian startup founded by former Flipkart executives that sought to bring “best of e-commerce” to...

Nvidia’s CEO defends his moat as AI labs change how they improve their AI models

Nvidia raked in more than $19 billion in net income during the last quarter, the company reported...

Snowflake snaps up data management company Datavolo

Cloud giant Snowflake has agreed to acquire Datavolo, a data pipeline management company, for an undisclosed sum....

Solar power magnate Gautam Adani and others indicted over alleged $250M bribery scheme

Billionaire Gautam Adani and several executives at his company, the Indian conglomerate Adani Group, have been indicted...

‘PDF to Brainrot’ study tools are a strange iteration on a TikTok trend

Several AI-based study tools are capitalizing on a “PDF to Brainrot” trend, which will read the text...