OpenAI’s ChatGPT Was Built on Concealed ‘Mass Piracy’, Authors Tell Court

Home > Piracy >

Book authors have asked a New York federal judge to rule that OpenAI built its AI models on “mass piracy”. Pointing to internal documents, a summary judgment motion alleges the AI giant downloaded books from LibGen, hid the evidence by renaming datasets, and designed its models to supplant human writers. OpenAI filed the mirror-image motion, stressing that its data harvesting qualifies as fair use.

openai logoOver the past three years, authors have filed a series of lawsuits accusing AI companies of training their models on pirated books.

Some of those cases have already produced rulings, with a bittersweet victory for Meta in California for example.

In New York, several other cases were bundled into a single proceeding where Judge Sidney Stein is overseeing claims against OpenAI and Microsoft.

This includes the Authors Guild’s class action, a case filed by a group of nonfiction writers who were the first to name Microsoft as a defendant, and the Tremblay and Silverman lawsuit, which started in California in 2023 and survived a partial dismissal before moving to New York.

This week, these authors filed a motion for summary judgment. Ahead of any trial, they want Judge Stein to rule that OpenAI copied their work without permission, and that this can’t qualify as fair use. The motion covers 194 titles and asks for a finding of liability, not damages.

“OpenAI’s GPT models pose an existential threat to those who write and publish books,” the brief states, while adding that “AI-generated books of all types are already flooding the market.”

Built on Mass Piracy

The authors start by accusing OpenAI of obtaining the book copies through unauthorized sources. While the filing is heavily redacted, OpenAI stands accused of using torrented copies downloaded from LibGen,

“OpenAI did not even buy the books it used. Instead, it began by torrenting [REDACTED] books from the notorious and illegal pirate library Library Genesis, also known as LibGen,” the motion reads.

At the time, LibGen had already been featured in the U.S. Trade Representative’s list of notorious piracy markets. According to the authors, OpenAI was well aware of the controversial nature of the site.

OpenAI “took steps to conceal their piracy from the public,” the motion notes, pointing to the paper that introduced GPT-3. In that paper, OpenAI relabeled book compilations it previously called “Libgen1” and “Libgen 2” as the more “nondescript” “Books1” and “Books2.”

“OpenAI employees understood at the time that they had sourced books from an illegal site,” the filing reads.

Concealed

concealed

The renaming was not the end of it. OpenAI “deleted its LibGen files in the summer of 2022 due to legal concerns,” the motion notes, adding that these are “the only two training corpuses OpenAI has ever deleted.”

Before deleting the books, OpenAI allegedly used them to train the early GPT models. Or as the authors write, the company “built the foundations of its business on mass piracy.”

Replacing George R.R. Martin

The torrenting and piracy angle is one part of the filing. The motion also alleged that OpenAI built its models to replace the human writers it copied, and as evidence it highlights controversial tweets from a key employee.

In 2022, OpenAI hired Tarun Gogineni to lead its work on the writing quality of its models. According to the motion, Gogineni knew the models he was training would displace authors but considered that “acceptable economic disruption.”

This is notable because Gogineni specifically mentioned one of the plaintiffs, author George R.R. Martin, known for writing A Song of Ice and Fire which the HBO series Game of Thrones was based on.

In 2025, nearly two years after Martin sued, Gogineni tweeted that his “research mission” was to have GPT models write the “last two books of [Martin’s] A Song of Ice and Fire.”

Even if…

martin

Even if Martin “dies early, GPT-5 will autocomplete his series,” he added, suggesting that AI can replace the author.

Not Fair Use

OpenAI and other AI companies argue that training models on books is fair use. Courts have partly agreed with this, but with an important caveat.

The authors cite Bartz v. Anthropic, the 2025 California ruling that classified model training as potentially fair use, while stressing that downloading from a pirate library was not. Pirating books that can be purchased legally is “inherently, irredeemably infringing,” that court found.

The authors also argue that the copying was avoidable for training purposes, as their books were not per se necessary to create a general-purpose model.

Broader Claims

The motion is not limited to OpenAI. It also asks the court to hold that Microsoft is vicariously liable for OpenAI’s copyright infringement, since Microsoft could supervise the conduct and profited from it.

Microsoft invested roughly $13 billion across three agreements signed in 2019, 2021, and 2023, the authors stress.

conclusion

OpenAI has yet to respond to the authors directly, but it clearly believes that the evidence points in its favor.

In a cross-motion for summary judgment, filed on the same day, the company argues that its use of the books was fair use as a matter of law and that any regurgitation is vanishingly rare.

The filings highlighted here are part of a much broader push. Over the past days, plaintiffs including The New York Times, Daily News, and the Center for Investigative Reporting all submitted a combined summary judgment motion of their own against OpenAI and Microsoft.

With many millions of dollars at stake, as well as the future of AI training, these cases will be fought tooth and nail, so we certainly haven’t heard the last of it.

A copy of the authors’ redacted motion for partial summary judgment is available here (pdf), filed at the U.S. District Court for the Southern District of New York.

Sponsors

proton vpn promo promotion

PIA logo and service promotion

NordVPN logo and service promotion