Authors including John Grisham, David Baldacci, Jodi Picoult and Jonathan Franzen sued OpenAI and Microsoft three years ago for copyright infringement. A court filing unsealed on September 17, 2026 details internal messages and testimony laying out how OpenAI employees—including executives—talked about the company’s use of pirated books to train an early ChatGPT model, as well as their technology’s potential impact on authors.
Starting in 2019, OpenAI began using books from file-sharing site Library Genesis—LibGen for short—to help train its large language model. The site, which provides free access to books and scholarly articles, had faced global accusations of pirating copyright material. ..Tom Brown, then a top GPT-3 engineer, and Ben Mann, another member of the technical staff, described LibGen as “sketchy [as fuck] AF”. A federal court ordered LibGen to shut down in 2015 and academic publisher Elsevier received a $15 million judgment against the site in 2017.
OpenAI staffers discussed removing references to LibGen from papers that could later be public, according to the filing. In one instance, Dario Amodei, then a senior researcher at OpenAI, who is now chief executive of rival Anthropic, asked in Slack whether it was “sketchy to call our corpuses ‘Books1’ and ‘Books2’ and not say what they are, particularly when in fact they are Libgen (which is a slightly sketchy source).” …Technical staffer Tarun Gogineni wrote on X in December 2022 that he knew artists were complaining about increased AI-generated competition, but viewed potential job losses as “acceptable economic disruption.”
Excerpt from Melissa Korn et al, ‘Sketchy AF’: What to Know About How OpenAI Staff Discussed Book-Pirating, WSJ, Sept. 19, 2026