NEW YORK / RankWire.AI / – Hachette Book Group, Cengage Learning, and Elsevier have initiated legal action against Google concerning its Gemini artificial intelligence platform. Author Scott Turow and his organization, S.C.R.I.B.E., have joined the class action lawsuit. The complaint was filed on July 10 in the U.S. District Court for the Southern District of New York. The plaintiffs accuse Google of copying millions of copyrighted books and journal articles without authorization during the development and training of Gemini. As of July 15, the court had yet to rule on the allegations or certify the class.

According to the complaint, Google obtained its material through Google Books, Google Play Books, and Google Scholar. Publishers and authors provided these works for specific purposes, including search, sales, and research. The plaintiffs argue that these agreements did not permit broader commercial AI training. They also contend that Google downloaded extensive web-scraped datasets containing copyrighted material. The filing states that some of this data originated from known pirate sources and paywalled services.
The 57-page complaint outlines four claims under federal law. Three relate to alleged reproduction through Google services, web scraping, and the development or training of Gemini. The fourth invokes the Digital Millennium Copyright Act. The plaintiffs allege Google removed or altered copyright management information from training datasets. The filing also references internal discussions about using publisher-provided books. One assessment estimated potential fines between $10 billion and $100 billion. These allegations have not yet been tested in court.
Class Action Includes Registered Works
The proposed class comprises owners of registered U.S. copyrights in qualifying books and journal articles. Eligible books must have an International Standard Book Number (ISBN), and eligible articles must have a Digital Object Identifier or International Standard Serial Number. The class includes works allegedly copied from Google services or obtained through web scraping, as well as those reproduced during Gemini’s development or training.
Eligibility also depends on registration timing. One criterion requires registration within five years of publication and before Google’s alleged reproduction or distribution. Another requires registration within three months after publication. The class excludes government entities, Google affiliates, certain court participants, and individuals who properly opt out. The court must approve the class designation before the case proceeds on behalf of the larger group.
Claims for Damages and Court-Ordered Accounting
The plaintiffs seek statutory damages or actual damages for proven infringements and request Google’s profits attributable to any confirmed copyright violations. They also ask for an injunction, recovery of legal costs, and a jury trial. The complaint does not specify a total damages amount but requests that Google disclose Gemini training data, collection methods, and known model capabilities through a court-ordered accounting.
This accounting aims to identify copyrighted works used in Gemini’s training, as well as detail how Google collected, copied, processed, and encoded those materials. The plaintiffs also seek court supervision for the destruction of any unauthorized copies under Google’s control. Earlier, Hachette and Cengage sought to join separate AI lawsuits against Google in California. The current case adds Elsevier, Turow, and S.C.R.I.B.E., while pursuing claims related to Google services, web scraping, and Gemini’s training process.
