Hachette Book Group, Cengage Learning, Elsevier, and author Scott Turow have filed a lawsuit against Google in a federal court in New York, alleging the tech giant committed copyright infringement while training its Gemini AI models.

The nearly 60-page complaint, filed on Friday, claims that "Google willfully sidestepped this longstanding system designed to protect copyrights and compensate authors and publishers through a series of deliberate choices to develop Gemini1."

The plaintiffs allege Google initially copied books as source material via Google Books, using them beyond the "strictly limited purposes" originally permitted for Google Books and other services.

Kirk Sigmon, a technology and IP law expert, told Al Jazeera that any fair use defense by Google could be undermined if the books were acquired unlawfully. He noted the complexity of proving what was included in AI training datasets.

Oli Huggins, CEO of ExpertEdge and VP of Partnerships at Packt Publishing, highlighted the difficulty in proving whether AI-generated outputs constitute copyright infringement once information has been used for training.

Separately, CNN filed a lawsuit against AI company Perplexity, alleging it illegally copied over 17,000 stories to train its models, which produced content "identical or substantially similar to CNN’s content," according to a complaint filed in May.

Legal experts also point out unresolved questions about liability, with one noting, "If the user is actively trying to get the model to infringe, that could mean the user is ultimately on the hook rather than the AI system. This is still an open question that courts in the United States haven’t really grappled with."

Sources