The case, explained

Copyright and AI Training: Current Trends in Content Protection for 2025

6 min read · Updated June 2026 · Editorial oversight: Avv. Federico Papa

The issue of using protected content for training large language models reached a critical turning point in late 2025. According to reports in the national and international press between late 2023 and late 2025, the clash between publishing giants and AI labs has evolved into a complex legal battle over the lawfulness of data extraction. This article analyzes the regulatory and judicial evolution, starting from well-known cases to provide an operational summary for professionals. By reconstructing the facts and examining the legal framework updated with the recent Law 132/2025, we will see how the Italian legal system is addressing the challenges of GenAI. We will conclude with a didactic twin case to illustrate the practical application of opt-out protections and possible defense strategies for copyright holders.

In brief

The article examines the conflict between copyright holders and AI developers, focusing on the New York Times v. OpenAI case and actions taken by RTI and Medusa Film in Italy. It analyzes EU Directive 2019/790, the AI Act, and Law 132/2025, explaining Text and Data Mining exceptions and the opt-out mechanism. Through an illustrative twin case, it outlines the burden of proof and potential civil remedies for copyright infringement.

  1. The facts

    Recent legal news is marked by the clash between content creators and tech companies. According to reports from Wired Italia and Il Messaggero, the pivotal case began on December 27, 2023, when the New York Times sued OpenAI and Microsoft in the Manhattan Federal District Court. The allegation is the use of millions of protected articles to train models like GPT-4 without a license.

    Currently, the proceedings are in the discovery phase, with the court having already rejected some early dismissal attempts. In Italy, outlets like Diritto.it highlighted similar actions launched in December 2025 by RTI (Mediaset) and Medusa Film against AI developers for unauthorized use of audiovisual content. Publishers argue that AI does not just learn but reproduces nearly identical textual fragments (regurgitation), engaging in direct commercial competition with the original sources.

  2. The laws in play

    The regulatory framework centers on Directive (EU) 2019/790 and Law 633/1941 (Italian Copyright Law, LDA). Article 4 of the Directive introduces the Text and Data Mining exception (TDM) for commercial purposes, implemented in Italy by Article 70-quater LDA. This rule allows data extraction unless the rights holder has expressed a reservation (opt-out) in a machine-readable format.

    Recently, Law 132/2025 introduced Article 70-septies LDA, specifying obligations for AI models and referencing the transparency required by EU Regulation 2024/1689 (AI Act). Violation of these rules can lead to injunctions against dataset use, damages under Article 158 LDA, and potentially, an order to destroy model weights trained unlawfully.

  3. Jurisprudence

    According to analyses by specialized publications like Diritto.it, European lower court jurisprudence is beginning to define the boundaries of TDM. Recent rulings clarified that the exception for scientific research purposes cannot be invoked by entities that, despite being non-profit, create datasets intended for use by commercial partners to develop competing products.

    It has also been specified that the reservation of rights must be respected not only if included in server configuration files (such as robots.txt) but also if indicated in the terms and conditions of the website, provided they are accessible. Judges have emphasized that when an AI system's output can replace consultation of the original work, the use falls squarely within a violation of reproduction rights.

  4. Analysis drafted and verified with edit.legal

    To verify the provisions cited in this article, we used edit.legal. Test our legal AI on official sources and apply it to your own matters.

    Try edit.legal AI
  5. Lessons for professionals

    1. Technical implementation: For legal practitioners, a contractual clause alone is insufficient; it is essential to implement machine-readable opt-out protocols.

    2. AI Act Transparency: Leverage the new transparency obligations to request access to training logs and summary documentation.

    3. Output analysis as evidence: Monitor the output of competing models to identify cases of memorization or reproduction of distinctive elements, providing strong evidence of direct infringement.

References: Direttiva (UE) 2019/790Regolamento (UE) 2024/1689 (AI Act)Legge 22 aprile 1941, n. 633 (LDA)Legge 132/2025

Avv. Federico Papa
Editorial oversight: Avv. Federico Papa·ICAM

Frequently asked questions

What penalties do those who train AI on protected data face?

Penalties are primarily civil remedies: damages, injunctions, and in some cases, judicial orders to delete the dataset or perform machine unlearning.

Is it enough to write "AI use prohibited" on one's website?

According to EU regulations, the reservation of rights must be expressed in a machine-readable format to be enforceable against commercial data mining activities.

Does the AI Act protect Italian copyrights?

The AI Act imposes transparency obligations on developers, requiring them to declare the data used for training, thereby assisting rights holders in enforcing the protections available under Italian law.

Verified legal research and drafting with edit.legal

Legal research and drafting with citations checked against official databases. edit.legal is free to try, no credit card.

Try edit.legal for free