News

Delhi High Court Holds AI Training to Be Prima Facie Fair Dealing under the Copyright Act

For decades, copyright law has largely been concerned with works that are copied, distributed or communicated without the author’s permission. Generative artificial intelligence shifts the enquiry to an earlier stage. The central question is no longer whether an AI-generated output reproduces a protected work, but whether the use of copyrighted material to train an artificial intelligence model itself engages the exclusive rights of a copyright owner. The Copyright Act, 1957 offers no express answer to this question.

The Hon’ble Delhi High Court in ANI Media Pvt. Ltd. v. Open AI OpCo LLC, CS(COMM) 1028/2024 (I.A. 45300/2024), decided on 24 July 2026 by Justice Amit Bansal, examined this gap: whether AI training could amount to copyright infringement and, if so, whether it may nevertheless fall within the fair dealing exception under Section 52 of the Copyright Act, 1957.

Understanding the functioning of LLMs
Before turning to the parties’ contentions, the Court examined how Large Language Models (LLMs) function, since the issues raised turned directly on the way these models operate. The Court explained that LLMs are trained on vast datasets collected from multiple sources. During training, the data is broken into smaller units or “tokens”, converted into numerical representations known as embeddings, and processed to identify linguistic patterns and relationships. Rather than storing or retrieving complete articles, the model learns statistical relationships between words and predicts the next most likely token in response to a user’s prompt. The Court also explained that responses may, in certain cases, be supplemented through Retrieval-Augmented Generation (“RAG”), where the model retrieves information from external sources before generating an answer.

The Jurisdiction Issue
The Court took up jurisdiction first, since the objection went to the root of the matter. It held, prima facie, that it had territorial jurisdiction under Section 62(2) of the Copyright Act and Section 20 of the Code of Civil Procedure, 1908: ANI’s principal place of business and registered office lie within its jurisdiction, OpenAI targets and offers its services to users across India, and the outputs complained of were generated there. The Court also rejected OpenAI’s contention that the Copyright Act could not reach training carried out on US servers, reasoning that storage abroad was only the terminal step in a chain of events beginning with access to the works from India, and that accepting the argument would allow infringers to evade Indian copyright law by shifting that last link offshore.

Two Different Copyright Questions
On the substantive claims, the Court noted that there were two distinct issues raised by ANI (i) whether the storage by OpenAI of ANI’s data for training ChatGPT would amount to infringement of ANI’s copyright; and (ii) whether the use of ANI’s copyrighted data to generate responses for users would amount to infringement of ANI’s copyright.

On examining the issue regarding ChatGPT’s responses, the Court compared ANI’s articles with ChatGPT’s responses and noted that even the titles in the response generated by ChatGPT were substantially different from the titles of the articles on ANI’s website. The Court further applied the test for infringement laid down in R.G. Anand v. Deluxe Films, (1978) 4 SCC 118 and stated that there was no substantial similarity. It further observed that the illustrative articles relied upon by ANI had been published after the completion of the relevant training period and, therefore, “there cannot be any question of memorization.”

Is AI Training Copyright Infringement?
The more significant question concerned OpenAI’s use of ANI’s articles during the training of its LLMs. The Court held that the storage of ANI’s original literary works in the course of training the LLMs underlying ChatGPT amounts to reproduction under Section 14(a)(i) of the Copyright Act. It went on to consider whether OpenAI’s use was protected as fair dealing for “private or personal use, including research” under Section 52(1)(a)(i) of the Copyright Act.

Noting that Indian courts have adopted no single uniform test for fair dealing, and that there was broad consensus among counsel that the US four-factor test does not apply in India, the Court formulated three factors suited to the facts before it: whether OpenAI’s use of ANI’s works is limited to training the LLMs underlying ChatGPT; whether that use results in economic competition and prejudices ANI’s legitimate interests, causing actual or potential damage; and whether the functions ChatGPT performs serve the overall public interest. The Court observed that these factors are broadly in line with the principles emerging from Article 9 of the Berne Convention. The Court also rejected ANI’s argument that a profit-making entity is automatically barred from relying on Section 52(1)(a)(i), reasoning that Parliament had expressly confined certain other exceptions in Section 52 to non-commercial use and had not done so here.

Applying the above principles, the Court found no material at the interim stage to conclude that OpenAI had used ANI’s works for any purpose other than training its LLMs. It further observed that there was neither any averment in the plaint nor material on record to show that OpenAI’s activities had caused ANI to lose subscribers or suffer loss in its news syndication business. The Court also treated public interest as an established fairness factor, holding that the societal benefits flowing from LLMs i.e. improving access to information, supporting education and scientific research, enabling translation, and creating tools for persons with disabilities are relevant considerations in assessing whether a dealing is fair. The Court also held that the exception is not confined to research undertaken by human beings: to say otherwise, it observed, would be a regressive view that limits societal progress, and “research” therefore extends to machine learning. Accordingly, the Court held that OpenAI had, prima facie, established that its use of ANI’s works for training its LLMs was protected under Section 52(1)(a)(i).

Key Takeaways

  • For AI developers: On the Court’s prima facie view, the storage of copyrighted literary works for the training of LLMs amounts to reproduction under Section 14(a)(i) of the Copyright Act. However, where the requirements of Section 52(1)(a) are satisfied, such acts of storage may fall within the fair dealing exception and would not amount to infringement under Section 51. Because these findings were made at the interim stage on a prima facie standard, they do not settle the position at trial or on appeal.
  • For copyright owners: Copyright infringement cannot be established merely because an AI model generates responses based on copyrighted material. The output must satisfy the test of substantial similarity, and, on the facts of the case, ANI failed to establish a prima facie case of copyright infringement in respect of ChatGPT’s responses.

What happens next?

The suit continues. The Court expressly recorded that its observations were made only for the purpose of deciding the interim application and would have no bearing on the final outcome. It also held that whether OpenAI permanently stores training data in order to memorise and regurgitate ANI’s works is a disputed question that can only be determined at trial on evidence. An appellate court could yet take a different view of Section 52(1)(a).