AI Training and Copyright under Indian Law: Decoding ANI v. OpenAI

The rapid development of artificial intelligence has raised several complex questions regarding the application of existing copyright laws to emerging technologies. One such question concerns whether the use of copyrighted works for training large language models (“LLMs”) constitutes copyright infringement or can be protected as fair dealing. The Delhi High Court’s detailed order in ANI Media Pvt. Ltd. v. OpenAI OpCo LLC, CS(COMM) 1028/2024, dated 24 July 2026, although rendered while adjudicating an interim injunction application, is significant because it addresses how existing copyright principles may apply to AI training and generated outputs in India.

The Case

The suit filed by ANI against OpenAI pertains to the alleged unauthorized use of ANI’s copyrighted works by OpenAI for the purpose of training its LLMs. ANI primarily claims copyright infringement on two grounds: first, the copying and storage of ANI’s copyrighted works for training the LLMs; and second, the reproduction of ANI’s copyrighted works in the outputs generated by OpenAI’s LLM through its application, ChatGPT. Accordingly, ANI sought an interim injunction against OpenAI, restraining it from any further use, copying, or reproduction of ANI’s copyrighted works.

The Hon’ble Delhi High Court, considering the nature of the suit, emphasised the need for the law to evolve alongside technological advancements and observed that the legislature could not have anticipated the rapid developments in the field of artificial intelligence. The Court further observed that, while navigating such uncharted areas of law, guidance may be drawn from decisions of foreign courts, provided that such decisions are considered in the context of Indian intellectual property law and are not contrary to the legislative intent underlying the Copyright Act, 1957 (the “Act”).

For ease of adjudication, the Court framed four main issues for the interim application:

i. Whether this Court has jurisdiction to entertain the suit, considering that OpenAI’s servers are located in the United States of America.

ii. Whether the use of copyrighted data to generate responses for users amounts to infringement.

iii. Whether storage of ANI’s copyrighted data for training purposes amounts to infringement.

iv. Whether OpenAI’s use of copyrighted data falls within the ambit of fair dealing under the Act.

Key Holdings

The Delhi High Court, at the prima facie stage, held that it had territorial jurisdiction over the dispute and examined whether OpenAI’s use of ANI’s copyrighted works for training and generating outputs amounted to copyright infringement. On the reproduction claim, the Court found insufficient material to establish that OpenAI’s models had memorised and substantially reproduced ANI’s works, noting that some of the examples relied upon by ANI post-dated the relevant model training cut-off dates and were instead attributable to Retrieval-Augmented Generation (RAG). On the training issue, the Court interpreted Section 52(1)(a) of the Act broadly and, applying the doctrine of updating construction, prima facie held that the use and storage of copyrighted works for training LLMs could fall within “private or personal use, including research.” The Court further found the use to satisfy the fairness test, considering that it was limited to training, did not prima facie compete with or cause market harm to ANI, and served broader public-interest purposes. Accordingly, the Court found the requirements of fair dealing under Section 52(1)(a) prima facie satisfied and declined to grant an injunction in ANI’s favour.

Analysis

I. Purpose-based Interpretation of Section 52

The purpose-based interpretation considered by the Court is a well-established approach to statutory interpretation. Courts have repeatedly relied on this approach to allow legislation to apply where the law has not evolved alongside the relevant subject matter. The Court’s approach in this case may be justified, as it expands the jurisprudence around copyright and the legislative history of Section 52 of the Act in an evolving field such as generative AI. This becomes relevant because the Act does not expressly provide for the use of copyrighted works for training large language models (“LLMs”), requiring courts to apply existing copyright principles to a substantially different technological context.

The purpose of copyright law is to encourage the creation and dissemination of knowledge, while also protecting authors against unlawful exploitation of their works. These objectives are inextricably linked. The Court took this relationship into consideration while interpreting Section 52 of the Act.

II. Balancing Copyright Protection and Access to Knowledge

This understanding finds support in Eastern Book Company v. D.B. Modak,[1] where the Supreme Court recognised that copyright law balances authors’ interests with the public interest in preserving the public domain. In its discussion of originality, the Court also considered the public interest in maximising the production and dissemination of intellectual works. The significance of this reasoning is that access to knowledge forms part of the justification for copyright protection itself. For example, authors, researchers and students examine copyrighted works to develop their understanding and produce further work. While Section 14 of the Act protects authors by conferring exclusive rights, Section 52 balances those rights by permitting specified uses without the copyright owner’s consent.

III. Interpreting “Research” to Include “Training” for AI

Section 52 of the Act accommodates several purposes, including private or personal use, research, criticism and review, news reporting, education, religious ceremonies, judicial or legislative use, and library preservation. For this order, the focus is Section 52(1)(a)(i) of the Act, which permits fair dealing for “private or personal use, including research”.

The Court’s reliance on an updated interpretation of “research” to accommodate technological advancements can be understood as consistent with the provision’s underlying objective. By applying the doctrine of updating construction, the Court did not restrict “research” to research undertaken directly by an individual. Instead, it considered whether research undertaken through an AI system, which ultimately serves and benefits humans, could also fall within the provision. This interpretation also raises an important cost-benefit question. LLMs are trained on vast quantities of copyrighted material, and requiring developers to identify copyright owners and obtain individual licences for every work used in training could involve substantial transaction costs and may make large-scale AI training commercially difficult. On the other hand, permitting such use under fair dealing without licensing may deprive copyright owners of the opportunity to commercially license their works for AI training, particularly as a market for such licences develops. The issue therefore requires a balance between the cost and practical feasibility of licensing training datasets, the economic interests of copyright owners, and the broader benefits arising from the development of LLMs. The Court’s interpretation provides one way of striking this balance under the existing framework, although whether such a balance should ultimately be determined through fair dealing or a specific statutory licensing framework remains open.

This approach finds support in the legislative history of the amendments to Section 52. The Department-related Parliamentary Standing Committee on Human Resource Development’s 227th Report on the Copyright (Amendment) Bill, 2010 records the Department’s explanation that the exceptions were being extended for education and research because works were increasingly available in digital formats and on the internet.[2] The Department further explained that the proposed provisions were intended to accommodate the introduction of new technologies. The order’s approach therefore finds support in this stated objective of accommodating technological change.

IV. Beyond AI Training: Implications of RAG and Generated Outputs

Another important aspect highlighted in the order is whether the protection available to copyrighted works used for training extends to outputs generated through Retrieval-Augmented Generation (“RAG”). While the Court considered whether the incorporation of copyrighted works as training data could qualify as “research” under the fair dealing framework, RAG introduces a separate dimension to the copyright analysis. Unlike model training, where copyrighted material is processed during the development of the model, RAG enables an AI system to retrieve information from an external source and use that information while generating a response. Therefore, even where copyrighted material was not contained in the training dataset, its retrieval and substantial reproduction in a generated output could raise an independent question of copyright infringement.

Accordingly, distinguishing RAG from training does not necessarily resolve the copyright question. The relevant inquiry at the output stage should focus on whether protected expression has been substantially reproduced and whether the resulting output substitutes for, or otherwise prejudices the legitimate interests associated with, the original work. Thus, while the order provides some clarity on the use of copyrighted works at the training stage, the treatment of copyrighted material retrieved and reproduced through RAG remains an important issue requiring separate consideration.

Implications

The order has implications for two stakeholder groups:

a) AI Developers

The decision provides no blanket protection for training activities or generated outputs. The Court’s jurisdictional reasoning also indicates that foreign AI companies may face proceedings in India despite locating their training servers abroad, where the relevant activities and allegedly infringing outputs have a sufficient connection with India.

Further, the order potentially reduces one of the major economic barriers associated with LLM training—the transaction costs of identifying right holders and negotiating licences across datasets containing millions of copyrighted works. A requirement to obtain individual licences for every work may disproportionately increase the cost of model development, particularly for smaller developers,[3] and could strengthen the position of established AI companies that possess the resources to acquire large, licensed datasets. However, the missing piece is legal and economic certainty regarding when access to copyrighted material should require payment. If a functioning market for licensing high-quality content for AI training develops, the argument that training does not affect the copyright owner’s market becomes more difficult, since the relevant market may itself be the market for training licences.[4] Developers therefore need clarity not merely on whether training can constitute “research”, but on whether lawful access, the availability of licences, the substitutability of training data, and the commercial value derived from particular datasets affect the fairness analysis. Greater transparency and standardised licensing mechanisms could reduce these transaction costs, but overly granular disclosure or licensing requirements may themselves increase compliance costs and barriers to entry.

b) Copyright Owners

On the other hand, for copyright owners, the order highlights the importance of evidence concerning substantial reproduction, market substitution or prejudice to their legitimate interests when contesting a fair-dealing defence. Establishing that protected works were used for training may not, by itself, defeat that defence. Copyright owners may therefore need to examine both how their works are used during training and whether generated outputs reproduce protected expression or substitute for their works. The prima facie nature of the findings leaves these issues open to further examination in subsequent proceedings in this case.

The economic concern extends beyond whether an AI service directly competes with the market for the original work. Copyrighted works may themselves constitute economically valuable inputs into AI development, particularly where the content is specialised, reliable or difficult to substitute. If a market for licensing such works for AI training develops, uncompensated use may potentially affect a distinct licensing opportunity even where the AI product does not directly substitute for the original work.[5] However, the missing piece for copyright owners is information and bargaining power. An author may be unable to determine whether their work formed part of a training dataset, how it contributed to the model, or whether its use generated economic value, making both licensing and enforcement difficult. Further, the benefits of an emerging licensing market may not be distributed equally: large publishers and content owners can aggregate valuable datasets and negotiate directly with AI developers, while individual authors may face considerably higher enforcement costs and weaker bargaining positions. The future economic question is therefore not simply whether copyright owners should be compensated, but how a licensing or rights-reservation system can ensure that compensation reaches individual creators without creating transaction costs that make large-scale AI development commercially impractical.

Next Hearing in ANI v. OpenAI

The case is next to be heard on 27 January 2027, as per the Delhi High Court website, while ANI’s appeal against the 24 July 2026 interim order is separately scheduled to be heard next on 8 December 2026.

Way Forward

While the order considers a technologically adaptive approach to copyright law, its interpretation of Section 52 remains open to challenge on appeal. In Union of India v. Hansoli Devi, the Supreme Court’s Constitution Bench held that where statutory language is plain and unambiguous, courts must give effect to the words used and cannot adopt a different construction merely because it appears more consistent with the supposed object and policy of the legislation.[6]

Applying this principle, an appellate court could read “including research” in Section 52(1)(a)(i) as qualified by the preceding words “private or personal use”, rather than as an independent category covering research of every kind. If this reading is accepted, the broader benefits of AI development would not, by themselves, justify extending the provision to training models for a publicly available commercial service. The order’s reasoning could therefore be narrowed or reversed if the appellate court concludes that such an interpretation exceeds the statutory wording and requires legislative intervention.

This uncertainty makes legislative clarification prudent. Legislators should consider amending Section 52 to expressly address text and data mining, including AI training, and specify the circumstances in which these activities may proceed without individual authorisation or consent. Such an amendment should address lawful access, commercial and non-commercial uses, retention of training copies, and safeguards against reproduction of protected expression in outputs.


References:

[1] Eastern Book Company v. D.B. Modak, (2008) 1 SCC 1.

[2] Department-related Parliamentary Standing Committee on Human Resource Development, Two Hundred Twenty-seventh Report on the Copyright (Amendment) Bill, 2010 (Rajya Sabha Secretariat, 23 November 2010), para 19.2, p. 51, available at https://sansad.in/getFile/rsnew/Committee_site/Committee_File/ReportFile/16/16/227_2016_6_15.pdf?source=rajyasabha.

[3] UK Department for Science, Innovation and Technology, Department for Culture, Media and Sport & Intellectual Property Office, Copyright and Artificial Intelligence: Impact Assessment (18 March 2026), paras 106–108.

[4] UK Department for Science, Innovation and Technology, Department for Culture, Media and Sport & Intellectual Property Office, Report on Copyright and Artificial Intelligence (18 March 2026), Section G (Licensing).

[5] United States Copyright Office, Copyright and Artificial Intelligence, Part 3: Generative AI Training 66–71 (2025).

[6]  Union of India v. Hansoli Devi, (2002) 7 SCC 273.

Image Credits:

Photo by Supatman on Canva

Legislators should consider amending Section 52 to expressly address text and data mining, including AI training, and specify the circumstances in which these activities may proceed without individual authorisation or consent. Such an amendment should address lawful access, commercial and non-commercial uses, retention of training copies, and safeguards against reproduction of protected expression in outputs.

POST A COMMENT