
Ever since ChatGPT first emerged on the scene in 2022, there has been a vociferous debate about whether the indexing (or “scraping”) of public content that AI companies do when they are training a large-language model should be considered an infringement of the copyright held by publishers and/or the authors of those books, or whether it should be covered by the “fair use” exemption in US copyright law. As some of you may know, I have consistently been on the latter side of the debate — in a piece for the Columbia Journalism Review and then an edition of The Torment Nexus, I argued that the scraping or indexing of public content by LLMs should be legally no different than the indexing of books that Google did in the early 2000s as part of its Google Books project. After a court case that lasted for a number of years, judge Denny Chin ruled in 2013 that Google’s indexing of content was covered by the fair-use exemption because he believed it to be a “transformative” use, which is one of the four factors that judges have to take into account when they are making a decision. As I wrote last year:
Judges have to balance the competing elements of the “four factor” test, namely: 1) What is the purpose of the use? In other words, is it intended as parody or satire, is it for scholarly research or journalism, etc. 2) What is the nature of the original work? Is it artistic in nature? Is it fiction or nonfiction? 3) How much of the original does the infringing use involve — is it an excerpt or the entire work? and 4) What impact does the infringing use have on the market for the original? In the Google Books case, the scanning of millions of books was not done for research or journalism, in many cases the books in question were creative works of fiction, the entire book was copied, and the Authors Guild argued that it would have a negative impact on the market. One element in Google’s favour, however, was that while its indexing process made copies of the whole book, its search engine never showed users the entire thing.”
As you can see from the four factors, a fair-use decision is effectively a balancing act between different and competing interests: the interests of the author and/or publisher, in protecting and making money from their works, and the interest of the public in having “transformative” uses of art available to them. This kind of balancing is necessary because copyright itself was designed as a balancing act, between the commercial interests of creators and the public benefit of freely available artistic work — to “promote the progress of science and useful arts,” as the US Constitution describes it. Some authors and publishers (but not all) believe that copyright’s sole purpose is to enrich creators, but that’s not accurate; revenue for creators is important, but so is society’s interest in having publicly available and usable art. Judge Chin decided that the scanning of books in order to make them searchable and provide excerpts was transformative enough that it outweighed the infringement of copyright and potential market impact.
Note: This is a version of my Torment Nexus newsletter, which I send out via Ghost, the open-source publishing platform. You can see other issues and sign up here.
Continue reading “Judge says AI engines can index books but can’t pirate them”




















