Record Labels File Suit Against Internet Archive for Copyright Infringement

Pride goeth before destruction, and an haughty spirit before a fall. (Proverbs 16:18 KJV)

Citing 2,749 works in suit, six of the major music labels (UMG, et al.) have filed a multi-count complaint against Internet Archive (IA), Brewster Kahle personally, Kahle’s foundation, and an audio digitizing service operated by an individual named George Blood. Total potential damage award with legal fees:  around a half-billion dollars. Likelihood of defendants’ success, assuming all factual allegations are well-founded:  less than zero. So, while I have no idea how much cash on hand Kahle has to burn, this suit highlights a question I have often asked myself—namely how eager is he to put his money where his anti-copyright mouth is?

As the outcome in the book publishers’ lawsuit, Hachette et al., makes clear, cockamamie theories about how the law works may find an audience in the blogosphere, but they make poor arguments in court. In that case, Internet Archive relied on a cockamamie theory called Controlled Digital Lending (CDL) and hitched that wagon to a belief that the practice was shielded by the doctrine of fair use. The defendant lost on every point, and a negotiated judgment is already filed with the court notwithstanding IA’s right to appeal.

In this new case with the record labels, IA does not even have the gossamer of an unfounded theory to weave into its response. Instead, the initiative IA calls “The Great 78 Project” is alleged to entail knowing evasion of compliance with clearly defined statute. Without going into each of the counts against each of the defendants, the crux of the matter is that Great 78 makes digital copies of sound recordings from 78RPM vinyl records and hosts those files for unlimited streaming or downloading. So, if any of those sound recordings are still under copyright, this implicates violation of three of the exclusive rights under Section 106 of the Copyright Act—reproduction, distribution, and public performance by digital audio transmission (§106(1), (3), & (6) respectively).

The Great 78 Project purports to make available rare and difficult-to-find sound recordings, and presumably, some portion of the collection comprises works in the public domain (PD) and/or truly rare works that are not commercially available. But headlining the more than 2,000 works in suit, the complaint cites recordings that are neither in the PD nor rare by any means. Popular recordings by Elvis Presley, Duke Ellington, Billie Holiday, Ray Charles, Chuck Berry, Frank Sinatra, Ella Fitzgerald, Louis Armstrong, and Hank Williams are named as prime examples that can be accessed by legal, commercial means, including major streaming services.

The reason commercial availability of the sound recordings is relevant in this case is that under the provisions of the Music Modernization Act (MMA) of 2018, a library/archive is permitted to make pre-1972 sound recordings available if, among other conditions, it makes a good-faith effort to determine that the recordings are not commercially available. That is a pared down description, but it’s the basic principle, which IA apparently chose to ignore. According to the complaint, IA made no effort to fulfill its obligation to comply with any of the following Copyright Office guidelines:

…a reasonable search for purposes of 17 U.S.C § 1401)(c) must include, among other things: (i) searching the Copyright Office’s database of indexed schedules listing right owners’ pre-1972 sound recordings; (ii) searching Google, Yahoo!, or Bing; (iii) searching at least one of the following streaming services: Amazon Music Unlimited, Apple Music, Spotify, or TIDAL; (iv) searching YouTube; (v) searching SoundExchange’s repertoire database; (vi) searching at least one major seller of physical product, namely Amazon.com.

Moreover, the complaint cites compelling evidence that the defendants understood their obligations under §1401 and that failure to comply would constitute copyright infringement of works like the sound recordings in suit. For instance, IA stated in a blog post about the MMA shortly after it was signed into law, “But, as we understand it, the MMA means that libraries can make some of these older recordings freely available to the public as long as we do a reasonable search to determine that they are not commercially available.” The only logical conclusion, therefore, is that defendants ignored the “reasonable search” guidelines because it is obvious that many of the sound recordings at issue can be found commercially available by a young child using Google.

Will Kahle’s Copyright Hubris Kill His Archive?

A significant distinction between this suit and the Hachette case is that Brewster Kahle and the Kahle/Austin Foundation are named defendants. The complaint alleges that Kahle is directly involved in IA policy, activities, and promotion, including the Great 78 Project, that he funds the foundation through his trust, and the foundation, in turn, funds the project. “At Kahle’s direction, the Foundation used the funds Kahle had contributed to sponsor the Internet Archive’s massive and growing infringement. The Foundation donated money to Internet Archive that Internet Archive used to pay costs in furtherance of its infringement…” the complaint states.

The Internet Archive and its friends will, no doubt, repeat populist claims that they are serving the public, behaving as a library should, and that they are being targeted by a greedy industry. But the conduct alleged in the UMG complaint reveals an even more brazen decision to circumvent copyright law than the CDL scheme underlying the Hachette suit. In the book publishers’ case, IA advanced a theory (albeit a poor one) that it was acting within the confines of the law, but here, it simply elected to evade clearly articulated statutory confines and take its chances. And this time, the cost could indeed be the whole operation.

I get why many people want to support IA, not the least being that a large part of the organization is both legal and highly useful. Among those who simply agree with Kahle et al. that copyright should not exist, that pre-1972 sound recordings should not be protected, etc., fine. That’s an opinion to which people are entitled. But beyond that general view, IA supporters should not be confused into thinking this case is about big bad industry beating up on a library doing library-like things. Assuming the factual allegations are correct, there is barely a distinction between the alleged infringing conduct in the Great 78 Project and The Pirate Bay. And to the extent the legitimate archive has been treated like a front for mass infringement projects, the blame for that decision rests with Kahle and his colleagues, not with the music or book publishers.

Since the first post I wrote about Internet Archive, I have acknowledged that the repository of public domain and truly rare material is an invaluable research tool. In fact, that post in October of 2017 asked directly whether the anti-copyright rhetoric was necessary to the organization, but since then, it has become clear that Kahle has used IA’s operation and reputation to engage in much more than rhetoric. And in this potentially costly litigation with the record labels, it is conceivable that this hubristic crusade against copyright law could, as the proverb says, lead to the collapse of an otherwise good enterprise.


Photo by: panoramaimages

Before Generative AI, Big Tech Taught Artists to Abdicate Copyright Rights

One of the more challenging aspects of copyright advocacy is the fact that many artists and creators are conflicted about enforcing their own rights, and from observation, the disconnect is ideological. For the last 30 years, copyright skepticism has been woven into political narratives rooted in criticism of corporations and the excesses of capitalism—popular themes among the political left, which encompasses most artists. Now that generative AI developers are turning creative works into “pink slime,” and artists are suddenly more interested in their rights, it might help to recognize that the industry deploying AI is the same one that taught creators to advocate against copyright in the first place.

The year 2011 was an extraordinary time to jump into the fray. It was immediately apparent that allegations of “copyright maximalism” were deeply intertwined with a sincere and animated belief that the internet would foster a new and potent form of direct democracy to confront a litany of injustices. Copyright enforcement was characterized as a barrier to that promise, and so, the Stop SOPA campaign (to kill anti-piracy legislation) became part of a larger, frenetic collage that included OWS protests, European pirate parties, Anonymous, Wikileaks, etc., all feeding an atmosphere of revolution that corresponded with headlines and memes claiming that “Hollywood” wanted to use copyright to break the internet and stifle speech.

But Big Tech’s promise to democratize everything was a Trojan Horse from which the AI bots have now emerged to ransack the village. Not only did promoters of the “free flow of information” elide the fact that their platforms were as likely to produce the January 6th insurrection as the “Pussy Hat” March, but the allegation that copyright was a barrier to information flow had nothing to do with liberating our speech and everything to do with limiting their liability.

Every time members of the creative community echoed the anti-copyright messages pumped out by Fight for the Future, the EFF, Public Knowledge, or the platforms themselves, what was really being advocated was a lack of accountability for online service providers. I never fully understood how one of the most exploitative industries in history managed to turn anti-corporatist sentiment to its advantage, but I assumed it was the gestalt of the internet. The illusion that social platforms belong to the people was a charade that enabled Google, Facebook, et al. to camouflage their interests as our rights.

That theme has aged about as well as the tobacco industry’s efforts to sell freedom to get smokers to ignore cancer, but it’s been almost two years since Big Tech’s “Big Tobacco moment,” and little has changed. Neither in Congress nor the courts have online service providers been held accountable for much of anything—and that’s with laws on the books. When we consider that, for almost three decades, the major platforms have acted in bad faith with their end of the DMCA bargain, and the courts have interpreted Section 230 as an unlimited liability shield, it is hard to feel hopeful about a legal framework for accountability for harms resulting from AI.

In fact, certain AI tools (e.g., LLMs) may imply a wider “neutral” buffer between potentially harmed parties and potentially liable parties. “Knowledge” and “intent” are key factors in establishing liability, and we have watched Big Tech play shell game with the concept of what they can “know” or “intentionally” control about activity on their platforms. AI tools could take these shenanigans to the next level, enabling new forms of harm with an even weaker nexus linking the machines to the people who design and operate them.

In the copyright world, platform operators have consistently circumvented their obligations under the DMCA with shrugging statements like We can’t police the internet, alluding to staggering volume while conjuring an association with authoritarianism. Now, the circumstances are different. It is a near certainty that every creative work made has been, or will be, ingested into one or more AI training models, and unless the courts find this to be an act of mass piracy and order disgorgement of the datasets, creators may have to accept that their work is being turned into pink slime.

While it is encouraging to see artists take a more active interest in copyright rights as a response to AI, it is also a bittersweet transition in light of all that has happened so far. Whatever comes next, I hope the creative community will recognize that copyright rights are the closest thing to labor rights the independent artist has. And these rights should not be weakened or abandoned for the sake of more billionaires making false promises about democracy and free speech.

Training AI With Protected Works:  Is Copyright Law Designed to Respond?

generative ai

Many creators feel very strongly that “training” AI models with unlicensed, copyrighted works is unjust—not least because generative AIs built on their creativities will put some creators out of business while enriching more tech moguls. It is both insult and injury to see one’s work used, without consideration, to underwrite the mechanism of one’s own obsolescence. But regardless of how we may feel about the practice of “machine learning” (ML) with unlicensed material, it remains to be seen whether and where current law provides any remedies. I’ll try to consider that topic in this post and the next post, beginning with the allegation that ML is mass copyright infringement.

Four class action lawsuits against generative AI developers have been filed thus far in the District Court for the Northern District of California, and all by the same law firm. Because all the complaints are similar, I will stick to the two that were filed first. In Andersen et al. v. Stability AI et al., a class of visual artists is suing Stability AI and Midjourney;[1] and in Tremblay et al. v. Open AI, a class of book authors suing OpenAI over the development of ChatGPT.[2] Both complaints allege direct and vicarious copyright infringement as well as unlawful removal of copyright management information (CMI). Both complaints also contain counts for violation of the derivative works right §106(2), and based on that theory, the Andersen complaint alleges unlawful making available of said derivative works in violation of 106(3), (4), & (5). The complaints also contain state law allegations, but I will discuss those in the next post.

Reproduction and the Battle of Analogies

The question of whether ML with copyrighted works constitutes an act of mass infringement will turn on the factual consideration as to whether any copying occurs in violation of the reproduction right (§106(1)). In Andersen and Tremblay, there is considerable focus on the potential of a generative AI to output an infringing work based on its training corpus. For instance, if the work of Karla Ortiz (one of the named plaintiffs in Andersen) is part of the ingested materials, then the assumption is that the AI model has the potential to produce a copy of an existing Ortiz work or a work that is substantially similar to an Ortiz work.

The reproduction inquiry may be different for each model and each type of work used for input. In Andersen, the complaint states, “Because a trained diffusion model can produce a copy of any of its Training Images—which could number in the billions—the diffusion model can be considered an alternative way of storing a copy of those images.” By contrast, the Tremblay complaint alleges that copying occurs, but it does not specifically describe how the ChatGPT training process entails reproduction. “During training, the large language model copies each piece of text in the training dataset and extracts expressive information from it,” the complaint states.

If the AI system produces any copies of any of its training materials, this is evidence that the system violates the reproduction right. Prompt the generator to make an image of Dr. Strange, and if Dr. Strange comes out, then nobody can doubt that Dr. Strange is a latent copy in the system and that this potential to copy is sufficient evidence of infringement at the input stage. Alternatively, if the system can only produce work “in the style of” Karla Ortiz, this raises different issues (and very serious concerns) but may not be considered sufficient evidence of “reproduction” in the input process. But the courts need not look at outputs, or even potential outputs, to find violation of the reproduction right.

It has been held (specifically in the 9th Circuit)[3] that even storing a copy in random access memory (RAM) is sufficient to find a violation of the reproduction right. The AI developers will seek to prove that their systems do not copy the works ingested in any sense, or that if they do, they copy only non-protected (i.e., factual) elements of the works. Using anthropomorphic words like observe, learn, study, etc. to describe ML, the argument from the developers will be that these models are designed to obtain information about the works but not copy the works anywhere in the system. Input an illustration, for example, and what the system allegedly stores are millions of data points about line weights, composition, colors, shading, etc. Then, combined with billions of other data points from billions of other works, the model generates probability algorithms which are then used to produce new visual works when users prompt the system with instructions.

AI developers like to compare “training” their models to the learning a human artist does when she experiences or studies works other than her own. In addition to being a reductive and dehumanizing analogy for the ways in which artists teach themselves a craft, this line of reasoning may be seen by the courts as smoke and mirrors. The factual question is whether the system retains a copy long enough to be perceived by the machine, which has been held to be violative of §106(1). Long-term storage of a copy is not required, and my understanding is that making a “more than fleeting” copy is unavoidable in any computer system—i.e., that there is no such thing as ingestion without reproduction.

Proving reproduction will be the whole ballgame insofar as litigation can address whether feeding a corpus of protected works is a violation of law. We shall see what the courts make of the facts presented, but without finding reproduction, the other copyright complaints likely fall. For instance, removal of CMI is not a stand-alone violation. Section 1202 of the DMCA states that removal is a violation if the party doing the removing knows or has reasonable grounds to know “that it will induce, enable, facilitate, or conceal an infringement of any right under this title.” Therefore, there must be a colorable claim of infringement for the CMI allegation to survive.

Derivative Works Allegations

Both the Andersen and Tremblay complaints allege that the AIs produce unlicensed derivative works in violation of §106(2), though the arguments are different in each case. In Andersen, the allegation arises from the premise that the system cannot produce anything outside the limitations of its data set composed of protected works. “The resulting image [output] is necessarily a derivative work, because it is generated exclusively from a combination of the conditioning data and the latent images, all of which are copies of copyrighted images.…a latent diffusion system…can never exceed the limitations of its Training Images.”

It’s an interesting theory, but I’m not sure anything in copyright law can support the argument that all potential outputs of the generative AI are unauthorized derivatives of the total corpus of works in the training set. To find an infringing derivative of a visual work (typically one image) requires a substantial similarity inquiry comparing a specific original with the follow-on work to determine what has been copied and whether that copying renders the second work a derivative of the first. This is difficult enough in the world of humans intentionally using a single visual work to produce a different visual work (see Goldsmith v. Warhol!!). So, it seems highly speculative to ask a court to find generally that billions of images output are, as a matter of law, derivatives of the billions of images input. I’m not certain the court has anywhere to look for guidance to consider this reading of the derivative works right.

If this derivative works theory is tough with images, it would be even harder with text—i.e., to allege that the textual outputs are derivatives of all the textual inputs is akin to saying that every book written is a derivative of every book read. This echoes a popular sentiment among the anti-copyright crowd that no work is “original,” a premise that should not be given any legal weight, even in the service of trying to protect creators from AI developers. 

In Tremblay, the allegation is not that the individual outputs of ChatGPT are derivatives of the corpus of books used in training, but that the entire model is a single derivative work of its corpus. “Because the OpenAI Language Models cannot function without the expressive information extracted from Plaintiffs’ works (and others) and retained inside them, the OpenAI Language Models are themselves infringing derivative works, made without Plaintiffs’ permission and in violation of their exclusive rights under the Copyright Act,” the complaint states. [Emphasis added]

Again, claiming that the entire LLM is a single derivative work of the millions of literary works fed into the system would seem to strain the derivative works right beyond the limit where any court can venture. In fact, this allegation could potentially bolster the inevitable fair use defense the AI developers will be arguing—namely that the finding of “transformative use” in Google Books favors fair use of the corpus of work used in ML.  

Fair Use & Google Books

Notably, these cases are brought in California, controlled by the Ninth Circuit and, therefore, not bound by the Second Circuit decision in Google Books, which many believe to be the strongest precedent favoring fair use for the AI developers. The comparison is a natural one. Google scanned whole books into a system to create a unique tool for searching the contents of books without providing any whole-copy substitutes for legally obtained copies. The court, noting that its decision “pushed the boundaries of fair use,” found under factor one that Google Books is “transformative” for its utility and found under factor four that it did not pose a threat to the market for the books used.

What the AI developers will try to argue under Google Books is that 1) their systems are highly “transformative” because they use protected works to create novel (even revolutionary) applications; and 2) their systems are designed to avoid outputting any copies that would serve as substitutes for the works in the data set. It is conceivable that courts or juries would find the comparison compelling, though the aforementioned capacity of a given AI to output Dr. Strange means that, unlike Google Books, the visual AI system at issue does make substitutes available and, therefore, the precedent is inapt.

By contrast, ChatGPT or other text-based application could have a stronger defense under Google Books if it is not possible, for instance, to have the system output an entire in-copyright literary work. The Tremblay complaint refers to the output of summaries, which is evidence that a whole book was ingested, but a summary is not generally an infringement and is certainly not a substitutional copy.

Meanwhile, other considerations should perhaps militate against finding fair use for generative AI model training. For instance, Google Books is a research tool for humans to learn about books written by other humans, including humans who write more books. Generative AIs are not necessarily comparable. For instance, Stable Diffusion does not provide a user with any information about an ingested work, and it poses an unprecedented threat to professional visual artists unlike any technology that has come before. Thus, the courts should consider the sui generis purpose of the generative AI at issue when citing Google Books or any other precedent to consider fair use.

In a May post, I proposed that unless the generative AI at issue can show that it promotes authorship, the court should decline to consider a fair use defense. To clarify, in Campbell, the Supreme Court states, “The fair use doctrine thus ‘permits [and requires] courts to avoid rigid application of the copyright statute when, on occasion, it would stifle the very creativity which that law is designed to foster.”[4] Until generative AI changed the landscape, there was no need to affirm that “the very creativity” fostered by copyright means “human creativity.” But today, that distinction is necessary. Although generative AI can produce volumes of “creative” material, only those works which can be protected by copyright are works of authorship. And just like it is indecent to exploit an artist’s work to build a machine that might end her career, it would be absurd to allow fair use (a component of copyright law) to defend a technology that would potentially annihilate copyright’s purpose.

Of course, that’s one man’s opinion, and one that would apply to some, but not all, works derived by generative AI. As these tools develop, and their uses are explored by various types of creators, there are examples, both in practice and in theory, where we can find that generative AI does foster new authorship. This gets into the complicated question of copyrightability of works that humans create with some AI used in the process, and because this is itself a new discussion, it is difficult to say which generative AIs, if any, can be said to “promote the progress” of authorship as a matter of law.

Legal experts, both pro and anti-copyright, will comment upon the strengths and weaknesses of Andersen, Tremblay et al. represented by the one firm that has taken the lead on these lawsuits. But even where these cases may be flawed, they can provide some insight into the question posed by this essay:  is copyright law an answer to the potential hazards of generative AI? I suspect that a fundamental difficulty arises because generative AI poses an existential threat to the future of authors, and some of the injustices and cultural calamities inherent to that threat may not be remedied (or entirely remedied) by the principles of copyright. Remedies sounding in other areas of law could loom larger, especially for certain types of creators, and that will be the subject of the next post.


[1] Deviant Art is also a named defendant being sued for breach of contract for providing works to Stability for ingestion.

[2] The same firm is now representing Sarah Silverman and another class of book authors, though the complaint is essentially the same as Tremblay.

[3] MAI Systems Corp. v. Peak Computer, Inc., 991 F.2d 511 (9th Cir. 1993).

[4] Citing Stewart v. Abend (1990).

Image by: idaakerblom