Latest Posts

AI Machine Learning:  Remedies Other Than Copyright Law?

In my last post, I discussed some of the allegations that “machine learning” (ML) with the use of copyrighted works constitutes mass infringement. Citing the class action lawsuits Andersen and Tremblay, I predicted that if the courts do not find that ML unavoidably violates the reproduction right (§106(1)), copyright law may not offer much relief to the creators of the works used for AI development. As of last week, it remains to be seen whether we’ll get to that question after Judge Orrick of the Northern District of California stated that he is tentatively prepared to dismiss the suit with leave to amend the complaint. The judge did indicate that a claim of direct infringement could survive, but we’ll have to see what comes of an amended complaint.

As mentioned in the last post, if the court does not find a valid claim of copyright infringement, the other allegations will likely fail as a result. Nevertheless, though the state allegations may be moot in the class cases filed thus far, I had intended in this post to look at whether any non-copyright remedies present much hope for creators. For instance, the Andersen complaint alleges violations of statutory and common law rights of publicity and violations of statutory unfair practice prohibitions in the State of California.

Right of Publicity and Works “in the style of…”

One of the most concerning aspects of generative AI is that it allows a user to prompt the system to make a work “in the style of [named artist].” Karla Ortiz, in her testimony to the Senate Judiciary Committee on July 12, stated, “[Artist] Greg Rutkowski, had his name used as a prompt between Midjourney, Stability AI and the porn generator Unstable Diffusion, about 400,000 times as of December 2022. (And these are on the lower side of estimates).” This is a deeply personal assault on the identity, work, and potential livelihood of an artist who has spent years mastering his craft and developing that distinctive style which can now be mimicked by a computer. But as a matter of doctrine, copyright does not protect style. So, can state laws like the right of publicity (ROP) offer any relief?

Only half the states in the U.S. have statutory rights of publicity, though most states recognize a common law ROP, which may prove more expansive in a litigation. The California statute, considered one of the strongest, prohibits the use, without consent, of a person’s likeness, name, voice, or signature for commercial purposes. In particular, advertising a product, service, or viewpoint so that a representation of the individual implies that person’s endorsement, is a paradigmatic violation of ROP and may infringe the speech right of the individual. But does prompting a generative AI to create an image “in the style of [artist]” implicate the artist’s ROP? Maybe and sometimes.

Because a user can prompt the production of a visual work “in the style of Greg Rutkowski,” one obvious implication is that there will be hundreds or thousands of “Rutkowskis” in the world which the artist did not create. If any of those AI-generated works are substantially similar to an existing work he did create, then he may have claims of copyright infringement—and potentially too many to contemplate addressing. But what about the works that do not look substantially similar to anything in the artist’s portfolio but to the observer, do look like “new Rutkowskis”?

In the fine-art trade, the duty to validate provenance and authenticity (i.e., not commit forgery) should offer some protection for those artists who sell their work in galleries etc. But in the commercial market, if a new vodka brand wants images in the style of, say, Molly Crabapple for its ad campaign but doesn’t want to hire Molly for the job, they could use generative AI to make something in her style. No question this implies that artists will lose gigs, but whether ROP offers a remedy to Crabapple herself in this hypothetical is questionable.

To begin, artwork in the style of the artist is not a “likeness.” Thus, Crabapple would have to prove that observers seeing the ads would perceive the images as her work and that the images, therefore, result in an unlicensed endorsement in violation of ROP. That can be a high bar to reach, let alone repeat in what could amount to multiple complaints by just one artist—and potentially in multiple states! Although a violation of ROP can stand alone without an underlying violation of some other law, it is a fact-intensive, case-by-case consideration and, therefore, hard to imagine how it can support allegations of harm to a plaintiff class as alleged in Andersen.

Moreover, some states only recognize celebrity ROP and not average citizen ROP. So, in the case of the visual artist, what is the threshold where he or she is famous enough to be considered a celebrity? Crabapple, Rutkowski, et al. are very well known in the art world, but they’re not movie-star well known to the general public. So, what constitutes “celebrity” in this context? We don’t know. The specific problems caused by generative AI are brand new.

New Federal ROP Law?

In testimony before the SJC along with Karla Ortiz, General Counsel for Universal Music Group Jeffrey Harleston broached the subject of adopting a federal right of publicity, and the idea was at least entertained by some of the senators. A federal ROP could theoretically address some of the new harms to artists caused by generative AI. Not only would a federal statute provide a uniform, national framework, but the new law would be written with an understanding of AI and its potential harms as a foundation of legislative intent. Further, it was raised in committee that ROP should apply to everybody and not just celebrities.

The Motion Picture Association (MPA) filed comments in response to this discussion, and these were focused largely on the fact that, historically, ROP has applied to commercial/promotional uses, but not to expressive ones. The MPA is right to point to some tricky considerations that would need to shape a new federal ROP in order to strike a balance between disenfranchising creators or performers through AI-generated replicas while allowing use of the technology to create expressions that are protected by the First Amendment.

In fact, in my book, I allude to a hypothetical future biopic about Carrie Fisher that (with the family’s permission, of course) might dramatize scenes using AI replicas. Whether this use of the technology would be an engaging choice in lieu of casting an actress play young Carrie is a question of aesthetics and culture, but not a question that can or should be addressed as a matter of law. Suffice to say, the contours of a prospective new federal ROP are complex enough to be subject of future posts.

Unfair Competition

Unlike an ROP complaint, unfair competition does not stand alone as an allegation. In general, these laws bar businesses from gaining unfair advantage by engaging in some form of prohibited conduct. In the Andersen et al. class action suits, the underlying conduct is alleged to be copyright infringement, which allegedly makes the AI developers unfair competitors with the plaintiff class of artists. This state allegation would seem to have merit if the court finds the developers liable for violation of §106(1) of the Copyright Act, and unlike ROP, I can see it surviving as a complaint for a whole class—i.e., as unfair competition against all artists. That would be encouraging. But if, for instance, the court finds that not all the named plaintiffs have standing (i.e., do not have registered works in suit), it’s hard to say what this does to the unfair competition complaint as argued.

What About Trademark?

It is tempting to wonder whether certain creators can find protection in trademark law. In addition to the cost of registering and maintaining a trademark, only certain artists would be able to make effective use of this form of intellectual property. In this instance, trademark only protects use of the artist’s name in commerce, and since it is already illegal to trade in forgeries, registering one’s name as a trademark may be redundant and little protection against the use of AI to produce “in the style of” works.

Relatedly, the Federal Trade Commission (FTC) may have a role to play to protect consumers against fraud stemming from the uses of generative AI. As to the FTC stepping in, presumably they could respond to or seek prophylactically to protect consumers from forgery at scale. Not unlike the rampant proliferation of counterfeit products sold via eCommerce sites, generative AI certainly presents the opportunity for some party to start generating mass forgeries of popular artists and selling those in the millions. As such, measures to restrict the use of artists’ names in generative AI may belong in the FTC’s wheelhouse.

In both this post and the last, my intent is not to advocate on behalf of the AI developers. Far from it. Instead, I am trying to kick the tires of existing law to ask whether the law is sufficient to the task of protecting authors of creative works. Because, overall, I’m not sure it is, though it is also essential to note that every type of work has different implications (i.e., voice actors’ rights are more likely to sound in ROP than visual artists’ rights).

One thing is certain. Generative AI is not comparable to the printing press, camera, phonograph, or any more recent changes to production and distribution enabled by digital technology. Debate about AI must be sequestered from discussions about technologies of the past because few, if any, of those revolutions are instructive to the moment. There is no doubt that AI implies new regulation in medicine, finance, IT, security, and just about everywhere else it will invade; and there is no reason why Congress cannot adopt the same posture in order to protect America’s creative culture and economy.


Image source by: idaakerblom

Training AI With Protected Works:  Is Copyright Law Designed to Respond?

generative ai

Many creators feel very strongly that “training” AI models with unlicensed, copyrighted works is unjust—not least because generative AIs built on their creativities will put some creators out of business while enriching more tech moguls. It is both insult and injury to see one’s work used, without consideration, to underwrite the mechanism of one’s own obsolescence. But regardless of how we may feel about the practice of “machine learning” (ML) with unlicensed material, it remains to be seen whether and where current law provides any remedies. I’ll try to consider that topic in this post and the next post, beginning with the allegation that ML is mass copyright infringement.

Four class action lawsuits against generative AI developers have been filed thus far in the District Court for the Northern District of California, and all by the same law firm. Because all the complaints are similar, I will stick to the two that were filed first. In Andersen et al. v. Stability AI et al., a class of visual artists is suing Stability AI and Midjourney;[1] and in Tremblay et al. v. Open AI, a class of book authors suing OpenAI over the development of ChatGPT.[2] Both complaints allege direct and vicarious copyright infringement as well as unlawful removal of copyright management information (CMI). Both complaints also contain counts for violation of the derivative works right §106(2), and based on that theory, the Andersen complaint alleges unlawful making available of said derivative works in violation of 106(3), (4), & (5). The complaints also contain state law allegations, but I will discuss those in the next post.

Reproduction and the Battle of Analogies

The question of whether ML with copyrighted works constitutes an act of mass infringement will turn on the factual consideration as to whether any copying occurs in violation of the reproduction right (§106(1)). In Andersen and Tremblay, there is considerable focus on the potential of a generative AI to output an infringing work based on its training corpus. For instance, if the work of Karla Ortiz (one of the named plaintiffs in Andersen) is part of the ingested materials, then the assumption is that the AI model has the potential to produce a copy of an existing Ortiz work or a work that is substantially similar to an Ortiz work.

The reproduction inquiry may be different for each model and each type of work used for input. In Andersen, the complaint states, “Because a trained diffusion model can produce a copy of any of its Training Images—which could number in the billions—the diffusion model can be considered an alternative way of storing a copy of those images.” By contrast, the Tremblay complaint alleges that copying occurs, but it does not specifically describe how the ChatGPT training process entails reproduction. “During training, the large language model copies each piece of text in the training dataset and extracts expressive information from it,” the complaint states.

If the AI system produces any copies of any of its training materials, this is evidence that the system violates the reproduction right. Prompt the generator to make an image of Dr. Strange, and if Dr. Strange comes out, then nobody can doubt that Dr. Strange is a latent copy in the system and that this potential to copy is sufficient evidence of infringement at the input stage. Alternatively, if the system can only produce work “in the style of” Karla Ortiz, this raises different issues (and very serious concerns) but may not be considered sufficient evidence of “reproduction” in the input process. But the courts need not look at outputs, or even potential outputs, to find violation of the reproduction right.

It has been held (specifically in the 9th Circuit)[3] that even storing a copy in random access memory (RAM) is sufficient to find a violation of the reproduction right. The AI developers will seek to prove that their systems do not copy the works ingested in any sense, or that if they do, they copy only non-protected (i.e., factual) elements of the works. Using anthropomorphic words like observe, learn, study, etc. to describe ML, the argument from the developers will be that these models are designed to obtain information about the works but not copy the works anywhere in the system. Input an illustration, for example, and what the system allegedly stores are millions of data points about line weights, composition, colors, shading, etc. Then, combined with billions of other data points from billions of other works, the model generates probability algorithms which are then used to produce new visual works when users prompt the system with instructions.

AI developers like to compare “training” their models to the learning a human artist does when she experiences or studies works other than her own. In addition to being a reductive and dehumanizing analogy for the ways in which artists teach themselves a craft, this line of reasoning may be seen by the courts as smoke and mirrors. The factual question is whether the system retains a copy long enough to be perceived by the machine, which has been held to be violative of §106(1). Long-term storage of a copy is not required, and my understanding is that making a “more than fleeting” copy is unavoidable in any computer system—i.e., that there is no such thing as ingestion without reproduction.

Proving reproduction will be the whole ballgame insofar as litigation can address whether feeding a corpus of protected works is a violation of law. We shall see what the courts make of the facts presented, but without finding reproduction, the other copyright complaints likely fall. For instance, removal of CMI is not a stand-alone violation. Section 1202 of the DMCA states that removal is a violation if the party doing the removing knows or has reasonable grounds to know “that it will induce, enable, facilitate, or conceal an infringement of any right under this title.” Therefore, there must be a colorable claim of infringement for the CMI allegation to survive.

Derivative Works Allegations

Both the Andersen and Tremblay complaints allege that the AIs produce unlicensed derivative works in violation of §106(2), though the arguments are different in each case. In Andersen, the allegation arises from the premise that the system cannot produce anything outside the limitations of its data set composed of protected works. “The resulting image [output] is necessarily a derivative work, because it is generated exclusively from a combination of the conditioning data and the latent images, all of which are copies of copyrighted images.…a latent diffusion system…can never exceed the limitations of its Training Images.”

It’s an interesting theory, but I’m not sure anything in copyright law can support the argument that all potential outputs of the generative AI are unauthorized derivatives of the total corpus of works in the training set. To find an infringing derivative of a visual work (typically one image) requires a substantial similarity inquiry comparing a specific original with the follow-on work to determine what has been copied and whether that copying renders the second work a derivative of the first. This is difficult enough in the world of humans intentionally using a single visual work to produce a different visual work (see Goldsmith v. Warhol!!). So, it seems highly speculative to ask a court to find generally that billions of images output are, as a matter of law, derivatives of the billions of images input. I’m not certain the court has anywhere to look for guidance to consider this reading of the derivative works right.

If this derivative works theory is tough with images, it would be even harder with text—i.e., to allege that the textual outputs are derivatives of all the textual inputs is akin to saying that every book written is a derivative of every book read. This echoes a popular sentiment among the anti-copyright crowd that no work is “original,” a premise that should not be given any legal weight, even in the service of trying to protect creators from AI developers. 

In Tremblay, the allegation is not that the individual outputs of ChatGPT are derivatives of the corpus of books used in training, but that the entire model is a single derivative work of its corpus. “Because the OpenAI Language Models cannot function without the expressive information extracted from Plaintiffs’ works (and others) and retained inside them, the OpenAI Language Models are themselves infringing derivative works, made without Plaintiffs’ permission and in violation of their exclusive rights under the Copyright Act,” the complaint states. [Emphasis added]

Again, claiming that the entire LLM is a single derivative work of the millions of literary works fed into the system would seem to strain the derivative works right beyond the limit where any court can venture. In fact, this allegation could potentially bolster the inevitable fair use defense the AI developers will be arguing—namely that the finding of “transformative use” in Google Books favors fair use of the corpus of work used in ML.  

Fair Use & Google Books

Notably, these cases are brought in California, controlled by the Ninth Circuit and, therefore, not bound by the Second Circuit decision in Google Books, which many believe to be the strongest precedent favoring fair use for the AI developers. The comparison is a natural one. Google scanned whole books into a system to create a unique tool for searching the contents of books without providing any whole-copy substitutes for legally obtained copies. The court, noting that its decision “pushed the boundaries of fair use,” found under factor one that Google Books is “transformative” for its utility and found under factor four that it did not pose a threat to the market for the books used.

What the AI developers will try to argue under Google Books is that 1) their systems are highly “transformative” because they use protected works to create novel (even revolutionary) applications; and 2) their systems are designed to avoid outputting any copies that would serve as substitutes for the works in the data set. It is conceivable that courts or juries would find the comparison compelling, though the aforementioned capacity of a given AI to output Dr. Strange means that, unlike Google Books, the visual AI system at issue does make substitutes available and, therefore, the precedent is inapt.

By contrast, ChatGPT or other text-based application could have a stronger defense under Google Books if it is not possible, for instance, to have the system output an entire in-copyright literary work. The Tremblay complaint refers to the output of summaries, which is evidence that a whole book was ingested, but a summary is not generally an infringement and is certainly not a substitutional copy.

Meanwhile, other considerations should perhaps militate against finding fair use for generative AI model training. For instance, Google Books is a research tool for humans to learn about books written by other humans, including humans who write more books. Generative AIs are not necessarily comparable. For instance, Stable Diffusion does not provide a user with any information about an ingested work, and it poses an unprecedented threat to professional visual artists unlike any technology that has come before. Thus, the courts should consider the sui generis purpose of the generative AI at issue when citing Google Books or any other precedent to consider fair use.

In a May post, I proposed that unless the generative AI at issue can show that it promotes authorship, the court should decline to consider a fair use defense. To clarify, in Campbell, the Supreme Court states, “The fair use doctrine thus ‘permits [and requires] courts to avoid rigid application of the copyright statute when, on occasion, it would stifle the very creativity which that law is designed to foster.”[4] Until generative AI changed the landscape, there was no need to affirm that “the very creativity” fostered by copyright means “human creativity.” But today, that distinction is necessary. Although generative AI can produce volumes of “creative” material, only those works which can be protected by copyright are works of authorship. And just like it is indecent to exploit an artist’s work to build a machine that might end her career, it would be absurd to allow fair use (a component of copyright law) to defend a technology that would potentially annihilate copyright’s purpose.

Of course, that’s one man’s opinion, and one that would apply to some, but not all, works derived by generative AI. As these tools develop, and their uses are explored by various types of creators, there are examples, both in practice and in theory, where we can find that generative AI does foster new authorship. This gets into the complicated question of copyrightability of works that humans create with some AI used in the process, and because this is itself a new discussion, it is difficult to say which generative AIs, if any, can be said to “promote the progress” of authorship as a matter of law.

Legal experts, both pro and anti-copyright, will comment upon the strengths and weaknesses of Andersen, Tremblay et al. represented by the one firm that has taken the lead on these lawsuits. But even where these cases may be flawed, they can provide some insight into the question posed by this essay:  is copyright law an answer to the potential hazards of generative AI? I suspect that a fundamental difficulty arises because generative AI poses an existential threat to the future of authors, and some of the injustices and cultural calamities inherent to that threat may not be remedied (or entirely remedied) by the principles of copyright. Remedies sounding in other areas of law could loom larger, especially for certain types of creators, and that will be the subject of the next post.


[1] Deviant Art is also a named defendant being sued for breach of contract for providing works to Stability for ingestion.

[2] The same firm is now representing Sarah Silverman and another class of book authors, though the complaint is essentially the same as Tremblay.

[3] MAI Systems Corp. v. Peak Computer, Inc., 991 F.2d 511 (9th Cir. 1993).

[4] Citing Stewart v. Abend (1990).

Image by: idaakerblom

Get AI Wrong and There Will Be Nothing to Forgive

We all know the mantra that says it’s better to ask forgiveness than permission. According to Quote Investigator, the earliest published version of this sentiment appeared in 1846, but QI’s editors believe the notion is older than that and cannot be attributed to any one source. Whatever its derivation or contexts in which it has been used over many decades, the phrase is presently associated with Silicon Valley and the heedless “move fast and break things” approach to technological development.

I was hardly alone in noticing that Ocean Gate CEO Stockton Rush tech-broed the design of his Titan submersible, dismissing warnings and safety regulations as barriers to innovation (one of Silicon Valley’s favorite refrains about pesky rules). Moreover, because the vessel imploded and the passengers were apparently killed before they knew what happened, Titan’s fate seems an apt harbinger of the technological singularity—its analogy to crossing the event horizon of a black hole conjuring an uncomfortable squeezing parallel to death by implosion.

For anyone unfamiliar with the term technological singularity, it is often described as a threshold in AI development when computers “wake up” and their intelligence surpasses human intelligence. The event horizon analog, credited to sci-fi author Vernor Vinge, describes two principles: 1) that we have no way to predict what happens beyond the capacity of human intelligence; and 2) that we won’t know when we’ve crossed the horizon.

Of course, we need not anthropomorphize computers or manifest the many fictions about sentient machines to approach the horizon, and some experts believe we are already inside the gravitational pull of singularity. For instance, in a May editorial for The Hill, McGill University scholar J. Mauricio Gaona, asserting that singularity is “already underway,” states …

The possibility of soon reaching a point of singularity is often downplayed by those who benefit most from its development, arguing that AI has been designed solely to serve humanity and make humans more productive.

Such a proposition, however, has two structural flaws. First, singularity should not be viewed as a specific moment in time but as a process that, in many areas, has already started. Second, developing gradual independence of machines while fostering human dependence through their daily use will, in fact, produce the opposite result: more intelligent machines and less intelligent humans.  

Gaona notes that the commercial potential of AI in medicine, finance, transportation et al. will require unsupervised learning algorithms (i.e., machines that effectively “train” themselves) and that granting even limited autonomy to these systems means we have already stepped over the threshold toward singularity. Further, he argues, once AI meets quantum computing, then “Crossing the line between basic optimization and exponential optimization of unsupervised learning algorithms is a point of no return that will inexorably lead to AI singularity.” Not to worry, though, the U.S. Congress is on the job.

On June 21, Senator Schumer, speaking at the Center for Strategic and International Studies (CSIS), discussed the SAFE Innovation Framework for Artificial Intelligence. “Change at such blistering speed may seem frightening to some—but if applied correctly, AI promises to transform life on Earth for the better. It will reshape how we fight disease, tackle hunger, manage our lives, enrich our minds, and ensure peace. But there are real dangers too: job displacement, misinformation, a new age of weaponry, and the risk of being unable to manage this technology altogether,” Sen. Schumer stated. The SAFE framework is outlined as follows:

  • Security. Necessary to protect national security for the U.S. and economic security for residents whose jobs may be displaced by automation.
  • Accountability. The providers of AI systems must deploy these systems in a transparent and responsible way. They must remain responsible for violations of the protections ultimately put in place by promoting misinformation, violating intellectual property rights, or when the AI is biased.
  • Foundations. AI algorithms and products must be developed in a way that promotes America’s foundations such as justice, freedom, and civil rights.
  • Explainability. The providers of AI systems must provide appropriate disclosures that inform the public about the system, the data it uses, and its contents.
  • Innovation. The overall guiding principle for any regulations or policy regarding AI should be to encourage, not quash, innovation so that the U.S. becomes and remains the global leader in this technology.

Is that all? Having worked for just over a decade on the edges of policymaking, I find it hard to believe that Congress can be nimble enough to address all those bullet points while keeping up with AI development itself. And that’s if Members agree about the framework’s principles. “Promotes … justice, freedom, and civil rights.”? Near as I can tell, there is not much consensus on the meaning of those words these days. Or what about “misinformation”? How many of Schumer’s colleagues on the right can plausibly subscribe to a common definition of “misinformation” while they carry Trump’s luggage through the gauntlet of his well-earned indictments? With millions of American voters willfully blinding themselves to old-school evidence of criminal conduct, are we anywhere near capable of addressing the unprecedented realism of AI-generated chicanery?

It is certainly conceivable that with the right controls in place, AI can be harnessed to make life better for humans, and, indeed, if that is not the goal, then why continue to build it? Unfortunately, the answer from many of those doing the building is “because we can.” And, thus, we are locked into taking this roller-coaster ride whether we want to or not. At least if we do cross the threshold toward singularity, the tech-bros won’t have to ask humanity for forgiveness, though they may have to ask their machines for mercy.


Image sources by: vchalupAgor2012