In Opposition to Copyright Protection for AI Works

This is a response to “Paradise Rejected: A Conversation about AI and Authorship with Dr. Ryan Abbott” hosted by Professor Sandra Aistars at the Center for the Protection of Intellectual Property (CPIP) at George Mason University School of Law. It was first published on the CPIP blog in conjunction with Professor Aistars’s post. 


On February 14, the U.S. Copyright Office confirmed its rejection of an application for a claim of copyright in a 2D artwork called “A Recent Entrance to Paradise.” The image, created by an AI designed by Dr. Stephen Thaler, was rejected by the Office on the longstanding doctrine which holds that in order for copyright to attach, a work must be the product of human authorship. Among the examples cited in the Copyright Office Compendium as ineligible for copyright protection is “a piece of driftwood shaped by the ocean,” a potentially instructive analog as the debate about copyright and AI gets louder in the near future.

What follows assumes that we are talking about autonomous AI machines producing creative works that no human envisions at the start of the process, other than perhaps the medium. So, the human programmers might know they are building a machine to produce music or visual works, but they do not engage in co-authorship with the AI to produce the expressive elements of the works themselves. Code and data go in, and something unpredictable comes out, much like nature forming the aesthetic piece of driftwood.

As a cultural question, I have argued many times that AI art is a contradiction in terms—not because an AI cannot produce something humans might enjoy, but because the purpose of art, at least in the human experience so far, would be obliterated in a world of machine-made works. It seems that what the AI would produce would be literally and metaphorically bloodless, and after some initial astonishment with the engineering, we may quickly become uninterested in most AI works that attempt to produce more than purely decorative accidents.

In this regard, I would argue that the question presented is not addressed by the “creative destruction” principle, which demands that we not stand in the way of machines doing things better than humans. “Better” is a meaningful concept if the job is microsurgery but meaningless in the creation or appreciation of art. Regardless, the copyrightability question does not need to delve too deeply into the nature or purpose of art because the human element in copyright is not just a paragraph about registration in the USCO Compendium but, in fact, runs throughout application of the law.

Doctrinal Oppositions to Copyright in AI Works

In the United States and elsewhere, copyright attaches automatically to the “mental conception” of a work the moment the conception is fixed in a tangible medium such that it can be perceived by an observer. So, even at this fundamental stage, separate from the Copyright Office approving an application, the AI is ineligible because it does not engage in “mental conception” by any reasonable definition of that term. We do not protect works made by animals, who possess consciousness that far exceeds anything that can be said to exist in the most sophisticated AI. (And if an AI attains true consciousness, we humans may have nothing to say about laws and policies on the other side of that event horizon.)

Next, the primary reason to register a claim of copyright with the USCO is to provide the author with the opportunity, if necessary, to file a claim of infringement in federal court. But to establish a basis for copying, a plaintiff must prove that the alleged infringer had access to the original work and that the secondary work is substantially or strikingly similar to the work allegedly copied. The inverse ratio rule applied by the courts holds that the more that access can be proven, the less similarity weighs in the consideration and vice-versa. But in all claims of copying, independent creation (i.e., the principle that two authors might independently create nearly identical works) nullifies any complaint. These are considerations not just about two works, but about human conduct.

If AIs do not interact with the world, listen to music, read books, etc. in the sense that humans do these things, then, presumably, all AI works are works of independent creation. If multiple AIs are fed the same corpus of works (whether in or out of copyright works) for the purpose of machine learning, and any two AIs produce two works that are substantially, or even strikingly, similar to one another, the assumption should still be independent creation. Not just independent, but literally mindless, unless again, the copyright question must first be answered by establishing AI consciousness.

In principle, AI Bob is not inspired by, or even aware of, the work of AI Betty. So, if AI Bob produces a work strikingly similar to a work made by AI Betty, any court would have to toss out BettyBot v. BobBot on a finding of independent creation. Alternatively, do we want human juries considering facts presented by human attorneys describing the alleged conduct of two machines?

If, on the other hand, an AI produces a work too similar to one of the in-copyright works fed into its database, this begs the question as to whether the AI designer has simply failed to achieve anything more than an elaborate Xerox machine. And hypothetical facts notwithstanding, it seems that there is little need to ask new copyright questions in such a circumstance.

The factual copying complication raises two issues. One is that if there cannot be a basis for litigation between two AI creators, then there is perhaps little or no reason to register the works with the Copyright Office. But more profoundly, in a world of mixed human and AI works, we could create a bizarre imbalance whereby a human could infringe the rights of a machine while the machine could potentially never infringe the rights of either humans or other machines. And this is because the arguments for copyright in AI works unavoidably dissociate copyright from the underlying meaning of authorship.

Authorship, Not Market Value, is the Foundation of Copyright

Proponents of copyright in AI works will argue that the creativity applied in programming (which is separately protected by copyright) is coextensive to the works produced by the AIs they have programmed. But this would be like saying that I have claim of co-authorship in a novel written by one of my children just because I taught them things when they were young. This does not negate the possibility of joint authorship between human and AI, but as stated above, the human must plausibly argue his own “mental conception” in the process as a foundation for his contribution.

Commercial interests vying for copyright in AI works will assert that the work-made-for-hire (WMFH) doctrine already implicates protection of machine-made works. When a human employee creates a protectable work in the course of his employment, the corporate entity, by operation of law, is automatically the author of that work. Thus, the argument will be made that if non-human entities called corporations may be legal authors of copyrightable works, then corporate entities may be the authors of works produced by the AIs they own. This analogizes copyrightable works to other salable property, like wines from a vineyard, but elides the fact that copyright attaches to certain products of labor, and not to others, because it is a fiction itself whose medium is the “personality of the author,” as Justice Holmes articulated in Bleistein.

The response to the WMFH argument should be that corporate-authored works are only protected because they are made by human employees who have agreed, under the terms of their employment, to provide authorship for the corporation. Authorship by the fictious entity does not exist without human authorship, and I maintain that it would be folly to remove the human creator entirely from the equation. We already struggle with corporate personhood in other areas of law, and we should ask ourselves why we believe that any social benefit would outweigh the risk of allowing copyright law to potentially exacerbate those tensions.

Alternatively, proponents of copyright for AI works may lobby for a sui generis revision to the Copyright Act with, perhaps, unique limitations for AI works. I will not speculate about the details of such a proposal, but it is hard to imagine one that would be worth the trouble, no matter how limited or narrow. If the purpose of copyright is to proscribe unlicensed copying (with certain limitations), we still run into the independent creation problem and the possible result that humans can infringe the rights of machines while machines cannot infringe the rights of humans. How does this produce a desirable outcome which does not expand the outsize role giant tech companies already play in society?

Moreover, copyright skeptics and critics, many with deep relationships with Big Tech, already advocate a rigidly utilitarian view of copyright law, which is then argued to propose new limits on exclusive rights and protections. The utilitarian view generally rejects the notion that copyright protects any natural rights of the author beyond the right to be “paid something” for the exploitation of her works, and this cynical, mercenary view of authors would likely gain traction if we were to establish a new framework for machine authorship.

Registration Workaround (i.e., lying)

In the meantime, as Stephen Carlisle predicts in his post on this matter, we may see a lot of lying by humans registering works that were autonomously created by their machines. This is plausible, but if the primary purpose of registration is to establish a foundation for defending copyrights in federal court, the prospect of a discovery process could militate against rampant falsification of copyright applications. Knowing misrepresentation on an application is grounds for invalidating the registration, subject to a fine of up to $2,500, and further implies perjury if asserted in court.

Of course, that’s only if the respondent can defend himself. A registration and threat of litigation can be enough to intimidate a party, especially if it is claimed by a big corporate tech company. So, instead of asking whether AI works should be protected, perhaps we should be asking exactly the opposite question: How do we protect human authorship against a technology experiment, which may have value in the world of data science, but which has nothing to do with the aim of copyright law?

 About the IP Clause

And with that statement, I have just implicated a constitutional argument because the purpose of copyright law, as stated in Article I Clause 8, is to “promote science.” Moreover, the first three subjects of protection in 1790—maps, charts, and books—suggest a view at the founding period that copyright’s purpose, twinned with the foundation for patent law, was more pragmatic than artistic.

Of course, nobody could reasonably argue that the American framers imagined authors as anything other than human or that copyright law has not evolved to encompass a great deal of art which does not promote the endeavor we ordinarily call “science.” So, we may see AI copyright proponents take this semantic argument out for a spin, but I do not believe it should withstand scrutiny for very long.

Perhaps, the more compelling question presented by the IP clause, with respect to this conversation, is what it means to “promote progress.” Both our imaginations and our experiences reveal technological results that fail to promote progress for humans. And if progress for people is not the goal of all law and policy, then what is? Surely, against the present backdrop in which algorithms are seducing humans to engage in rampant, self-destructive behavior, it does seem like a mistake to call these machines artists.

Addressing Fair Use Rhetoric in Debate Over SMART Act

On March 18th, Senators Tillis and Leahy of the IP Subcommittee introduced the SMART Copyright Act. The major functions of the bill, as codified in a proposed new Section 514, would empower the Librarian of Congress to approve designated technical measures (DTM) for identifying infringing material via a triennial rulemaking process. For a detailed description of the proposed rules and remedies in the bill, see Copyright Alliance CEO Keith Kupferschmid’s post. In this post, I wanted to respond to one criticism of technical measures for copyright enforcement—namely that they cannot account for fair use—but first, a recap of the background.

The SMART Act is a legislative response to the fact that after almost 25 years, the OSPs have rarely held up their end of the bargain under the terms of the Digital Millennium Copyright Act (DMCA), ratified in 1998. The foundation for all of Section 512 was predicated on the argument by the service providers of that period that they needed a liability shield against civil litigation stemming from the inevitability that customers would post copyright infringing material to their platforms. Thus, the statute lays out the conditions under which a platform can maintain the “safe harbor” shield, and as many copyright owners know, the major OSPs since then have not always complied with these conditions in good faith.

For the past quarter century, the major platforms have consistently avoided compliance with the notice-and-takedown process by, for instance, erecting unnecessary roadblocks for copyright owners to submit requests. Or, as a Virginia federal court recently affirmed, COX Communications remains on the hook for a billion-dollar damage award due to its failure to comply with the DMCA condition requiring the removal of repeat infringers.

These and many other examples paint a picture of a service platform industry that has fostered a culture of turning a blind eye to infringement and a reluctant, scattershot compliance with the statues they themselves lobbied to write. But specifically in regard to §512(i), which requires collaboration with copyright owners to develop standard technical measures (STMs) for identifying infringing material, Big Tech straight-up ghosted on the matter.

Remember that when the DMCA was being debated and drafted, it was the OSPs who presented their own technological capabilities as an implied promise that STMs could be developed to identify and remove infringing material. But not only did those service providers, and the bigger ones who followed, never engage with copyright owners to develop technical measures, they also funded a network of organizations (you know their names)[1] to promote the general theme that online copyright enforcement is fundamentally bad for society.

Kevin Madigan, VP, Policy and/ Copyright Counsel at Copyright Alliance, describes in a new post how The Network predictably repeats the same unfounded talking points no matter what proposal is on the table. And they have certainly dragged these orcs out of the mud once again in response to the SMART Act. But even as we examine the pros and cons of the bill itself, we should not lose sight of the fact that SMART is a legislative response to the OSPs’ refusal to comply with the conditions their own industry negotiated in the days of Web 1.0. So, maybe we can put the hyperbole in the drawer and sit down like adults? I wouldn’t hold my breath.

Technical Measures Can Never Account for Fair Use?

One gremlin The Network likes to call upon whenever technical measures are discussed for identifying infringing material is that an algorithm can never identify fair use. It is an oddly defeatist argument coming from representatives for an industry that makes bold promises about AI, and hardly allows the lack of perfection to stop them from experimenting with new products. But even if it is true that no algorithm could ever account for fair use, I’ll be blunt and say that neither can most of the platform users The Network claims to represent.

Let’s be real. The minority of professional creators who endeavor to be informed and engaged on copyright matters struggle with fair use; the experts who have worked in copyright law their entire careers struggle with fair use; and the courts struggle with fair use. Naturally, The Network exploits this uncertain landscape to imply that we could never hope for an algorithm to get fair use right. But this logical leap, which is meant to end discussion, also obscures the fact that the average social platform user doesn’t get fair use right either. That is if he even considers the question at all.

The term fair use is bandied about by The Network to promote the idea that it is the default status of most uses of protected works. It is a rhetorical strategy that, when paired with the message that copyright enforcement is inherently a form of censorship, promotes an ideological agenda seeking to elevate the fair use exception to the status of a civil right. But although it is true that the fair use doctrine evolved in U.S. law partly in support of the speech right, it remains an affirmative defense to a claim of copyright infringement. And the distinction matters.

Not only does The Network consistently ignore the literal First Amendment safeguards in the DMCA (which would not be disturbed by the SMART Act), but when they assert that no AI could ever account for fair use, I suspect they are alluding to a much larger constellation of presumed fair uses than actually exists. In other words, the argument against STMs, on the basis of fair use, almost certainly encompasses the effort to expand the volume and types of uses that copyright critics believe should be excepted under the doctrine.

I cannot prove this assumption and would not claim that, if true, it necessarily simplifies the technological challenge at hand. But it if we are going to take the matter seriously at all, it is important to know which definition of “fair use” is being applied—one grounded in case law, or one which the copyright critics would like to revise as they see fit?

As a practical matter, if the average user of a protected work on a social platform makes a fair use of that work, it is more likely the result of dumb luck than a well-informed and carefully considered decision. This is a common-sense assumption based on the low probability that the average user knows anything about the fair use doctrine. And if that is not correct, then perhaps the entire foundation for the liability shield codified in Section 512 should be reconsidered. Because the premise for this whole conversation was, and remains, that the average user of the internet does not know much of anything about copyright law.

Moreover, we forget that the presumed neutrality of the service provider is in contradiction with the idea that a fair use analysis should be a component of technical measures in the first place. The role of platform management was anticipated by the DMCA to be somewhat deaf and dumb in the process. Infringing material would be removed upon receipt of a valid notice, and if the uploader of the material believed the use to be a fair use, he could file a counter-notice to that effect. The human actors on the two sides of this equation were always anticipated to play active roles, and although The Network insists that the low rate of counter-notice filing is predicated on fear alone, it is also plausible that it is the result of many uses of works that are not defensibly fair uses.

What know for sure is that tens of millions of infringing uses occur every day and that only the large, corporate copyright owners have anything close to the resources necessary to mitigate the scope of piracy online. Large platforms like YouTube have deployed their own technical measures (e. g., Content ID and Copyright Match) insofar as they serve the platform’s bottom line. But these systems do little or nothing for independent and small business creators, and leaving this class of professionals in the digital dust was not the intent of the DMCA. Whether some version of the SMART Act can address the problem remains to be seen. But these rhetorical arguments against even trying are as tedious as they are hollow.


[1] In case you don’t, the Electronic Frontier Foundation, PublicKnowledge, Re:Create Coalition, Fight for the Future, Library Copyright Alliance, Authors Alliance, and a host of legal academics.

Image source by: idaakerblom

What Problem Do Those eBook Bills Address Anyway?

In late December, New York Governor Kathy Hochul vetoed the state’s library ebook bill, acknowledging that the law would be preempted by the Copyright Act. In mid-February, a district court in the State of Maryland, responding to a lawsuit filed by the Association of American Publishers (AAP), ordered a preliminary injunction suspending that state’s ebook law, also on preemption grounds. Recognizing which way the wind was blowing, Kyle Courtney of Library Futures Foundation drafted a letter on February 1 to the House Committee on Corporations of the Rhode Island State legislature proposing amendment to that state’s bill, writing:

…we are advising, based on the current landscape involving litigation and vetoes of similar eBooks laws in other states, that you consider friendly amendments below that will effectuate enough changes in H7113 to help avoid running afoul of the challenges documented below with respect to activities in other states.

What follows is a recommendation that Rhode Island remove one paragraph demanding that publishers license to libraries et al., which the footnote describes as the language in direct conflict with federal law. However, the remaining provisions of the bill still invite a preemption challenge because they presume to dictate terms and pricing models to publishers in conflict with the principle that copyright protects the author/owner’s right to decide the manner in which a work is made available. Hence, the provisions that would remain in the RI bill, as well as nearly identical bills in five other states, may still be construed as unconstitutional state compulsory licenses.

As Courtney’s letter emphasizes, the strategic approach taken by the various lobbying organizations pushing for these bills is to present the subject in the context of state contracts while seeking to remedy a consumer protection problem—namely, the alleged “unconscionability in licensing” practices by the publishers. But so far, the organizations lobbying for these bills have yet to support the accusation that current ebook licensing regimes are extortionate and/or that they are causing a disruption in a library system’s ordinary capacity to serve its community. And that’s to say nothing of presenting a compelling case in every state in which these bills have been introduced.

It is no surprise the American Library Association (ALA) et al. have not presented a thorough argument, because it would be a hell of lot of work. To assess whether a given market is underserved (in any context) requires a considerable amount of research and evidence, including counterfactuals, polling, budget analysis, etc. In this instance, it would be a rather large data-science project to manage and model all the relevant inputs, like overall reading trends, library-use trends, preferences for digital vs. physical materials, and cultural and economic data, to determine whether, and where, the ebook borrowing market is underserved and conclude that the licensing models are the cause.

Instead of doing any of that homework, what associations like LFF and the ALA have done instead is to compare the consumer price of an ebook purchase (e.g., $18) to a library price of an ebook license (e.g., $55 for 2 years), then cry foul and draft legislation to resolve this apparent injustice. But if state lawmakers are going to accuse the publishers of unfair practices to justify a law that flies in the face of the Copyright Act, it should demand more evidence than these two numbers alone. Or if state lawmakers are going to elide all complexity in favor of blunt metrics, then why not simply recognize that three times the price to make an ebook available to fifty times the readers hardly sounds like extortion by any reasonable definition?

The Mid-Hudson Library System

Although I certainly do not have the resources or data-science chops to do the kind of research mentioned above, I did a little digging into the Mid-Hudson Library System (MHLS), which serves my home region, just to see what I could learn.

One of 23 systems in New York State, MHLS comprises 76 small-town and public-school libraries in five counties with a total population of more than 686,000 (~ 258,000 households) earning a median income of about $76,000/year. The 2021 budget for the library system was just under $4 million, a little more than half of which comes from statewide and local taxpayers. In 2021, MHLS spent about $90,000 (2.25% of its budget) on digital lending materials, through a few different marketplaces, and presumably using more than one licensing model.

For example, OverDrive, one of the major marketplace platforms where librarians license digital materials, makes ebooks available under three different licensing models. Through Simultaneous Access, certain publishers offer package deals for multiple titles up to a certain number of loans. In the One Customer One Use model, presumably for back catalog or less popular books, the licenses never expire. And the model most often used by the major publishers for the most popular books is Metered Lending, which offers one or two-year licenses and/or limits the number of loans per license.

In 2021, MHLS ebook circulation was ~ 314,000, and the first three months of 2022 are tracking toward a similar total. Even at the unrealistic frequency of one book per unique patron, that would be less than 1/3 of the total population in the system, which likely says more about demand than it does about supply. In fact, at the national level, although ebooks and audiobooks continue to occupy a greater percentage of a library’s collection, print book borrowing is still 518.92% higher than ebook borrowing as of 2019.

Looking at the catalog, it appears that MHLS offers about 10,000 ebooks (70% fiction/30% nonfiction), presumably under more than one licensing model. But even if all 10,000 were licensed under Metered Lending at a rate of $55 for two years, this amounts to a cost of about $1.07/year per household in the system. Alternatively, we can estimate that a two-year license of $55, at a maximum rate of one loan every two weeks ($55 / 52 readers), is a Cost Per Loan (CPL) of about $1.06.

So, the numbers available do not seem to justify even a hypothesis that ebook licensing is unduly burdensome or is resulting in underserving the MHLS community. And the overall demand nationwide for borrowed ebooks hardly justifies the rhetoric of the lobbyists, who would have us believe that a literature-starved public is suffering on the libraries’ virtual steps at the mercy of the big publishers. When an expenditure is just over two percent of the operating budget, one must step back and look more holistically at the question presented.

Collections Are a Fraction of a Library’s Expense

The data collected in the Institute of Museum and Public Services (IMLS) Public Library Survey reveals that libraries’ costs are increasing for personnel and general operating expenses while costs are trending downward for collection materials—especially the cost of ebooks and audiobooks. Noting that most libraries spend an average 10% of their annual budgets on their collections overall, an article in Wordsrated summarizing the IMLS Survey states, “The drop in price per item is due to library collections becoming increasingly digital. This is because the price per digital item has declined significantly. All while the average cost per book increased 10% since 2003.”

The statistical trends in the IMLS Survey suggest that libraries are going through a lot of transition these days—as collections become more digital, as physical spaces are adapted to provide more programs and services, and as overall reading and borrowing habits continue to shift in the market. Change in any system presents both opportunities and challenges, and it is a safe bet that not every local library will, or can, adapt in the same way. But if the data show that ebooks are, as of 2019, “the cheapest material in a library’s collection,” then why on Earth is this the moment to lobby for these ebook bills in the states?

The answer to that cannot be, “Well, if the prices were even lower, we could do more.” Yeah. That’s how everything in life works. But for one thing, as much as publishers and authors care quite a bit about library patrons, it is not incumbent upon them to outright subsidize the libraries as they navigate the changing landscape—let alone by mandating that the publishers remain bound by old models so that libraries can adapt to new ones. That’s not a symbiotic relationship.

Looking forward, neither the libraries nor the publishers can say what the trends will be in five or ten years, but the libraries should be cautious about putting too many eggs in the ebooks basket. What happens to the relevance of the seventy or so local libraries in MHLS if the system plays an outsized role as a conduit for ebook lending? Don’t at least some taxpayers or prospective donors in each town begin to wonder why they need to keep paying the librarians and maintaining the buildings? Perhaps the local librarians should look at the data and ask whether ALA, LFF et al are doing them any favors.

Of course, knowing the track records of the people behind these ebook bills, it is fair to doubt that they are trying to solve a problem at all but are instead pursuing a broad, anti-copyright agenda. The tone of Courtney’s letter, for instance, makes clear that he (and his colleagues) object to the legal doctrines on which the NY and MD bills were opposed and that his recommendations to RI are a begrudging pivot in strategy to achieve the same ends by a slightly amended rationale.

But to oblige any copyright owner to make a work available under terms mandated by state law invites substantial conflict with federal law and the authority of Congress alone to amend that law. Consequently, no state legislature should embark on such an adventure without a compelling and thorough analysis of the problem allegedly being solved. And so far, the lobbyists for these ebook bills have presented little more than a melodrama barely worth reading at any price.