Books are Not Floor Wax and Road Salt

One would think this is obvious, particularly to a librarian, but perhaps not to Douglas Lord, President of the Connecticut Library Association (CLA). In a letter addressed to the state assembly advocating passage of H.B. 6829, Connecticut’s version of similar bills proposed (and shot down) in other states to address alleged unfairness in eBook licensing to libraries, Lord writes:

It is very important to note that this legislation has nothing to do with copyright, it is a matter of contract law. In the same way that taxpayer funds are treated preferentially with all other state contracts – from floor wax to vehicles to road salt – the same should be true for electronic content. [Emphasis included]

Although the Connecticut bill does not require publishers to license to libraries in the state, it contains several provisions defining various publishers’ licensing models as “unfair trade practice,” which is tantamount to a state compulsory license, which means H.B. 6829 is preempted by the Copyright Act. So, it has something to do with copyright law. In fact, although I am sure Lord does not sincerely equate books to floor wax and road salt, his disregard for the unique cultural value of the former may explain his absurd allegation that copyright law is not implicated in a state bill about contracts. Every contract negotiated for the use of copyrightable works rests upon the author’s exclusive rights enumerated in Section 106 of Title 17. So, Lord’s declaration is either intentionally misleading or naively misguided.

Notably, Lord’s letter reiterates the ambiguous rationale that has been proffered by every advocate of these bills in every state so far—i.e., the difference between the retail price of an eBook purchase compared to the licensing models that publishers offer to libraries. He states, “Consumers pay, on average, $12.77 for eBooks from retailers like Amazon. The average cost for a public library for the exact same product is $45.75.” Indeed, if one does not look beyond those two numbers or gather any relevant community information, the price comparison looks outrageous, even extortionate.

But to address this issue, I did my best to examine the market in my own region served by the Mid-Hudson Library System and found that a) less than one-third of the MHLS community accesses the library system for books of any kind; and b) that the average eBook cost per read is ~$1.06. And apropos the big picture for the taxpayer, it is notable that maintaining a library’s collection—both physical and electronic—is usually a fraction of its operating costs. To quote my post looking at MHLS:

The data collected in the Institute of Museum and Public Services (IMLS) Public Library Survey reveals that libraries’ costs are increasing for personnel and general operating expenses while costs are trending downward for collection materials—especially the cost of ebooks and audiobooks. Noting that most libraries spend an average 10% of their annual budgets on their collections overall, an article in Wordsrated summarizing the IMLS Survey states, “The drop in price per item is due to library collections becoming increasingly digital. This is because the price per digital item has declined significantly. All while the average cost per book increased 10% since 2003.”

While $1.06 per read does not seem extortionate, I do not claim to know whether that price is “fair to the taxpayer” in New York or Connecticut or anywhere else. But that’s my point. No advocate of these eBook bills, to my knowledge, has attempted to demonstrate a critical need for this legislation based on cost/benefit numbers, which is odd when one is alleging unfair use of public funds. And I suspect that’s because these bills are not directed at solving a real problem but are instead the hobbyhorse of anti-copyright activists like Jonathan Band and Kyle Courtney. Consequently, it is no surprise that advocacy of these bills, including this letter from CLA, repeats the vague tautology that publishers are extortionate and usually ignores the interests of authors.

Here, Lord goes a step further and claims that “Authors get no added royalties or income from these sales.” Not true. Authors’ contracts include revenue from eBook licensing to libraries, and the author’s percentage of eBook revenue is usually higher than her cut from physical book sales. Plus, those percentages typically increase as sales go up, advances are covered, etc. So, I am not sure whence Lord’s assertion comes, but it is consistent with the logic behind this bill—that books are like other commodities, and the author’s pecuniary interests are not directly associated with her copyright rights.

As I’ve repeated in nearly every post on this topic, the libraries should be careful what they wish for when it comes to eBook licensing and, if they hope to remain relevant, should avoid putting too many eggs in the digital basket. The logic is not hard to follow. If 90% of the cost of keeping libraries open is not about the collection, and the digital collection grows too large, how long before taxpayers figure out that facilitating eBook loans can be done with a website and without those expensive buildings and librarians? After all, some taxpayers may think that a former library would be a handy place to stockpile floor wax and road salt.


Photo by: AndreyPopov

When the State Steals Your Work – Podcast with Rick Allen

In March 2020, the Supreme Court delivered its opinion in the case Allen v. Cooper. The outcome was not surprising because the Court affirmed precedent ruling from the late 1990s which held that the 11th Amendment bars suing a state or state actors for damages stemming from intellectual property infringement.

Thus far, I’ve explored the murky waters of state sovereign immunity as it relates to Allen v. Cooper and other cases, including author Michael Bynum and photographer Jim Olive’s lawsuits filed in the State of Texas. So far, my focus in this area has been academic. But on February 8th, Rick Allen filed an amended complaint in North Carolina, and after I read that narrative, I wanted to invite Rick back to the podcast to talk more personally about his story, what it means to him, and what it should mean to anyone who hears it.


Show Contents

  • 1:15 Becoming an Underwater Cameraman
  • 11:06 Queen Anne’s Revenge Opportunity of Lifetime
  • 15:06 Wreck Diving and Filming
  • 27:19 Personal Investment
  • 37:20 Rare Cooperation Between Treasure Hunters and Archeologists
  • 43:00 A Near-Fatal Accident
  • 48:00 State Infringements
  • 59:00 Blackbeard’s Law
  • 1:05:00 Suing the State of North Carolina
  • 1:16:00 Implications for All Creators
  • 1:29:00 Overlap with Censorship

Photo of Rick Allen by Cindy Burnham.

Why Machine Training AI with Protected Works is Not Fair Use

As most copyright watchers already know, two lawsuits were filed at the start of the new year against AI visual works companies. In the U.S., a class-action was filed by visual artists against DeviantArt, Midjourney, and Stability AI; and in the UK, Getty Images is suing Stability AI. Both cases allege infringing use of large volumes of protected works fed into the systems to “train” the algorithms. Regardless of how these two lawsuits might unfold, I want to address the broad defense, already being argued in the blogosphere, that training generative AIs with volumes of protected works is fair use. I don’t think so.

Copyright advocates, skeptics, and even outright antagonists generally agree that the fair use exception, correctly applied, supports the broad aim of copyright law to promote more creative work. In the language of the Constitution, copyright “promotes the progress of science,” but a more accurate, modern description would be that copyright promotes new “authorship” because we do not tend to describe literature, visual arts, music, etc. as “science.”

The fair use doctrine, codified in the federal statute in 1976, originated as judge-made law, and from the seminal Folsom v. Marsh to the contemporary AWF v. Goldsmith, the courts have restated, in one way or another, their responsibility to balance the first author’s exclusive rights with a follow-on author’s interest in creating new expression. And as a matter of general principle, it is held that the public benefits from this balancing act because the result is a more diverse market of creative and cultural works.

Fair use defenses are case-by-case considerations and while there may be specific instances in which an AI purpose may be fair use, there are no blanket exceptions. More broadly, though, if the underlying goal of copyright’s exclusive rights and the fair use exception is to promote new “authorship,” this is doctrinally fatal to the proposal that training AIs on volumes of protected works favors a finding of fair use. Even if a court holds that other limiting doctrines render this activity by certain defendants to be non-infringing, a fair use defense should be rejected at summary judgment—at least for the current state of the technology, in which the schematic encompassing AI machine, AI developer, and AI user does nothing to promote new “authorship” as a matter of law.

The definition of “author” in U.S. copyright law means “human author,” and there are no exceptions to this anywhere in our history. The mere existence of a work we might describe as “creative” is not evidence of an author/owner of that work unless there is a valid nexus between a human’s vision and the resulting work fixed in a tangible medium. If you find an anonymous work of art on the street, absent further research, it has no legal author who can assert a claim of copyright in the work that would hold up in any court. And this hypothetical emphasizes the point that the legal meaning of “author” is more rigorous than the philosophical view that art without humans is oxymoronic. (Although it is plausible to find authorship in a work that combines human creativity with AI, I address that subject below.)

As a matter of law, the AI machine itself is disqualified as an “author” full stop. And the although the AI owner/developer and AI user/customer are presumably both human, neither is defensibly an “author” of the expressions output by the AI. At least with the current state of technologies making headlines, nowhere in the process—from training the AI, to developing the algorithm, to entering prompts into the system—is there an essential link between those contributions and the individual expressions output by the machine. Consequently, nothing about the process of ingesting protected works to develop these systems in the first place can plausibly claim to serve the purpose of promoting new “authorship.”

But What About the Google Books Case?

Indeed. In the fair use defenses AI developers will present, we should expect to see them lean substantially on the holding in Authors Guild v. Google Books—a decision which arguably exceeds the purpose of fair use to promote new authorship. The Second Circuit, while acknowledging that it was pushing the boundaries of fair use, found the Google Books tool to be “transformative” for its novel utility in presenting snippets of books; and because that utility necessitates scanning whole books into its database, a defendant AI developer will presumably want to make the comparison. But a fair use defense applied to training AIs with volumes of protected works should fail, even under the highly utilitarian holding in Google Books.

While people of good intent can debate the legal merits of that decision, the utility of the Google Books search engine does broadly serve the interest of new authorship with a useful research tool—one I have used many times myself. Google Books provides a new means by which one author may research the works of another author, and this is immediately distinguishable from the generative AI which may be trained to “write books” without authors. Thus, not only does the generative AI fail to promote authorship of the individual works output by the system, but it fails to promote authorship in general.

Although the technology is primitive for the moment, these AIs are expected to “learn” exponentially and grow in complexity such that AIs will presumably compete with or replace at least some human creators in various fields and disciplines. Thus, an enterprise which proposes to diminish the number of working authors, whether intentionally or unintentionally, should only be viewed as devastating to the purpose of copyright law, including the fair use exception.

AI proponents may argue that “democratizing” creativity (i.e., putting these tools in every hand) promotes authorship by making everyone an author. But aside from the cultural vacuum this illusion of more would create, the user prompting the AI has a high burden to prove authorship, and it would really depend on what he is contributing relative to the AI. As mentioned above, some AIs may evolve as tools such that the human in some way “collaborates” with the machine to produce a work of authorship. But this hypothetical points to the reason why fair use is a fact-specific, case-by-case consideration. AI Alpha, which autonomously creates, or creates mostly without human direction, should not benefit from the potential fair use defense of AI Beta, which produces a tool designed to aid, but not replace, human creativity.

Broadly Transformative? Don’t Even Go There

Returning to the constitutional purpose of copyright law to “promote science,” the argument has already been floated as a talking point that training AI systems with protected works promotes computer science in general and is, therefore, “transformative” under fair use factor one for this reason. But this argument should find no purchase in court. To the extent that one of these neural networks might eventually spawn revolutionary utility in medicine or finance etc., it would be unsuitable to ask a court to hold that such voyages of general discovery fit the purpose of copyright, to say nothing of the likelihood that the adventure strays inevitably into patent law. Even the most elastic fair use findings to date reject such a broad defense.

It may be shown that no work(s) output by a particular AI infringes (copies) any of the works that went into its training. It may also be determined that the corpus of works fed into an AI is so rapidly atomized into data that even fleeting “reproduction” is found not to exist, and, thus, the 106(1) right is not infringed. Those questions are going to be raised in court before long, and we shall see where they lead. But to presume fair use as a broad defense for AI “training” is existentially offensive to the purpose of copyright, and perhaps to law in general, because it asks the courts to vest rights in non-humans, which is itself anathema to caselaw in other areas.[1]

It is my oft-stated opinion that creative expression without humans is meaningless as a cultural enterprise, but it is a matter of law to say that copyright is meaningless without “authors” and that there is no such thing as non-human “authors.” For this reason, the argument that training AIs on protected works is inherently fair use should be denied with prejudice.


[1] Cetaceans v. Bush holding that animals do not have standing in court was the basis for rejecting PETA’S complaint against photographer Slater for infringing the copyright rights of the monkey in the “Monkey Selfie” fiasco.