Latest Posts

Record Labels File Suit Against Internet Archive for Copyright Infringement

Pride goeth before destruction, and an haughty spirit before a fall. (Proverbs 16:18 KJV)

Citing 2,749 works in suit, six of the major music labels (UMG, et al.) have filed a multi-count complaint against Internet Archive (IA), Brewster Kahle personally, Kahle’s foundation, and an audio digitizing service operated by an individual named George Blood. Total potential damage award with legal fees:  around a half-billion dollars. Likelihood of defendants’ success, assuming all factual allegations are well-founded:  less than zero. So, while I have no idea how much cash on hand Kahle has to burn, this suit highlights a question I have often asked myself—namely how eager is he to put his money where his anti-copyright mouth is?

As the outcome in the book publishers’ lawsuit, Hachette et al., makes clear, cockamamie theories about how the law works may find an audience in the blogosphere, but they make poor arguments in court. In that case, Internet Archive relied on a cockamamie theory called Controlled Digital Lending (CDL) and hitched that wagon to a belief that the practice was shielded by the doctrine of fair use. The defendant lost on every point, and a negotiated judgment is already filed with the court notwithstanding IA’s right to appeal.

In this new case with the record labels, IA does not even have the gossamer of an unfounded theory to weave into its response. Instead, the initiative IA calls “The Great 78 Project” is alleged to entail knowing evasion of compliance with clearly defined statute. Without going into each of the counts against each of the defendants, the crux of the matter is that Great 78 makes digital copies of sound recordings from 78RPM vinyl records and hosts those files for unlimited streaming or downloading. So, if any of those sound recordings are still under copyright, this implicates violation of three of the exclusive rights under Section 106 of the Copyright Act—reproduction, distribution, and public performance by digital audio transmission (§106(1), (3), & (6) respectively).

The Great 78 Project purports to make available rare and difficult-to-find sound recordings, and presumably, some portion of the collection comprises works in the public domain (PD) and/or truly rare works that are not commercially available. But headlining the more than 2,000 works in suit, the complaint cites recordings that are neither in the PD nor rare by any means. Popular recordings by Elvis Presley, Duke Ellington, Billie Holiday, Ray Charles, Chuck Berry, Frank Sinatra, Ella Fitzgerald, Louis Armstrong, and Hank Williams are named as prime examples that can be accessed by legal, commercial means, including major streaming services.

The reason commercial availability of the sound recordings is relevant in this case is that under the provisions of the Music Modernization Act (MMA) of 2018, a library/archive is permitted to make pre-1972 sound recordings available if, among other conditions, it makes a good-faith effort to determine that the recordings are not commercially available. That is a pared down description, but it’s the basic principle, which IA apparently chose to ignore. According to the complaint, IA made no effort to fulfill its obligation to comply with any of the following Copyright Office guidelines:

…a reasonable search for purposes of 17 U.S.C § 1401)(c) must include, among other things: (i) searching the Copyright Office’s database of indexed schedules listing right owners’ pre-1972 sound recordings; (ii) searching Google, Yahoo!, or Bing; (iii) searching at least one of the following streaming services: Amazon Music Unlimited, Apple Music, Spotify, or TIDAL; (iv) searching YouTube; (v) searching SoundExchange’s repertoire database; (vi) searching at least one major seller of physical product, namely Amazon.com.

Moreover, the complaint cites compelling evidence that the defendants understood their obligations under §1401 and that failure to comply would constitute copyright infringement of works like the sound recordings in suit. For instance, IA stated in a blog post about the MMA shortly after it was signed into law, “But, as we understand it, the MMA means that libraries can make some of these older recordings freely available to the public as long as we do a reasonable search to determine that they are not commercially available.” The only logical conclusion, therefore, is that defendants ignored the “reasonable search” guidelines because it is obvious that many of the sound recordings at issue can be found commercially available by a young child using Google.

Will Kahle’s Copyright Hubris Kill His Archive?

A significant distinction between this suit and the Hachette case is that Brewster Kahle and the Kahle/Austin Foundation are named defendants. The complaint alleges that Kahle is directly involved in IA policy, activities, and promotion, including the Great 78 Project, that he funds the foundation through his trust, and the foundation, in turn, funds the project. “At Kahle’s direction, the Foundation used the funds Kahle had contributed to sponsor the Internet Archive’s massive and growing infringement. The Foundation donated money to Internet Archive that Internet Archive used to pay costs in furtherance of its infringement…” the complaint states.

The Internet Archive and its friends will, no doubt, repeat populist claims that they are serving the public, behaving as a library should, and that they are being targeted by a greedy industry. But the conduct alleged in the UMG complaint reveals an even more brazen decision to circumvent copyright law than the CDL scheme underlying the Hachette suit. In the book publishers’ case, IA advanced a theory (albeit a poor one) that it was acting within the confines of the law, but here, it simply elected to evade clearly articulated statutory confines and take its chances. And this time, the cost could indeed be the whole operation.

I get why many people want to support IA, not the least being that a large part of the organization is both legal and highly useful. Among those who simply agree with Kahle et al. that copyright should not exist, that pre-1972 sound recordings should not be protected, etc., fine. That’s an opinion to which people are entitled. But beyond that general view, IA supporters should not be confused into thinking this case is about big bad industry beating up on a library doing library-like things. Assuming the factual allegations are correct, there is barely a distinction between the alleged infringing conduct in the Great 78 Project and The Pirate Bay. And to the extent the legitimate archive has been treated like a front for mass infringement projects, the blame for that decision rests with Kahle and his colleagues, not with the music or book publishers.

Since the first post I wrote about Internet Archive, I have acknowledged that the repository of public domain and truly rare material is an invaluable research tool. In fact, that post in October of 2017 asked directly whether the anti-copyright rhetoric was necessary to the organization, but since then, it has become clear that Kahle has used IA’s operation and reputation to engage in much more than rhetoric. And in this potentially costly litigation with the record labels, it is conceivable that this hubristic crusade against copyright law could, as the proverb says, lead to the collapse of an otherwise good enterprise.


Photo by: panoramaimages

Before Generative AI, Big Tech Taught Artists to Abdicate Copyright Rights

One of the more challenging aspects of copyright advocacy is the fact that many artists and creators are conflicted about enforcing their own rights, and from observation, the disconnect is ideological. For the last 30 years, copyright skepticism has been woven into political narratives rooted in criticism of corporations and the excesses of capitalism—popular themes among the political left, which encompasses most artists. Now that generative AI developers are turning creative works into “pink slime,” and artists are suddenly more interested in their rights, it might help to recognize that the industry deploying AI is the same one that taught creators to advocate against copyright in the first place.

The year 2011 was an extraordinary time to jump into the fray. It was immediately apparent that allegations of “copyright maximalism” were deeply intertwined with a sincere and animated belief that the internet would foster a new and potent form of direct democracy to confront a litany of injustices. Copyright enforcement was characterized as a barrier to that promise, and so, the Stop SOPA campaign (to kill anti-piracy legislation) became part of a larger, frenetic collage that included OWS protests, European pirate parties, Anonymous, Wikileaks, etc., all feeding an atmosphere of revolution that corresponded with headlines and memes claiming that “Hollywood” wanted to use copyright to break the internet and stifle speech.

But Big Tech’s promise to democratize everything was a Trojan Horse from which the AI bots have now emerged to ransack the village. Not only did promoters of the “free flow of information” elide the fact that their platforms were as likely to produce the January 6th insurrection as the “Pussy Hat” March, but the allegation that copyright was a barrier to information flow had nothing to do with liberating our speech and everything to do with limiting their liability.

Every time members of the creative community echoed the anti-copyright messages pumped out by Fight for the Future, the EFF, Public Knowledge, or the platforms themselves, what was really being advocated was a lack of accountability for online service providers. I never fully understood how one of the most exploitative industries in history managed to turn anti-corporatist sentiment to its advantage, but I assumed it was the gestalt of the internet. The illusion that social platforms belong to the people was a charade that enabled Google, Facebook, et al. to camouflage their interests as our rights.

That theme has aged about as well as the tobacco industry’s efforts to sell freedom to get smokers to ignore cancer, but it’s been almost two years since Big Tech’s “Big Tobacco moment,” and little has changed. Neither in Congress nor the courts have online service providers been held accountable for much of anything—and that’s with laws on the books. When we consider that, for almost three decades, the major platforms have acted in bad faith with their end of the DMCA bargain, and the courts have interpreted Section 230 as an unlimited liability shield, it is hard to feel hopeful about a legal framework for accountability for harms resulting from AI.

In fact, certain AI tools (e.g., LLMs) may imply a wider “neutral” buffer between potentially harmed parties and potentially liable parties. “Knowledge” and “intent” are key factors in establishing liability, and we have watched Big Tech play shell game with the concept of what they can “know” or “intentionally” control about activity on their platforms. AI tools could take these shenanigans to the next level, enabling new forms of harm with an even weaker nexus linking the machines to the people who design and operate them.

In the copyright world, platform operators have consistently circumvented their obligations under the DMCA with shrugging statements like We can’t police the internet, alluding to staggering volume while conjuring an association with authoritarianism. Now, the circumstances are different. It is a near certainty that every creative work made has been, or will be, ingested into one or more AI training models, and unless the courts find this to be an act of mass piracy and order disgorgement of the datasets, creators may have to accept that their work is being turned into pink slime.

While it is encouraging to see artists take a more active interest in copyright rights as a response to AI, it is also a bittersweet transition in light of all that has happened so far. Whatever comes next, I hope the creative community will recognize that copyright rights are the closest thing to labor rights the independent artist has. And these rights should not be weakened or abandoned for the sake of more billionaires making false promises about democracy and free speech.

EFF to Honor Scientific Paper Pirate Sci-Hub

The Electronic Frontier Foundation (EFF) announced that among the 2023 recipients of the EFF Award (formerly the Pioneer Award), it will honor Sci-Hub founder Alexandra Asanova Elbakyan this September. The Russian-based Sci-Hub is an enterprise-scale pirate site specifically built to host scientific papers about which the EFF states:

Through Sci-Hub, Elbakyan has strived to shatter academic publishing’s monopoly-like mechanisms in which publishers charge high prices even though authors of articles in academic journals receive no payment. She has been targeted by many lawsuits and government actions, and Sci-Hub is blocked in some countries, yet she still stands tall for the idea that restricting access to information and knowledge violates human rights. 

In addition to the EFF’s usual flare for the dramatic, the organization continues to flaunt its unwavering hostility toward all copyright rights as a foundational raison d’etre. Note the word shatter in that quote above. To the ideologues at EFF et al., Sci-Hub is not merely a response to subscription fees but should be revered for its assault on the very idea that journal publishers ought to exist in the first place. In fact, the timing of this award is telling in that it comes two years after a landmark, open access agreement was negotiated by the University of California (UC) with publishing giant Elsevier.

Although subscription cost has often been a point of contention for many in the academic community, even contributing authors, the UC has been negotiating open access agreements with academic publishers on behalf of colleagues at smaller institutions with more limited resources. “We refer to these agreements as transformative open access agreements because they convert subscription payments into payments for open access publishing (with reading provided for free). It is a new approach we helped develop with other leading institutions a few years ago, in large part through the OA2020 initiative,” says librarian and economics professor Jeffrey MacKie-Mason of UC Berkeley.

Settled in March of 2021, the Elsevier deal was the ninth open access agreement the UC negotiated with academic publishers, which suggests that many authors of scientific research papers do not view a pirate site like Sci-Hub as a viable “solution” to whatever criticisms they have of the commercial publishers. While I do not presume to have expertise about the complex world of scientific journal publishing, a 2018 article by industry consultant Kent Andersen lists “102 things journal publishers do,” and it’s a lot more than hosting PDFs on a website. Notably, even among some of the critical comments on that article, which appear to be written by academic authors, there is no mention of Sci-Hub in particular, or piracy in general, as obviating the role played by journal publishers in the industry.

Although it is true (as the EFF emphasizes) that the authors of these papers are not paid for their writing, the academic publisher is more like a venue operator than a trade book publisher—a venue operator that, at best, serves as a neutral party to control quality. These journals invest substantial resources to review millions of submissions, prepare documents, maintain databases, check for plagiarism, organize peer review, etc. And although these investments need to be recovered profitably for the publisher to exist, the UC deals indicate that there is room for negotiation, which leaves EFF’s panegyric to Sci-Hub sounding as hollow as it is untimely.

Speaking of timing, with academics, policymakers, journalists, artists, and just about everyone else wondering how badly generative AI might exacerbate the misinformation problem, could there be a worse moment to award a pirate of scientific journals? How is Sci-Hub not the natural place for a generative AI developer to harvest scientific writing to train an algorithm to, perhaps, “write” papers without scientists? Notably, in the class-action case Tremblay et al. v. OpenAI, the plaintiffs allege that the defendant obtained literary works for machine learning (ML) from “shadow libraries,” (i.e., pirate sites like Z-Library). So, by the same logic, Sci-Hub would seem to be a natural source where an AI developer can scrape scientific papers.

I am neither motivated nor qualified to critique the entire scientific publishing ecosystem, let alone to dispute complaints among some academics about cost, et al. I would grant Elbakyan the benefit of the doubt that her intent is at least distinguishable from the typical entertainment media pirate whose only motive is financial, and I recognize that scientists and academics in various regions have access Sci-Hub for what may be difficult to obtain information. Nevertheless, the worn out view that piracy is a solution to imperfections in a given system is, at best, narrowly focused on distribution while ignoring the means and motives for production.

Not unlike Peter Sunde’s mourning the lost Marxist idealism he saw in the The Pirate Bay, Elbakyan echoed this same naivete when she told the Washington Post in 2016, “On my website, any person can read as many papers as they want for free, and sending donations is their free will. Why Elsevier cannot work like this, I wonder?” Indeed. The alleged “white hat” pirate never seems to grasp that there is always a cost to production and that, whatever system covers that cost, it won’t be a damn tip jar, and it will rely on copyright in some form. As the court stated in 2015 when Elsevier successfully sued Sci-Hub for infringement, “Elbakyan’s solution to the problems she identifies, simply making copyrighted content available for free via a foreign website, disserves the public interest.”

As for the EFF Award, it’s worth asking what Sci Hub’s agenda is in 2023, if indeed traditional publishers are adopting open access agreements and academics are still willing to work with those publishers? Is it truly Elbakyan’s mission to “shatter” the entire scientific publishing ecosystem and, with it, essential processes like peer review? Or is that just the EFF’s hyperbole? Presumably, it’s both. And by honoring Sci-Hub, the EFF proves once again that it will promote any anti-copyright agenda—legal or otherwise—with the zeal of a conspiracy theorist watching “chemtrails” fill the sky.