Is “Machine Learning” Copying or Reading?

machine reading

I recently attended a round-table discussion on the subject of artificial intelligence and copyright.  The first of several engaging topics I thought warranted a post was the question of “machine learning,” which I put in quotes here with respect to one scholar who admonished against anthropomorphizing AI by using words for human activities to describe the actions of computers.  I think that view is fundamentally correct, though there is also grounds for analogy, as will be made clear by the following premise:

When you read a book, even if we might say, by way of analogy, that you are “copying” the content of that book onto your brain, this clearly does not infringe §106(1) of the copyright law proscribing unauthorized copying.  Since the author naturally hopes that you will read her book, such a prohibition would be absurd, even if you had an eidetic memory and could, if prompted, recite the entire work verbatim.  But if you used that gift to type from memory the entire book and made that document available, you would then violate more than one statute under the copyright law.

So, the question raised in regard to “machine learning” is whether the computer scientist who wishes to feed a corpus of books—say the anthology of American literature—into an AI should be required to obtain licenses for the works still under copyright.  Thus, the first analysis is whether the act of “copying” can be said to occur in this circumstance any more than it would be for the human reader who consumes the same body of literature.

It strikes me that if what the AI does in this case is ingest the corpus of books and almost instantly deconstructs those works by synthesizing them through a neural network, then the computer scientist has a pretty solid argument that no copying has taken place.   If the machine does not retain intact copies of works—or even large sections of works—-with the purpose of making those intact copies available to the human market, then this “machine reading” process is arguably analogous to the human whose reading does not infringe §106(1) of the copyright law.

That said, intent of the computer scientist may be a significant factor.  For instance, if the training of the AI will have a commercial purpose, this may suggest a requirement to license the works under copyright.  But intent can be very tricky on the leading edge of science because it is neither realistic, nor even desirable, to insist that every researcher know exactly where his experiments will lead.  This would nullify the process of discovery whence many great achievements have been made; hence, discovery is justification itself, and I suspect the tech companies would appeal to this rationale in regard to “machine learning.”

If the computer scientist’s goal is to see whether he can get his AI to “learn” about the American experience through literature, but he does not have a particular product or service in mind at the outset, it seems that copyright owners would be on fairly shaky ground to enjoin his use of the books.  As long as nothing that comes out the other end looks like any of the products that went in, it strikes me that this experiment exists beyond the statutory framework of copyright law.

Of course this portrait of the individual scientist beavering away in his modest lab to see what he may discover is not what is taking place in reality. We know perfectly well that major AI experimentation occurs in the R&D labs of companies like Google and Facebook, who are well shielded by trade-secret law from divulging what they are working on or for what purpose.  Like any other corporations, they are free to announce a new product or service without telling the public how they arrived at the latest result.

So, even if the use of copyrighted works as source material resulting in a commercial end might recommend some type of licensing regime, it may be very difficult to identify the threshold when the blind process of scientific discovery becomes a clear intent to exploit a commercial opportunity.  And, as mentioned, these companies would be under no obligation to divulge that eureka moment to anyone.  

On the other hand, the moment Google or Facebook did announce that new product, rightsholders could justifiably complain that a massive, highly-profitable corporation has used potentially billions of dollars worth of material without paying for any of it.  As one scholar at the round-table noted, tech companies may not use raw silicon for free, so why should they get to exploit millions of creative works for free, no matter what they’re turning that data into?

It’s a good question.  One that would seem to suggest a new subsection of the copyright law, and this would certainly be consistent with the fact that new forms of exploitation of works may demand equally new forms of compensation.  If nothing else, that type of statutory response could spare us all the tedious and false harangue that insists “copyright owners just want to stand in the way of innovation.”

That argument prevailed for far too long, and now the so-called innovators have a lot of splainin’ to do about their culture of blind disruption for the sake of disruption. Especially in light of the fact that AI may have some very profound effects on society as we know it, maybe this time around the copyright owners should be treated like experienced voices in the conversation rather than canaries wasting their breath in the proverbial coal mine.

“The internet has failed.”

T Bone Burnett has chosen to remove the video of his excellent keynote address at SXSW 2019 but has graciously made the text available to Illusion of More. Read the full speech here.

” … today there is a growing understanding that the internet has morphed into an insidious surveillance and propaganda machine.”

The Internet is Not (and never was) Paradise

I was reading an editorial the other day written by Stephen Witt for NPR shortly after the passing of John Parry Barlow in 2018; and it occurred to me that internet activists seem to fit one of two profiles—Mourners and Evangelicals. And both are full of shit.

Witt does an excellent job summarizing the early barefoot wanderings of the college-dropout, Grateful Dead lyricist, turned techno-libertarian prophet who would eventually co-found the Electronic Frontier Foundation …

It was 1985, and Barlow, not a computer person, did not know what “online” was. But he wangled an Internet account out of a Stanford academic — they were not available to the general public at the time — and began to anonymously visit Deadhead forums on Usenet, one of the earliest hosts for Internet discussion. Despite an apparently fatal lack of any STEM education, Barlow grasped the technology’s potential. “I had a religious experience upon encountering what was a very small online environment,” he said. “I felt that what I was looking at was something profoundly different than anything that had happened in the history of the human race.

The spirit of Witt’s article Tech Utopianism And Our Walled Gardens: Is It Time For A Jailbreak? places it among the many laments for the internet as a paradise lost.  Like other articles of its kind, Witt’s homage to Barlow harkens to an ideal that never existed—a cybernetic Eden, where the purity of human mind and spirit might have remained unsullied had it not been for the original sin of commerce that cast us into the hyper-monetized, surveillance-capitalized, barely-civilized landscape dominated by today’s billion-dollar platforms.  

Not surprisingly, Witt alludes to the fact that copyright infringement was a foundational rite of the new cyber-religion evangelized by the prophets; and it is just a little too perfect that, as an ambassador of the Dead (the most famous band to encourage bootlegging its live performances) Barlow and disciples viewed intellectual property theft as a pathway to the promised land …

… if information was instantly reproducible at no cost, only by creating barriers to open communication between private individuals could the now-artificial scarcity of copyright be maintained.  A true cyberlibertarian — and perhaps we should call him an anarchist — Barlow took the extreme position, denying that the state had the authority to limit peer-to-peer communication. This necessitated an abandonment of the concept of intellectual property, even if that proved corrosive to both the profit margins of large corporations and the meager income streams of small songwriters, including Barlow’s own.

I will admit that my cynicism here is colored by the fact that a world resembling an endless Dead show is my own version of Hell, but personal taste is also germane to the broader point that utopias always fail because they presume to impose a monolithic world view on everyone.  (One man’s Paradise is always another’s Purgatory.)  And that presumptuousness is certainly a running theme wherever digital activism embraces the anti-copyright agenda—too often insisting that all artists must adopt the “sharing” attitude espoused by The Grateful Dead, overlooking the nagging bugaboo that choice is the foundation of liberty.  

So, in regard to the internet writ large, Witt’s elegy fits the profile of the Mourner’s view of cyberspace—a resignation to the fact that utopia is gone and can never be rediscovered, and that any hope of building Paradise anew should be abandoned.  We cannot return and so might as well unplug. 

But while the Mourners have discarded the hope of returning to the Eden that never existed, their idealistic rhetoric remains in Activist 2.0—the Evangelicals, who now defend the status quo of the corporatized internet despite the fact that it allegedly destroyed the original garden in the first place.  The Evangelical is easy to spot.  She still clings to that original Barlowian sacrament of “sharing” content and responds to any proposal to protect copyright owners by declaring that [Insert policy here] will destroy the internet as we know it! 

Of course, the whole narrative is a lie—from Barlow’s catharsis to the present battle over the “soul” of the web.  As investigative reporter Yasha Levine states very pointedly…

…the truth is that EFF is a corporate front. It is America’s oldest and most influential internet business lobby—an organization that has played a pivotal role in shaping the commercial internet as we know it and, increasingly, hate it. That shitty internet we all inhabit today? That system dominated by giant monopolies, powered by for-profit surveillance and influence, and lacking any democratic oversight? EFF is directly responsible for bringing it into being.

Hence, the too-common refrain that we might “destroy the internet as we know it” is an odd rhetorical tactic insofar as it is not at all clear, from any point of view, why the internet we have is something worth preserving.  As a general observation, why is it rational to assume that the function of the internet, which has largely been ceded to the management of Google, Facebook, Twitter, et al, is exactly perfect as is and should never be changed?  By what measure, other than Big Tech’s profits, have we supposedly achieved our digital apotheosis?

Never mind the fact that protests against any type of copyright proposal invariably resort to hyperbole and disinformation (see claims that Article 13 will “kill memes”), but even if some new proposal were to change the internet, so what? As naive as I think the Barlow-worshipping purists were/are in the first place, we can at least all agree that their internet is not the internet we have, that the internet we have is dominated by big corporations and, therefore, hardly sacred.

That being the case, contemporary digital activists should drop the quasi-religious overtones when debating policy—stop talking about the internet as though it were holy ground that cannot be disturbed.  It is worth keeping in mind that every time the artists and creators have inveighed against their rights being trampled by the big internet platforms, the digerati have presumptuously lectured them that “change is good.”  Indeed it can be good.  And right now, what needs changing is the internet as we know it.