Comparing AI Prompts to Button-Pushing on a Camera

Plenty is being said about AI systems that generate visual works, written works, music, etc. And plenty more will be said, especially now that lawsuits have been filed against some of the AI-generated image companies. In this post, I want to address a misconception about authorship in copyright law that may be warping the AI conversation. As I understand the argument, some AI proponents allege that the act of writing prompts is comparable to the act of pushing the button on a camera and, therefore, vests copyright rights in the proverbial “button pusher.”

Although it is possible to conceive a scenario in which this analogy might apply, it is important to first understand that the underlying premise (i.e., that button pushing establishes authorship in a photograph) is wrong. In fact, when photography emerged as the first machine-made work, it posed a challenge to copyright law that still provides an ideal context for discussing what it means to say that copyright protects creative expression the moment the author causes that expression to be fixed in a tangible medium. Note that the key ingredients are expression, an author, and fixation, and inherent to the process binding all three is an interval of human effort enabling the author’s concept (or vision) of the expression to be manifest as fixation.

With photography, the interval of effort may be stately or a mere fraction of a second, but copyright law does not discriminate between the photographer who carries a vision in her mind for weeks of preparation and arrangement and the photographer who captures a fleeting moment from real life. In both cases, triggering the shutter is the proximate cause of fixation,[1] but vesting copyright rights in the photographer is predicated on an assumption that, even in a fraction of a second, she made creative choices sufficient to find a modicum of original expression in the image.

Various Scenarios in Which It Is Not About the Button

In the case of a studio shoot with a lot of preparation, lighting, props, wardrobe, etc., the photographer may not even touch the camera very often. It may be mounted on a tripod with an assistant triggering the shutter from a computer or remote control while the photographer directs all the creative aspects that comprise the resulting images. Copyright holds unequivocally that this individual is the author of the photographs because it is his expression that is being fixed in each image, but the mechanical “button-pushing” is irrelevant except as a purely mechanical step in fixation.[2]

For the street photographer or photojournalist, the same principles apply, but copyright allows for the arguably metaphysical assumption that even in the tiny interval between seeing the real-life subject and capturing it, the photographer makes subtle choices that imbue the work with sufficient expression to be protected. Again, the button causes fixation but is not the basis of authorship, and this would be evident in the analysis of the content and qualities of the photograph, if it were to become the subject of a copyright infringement lawsuit.

By contrast, if a truly accidental photograph is captured (e.g., by a camera accidentally dropped from the Eiffel Tower), there is no authorship in that image—not because a human did not push the button, but because there is no colorable nexus between the human’s mental conception and the resulting photograph. On the other hand, if a photographer intentionally drops a camera from the Eiffel Tower and triggers the shutter by remote on its way down, copyright attaches to those images—not because a human pushed the button, but because a human conceived of the series of falling photographs and arranged the circumstances by which they could be made.

Although it is important to note that cameras are not machines trained with a corpus of existing photographs, this last example may be the closest analogy to the prompt directing the AI generator (in its current state) to make an image. If the prompt writer has a general sense of the image she wants to produce, but there is still an element of chance about what the machine will make, the prompt writer may argue that she is no less an author than the photographer who intentionally allows some element of chance into the process of making his images.

While this premise sounds reasonable as a general proposition, what it really implies is a case-by-case consideration as to how much human expression exists in the resulting works. Even in the example of the camera tossed intentionally off the Eiffel Tower, the photographer can control certain qualities in the images and may even have a vision for how they are to be used, displayed, or distributed. He knows the characteristics of the camera and lens and can select settings with the intent to control some of the qualitative results in the final photos.

By contrast, the prompter directing the image-generating AI is arguably not in control of enough of the qualitative elements in the final image to claim authorship—at least not at the current state of the technology. Entering the prompt “A mermaid wrestling a sea lion in outer space in the style of Cartier-Bresson” may produce an image that checks each of those boxes, but the prompt writer is not controlling the qualitative choices that comprise the result. Composition, line weight, shading, lighting, texture, scale, proportion, etc. are all “selected” by the AI based on what it has “learned” from the millions of visual works fed into its code, so there is a critical disconnect between the human’s vision of “A mermaid wresting a sea lion in outer space in the style of Cartier-Bresson” and the interval of effort that fixes the image in a tangible medium.

At some future state of the technology, the human may prompt a draft image to be made and then prompt changes to the qualitative elements, at which point it may be tough to deny that there is authorship in the resulting work. If these technologies develop in this way—such that the prompter is essentially painting with words instead of a stylus—this anticipates that, for instance, a disabled individual could truly create visual works with her mind akin to the way Stephen Hawking wrote books. But in this paradigm, the AI does not present a unique challenge to the concept of authorship because the human is in control of sufficient expression in the work.

Dynamic Ethical Standards

Of course, this theoretical discussion assumes integrity among individuals who claim authorship in various works. The guy whose camera accidentally snaps a photo does not have to admit he played no role in its making, and AI currently presents a similar challenge. The issue of integrity is a hot conversation we’re having in response to generative AI—especially in academia where ChatGPT is already “writing” papers for students. Notably, few people would question the judgment that the student who turns in a paper “written” by an AI is a cheat deserving the same sanctions as if he were caught plagiarizing. Yet, somehow, when the material is a “creative” work, AI advocates argue that the prompter is an author of a visual work comparable to a photographer using a camera.

This dichotomy can only be reconciled by confronting the fact that certain uses of AIs are not only not authorship but are needlessly destructive to the very purpose of intellectual and cultural endeavor. The student who shirks writing his own paper learns nothing and so, potentially graduates from a program unqualified. Likewise, the prompter using an image-generating AI is not an artist and contributes nothing to the purpose of art. Thus, while there may be uses for these systems, their potential cultural value depends on more than technological development for its own sake.

Because these technologies are still new and still primitive relative to their expected capabilities, it is hard to predict where the more serious aspects of the narrative will lead. Some of the generative AIs are barely more than toys at the moment (e.g., turning profile pics into oil paintings), but what they will do a year from now, let alone five years, will inform how we address the issues—cultural, legal, and ethical. For now, though, I insist that no, prompting is not equivalent to button-pushing with a camera, even if button-pushing were as significant as many people think it is.


[1] This is true with digital photography. With film, one could argue that the latent image on the negative is not fixation until it is at least developed because it cannot be perceived by either human or machine reader.

[2] And there are likely to be further steps like retouching or printing, which may fix the final version of the image.

Photo by author.

Sen. Cruz Brief Wrongly Portrays Section 230 as a Neutrality Law

Among the briefs filed in Gonzalez v. Google asking the Supreme Court to properly read Section 230 of the Communications Decency Act is one filed by Sen. Ted Cruz, Rep. Mike Johnson, and fifteen other Republican Members of Congress. Presenting similar textual arguments as the brief filed by Cyber Civil Rights Initiative (CCRI), highlighted here in a recent post, Sen. Cruz et al. petition the Court to address a matter that has nothing to do with Section 230—a politically motivated complaint summarized as follows:

Confident in their ability to dodge liability, platforms have not been shy about restricting access and removing content based on the politics of the speaker, an issue that has persistently arisen as Big Tech companies censor and remove content espousing conservative political views, despite the lack of immunity for such actions in the text of §230.

Allegations of viewpoint bias are inaptly raised in Gonzalez, or indeed any case addressing Section 230. Even if it could be shown that a social platform actively engages in true political bias (i.e., moderating ideas and speakers rather than extremism), this is not a Section 230 issue because neither political nor any other form of bias necessarily implicates civil liability for an online platform any more than it does for a newspaper or TV network.

The First Amendment protects bias, and Section 230 does not alter this fact. Hence, the Cruz brief strays far from the purpose of the Court’s review in Gonzalez by erroneously implying that bias is inherently grounds for litigation when it alleges that the overbroad interpretation of Section 230 immunity causes or sustains politically motivated censorship. But Section 230 is not and never has been a viewpoint neutrality law. Cruz et al. are asking the Court for a misreading that has lived in the PR of the platforms and the rhetoric of tech-utopianism, but is nowhere in the statute.

Specifically, the Cruz brief alleges that the platforms have been shielded in censorious conduct by a poor statutory reading of their right to “good faith” removal of material that is “otherwise objectionable.” The amici argue that those words must be read in balance with the preceding words in the statute providing immunity where platforms remove or restrict access to third-party content that is “obscene, lewd, lascivious, filthy, excessively violent, harassing, or otherwise objectionable.” The brief then asserts (a bit wryly) that “…conservative viewpoints on social and political matters do not rise to the level of being ‘obscene, lewd, lascivious, filthy, excessively violent, harassing, or otherwise objectionable.’”

Notably, the brief both elides a definition of “conservative” in context to its argument and asks the Court to read Section 230 as a mandate that platforms leave all material online that does not meet a very narrow, statutory definition of “objectionable.” This is false. Section 230 was written to encourage platforms to adopt and enforce their own community standards (i.e., decide what is objectionable), which does not disturb the general right to host a platform which may be politically biased in any direction. A proper reading of 230 simply means that platforms shall not be unconditionally immunized against potential liability for hosting content that results in some form of harm which may be remedied through civil litigation.

The Cruz brief does not distinguish between amici’s political bona fides and the broad spectrum of hate-speech and violence-inciting material that some Americans now call “conservative,” and which is indeed problematic for platforms. For instance, as the Alex Jones verdicts or the January 6th convictions make clear, material that certain people are willing to label “conservative” may, as a matter of law, be libel, defamation, or disinformation that results in individuals being harassed and threatened, or which leads to violence—or even an insurrection. And it is precisely in this context (i.e., blurring the line between political views and actionable conduct) that the complaint in the Cruz brief is so inaptly raised in Gonzalez.

Petitioner Gonzalez alleges that Google’s “recommendation” algorithms contributed to fostering terrorist activity by promoting ISIS recruiting videos in a manner that predictably roused a latent terrorist, who then acted on those emotional triggers. Regardless of whether that complaint prevails in context to the anti-terrorism statutes at issue, the general allegation about the platform’s role in Gonzalez is indistinguishable from an algorithm detecting that a user likes InfoWars and, therefore, “recommends” QAnon videos or some other tinfoil-hat material with the foreseeable result that some domestic terrorist will assault a family in Sandy Hook, ransack the Capitol, conspire to kidnap a sitting governor, etc.

Thus, if the Court agrees that Google is not shielded from litigation in Gonzalez, the allegations of liability presented, even if they do not prevail in that instance, are directly analogous to an Alex Jones or an election-denier issue for a platform moderation team. Even if Google is ultimately not found to be liable for the ISIS-related killing of Nohemi Gonzalez, allowing the case to proceed past the Section 230 veil will demonstrate that there is a plausible, common-sense nexus between amplification of certain material and harmful conduct.

Under a correct reading of 230 (i.e., no unconditional immunity for platforms), the platforms may be more effective in addressing inciting material—a goal that should have bipartisan support from lawmakers interested in both a proper reading of the statute and the general welfare of the nation. Unfortunately, this political monkey wrench in the Section 230 issue is part of a broader narrative in which social platforms have allegedly tilted the scales—but in favor of extreme right-wing material calling itself “conservative.” For instance, in February 2021, BuzzFeed reported:

Internal documents obtained by BuzzFeed News and interviews with 14 current and former employees show how the company’s policy team — guided by [Republican lobbyist and conspiracy promoter] Joel Kaplan, the vice president of global public policy, and Zuckerberg’s whims — has exerted outsize influence while obstructing content moderation decisions, stymieing product rollouts, and intervening on behalf of popular conservative figures who have violated Facebook’s rules.

This and other reports, including testimony before Congress, reveal a pattern of (if anything) pro right-wing bias at Facebook and other platforms, including evidence that the “anti-conservative” story itself is a fiction promoted by individuals like Kaplan. More specifically, Zuckerberg’s apparent resistance to remove Alex Jones from the platform demonstrates how the chronic misreading of Section 230 would only benefit a Trumpianized GOP that embraces every extremist willing to wear a red hat.

A correct reading of 230 opens the possibility that Facebook could be liable for hosting or “recommending” InfoWars, while an incorrect reading forecloses that possibility at summary judgment. Only one of these interpretations benefits those elements of the GOP who choose to align themselves more closely with that brand of “conservatism.” Thus, consistent with the Trumpian tactic of weaponizing alleged victimhood, the comparatively mild complaint of viewpoint bias in the Cruz brief is political theater—blaming social media platforms for actions that a) they have not taken; b) they have a constitutional right to take, if they want to; c) are unrelated to Section 230 immunity; and d) detract from an important legal question for real victims barred from pursuing relief by misreading the statute.

Cruz and his fellow amici have heard or read testimony from witnesses like whistleblower Frances Haugen, who explained to the Senate Commerce Committee how Facebook consistently put profits ahead of safety, adding, “The result has been a system that amplifies division, extremism, and polarization — and undermining societies around the world. In some cases, this dangerous online talk has led to actual violence that harms and even kills people.” Specifically, Haugen and other former insiders have repeated the theme that extremism has been good for social platforms—that angry users are active users, and active users translates to profit for these companies.

A proper reading of Section 230 will not solve every problem fostered by social platforms, but it can have the effect of forcing platform operators to identify when speech is reasonably linked to harmful conduct and to acknowledge a nexus between addictive algorithm design and illegal activity—from terrorism to “revenge porn.” Very real harms have been shielded and exacerbated by misreading Section 230, and it is this error of law which the Court should resolve. In the process, it should decline to address the subject of viewpoint neutrality as the inappropriate, political side show it is.

AI “Art” is Boring

Adam was bored alone; then Adam and Eve were bored together; then Adam and Eve and Cain and Abel were bored en famille; then the population of the world increased, and the peoples were bored en masse. To divert themselves they conceived the idea of constructing a tower high enough to reach the heavens. This idea is itself as boring as the tower was high, and constitutes a terrible proof of how boredom gained the upper hand. – Soren Kierkegaard (1843) –

I had not thought about Kierkegaard writing on the subject of boredom in years. The essay from which the above quote is extracted was a favorite in college for its biting humor, but something about Rogers Brubaker’s excellent article about democratizing culture sent me in search of my 38-year-old (ouch) copy of The Kierkegaard Anthology, and I think it was this paragraph of Brubaker’s which triggered the thought:

But the question is not just how many people engage in cultural production — it’s how people engage. The AI music company Amper promises to help customers “create your own original music in seconds.” The creativity involved is rather attenuated, amounting to editing and tweaking the music generated by the AI, but that didn’t stop Amper co-founder Drew Silverstein from evangelizing in a TED talk about how AI can “democratize music” by enabling “anyone to express their creativity through music.” 

That promise to “create your own original music in seconds” was the portkey back to Kierkegaard. “In the case of children, the ruinous character of boredom is universally acknowledged,” he writes, and, indeed, I maintain that boredom is the inevitable outcome of AI toys promising to make music, visual art, poetry, etc. We have all experienced as children and witnessed as adults that transition between playing with a new toy and rapid disenchantment because the toy fails to engage the imagination. I am not the only Gen-X parent, for instance, to notice that when LEGO began selling kits to build branded objects like Star Wars spaceships, my own children would usually complete the assembly once and then be done with the toy forever. By contrast, my contemporaries and I spent hours with sets composed of bricks and no predetermined design.

Kierkegaard proposes that the plebian bores others and amuses himself while the aristocrat amuses others and bores himself—a dialectic perhaps well suited to describe the inevitable use of AI machines to “make one’s own music or art.” At the current state of the technology, the input of the human user is barely creative—little more than dropping a coin in a jukebox—and thus, all users similarly situated are plebian bores for the time being. The works resulting from their prompts may amuse them (for a while), but they will mostly bore others who will only be interested in “making their own music” with the same toys. Before long, a million individual users of the music generating AI will achieve a collective homeostatic boredom—a two-dimensional Babel leading nowhere.

Perhaps one of these accidental works will reach escape velocity, break through the gravitational force of mass boredom and “go viral” for a fleeting period. Some AI-generated ditty might be next year’s “Baby Shark” or even share the apotheotic luminance of a “Gagnam Style.” Someone will choreograph a short dance to accompany the tune, and TikTokers will fall in line to perform their versions, and Big Tech will look down and see that it is good, and their disciples will proclaim, “Behold the new culture! The human songwriter is an anachronism.” And it will all be as boring as it is ephemeral.

It is possible, of course, that generative AIs will become sophisticated enough to be collaborative tools wielded by the human artists—that the human still selects and arranges the creative elements to achieve her vision while the AI “helps” in some way. If and when we get there, we shall see. But in the meantime, it is clear that AIs do not need to be more sophisticated to replace some creative human work right now. My good friend Marco North writes on Facebook to me, “A full roster of AI voice talent costs less than $100 a month, works 24/7 and [will] do endless revisions….Voice work is perfect gig work for actors, say goodbye to lots of that.”

A gifted polymath in film, photography, music, poetry, and prose—Marco writes a weekly blog called Impressions of an Expat. Initially written from Moscow, he now writes from Tblisi, and in his latest post, he describes a happenstance encounter with the statue of Georgian poet Vazha-Pshavela (Luka Razikashvili) and his feelings about AI “art.” He asks:

Who will be the subject of the next statue? An algorithm? Will there be streets named after TikTok? Will we name a playground after a Spotify playlist curator? These are the people that tell our stories now. Midjourney highway will take you there. Take a left at ChatGPT square, you can’t miss it.

Yes. That is a vision of a possible future. Of course, if the tech giants can make the world just boring enough, then certain humans will do what certain humans do. They will disassemble the unengaging toy and turn it into something else—something called art. And then, the world will start to be interesting again.