Published in the Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT '21), March 3–10, 2021, by Emily M. Bender (University of Washington), Timnit Gebru (Black in AI), Angelina McMillan-Major (University of Washington), and "Shmargaret Shmitchell" (The Aether) — the pseudonym adopted amid the Google dispute. DOI 10.1145/3442188.3445922. See Stochastic Parrots (Bender, Gebru, McMillan-Major, Mitchell, 2021) and the Octopus Test.
The question it asks
The paper's framing is a cost-benefit question posed to a research field, not a claim about what models cannot do: "we take a step back and ask: How big is too big? What are the possible risks associated with this technology and what paths are available for mitigating those risks?"
Its target is the specific competitive dynamic of 2018–2021 — BERT and variants, GPT-2, T-NLG, GPT-3, and Switch-C, "with institutions seemingly competing to produce ever larger LMs." The authors grant the value of the work directly: "investigating properties of LMs and how they change with size holds scientific interest, and large LMs have shown improvements on various tasks." The question is "whether enough thought has been put into the potential risks associated with developing them."
Environmental and financial cost
The figures cited are Strubell et al.'s: against an average human's estimated 5 tonnes of CO₂e per year, training a Transformer (big) model with neural architecture search "emitted 284t of CO₂," and training a single BERT base model without hyperparameter tuning was estimated "to require as much energy as a trans-American flight."
The argument the paper builds on this is distributional rather than absolute: "we must keep in mind how the risks and benefits are distributed, because they do not accrue to the same people." Its formulation is a question: "Is it fair or just to ask, for example, that the residents of the Maldives (likely to be underwater by 2100) or the 800,000 people in Sudan affected by drastic floods pay the environmental price of training and deploying ever larger English LMs, when similar large-scale models aren't being produced for Dhivehi or Sudanese Arabic?"
A qualification often lost in citation: the authors note that in industrial deployment "the cost of inference might greatly outweigh that of training in the long run," so "it may be more appropriate to deploy models with lower energy costs during inference even if their training costs are high."
Unfathomable training data
The section rebuts the inference from scale to representativeness: it is "easy to imagine that very large datasets, such as Common Crawl… must therefore be broadly representative of the ways in which different people view the world," but several factors narrow internet participation, and "in all cases, the voices of people most likely to hew to a hegemonic viewpoint are also more likely to be retained."
Three mechanisms are given. Harassment drives groups off platforms — the paper cites harassment on Twitter experienced by "a wide range of overlapping groups including domestic abuse victims, sex workers, trans people, queer people, immigrants, medical patients (by their providers), neurodivergent people, and visibly or vocally disabled people" — producing "a feedback loop that lessens the impact of data from underrepresented populations." Alternative fora are less likely to be crawled: anti-ageist discussion in older adults' blogging communities is "less likely to be found than other blogs that have more incoming and outgoing links." And filtering compounds it: GPT-3's training set was a Common Crawl filter trained to select documents resembling GPT-2's training data — documents linked from Reddit, plus Wikipedia and books — so the filter inherits Reddit's demographics.
The paper's broader point about auditing is that identifying bias presupposes knowing what to look for: work on bias "generally start[s] from US protected attributes such as race and gender (as understood within the US)," but "salient identity characteristics and expressions of bias are also culture-bound," so "we may miss marginalized identities if we don't know what to audit for." Verifying safety, it argues, "requires engaging with the systems of power that lead to the harmful outcomes," and any operationalization of shifting social norms into an algorithm "is necessarily political (whether or not developers choose the path of maintaining the status quo ante)."
Misdirected research effort
The third harm is opportunity cost in the research programme itself, following BERT's results on GLUE, SQuAD, and SWAG — "all datasets designed to test language understanding and/or commonsense reasoning." The cost is stated in two parts: "time not spent applying meaning capturing approaches to meaning sensitive tasks, and… time not spent exploring more effective ways of building technology with datasets of a size that can be carefully curated and available for a broader set of languages."
The stochastic parrot
The paper's most quoted passage defines the term:
"Contrary to how it may seem when we observe its output, an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot."
The supporting argument is about the reader, not the model. Human communication "takes place between individuals who share common ground and are mutually aware of that sharing," who "have communicative intents which they use language to convey, and who model each others' mental states." Even reading text from an unknown author, "we build a partial model of who they are and what common ground we think they share with us, and use this in interpreting their words."
So: "Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader's state of mind. It can't have been, because the training data never included sharing thoughts with a listener." The paper is careful that the illusion is located in the audience: "coherence is in fact in the eye of the beholder," and "if one side of the communication does not have meaning, then the comprehension of the implicit meaning is an illusion arising from our singular human understanding of language (independent of the model)." A footnote distinguishes controlled generation from communicative intent, so the claim is not that steering output is impossible.
The stated risks follow from that gap: "the mix of human biases and seemingly coherent language heightens the potential for automation bias, deliberate misuse, and amplification of a hegemonic worldview." Abusive output additionally risks "producing more (synthetic) abusive language that may be included in the next iteration of large-scale training data collection" — a feedback loop identified in 2021.
The section's summary places accountability at the centre: "the risks associated with synthetic but seemingly coherent text are deeply connected to the fact that such synthetic text can enter into conversations without any person or entity being accountable for it. This accountability both involves responsibility for truthfulness and is important in situating meaning." The paper quotes Maggie Nelson: "Words change depending on who speaks them; there is no cure."
Paths forward
The recommendations are procedural rather than prohibitive, and the paper is explicit that it is not calling for abandonment. It urges planning "before starting to build either datasets or systems trained on datasets," proposes pre-mortems — reverse-engineering hypothetical failures — to build "an evaluation culture that considers not only average-case performance… and best-case performance (cherry-picked examples), but also worst-case performance," and recommends value sensitive design methods including envisioning cards, value scenarios, and panels of experiential experts, applied "early in the development process… rather than as a post-hoc discovery of risks."
It also considers the strongest objection to itself — that backing off large models sacrifices benefits to marginalized populations — through automatic speech recognition and captioning for Deaf and hard-of-hearing users. Its answer has two parts: the largest LMs "typically are too large and too slow for the near real-time needs of ASR systems" anyway, and where large models are genuinely critical the situation is "an instance of a dual use problem," raising the question of whether models could be built "in such a way that synthetic text generated with them would be watermarked and thus detectable" and whether "policy approaches… could effectively regulate their use."
Conclusion
The costs the paper enumerates are "environmental costs (borne typically by those not benefiting from the resulting technology); financial costs, which in turn erect barriers to entry, limiting both who can contribute to this research area and which languages can benefit from the most advanced techniques; opportunity cost… and the risk of substantial harms, including stereotyping, denigration, increases in extremist ideology, and wrongful arrest, should humans encounter seemingly coherent LM output and take it for the words of some person or organization who has accountability for what is said."
Its closing claim is the one most consequential for policy: "we call on the field to recognize that applications that aim to believably mimic humans bring risk of extreme harms. Work on synthetic human behavior is a bright line in ethical AI development, where downstream effects need to be understood and modeled in order to block foreseeable harm to society and different social groups."
Relationships
- supports: Stochastic Parrots (Bender, Gebru, McMillan-Major, Mitchell, 2021) and the Octopus Test — the primary text behind the concept
- related: AI and Misinformation — the synthetic-text accountability argument
- related: Distributed AI Research Institute (DAIR) — the institute Gebru founded after the dispute over this paper
- related: Web Scraping for AI Training — the corpus-composition argument
- related: AI Environmental Impact, AI Bias and Discrimination, AI Benchmarks and Evaluation, Emily M. Bender, Timnit Gebru