Bender 2021 stochastic parrots

A position paper asking how big is too big, written when the size race was still young. It names the central concept of this cluster, the stochastic-parrot, and introduces documentation-debt and value-lock. Its warning about synthetic text re-entering training data, in section 6.2, predates the proof literature behind model-collapse by three years.

  • Paper: “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜”
  • Authors: emily-bender and timnit-gebru as joint first authors, plus angelina-mcmillan-major and shmargaret-shmitchell. Affiliations per the byline: University of Washington, Black in AI, University of Washington, The Aether.
  • Venue: facc-t ‘21, March 3 to 10 2021, virtual event, Canada. Pages 610 to 623, DOI 10.1145/3442188.3445922, licensed CC BY 4.0 per the ACM record.

The state of the art it describes

Section 2 defines a language model as a system trained on string prediction tasks, then Table 1 charts the scale race: BERT at 340M parameters on 16GB of data, GPT-3 at 175B on 570GB, GShard at 600B, Switch-C at 1.57T on 745GB. The paper notes 7% of GPT-3’s training data was non-English, and cites work finding over 90% of the world’s languages used by more than a billion people have little to no language technology support.

The four risk families

  1. Environmental and financial cost, section 3. One Transformer (big) model trained with neural architecture search emitted an estimated 284 tonnes of CO2e, against roughly 5 tonnes per person per year. Training one BERT base model used about the energy of a trans-American flight. Compute for the largest deep learning models grew 300,000-fold in six years, outpacing Moore’s Law. The paper’s equity argument: climate harms land on marginalized communities first while the models being trained are almost all English-language.
  2. Unfathomable training data, section 4. Internet participation is skewed. Reddit users in the US were 67% men and 64% aged 18 to 29 per a 2016 Pew survey the paper cites; Wikipedians are 8.8 to 15% women or girls. Moderation practices push marginalized voices off mainstream platforms, and the filters on the remaining data cut further: the Colossal Clean Crawled Corpus discards any page containing one of roughly 400 “Dirty, Naughty, Obscene or Otherwise Bad Words”, which suppresses LGBTQ online spaces along with porn. Hegemonic viewpoints survive at every step. Size does not buy diversity.
  3. Misdirected research effort, section 5, titled Down the garden path. Following Bender and Koller 2020, languages are systems of signs pairing form with meaning, and LM training data contains only form. No training objective gives the system access to meaning, so benchmark success on tasks solvable by form manipulation is evidence about form manipulation. Around 26% of a sample of ACL, NAACL and EMNLP papers since 2018 cite BERT, and test manipulation shows the systems ride spurious cues. Karen Spärck Jones 2004 supplies the epistemological framing.
  4. Risks of seeming coherence, section 6. Coherence is in the eye of the beholder: readers model an interlocutor with beliefs and intent because human communication is co-constructed, and nothing dismisses that reflex when the text has no author behind it. Harms catalogued: reproducing and amplifying training-data bias, psychological harm to stereotype subjects, allocational and reputational harm when LMs or their embeddings serve classification and query expansion, cheap synthetic text populating extremist recruitment boards, machine translation errors attributed to the source author, including the arrest of a Palestinian man whose Arabic “good morning” became “attack them”, and personally identifiable information extraction that grows more effective with model size. Synthetic text enters conversations with no person accountable for its truth.

The recommendations, section 7

Weigh environmental and financial cost before starting a project. Budget for curation and documentation and collect only as much data as that budget can document. Adopt the documentation frameworks the group helped build: data statements, datasheets for datasets, model cards. Run pre-mortems and value sensitive design exercises early, not post hoc. Treat synthetic human behavior as a bright line in ethical AI development. For genuinely useful cases like automatic captioning, search for paths that do not require ever-larger models, or treat the large model as a dual-use tool and work on watermarking and policy.

In this vault

Section 6.2 is the pre-proof: LM outputs reinforcing and propagating stereotypes into “future LMs trained on training sets that ingested the previous generation LM’s output”. shumailov-2024-model-collapse names and proves that recursion, gerstgrasser-2024-accumulating-data answers the scale critique with accumulation bookkeeping, and padmakumar-2024-content-diversity measures the corpus-level drift the paper predicted. The form-without-meaning argument also describes phantom-citations: a system that stitches plausible references together has no more access to truth than one that stitches plausible sentences.

Sources

  • ACM DOI record, Crossref metadata read 2026-10-04 for title, authors, pages, license.
  • Camera-ready PDF read in full 2026-10-04. The publisher page returned 403 to direct fetch, so the PDF was taken from the ACM mirror https://s10251.pcdn.co/pdf/2021-bender-parrots.pdf and filed at Inbox/assets/2026-10-04-bender-2021-stochastic-parrots.pdf.