Documentation debt
Definition (Bender et al. 2021, section 4.4, as applied to data): the situation where datasets are both undocumented and too large to document post hoc. The paper adapts a term from the software literature, where documentation debt names code that outran its docs.
The debt accrues silently. A team ingests web text at scale because scale correlates with benchmark scores, and the cost of describing what actually went in is deferred. The deferral compounds: by the time anyone needs the record, the corpus is unfathomable, in the paper’s word, and reconstruction is impossible at any reasonable budget.
The accountability chain
Documentation is the substrate of accountability. The paper’s chain: audits of a model’s bias presuppose knowledge of its training distribution, and remedies for harm, like appeals against a discriminatory classification, require a describable basis. Undocumented training data perpetuates harm without recourse, which is the paper’s own phrase.
GPT-3 is the example sitting closest at hand in 2021. The model card noted skewed participation in the source data, the paper reports, yet the filtered Common Crawl subset was never fully documented, so what the filtering removed besides “unintelligible” text is unmeasured and unknown. The GPT-3 training set itself is not openly available. Auditors had to reconstruct from GPT-2’s data instead.
The prescription
Budget for curation and documentation at the start of a project, and collect only as much data as that budget can thoroughly document. The size of the dataset is capped by the labor available to describe it, not by storage. The paper points at archival history as the model for how much resourcing real collection takes, citing Jo and Gebru 2020, and calls the expectation that raw ingestion of the world’s text yields neutral input a fantasy, quoting Birhane and Prabhu via Ruha Benjamin: feeding AI systems on the world’s beauty, ugliness, and cruelty, but expecting it to reflect only the beauty is a fantasy.
The documentation frameworks the same authors built operationalize this: data statements for NLP (Bender and Friedman 2018), datasheets for datasets (Gebru et al. 2018/2020), model cards for model reporting (Mitchell et al. 2019).
Related
- stochastic-parrot: the sibling concept from the same paper.
- value-lock: what static undocumented data does to social meaning over time.
- bender-2021-stochastic-parrots section 4 for the full data critique.
Sources
- bender-2021-stochastic-parrots sections 4.1 and 4.4.