Slop Code Bench

Definition: benchmark for iterative coding ability (gives model a greenfield build task, then asks it to add features to that same codebase without knowing the full problem upfront).

Origin

2026-10-04-youtube-ai-skeptic-sabbath: cited as experimental evidence that “best models fail catastrophically at production software engineering.”

Key finding

  • 15% success rate for best models (GPT-5.5 Codex, tested May 2024) on iterative feature addition
  • Quality degrades over time as models continue working on same codebase
  • Failed test count trends upward across all models

Significance

Demonstrates the gap between “one-shot automation awe” (greenfield success) and real software engineering (iterative, discovery-driven, inductive, pattern discrimination in massive codebases). The speaker argues this is “not just emotional opinion” but “experimentally proven.”

Contrast

Greenfield benchmarks (HumanEval, MBPP) measure different capability. Slop Code Bench measures the iterative, context-heavy work that is actual production engineering.

  • ai-psychosis: the false confidence from greenfield-only evaluation
  • skill-atrophy-ai: why humans must maintain the skills models lack
  • human-ai-synergy: meta-analysis showing combinations beat humans alone but lose to best single performer