Entropy collapse (LLM training)
Definition here: the reduction of a model’s output or solution diversity caused by the training pipeline and the prompting scaffold, independent of any synthetic-data recursion. Not to be confused with posterior collapse in VAEs or neural collapse in classifier geometry, which are different literatures that share the word. This is mechanism C in llm-degradation-toward-the-mean.
Channels
- Supervised fine-tuning. Cross-entropy SFT reduces output diversity; game-theoretic entropy maximization is the proposed fix (GEM, arXiv:2408.16673, ICLR 2025).
- Preference optimization (RLHF, DPO). The KL regularizer against the reference policy sharpens the distribution. Consensus statement in Diverse Preference Learning (arXiv:2511.08594): aligned models “generate text with repetitive structure and word choice,” “approach problems in more uniform ways,” and reflect “a narrower range of societal perspectives.” DivPO (arXiv:2501.18101) covers the same sharpening across RLHF, DPO, and SFT.
- RL with verifiable rewards. cui-2025-entropy-mechanism: entropy drops sharply and early across RL runs without intervention, the mechanism is positive covariance between action probability and logit change under advantage weighting, and the empirical law R = -a·e^H + b prices diversity in accuracy units. The sharpest symptom, from “The Choice of Divergence” (arXiv:2509.07430): Pass@k falls while Pass@1 rises. One-attempt accuracy improves by spending the spread of possible solutions. DAPO’s Clip-Higher (arXiv:2503.14476), CE-GPPO (arXiv:2509.20712), and SCOPE-RL (arXiv:2510.08141) are the mitigation line; Cui’s own are Clip-Cov and KL-Cov.
- Inference-time format. yun-2025-price-of-format: chat-template structural tokens act as behavioral anchors and collapse open-ended diversity even at high temperature. Temperature does not fix it; structure removal does.
Reading
When a user says “models all sound the same,” the honest attributions, in order of evidence strength: pipeline entropy collapse (this page), template effects, then the recursion story (model-collapse), which needs the beta = gamma = 0 premise to be even the main character.
The recursion side has its own formal entropy proof now, so the split can be stated precisely. Zenil v2 Theorem 2 (zenil-2026-self-improvement-limits, full-text read 2026-10-04) makes H(Qₜ) a supermartingale under closed-loop retraining on finite samples: entropy decays because resampling drops the tails. The channels above are pipeline effects (post-training and scaffolding); none of them needs synthetic-data recursion. Same symptom word, two mechanisms, different fixes.
Sources
Per-paper links on cui-2025-entropy-mechanism and yun-2025-price-of-format; homogenization worker note
under .pi-subagents/artifacts/outputs/22bbe52f-7ffe-46a9-80fd-4e911e9fa29f/notes/.