Vaccaro 2024 human-AI meta-analysis
A preregistered systematic review and meta-analysis of experiments comparing humans, AI, and their combination. Ryan clipped the roadmap passage into 2026-09-20-when-combinations-of-humans-and-ai-are-useful before this page existed.
- Paper: “When combinations of humans and AI are useful”
- Authors: michelle-vaccaro, abdullah-almaatouq, thomas-malone, MIT Center for Collective Intelligence, Sloan School of Management. Vaccaro also holds the Institute for Data, Systems, and Society affiliation at Schwarzman College of Computing.
- Venue: nature-human-behaviour, volume 8, pages 2293 to 2303, published open access 2024-10-28.
Method
106 experiments from 74 papers, 370 effect sizes, covering studies published 2020-01-01 to 2023-06-30. The search ran across the ACM Digital Library, the Web of Science, and the Association for Information Systems eLibrary. Inclusion demanded an original human-participants experiment reporting the performance of humans alone, AI alone, and the combination. The synthesis used a three-level meta-analytic model, and the reproduction materials are in an Open Science Framework repository.
Findings
- Synergy, measured against the better of human-alone or AI-alone: negative on average, Hedges’ g = −0.23 (95% CI −0.39 to −0.07). The combinations lost to the best single performer.
- Augmentation, measured against humans alone: positive, g = 0.64 (95% CI 0.53 to 0.74). The AI did help humans.
- Task type moderates. Decision tasks, choosing among a finite set of options, carried a significant loss (g = −0.27). Creation tasks, open-response content, carried a positive average (g = 0.19, not significant against zero on its own n = 34, but a significantly different result from decision tasks). About 85% of the effect sizes were decision tasks.
- Relative baseline moderates. When the human alone beat the AI alone, the combination gained synergy (g = 0.46). When the AI alone beat the human, the combination lost against the AI (g = −0.54). The authors hypothesize that the better party is also better at deciding when to trust the other.
- Explanations and confidence displays, the most-studied levers, did not affect the outcome. The authors suggest the field shift attention to task type and to which side is better alone. The division of labor deserves more study than it has received.
- Pre-divided subtasks appeared in only 3 experiments. Their pooled g = 0.22 was positive but not significant.
Limitations the authors state
Possible publication bias, which they consider unlikely to explain the negative headline result because the bias would favor positive findings. High heterogeneity (I² = 97.7%) with much of it unexplained. Variation in designs, participant pools, and measurement across studies. Laboratory configurations that may not match practical ones.
In this vault
The average loss for combinations is the empirical case for the division of labor in AGENTS: the agent does the fetching and filing, and Ryan decides. See human-ai-synergy for the definitions and moderators.
Sources
- Nature Human Behaviour, fetched and read 2026-09-21.