Abstract
Scientists and practitioners are aggressively moving to deploy digital twins— LLM-based models of real individuals—across social science and policy research. We conducted 19 pre-registered studies with 164 diverse outcomes (e.g., attitudes towards hiring algorithms, intention to share misinformation) and compared human responses to those of their digital twins (trained on each person’s previous answers to over 500 questions). We establish an empirical benchmark for digital twin performance: digital twins’ answers are only modestly more accurate than those from the (homogeneous) base LLM and correlate weakly with human responses (average r = 0.20). To guide future twin development, we further document five ways in which digital twins distort human behavior: (i) insufficient individuation, (ii) stereotyping, (iii) representation bias, (iv) ideological biases, (v) hyper-rationality. Finally, we make our full dataset and code public as a standardized testbed for novel digital twin methodologies. Together, our results caution against premature deployment while laying the groundwork for the transparent, replicable, and iterative science necessary for responsible deployment of twins.
One-sentence summary: Despite bold industry promises for digital twins, a large-scale mega-study finds they perform only marginally better than generic AI personas and documents five key distortions that serve as a diagnostic roadmap for improvement.