Abstract
Synthetic data, and in particular digital twins, have captured the interest and imagination of practitioners and academics across disciplines. Indeed, the ability to simulate human behavior and responses in never-seen-before scenarios and questions, using both structured and unstructured data as both input and output, opens powerful possibilities. But a review of the empirical evidence available to date suggests that despite some encouraging results, synthetic data currently do not predict well overall, and systematically distort, human behavior. Many research opportunities exist in improving the accuracy and usefulness of synthetic data and developing novel use cases for digital twins. But researchers and practitioners also need to properly calibrate their expectations of how predictable human behavior is, i.e., on the maximum predictivity achievable by synthetic data.