Key Takeaway: Researchers from Columbia Business School built the public blueprint that lets anyone create digital twins, then ran the most rigorous test of them to date. As stand-ins for human survey respondents, the twins came only modestly close to real people's answers. But in a real-world field experiment with a global bank, twins built from the right data lifted marketing results. Predicting human behavior is genuinely hard, and digital twins only pay off when they're fed data suited to the job.
Startups that build AI models of real people have raised hundreds of millions of dollars, and some are already worth near or above $1 billion, on the promise that a simulated customer can stand in for a real one: a digital twin. These sorts of AI models are fed enough information about a person so that they can, theoretically, answer questions the way that person would. Ask the model whether it would switch banks, buy a gadget, or click an ad, and it offers a reply..
Olivier Toubia, Glaubinger Professor of Business at Columbia Business School, and a team of CBS faculty have spent the past year putting that promise to the test in more than 35 different studies across the School’s Digital Twins Lab. Their findings cut two ways. First, digital twins fall short when built on generic data, they found. But, second, digital twins can deliver real value when given rich, relevant data.
Not just another focus group
Prior to the advent of generative AI, researchers who couldn't easily observe large groups of people relied on convenient stand-ins, like students in a psychology class, volunteers in a lab, or an online panel. While these groups can serve as useful proxies, they often cannot match the real thing.
Toubia puts digital twins in the same company. They give you usable input, he notes, but remain "only the approximation of the real thing."
To build the twins, the CBS team first surveyed more than 2,000 people and asked respondents over 500 questions each, covering personality, preferences, and reasoning. They then released the entire dataset as a blueprint anyone could use to build twins of their own.
‘It's just very hard to predict’
Across 19 studies and 164 different outcomes, the researchers compared what the twins “said” to what real people actually answered. A twin loaded with 500 answers about a person barely outperformed a bare-bones profile built from nothing but demographics like age and income. The twins captured the broad direction of a person's answers, according to Toubia, but rarely landed on the specifics. They provided more of an educated guess than a genuine read.
While that might sound like a verdict against digital twins, the reality is more complex. The team ran the same challenge a different way, building a traditional statistical model with no AI at all. They trained it on 500 real people and then used it to predict the rest. It, too, hit a low ceiling.
"It's not really an issue of digital twins," Toubia says. "It's just an issue of trying to predict something that is just very hard to predict."
‘No magic here’
Change the data and the task, and the picture of digital twins changes too. In a separate field study with the management consulting firm Oliver Wyman, Toubia and colleagues worked with a global bank to build twins of roughly 330,000 real customers. Instead of using surveys, the team built the twins using customers' actual accounts and spending history. Crucially, the twins sat on top of the bank's existing prediction models rather than replacing them.
The twins' objective was to sharpen the bank's marketing — how credit card offers should be framed and presented to customers, and which customers should hear about a particular credit card. Here, the twins had more success.
While the bank's own models still decided which customers were worth contacting, the twins tackled the narrower task of approaching each one, ranking which offer and which tone would most likely land. Against the bank's usual one-size-fits-all messaging, that personalization lifted credit card sign-ups by 6 percent and payroll transfers by 25 percent in an early field test. The extra sign-ups also skewed toward higher-value customers — though the credit card result was only marginally significant, and the personalized group saw more message variety than the comparison group, so some of the lift may owe to variety rather than the twins alone.
"There's no magic here," Toubia says. To predict whether someone takes a credit card offer, you feed the twin data about that behavior, like past offers and spending patterns, instead of a personality quiz. The model, Toubia notes, is only as good as the data you put into it.
The bottom line: Trust, but verify
Toubia and his team point to the same practical rules when working with any sort of digital twins application: Results depend on using relevant data, setting honest benchmarks, and validating findings before you accept them. When a vendor selling digital twins advertises 80 or 90 percent accuracy, your first question should be, "Compared to what?" A number means little without knowing what a simpler approach would have achieved, or how the model errs in the cases it gets wrong.
Toubia also draws a sharp line between a back test, which just replays the past, and a genuine forward test that predicts a behavior before it happens. Only the latter, he argues, deserves your confidence.
To make that easier, CBS has launched a free, nonprofit platform, ExploraTwin, where anyone can field a survey of hundreds of twins and compare the results against real data. The most promising uses, according to the School’s researchers, may lie beyond swapping twins in for survey takers altogether — gathering richer open-ended feedback, asking questions you couldn't ethically put to people, or building AI agents that act on a customer's behalf.
Open source by design
Seventeen CBS faculty members across departments as varied as marketing, management, operations, and finance assessed how twins performed: Tianyi Peng, Melanie Brucks, George Gui, Malek Ben Sliman, Eric J. Johnson, Silvia Bellezza, Dante Donati, Hortense Fong, Elizabeth Friedman, Mohamed Hussein, Kinshuk Jerath, Bruce Kogut, Kristen Lane, Hannah Li, Vicki Morwitz, Oded Netzer, and Toubia.
Toubia credits the School for making that collaboration possible through funding for an ambitious slate of studies, the computing and research support behind them, and a culture where colleagues pull each other into new work over lunch. This culture shows, he says, in the open way the team both assembles and shares its findings, and in the strength of the findings themselves.