Beta The Briev beta is out. Free on iPhone via TestFlight — install it in under a minute.

Join the beta ↗
Briev
Live
Technology

AI digital twins fall short in mimicking human behavior for social research

A new study finds that AI-generated digital twins often misrepresent individual responses, limiting their usefulness as substitutes for human participants in behavioral experiments.

In a September 2 paper in Science Advances, a team led by Olivier Toubia built AI “digital twins” by inputting detailed survey data from more than 2,000 Americans into a large language model. The twins were tested in 19 social-science experiments, ranging from reactions to political donations to attitudes toward algorithmic hiring. While the twins outperformed chance and matched the variability of real respondents better than chatbots that only knew demographics, they still erred on average about 25% of the time and produced overly uniform, stereotype-aligned answers.

Their performance improved with participants who were wealthier and more educated, and they showed systematic biases such as greater trust in others and less worry about technological risks. Hadi Hosseini noted that the models tend to rationalize human judgment, and suggested that continuous, conversational training could reduce distortions. Toubia sees limited but valuable roles for such twins, like pre-testing surveys or generating lengthy responses when human participants are fatigued, but stresses the need for realistic expectations about synthetic data’s capabilities.

Why it matters

Understanding AI twins' limits helps researchers gauge when synthetic participants can aid studies without compromising data quality.

In this story

AI digital twinsbehavioral researchlarge language modelsurvey biassynthetic datasocial science experiments
Get the beta ↗