Human testing of AI products is slow and expensive. This paper builds the other option at planetary scale: a simulated population you can put in front of a survey, a chatbot, a website, or an app before any real person sees it. Organized by Xiaomin Li and Yuexing Hao, with contributors and advisors across dozens of universities and labs.
Every record answers 1,290 categorical questions, grouped into five areas. Lifestyle is the biggest slice, capability close behind.
grounded in UN population data, World Bank, ILOSTAT and public surveys
400,000 released personas are drawn from a directed graph over all 1,290 dimensions. Each attribute is sampled conditional on its parents, so correlated traits stay correlated. A persona whose primary language is English cannot have an English proficiency of none.
599,847 released personas come from six real-world sources, mapped into the same schema.
Persona 8B mostly stays behind the curtain. The public coreset holds 999,847 records, contradiction-checked, deduplicated, and calibrated to published demographics for age, region, gender and urbanicity.
Humans rated a 100-persona sample of the extraction quality. Then the models were scored on how often they landed within one point of the human mean.
agreement with the human mean, share of comparisons within one point
The MatrAIx Playground runs the same persona through four kinds of product surface. The 1,010 tasks skew heavily toward the two cheap environments; native web and app runs stay small because every trial is expensive.
completes questionnaires, records answers and rationales
converses with a bot, records the dialogue, tool calls, resolution
browses a site, records pages, actions, screenshots, submission
operates Linux, macOS and iOS apps, records final state and file changes
1,010 tasks across more than 25 domains
the other bucket covers 20+ domains: travel, legal, insurance, education, entertainment, food, real estate, games
The adherence study: ten behavioral attributes, four environments, five personas declaring each pole. 400 trials. Each dot below is one trial; a coral dot means the declared behavior was expressed, or correctly suppressed.
the App environment is the hard one: only 6 of 10 behaviors held on both poles
On the OpenBB task, a persona's declared trust level separated subgroups under all three agent models, every effect significant at q < 10−8, and all three models put the four trust groups in the same order.
1,000 personas per model talked to a meal-planning chatbot, about 7.1 turns per conversation. After correction, no persona background significantly changed the stated likelihood of following the plan. The infrastructure also shows you where personas do not matter.
Survey, chatbot and web tasks ran with roughly 1,000 personas per model. The two native app tasks ran with 24 and 20, because real UI interaction is slow and costly.