KnowledgeSketches|← All papersLenny’s PodcastResearch papers
Knowledge Sketches · arXiv 2608.04205 · August 2026

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

Human testing of AI products is slow and expensive. This paper builds the other option at planetary scale: a simulated population you can put in front of a survey, a chatbot, a website, or an app before any real person sees it. Organized by Xiaomin Li and Yuexing Hao, with contributors and advisors across dozens of universities and labs.

0
persona records in Persona 8B, one per human on Earth, give or take
1,290dimensions
4environments
1,010tasks
18,189trials run
human-grounded synthetic
drag to spin. 24,000 dots standing in for 8,300,000,000 records

What a persona is made of

Every record answers 1,290 categorical questions, grouped into five areas. Lifestyle is the biggest slice, capability close behind.

grounded in UN population data, World Bank, ILOSTAT and public surveys

Two ways a persona is born

Sampled from a graph

region language career proficiency

400,000 released personas are drawn from a directed graph over all 1,290 dimensions. Each attribute is sampled conditional on its parents, so correlated traits stay correlated. A persona whose primary language is English cannot have an English proficiency of none.

Extracted from real people

599,847 released personas come from six real-world sources, mapped into the same schema.

  • Wikipedia biographies323,438
  • Stack Overflow developer survey113,120
  • Amazon review histories97,915
  • General Social Survey63,532
  • PRISM alignment profiles1,487
  • MatrAIx persona survey, consented355

The million they released

Persona 8B mostly stays behind the curtain. The public coreset holds 999,847 records, contradiction-checked, deduplicated, and calibrated to published demographics for age, region, gender and urbanicity.

human-grounded, 599,847synthetic, 400,000

How good are the extractions?

Humans rated a 100-persona sample of the extraction quality. Then the models were scored on how often they landed within one point of the human mean.

4.135/5mean human quality score
89.1%two LLM judges within one point of each other

agreement with the human mean, share of comparisons within one point

Four places a persona goes to work

The MatrAIx Playground runs the same persona through four kinds of product surface. The 1,010 tasks skew heavily toward the two cheap environments; native web and app runs stay small because every trial is expensive.

Survey
621 tasks

completes questionnaires, records answers and rationales

AI Chatbot
371 tasks

converses with a bot, records the dialogue, tool calls, resolution

Web
12 tasks

browses a site, records pages, actions, screenshots, submission

App
6 tasks

operates Linux, macOS and iOS apps, records final state and file changes

1,010 tasks across more than 25 domains

the other bucket covers 20+ domains: travel, legal, insurance, education, entertainment, food, real estate, games

Do personas stay in character?

The adherence study: ten behavioral attributes, four environments, five personas declaring each pole. 400 trials. Each dot below is one trial; a coral dot means the declared behavior was expressed, or correctly suppressed.

91.5% 366 of 400 trials held character

the App environment is the hard one: only 6 of 10 behaviors held on both poles

What moved results, what did not

Trust in AI moved a finance task Cramér's V 0.228 to 0.363

On the OpenBB task, a persona's declared trust level separated subgroups under all three agent models, every effect significant at q < 10−8, and all three models put the four trust groups in the same order.

Meal planning moved nothing best q = 0.51

1,000 personas per model talked to a meal-planning chatbot, about 7.1 turns per conversation. After correction, no persona background significantly changed the stated likelihood of following the plan. The infrastructure also shows you where personas do not matter.

Scale is uneven by design

Survey, chatbot and web tasks ran with roughly 1,000 personas per model. The two native app tasks ran with 24 and 20, because real UI interaction is slow and costly.

Read it with these caveats

  • The persona agents were powered by Claude Opus 4.8, GPT 5.5 and Claude Haiku 4.5. The model is part of the experiment, so every result should be reported with the model that produced it.
  • Important findings should be checked with more than one model before they guide a product decision.
  • Simulated users do not retire real ones. Human studies remain necessary before applying conclusions to real populations or consequential decisions.
Based on the paper MatrAIx: Simulating the World with 8.3 Billion Persona Agents (arXiv 2608.04205), Li, Hao, et al., August 2026. Every number on this page is from the paper.Read the paperDownload the PDF
This sketch began as a paper.
We draw podcast episodes too. Send us a transcript. $1 a sketch, or $49 for all 11 formats.
Sketch mine →