A Stanford team read the privacy policies of the six biggest US chatbot makers, 28 documents in all, and coded them line by line against California's privacy law. This is what those policies actually say about your chats. By Jennifer King, Kevin Klyman, Emily Capstick, Tiffany Saade and Victoria Hsieh.
Sources of training data, as disclosed. A coral dot means the policy says yes, teal means it explicitly says no, an empty dot means the policy does not address it or the wording is ambiguous.
five of the six admit to scraping the public web. Amazon's policy does not mention AI training at all
Retention as stated in policy. Three of the six appear to keep at least some chat data forever.
Google's human-reviewed chats live 3 years, disconnected from your account. Anthropic keeps opted-out users' chats 30 days, everyone else's 5 years
Google, Meta, Microsoft and OpenAI all allow accounts for ages 13 to 18. Amazon, Meta and OpenAI do not say they treat teens' chat data differently, indicating they likely train on it by default.
Google expanded Gemini to children under thirteen in May 2025, and will train on data from children 13 to 18, but only if those children opt in.
Anthropic neither collects data from nor accepts users under eighteen, though it does not age-verify. Microsoft collects data from under-18 users but states it does not use it to build language models.
Enterprise customers are excluded from training by default. OpenAI states that organizations are opted out of data sharing unless they explicitly opt in. The exact opposite of the consumer default.
OpenAI's interface pitches data sharing as "Improve the model for everyone", a social appeal the paper links to a dark-pattern technique the literature calls guiltshaming. Defaults are sticky; most users never change them.
Microsoft is alone in stating it removes names, phone numbers, device and account identifiers, sensitive personal data, physical addresses and email addresses before training. Anthropic and OpenAI instead train models not to repeat personal data they may contain.
Google and OpenAI both use human reviewers on chats. Reporters found contract workers reviewing Meta AI chats routinely saw identifiable information, and could identify at least one individual from the data.
Material facts sit in sub-policies and FAQs, not the main privacy policy. OpenAI's practices alone took six documents to piece together. The team coded everything by hand in AtlasTI and deliberately used no LLMs, citing hallucination risk.
Five recommendations, aimed at policymakers and the developers themselves.
The CCPA only helps after data is collected, and excludes publicly available information. Federal law should close the LLM-shaped gaps.
Ask users before training on their chats, not after. Children should never be opted in by default.
Clarify how chat data is collected, stored and processed for training, in the main policy, plus standardized artifacts like datasheets and model cards.
Strip sensitive data like social security numbers and health data from chats and uploads before anything else happens, regardless of retention policy.
Temporary chats already exist, stored 30 days at OpenAI and 3 days at Google. Apple trains on no user data, Proton keeps no logs. The paper says frontier developers should build on these patterns rather than treat privacy as an afterthought.