KnowledgeSketches|← All papersLenny’s PodcastResearch papers
Knowledge Sketches · arXiv 2509.05382 · AIES 2025

Six Privacy Policies, Read So You Never Have To

A Stanford team read the privacy policies of the six biggest US chatbot makers, 28 documents in all, and coded them line by line against California's privacy law. This is what those policies actually say about your chats. By Jennifer King, Kevin Klyman, Emily Capstick, Tiffany Saade and Victoria Hsieh.

0 of 6
developers train on your chats by default. Anthropic was the last holdout, until its switch to opt-out in September 2025
28documents read
700MChatGPT users, Aug 2025
~90%US chatbot market covered
0LLMs used to code them
What the six policies say
Each cell is the paper's reading of the company's own policy, as of May 2025, with Anthropic's later switch accounted for. Hover any cell.
worse for your privacy better for your privacy policy does not say
"not specified" is doing a lot of work in Amazon's column. The paper's all-six claim rests on the Nova interface notice, since Amazon's written policy says nothing

What they say they train on

Sources of training data, as disclosed. A coral dot means the policy says yes, teal means it explicitly says no, an empty dot means the policy does not address it or the wording is ambiguous.

five of the six admit to scraping the public web. Amazon's policy does not mention AI training at all

How long they keep your chats

Retention as stated in policy. Three of the six appear to keep at least some chat data forever.

Google's human-reviewed chats live 3 years, disconnected from your account. Anthropic keeps opted-out users' chats 30 days, everyone else's 5 years

The children's data problem

4 of 6 allow teen accounts and appear to train 13 to 18

Google, Meta, Microsoft and OpenAI all allow accounts for ages 13 to 18. Amazon, Meta and OpenAI do not say they treat teens' chat data differently, indicating they likely train on it by default.

Google went younger under 13

Google expanded Gemini to children under thirteen in May 2025, and will train on data from children 13 to 18, but only if those children opt in.

Two hold the line Anthropic, Microsoft

Anthropic neither collects data from nor accepts users under eighteen, though it does not age-verify. Microsoft collects data from under-18 users but states it does not use it to build language models.

Findings between the rows

Consumers pay with data, businesses do not two-tiered

Enterprise customers are excluded from training by default. OpenAI states that organizations are opted out of data sharing unless they explicitly opt in. The exact opposite of the consumer default.

The opt-out is framed as guilt guiltshaming

OpenAI's interface pitches data sharing as "Improve the model for everyone", a social appeal the paper links to a dark-pattern technique the literature calls guiltshaming. Defaults are sticky; most users never change them.

Only Microsoft says it strips personal data 1 of 6

Microsoft is alone in stating it removes names, phone numbers, device and account identifiers, sensitive personal data, physical addresses and email addresses before training. Anthropic and OpenAI instead train models not to repeat personal data they may contain.

Humans read your chats identifiable

Google and OpenAI both use human reviewers on chats. Reporters found contract workers reviewing Meta AI chats routinely saw identifiable information, and could identify at least one individual from the data.

The answers hide in branch policies 28 documents

Material facts sit in sub-policies and FAQs, not the main privacy policy. OpenAI's practices alone took six documents to piece together. The team coded everything by hand in AtlasTI and deliberately used no LLMs, citing hallucination risk.

Google's own advice to its users"Please don't enter confidential information in your conversations or any data you wouldn't want a reviewer to see or Google to use to improve our products, services, and machine-learning technologies."

What the authors want done

Five recommendations, aimed at policymakers and the developers themselves.

Adopt comprehensive federal privacy regulation

The CCPA only helps after data is collected, and excludes publicly available information. Federal law should close the LLM-shaped gaps.

Require opt-in for model training

Ask users before training on their chats, not after. Children should never be opted in by default.

Put AI practices in the privacy policy itself

Clarify how chat data is collected, stored and processed for training, in the main policy, plus standardized artifacts like datasheets and model cards.

Filter personal information from inputs by default

Strip sensitive data like social security numbers and health data from chats and uploads before anything else happens, regardless of retention policy.

Build privacy-preserving AI

Temporary chats already exist, stored 30 days at OpenAI and 3 days at Google. Apple trains on no user data, Proton keeps no logs. The paper says frontier developers should build on these patterns rather than treat privacy as an afterthought.

Read it with these caveats

  • This is an analysis of policies as written, not of what companies actually do. Practices can be better or worse than the paper trail.
  • It is a snapshot of May 2025, with Anthropic's September opt-out switch folded in. Policies change often, and sub-policy updates usually arrive without notice to users.
  • Scope is the six largest US developers and their consumer chatbots. Startups, open-weight distributors and non-US firms are out of frame.
  • Not every sub-policy was coded. The authors followed links a user could reasonably find from the main policy, onboarding and chat interface.
Based on the paper User Privacy and Large Language Models: An Analysis of Frontier Developers’ Privacy Policies (arXiv 2509.05382), King, Klyman, Capstick, Saade, Hsieh. Stanford University, AIES 2025. Every claim on this page is from the paper.Read the paperDownload the PDF
This sketch began as a paper.
We draw podcast episodes too. Send us a transcript. $1 a sketch, or $49 for all 11 formats.
Sketch mine →