We still don’t know how people are really using AI
2026-08-20 · MIT Technology Review
We Still Don't Know How People Really Use AI
Corporate Reports Lack Independent Verification
AI companies such as Anthropic and OpenAI regularly publish reports on how people use their models like Claude and ChatGPT. However, AI researchers argue that these reports only contain the data the companies want the public to see. There is currently no independent source to verify their claims.
Anka Reuel, a computer science PhD candidate at Stanford’s Trustworthy AI Research (STAIR) Lab, is co-leading a new initiative called the AI Observatory. The project aggregates and analyzes real user conversations with popular AI models that were collected with consent from seven existing datasets. Its goal is to provide independent data to help researchers and policymakers better understand generative AI usage patterns.
Reuel notes that highly consequential decisions about AI’s benefits and risks are currently being made based on very limited data.
Key Findings from the AI Observatory
The research reveals that AI usage differs significantly across models and has evolved over time. Compared to reports from major AI companies, which focus primarily on work-related applications, the AI Observatory identified substantially more sensitive behaviors.
When researchers applied Anthropic’s filtering methodology (which focuses on economic and productivity uses) to the Observatory dataset, they found that 48% of conversations would have been filtered out. These non-work conversations showed markedly higher rates of:
- Health and relationship topics (44.2% vs 31.2% in Anthropic’s data)
- Adult or illicit topics (7.9% vs 2.1%)
- Harassment and hate (27.5% vs 5.66%)
- Sexual content (16.7% vs 2.4%)
OpenAI’s 2025 report similarly found that only 30% of consumer usage of ChatGPT was work-related.
Changes Over Time
Analysis of datasets such as WildChat showed clear temporal trends between 2023 and 2025:
- Conversations became longer, with increasing prompt tokens, response tokens, and conversation turns.
- Small talk increased significantly, suggesting growing use of AI for companionship.
- AI models’ self-disclosure (admitting to being a chatbot) decreased.
- Sensitive exchanges — including sexual harassment and hate speech — became less frequent, possibly indicating improved safety guardrails.
Distinct Usage Patterns Across Models
The study found notable differences in how people interact with different AI systems:
- Grok and Gemini were used more frequently for information retrieval. Grok was especially popular for news and politics but also concentrated more misinformation.
- Claude saw heavier use for coding assistance.
- Gemini was more commonly used for social interaction and roleplay.
- ChatGPT was frequently turned to for homework help.
Differences existed even between versions of the same model. Conversations with GPT-3.5-powered ChatGPT tended to be shorter, while those with GPT-4o were longer and more iterative.
Shayne Longpre, a recent MIT Media Lab PhD graduate and co-lead of the research, stated that “no single company report tells the whole story.”
Scale and Limitations of the Dataset
The AI Observatory analyzed 85,633 conversational turns across 24,521 conversations from approximately 5,000 users interacting with 52 different models between 2023 and 2025. While valuable, this remains modest compared to corporate datasets — Anthropic’s latest index analyzed 1 million Claude conversations, and OpenAI examined 1.5 million.
Because the data comes from voluntarily shared conversations, sensitive uses are likely underrepresented. The researchers caution that their findings should not be interpreted as representative of all AI usage.
Anthropic stated that its published research reflects its teams’ specific questions and that it supports independent external research. OpenAI did not respond to requests for comment.
The AI Observatory provides an important independent perspective on how people actually interact with today’s AI systems, revealing nuances that corporate reports tend to miss.