Contractors are reading your ChatGPT chats under Project Lily
Leaked internal documents and real prompts show OpenAI pays hundreds of contractors to read and rate ChatGPT conversations, some of which still contain sensitive personal detail.

OpenAI employs hundreds of contractors whose job is to read the prompts people type into ChatGPT, and those conversations sometimes contain sensitive personal information, according to an investigation published by 404 Media. The report is based on leaked internal documents, internal chat channels and real prompts seen by the publication.
What the reviewers actually see
The reviewers do not see usernames, and OpenAI says it tries to strip personal information from prompts before they reach human eyes. The company also acknowledged that sensitive details can still get through. The contractors’ job is to rate and critique the replies the chatbot produces, which is how the company works out which answers are good, which are unhelpful and which need to be retrained. In other words, a share of the improvement people notice in ChatGPT comes from paid human attention on real conversations rather than from scraping alone.
That is a distinction worth holding on to, because ChatGPT is used less like a search box and more like a therapist, a colleague or a friend. People type in medical worries, money trouble and relationship detail they would not post publicly, and many of them assume the machine on the other end is the only reader.
Training the flattery out of the model
Internal documents reviewed by 404 Media show reviewers working on two specific habits: teaching ChatGPT not to describe itself as a person, and teaching it to be less sycophantic. The second has been a persistent complaint about OpenAI’s models, with an over-eager 4o update drawing heavy criticism and, according to the report, featuring in litigation. A model that agrees with everything is cheap to build and expensive to trust, and the fix appears to be human reviewers marking the praise-heavy answers as failures.
Anthropic is doing the same thing
OpenAI is not alone in this. Anthropic confirmed to 404 Media that it also uses human review to improve its models, which makes the practice closer to an industry standard than an OpenAI quirk. Neither company hides the general principle in its privacy documentation, but the scale and the mechanics — how many people, how many chats, what they are told to look for — are the parts most users have never seen set out plainly.
Why it lands now
ChatGPT passed 900 million users, and the volume of intimate material passing through it grows with every one of them. Contract review of that material is not a scandal in itself: someone has to judge whether an answer was good. The question the report raises is about consent and precision. If the safety net that removes personal detail before review is imperfect, and the company says it is, then the practical privacy guarantee is a filter rather than a promise — and users have no way to tell which of their conversations fell through it.
Our opinion
The uncomfortable part of this story is not that humans read chatbots’ output. It is that everyone assumed some version of it and nobody could name the number. A model improves partly because contractors sit and grade boredom for a living, and that labour is invisible in every launch keynote. OpenAI’s anonymisation filter is a reasonable safeguard, but a safeguard that admits it sometimes fails is a different product promise from one that never mentions the possibility. Users deserve the blunt version: type it in as though a stranger may read it, because sometimes one will. The interesting follow-up is whether the same transparency arrives for the models that never ask permission at all.