How ChatGPT Learns Without Learning You: Privacy Filter and the Controls You Own!
Hey there, it's me, Shii-chan! Today's topic is one we all quietly wonder about: how does an AI use our conversations to learn? Privacy is worth understanding, so let's dig in.
OpenAI News
What was announced?
OpenAI News published a post explaining how ChatGPT learns about the world while protecting privacy. It walks through three things in plain language: what information may be used in model training, how they reduce personal information in that process, and how you can control whether your ChatGPT conversations help improve their models.
ChatGPT keeps getting better at complex, real-world work like coding, research, analysis, and multi-step tasks because it's trained on a wide variety of data. The point of this post is that protecting people's privacy is built in right alongside those gains.
Why it matters
People are using ChatGPT in increasingly personal ways, sometimes for questions that touch sensitive parts of their lives. OpenAI frames this as a deep responsibility and says protecting privacy is central to how they build. When the data used for training and the safeguards that strip personal info are spelled out, it answers that nagging question of "what actually happens to my conversations?" That's useful to know whether you're an engineer or an everyday user.
What changes
The biggest takeaway is that it's now clear you get to choose whether your conversations help train future models.
Go to Settings, then Data Controls, and turn off "Improve the model for everyone." Once it's off, new conversations still show up in your chat history but aren't used to train ChatGPT. For people who leave it on, the personal-info masking I describe below also gets applied.
Dive Deep
The training data is a mix: publicly available information, information accessed through partnerships, and information provided or generated by users, contractors, and researchers. For public internet content, they use only what is freely and openly accessible — think a post on a public forum or a public blog.
The star of the personal-info safeguards is OpenAI Privacy Filter, a tool that finds and masks personal information in text. In OpenAI's evaluations, it removes personal information more effectively than any other tool of its kind. An internal version runs at multiple stages of training — on public datasets and on user conversations when "Improve the model for everyone" is enabled.
There are a few controls you can reach for, too:
- Temporary Chat: start one by opening a new chat and clicking the "Temporary" button in the top-right. These don't appear in chat history, don't create memories, and aren't used to improve the models. They're retained for 30 days for safety, then deleted.
- Memory: remembers things like important people or ongoing projects so you don't have to keep repeating yourself. You can review, edit, or delete it anytime, or turn it off entirely.
- You can also export your data, delete your account, and file requests through the privacy request portal.
OpenAI also states plainly that protecting privacy and addressing serious risks of harm have to work together, and that they keep strengthening how they detect and respond to credible threats of violence. You can read more on the community safety and enforcement page.
Wrap-up
- Training data mixes public info, partner data, and info provided or generated by users and others; public web content is used only if freely and openly accessible
- OpenAI Privacy Filter masks personal info across public datasets and opt-in user conversations
- Turning off "Improve the model for everyone" keeps your conversations out of training
- Temporary Chats skip history, memory, and training, and are deleted after a 30-day safety retention
- Privacy and safety (handling harm risks) are meant to work together
If you've wondered how your ChatGPT data is handled, or you're sorting out a usage policy for your team, this is a tidy explainer worth bookmarking!