What a company does with your data when its AI is free

One query to a large model costs real money in electricity and graphics cards. If you are not paying it, the interesting question is not «why is this free?» but «what does the company get instead?». The answer is not a secret: it is in the terms, written in a language nobody reads.
Producing an answer from a large model burns electricity, memory and time on graphics cards that cost tens of thousands. None of that is free. When a service hands it to you without charging, somebody is footing the bill and expects something back.
Worth saying without melodrama: at the large companies, that something is not usually selling your conversations to third parties. It is something else — less scandalous, and more useful to them.
The four things they do get
1. Material to improve the model. This is the main one. Your conversations — especially the ones you correct, retry or rate — help tune the system. On free tiers this normally arrives switched on, with an option to turn it off in settings. How to do that is covered in how to delete your ChatGPT history.
2. Knowing what people ask for. What gets asked, in which language, at what hour, which answers make users give up. Aggregated, that is worth a fortune when deciding what to build next.
3. Habit. The most underrated one. A genuinely good free tier builds the habit, and the habit is what turns into a subscription one day. A good deal of what we call “free” is marketing with very good numbers behind it.
4. Account data. Email, device, rough location by IP, how often you show up. The same as any cloud service, no more and no less.
The part that surprises people: humans read some of it
This is the one that stings when you find it, and it is not hidden — it is in the policies.
The large providers reserve the ability for human reviewers to read a sample of conversations, to assess answer quality and catch abuse. Google states it outright in the Gemini help pages, along with the obvious advice: do not put confidential information in there.
It does not mean somebody is reading yours right now. It means you cannot rule it out, and any risk calculation has to start from that.
Free, paid and business are not the same deal
The difference that matters is not the price, it is the type of contract:
| Trains on your data | Human review | Stored | |
|---|---|---|---|
| Free tier | Usually yes, can be turned off | Possible | Yes |
| Personal paid | Depends on the setting | Possible | Yes |
| Business / Education | Usually not, by default | Limited | Per contract |
| Developer API | Usually not, by default | Limited | Short retention |
What stands out in that table is that paying for a personal plan does not automatically change the deal. It changes your usage limits and speed. Training remains a toggle you have to go and look at.
Where the framework genuinely changes is business accounts and the API, because there a signed contract with specific obligations exists.
Where the risk stops being theoretical
With a large, well-known service, the realistic worry is not that a stranger reads your things. It is that a sensitive text ends up stored for longer than you assumed, in a jurisdiction that is not yours, with a small but real chance of crossing a reviewer’s screen.
With a small, unknown service — or one installed as a browser extension — the picture changes entirely. There is no reputation at stake, the privacy policy may be a copied template, and there is no way to verify anything. It is the same reasoning that applies to any network you do not control, set out in public Wi-Fi: which risks are real: the danger is not the technology, it is who is on the other end.
If you cannot name the company behind an AI, do not paste anything into it you would not write on a postcard.
The useful question
Instead of “is this AI safe?”, which has no answer, this one works better: if what I am about to paste turned up in six months somewhere I do not control, what would happen?
- A complaint letter to a phone company: nothing happens.
- Marketing copy: nothing happens.
- A contract with names, ID numbers and a bank account: quite a lot happens.
- A relative’s medical report: a great deal happens, and it was not yours to share.
For the last two, the fix is not to stop using the tool but to strip out what identifies people. The full method is in why you should never paste private data into an AI.
What you can do today, in ten minutes
- Open the privacy settings of whichever AI you use daily and check the training toggle.
- Decide which account is for what: the personal one for trivia, the work one for work.
- Uninstall any AI browser extension whose maker you cannot name.
- Get into the habit of swapping out names and numbers before you paste.
The short version
- Free means the company is paid in data, in learning and in habit, not in cash.
- The norm is not selling conversations but using them to train, plus sampled review.
- Human review exists and is written into the policies.
- Paying for a personal plan does not switch training off by itself — check the setting.
- The big risk is not the AI you have heard of. It is the one you have not.
Frequently asked
The questions that keep coming up
Do free AI services sell my conversations?
The large companies do not sell them as a product. The normal use is training and improving the model, plus letting human reviewers read a sample for quality and abuse detection. With small or unknown services, the guarantees are far weaker.
Can an employee read my conversation?
Yes, in certain circumstances. The major providers' policies allow human review of samples of conversations for safety and quality. Google says so explicitly for Gemini.
Does a paid plan protect me more?
Generally yes, but not because of the payment — because of the account type. Business plans and APIs usually exclude your data from training by default. On personal paid plans it is worth checking the setting, because it does not always change on its own.
What is stored besides what I type?
Email, device, approximate IP location, usage times and, if the app allows uploads, the files you send. It is the same set any cloud service collects.


