AI Basics
Training data
The large collection of text, images, or other content an AI model learns from, which shapes what it can say and which patterns it repeats.
- Parents
- Educators and school leaders
- Policymakers and staff
What parents should know
Training data is the material a model learned from before your child opened the app. Your child's chat is a separate question: some companies use later conversations to train again, and some promise they do not. Ask that question out loud. The answer belongs in the contract, not only in an ad.
On this page
What is training data?
A generative model is adjusted on a huge set of examples. Those examples are the training data. They can include books, web pages, code, photos, and conversations. The model does not store them like a library card. It stores patterns. Those patterns are why it can write a sentence and why it can repeat a bias or a myth that was common in the pile.
After launch, a company may train further on new chats. That is a product decision. Not used to train models means the child's conversation is not added to that pile. It does not mean the original model was trained on nothing, and it does not mean the chat is invisible to the parent account.
Why training data matters
Kids type secrets, jokes, and other people's names. If that text trains a future model, it can leave the family account in a way a delete button does not fully explain. If it does not train a model, you still need to know who inside the company or the school can read it.
Training data also explains unfair answers. A model that rarely saw a kind of student, accent, or family may stumble on purpose or by omission. You will not get the dataset. You can get a straight sentence about whether today's chats join it.
How it shows up in practice
- A district's vendor agreement says student prompts are not used to train a general model.
- A consumer app's settings bury a improve the model switch that defaults to on.
- A parent asks the question in a sales call and writes down the answer.
- A class talks about why a fluent answer can still repeat a stereotype.
How HeyOtto helps
HeyOtto does not use conversations to train models. Kids' chats are not sold. A parent can still read the transcript in the dashboard, and a school deployment gives teachers and advisors a view of student sessions. Those are different from training. Deleting a memory in the parent tools removes what Otto saved for personalization. It is not a claim about the original model's training set.
- There are no ads, and chats are not sold.
- Parents can view, edit, or delete memories Otto saved.
- Ask the school for the agreement if the deployment is a district contract.
For families
Try freeFAQs
If chats are not used to train models, who can still read them?
On HeyOtto, the parent can read the family chats. In a school deployment, teachers or advisors can see student use. Not training a model is a promise about the dataset. It is not a promise that no adult can open the transcript. Tell your child which adults can.
Can I see the training data?
Usually no. Companies do not hand families the pile a foundation model learned from. What you can ask for is a clear line on your child's chats going forward. If that line is missing, treat it as a no.
Does training data make the bot biased?
It can. Patterns in the data become patterns in the replies. That is algorithmic bias. A school conversation works better when students can see the actual prompt and reply, which a teacher dashboard allows, than when they only hear that models are trained on the internet.
Is a memory the same as training?
No. A memory is a note the product keeps so the next chat can be more relevant. Training changes the model for many users. HeyOtto lets a parent view, edit, or delete memories. Conversations are not used to train models.
Sources
Last reviewed September 26, 2026. This entry is reviewed twice a year.
