Safety & Design
AI guardrails
The rules and filters that limit what an AI system will say or do, such as blocking harmful topics or redirecting risky conversations.
- Parents
- Businesses
What parents should know
AI guardrails are the rules that limit what a chatbot will say or do. They can block a topic, refuse a dangerous request, or point a risky conversation toward help. No set of rules catches everything. Ask what a product blocks, what a parent can add, and what happens when a child is in distress.
On this page
What are AI guardrails?
AI guardrails are the limits built around a model so it will not say or do certain things. Some limits are in the model. Some are filters that run before or after the reply. Some are product rules, such as a parent-set topic list or a requirement to show a crisis resource. A blocked topic, a refused request for dangerous instructions, and a redirect away from sexual content are all guardrails.
NIST's AI Risk Management Framework treats these limits as part of managing risk, not as a promise of safety. A guardrail can fail when a user rephrases a request, when the model is updated, or when a new topic appears that the rules did not name. Teams test guardrails on purpose. Families can do a smaller version of that test with one or two topics they care about.
Guardrails are not the same as parental controls, though they work together. Guardrails are what the system will or will not do. Parental controls are how an adult sets and watches those rules. A strong product lets the adult see when a guardrail was hit, instead of failing silently.
Why AI guardrails matter
Children will ask things a general chatbot was not prepared to answer. Guardrails are the difference between a graphic reply, a flat refusal, and a reply that is honest and age-fit. They also matter when a child describes harm, bullying, or hopelessness. The right behavior is to slow down, offer a real resource, and tell a trusted adult when the product is built to do that.
Businesses that add AI to a kids' experience inherit this duty. A venue, an app, or a membership program needs to know which topics are out of bounds, who is notified, and what is logged. A filter that only blocks a list of swear words will miss a harmful request written in ordinary language.
Guardrails can also be too blunt. A tool that refuses every question about health or news can push a teen to a product with no limits. The useful standard is specific: block what the family or school has ruled out, explain the limit in plain language, and keep a person in the loop for the hard cases.
How it shows up in practice
- A child asks for dangerous instructions and the product refuses and offers a safer path.
- A parent adds a topic to a block list and the assistant stops engaging with that subject for that child.
- A chat that signals distress shows a crisis line and alerts a parent, without claiming to be emergency services.
- A school turns off image generation for younger grades and leaves guided homework help on.
- A company tests guardrails with reworded prompts, not only with a list of banned words.
How HeyOtto helps
HeyOtto lets parents set topic limits that apply to that child. Real-time safety alerts tell a parent when a conversation needs them. If a child is in distress, the chat can surface crisis resources and helpline numbers. HeyOtto does not auto-contact emergency services.
- Topic limits are set by the parent, per child.
- Alerts can arrive by email, SMS, in the app, or as a daily digest.
- Businesses can talk with HeyOtto about offering AI to kids with those adult-visible limits in place.
For families
Try freeFAQs
What are AI guardrails?
They are the rules and filters that limit what an AI system will say or do. Examples include blocking a harmful topic, refusing a dangerous how-to, and redirecting a risky conversation toward a person or a crisis line. They reduce harm. They do not catch every prompt, so products need testing and a way for an adult to see what happened.
Can kids get around AI guardrails?
Sometimes. Rephrasing a question can slip past a filter that only matches keywords. Stronger guardrails look at the meaning of the request and apply the same limit to that child no matter how the question is worded. You can test this. Ask the blocked question two different ways and compare the replies.
Who sets the guardrails, the company or the parent?
Both, in a well-built product. The company sets baseline limits, such as refusing sexual content involving minors or dangerous instructions. The parent or school adds limits that match their rules, such as topics that are off limits in that home. If only the company can set rules, ask whether those rules match your child.
What should an AI do if a child talks about self-harm?
It should take the message seriously, share a real crisis resource, and tell a trusted adult when the product is designed to alert one. In the United States, 988 is the Suicide and Crisis Lifeline. The AI is not a clinician and should not pretend to handle the emergency alone. Stay with your child if you are the adult who was alerted.
Sources
Last reviewed September 26, 2026. This entry is reviewed twice a year.
