AI Basics
Multimodal AI
AI that can take in or produce more than one kind of content, such as text, images, voice, and video, in the same session.
- Parents
- Educators and school leaders
What parents should know
Multimodal AI works with more than words. A child might type a question, hear a reply, and get a picture. Each mode is another thing to store and another thing a parent may want to see. A text transcript that hides the image is an incomplete record.
On this page
What is multimodal AI?
A single-mode tool only reads and writes text. A multimodal tool can also accept a photo, speak an answer, or draw a picture from a description. Kids meet this in camera features, storybook makers, and chat windows with a microphone. The model is still guessing a plausible result. A picture can be false in the same way a sentence can.
The privacy difference is the file. A photo of a classmate is personal information in a way a math question is not. Voice can be copied. Video can be shared. The parent control that matters is whether those tools are on, and whether the result can leave the account.
Why multimodal AI matters
A child who would never type a cruel sentence may still generate an image of a classmate. Schools feel this first. Turning image generation off for younger grades, while leaving step-by-step homework on, is a multimodal decision, not a ban on AI.
Voice feels private and is easy to record. If a product can clone a voice from a short clip, do not treat the microphone as a toy. HeyOtto's voice feature is speech in and speech out for the assistant. It is not a tool for copying someone else's voice.
How it shows up in practice
- A student adds a drawing to a history project inside an account the teacher can see.
- A parent turns image generation off after a mean picture, and leaves chat on.
- A voice reply is still visible later as part of the conversation record.
- A public story waits for parent approval before anyone outside the family can open it.
How HeyOtto helps
HeyOtto combines chat, age-filtered image generation, stories, and voice in one parent-visible account. Voice is available, and HeyOtto is not voice-first. Kids type and read first. Parents set tool permissions per child, so images can be off while homework stays on. A public share of a creation needs parent approval and can be revoked. The image is still a generated picture, not a photograph of an event.
- Image tools are age-filtered. They are not a face-swap or nudify product.
- Schools can turn image generation and other student tools on or off.
- Parents can read the chat that asked for the image.
For families
Try freeFAQs
Is a picture from AI evidence?
No. A generated image can show a place, a person, or an event that was never photographed. Treat it as a drawing the model made. If a photo is supposed to prove something, you need the source, not the generation.
Can I turn off images and keep chat?
On HeyOtto, tool permissions are per child. You can lock image generation without deleting the account. Schools can make the same choice by grade. Homework help does not require the image tool.
Does voice mean the bot has a body?
No. Voice is a way to hear and speak the same assistant. HeyOtto is not voice-first. A spoken reply can still be wrong, and you should still be able to review what was said.
Should kids upload photos of friends?
No. Use a description or a drawing. A real photo of another child is their privacy, and it can be misused outside the assignment. Parent approval for a public link does not undo a face that should not have been uploaded.
Sources
Last reviewed September 26, 2026. This entry is reviewed twice a year.
