For the first wave of chatbots, everything was text. You typed; it typed back. Today’s leading models are multimodal: they can take in images, documents, audio and in some cases video, and produce speech and images as well as text. This shift has quietly expanded what AI assistants can do.
Updated September 2026: we added links to the research behind how multimodal models see images and a real-world accessibility example.
What multimodal means
A modality is a type of data: text, images, audio or video. A multimodal model can process more than one. In practice this means you can:
- Upload a photo and ask questions about it
- Share a chart, diagram or screenshot and get an explanation
- Upload a scanned document or handwritten notes and extract the content
- Have a spoken conversation, as described in our voice mode tips
- In some tools, analyze video or generate images and speech
How it works, simply
In a text-only model, words are split into tokens and processed. A multimodal model also converts other inputs into a form it can process alongside those tokens. An image, for example, is broken into small patches that are each turned into a numerical representation, an approach popularized by Google’s Vision Transformer research in 2020. Training on huge numbers of images paired with text teaches the model how visual features relate to words.
Increasingly, models are trained on multiple modalities from the start, rather than having separate systems bolted together. Google’s Gemini technical report, for example, describes models built to understand image, audio, video and text together. That tends to make them better at tasks that genuinely combine types of information, such as explaining a diagram.
Practical uses
- Troubleshooting: photograph an error message, a wiring panel or a dashboard warning light.
- Learning: snap a textbook problem and ask for a step-by-step explanation.
- Work documents: extract tables from PDFs, describe charts and compare slide decks.
- Accessibility: describe surroundings, read labels and signs, and explain images for people who are blind or have low vision. The Be My Eyes app launched Be My AI, a GPT-4-powered visual assistant, for exactly this in 2023.
- Shopping and travel: translate menus and signs, identify products and compare labels.
Try it Next time a chart in a report confuses you, take a screenshot and ask: “What is this chart showing, what is the main takeaway, and is anything about it misleading?”
Current limitations
- Fine detail and counting. Models can miss small text, miscount objects or misread precise values on charts.
- Spatial reasoning. Understanding exact positions, measurements and layouts remains challenging.
- Confident errors. Just as with text, models can describe things that are not there. See our hallucinations guide.
- Privacy. Photos and screenshots can contain more personal information than you realize, such as faces, addresses and account details in the background.
Why it matters
Multimodality makes AI useful in the physical world, not just on the screen. It is also a foundation for agents that operate computers by looking at the screen, and for devices such as smart glasses that see what you see. As we predicted in our 2026 trends piece, multimodal is now the default for leading assistants.
Show, don’t describe
When AI can see and hear, you can show it problems instead of describing them. That is a big usability leap, as long as you double-check fine details and think about what is in the frame before you share it.
Sources
- An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale, arXiv, 2020
- Gemini: A Family of Highly Capable Multimodal Models, arXiv, December 2023
- Introducing: Be My AI, Be My Eyes, August 2023



