Multimodal AI
Multimodal AI models accept or produce more than one kind of data - for example text and images, or audio and video - rather than text alone.
A multimodal model can read a screenshot, describe a photo, answer questions about a chart or a PDF, or generate images and speech. Input and output modalities are listed separately: a model may read images but only write text.
Examples tracked on AI Stats Live
| Tool | Company | Category | From | Free tier | Popularity |
|---|---|---|---|---|---|
| ElevenLabs | ElevenLabs | Voice & speech | $5/mo | Yes | #8 |
| Gemini | General assistants | $19.99/mo | Yes | #9 | |
| Whisper | OpenAI | Voice & speech | — | Unknown | #18 |
| Grok | xAI | General assistants | $30/mo | Yes | #19 |
| ComfyUI | Comfy Org | Image generation | — | Unknown | #43 |
| Midjourney | Midjourney | Image generation | $10/mo | No | #48 |
| Qwen Chat | Alibaba (Qwen) | General assistants | — | Yes | #53 |
| Stable Diffusion | Stability AI | Image generation | — | Unknown | #55 |
| Sora | OpenAI | Video generation | — | Unknown | #56 |
| Veo & Flow | Video generation | — | Unknown | #57 |
Popularity is measured attention (Wikipedia, Hacker News, package downloads), not quality. Prices are the cheapest paid individual plan; each profile shows whether the price is verified. Data refreshed 2h ago.
Related ranking: Image generation models
- 1.Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) Google30 Jun 2026
- 2.Nano Banana 2 (Gemini 3.1 Flash Image) Google18 Jun 2026
- 3.Nano Banana Pro (Gemini 3 Pro Image) Google18 Jun 2026
- 4.GPT-5.4 Image 2 OpenAI21 Apr 2026
- 5.Nano Banana 2 (Gemini 3.1 Flash Image Preview) Google26 Feb 2026
Models that output images, newest first (image quality has no shared benchmark in our sources yet). Source: OpenRouter. Full ranking
Browse categories
Related terms
Explanation written by AI Stats Live editors; last reviewed 29 Sep 2026. Examples and rankings update automatically from the sources listed on the methodology page.