Skip to content

Multimodal AI

Multimodal AI models accept or produce more than one kind of data - for example text and images, or audio and video - rather than text alone.

A multimodal model can read a screenshot, describe a photo, answer questions about a chart or a PDF, or generate images and speech. Input and output modalities are listed separately: a model may read images but only write text.

Examples tracked on AI Stats Live

Multimodal AI: example tools with company, category, price and popularity rank
ToolCompanyCategoryFromFree tierPopularity
ElevenLabsElevenLabsVoice & speech$5/moYes#8
GeminiGoogleGeneral assistants$19.99/moYes#9
WhisperOpenAIVoice & speech—Unknown#18
GrokxAIGeneral assistants$30/moYes#19
ComfyUIComfy OrgImage generation—Unknown#43
MidjourneyMidjourneyImage generation$10/moNo#48
Qwen ChatAlibaba (Qwen)General assistants—Yes#53
Stable DiffusionStability AIImage generation—Unknown#55
SoraOpenAIVideo generation—Unknown#56
Veo & FlowGoogleVideo generation—Unknown#57

Popularity is measured attention (Wikipedia, Hacker News, package downloads), not quality. Prices are the cheapest paid individual plan; each profile shows whether the price is verified. Data refreshed 2h ago.

Related ranking: Image generation models

  1. 1.Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) Google30 Jun 2026
  2. 2.Nano Banana 2 (Gemini 3.1 Flash Image) Google18 Jun 2026
  3. 3.Nano Banana Pro (Gemini 3 Pro Image) Google18 Jun 2026
  4. 4.GPT-5.4 Image 2 OpenAI21 Apr 2026
  5. 5.Nano Banana 2 (Gemini 3.1 Flash Image Preview) Google26 Feb 2026

Models that output images, newest first (image quality has no shared benchmark in our sources yet). Source: OpenRouter. Full ranking

Browse categories

Related terms

Explanation written by AI Stats Live editors; last reviewed 29 Sep 2026. Examples and rankings update automatically from the sources listed on the methodology page.