Skip to content
other benchmark · included in ECI

ScienceQA

ScienceQA benchmark (score column: Score).
Results
22
Random baseline
25.0%
Score ceiling
100%
Released
20 Sep 2022

Score by model release date

Best score by organisation

#ModelScoreRelativeEvidence
1Phi-3.5-vision-instruct
91.3%Source ↗
2GPT-4o (May 2024)
OpenAI
88.5%Source ↗
3Gemini 1.0 Pro Vision
Google DeepMind
79.7%Source ↗
4falcon-11B-vlm
74.9%Source ↗
5BLIP-2 (Q-Former) Open source
Salesforce Research
74.2%Source ↗
6llama3-llava-next-8b
73.7%Source ↗
7llava-v1.6-vicuna-13b
73.6%Source ↗
8llava-v1.6-mistral-7b
72.8%Source ↗
9MM1-7B-Chat
72.6%Source ↗
10Claude 3 Haiku
Anthropic
72.0%Source ↗
11llava-v1.6-vicuna-7b
70.6%Source ↗
12InternVL-Chat-ViT-6B-Vicuna-13B
70.1%Source ↗
13MM1-3B-Chat
69.4%Source ↗
14Qwen-VL-Chat
68.2%Source ↗
15llava-v1.5-7b
66.8%Source ↗
16InternVL-Chat-ViT-6B-Vicuna-7B
66.2%Source ↗
17instructblip-vicuna-13b
63.1%Source ↗
18instructblip-vicuna-7b
60.5%Source ↗
19Llama 2-13B Open source
Meta AI
55.8%Source ↗
20LLaMA-13B Open source
Meta AI
43.3%Source ↗
21Llama 2-7B Open source
Meta AI
43.1%Source ↗
22LLaMA-7B Open source
Meta AI
36.2%Source ↗

Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.