other benchmark · included in ECI
ScienceQA
ScienceQA benchmark (score column: Score).
Results
22
Random baseline
25.0%
Score ceiling
100%
Released
20 Sep 2022
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Phi-3.5-vision-instruct | 91.3% | — | 16 Aug 2024 | — | Source ↗ | |
| 2 | GPT-4o (May 2024) OpenAI | 88.5% | — | 13 May 2024 | — | Source ↗ | |
| 3 | Gemini 1.0 Pro Vision Google DeepMind | 79.7% | — | 4 Jan 2024 | — | Source ↗ | |
| 4 | falcon-11B-vlm | 74.9% | — | 21 May 2024 | — | Source ↗ | |
| 5 | BLIP-2 (Q-Former) Open source Salesforce Research | 74.2% | — | 6 Feb 2023 | — | Source ↗ | |
| 6 | llama3-llava-next-8b | 73.7% | — | 20 Apr 2024 | — | Source ↗ | |
| 7 | llava-v1.6-vicuna-13b | 73.6% | — | 31 Jan 2024 | — | Source ↗ | |
| 8 | llava-v1.6-mistral-7b | 72.8% | — | 31 Jan 2024 | — | Source ↗ | |
| 9 | MM1-7B-Chat | 72.6% | — | 14 Mar 2024 | — | Source ↗ | |
| 10 | Claude 3 Haiku Anthropic | 72.0% | — | 7 Mar 2024 | — | Source ↗ | |
| 11 | llava-v1.6-vicuna-7b | 70.6% | — | 31 Jan 2024 | — | Source ↗ | |
| 12 | InternVL-Chat-ViT-6B-Vicuna-13B | 70.1% | — | 16 Aug 2024 | — | Source ↗ | |
| 13 | MM1-3B-Chat | 69.4% | — | 14 Mar 2024 | — | Source ↗ | |
| 14 | Qwen-VL-Chat | 68.2% | — | 20 Aug 2023 | — | Source ↗ | |
| 15 | llava-v1.5-7b | 66.8% | — | 5 Oct 2023 | — | Source ↗ | |
| 16 | InternVL-Chat-ViT-6B-Vicuna-7B | 66.2% | — | 25 Dec 2023 | — | Source ↗ | |
| 17 | instructblip-vicuna-13b | 63.1% | — | 25 Dec 2023 | — | Source ↗ | |
| 18 | instructblip-vicuna-7b | 60.5% | — | 22 May 2023 | — | Source ↗ | |
| 19 | Llama 2-13B Open source Meta AI | 55.8% | — | 18 Jul 2023 | — | Source ↗ | |
| 20 | LLaMA-13B Open source Meta AI | 43.3% | — | 24 Feb 2023 | — | Source ↗ | |
| 21 | Llama 2-7B Open source Meta AI | 43.1% | — | 18 Jul 2023 | — | Source ↗ | |
| 22 | LLaMA-7B Open source Meta AI | 36.2% | — | 24 Feb 2023 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.