▤Benchmark19d ago
Epoch AI evaluates Claude Opus 4.7Furniture Assembly 33.3%
Epoch AI Benchmarking HubClaude Opus 4.7Anthropic ▤Benchmark19d ago
Epoch AI evaluates Claude Opus 4.6Furniture Assembly 28.3%
Epoch AI Benchmarking HubClaude Opus 4.6Anthropic ▤Benchmark19d ago
Epoch AI evaluates Claude Opus 4.5 (64k thinking)Furniture Assembly 28.3%
▤Benchmark19d ago
Epoch AI evaluates GPT-5.6 LunaFurniture Assembly 42.5%
Epoch AI Benchmarking HubGPT-5.6 LunaOpenAI ▤Benchmark19d ago
Epoch AI evaluates GPT-6 AstraFurniture Assembly 80.0%
Epoch AI Benchmarking HubGPT-6 AstraOpenAI ▤Benchmark19d ago
Epoch AI evaluates GPT-5.6 SolFurniture Assembly 56.7%
Epoch AI Benchmarking HubGPT-5.6 SolOpenAI ▤Benchmark19d ago
Epoch AI evaluates GPT-5.6 TerraFurniture Assembly 54.2%
Epoch AI Benchmarking HubGPT-5.6 TerraOpenAI ▤Benchmark19d ago
Epoch AI evaluates GPT-5.5 (xhigh)Furniture Assembly 44.2%
Epoch AI Benchmarking HubGPT-5.5 (xhigh)OpenAI ▤Benchmark19d ago
Epoch AI evaluates GPT-5.2 (xhigh)Furniture Assembly 38.3%
Epoch AI Benchmarking HubGPT-5.2 (xhigh)OpenAI ▤Benchmark19d ago
Epoch AI evaluates GPT-5.4 (xhigh)Furniture Assembly 37.5%
Epoch AI Benchmarking HubGPT-5.4 (xhigh)OpenAI ◇Open release20d ago
DeepSeek releases DeepSeek V4.1 FlashDeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
◇Open release20d ago
DeepSeek releases DeepSeek V4.1 Flash (unknown)◆Model launch21d ago
Inception releases Mercury 2.5Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
◆Model launch21d ago
Inception Labs releases Mercury 2.5 (unknown)▤Benchmark21d ago
Epoch AI evaluates Claude Fable 5.1EBR-bench 57.1%
Epoch AI Benchmarking HubClaude Fable 5.1Anthropic ◇Open release21d ago
Nex AGI releases Nex-N2.5-ProNex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...
OpenRouter models APINex-N2.5-ProNex AGI ◇Open release21d ago
Nex AGI releases Nex-N2.5-MiniNex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...
OpenRouter models APINex-N2.5-MiniNex AGI ▤Benchmark24d ago
Epoch AI evaluates Claude Opus 5EBR-bench 45.7%
Epoch AI Benchmarking HubClaude Opus 5Anthropic ◆Model launch25d ago
OpenAI releases GPT-6 Astra ProGPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks.
Learn more in OpenAI's do
OpenRouter models APIGPT-6 Astra ProOpenAI ◆Model launch25d ago
OpenAI releases GPT-6 Astra (batch)GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-hor
◆Model launch25d ago
OpenAI releases GPT-6 Astra Pro (batch)GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks.
Learn more in OpenAI's do
▤Benchmark25d ago
Epoch AI evaluates GPT-5.6 SolEBR-bench 44.8%
Epoch AI Benchmarking HubGPT-5.6 SolOpenAI ◆Model launch26d ago
OpenAI releases GPT-6 AstraGPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-hor
Epoch AI Benchmarking HubGPT-6 AstraOpenAI ◆Model launch26d ago
OpenAI releases GPT-6 Astra (none)Epoch AI Benchmarking HubGPT-6 Astra (none)OpenAI ◆Model launch26d ago
OpenAI releases GPT-6 Astra (low)Epoch AI Benchmarking HubGPT-6 Astra (low)OpenAI ◆Model launch26d ago
OpenAI releases GPT-6 Astra (medium)◆Model launch26d ago
OpenAI releases GPT-6 Astra (xhigh)Epoch AI Benchmarking HubGPT-6 Astra (xhigh)OpenAI ◆Model launch26d ago
OpenAI releases GPT-6 Astra (pro, max)◆Model launch26d ago
OpenAI releases GPT-6 Astra (unknown thinking)◆Model launch27d ago
Meta releases Muse Spark 1.3Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...
Epoch AI Benchmarking HubMuse Spark 1.3Meta ◆Model launch27d ago
Google releases Gemini 3.8 FlashGemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Epoch AI Benchmarking HubGemini 3.8 FlashGoogle ◆Model launch27d ago
Google releases Gemini 3.8 Flash (batch)Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
◆Model launch27d ago
Meta releases Muse Spark 1.3 ContributorMuse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track in
◆Model launch27d ago
Meta AI releases Muse Spark 1.3 (xhigh)◆Model launch27d ago
Google DeepMind releases Gemini 3.8 Flash (unknown)◆Model launch27d ago
Meta AI releases Muse Spark 1.3 (unknown)◆Model launch27d ago
Google DeepMind releases Gemini 3.8 Flash (medium)◆Model launch27d ago
Google DeepMind releases Gemini 3.8 Flash (low)▤Benchmark27d ago
Epoch AI evaluates Qwen3.8 Max (0902) (xhigh)GPQA diamond 92.3% · FrontierMath-Tiers-1-3-v2-Private 65.6% · FrontierMath-Tier-4-v2-Private 34.1% · OTIS Mock AIME 2024-2025 100.0%
▤Benchmark27d ago
Epoch AI evaluates Gemini 3.8 FlashGPQA diamond 95.4% · FrontierMath-Tiers-1-3-v2-Private 68.4% · FrontierMath-Tier-4-v2-Private 22.0% · OTIS Mock AIME 2024-2025 98.9%
Epoch AI Benchmarking HubGemini 3.8 FlashGoogle