AI news & updates
803 events
- BenchmarkEpoch AI evaluates Claude Opus 5
EBR-bench 45.7%
- Model launchMajorOpenAI releases GPT-6 Astra Pro
GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's do
- Model launchMajorOpenAI releases GPT-6 Astra (batch)
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-hor
- Model launchMajorOpenAI releases GPT-6 Astra Pro (batch)
GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's do
- BenchmarkEpoch AI evaluates GPT-5.6 Sol
EBR-bench 44.8%
- Model launchMajorOpenAI releases GPT-6 Astra
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-hor
- Model launchMajorOpenAI releases GPT-6 Astra (none)
- Model launchMajorOpenAI releases GPT-6 Astra (low)
- Model launchMajorOpenAI releases GPT-6 Astra (medium)
- Model launchMajorOpenAI releases GPT-6 Astra (xhigh)
- Model launchMajorOpenAI releases GPT-6 Astra (pro, max)
- Model launchMajorOpenAI releases GPT-6 Astra (unknown thinking)
- Model launchMajorMeta releases Muse Spark 1.3
Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...
- Model launchMajorGoogle releases Gemini 3.8 Flash
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
- Model launchMajorGoogle releases Gemini 3.8 Flash (batch)
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
- Model launchMajorMeta releases Muse Spark 1.3 Contributor
Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track in
- Model launchMajorMeta AI releases Muse Spark 1.3 (xhigh)
- Model launchMajorGoogle DeepMind releases Gemini 3.8 Flash (unknown)
- Model launchMajorMeta AI releases Muse Spark 1.3 (unknown)
- Model launchMajorGoogle DeepMind releases Gemini 3.8 Flash (medium)
- Model launchMajorGoogle DeepMind releases Gemini 3.8 Flash (low)
- BenchmarkEpoch AI evaluates Qwen3.8 Max (0902) (xhigh)
GPQA diamond 92.3% · FrontierMath-Tiers-1-3-v2-Private 65.6% · FrontierMath-Tier-4-v2-Private 34.1% · OTIS Mock AIME 2024-2025 100.0%
- BenchmarkEpoch AI evaluates Gemini 3.8 Flash
GPQA diamond 95.4% · FrontierMath-Tiers-1-3-v2-Private 68.4% · FrontierMath-Tier-4-v2-Private 22.0% · OTIS Mock AIME 2024-2025 98.9%
- Model launchMajorQwen releases Qwen3.8 Max (0902)
Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...
- Model launchMajorAnthropic releases Claude Fable 5.1
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
- Model launchMajorAnthropic releases Claude Fable 5.1 (batch)
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
- Model launchMajorAlibaba releases Qwen3.8 Max (0902) (xhigh)
- Model launchMajorAnthropic releases Claude Fable 5.1 (xhigh)
- Model launchMajorAnthropic releases Claude Fable 5.1 (medium)
- Model launchMajorAnthropic releases Claude Fable 5.1 (low)
- Model launchMajorAnthropic releases Claude Fable 5.1 (unknown)
- BenchmarkEpoch AI evaluates Claude Fable 5.1
FrontierMath-Tiers-1-3-v2-Private 90.2% · FrontierMath-Tier-4-v2-Private 87.8% · OTIS Mock AIME 2024-2025 100.0% · SimpleQA Verified 70.8%
- Open releaseIBM releases Granite 4.2 8B
Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort,...
- BenchmarkEpoch AI evaluates GPT-4o-mini
SimpleQA Verified 8.3%
- BenchmarkEpoch AI evaluates GPT-4o (Aug 2024)
SimpleQA Verified 26.0%
- BenchmarkEpoch AI evaluates GPT-4.1 Nano
SimpleQA Verified 6.0%
- BenchmarkEpoch AI evaluates GPT-4.1 Mini
SimpleQA Verified 12.7%
- BenchmarkEpoch AI evaluates GPT-4.1
SimpleQA Verified 31.1%
- BenchmarkEpoch AI evaluates o3 Mini
SimpleQA Verified 15.3%
- BenchmarkEpoch AI evaluates o1
SimpleQA Verified 41.1%