AI news & updates
803 events
- BenchmarkEpoch AI evaluates DeepSeek v4 Pro (unknown thinking)
Mystery Game Puzzles 17.0%
- BenchmarkEpoch AI evaluates GPT-5.4 nano (low)
Mystery Game Puzzles 3.0%
- BenchmarkEpoch AI evaluates Nemotron 3 Ultra
Mystery Game Puzzles 20.0%
- BenchmarkEpoch AI evaluates GPT-5.6 Terra (none)
Mystery Game Puzzles 14.0%
- BenchmarkEpoch AI evaluates GPT-5 nano (medium)
Mystery Game Puzzles 9.0%
- BenchmarkEpoch AI evaluates GPT-5.1 (no thinking)
Mystery Game Puzzles 15.0%
- BenchmarkEpoch AI evaluates GPT-5.4 mini (none)
Mystery Game Puzzles 11.0%
- BenchmarkEpoch AI evaluates GPT-5.4 nano (no thinking)
Mystery Game Puzzles 9.0%
- Open releaseMajorQwen releases Qwen3.8 Flash
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video
- Open releaseZ.ai releases GLM 5.3 Flash (batch)
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
- BenchmarkEpoch AI evaluates GLM 5.3 Flash
GPQA diamond 90.2% · OTIS Mock AIME 2024-2025 93.9% · Chess Puzzles 14.0%
- BenchmarkEpoch AI evaluates GLM 5.3
FrontierMath-Tiers-1-3-v2-Private 68.8% · FrontierMath-Tier-4-v2-Private 29.3%
- BenchmarkEpoch AI evaluates GLM 5.3
GPQA diamond 90.9% · OTIS Mock AIME 2024-2025 91.1% · Chess Puzzles 21.0%
- Model launchMajorMeta releases Muse Spark 1.2 Contributor
Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...
- Open releaseMajorDeepSeek releases DeepSeek V4 Flash Vision Exp
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base
- Open releaseZ.ai releases GLM 5.3 Flash
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
- Open releaseTencent releases Hy-MT2-1.8B
Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glos
- Open releaseTencent releases Hy-MT2-30B-A3B
Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual,
- Open releaseZ.ai (Zhipu AI) releases GLM-5.3-Flash (unknown)
- Open releaseTencent releases Hy-MT2-7B
Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based,
- BenchmarkEpoch AI evaluates DeepSeek V4 Pro 0813
FrontierMath-Tiers-1-3-v2-Private 64.6% · FrontierMath-Tier-4-v2-Private 26.8% · Mystery Game Puzzles 43.0%
- Open releaseZ.ai releases GLM 5.3 (batch)
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
- BenchmarkEpoch AI evaluates DeepSeek V4 Pro 0813
GPQA diamond 91.7% · OTIS Mock AIME 2024-2025 98.6% · Chess Puzzles 47.0%
- BenchmarkEpoch AI evaluates Grok 4.6 (xhigh)
EBR-bench 30.5%
- BenchmarkEpoch AI evaluates Inkling Small (xhigh)
Mystery Game Puzzles 6.0%
- Open releaseZ.ai releases GLM 5.3
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
- Open releaseMajorQwen releases Qwen3.8 27B
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...
- Open releaseMajorAlibaba releases Qwen 3.8 27B
- Model launchZ.ai (Zhipu AI) releases GLM-5.3 (unknown thinking)
- Open releaseMajorAlibaba releases Qwen3.8 27B (xhigh)
- BenchmarkEpoch AI evaluates Gemini 3.7 Flash
GPQA diamond 94.8% · FrontierMath-Tiers-1-3-v2-Private 71.6% · FrontierMath-Tier-4-v2-Private 36.6% · OTIS Mock AIME 2024-2025 97.2%
- BenchmarkEpoch AI evaluates Inkling Small (xhigh)
GPQA diamond 88.5% · FrontierMath-Tiers-1-3-v2-Private 46.3% · FrontierMath-Tier-4-v2-Private 17.1% · OTIS Mock AIME 2024-2025 90.0%
- BenchmarkEpoch AI evaluates Grok 4.6 (xhigh)
GPQA diamond 93.2% · FrontierMath-Tiers-1-3-v2-Private 66.0% · FrontierMath-Tier-4-v2-Private 31.7% · OTIS Mock AIME 2024-2025 99.2%
- Model launchMajorGoogle releases Gemini 3.7 Flash
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
- Model launchMajorGoogle releases Gemini 3.7 Flash (batch)
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
- Open releaseMajorDeepSeek releases DeepSeek V4 Pro 0813
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
- Model launchMajorGoogle DeepMind releases Gemini 3.7 Flash (medium)
- Open releaseMajorDeepSeek releases DeepSeek V4 Pro 0813 (low)
- Model launchMajorGoogle DeepMind releases Gemini 3.7 Flash (low)
- Open releaseMajorDeepSeek releases DeepSeek V4 Pro 0813 (none)