Skip to content

GLM 5.3 Flash

ModelOpen sourceActive
by Z.ai (Zhipu AI) · family “glm-flash”

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while... (description from the OpenRouter listing)

in: textin: imagein: videoout: textReasoningTool useStructured output
Epoch Capabilities Index
151.9 #41 of 268
90% CI 149.6 – 154.2
Listing ↗
Released
20 Aug 2026
Context window
1.31M tokens
Max output
944K tokens
Input $ / 1M tokens
$0.15
Output $ / 1M tokens
$0.50
Cached input $ / 1M
$0.030
Each benchmark shown separately with its own source

Benchmark results (10)

BenchmarkDomainScorevs best recordedSettingRunSource
Mystery Game Puzzlesgames8.0% ±2.7
10%
max28 Aug 2026Epoch ↗
Surface Evolver Benchscience52.5%
55%
max—External ↗
FrontierMath-Tiers-1-3-v2-Privatemath55.8% ±2.9
60%
max27 Aug 2026Epoch ↗
FrontierMath-Tier-4-v2-Privatemath17.1% ±5.9
17%
max27 Aug 2026Epoch ↗
FrontierCodecoding31.8%
60%
max—External ↗
DeepSWEcoding63.4%
86%
max—External ↗
ProofBenchmath21.0%
21%
max—External ↗
Chess Puzzlesgames14.0% ±3.5
19%
max26 Aug 2026Epoch ↗
OTIS Mock AIME 2024-2025math93.9% ±2.2
94%
max26 Aug 2026Epoch ↗
GPQA diamondscience90.2% ±1.7
94%
max26 Aug 2026Epoch ↗

Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.

API list price over time

Price history

$ per 1M tokens

3 recorded prices on 2 Sep 2026 (OpenRouter listing + Internet Archive snapshots); re-read on every data refresh, most recently 7h ago. Steps show when the price changed.

Events