Skip to content
Benchmark

Epoch AI evaluates GPT-3.5 Turbo (Jan 2024)

FrontierMath-Tiers-1-3-v2-Private 0.0% · Mystery Game Puzzles 3.0%

Type
BENCHMARK_RESULT
Detection
benchmark
Confidence
high
Status

Primary source

Epoch AI Benchmarking Hub · logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F ↗

Independent benchmark runs (GPQA Diamond, SWE-bench Verified, FrontierMath, ...) plus externally reported leaderboards and the Epoch Capabilities Index (ECI).