Benchmark
Epoch AI evaluates GPT-3.5 Turbo (Jan 2024)
FrontierMath-Tiers-1-3-v2-Private 0.0% · Mystery Game Puzzles 3.0%
- Type
- BENCHMARK_RESULT
- Detection
- benchmark
- Confidence
- high
- Status
Related entities
FrontierMath-Tiers-1-3-v2-Private 0.0% · Mystery Game Puzzles 3.0%
Independent benchmark runs (GPQA Diamond, SWE-bench Verified, FrontierMath, ...) plus externally reported leaderboards and the Epoch Capabilities Index (ECI).