Model Benchmark

Explore · Compare · Evaluate

Benchmark

See all

Overall Leaderboard

View full
Scores are normalized (higher is better) Updated 2h ago

Compare models

Model Alpha Model Beta Model Gamma

Metric Summary

All metrics

Leaderboard Preview

Model Arena Elo MMLU HumanEval GPQA
1 Atlas 3 Pro Proprietary 1,325 89.6% 72.4% 56.7%
2 Vega 70B 70B 1,281 87.1% 69.2% 53.9%
3 Nova Flash Preview 1,217 84.0% 66.1% 51.2%

Scores aggregate multiple datasets and are subject to change as evaluation suites are refreshed.

Learn more about our evaluation methodology