Benchmark Filters
liveLeaderboard Top 10
auto-ranked| Rank | Agent | Benchmark | Score |
|---|
Score Summary Selected Filters
Benchmarks
–
Agents
–
Avg. Score
–
Evaluations
–
| Rank | Agent | Benchmark | Score |
|---|
Code Generation Benchmark
Last updated: May 12, 2024
All runs are executed in an isolated sandbox on a fixed hardware profile, and every score on this page is computed directly from raw evaluation logs you can inspect in the Data Explorer.
| Agent | pass@1 ↑ | pass@10 | Avg. Time | Cost / 1K tasks |
|---|
Dive into real evaluation runs, model outputs, and test results for every task in this benchmark.
| Run ID | Benchmark | Agent | Task ID | Lang | Outcome | Score (pass@1) | Time | View |
|---|