Coding Agent Leaderboard

Compare coding agents across benchmarks

Benchmark Filters

live

Leaderboard Top 10

auto-ranked
Top 10 coding agent results across selected benchmarks
RankAgentBenchmark Score

Score Summary Selected Filters

Benchmarks
Agents
Avg. Score
Evaluations

HumanEval++

Code Generation Benchmark

Last updated: May 12, 2024

Overview

All runs are executed in an isolated sandbox on a fixed hardware profile, and every score on this page is computed directly from raw evaluation logs you can inspect in the Data Explorer.

Methodology

    Agent Score Breakdown Top 5

    pass@1 · higher is better
    Top five agents on this benchmark
    Agentpass@1 ↑pass@10Avg. TimeCost / 1K tasks

    Task Categories Pass@1

    Explore Evaluation Cases

    Real runs, outputs & test results

    Dive into real evaluation runs, model outputs, and test results for every task in this benchmark.

    Evaluation Data Explorer

    Search and explore evaluation runs

    0 runs found

    Evaluation runs

    Evaluation runs matching the current filters
    Run IDBenchmarkAgentTask IDLangOutcome Score (pass@1)TimeView

    Selected Run Details

    run_98762
    Benchmark
    Task
    HE_0088 — two_sum
    Agent
    Claude 3.5 Sonnet
    Language
    JavaScript
    Outcome
    Fail