← Benchmark leaderboardAll leaderboards

humaneval / Benchmark deep dive

HumanEval: inside the results.

Compare the outcomes, understand the gaps, and inspect the work behind every score.

10selected tasks per model
3model configurations
29/30primary-verified attempts
82recorded tool actions

What the results tell us.

01

Completed attempts clear the tests.

29 of 30 selected attempts passed the declared primary evaluation. 1 ended with an execution error or timeout. Every selected attempt stays in the report.

02

Read the tie in context.

10 selected tasks per model cover a small part of 164 benchmark tasks. One attempt per task measures this run; repeated-run consistency and full-suite leadership are not established.

03

Cost is still an open measurement.

Dollar cost is recorded for 0 of 30 attempts. Subscription usage and missing provider billing are not converted into an invented per-task price. Latency includes the scope declared in the protocol below.

Performance in context.

Scores use the same selected tasks and the report’s declared primary evaluator.

Native executable task success · one attempt per task
Model / agentSuccessExecution errorsMean latencyCost / task
ChatGPT subscriptiongpt-6-astra90.0%9/10 selected tasks
1135.8sNot reported
Claude subscriptionclaude-sonnet-5100.0%10/10 selected tasks
048.1sNot reported
DeepSeekdeepseek-v4-pro100.0%10/10 selected tasks
049.9sNot reported

Pilot coverage: 10/164 source tasks. Small fixed selections do not establish a full-benchmark ranking. Missing grades remain visible.

Where attempts fall short.

Execution failure and an incorrect result are separate outcomes.

ChatGPT subscription

Verified successes
9
Graded failures
0
Execution errors / timeouts
1
Completed, ungraded
0

TimeoutError:

Claude subscription

Verified successes
10
Graded failures
0
Execution errors / timeouts
0
Completed, ungraded
0

No failing task was recorded in this selected sample.

DeepSeek

Verified successes
10
Graded failures
0
Execution errors / timeouts
0
Completed, ungraded
0

No failing task was recorded in this selected sample.

The work behind the score

Inspect a task.

Original instructions, observed actions, and the verdict for each attempt.

Pass

gpt-6-astra · 73.0s · 3 recorded actions

Original task

Complete the following Python snippet. Write only the completion (the code after this snippet) to /logs/artifacts/answer.py. Do not repeat the provided code. def truncate_number(number: float) -> float: """ Given a positive floating point number, it can be decomposed into and integer part (largest integer smaller than given number) and decimals (leftover part always smaller than 1). Return the decimal part of the number. >>> truncate_number(3.5) 0.5 """

Grading criteria & verdicts

Executable checksPass

Recorded benchmark evaluator · 6d43fb980f9fee3c892a914eda09951f772ad10d:original-tests-seed0

reward=1; pass threshold=1.0

Rubric reviewUngraded

Recorded benchmark evaluator · none

Rubric grading is not configured for this task.

Screenshot reviewUngraded

Recorded benchmark evaluator · none

Vision grading is not configured for this task.

Link to this attempt ↗
Source checksums

d5d59a402ec4a05558c257c689dc669746d943a09bdf323844b5f5a203d2face.tar.gz
f7678544d83b3ab5160ec7b814dde7af9aa5e726559327da53a229ee122666d4

raw_outputs.jsonl
61ad44b3d43088eb3b3fc1732970a21758f7a330177c66a51e7788d5a73fb75e

Reproducibility

The evaluation protocol.

Measured ten-task sample from 164 public evaluation tasks. One attempt per task with a common terminal scaffold; original tests, RNG seed 0. This is an agentic evaluation, not the original code-only sampling protocol. Source: https://github.com/openai/human-eval/tree/6d43fb980f9fee3c892a914eda09951f772ad10d. Registered source archive SHA-256: f373485244ee5745ff98af9f9e20ff824f5b4b68bb1f37a61dd66e20e2cb4581. Selected tasks: humaneval-2, humaneval-39, humaneval-48, humaneval-50, humaneval-61, humaneval-64, humaneval-75, humaneval-79, humaneval-100, humaneval-136. ChatGPT subscription: harbor-codex-subscription; observed chatgpt_subscription / 0.154.0 / max. Claude subscription: harbor-claude-subscription; observed claude_subscription / 2.1.273 / max. DeepSeek: harbor-deepseek; observed deepseek_api / max. 3 concurrent task attempts; 600 seconds and 100 terminal tool calls per task. Account queues are outside task clocks. Latency includes environment setup, agent execution and native verification. Subscription cost is unmeasured. Full traces and pinned task/world images are retained in raw_outputs.jsonl.

Task selection
10 frozen tasks, one attempt for each model. 6d43fb980f9f:pilot-1:47adf872a27bf6ae875d4ac4d874e151532770864c2affcaff4fd586d1ca87f6
Tools & environment
GKE gVisor; images pinned in manifest. Provider driver versions are recorded; matching task tools do not imply identical provider implementations.
Budget
600s/task; 100 max turns; no automatic retries
Independent grading
Native executable checks in this original publication. Astra supplements are separately identified when available.. Original grader versions and missing verdicts are preserved.
Evidence
Task prompts, tool inputs, observed results, final answers and grader receipts. Each trajectory links to its original artifact checksums.
Limits
One attempt per task. No pass@k, repeated-run reliability or expert agreement claim. Unreported costs stay unknown. Browser Use’s managed agent belongs to its own harness cohort.
Open the original report and downloads
HumanEval · agent evaluation3 runs · Published evaluation