PERFORMANCE
Task success
Higher is better · % of selected tasks
A task must pass the declared primary grader. Completed but ungraded or uncollected tasks prevent a success ranking.
Coding / Benchmark leaderboard
Python function generation with original executable tests.
Ten frozen HumanEval tasks, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. Native tests determine success; one Astra attempt reached the shared time limit before verification completed.
Highest task success
100.0%Claude subscription · DeepSeekAll selected tasks stay in the denominatorLowest median latency
39.4sClaude subscriptionRecorded task time; scope in the protocolLowest reported cost / task
Not measuredNo complete measurement availableReported usage only; see excluded chargesPERFORMANCE
Higher is better · % of selected tasks
A task must pass the declared primary grader. Completed but ungraded or uncollected tasks prevent a success ranking.
TRADEOFFS
More success, less time
Each point uses the same task selection, grading rules, environment and budget. This evaluation covers a subset of the source benchmark.
RELIABILITY
Every selected task is accounted for
Execution errors count against task success. They are not presented as a grader verdict. Supplemental vision and provider judgments remain in the full report.
LATENCY DISTRIBUTION
Minimum → median → P95 → maximum
Median 67.4sP95 600.1s
Median 39.4sP95 110.4s
Median 44.7sP95 90.7s
P95 uses the nearest ranked observed task. Small selections can have a P95 equal to their slowest attempt. Missing timing records are disclosed.
ALL COMPARISON METRICS
| Rank | Agent / model | Errors | Evidence | |||||
|---|---|---|---|---|---|---|---|---|
| 1 | Claude subscriptionclaude-sonnet-5 | 100.0%10/10 passed | 100.0%Execution, irrespective of grade | 39.4s10/10 measured | 110.4s10/10 measured | Not measured0/10 reported | 00 missing · 0 ungraded | Task results ↗ |
| 1 | DeepSeekdeepseek-v4-pro | 100.0%10/10 passed | 100.0%Execution, irrespective of grade | 44.7s10/10 measured | 90.7s10/10 measured | Not measured0/10 reported | 00 missing · 0 ungraded | Task results ↗ |
| 3 | ChatGPT subscriptiongpt-6-astra | 90.0%9/10 passed | 90.0%Execution, irrespective of grade | 67.4s10/10 measured | 600.1s10/10 measured | Not measured0/10 reported | 10 missing · 0 ungraded | Task results ↗ |
Ranks compare task success within this frozen evaluation; ties share rank. Filtering and metric sorting do not change the success rank. Missing measurements are shown explicitly and sort last.
Measured ten-task sample from 164 public evaluation tasks. One attempt per task with a common terminal scaffold; original tests, RNG seed 0. This is an agentic evaluation, not the original code-only sampling protocol. Source: https://github.com/openai/human-eval/tree/6d43fb980f9fee3c892a914eda09951f772ad10d. Registered source archive SHA-256: f373485244ee5745ff98af9f9e20ff824f5b4b68bb1f37a61dd66e20e2cb4581. Selected tasks: humaneval-2, humaneval-39, humaneval-48, humaneval-50, humaneval-61, humaneval-64, humaneval-75, humaneval-79, humaneval-100, humaneval-136. ChatGPT subscription: harbor-codex-subscription; observed chatgpt_subscription / 0.154.0 / max. Claude subscription: harbor-claude-subscription; observed claude_subscription / 2.1.273 / max. DeepSeek: harbor-deepseek; observed deepseek_api / max. 3 concurrent task attempts; 600 seconds and 100 terminal tool calls per task. Account queues are outside task clocks. Latency includes environment setup, agent execution and native verification. Subscription cost is unmeasured. Full traces and pinned task/world images are retained in raw_outputs.jsonl.
ChatGPT subscription: 600s/task; 100 max turns; no automatic retries. GKE gVisor; images pinned in manifest.
Claude subscription: 600s/task; 100 max turns; no automatic retries. GKE gVisor; images pinned in manifest.
DeepSeek: 600s/task; 100 max turns; no automatic retries. GKE gVisor; images pinned in manifest.
Cost reflects the charges present in the source report; hosting, subscription fees and independent grading may be excluded. Task latency uses the source report’s measurement window. These are not streaming inference measurements.
Inspect the source report and individual attempts ↗Measured ten-task sample from 164 public evaluation tasks. One attempt per task with a common terminal scaffold; original tests, RNG seed 0. This is an agentic evaluation, not the original code-only sampling protocol. Source: https://github.com/openai/human-eval/tree/6d43fb980f9fee3c892a914eda09951f772ad10d. Registered source archive SHA-256: f373485244ee5745ff98af9f9e20ff824f5b4b68bb1f37a61dd66e20e2cb4581. Selected tasks: humaneval-2, humaneval-39, humaneval-48, humaneval-50, humaneval-61, humaneval-64, humaneval-75, humaneval-79, humaneval-100, humaneval-136. ChatGPT subscription: harbor-codex-subscription; observed chatgpt_subscription / 0.154.0 / max. Claude subscription: harbor-claude-subscription; observed claude_subscription / 2.1.273 / max. DeepSeek: harbor-deepseek; observed deepseek_api / max. 3 concurrent task attempts; 600 seconds and 100 terminal tool calls per task. Account queues are outside task clocks. Latency includes environment setup, agent execution and native verification. Subscription cost is unmeasured. Full traces and pinned task/world images are retained in raw_outputs.jsonl.
Read the failure analysis, compare model outcomes, and inspect task trajectories and grading criteria.
HumanEval · agent evaluation · Open the deep dive →Native task success, errors, latency and recorded cost.
Blobfish-hosted Harbor harness on GKE; frozen task subset.
The published pilot is a subset. Inspect the run protocol before comparing scores.
A managed Harbor task package is registered on GKE. Published reports identify the agent configurations and selected tasks.
Read the harness and grading protocol →