Coding / Benchmark leaderboard

MBPP

Python programming tasks evaluated against source-defined tests.

Harbor · GKE · Native tests

MBPP · agent evaluation

View full report ↗

Ten matching tasks from the official MBPP test split, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. All three passed the original native tests on these selected tasks.

Matched evaluation10 selected / 500 source tasksNative tests2026-09-16

Highest task success

100.0%ChatGPT subscription · Claude subscription · DeepSeekAll selected tasks stay in the denominator

Lowest median latency

33.3sDeepSeekRecorded task time; scope in the protocol

Lowest reported cost / task

Not measuredNo complete measurement availableReported usage only; see excluded charges
Compare agents

PERFORMANCE

Task success

Higher is better · % of selected tasks

ChatGPT subscription100.0%

gpt-6-astra10 pass · 0 errors

Claude subscription100.0%

claude-sonnet-510 pass · 0 errors

DeepSeek100.0%

deepseek-v4-pro10 pass · 0 errors

A task must pass the declared primary grader. Completed but ungraded or uncollected tasks prevent a success ranking.

TRADEOFFS

Success vs. latency

More success, less time

0%25%50%75%100%0.0s19.3s38.6s57.9s77.2sChatGPT subscription · gpt-6-astra: 100.0%; 66.5sClaude subscription · claude-sonnet-5: 100.0%; 41.9sDeepSeek · deepseek-v4-pro: 100.0%; 33.3s
ChatGPT subscriptionClaude subscriptionDeepSeek

Each point uses the same task selection, grading rules, environment and budget. This evaluation covers a subset of the source benchmark.

RELIABILITY

What happened to every task

Every selected task is accounted for

PassedFailedExecution errorUngraded / missing
ChatGPT subscription10/10 recorded
Claude subscription10/10 recorded
DeepSeek10/10 recorded

Execution errors count against task success. They are not presented as a grader verdict. Supplemental vision and provider judgments remain in the full report.

LATENCY DISTRIBUTION

Typical time and the long tail

Minimum → median → P95 → maximum

ChatGPT subscription10/10 timed

Median 66.5sP95 110.7s

Claude subscription10/10 timed

Median 41.9sP95 84.3s

DeepSeek10/10 timed

Median 33.3sP95 127.2s

P95 uses the nearest ranked observed task. Small selections can have a P95 equal to their slowest attempt. Missing timing records are disclosed.

ALL COMPARISON METRICS

Inspect the numbers.

MBPP · agent evaluation · 10 selected tasks · 2026-09-16
RankAgent / modelErrorsEvidence
1ChatGPT subscriptiongpt-6-astra100.0%10/10 passed100.0%Execution, irrespective of grade66.5s10/10 measured110.7s10/10 measuredNot measured0/10 reported00 missing · 0 ungradedTask results ↗
1Claude subscriptionclaude-sonnet-5100.0%10/10 passed100.0%Execution, irrespective of grade41.9s10/10 measured84.3s10/10 measuredNot measured0/10 reported00 missing · 0 ungradedTask results ↗
1DeepSeekdeepseek-v4-pro100.0%10/10 passed100.0%Execution, irrespective of grade33.3s10/10 measured127.2s10/10 measuredNot measured0/10 reported00 missing · 0 ungradedTask results ↗

Ranks compare task success within this frozen evaluation; ties share rank. Filtering and metric sorting do not change the success rank. Missing measurements are shown explicitly and sort last.

Measurement definitions, configuration & protocol

Measured ten-task sample from 500 public evaluation tasks. One attempt per task with a common terminal scaffold; original tests, RNG seed 0. This is an agentic evaluation, not the original code-only sampling protocol. Source: https://github.com/google-research/google-research/tree/f46ca8374b4cddef97ca4208ad986049d74d296a. Registered source archive SHA-256: d2043b7ed0e7cf551d80686875d8f5de146a4382d1259081a6fe13f876f38253. Selected tasks: mbpp-14, mbpp-34, mbpp-42, mbpp-182, mbpp-273, mbpp-322, mbpp-352, mbpp-398, mbpp-496, mbpp-505. ChatGPT subscription: harbor-codex-subscription; observed chatgpt_subscription / 0.154.0 / max. Claude subscription: harbor-claude-subscription; observed claude_subscription / 2.1.273 / max. DeepSeek: harbor-deepseek; observed deepseek_api / max. 3 concurrent task attempts; 600 seconds and 100 tool calls per task; 7200 seconds per run. Same-provider task queues within this run are outside task clocks. Latency includes environment setup, agent execution and native verification. Subscription cost is unmeasured. Full traces and pinned task/world images are retained in raw_outputs.jsonl. Public text exports remove unrelated subscription-account email identities; their original and public checksums are verified separately. Original private evidence is preserved. 26 original task archives are public; 4 containing account identities remain private. Public raw output retains every task and its recorded trace.

Harness
Harbor · GKE
Primary evaluator
Native tests
Source coverage
10 / 500 tasks

ChatGPT subscription: 600s/task; 100 max turns; no automatic retries. GKE gVisor; images pinned in manifest.

Claude subscription: 600s/task; 100 max turns; no automatic retries. GKE gVisor; images pinned in manifest.

DeepSeek: 600s/task; 100 max turns; no automatic retries. GKE gVisor; images pinned in manifest.

Cost reflects the charges present in the source report; hosting, subscription fees and independent grading may be excluded. Task latency uses the source report’s measurement window. These are not streaming inference measurements.

Inspect the source report and individual attempts ↗
Evaluation protocol and downloads

Measured ten-task sample from 500 public evaluation tasks. One attempt per task with a common terminal scaffold; original tests, RNG seed 0. This is an agentic evaluation, not the original code-only sampling protocol. Source: https://github.com/google-research/google-research/tree/f46ca8374b4cddef97ca4208ad986049d74d296a. Registered source archive SHA-256: d2043b7ed0e7cf551d80686875d8f5de146a4382d1259081a6fe13f876f38253. Selected tasks: mbpp-14, mbpp-34, mbpp-42, mbpp-182, mbpp-273, mbpp-322, mbpp-352, mbpp-398, mbpp-496, mbpp-505. ChatGPT subscription: harbor-codex-subscription; observed chatgpt_subscription / 0.154.0 / max. Claude subscription: harbor-claude-subscription; observed claude_subscription / 2.1.273 / max. DeepSeek: harbor-deepseek; observed deepseek_api / max. 3 concurrent task attempts; 600 seconds and 100 tool calls per task; 7200 seconds per run. Same-provider task queues within this run are outside task clocks. Latency includes environment setup, agent execution and native verification. Subscription cost is unmeasured. Full traces and pinned task/world images are retained in raw_outputs.jsonl. Public text exports remove unrelated subscription-account email identities; their original and public checksums are verified separately. Original private evidence is preserved. 26 original task archives are public; 4 containing account identities remain private. Public raw output retains every task and its recorded trace.

Inside the results

Read the failure analysis, compare model outcomes, and inspect task trajectories and grading criteria.

MBPP · agent evaluation · Open the deep dive →

What this benchmark measures

Metrics

Native task success, errors, latency and recorded cost.

Execution requirements

Blobfish-hosted Harbor harness on GKE; frozen task subset.

Scope and limitations

The published pilot is a subset. Inspect the run protocol before comparing scores.

Evaluation availability

A managed Harbor task package is registered on GKE. Published reports identify the agent configurations and selected tasks.

Read the harness and grading protocol →

Sources and company fit