← Benchmark leaderboardAll leaderboards

mbpp / Benchmark deep dive

MBPP: inside the results.

Compare the outcomes, understand the gaps, and inspect the work behind every score.

10selected tasks per model
3model configurations
30/30primary-verified attempts
77recorded tool actions

What the results tell us.

01

Completed attempts clear the tests.

30 of 30 selected attempts passed the declared primary evaluation. 0 ended with an execution error or timeout. Every selected attempt stays in the report.

02

Read the tie in context.

10 selected tasks per model cover a small part of 500 benchmark tasks. One attempt per task measures this run; repeated-run consistency and full-suite leadership are not established.

03

Cost is still an open measurement.

Dollar cost is recorded for 0 of 30 attempts. Subscription usage and missing provider billing are not converted into an invented per-task price. Latency includes the scope declared in the protocol below.

Performance in context.

Scores use the same selected tasks and the report’s declared primary evaluator.

Native executable task success · one attempt per task
Model / agentSuccessExecution errorsMean latencyCost / task
ChatGPT subscriptiongpt-6-astra100.0%10/10 selected tasks
070.9sNot reported
Claude subscriptionclaude-sonnet-5100.0%10/10 selected tasks
048.2sNot reported
DeepSeekdeepseek-v4-pro100.0%10/10 selected tasks
043.2sNot reported

Pilot coverage: 10/500 source tasks. Small fixed selections do not establish a full-benchmark ranking. Missing grades remain visible.

Where attempts fall short.

Execution failure and an incorrect result are separate outcomes.

ChatGPT subscription

Verified successes
10
Graded failures
0
Execution errors / timeouts
0
Completed, ungraded
0

No failing task was recorded in this selected sample.

Claude subscription

Verified successes
10
Graded failures
0
Execution errors / timeouts
0
Completed, ungraded
0

No failing task was recorded in this selected sample.

DeepSeek

Verified successes
10
Graded failures
0
Execution errors / timeouts
0
Completed, ungraded
0

No failing task was recorded in this selected sample.

The work behind the score

Inspect a task.

Original instructions, observed actions, and the verdict for each attempt.

Pass

gpt-6-astra · 62.3s · 2 recorded actions

Original task

Write a complete Python solution to /logs/artifacts/answer.py for the following task. Write a python function to find the volume of a triangular prism. Tests supplied by the original MBPP prompt: assert find_Volume(10,8,6) == 240 assert find_Volume(3,2,2) == 6 assert find_Volume(1,2,1) == 1

Grading criteria & verdicts

Executable checksPass

Recorded benchmark evaluator · f46ca8374b4cddef97ca4208ad986049d74d296a:original-tests-seed0

reward=1; pass threshold=1.0

Rubric reviewUngraded

Recorded benchmark evaluator · none

Rubric grading is not configured for this task.

Screenshot reviewUngraded

Recorded benchmark evaluator · none

Vision grading is not configured for this task.

Link to this attempt ↗
Source checksums

06861430b4d888f0393fdf928b70b74520e1bcc8e6681b64ebf047524ce1bd1b.tar.gz
3c398669609a9862ce00f5ab18758f017aca4c9a4a66b1bbc271a9846509f898

raw_outputs.jsonl
70352566c67b322e5d50702edd192a8ed06c5d4085fc8e65468e34ede11af423

Reproducibility

The evaluation protocol.

Measured ten-task sample from 500 public evaluation tasks. One attempt per task with a common terminal scaffold; original tests, RNG seed 0. This is an agentic evaluation, not the original code-only sampling protocol. Source: https://github.com/google-research/google-research/tree/f46ca8374b4cddef97ca4208ad986049d74d296a. Registered source archive SHA-256: d2043b7ed0e7cf551d80686875d8f5de146a4382d1259081a6fe13f876f38253. Selected tasks: mbpp-14, mbpp-34, mbpp-42, mbpp-182, mbpp-273, mbpp-322, mbpp-352, mbpp-398, mbpp-496, mbpp-505. ChatGPT subscription: harbor-codex-subscription; observed chatgpt_subscription / 0.154.0 / max. Claude subscription: harbor-claude-subscription; observed claude_subscription / 2.1.273 / max. DeepSeek: harbor-deepseek; observed deepseek_api / max. 3 concurrent task attempts; 600 seconds and 100 tool calls per task; 7200 seconds per run. Same-provider task queues within this run are outside task clocks. Latency includes environment setup, agent execution and native verification. Subscription cost is unmeasured. Full traces and pinned task/world images are retained in raw_outputs.jsonl. Public text exports remove unrelated subscription-account email identities; their original and public checksums are verified separately. Original private evidence is preserved. 26 original task archives are public; 4 containing account identities remain private. Public raw output retains every task and its recorded trace.

Task selection
10 frozen tasks, one attempt for each model. f46ca8374b4c:pilot-1:28cbcf6c4d6a792dd22b7441b347c709e292c4998ec7e3bb9e4e2a97d1ca97eb
Tools & environment
GKE gVisor; images pinned in manifest. Provider driver versions are recorded; matching task tools do not imply identical provider implementations.
Budget
600s/task; 100 max turns; no automatic retries
Independent grading
Native executable checks in this original publication. Astra supplements are separately identified when available.. Original grader versions and missing verdicts are preserved.
Evidence
Task prompts, tool inputs, observed results, final answers and grader receipts. Each trajectory links to its original artifact checksums.
Limits
One attempt per task. No pass@k, repeated-run reliability or expert agreement claim. Unreported costs stay unknown. Browser Use’s managed agent belongs to its own harness cohort.
Open the original report and downloads
MBPP · agent evaluation3 runs · Published evaluation