Completed attempts clear the tests.
30 of 30 selected attempts passed the declared primary evaluation. 0 ended with an execution error or timeout. Every selected attempt stays in the report.
mbpp / Benchmark deep dive
Compare the outcomes, understand the gaps, and inspect the work behind every score.
30 of 30 selected attempts passed the declared primary evaluation. 0 ended with an execution error or timeout. Every selected attempt stays in the report.
10 selected tasks per model cover a small part of 500 benchmark tasks. One attempt per task measures this run; repeated-run consistency and full-suite leadership are not established.
Dollar cost is recorded for 0 of 30 attempts. Subscription usage and missing provider billing are not converted into an invented per-task price. Latency includes the scope declared in the protocol below.
Scores use the same selected tasks and the report’s declared primary evaluator.
| Model / agent | Success | Execution errors | Mean latency | Cost / task |
|---|---|---|---|---|
| ChatGPT subscriptiongpt-6-astra | 100.0%10/10 selected tasks | 0 | 70.9s | Not reported |
| Claude subscriptionclaude-sonnet-5 | 100.0%10/10 selected tasks | 0 | 48.2s | Not reported |
| DeepSeekdeepseek-v4-pro | 100.0%10/10 selected tasks | 0 | 43.2s | Not reported |
Pilot coverage: 10/500 source tasks. Small fixed selections do not establish a full-benchmark ranking. Missing grades remain visible.
Execution failure and an incorrect result are separate outcomes.
No failing task was recorded in this selected sample.
No failing task was recorded in this selected sample.
No failing task was recorded in this selected sample.
The work behind the score
Original instructions, observed actions, and the verdict for each attempt.
gpt-6-astra · 62.3s · 2 recorded actions
Write a complete Python solution to /logs/artifacts/answer.py for the following task. Write a python function to find the volume of a triangular prism. Tests supplied by the original MBPP prompt: assert find_Volume(10,8,6) == 240 assert find_Volume(3,2,2) == 6 assert find_Volume(1,2,1) == 1
Recorded benchmark evaluator · f46ca8374b4cddef97ca4208ad986049d74d296a:original-tests-seed0
reward=1; pass threshold=1.0
Recorded benchmark evaluator · none
Rubric grading is not configured for this task.
Recorded benchmark evaluator · none
Vision grading is not configured for this task.
Reproducibility
Measured ten-task sample from 500 public evaluation tasks. One attempt per task with a common terminal scaffold; original tests, RNG seed 0. This is an agentic evaluation, not the original code-only sampling protocol. Source: https://github.com/google-research/google-research/tree/f46ca8374b4cddef97ca4208ad986049d74d296a. Registered source archive SHA-256: d2043b7ed0e7cf551d80686875d8f5de146a4382d1259081a6fe13f876f38253. Selected tasks: mbpp-14, mbpp-34, mbpp-42, mbpp-182, mbpp-273, mbpp-322, mbpp-352, mbpp-398, mbpp-496, mbpp-505. ChatGPT subscription: harbor-codex-subscription; observed chatgpt_subscription / 0.154.0 / max. Claude subscription: harbor-claude-subscription; observed claude_subscription / 2.1.273 / max. DeepSeek: harbor-deepseek; observed deepseek_api / max. 3 concurrent task attempts; 600 seconds and 100 tool calls per task; 7200 seconds per run. Same-provider task queues within this run are outside task clocks. Latency includes environment setup, agent execution and native verification. Subscription cost is unmeasured. Full traces and pinned task/world images are retained in raw_outputs.jsonl. Public text exports remove unrelated subscription-account email identities; their original and public checksums are verified separately. Original private evidence is preserved. 26 original task archives are public; 4 containing account identities remain private. Public raw output retains every task and its recorded trace.