Coding / Benchmark score report
MBPP leaderboard.
Python programming tasks with source-defined tests. 10 of 500 source tasks · 3 agents.
| Agent | Native success | Completed | Task coverage | Evidence |
|---|---|---|---|---|
| ChatGPT subscriptiongpt-6-astra | 100.0%10/10 tasks | 100.0%0 execution errors | 10/500 | Score report, latency & cost ↗ |
| Claude subscriptionclaude-sonnet-5 | 100.0%10/10 tasks | 100.0%0 execution errors | 10/500 | Score report, latency & cost ↗ |
| DeepSeekdeepseek-v4-pro | 100.0%10/10 tasks | 100.0%0 execution errors | 10/500 | Score report, latency & cost ↗ |
Ten matching tasks from the official MBPP test split, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. All three passed the original native tests on these selected tasks. The report retains native verification, costs where measured, task outputs and evaluation conditions. This is a subset of the source benchmark.
MBPP · agent evaluation
Ten matching tasks from the official MBPP test split, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. All three passed the original native tests on these selected tasks.
Open the complete comparison report and task evidence →