← Blobfish AI / LeaderboardRun your endpoint →

Coding / Benchmark score report

MBPP leaderboard.

Python programming tasks with source-defined tests. 10 of 500 source tasks · 3 agents.

MBPP · 10-task matched pilot
AgentNative successCompletedTask coverageEvidence
ChatGPT subscriptiongpt-6-astra100.0%10/10 tasks100.0%0 execution errors10/500Score report, latency & cost ↗
Claude subscriptionclaude-sonnet-5100.0%10/10 tasks100.0%0 execution errors10/500Score report, latency & cost ↗
DeepSeekdeepseek-v4-pro100.0%10/10 tasks100.0%0 execution errors10/500Score report, latency & cost ↗

Ten matching tasks from the official MBPP test split, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. All three passed the original native tests on these selected tasks. The report retains native verification, costs where measured, task outputs and evaluation conditions. This is a subset of the source benchmark.

MBPP · agent evaluation

Ten matching tasks from the official MBPP test split, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. All three passed the original native tests on these selected tasks.

Open the complete comparison report and task evidence →