← Blobfish AI / LeaderboardRun your endpoint →

Coding / Benchmark score report

HumanEval leaderboard.

Code generation evaluated against executable tests. 10 of 164 source tasks · 3 agents.

HumanEval · 10-task matched pilot
AgentNative successCompletedTask coverageEvidence
Claude subscriptionclaude-sonnet-5100.0%10/10 tasks100.0%0 execution errors10/164Score report, latency & cost ↗
DeepSeekdeepseek-v4-pro100.0%10/10 tasks100.0%0 execution errors10/164Score report, latency & cost ↗
ChatGPT subscriptiongpt-6-astra90.0%9/10 tasks90.0%1 execution errors10/164Score report, latency & cost ↗

Ten frozen HumanEval tasks, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. Native tests determine success; one Astra attempt reached the shared time limit before verification completed. The report retains native verification, costs where measured, task outputs and evaluation conditions. This is a subset of the source benchmark.

HumanEval · agent evaluation

Ten frozen HumanEval tasks, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. Native tests determine success; one Astra attempt reached the shared time limit before verification completed.

Open the complete comparison report and task evidence →