Coding / Benchmark score report
HumanEval leaderboard.
Code generation evaluated against executable tests. 10 of 164 source tasks · 3 agents.
| Agent | Native success | Completed | Task coverage | Evidence |
|---|---|---|---|---|
| Claude subscriptionclaude-sonnet-5 | 100.0%10/10 tasks | 100.0%0 execution errors | 10/164 | Score report, latency & cost ↗ |
| DeepSeekdeepseek-v4-pro | 100.0%10/10 tasks | 100.0%0 execution errors | 10/164 | Score report, latency & cost ↗ |
| ChatGPT subscriptiongpt-6-astra | 90.0%9/10 tasks | 90.0%1 execution errors | 10/164 | Score report, latency & cost ↗ |
Ten frozen HumanEval tasks, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. Native tests determine success; one Astra attempt reached the shared time limit before verification completed. The report retains native verification, costs where measured, task outputs and evaluation conditions. This is a subset of the source benchmark.
HumanEval · agent evaluation
Ten frozen HumanEval tasks, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. Native tests determine success; one Astra attempt reached the shared time limit before verification completed.
Open the complete comparison report and task evidence →