# HumanEval evaluation

Run: `br-0a0bae89edc579fb1dad8521b6599705`
Dataset version: `6d43fb980f9f:pilot-1`
Frozen manifest SHA-256: `47adf872a27bf6ae875d4ac4d874e151532770864c2affcaff4fd586d1ca87f6`
Source SHA-256: `f373485244ee5745ff98af9f9e20ff824f5b4b68bb1f37a61dd66e20e2cb4581`

Selected tasks per scaffold: 10. Registered: 10. Official total: 164.
These results cover the selected tasks and the recorded execution budgets only.

Native, rubric and vision verdicts are separate. Unavailable graders are not passes; a pass rate is reported only with complete grading coverage. No retries are silently substituted.

## ChatGPT subscription

Model: `gpt-6-astra`; scaffold: `harbor-codex-subscription`; scaffold revision: `terminal-mcp-v1`.

| Site | Tasks | Completed | Native passes / graded | Rubric passes / graded | Vision passes / graded | Mean latency |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| Python | 10 | 9 | 9 / 9 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 135.83s |
| Overall | 10 | 9 | 9 / 9 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 135.83s |

P95 latency: 600.14s. Model cost: not fully measured.
Task budget: 600s; max agent steps: 100.

## Claude subscription

Model: `claude-sonnet-5`; scaffold: `harbor-claude-subscription`; scaffold revision: `terminal-mcp-v1`.

| Site | Tasks | Completed | Native passes / graded | Rubric passes / graded | Vision passes / graded | Mean latency |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| Python | 10 | 10 | 10 / 10 (100.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 48.12s |
| Overall | 10 | 10 | 10 / 10 (100.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 48.12s |

P95 latency: 110.37s. Model cost: not fully measured.
Task budget: 600s; max agent steps: 100.

## DeepSeek

Model: `deepseek-v4-pro`; scaffold: `harbor-deepseek`; scaffold revision: `terminal-mcp-v1`.

| Site | Tasks | Completed | Native passes / graded | Rubric passes / graded | Vision passes / graded | Mean latency |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| Python | 10 | 10 | 10 / 10 (100.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 49.94s |
| Overall | 10 | 10 | 10 / 10 (100.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 49.94s |

P95 latency: 90.69s. Model cost: not fully measured.
Task budget: 600s; max agent steps: 100.

