# DevOpsBench-100 evaluation

Run: `br-8660e046ae07a0ac8eebdb65880598fa`
Dataset version: `3.2.7:revision8:pilot-1`
Frozen manifest SHA-256: `09fa6e0c38eedf4223f9262166bfac29c9497080e2b575374b0474ca9b7ecfb2`
Source SHA-256: `9fc52450e746c2de454a1662a3053f9afa6616ec9f5feea6ca31012ce871e27f`

Selected tasks per scaffold: 10. Registered: 10. Official total: 100.
These results cover the selected tasks and the recorded execution budgets only.

Native, rubric and vision verdicts are separate. Unavailable graders are not passes; a pass rate is reported only with complete grading coverage. No retries are silently substituted.

## ChatGPT subscription

Model: `gpt-6-astra`; scaffold: `harbor-codex-subscription`; scaffold revision: `terminal-mcp-v3`.

| Site | Tasks | Completed | Native passes / graded | Rubric passes / graded | Vision passes / graded | Mean latency |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| aiops_detection | 1 | 1 | 0 / 1 (0.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 369.60s |
| aiops_localization | 1 | 1 | 0 / 1 (0.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 553.93s |
| cross_system | 3 | 3 | 0 / 3 (0.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 481.30s |
| error_rate_reduction | 1 | 1 | 0 / 1 (0.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 533.93s |
| latency_optimization | 1 | 1 | 0 / 1 (0.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 553.02s |
| multi_service_rollout | 2 | 2 | 0 / 2 (0.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 501.03s |
| reconciliation | 1 | 0 | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 222.30s |
| Overall | 10 | 9 | 0 / 9 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 467.87s |

P95 latency: 587.34s. Model cost: not fully measured.
Task budget: 600s; max agent steps: 100.

## Claude subscription

Model: `claude-sonnet-5`; scaffold: `harbor-claude-subscription`; scaffold revision: `terminal-mcp-v3`.

| Site | Tasks | Completed | Native passes / graded | Rubric passes / graded | Vision passes / graded | Mean latency |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| aiops_detection | 1 | 0 | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 602.75s |
| aiops_localization | 1 | 0 | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 602.67s |
| cross_system | 3 | 2 | 0 / 2 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 488.33s |
| error_rate_reduction | 1 | 0 | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 602.67s |
| latency_optimization | 1 | 0 | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 602.72s |
| multi_service_rollout | 2 | 0 | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 602.65s |
| reconciliation | 1 | 0 | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 602.71s |
| Overall | 10 | 2 | 0 / 2 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 568.38s |

P95 latency: 602.75s. Model cost: not fully measured.
Task budget: 600s; max agent steps: 100.

## DeepSeek

Model: `deepseek-v4-pro`; scaffold: `harbor-deepseek`; scaffold revision: `terminal-mcp-v3`.

| Site | Tasks | Completed | Native passes / graded | Rubric passes / graded | Vision passes / graded | Mean latency |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| aiops_detection | 1 | 1 | 0 / 1 (0.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 558.95s |
| aiops_localization | 1 | 0 | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 602.71s |
| cross_system | 3 | 1 | 0 / 1 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 550.77s |
| error_rate_reduction | 1 | 0 | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 602.72s |
| latency_optimization | 1 | 1 | 0 / 1 (0.0%) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 577.29s |
| multi_service_rollout | 2 | 1 | 0 / 1 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 571.11s |
| reconciliation | 1 | 0 | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 602.60s |
| Overall | 10 | 4 | 0 / 4 (incomplete grading) | 0 / 0 (incomplete grading) | 0 / 0 (incomplete grading) | 573.88s |

P95 latency: 602.72s. Model cost: not fully measured.
Task budget: 600s; max agent steps: 100.

