# DevOpsBench-100 evaluation

Report: `DevOpsBench-100 · agent evaluation`
Dataset version: `3.2.7:revision8:pilot-1:09fa6e0c38eedf4223f9262166bfac29c9497080e2b575374b0474ca9b7ecfb2`
Primary evaluator: `native` — Native + configured rubric/vision; 1692eac726f56d596cbfd3a15a5ae670811556900f661efb2974cda3889d340c
Environment: GKE gVisor; images pinned in manifest
Budget: 600s/task; 100 max turns; no automatic retries

Selected tasks per agent: 10. Official total: 100.
These results cover the selected tasks and the recorded execution budgets only. Native, rubric and vision verdicts are separate. Unavailable graders are not passes; a pass rate is reported only with complete grading coverage. No retries are silently substituted.

Protocol: Ten matching tasks from the exact published revision 8 (2 source-heldout, 8 source-train). Original main/world images and original deterministic state verifiers. Pilot time/command budgets are recorded per run. Source: https://blobfish.ai/benchmarks/devopsbench-100. Registered source archive SHA-256: 9fc52450e746c2de454a1662a3053f9afa6616ec9f5feea6ca31012ce871e27f. Selected tasks: dob100-010-detect-status-page-recurrence, dob100-014-localize-checkout-latency, dob100-035-port-close-backlog-issues, dob100-036-port-close-blocked-issues, dob100-049-payments-retry, dob100-062-w6-copy-13, dob100-076-gateway-pool-reuse, dob100-081-backorders, dob100-085-notification-templates, dob100-089-rcn-customer-facing-incidents. ChatGPT subscription: harbor-codex-subscription; observed chatgpt_subscription / 0.154.0 / max. Claude subscription: harbor-claude-subscription; observed claude_subscription / 2.1.273 / max. DeepSeek: harbor-deepseek; observed deepseek_api / max. 2 concurrent task attempts; 600 seconds and 100 tool calls per task; 10800 seconds per run. Same-provider task queues within this run are outside task clocks. Latency includes environment setup, agent execution and native verification. Subscription cost is unmeasured. Full traces and pinned task/world images are retained in raw_outputs.jsonl. Controller timing receipts verify the complete task window for all 30 attempts. Distributed subscription account waits finished before their task clocks started; queue time is recorded separately. Every trusted solver was submitted within its task deadline and retired before releasing its account. The source declares a 2400-second agent budget. This pilot caps setup, solver execution and native verification together at 600 seconds. Eight selected tasks come from the training split and two from the heldout split of the released source. These measurements describe this shorter-budget, single-attempt ten-task pilot. Individual native assertions are available for 15 graded task attempts in the task explorer and native_verifier_details.json, with original archive and verifier-file checksums. Partial rewards are separate from the recorded pass/fail verdict.

## ChatGPT subscription

Model: `gpt-6-astra`; run: `br-8660e046ae07a0ac8eebdb65880598fa:chatgpt`; source: hosted; 2026-09-16T08:02:52.176244+00:00 → 2026-09-16T10:26:15.055087+00:00.

| Site / category | Tasks | Completed | Primary passes | Native passes / graded | Rubric passes / graded | Vision passes / graded | Mean latency |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| aiops_detection | 1 | 1 | 0 (0.0%) | 0 / 1 (0.0%) | — | — | 369.60s |
| aiops_localization | 1 | 1 | 0 (0.0%) | 0 / 1 (0.0%) | — | — | 553.93s |
| cross_system | 3 | 3 | 0 (0.0%) | 0 / 3 (0.0%) | — | — | 481.30s |
| error_rate_reduction | 1 | 1 | 0 (0.0%) | 0 / 1 (0.0%) | — | — | 533.93s |
| latency_optimization | 1 | 1 | 0 (0.0%) | 0 / 1 (0.0%) | — | — | 553.02s |
| multi_service_rollout | 2 | 2 | 0 (0.0%) | 0 / 2 (0.0%) | — | — | 501.03s |
| reconciliation | 1 | 0 | 0 (0.0%) | 0 / 0 (incomplete grading) | — | — | 222.30s |
| Overall | 10 | 9 | 0 (0.0%) | 0 / 9 (incomplete grading) | — | — | 467.87s |

P95 latency: 587.34s (10/10 timed). Model cost: recorded for 0/10 tasks; not fully measured.
Execution errors or timeouts: 1. Completed but ungraded: 0. Not attempted: 0.

Notes: The report preview shortens long outputs. Download raw_outputs.jsonl for complete records. Scaffold harbor-codex-subscription at terminal-mcp-v3.

## Claude subscription

Model: `claude-sonnet-5`; run: `br-8660e046ae07a0ac8eebdb65880598fa:claude`; source: hosted; 2026-09-16T08:02:52.176244+00:00 → 2026-09-16T10:26:15.055087+00:00.

| Site / category | Tasks | Completed | Primary passes | Native passes / graded | Rubric passes / graded | Vision passes / graded | Mean latency |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| aiops_detection | 1 | 0 | 0 (0.0%) | 0 / 0 (incomplete grading) | — | — | 602.75s |
| aiops_localization | 1 | 0 | 0 (0.0%) | 0 / 0 (incomplete grading) | — | — | 602.67s |
| cross_system | 3 | 2 | 0 (0.0%) | 0 / 2 (incomplete grading) | — | — | 488.33s |
| error_rate_reduction | 1 | 0 | 0 (0.0%) | 0 / 0 (incomplete grading) | — | — | 602.67s |
| latency_optimization | 1 | 0 | 0 (0.0%) | 0 / 0 (incomplete grading) | — | — | 602.72s |
| multi_service_rollout | 2 | 0 | 0 (0.0%) | 0 / 0 (incomplete grading) | — | — | 602.65s |
| reconciliation | 1 | 0 | 0 (0.0%) | 0 / 0 (incomplete grading) | — | — | 602.71s |
| Overall | 10 | 2 | 0 (0.0%) | 0 / 2 (incomplete grading) | — | — | 568.38s |

P95 latency: 602.75s (10/10 timed). Model cost: recorded for 0/10 tasks; not fully measured.
Execution errors or timeouts: 8. Completed but ungraded: 0. Not attempted: 0.

Notes: The report preview shortens long outputs. Download raw_outputs.jsonl for complete records. Scaffold harbor-claude-subscription at terminal-mcp-v3.

## DeepSeek

Model: `deepseek-v4-pro`; run: `br-8660e046ae07a0ac8eebdb65880598fa:deepseek`; source: hosted; 2026-09-16T08:02:52.176244+00:00 → 2026-09-16T10:26:15.055087+00:00.

| Site / category | Tasks | Completed | Primary passes | Native passes / graded | Rubric passes / graded | Vision passes / graded | Mean latency |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| aiops_detection | 1 | 1 | 0 (0.0%) | 0 / 1 (0.0%) | — | — | 558.95s |
| aiops_localization | 1 | 0 | 0 (0.0%) | 0 / 0 (incomplete grading) | — | — | 602.71s |
| cross_system | 3 | 1 | 0 (0.0%) | 0 / 1 (incomplete grading) | — | — | 550.77s |
| error_rate_reduction | 1 | 0 | 0 (0.0%) | 0 / 0 (incomplete grading) | — | — | 602.72s |
| latency_optimization | 1 | 1 | 0 (0.0%) | 0 / 1 (0.0%) | — | — | 577.29s |
| multi_service_rollout | 2 | 1 | 0 (0.0%) | 0 / 1 (incomplete grading) | — | — | 571.11s |
| reconciliation | 1 | 0 | 0 (0.0%) | 0 / 0 (incomplete grading) | — | — | 602.60s |
| Overall | 10 | 4 | 0 (0.0%) | 0 / 4 (incomplete grading) | — | — | 573.88s |

P95 latency: 602.72s (10/10 timed). Model cost: recorded for 0/10 tasks; not fully measured.
Execution errors or timeouts: 6. Completed but ungraded: 0. Not attempted: 0.

Notes: The report preview shortens long outputs. Download raw_outputs.jsonl for complete records. Scaffold harbor-deepseek at terminal-mcp-v3.

