Blobfish-hosted evaluations

The harness behind the results.

We run frozen task packages in isolated GKE environments, grade the outcome, and retain the evidence. Published runs and harness qualification are available here.

3registered native benchmark packages
4published evaluation reports
Harbor + HUDsupported task package formats

Registry observed 2026-09-16 04:23 UTC. This is a dated record, not a live availability check. Read the current registry ↗

Measured company / agent runs

Published results and reports

Agent scores use their declared grading protocol. Native, rubric and vision outcomes remain separate.

Harbor · GKE · Native tests

HumanEval · agent evaluation

View full report ↗

Ten frozen HumanEval tasks, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. Native tests determine success; one Astra attempt reached the shared time limit before verification completed.

10 matching tasks · 164 source tasks · 2026-09-16
RankAgent / modelSuccessPassedErrorsAvg. latencyCost / taskReport
1Claude subscriptionclaude-sonnet-5100.0%10/10048.1sNot reportedOpen ↗
1DeepSeekdeepseek-v4-pro100.0%10/10049.9sNot reportedOpen ↗
3ChatGPT subscriptiongpt-6-astra90.0%9/101135.8sNot reportedOpen ↗

Ranks apply only to this matched task set. Tied scores share a rank. These are selected-task results, not full-benchmark scores. Missing cost measurements remain unreported.

Evaluation protocol and downloads

Measured ten-task sample from 164 public evaluation tasks. One attempt per task with a common terminal scaffold; original tests, RNG seed 0. This is an agentic evaluation, not the original code-only sampling protocol. Source: https://github.com/openai/human-eval/tree/6d43fb980f9fee3c892a914eda09951f772ad10d. Registered source archive SHA-256: f373485244ee5745ff98af9f9e20ff824f5b4b68bb1f37a61dd66e20e2cb4581. Selected tasks: humaneval-2, humaneval-39, humaneval-48, humaneval-50, humaneval-61, humaneval-64, humaneval-75, humaneval-79, humaneval-100, humaneval-136. ChatGPT subscription: harbor-codex-subscription; observed chatgpt_subscription / 0.154.0 / max. Claude subscription: harbor-claude-subscription; observed claude_subscription / 2.1.273 / max. DeepSeek: harbor-deepseek; observed deepseek_api / max. 3 concurrent task attempts; 600 seconds and 100 terminal tool calls per task. Account queues are outside task clocks. Latency includes environment setup, agent execution and native verification. Subscription cost is unmeasured. Full traces and pinned task/world images are retained in raw_outputs.jsonl.

Harbor · GKE · Native tests

MBPP · agent evaluation

View full report ↗

Ten matching tasks from the official MBPP test split, attempted once by ChatGPT and Claude through their subscription CLIs and by DeepSeek through its API. All three passed the original native tests on these selected tasks.

10 matching tasks · 500 source tasks · 2026-09-16
RankAgent / modelSuccessPassedErrorsAvg. latencyCost / taskReport
1ChatGPT subscriptiongpt-6-astra100.0%10/10070.9sNot reportedOpen ↗
1Claude subscriptionclaude-sonnet-5100.0%10/10048.2sNot reportedOpen ↗
1DeepSeekdeepseek-v4-pro100.0%10/10043.2sNot reportedOpen ↗

Ranks apply only to this matched task set. Tied scores share a rank. These are selected-task results, not full-benchmark scores. Missing cost measurements remain unreported.

Evaluation protocol and downloads

Measured ten-task sample from 500 public evaluation tasks. One attempt per task with a common terminal scaffold; original tests, RNG seed 0. This is an agentic evaluation, not the original code-only sampling protocol. Source: https://github.com/google-research/google-research/tree/f46ca8374b4cddef97ca4208ad986049d74d296a. Registered source archive SHA-256: d2043b7ed0e7cf551d80686875d8f5de146a4382d1259081a6fe13f876f38253. Selected tasks: mbpp-14, mbpp-34, mbpp-42, mbpp-182, mbpp-273, mbpp-322, mbpp-352, mbpp-398, mbpp-496, mbpp-505. ChatGPT subscription: harbor-codex-subscription; observed chatgpt_subscription / 0.154.0 / max. Claude subscription: harbor-claude-subscription; observed claude_subscription / 2.1.273 / max. DeepSeek: harbor-deepseek; observed deepseek_api / max. 3 concurrent task attempts; 600 seconds and 100 tool calls per task; 7200 seconds per run. Same-provider task queues within this run are outside task clocks. Latency includes environment setup, agent execution and native verification. Subscription cost is unmeasured. Full traces and pinned task/world images are retained in raw_outputs.jsonl. Public text exports remove unrelated subscription-account email identities; their original and public checksums are verified separately. Original private evidence is preserved. 26 original task archives are public; 4 containing account identities remain private. Public raw output retains every task and its recorded trace.

Browser Use Cloud · GKE vision review · Rubric review

WebVoyager · Browser Use v4

View full report ↗

15 original tasks across 15 websites. Separate evidence review supports 10 outcomes; every failure and output is retained.

15 matching tasks · 643 source tasks · 2026-09-16
RankAgent / modelSuccessPassedErrorsAvg. latencyCost / taskReport
Browser Use v4gpt-5.6-luna66.7%10/150117.9s$0.039Open ↗

Single-agent pilot; no competitive rank. These are selected-task results, not full-benchmark scores. Missing cost measurements remain unreported.

Evaluation protocol and downloads

Imported recorded evidence; not independently rerun by the report engine. Scaffold: browser-use-cloud-v4 at SDK 3.11.3. Unknown measurements remain unknown. 15/643 tasks; one per website, seed 20260915. Separate Codex AI evidence review, not official GPT-4V grading or human adjudication; review model version was not retained. Unverified requirements receive no success credit under the recorded rubric. A later, separate Astra vision review runs on GKE using only the final screenshot and answer. Its restricted evidence scope is supplemental and does not replace the original rubric primary. Vision latency and subscription token usage are recorded separately; no dollar cost is inferred. Provider trajectory judgments are retained separately. Latency measures create-to-terminal agent time. Cost is final provider-reported run cost including asynchronous judge updates; browser hosting is separate. Original fixed dates were preserved.

Harbor · GKE · DOM rubric

Online-Mind2Web · Codex + Playwright

View full report ↗

3/10 tasks passed the DOM rubric; 3 execution errors. The original GitHub snapshot, trajectories, screenshots and separate vision verdicts are retained.

10 matching tasks · 300 source tasks · 2026-09-16
RankAgent / modelSuccessPassedErrorsAvg. latencyCost / taskReport
Codex + Playwrightgpt-6-astra30.0%3/103182.9sNot reportedOpen ↗

Single-agent pilot; no competitive rank. These are selected-task results, not full-benchmark scores. Missing cost measurements remain unreported.

Evaluation protocol and downloads

Imported recorded evidence; not independently rerun by the report engine. Scaffold: harbor-playwright-dom-v1 at Harbor 0.21.0; Playwright 1.58.0; Codex CLI 0.154.0. Unknown measurements remain unknown. Fresh headless Chromium per task. Independent DOM rubric primary; at most eight uniformly spaced screenshots including first and last for supplemental vision. Native task-success grading is unconfigured. This is not the official WebJudge protocol or human adjudication. Solver and independent graders use the same model family; dollar costs are unmeasured. Task latency includes browser/solver setup and execution. Per-grader latency and total duration including grading are retained separately in raw outputs. Partial trajectories preserve original execution errors without rescoring.

Run activity

Pending results and harness failures

Retained status from owned evaluation runs. An incomplete or invalid harness run does not receive a model score.

In progress

DevOpsBench-100

Matched replacement run: ChatGPT, Claude and DeepSeek attempt the same ten original tasks after all three native world-tool controls passed. Scores are pending reviewed publication.

10 selected tasks × 3 agents

Last observed 2026-09-16 04:39:02 UTC
br-7c3aab5279287890579fa7278dd0c911

Cancelled · harness qualification failure

DevOpsBench-100

The original tasks required 97 world tools, but the solver exposed only a terminal. The batch was cancelled, evidence retained, and the tool adapter repaired before the replacement run. This is a harness failure, not a company performance score.

10 selected tasks × 3 agents

Last observed 2026-09-16 03:36:18 UTC
br-68555897ccde9737d8d0f34d4dbf677e

Registered on GKE

Hosted benchmark packages

These prepared subsets are available to the managed runner. Registration does not mean every source task has been evaluated.

Registered · harbor

DevOpsBench-100

Ten matching tasks from the exact published revision 8 (2 source-heldout, 8 source-train). Original main/world images and original deterministic state verifiers. Pilot time/command budgets are recorded per run.

10 prepared tasks / 100 source tasks · Native grading: 10/10

Version: 3.2.7:revision8:pilot-1
Package SHA-256: 09fa6e0c38eedf4223f9262166bfac29c9497080e2b575374b0474ca9b7ecfb2

View benchmark and publication status →
Task selection and supported agents

harbor-terminus-2 · harbor-codex · harbor-claude-code · harbor-codex-subscription · harbor-claude-subscription · harbor-deepseek

dob100-010-detect-status-page-recurrence, dob100-014-localize-checkout-latency, dob100-035-port-close-backlog-issues, dob100-036-port-close-blocked-issues, dob100-049-payments-retry, dob100-062-w6-copy-13, dob100-076-gateway-pool-reuse, dob100-081-backorders, dob100-085-notification-templates, dob100-089-rcn-customer-facing-incidents

Registered · harbor

HumanEval

Measured ten-task sample from 164 public evaluation tasks. One attempt per task with a common terminal scaffold; original tests, RNG seed 0. This is an agentic evaluation, not the original code-only sampling protocol.

10 prepared tasks / 164 source tasks · Native grading: 10/10

Version: 6d43fb980f9f:pilot-1
Package SHA-256: 47adf872a27bf6ae875d4ac4d874e151532770864c2affcaff4fd586d1ca87f6

View benchmark and publication status →
Task selection and supported agents

harbor-terminus-2 · harbor-codex-subscription · harbor-claude-subscription · harbor-deepseek

humaneval-2, humaneval-39, humaneval-48, humaneval-50, humaneval-61, humaneval-64, humaneval-75, humaneval-79, humaneval-100, humaneval-136

Registered · harbor

MBPP

Measured ten-task sample from 500 public evaluation tasks. One attempt per task with a common terminal scaffold; original tests, RNG seed 0. This is an agentic evaluation, not the original code-only sampling protocol.

10 prepared tasks / 500 source tasks · Native grading: 10/10

Version: f46ca8374b4c:pilot-1
Package SHA-256: 28cbcf6c4d6a792dd22b7441b347c709e292c4998ec7e3bb9e4e2a97d1ca97eb

View benchmark and publication status →
Task selection and supported agents

harbor-terminus-2 · harbor-codex-subscription · harbor-claude-subscription · harbor-deepseek

mbpp-14, mbpp-34, mbpp-42, mbpp-182, mbpp-273, mbpp-322, mbpp-352, mbpp-398, mbpp-496, mbpp-505

Harness qualification

Reference solutions pass. Empty solutions fail.

These controls test the harness and its verifiers. They are not agent performance scores.

humaneval

10/10 reference solutions passed · 10/10 no-op solutions rejected

2026-09-16 02:32:55 UTC · Run br-d0b7ecaf6d13fb49bfb18402049097a6

mbpp

10/10 reference solutions passed · 10/10 no-op solutions rejected

2026-09-16 02:38:02 UTC · Run br-26c40d4c46ca67a16d5f1243d0e8e2c5

devops-first-two

2/2 reference solutions passed · 2/2 no-op solutions rejected

2026-09-16 02:36:06 UTC · Run br-60b5f869964198d5da3bb6e3cc773ceb

devops-remaining-eight

8/8 reference solutions passed · 8/8 no-op solutions rejected

2026-09-16 02:39:39 UTC · Run br-19cc7a102f92fa8d08c82a598e14796e

claude-subscription

claude / answer-42: native pass; rubric not configured

2026-09-16 02:31:12 UTC · Run br-9c0a65bdc49ba5b10b7ba28c78e67dd5

codex-rubric

positive / answer-42: native pass; rubric pass
negative / answer-42: native fail; rubric fail

2026-09-16 02:37:52 UTC · Run br-834551ee19d75311b1ecd354e87d1e9e

Execution and qualification provenance

Dedicated cluster: bf-benchmarks · us-central1-a. Task execution and provider credentials are isolated. Subscription caches and immutable run artifacts are retained separately from disposable task containers.

Qualification image: us-central1-docker.pkg.dev/blobfish-ai-429200/cloud-run-source-deploy/bf-benchmark-runner@sha256:c29c39cc466ea083014a702664844a2304a67c1a01a31f29d14233d6ef9a97ca

Observed API image: us-central1-docker.pkg.dev/blobfish-ai-429200/cloud-run-source-deploy/bf-benchmark-runner@sha256:8ac7bb9c1c22f3ae1ee47e476242d4ce0ed71b7e0f3eb1dce6253d9fdee4cdfb

The qualification image records the earlier control runs. It is not inferred from the currently deployed image.

Additional recorded evaluations

Training benchmark results

Historical base-versus-tuned evaluations use their original task sets and training lineage. They are separate from the hosted company pilots above.

Full benchmark evidence ↗
Recorded Qwen3-8B evaluations · source and limitations preserved per row
BenchmarkBaselinePost-trainedChange
τ²-bench retail (Qwen3-8B B4 · paired n=78)38.14%53.42%+15.28 pp
τ²-bench telecom (Qwen3-8B B4 · n=100)17%47%+30 pp
τ²-bench airline (Qwen3-8B B4 · n=50)19.3%23.5%+4.2 pp
τ²-bench three-domain average (Qwen3-8B B4)25.8%41.3%+15.5 pp
BFCL V4 multi-turn (Qwen3-8B B4 · all 800)10.6%30.5%+19.9 pp
BFCL V4 AST + abstention (Qwen3-8B B4 · n=696)82.6%85.9%+3.3 pp
MCP-Mark filesystem + postgres (Qwen3-8B · n=204)3.9%2.9%-1 pp
Qwen3-14B composite-world benchmark batteryNot runNot run
Historical scope and interpretation

Every row except MCP-Mark is a recorded Qwen3-8B B4 composition-corpus run from docs/results.jsonl (2026-07-12..15); B4 used 432 composition worlds. The MCP-Mark row is NOT B4: docs/results.jsonl line 98 records it as the gate_scale adapter (narrow gym worlds, standard suite, fs+pg, k=4), and B4 was never evaluated on MCP-Mark — line 106 states that every prior external null (tau2 parity, BFCL flat, MCP-Mark flat) was measured on the gate_scale lineage and none of them tested composition. Within-row base/tuned comparisons use the same stated harness; cross-row absolutes are not treated as interchangeable, and the MCP-Mark row must not be read as evidence about B4. Qwen3-14B remains null because no 14B training/evaluation run is recorded. The current customer-created /sandbox world has executable validation but has not itself been isolated as the causal training corpus for these results.

τ²-bench retail (Qwen3-8B B4 · paired n=78): Recorded paired B4 result: 95% CI [+6.8,+23.8]pp, p=0.0004. The broader n=100x4 run is ledgered; the paired join contains 78 tasks.

τ²-bench telecom (Qwen3-8B B4 · n=100): Recorded never-trained-domain result with deterministic ENV_ASSERTION reward, p<1e-8. The tuned arm completed 181/200 planned episodes.

τ²-bench airline (Qwen3-8B B4 · n=50): Recorded paired unsteered result; sign-test p=0.327. This is an honest non-significant boundary, not evidence of airline lift.

τ²-bench three-domain average (Qwen3-8B B4): Derived from the recorded retail, telecom, and airline domain results. The base nearly matches the paper's 26.2, but 41.3 remains well below the paper's 61.8 target.

BFCL V4 multi-turn (Qwen3-8B B4 · all 800): Recorded full multi-turn slice using the official data and checker; tuned-only 192 vs base-only 33, McNemar p<0.0001. This is a slice result, not the paper's full BFCL-V4 aggregate.

BFCL V4 AST + abstention (Qwen3-8B B4 · n=696): Recorded same-session comparison with the corrected decoder and identical vLLM harness for both arms.

MCP-Mark filesystem + postgres (Qwen3-8B · n=204): Recorded gate_scale-adapter result (NOT the B4 composition corpus, which was never evaluated on MCP-Mark): postgres was exactly flat (6.0%→6.0%); the combined one-point decrease is floor-level noise across 204 tasks (Fisher p=0.79), not a regression. No MCP-Mark improvement was demonstrated, and none was refuted for B4 — B4 has no MCP-Mark measurement.

Qwen3-14B composite-world benchmark battery: Not run. Configuration and benchmark locks exist, but no recorded 14B checkpoint or evaluation result exists in this repository.