Browser automation / Benchmark leaderboard

Online-Mind2Web

300 real-world tasks spanning 136 websites; maintainers update invalid tasks.

Harbor · GKE · DOM rubric

Online-Mind2Web · Codex + Playwright

View full report ↗

3/10 tasks passed the DOM rubric; 3 execution errors. The original GitHub snapshot, trajectories, screenshots and separate vision verdicts are retained.

Single-agent evaluation10 selected / 300 source tasksDOM rubric2026-09-16

Task success

30.0%Codex + PlaywrightAll selected tasks stay in the denominator

Median task latency

134.0sCodex + PlaywrightRecorded task time; scope in the protocol

Mean reported cost / task

Not measuredNo complete measurement availableReported usage only; see excluded charges
Compare agents

PERFORMANCE

Task success

Higher is better · % of selected tasks

Codex + Playwright30.0%

gpt-6-astra3 pass · 3 errors

A task must pass the declared primary grader. Completed but ungraded or uncollected tasks prevent a success ranking.

TRADEOFFS

Success vs. latency

More success, less time

0%25%50%75%100%0.0s38.9s77.7s116.6s155.4sCodex + Playwright · gpt-6-astra: 30.0%; 134.0s
Codex + Playwright

One measured agent; additional matched runs are needed for a comparison. This evaluation covers a subset of the source benchmark.

RELIABILITY

What happened to every task

Every selected task is accounted for

PassedFailedExecution errorUngraded / missing
Codex + Playwright10/10 recorded

Execution errors count against task success. They are not presented as a grader verdict. Supplemental vision and provider judgments remain in the full report.

LATENCY DISTRIBUTION

Typical time and the long tail

Minimum → median → P95 → maximum

Codex + Playwright10/10 timed

Median 134.0sP95 335.4s

P95 uses the nearest ranked observed task. Small selections can have a P95 equal to their slowest attempt. Missing timing records are disclosed.

ALL COMPARISON METRICS

Inspect the numbers.

Online-Mind2Web · Codex + Playwright · 10 selected tasks · 2026-09-16
RankAgent / modelErrorsEvidence
Codex + Playwrightgpt-6-astra30.0%3/10 passed70.0%Execution, irrespective of grade134.0s10/10 measured335.4s10/10 measuredNot measured0/10 reported30 missing · 0 ungradedTask results ↗

No competitive ranks: a ranking requires multiple agents on the identical tasks, budget, environment and grading protocol. Missing measurements are shown explicitly and sort last.

Measurement definitions, configuration & protocol

Imported recorded evidence; not independently rerun by the report engine. Scaffold: harbor-playwright-dom-v1 at Harbor 0.21.0; Playwright 1.58.0; Codex CLI 0.154.0. Unknown measurements remain unknown. Fresh headless Chromium per task. Independent DOM rubric primary; at most eight uniformly spaced screenshots including first and last for supplemental vision. Native task-success grading is unconfigured. This is not the official WebJudge protocol or human adjudication. Solver and independent graders use the same model family; dollar costs are unmeasured. Task latency includes browser/solver setup and execution. Per-grader latency and total duration including grading are retained separately in raw outputs. Partial trajectories preserve original execution errors without rescoring.

Harness
Harbor · GKE
Primary evaluator
DOM rubric
Source coverage
10 / 300 tasks

Codex + Playwright: One attempt; 300 solver seconds; 60 browser actions; one task at a time; 7200s Job deadline. GKE bf-benchmarks; fresh gVisor Chromium per task; no model credentials in browser guest.

Cost reflects the charges present in the source report; hosting, subscription fees and independent grading may be excluded. Task latency uses the source report’s measurement window. These are not streaming inference measurements.

Inspect the source report and individual attempts ↗
Evaluation protocol and downloads

Imported recorded evidence; not independently rerun by the report engine. Scaffold: harbor-playwright-dom-v1 at Harbor 0.21.0; Playwright 1.58.0; Codex CLI 0.154.0. Unknown measurements remain unknown. Fresh headless Chromium per task. Independent DOM rubric primary; at most eight uniformly spaced screenshots including first and last for supplemental vision. Native task-success grading is unconfigured. This is not the official WebJudge protocol or human adjudication. Solver and independent graders use the same model family; dollar costs are unmeasured. Task latency includes browser/solver setup and execution. Per-grader latency and total duration including grading are retained separately in raw outputs. Partial trajectories preserve original execution errors without rescoring.

Inside the results

Read the failure analysis, compare model outcomes, and inspect task trajectories and grading criteria.

Online-Mind2Web (GitHub snapshot pilot) · agent evaluation · Open the deep dive →

What this benchmark measures

Metrics

Task success under the pinned evaluation protocol; human review or specified WebJudge.

Execution requirements

Live browser, action/screenshot trace and dated task snapshot.

Scope and limitations

Pin updated task IDs and evaluator. It is different from offline Mind2Web and Mind2Web 2.

Evaluation availability

Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.

Sources and company fit

Companies whose products may fit

Research recommendations based on product capabilities. These companies have not necessarily run this benchmark or integrated with Blobfish.

Compatibility notes for each company

Browser Use: Capability-aligned; adapter/access to qualify

Browserbase / Stagehand: Capability-aligned; adapter/access to qualify

Firecrawl: Composite system

Microsoft Playwright: Composite system

Skyvern: Capability-aligned; adapter/access to qualify

Steel: Composite system

TinyFish: Capability-aligned; adapter/access to qualify