PERFORMANCE
Task success
Higher is better · % of selected tasks
A task must pass the declared primary grader. Completed but ungraded or uncollected tasks prevent a success ranking.
Browser automation / Benchmark leaderboard
300 real-world tasks spanning 136 websites; maintainers update invalid tasks.
3/10 tasks passed the DOM rubric; 3 execution errors. The original GitHub snapshot, trajectories, screenshots and separate vision verdicts are retained.
Task success
30.0%Codex + PlaywrightAll selected tasks stay in the denominatorMedian task latency
134.0sCodex + PlaywrightRecorded task time; scope in the protocolMean reported cost / task
Not measuredNo complete measurement availableReported usage only; see excluded chargesPERFORMANCE
Higher is better · % of selected tasks
A task must pass the declared primary grader. Completed but ungraded or uncollected tasks prevent a success ranking.
TRADEOFFS
More success, less time
One measured agent; additional matched runs are needed for a comparison. This evaluation covers a subset of the source benchmark.
RELIABILITY
Every selected task is accounted for
Execution errors count against task success. They are not presented as a grader verdict. Supplemental vision and provider judgments remain in the full report.
LATENCY DISTRIBUTION
Minimum → median → P95 → maximum
Median 134.0sP95 335.4s
P95 uses the nearest ranked observed task. Small selections can have a P95 equal to their slowest attempt. Missing timing records are disclosed.
ALL COMPARISON METRICS
| Rank | Agent / model | Errors | Evidence | |||||
|---|---|---|---|---|---|---|---|---|
| — | Codex + Playwrightgpt-6-astra | 30.0%3/10 passed | 70.0%Execution, irrespective of grade | 134.0s10/10 measured | 335.4s10/10 measured | Not measured0/10 reported | 30 missing · 0 ungraded | Task results ↗ |
No competitive ranks: a ranking requires multiple agents on the identical tasks, budget, environment and grading protocol. Missing measurements are shown explicitly and sort last.
Imported recorded evidence; not independently rerun by the report engine. Scaffold: harbor-playwright-dom-v1 at Harbor 0.21.0; Playwright 1.58.0; Codex CLI 0.154.0. Unknown measurements remain unknown. Fresh headless Chromium per task. Independent DOM rubric primary; at most eight uniformly spaced screenshots including first and last for supplemental vision. Native task-success grading is unconfigured. This is not the official WebJudge protocol or human adjudication. Solver and independent graders use the same model family; dollar costs are unmeasured. Task latency includes browser/solver setup and execution. Per-grader latency and total duration including grading are retained separately in raw outputs. Partial trajectories preserve original execution errors without rescoring.
Codex + Playwright: One attempt; 300 solver seconds; 60 browser actions; one task at a time; 7200s Job deadline. GKE bf-benchmarks; fresh gVisor Chromium per task; no model credentials in browser guest.
Cost reflects the charges present in the source report; hosting, subscription fees and independent grading may be excluded. Task latency uses the source report’s measurement window. These are not streaming inference measurements.
Inspect the source report and individual attempts ↗Imported recorded evidence; not independently rerun by the report engine. Scaffold: harbor-playwright-dom-v1 at Harbor 0.21.0; Playwright 1.58.0; Codex CLI 0.154.0. Unknown measurements remain unknown. Fresh headless Chromium per task. Independent DOM rubric primary; at most eight uniformly spaced screenshots including first and last for supplemental vision. Native task-success grading is unconfigured. This is not the official WebJudge protocol or human adjudication. Solver and independent graders use the same model family; dollar costs are unmeasured. Task latency includes browser/solver setup and execution. Per-grader latency and total duration including grading are retained separately in raw outputs. Partial trajectories preserve original execution errors without rescoring.
Read the failure analysis, compare model outcomes, and inspect task trajectories and grading criteria.
Online-Mind2Web (GitHub snapshot pilot) · agent evaluation · Open the deep dive →Task success under the pinned evaluation protocol; human review or specified WebJudge.
Live browser, action/screenshot trace and dated task snapshot.
Pin updated task IDs and evaluator. It is different from offline Mind2Web and Mind2Web 2.
Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.
Research recommendations based on product capabilities. These companies have not necessarily run this benchmark or integrated with Blobfish.
Browser Use: Capability-aligned; adapter/access to qualify
Browserbase / Stagehand: Capability-aligned; adapter/access to qualify
Firecrawl: Composite system
Microsoft Playwright: Composite system
Skyvern: Capability-aligned; adapter/access to qualify
Steel: Composite system
TinyFish: Capability-aligned; adapter/access to qualify