Browser automation / Benchmark leaderboard

WebVoyager

End-to-end navigation and task completion on live websites; original release has 643 tasks across 15 sites.

Browser Use Cloud · GKE vision review · Rubric review

WebVoyager · Browser Use v4

View full report ↗

15 original tasks across 15 websites. Separate evidence review supports 10 outcomes; every failure and output is retained.

15 matching tasks · 643 source tasks · 2026-09-16
RankAgent / modelSuccessPassedErrorsAvg. latencyCost / taskReport
Browser Use v4gpt-5.6-luna66.7%10/150117.9s$0.039Open ↗

Single-agent pilot; no competitive rank. These are selected-task results, not full-benchmark scores. Missing cost measurements remain unreported.

Evaluation protocol and downloads

Imported recorded evidence; not independently rerun by the report engine. Scaffold: browser-use-cloud-v4 at SDK 3.11.3. Unknown measurements remain unknown. 15/643 tasks; one per website, seed 20260915. Separate Codex AI evidence review, not official GPT-4V grading or human adjudication; review model version was not retained. Unverified requirements receive no success credit under the recorded rubric. A later, separate Astra vision review runs on GKE using only the final screenshot and answer. Its restricted evidence scope is supplemental and does not replace the original rubric primary. Vision latency and subscription token usage are recorded separately; no dollar cost is inferred. Provider trajectory judgments are retained separately. Latency measures create-to-terminal agent time. Cost is final provider-reported run cost including asynchronous judge updates; browser hosting is separate. Original fixed dates were preserved.

What this benchmark measures

Metrics

Task success; supplemented with cost per success and elapsed time.

Execution requirements

Browser agent with live-web access and recorded trajectories.

Scope and limitations

Website drift, blocked pages and judge configuration affect comparability; distinguish the dataset from its reference agent.

Evaluation availability

Published browser evaluations retain the original tasks, tool evidence, grader receipts and measured costs. Each report identifies its browser harness and model configuration.

Sources and company fit

Companies whose products may fit

Research recommendations based on product capabilities. These companies have not necessarily run this benchmark or integrated with Blobfish.

Compatibility notes for each company

Browser Use: Capability-aligned; adapter/access to qualify

Browserbase / Stagehand: Capability-aligned; adapter/access to qualify

Firecrawl: Composite system

Microsoft Playwright: Composite system

Skyvern: Capability-aligned; adapter/access to qualify

Steel: Composite system

TinyFish: Capability-aligned; adapter/access to qualify