Browser automation / Benchmark leaderboard
WebVoyager
End-to-end navigation and task completion on live websites; original release has 643 tasks across 15 sites.
WebVoyager · Browser Use v4
15 original tasks across 15 websites. Separate evidence review supports 10 outcomes; every failure and output is retained.
| Rank | Agent / model | Success | Passed | Errors | Avg. latency | Cost / task | Report |
|---|---|---|---|---|---|---|---|
| — | Browser Use v4gpt-5.6-luna | 66.7% | 10/15 | 0 | 117.9s | $0.039 | Open ↗ |
Single-agent pilot; no competitive rank. These are selected-task results, not full-benchmark scores. Missing cost measurements remain unreported.
Evaluation protocol and downloads
Imported recorded evidence; not independently rerun by the report engine. Scaffold: browser-use-cloud-v4 at SDK 3.11.3. Unknown measurements remain unknown. 15/643 tasks; one per website, seed 20260915. Separate Codex AI evidence review, not official GPT-4V grading or human adjudication; review model version was not retained. Unverified requirements receive no success credit under the recorded rubric. A later, separate Astra vision review runs on GKE using only the final screenshot and answer. Its restricted evidence scope is supplemental and does not replace the original rubric primary. Vision latency and subscription token usage are recorded separately; no dollar cost is inferred. Provider trajectory judgments are retained separately. Latency measures create-to-terminal agent time. Cost is final provider-reported run cost including asynchronous judge updates; browser hosting is separate. Original fixed dates were preserved.
What this benchmark measures
Metrics
Task success; supplemented with cost per success and elapsed time.
Execution requirements
Browser agent with live-web access and recorded trajectories.
Scope and limitations
Website drift, blocked pages and judge configuration affect comparability; distinguish the dataset from its reference agent.
Evaluation availability
Published browser evaluations retain the original tasks, tool evidence, grader receipts and measured costs. Each report identifies its browser harness and model configuration.
Sources and company fit
- Original benchmark ↗Established
Companies whose products may fit
Research recommendations based on product capabilities. These companies have not necessarily run this benchmark or integrated with Blobfish.
Compatibility notes for each company
Browser Use: Capability-aligned; adapter/access to qualify
Browserbase / Stagehand: Capability-aligned; adapter/access to qualify
Firecrawl: Composite system
Microsoft Playwright: Composite system
Skyvern: Capability-aligned; adapter/access to qualify
Steel: Composite system
TinyFish: Capability-aligned; adapter/access to qualify