Leaderboards / WebBench-Interact Browser automation
WebBench-Interact Web agents given one-line human requests, each needing one UI mechanic (52 distinct: sliders, native date and time fields, iframes, shadow DOM, hover and context menus, blocking dialogs, virtualised lists, drag and drop, typeahead, canvas controls, sign-in …); the server checks the state the agent submitted.
73.1 First-time pass 40 / 52 strict passes · 50 graded · 2 execution errors
Full benchmark 52 tasks attempted
View run Run details Evaluation WebBench-Interact · TinyFish Web Agent API Environment webbench.interact_bench v1 (hosted protocol, public tun… Grader WebBench suite scorer Execution errors 2 Coverage 52 of 52 source tasks attempted; 50 graded. Open the run record (per-task verdicts) About WebBench-Interact What it measures First-time pass (score), tasks passed (final submission equals the target state), rejected submissions, steps, wall time. No judge.
Scope & limitations Each task isolates one mechanic; long multi-application workflows are the WebBench suite itself.
Evaluation environment One page per challenge served through a throw-away public tunnel to the vendor's hosted agent; one attempt per task, provider defaults, 300 s budget.
Source Original benchmark Dataset & task catalog ↗