Skip to content

Browser automation

WebBench-Interact

Web agents given one-line human requests, each needing one UI mechanic (52 distinct: sliders, native date and time fields, iframes, shadow DOM, hover and context menus, blocking dialogs, virtualised lists, drag and drop, typeahead, canvas controls, sign-in …); the server checks the state the agent submitted.

Browser automation

Model performance

1 recorded configuration · 52 source tasks

Updated
Higher is better · 0–100 · execution errors score 0
T
TinyFishWeb Agent API
73.1 First-time pass40 / 52 strict passes · 50 graded · 2 execution errors
Full benchmark52 tasks attempted
View run
Run details
Evaluation
WebBench-Interact · TinyFish Web Agent API
Environment
webbench.interact_bench v1 (hosted protocol, public tun…
Grader
WebBench suite scorer
Execution errors
2
Coverage
52 of 52 source tasks attempted; 50 graded.
Open the run record (per-task verdicts)

First-time pass is the mean per-task score from the release's own verifier over the full roster. Strict passes are that verifier's all-or-nothing verdicts (version under Run details); the per-task verdicts are in the run record under Run details.

About WebBench-Interact

What it measures

First-time pass (score), tasks passed (final submission equals the target state), rejected submissions, steps, wall time. No judge.

Scope & limitations

Each task isolates one mechanic; long multi-application workflows are the WebBench suite itself.

Evaluation environment

One page per challenge served through a throw-away public tunnel to the vendor's hosted agent; one attempt per task, provider defaults, 300 s budget.

Source

Original benchmark

Dataset & task catalog ↗