Skip to content

Web data extraction

WebBench-Fetch

Fetch APIs on generated pages behind real-web friction (heavy chrome, consent overlays, client-side rendering, late content, tables): does the returned text contain the page's content, and how much else?

Web data extraction

Model performance

3 recorded configurations · 96 source tasks

Updated
Higher is better · 0–100 · execution errors score 0
T
TinyFishFetch API
48.3 Content F146 / 96 strict passes · 96 graded · 0 execution errors
Full benchmark96 tasks attempted
View run
Run details
Evaluation
WebBench-Fetch · TinyFish Fetch API
Environment
webbench.fetch_bench (hosted protocol, public tunnel)
Grader
WebBench suite scorer
Coverage
96 of 96 source tasks attempted; 96 graded.
Open the run record (per-task verdicts)
J
Jina AIReader
42.8 Content F141 / 96 strict passes · 96 graded · 0 execution errors
Full benchmark96 tasks attempted
View run
Run details
Evaluation
WebBench-Fetch · Jina AI Reader
Environment
webbench.fetch_bench (hosted protocol, public tunnel)
Grader
WebBench suite scorer
Coverage
96 of 96 source tasks attempted; 96 graded.
Open the run record (per-task verdicts)
B
BrowserbaseFetch API
18.7 Content F121 / 96 strict passes · 96 graded · 0 execution errors
Full benchmark96 tasks attempted
View run
Run details
Evaluation
WebBench-Fetch · Browserbase Fetch API
Environment
webbench.fetch_bench (hosted protocol, public tunnel)
Grader
WebBench suite scorer
Coverage
96 of 96 source tasks attempted; 96 graded.
Open the run record (per-task verdicts)

Content F1 is the mean per-task score from the release's own verifier over the full roster. Strict passes are that verifier's all-or-nothing verdicts (version under Run details); the per-task verdicts are in the run record under Run details.

About WebBench-Fetch

What it measures

Content F1 (harmonic mean of content recall and signal ratio), usable pages (recall ≥ 0.9 and signal ≥ 0.3), table-row alignment, latency — computed against exact ground truth, no LLM judge.

Scope & limitations

Generated pages, not a crawl of the live web (anti-bot and paywalls are not measured). A rendered-browser reference ceiling and a tag-strip floor are recorded beside each run and are not ranked.

Evaluation environment

Pages generated from a WebBench world and served through a throw-away public tunnel; one request per page, provider defaults, no retries; every fetcher in one sitting on the same pages.

Source

Original benchmark

Dataset & task catalog ↗