Fetch APIs on generated pages behind real-web friction (heavy chrome, consent overlays, client-side rendering, late content, tables): does the returned text contain the page's content, and how much else?
Content F1 (harmonic mean of content recall and signal ratio), usable pages (recall ≥ 0.9 and signal ≥ 0.3), table-row alignment, latency — computed against exact ground truth, no LLM judge.
Scope & limitations
Generated pages, not a crawl of the live web (anti-bot and paywalls are not measured). A rendered-browser reference ceiling and a tag-strip floor are recorded beside each run and are not ranked.
Evaluation environment
Pages generated from a WebBench world and served through a throw-away public tunnel; one request per page, provider defaults, no retries; every fetcher in one sitting on the same pages.