← Blobfish AI / LeaderboardRun an evaluation →

Measured pilot · September 16, 2026 UTC

Browser Use.
Fifteen real web tasks.

One fresh browser per task. The original instructions. Every answer and the reasoning behind its review.

10/15Evidence-supported outcomesSeparate Codex AI review
15/15Provider-completed sessionsCompletion is separate from success
15/643Source-task coverageOne task from each of 15 websites

What this establishes. A measured Browser Use Cloud v4 pilot using gpt-5.6-luna. The provider’s trajectory judge passed 12/15; a separate Codex evidence review supported 10/15. The report preserves both judgments and their disagreement.

How to read the result. Four tasks did not meet the original request and one material requirement could not be verified. The recorded rubric assigns no success credit to that unverified requirement. Two tasks still request travel dates in 2024. This pilot is not a full-suite score, competitor ranking, official WebVoyager GPT-4V evaluation or human adjudication.

Measured cost. $0.588206 in final provider-reported run costs, including asynchronous provider judging; $0.012333 in separately reported browser hosting. The supervising review’s cost was not measured. Vision grading used a subscription account; its token usage is retained without estimating a dollar cost.

The browser pilot was not rerun. The original outputs and rubric review are preserved; the new vision receipts are published as a dated supplement. Native execution grading remains unconfigured. Original publication manifest ↗

WebVoyager · 15-task Browser Use pilot · agent evaluation1 runs · Published evaluation