← Benchmark leaderboardAll leaderboards

online-mind2web / Benchmark deep dive

Online-Mind2Web (GitHub snapshot pilot): inside the results.

Compare the outcomes, understand the gaps, and inspect the work behind every score.

10selected tasks per model
1model configurations
3/10primary-verified attempts
89recorded tool actions

What the results tell us.

01

A completed run still needs evidence.

3 of 10 selected attempts passed the declared primary evaluation. 3 ended with an execution error or timeout. Every selected attempt stays in the report.

02

Evidence scope changes the verdict.

2 task verdicts differ between the rubric and screenshot reviews. The rubric can use captured page text; vision is limited to the recorded frames. Open the task to see what each judgment actually supports.

03

Cost is still an open measurement.

Dollar cost is recorded for 0 of 10 attempts. Subscription usage and missing provider billing are not converted into an invented per-task price. Latency includes the scope declared in the protocol below.

Performance in context.

Scores use the same selected tasks and the report’s declared primary evaluator.

Evidence-supported task success · one attempt per task
Model / agentSuccessExecution errorsMean latencyCost / task
Codex + Playwrightgpt-6-astra30.0%3/10 selected tasks
3182.9sNot reported

Pilot coverage: 10/300 source tasks. Small fixed selections do not establish a full-benchmark ranking. Missing grades remain visible.

Where attempts fall short.

Execution failure and an incorrect result are separate outcomes.

Codex + Playwright

Verified successes
3
Graded failures
4
Execution errors / timeouts
3
Completed, ungraded
0

RuntimeError: Server error '502 Bad Gateway' for url 'http://10.40.192.188:8765/solve' For more information check: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/502

The work behind the score

Inspect a task.

Original instructions, observed actions, and the verdict for each attempt.

Execution error

gpt-6-astra · 335.4s · 15 recorded actions

Original task

Search for a new internal M2 Samsung SSD drive between $25 and $200.

Grading criteria & verdicts

Executable checksUngraded

Recorded benchmark evaluator · not-configured

Evaluator was not configured

Rubric reviewUngraded

gpt-6-astra · browser-dom-constraints-v1

RuntimeError: Server error '502 Bad Gateway' for url 'http://10.40.192.188:8765/solve' For more information check: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/502

Screenshot reviewUngraded

gpt-6-astra · browser-uniform-eight-frames-v1

RuntimeError: Server error '502 Bad Gateway' for url 'http://10.40.192.188:8765/solve' For more information check: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/502

Link to this attempt ↗
Source checksums

515f2e5811cfdd5e0e669e40f17886d8.tar.gz
4aefdbcd8155a827db7abc968b970ad097edeb68d63290ab958110f646719899

raw_outputs.jsonl
df04693795f8642cd5fe222688f4c47f3e57b122748b0ec2cb99fcb1c334811b

Reproducibility

The evaluation protocol.

Imported recorded evidence; not independently rerun by the report engine. Scaffold: harbor-playwright-dom-v1 at Harbor 0.21.0; Playwright 1.58.0; Codex CLI 0.154.0. Unknown measurements remain unknown. Fresh headless Chromium per task. Independent DOM rubric primary; at most eight uniformly spaced screenshots including first and last for supplemental vision. Native task-success grading is unconfigured. This is not the official WebJudge protocol or human adjudication. Solver and independent graders use the same model family; dollar costs are unmeasured. Task latency includes browser/solver setup and execution. Per-grader latency and total duration including grading are retained separately in raw outputs. Partial trajectories preserve original execution errors without rescoring.

Task selection
10 frozen tasks, one attempt for each model. bab3fc3f1bff5b42c0624c073b68ddf9ad651d91
Tools & environment
GKE bf-benchmarks; fresh gVisor Chromium per task; no model credentials in browser guest. Provider driver versions are recorded; matching task tools do not imply identical provider implementations.
Budget
One attempt; 300 solver seconds; 60 browser actions; one task at a time; 7200s Job deadline
Independent grading
gpt-6-astra. Original grader versions and missing verdicts are preserved.
Evidence
Task prompts, tool inputs, observed results, final answers and grader receipts. Each trajectory links to its original artifact checksums.
Limits
One attempt per task. No pass@k, repeated-run reliability or expert agreement claim. Unreported costs stay unknown. Browser Use’s managed agent belongs to its own harness cohort.
Open the original report and downloads
Online-Mind2Web (GitHub snapshot pilot) · agent evaluation1 runs · Published evaluation