Measured GKE pilot · September 16, 2026 UTC
Codex + Playwright.
Online-Mind2Web, with the evidence.
Ten frozen tasks from a public GitHub snapshot. A fresh browser for each attempt. The actions, screenshots and judgments behind every outcome.
What ran. gpt-6-astra through the Codex CLI with maximum reasoning, using Playwright 1.58.0 and Chromium in fresh gVisor containers. Harbor 0.21.0 manages task execution on GKE. The controller allows browser actions and records their resulting pages; model credentials stay in separate worker pods.
The task source. The original Online-Mind2Web instructions, starting URLs and human reference lengths come from AllenAI’s pinned public GitHub snapshot. Ten task IDs were selected before execution using seed 20260916. This snapshot is identified separately from the current official dataset. Dated requests, unavailable items and access barriers remain in the results.
How success is judged. Every task has frozen requirements. The primary rubric reviewer uses captured page text and factual actions. A separate vision reviewer uses up to eight evenly sampled screenshots, including the first and last. Both use Astra in fresh GKE grader pods and receive no earlier verdict. A final answer alone cannot establish success.
Scope and limits. This is a ten-task pilot using a custom evaluation protocol, not the official o4-mini WebJudge score, human adjudication or a competitive ranking. The solver and graders share a model family. Native task-success grading is unconfigured. Each task allows one attempt, 300 solver seconds and 60 browser actions. Task latency includes browser and solver setup through execution; each grader’s time and the total duration are retained separately. Subscription token usage is retained; dollar cost is unknown.
Expand a task below for its output and grader explanations. Evidence links include the final recorded screenshot, factual trajectory and browser archive. 3 failed attempts retain partial trajectories with their original error verdicts. Raw model event archives remain private. Dataset attribution: Xue et al., COLM 2025; Deng et al., NeurIPS 2023. Dataset license: CC BY 4.0.