Test the work that matters
Choose a registered HUD or Harbor benchmark. Run different agent scaffolds in isolated environments, with a grader for every task.
Agent evaluation, with the evidence.
Run the tests. Understand the failures. Build a better agent.
Benchmark agents on the same tasks and turn their results into private reports—with the detail your engineers need and the clarity your customers expect.
Browser workflows · 12 example tasks
Your next improvement is in the details.Open any failed task to see the output, evaluator explanation, and execution trace.
From evaluation to improvement
Choose a registered HUD or Harbor benchmark. Run different agent scaffolds in isolated environments, with a grader for every task.
Separate incorrect outcomes from timeouts and execution errors. See the category breakdown, then inspect the exact task evidence.
Compare configurations under matching conditions. Give engineering a focused starting point and share a report when you’re ready.
Two views. One body of evidence.
Go from an overall score to the exact task where your agent got stuck.
Compare agents on matching task sets, with the context behind the ranking.
One workflow, many benchmarks
Run portable environments with OpenAI or Claude agent scaffolds. Keep the native task grade and full trajectory.
Prepared environment images · pinned SDKExplore the workflowUse task.toml, instructions, environment images, and tests/test.sh. Compare Terminus 2, Codex, and Claude Code.
A fresh task environment for every attemptExplore the workflowRegister browser tasks with screenshot evidence and task-specific rubrics. Export verdicts and reasoning for each task.
WebVoyager and Mind2Web require prepared task packagesExplore the workflowThe live registry shows exactly which task packages are available. Infrastructure controls are labeled separately from capability benchmarks. Every report retains its dataset version, task coverage, scaffold, and grading rules.
Your next release, backed by evidence.
Start an evaluation or turn an existing run into your first report.