Skip to content

Agent evaluation, with the evidence.

Know how your agent really performs.

Run the tests. Understand the failures. Build a better agent.

Benchmark agents on the same tasks and turn their results into private reports—with the detail your engineers need and the clarity your customers expect.

Private by default Evidence for every result
BLOBFISH / REPORTSSAMPLE DATA

Browser workflows · 12 example tasks

Competitive landscape

3agent configurations4task categories
AgentTask successCost / task
Atlas83.3%$0.101
Your agentYOU66.7%$0.071
Orbit50.0%$0.046

Your next improvement is in the details.Open any failed task to see the output, evaluator explanation, and execution trace.

Explore the full sample
Task success & reliability Cost & latency Failure analysis Raw outputs, grader verdicts & reports

From evaluation to improvement

A score is the start.
The report tells you what’s next.

Read our methodology
01

Test the work that matters

Choose a registered HUD or Harbor benchmark. Run different agent scaffolds in isolated environments, with a grader for every task.

02

Find the failure patterns

Separate incorrect outcomes from timeouts and execution errors. See the category breakdown, then inspect the exact task evidence.

03

Make the next release better

Compare configurations under matching conditions. Give engineering a focused starting point and share a report when you’re ready.

Two views. One body of evidence.

Reports built to be used.

Explore illustrative examples below

One workflow, many benchmarks

Your benchmark.
Your scaffold. Clear results.

Open the live registry
HUD FORMAT

HUD tasksets

Run portable environments with OpenAI or Claude agent scaffolds. Keep the native task grade and full trajectory.

Prepared environment images · pinned SDKExplore the workflow
HARBOR FORMAT

Harbor benchmarks

Use task.toml, instructions, environment images, and tests/test.sh. Compare Terminus 2, Codex, and Claude Code.

A fresh task environment for every attemptExplore the workflow
VISION + RUBRIC

Browser evaluation

Register browser tasks with screenshot evidence and task-specific rubrics. Export verdicts and reasoning for each task.

WebVoyager and Mind2Web require prepared task packagesExplore the workflow

The live registry shows exactly which task packages are available. Infrastructure controls are labeled separately from capability benchmarks. Every report retains its dataset version, task coverage, scaffold, and grading rules.

Your next release, backed by evidence.

Find the gap.
Then close it.

Start an evaluation or turn an existing run into your first report.

Create your report Talk to us about a custom evaluation