Skip to content

Blobfish AI / Independent evaluations

Agent leaderboards.

Different benchmarks. Real task outcomes.
Compare models, understand failures, and inspect every task.

Matched tasks and tools. Independent Astra grading. Recorded evidence.

Measured evaluations

Scores you can inspect.

Every result links to its task coverage, grading protocol and evidence.

See hosted runs

We host the harness.

Frozen tasks, isolated environments and retained execution evidence.

Models get the same tools.

Compare Claude, ChatGPT and DeepSeek on the same task environments and budgets.

Open the complete report.

Inspect outcomes, latency, recorded costs and why individual tasks failed.

The benchmark directory

Find the right test.

Browse by capability, company fit, or source catalog.

Download directory
122 research profiles303 Harbor entries47 HF official-filter cards15 Blobfish suites990 company recommendations

Catalog snapshot: 2026-09-15. Exact source matches are joined; versions and ports remain distinct. This includes every entry reviewed in our research, not every dataset on Hugging Face. Catalog inclusion and company fit do not imply a hosted integration or a measured result.

428 benchmark entries

Browser automationResults published

WebVoyager

End-to-end navigation and task completion on live websites; original release has 643 tasks across 15 sites.

For Browser Use, Browserbase / Stagehand, Firecrawl +4

View leaderboard
Browser automationResults published

Online-Mind2Web

300 real-world tasks spanning 136 websites; maintainers update invalid tasks.

For Browser Use, Browserbase / Stagehand, Firecrawl +4

View leaderboard
Browser automationCatalog

WebArena

Multi-step commerce, forum, CMS and developer-site workflows in resettable websites.

For Automation Anywhere, Browser Use, Browserbase / Stagehand +6

View benchmark
Browser automationCatalog

VisualWebArena

Browser tasks requiring visual understanding, including images and visually grounded shopping.

For Anthropic, Browser Use, Browserbase / Stagehand +3

View benchmark
Browser automationCatalog

WorkArena / WorkArena++

ServiceNow knowledge-work tasks and compositional enterprise workflows.

For Automation Anywhere, Browser Use, Browserbase / Stagehand +3

View benchmark
Desktop and mobile agentsCatalog

OSWorld-Verified

Desktop tasks involving applications, files and cross-application workflows.

For Anthropic, Automation Anywhere, Google DeepMind +2

View benchmark
Desktop and mobile agentsCatalog

OSWorld 2.0

Long-horizon computer-use tasks with coordinated code, assets and mocked websites.

For Anthropic, Automation Anywhere, Google DeepMind +2

View benchmark
Desktop and mobile agentsCatalog

ScreenSpot-Pro

Locate the correct UI target in professional, high-resolution screenshots.

For Anthropic, Automation Anywhere, Google DeepMind +2

View benchmark
Desktop and mobile agentsCatalog

AndroidWorld

116 parameterized task templates across 20 Android apps.

For Anthropic, Google DeepMind, OpenAI +2

View benchmark
Coding and terminal agentsCatalog

SWE-bench Verified

Resolve 500 human-validated GitHub issues.

For All Hands / OpenHands, Augment Code, Cline +6

View benchmark
Coding and terminal agentsCatalog

SWE-bench Pro

Long-horizon repository-level software tasks; public Harbor snapshot has 731 tasks.

For All Hands / OpenHands, Augment Code, Cline +7

View benchmark
Coding and terminal agentsCatalog

SWE-bench Multilingual

Repository issue resolution across nine programming languages, 300 test instances.

For All Hands / OpenHands, Augment Code, Cline +7

View benchmark
Coding and terminal agentsCatalog

Terminal-Bench 2.1

89 difficult terminal tasks across engineering, systems and scientific work.

For All Hands / OpenHands, Anthropic, Augment Code +9

View benchmark
Coding and terminal agentsCatalog

Terminal-Bench 4.0

Current continuous Terminal-Bench release with task and resource revisions; catalog snapshot lists 66 tasks.

For All Hands / OpenHands, Anthropic, Augment Code +9

View benchmark
Coding and terminal agentsCatalog

LiveCodeBench

Competition-style coding, with date-bounded problem releases.

For Alibaba Qwen, Anthropic, Cognition / Devin +7

View benchmark
Coding and terminal agentsCatalog

Aider Polyglot

225 programming exercises across six languages.

For Augment Code, Cline, Cognition / Devin +4

View benchmark
Coding and terminal agentsCatalog

BigCodeBench / BigCodeBench-Hard

Python programming tasks using real libraries and APIs.

For All Hands / OpenHands, Augment Code, Cline +8

View benchmark
Coding and terminal agentsCatalog

HumanEval+ / MBPP+ (EvalPlus)

Small Python function synthesis with expanded tests.

For Alibaba Qwen, All Hands / OpenHands, Anthropic +17

View benchmark
Coding and terminal agentsCatalog

DeepSWE 1.1

113 original long-horizon repository tasks across five programming languages.

For All Hands / OpenHands, Augment Code, Cline +7

View benchmark
Coding and terminal agentsCatalog

Android Bench

100 software engineering tasks on Android repositories.

For Augment Code, Cognition / Devin, Cursor +3

View benchmark
Coding and terminal agentsCatalog

Long-Horizon Terminal-Bench

46 tasks testing sustained useful work over hundreds of steps.

For All Hands / OpenHands, Augment Code, Cline +6

View benchmark
Customer service and tool useCatalog

τ-bench / τ²-bench

Policy-following conversations, tool actions and cooperative customer-service tasks.

For Ada, Decagon, Forethought +5

View benchmark
Customer service and tool useCatalog

τ³-bench: text and knowledge

Customer service in airline, retail, telecom and banking-knowledge domains; Harbor lists 375 tasks.

For Ada, Decagon, Forethought +6

View benchmark

For teams building agents

Put your agent on the board.

We help teams evaluate browser, coding, voice, legal, search and long-horizon agents against real work.