Skip to content

Leaderboard / Benchmark directory

Find the right test.

Browse by capability, company fit, or source catalog. Benchmarks with published results open their leaderboard; the rest open their catalog page.

122 research profiles303 Harbor entries47 HF official-filter cards15 Blobfish suites990 company recommendations

Catalog snapshot: 2026-09-15. Exact source matches are joined; versions and ports remain distinct. This includes every entry reviewed in our research, not every dataset on Hugging Face. Catalog inclusion and company fit do not imply a hosted integration or a measured result.

5 benchmarks with results

Browser automationResults published

WebVoyager

End-to-end navigation and task completion on live websites; original release has 643 tasks across 15 sites.

For Browser Use, Browserbase / Stagehand, Firecrawl +4

View leaderboard
Browser automationResults published

Online-Mind2Web

300 real-world tasks spanning 136 websites; maintainers update invalid tasks.

For Browser Use, Browserbase / Stagehand, Firecrawl +4

View leaderboard
Browser automationResults published

WebArena

Multi-step commerce, forum, CMS and developer-site workflows in resettable websites.

For Automation Anywhere, Browser Use, Browserbase / Stagehand +6

View leaderboard
CodingResults published

HumanEval

Python function generation with original executable tests.

View leaderboard
CodingResults published

MBPP

Python programming tasks evaluated against source-defined tests.

View leaderboard