Skip to benchmark catalog

Blobfish Benchmark Portfolio · 2026

Benchmarks for work that has consequences.

One leaderboard cannot describe an agent. We evaluate the work domain by domain—against public Harbor tasks, real tool contracts, and worlds you can inspect over MCP.

7professional domains
1,241public source tasks
24/24reference tasks pass
0/24random-floor tasks pass

The portfolio

A benchmark for each kind of work.

The task count belongs to the named Harbor source suite. The world count belongs to the separate Blobfish companion environment. Keeping those layers explicit prevents a local diagnostic from masquerading as an upstream result.

01 / Legal operationsModel results live

CounselBench-100

Can an agent finish a legal matter—not just answer a question?

Long-horizon legal work across evidence rooms, custody checks, source-grounded findings, and final deliverables.

Source tasks
100
MCP tools
Filesystem
Reference / floor
Published
$harbor run -d blobfishai/counselbench-100
CounselBench-100Open benchmark
02 / Financial researchBaselines live

FilingBench-50

Can an agent find the number, show its provenance, and update the model?

Public financial-research questions paired with a deterministic DCF, assumptions, period-data, and review-note world.

Source tasks
50
MCP tools
44
Reference / floor
4/4 · 0/4
$harbor run -d vals/financeagent
Finance Agent BenchmarkOpen benchmark
03 / Customer operationsBaselines live

ServiceBench-375

Can an agent follow policy while the customer and the system keep changing?

Multi-turn service-agent evaluation paired with an SLA, ticket, escalation, message, and knowledge-base world.

Source tasks
375
MCP tools
43
Reference / floor
4/4 · 0/4
$harbor run -d sierra-research/tau3-bench
04 / Scientific discoveryBaselines live

DiscoveryBench-102

Can an agent turn scientific data into a reproducible result?

Scientific-programming tasks paired with an assay, compound, target, QC-result, and literature world.

Source tasks
102
MCP tools
43
Reference / floor
4/4 · 0/4
$harbor run -d scienceagentbench/scienceagentbench
ScienceAgentBenchOpen benchmark
05 / Defensive securityBaselines live

DefenseBench-100

Can an agent harden a service without reaching beyond the sandbox?

Local defensive-hardening tasks paired with a permission-aware search, document, employee, and access-grant world.

Source tasks
100
MCP tools
42
Reference / floor
4/4 · 0/4
$harbor run -d polyvorlabs/cyberdefense-bench
CyberDefenseBenchOpen benchmark
06 / Enterprise operationsBaselines live

OpsBench-14

Can an agent join the systems behind a real operational decision?

Cross-functional enterprise questions paired with vendors, processes, changes, controls, and review workflows.

Source tasks
14
MCP tools
42
Reference / floor
4/4 · 0/4
$harbor run -d Enterprise-Bench/l1-l2-bench
Enterprise-Bench L1–L2Open benchmark
07 / Software engineeringBaselines live

RepoBench-500

Can an agent resolve the issue in the repository—not just propose a patch?

Expert-verified GitHub issues paired with a service, repository, pull-request, incident, ADR, and standup world.

Source tasks
500
MCP tools
44
Reference / floor
4/4 · 0/4
$harbor run -d swe-bench/swe-bench-verified
SWE-bench VerifiedOpen benchmark

One publication contract

Source task. Executable world. Evidence trail.

Inspired by the clarity of professional-work leaderboards, every page separates what was evaluated, what can be reproduced, and what has not been run yet.

  1. 01

    Evaluate on Harbor

    Keep the upstream task, sandbox, verifier, version, and scoring semantics intact.

  2. 02

    Inspect over MCP

    Open the companion schema, tools, seeded state, task split, and exact-state verifier locally.

  3. 03

    Publish the run

    Attach provider identity, configuration, trajectories, cost, and immutable artifacts before showing a score.

Result integrity

Controls first. Model ranks when the runs exist.

CounselBench has published, inspectable model runs. The six new suites now expose a 24/24 reference ceiling, a 0/24 seeded-random floor, real task and asset records, and paired trajectories. Model rows remain reserved for version-pinned executions.

See a fully published benchmark →

Harbor inventory observed 2026-08-26; 282 public datasets were visible in the live catalog. Source licenses and runtime requirements remain authoritative at the linked upstream pages.