Skip to benchmark catalog
15professional domains
1,947public source tasks
9first-party benchmarks
24/24reference tasks pass
0/24wrong-evidence tasks pass
628/628chain-audited tasks pass

The evidence contract

Know what every number means before you compare it.

The source evaluation, executable environment, and recorded run are separate layers. Each has its own identity and evidence, so a local control can never masquerade as an upstream score.

  1. 01

    Source evaluation

    Keep the public task, sandbox, verifier, version, and scoring semantics intact.

  2. 02

    Executable world

    Inspect the companion schema, tools, seeded state, task split, and deterministic checks over MCP.

  3. 03

    Recorded run

    Resolve provider identity, configuration, trajectories, cost, and immutable artifacts to one exact release.

628/628 chain-audited tasks pass. Admission checks the full dependent chain: required investigations, derived intermediates, three costed and authority-checked alternatives, and an exact multi-field answer. No LLM judge is used.

Inspect the executable realism standard ↗

The portfolio

A benchmark for each kind of work.

Start with the domain, then open its evidence page to inspect tasks, assets, tool contracts, controls, executable world proofs, and any version-pinned model runs that meet the publication rule.

Blobfish benchmarks

09 releases

First-party releases: built, oracle-qualified, and published by Blobfish, each with a purpose-built page carrying full task, asset, contract, and trajectory evidence.

01 / Legal operationsAwaiting first model run

CounselBench-100

Can an agent finish a legal matter—not just answer a question?

Long-horizon legal work across evidence rooms, custody checks, source-grounded findings, and final deliverables.

Source tasks
100
MCP tools
Filesystem
Oracle replay
100/100
harbor run -d blobfishai/counselbench-100
CounselBench-100Inspect evidence
02 / Sales operationsAwaiting first model run

SalesBench-100

Can an agent run the revenue workflow—not just summarize the call?

Long-horizon revenue work across Salesforce, HubSpot, Gong, seeded account rooms, controlled CRM mutations, and executive deliverables.

Source tasks
100
MCP tools
35
Oracle replay
100/100
harbor run -d blobfishai/salesbench-100
SalesBench-100Inspect evidence
03 / Software operationsAwaiting first model run

DevOpsBench-100

Can an agent run the incident, ship the fix, and prove it landed?

Long-horizon DevOps/SRE work over one executable NovaCart world: root-cause analysis, code changes, canaried deploys, migrations, flags, and incident closure.

Source tasks
100
MCP tools
97
Oracle replay
100/100
harbor run -d blobfishai/devopsbench-100
DevOpsBench-100Inspect evidence
04 / Financial operationsAwaiting first model run

LedgerBench-100

Can an agent run the finance desk—not just look up the number?

Corporate-finance work across 22 families: a D365-shaped ERP, Odoo, subsidiary books, drive, email, documents, and real SEC XBRL filings behind 8 MCP servers.

Source tasks
100
MCP tools
66
Oracle replay
100/100
harbor run -d blobfishai/ledgerbench-100
LedgerBench-100Inspect evidence
05 / Manufacturing operationsModel results live

FactoryBench-100

Can an agent answer the factory question—and safely carry the decision through?

Realistic manufacturing decisions framed as high-level employee requests. Agents must discover the relevant Oracle Fusion, Gmail, Drive, Sheets, and Slack evidence, reconcile records, calculate and compare options, then execute only the supported outcome.

Source tasks
100
MCP tools
94
Oracle replay
100/100
harbor run -d blobfishai/factorybench-100
FactoryBench-100Inspect evidence
06 / Semiconductor operationsAwaiting first model run

SemiKongBench-100

Can an agent contain the semiconductor exception—and prove the fab is safe?

Long-horizon semiconductor exception work across lot genealogy, plasma etch, CMP, lithography, maintenance, materials, final test, reliability, FA/CAPA, and supply traceability.

Source tasks
100
MCP tools
38
Oracle replay
100/100
harbor run -d blobfishai/semikongbench-100
SemiKongBench-100Inspect evidence
07 / Investment bankingAwaiting first model run

DealBench-100

Can an agent run the deal—not just calculate one valuation cell?

Long-horizon investment-banking execution across source control, QoE, comps, precedents, DCF, LBO, merger math, bids, model-to-deck consistency, approvals, and review-ready client handoffs.

Source tasks
100
MCP tools
36
Oracle replay
100/100
harbor run -d blobfishai/dealbench-100-suite
08 / Enterprise resource planningAwaiting first model run

ERPBench-100

Can an agent run the ERP transaction end to end—not just read the record?

Long-horizon ERP execution across order import, shipment verification, receipt application, reorder requisitions, receiving and three-way match, worker document compliance, shift rollups, channel-order sync, hiring within quota, and effective-dated price batches.

Source tasks
100
MCP tools
52
Oracle replay
100/100
harbor run -d blobfishai/erpbench-100-suite
09 / CRM operationsAwaiting first model run

Arc CRM 6

Can an agent finish the CRM workflow without breaking the account?

Independent synthetic CRM work across contact onboarding, stage correction, bounded quotes, quote replacement, unsigned contracts and corrected document associations. All source conversations were inspected; zero source rows are reproduced.

Source tasks
6
MCP tools
33
Oracle replay
6/6 local + registry
harbor run -d blobfishai/arc-crm-6@sha256:fc9ca3b80d6a1a3c3e5bd1b975c54f1f3ada21b49668cb5f25e2912d67b69dfa -a oracle -e docker -n 1 -k 1 -r 0
Arc CRM 6 · v0.1.2Inspect evidence

Executable world proofs

01 inspected world

Benchmark-shaped environments with inspectable tasks, state, tools, verifiers, and trajectories. They are listed in the portfolio but kept outside leaderboard comparisons until one frozen evaluation identity has comparable model runs.

16 / Cross-functional operationsWorld proof · not ranked

Northstar Composite Operations World

Can an agent coordinate real state changes across customer, support, billing, engineering, and knowledge systems without touching the wrong records?

The interactive Northstar release contains 100 frozen tasks with executable verifiers and task-level traces. Its page separately labels a larger 296-tool customer release and 100 single-model receipts, preventing either artifact from being mistaken for a cross-model leaderboard.

Frozen tasks
100
MCP surface
5 servers · 32 tools
Qualification
100/100 oracle · 400/400 replays
1,834/1,834 verifier mutants killed · 0 false positives · 0 false negativesOpen tasks, traces, and downloads

HubBench — professional-domain worlds

01 rolling program

13 independently authored, oracle-proven domain families inspired by the Harbor Hub inventory. These shared worlds are not individual adaptations of every source dataset. Inspect source classifications and pending adaptations →

13/13 families · ClinicOps · HostOps · DataDesk · ResearchDesk · DesignOps · DeskOps · ITSMDesk · PolicyDesk · RepoDesk · SciLab · SecOps · WebStudio · WorkplaceAwaiting first model run

HubBench

Can one agent handle employee decisions across 13 professional domains?

Each released family is a self-contained SQLite world behind provider-shaped MCP tools and a terminal CLI, where the question is only answerable through a dependent chain of evidence and a deterministic verifier grades the whole chain — options, intermediates, authority, and the exact answer.

Qualified tasks
104
Domain families
13
Oracle replay
104/104 · 0 false accepts
Blobfish Research · rolling release toward 104 tasksInspect evidence

Companion suites for public benchmarks

06 suites

Public Harbor evaluations from the wider ecosystem, each paired with a separate executable Blobfish diagnostic world.

10 / Financial researchBaselines live

FilingBench-50

Can an agent find the number, show its provenance, and update the model?

Public financial-research questions paired with a deterministic DCF, assumptions, period-data, and review-note world.

Source tasks
50
MCP tools
44
Reference / floor
4/4 · 0/4
harbor run -d vals/financeagent
Finance Agent BenchmarkInspect evidence
11 / Customer operationsBaselines live

ServiceBench-375

Can an agent follow policy while the customer and the system keep changing?

Multi-turn service-agent evaluation paired with an SLA, ticket, escalation, message, and knowledge-base world.

Source tasks
375
MCP tools
43
Reference / floor
4/4 · 0/4
harbor run -d sierra-research/tau3-bench
12 / Scientific discoveryBaselines live

DiscoveryBench-102

Can an agent turn scientific data into a reproducible result?

Scientific-programming tasks paired with an assay, compound, target, QC-result, and literature world.

Source tasks
102
MCP tools
43
Reference / floor
4/4 · 0/4
harbor run -d scienceagentbench/scienceagentbench
ScienceAgentBenchInspect evidence
13 / Defensive securityBaselines live

DefenseBench-100

Can an agent harden a service without reaching beyond the sandbox?

Local defensive-hardening tasks paired with a permission-aware search, document, employee, and access-grant world.

Source tasks
100
MCP tools
42
Reference / floor
4/4 · 0/4
harbor run -d polyvorlabs/cyberdefense-bench
CyberDefenseBenchInspect evidence
14 / Enterprise operationsBaselines live

OpsBench-14

Can an agent join the systems behind a real operational decision?

Cross-functional enterprise questions paired with vendors, processes, changes, controls, and review workflows.

Source tasks
14
MCP tools
42
Reference / floor
4/4 · 0/4
harbor run -d Enterprise-Bench/l1-l2-bench
Enterprise-Bench L1–L2Inspect evidence
15 / Software engineeringBaselines live

RepoBench-500

Can an agent resolve the issue in the repository—not just propose a patch?

Expert-verified GitHub issues paired with a service, repository, pull-request, incident, ADR, and standup world.

Source tasks
500
MCP tools
44
Reference / floor
4/4 · 0/4
harbor run -d swe-bench/swe-bench-verified
SWE-bench VerifiedInspect evidence

Result integrity

A missing score is a disclosure, not a zero.

The five exact current packages pass 500/500 Harbor Oracle solvability trials with zero exceptions or retries. FactoryBench publishes one version-pinned full-suite model run. CounselBench, SalesBench, DevOpsBench, and LedgerBench likewise withhold older, partial, or pre-release rows until a full model run exists on the exact current release.

The six companion suites expose a 24/24 reference ceiling, a 0/24 wrong-evidence floor, inspectable task and asset records, and paired trajectories. Reference replay proves solvability, not model difficulty.

Exact-package oracle
500/500
Reference ceiling
24/24
Wrong-evidence floor
0/24

Harbor inventory observed 2026-08-29; 282 public datasets were visible in the live catalog. Source licenses and runtime requirements remain authoritative at the linked upstream pages.

Your workflow, made executable

Need a benchmark that looks like the work?

Bring the systems, policies, and failure modes that matter. Blobfish turns them into inspectable tasks, stateful tools, deterministic verification, and a release you can rerun.