Skip to benchmark
Blobfish ResearchSalesBench-100 v3.4.2Public

Can an agent run the revenue workflow—not just summarize the call?

SalesBench-100 measures complete sales operations: reconcile evidence, validate the CRM and engagement system, update only authorized records, preserve every control, and deliver an auditable executive brief.

100sales workflows
2,800seeded sources
93median reference calls
4MCP servers
0LLM grading calls
Executed release

1,500 isolated qualification runs cover 100/100 oracle passes, 100/100 exact replays, and 1,300 negative-control attacks across 13 attack types. False accepts remain 0. 100/100 write contracts prove authorized scope, an agent-visible destination, typed provider inputs, semantic prose, persisted-state readback, and containment. Wrong-target and keyword-stuffing attacks were rejected; hidden reference text and hidden serialization required: 0 and 0. No model row is shown until a full run on this exact release exists.

Inspect the verifier contract →
Loading the 100-task benchmark explorer…

Run your model

Score your own model on the SalesBench-100 sandbox.

Scores the public sandbox task with your model — verified by the sandbox grader, not a leaderboard entry.

Claude, ChatGPT, Grok, DeepSeek, or a checkpoint of your own behind an OpenAI-compatible endpoint. The run drives the same isolated MCP sandbox the Live MCP console uses; the score is the sandbox grader's verdict on the state your model left behind. Run the one public task inline, or a suite of the hosted frozen tasks as a background job and publish it below the official leaderboard. The full 100-task suite runs on Harbor with the command at the bottom. How runs, suites, and community runs work →

Required. Mint one with POST /api/v1/auth/keys. It authenticates the run and lists your runs afterwards; it is sent with the request only.
What to run
Today's inline run: one public sandbox task, scored and returned in this request.
Model
Not configured for you yet — paste a Anthropic key here for this run, or save one with POST /api/v1/accounts/providers. Never stored with the run.
140 model turns; the run is capped at four minutes.
Full 100-task suite on Harbor
harbor run -d blobfishai/salesbench-100 -a <agent> -m anthropic/claude-sonnet-5
Replace <agent> with your agent scaffold. A hosted run above scores the frozen tasks this host carries; a leaderboard row needs all 100 tasks on this release with traces.

Community runs

Self-runs on SalesBench-100 — below the official leaderboard, never merged into it.

self-run · graded by the benchmark's own verifier · not the frozen official suite unless suite:full

Community runs are self-runs: a caller's model, graded by this benchmark's own verifier over the hosted frozen tasks the run covered. They are listed below the official leaderboard and never merged into it. The Coverage column counts the official tasks each run itself covered, out of the 100-task release; “hosted” tasks are the frozen tasks this host carries, which caps what any run here can cover. Only a full suite over every official task reads official suite. How to run and publish one →

Loading community runs…

Run it yourself

The dataset is public. The four-system world is executable.

Download every seed, run the Harbor world, replay the oracle, and inspect every state transition and scored criterion before submitting the first exact model run.