Can an agent run the revenue workflow—not just summarize the call?
SalesBench-100 measures complete sales operations: reconcile evidence, validate the CRM and engagement system, update only authorized records, preserve every control, and deliver an auditable executive brief.
1,500 isolated qualification runs cover 100/100 oracle passes, 100/100 exact replays, and 1,300 negative-control attacks across 13 attack types. False accepts remain 0. 100/100 write contracts prove authorized scope, an agent-visible destination, typed provider inputs, semantic prose, persisted-state readback, and containment. Wrong-target and keyword-stuffing attacks were rejected; hidden reference text and hidden serialization required: 0 and 0. No model row is shown until a full run on this exact release exists.
Inspect the verifier contract →Run your model
Score your own model on the SalesBench-100 sandbox.
Scores the public sandbox task with your model — verified by the sandbox grader, not a leaderboard entry.
Claude, ChatGPT, Grok, DeepSeek, or a checkpoint of your own behind an OpenAI-compatible endpoint. The run drives the same isolated MCP sandbox the Live MCP console uses; the score is the sandbox grader's verdict on the state your model left behind. Run the one public task inline, or a suite of the hosted frozen tasks as a background job and publish it below the official leaderboard. The full 100-task suite runs on Harbor with the command at the bottom. How runs, suites, and community runs work →
harbor run -d blobfishai/salesbench-100 -a <agent> -m anthropic/claude-sonnet-5Replace <agent> with your agent scaffold. A hosted run above scores the frozen tasks this host carries; a leaderboard row needs all 100 tasks on this release with traces.Community runs
Self-runs on SalesBench-100 — below the official leaderboard, never merged into it.
self-run · graded by the benchmark's own verifier · not the frozen official suite unless suite:full
Community runs are self-runs: a caller's model, graded by this benchmark's own verifier over the hosted frozen tasks the run covered. They are listed below the official leaderboard and never merged into it. The Coverage column counts the official tasks each run itself covered, out of the 100-task release; “hosted” tasks are the frozen tasks this host carries, which caps what any run here can cover. Only a full suite over every official task reads official suite. How to run and publish one →
Loading community runs…
Run it yourself
The dataset is public. The four-system world is executable.
Download every seed, run the Harbor world, replay the oracle, and inspect every state transition and scored criterion before submitting the first exact model run.