CRM operations / Benchmark leaderboard

Arc CRM 6

Independent synthetic CRM work across contact onboarding, stage correction, bounded quotes, quote replacement, unsigned contracts and corrected document associations. All source conversations were inspected; zero source rows are reproduced.

No published company result yet.

The existing Blobfish benchmark page retains its task definitions, baseline evidence and release status. Company comparisons will appear here after a reviewed evaluation.

View existing benchmark evidence and release status →

Arrange an evaluation for your company →

What this benchmark measures

Metrics

Deterministic evidence, post-write readback, outcome, containment, handoff and structured-answer checks over saved SQLite state. All 167 authoring negative controls reject. Reference controls establish solvability, not model performance.

Execution requirements

Package 0.1.2 exposes 31 domain/evidence contracts plus context and answer controls through CLI, real HTML forms, REST and MCP sharing an isolated episode-local world. A separate networkless verifier grades the collected state. Six source-inspired workflows and 33 synthetic evidence files (PDF, XLSX, EML and Markdown) are partial coverage, not 1,200 converted conversations or upstream API parity.

Scope and limitations

This Blobfish suite has its own task and environment definitions. Upstream scores and reference controls are not company performance scores.

Evaluation availability

Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.

Sources and company fit