CRM operations / Benchmark leaderboard
Arc CRM 6
Independent synthetic CRM work across contact onboarding, stage correction, bounded quotes, quote replacement, unsigned contracts and corrected document associations. All source conversations were inspected; zero source rows are reproduced.
No published company result yet.
The existing Blobfish benchmark page retains its task definitions, baseline evidence and release status. Company comparisons will appear here after a reviewed evaluation.
View existing benchmark evidence and release status →
Arrange an evaluation for your company →What this benchmark measures
Metrics
Deterministic evidence, post-write readback, outcome, containment, handoff and structured-answer checks over saved SQLite state. All 167 authoring negative controls reject. Reference controls establish solvability, not model performance.
Execution requirements
Package 0.1.2 exposes 31 domain/evidence contracts plus context and answer controls through CLI, real HTML forms, REST and MCP sharing an isolated episode-local world. A separate networkless verifier grades the collected state. Six source-inspired workflows and 33 synthetic evidence files (PDF, XLSX, EML and Markdown) are partial coverage, not 1,200 converted conversations or upstream API parity.
Scope and limitations
This Blobfish suite has its own task and environment definitions. Upstream scores and reference controls are not company performance scores.
Evaluation availability
Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.
Sources and company fit
- Original benchmark ↗Blobfish benchmark suite
- Dataset / catalog ↗