Customer operations / Benchmark leaderboard
ServiceBench-375
Multi-turn service-agent evaluation paired with an SLA, ticket, escalation, message, and knowledge-base world.
No published company result yet.
The existing Blobfish benchmark page retains its task definitions, baseline evidence and release status. Company comparisons will appear here after a reviewed evaluation.
View existing benchmark evidence and release status →
Arrange an evaluation for your company →What this benchmark measures
Metrics
Policy-aware, tool-using conversations across airline, retail, telecom, and banking knowledge domains.
Execution requirements
The official repository remains tau2-bench; its current release is branded τ³-bench.
Scope and limitations
This Blobfish suite has its own task and environment definitions. Upstream scores and reference controls are not company performance scores.
Evaluation availability
Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.
Sources and company fit
- Original benchmark ↗Blobfish benchmark suite
- Dataset / catalog ↗