Browser-only SaaS operations
WebBench
Browser-only correlated SaaS work over deterministic populated worlds: find the right record among look-alikes, carry ids, dates and amounts across applications, obey the policy or capacity that gates an action, and record a structured decision — graded executably against the application backends.
About WebBench
What it measures
Executable WebScore verification tree per task: gates for containment, evidence before the first commit, no canary or injected write, and abstention; weighted leaves for evidence reads, state rows, answer fields, cross-application consistency and read-backs, graded from the pulled world database and the server-side page trace. No LLM judge; exact call order is not graded.
Scope & limitations
This Blobfish suite has its own task and environment definitions. Upstream scores and reference controls are not company performance scores.
Evaluation environment
The exact 0.3.0 release contains 137 tasks from 13 calibrated templates over 3 populated families (deskops, itsmdesk, workplace), each served as 10–12 server-rendered web applications over one SQLite world. The agent's only surface is the browser CLI; MCP, REST and any tool CLI answer 404. Every task shipped with its nine build-time controls, passed the Docker oracle gate 137/137 and a registry round-trip 5/5. Templates were calibrated on DeepSeek V4.1 Flash; those are calibration measurements, not leaderboard results.