Manufacturing operations / Benchmark leaderboard
FactoryBench-100
Realistic manufacturing decisions framed as high-level employee requests. Agents must discover the relevant Oracle Fusion, Gmail, Drive, Sheets, and Slack evidence, reconcile records, calculate and compare options, then execute only the supported outcome.
No published company result yet.
The existing Blobfish benchmark page retains its task definitions, baseline evidence and release status. Company comparisons will appear here after a reviewed evaluation.
View existing benchmark evidence and release status →
Arrange an evaluation for your company →What this benchmark measures
Metrics
Consult the upstream task verifier.
Execution requirements
Dataset-specific harness qualification required.
Scope and limitations
Listed in the catalog; not treated as independently established or famous. Catalog task counts do not prove validated coverage.
Evaluation availability
Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.
Sources and company fit
- Original benchmark ↗Public description screening
- Harbor · blobfishai/factorybench-100 ↗100 source tasks · Public · Public description screening