Scientific discovery / Benchmark leaderboard

DiscoveryBench-102

Scientific-programming tasks paired with an assay, compound, target, QC-result, and literature world.

No published company result yet.

The existing Blobfish benchmark page retains its task definitions, baseline evidence and release status. Company comparisons will appear here after a reviewed evaluation.

View existing benchmark evidence and release status →

Arrange an evaluation for your company →

What this benchmark measures

Metrics

102 scientific-programming tasks: 38 deterministic and 64 visual/LLM-judge tasks.

Execution requirements

Three tasks require GPUs; visual tasks require the verifier model configured by the upstream package.

Scope and limitations

This Blobfish suite has its own task and environment definitions. Upstream scores and reference controls are not company performance scores.

Evaluation availability

Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.

Sources and company fit