Scientific discovery / Benchmark leaderboard
DiscoveryBench-102
Scientific-programming tasks paired with an assay, compound, target, QC-result, and literature world.
No published company result yet.
The existing Blobfish benchmark page retains its task definitions, baseline evidence and release status. Company comparisons will appear here after a reviewed evaluation.
View existing benchmark evidence and release status →
Arrange an evaluation for your company →What this benchmark measures
Metrics
102 scientific-programming tasks: 38 deterministic and 64 visual/LLM-judge tasks.
Execution requirements
Three tasks require GPUs; visual tasks require the verifier model configured by the upstream package.
Scope and limitations
This Blobfish suite has its own task and environment definitions. Upstream scores and reference controls are not company performance scores.
Evaluation availability
Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.
Sources and company fit
- Original benchmark ↗Blobfish benchmark suite
- Dataset / catalog ↗