Skip to content

Typed decisions

Typed Decision Bench

One suite per System One use-case family (workflow control, retrieval and knowledge, safety and quality, data and operations, real-time and agents) across 13 industry verticals: yes/no, choice and ordinal-score decisions on public labelled data and seeded generators, scored as calibrated probabilities.

Typed decisions

Model performance

1 recorded configuration · 6,162 source tasks

Updated
Higher is better · 0–100 · execution errors score 0
O
OpenJevopenjev-latest
75.4 DecisionScore4394 / 6162 strict passes · 6162 graded · 0 execution errors
Full benchmark6162 tasks attempted
View run
Run details
Evaluation
Typed Decision Bench · v0.2.0
Environment
jevfish_bench runner · hosted API
Grader
Ground-truth labels
Verifier
typed-decision-bench:0.2.0
Coverage
6162 of 6162 source tasks attempted; 6162 graded.
Open the run record (per-task verdicts)

DecisionScore is the mean per-task score from the release's own verifier over the full roster. Strict passes are that verifier's all-or-nothing verdicts (version under Run details); the per-task verdicts are in the run record under Run details.

About Typed Decision Bench

What it measures

DecisionScore: the mean per-item proper score (1 - normalised Brier) over the full roster, unanswered items as 0; accuracy, ROC-AUC, ECE, log loss and Brier per suite.

Scope & limitations

Public source datasets overlap with some systems' training data; every system and suite carries a contamination tier on the benchmark page. Latency is not ranked: hosted APIs and local open-weight runs have different measurement boundaries.

Evaluation environment

Each system receives the identical compact-JSON state and typed questions through the System One wire format; structurally unsupported items are recorded as unanswered.

Source

Original benchmark

Dataset & task catalog ↗