One suite per System One use-case family (workflow control, retrieval and knowledge, safety and quality, data and operations, real-time and agents) across 13 industry verticals: yes/no, choice and ordinal-score decisions on public labelled data and seeded generators, scored as calibrated probabilities.
DecisionScore: the mean per-item proper score (1 - normalised Brier) over the full roster, unanswered items as 0; accuracy, ROC-AUC, ECE, log loss and Brier per suite.
Scope & limitations
Public source datasets overlap with some systems' training data; every system and suite carries a contamination tier on the benchmark page. Latency is not ranked: hosted APIs and local open-weight runs have different measurement boundaries.
Evaluation environment
Each system receives the identical compact-JSON state and typed questions through the System One wire format; structurally unsupported items are recorded as unanswered.