Typed Decision Bench · v0.3.0 · tasks
25 tasks, every system's result on each
Each task is one decision type drawn from one public labelled source. Open a task to see sample items turn by turn: the request every system receives, each system's response (its probability distribution and top pick) and the grade against the expected answer. Scores are DecisionScore (mean per-item proper score, 0–100, unanswered = 0) with top-pick accuracy beneath. Run report: every published item, request to verdict → Leaderboard →
Retrieval and knowledge
| Task | Use case · vertical | Type | Items | Jev | JevFish | NanoJev | OpenJev | System One Scorer | decider-2b | Data |
|---|---|---|---|---|---|---|---|---|---|---|
abt-buy-product-match | entity alignment · e commerce | score | 200 | 94.9 93% | 82.7 89% | 74.5 67% | 92.9 92% | 85.4 87% | 82.7 91% | JSON |
casehold-cited-holding | citation verification · legal | choice | 200 | 87.0 80% | 75.8 66% | 61.2 34% · 198/200 | 74.7 67% | 69.5 52% | 80.0 71% | JSON |
esci-product-rerank | semantic reranking · e commerce | noul | 490 | 75.0 73% | 68.2 57% | 71.1 59% · 489/490 | 71.8 72% | 75.5 62% | 74.2 62% | JSON |
hotpot-context-filter | rag passage filtering · knowledge work | noul | 200 | 96.3 95% | 87.6 80% | 0.0 — · 0/200 | 85.7 93% | 75.4 59% | 86.7 84% | JSON |
trialgpt-criterion-support | citation verification · healthcare | choice | 159 | 73.8 64% | 70.5 53% | 58.7 13% | 60.6 53% | 70.6 53% | 73.3 58% | JSON |
Data and operations
| Task | Use case · vertical | Type | Items | Jev | JevFish | NanoJev | OpenJev | System One Scorer | decider-2b | Data |
|---|---|---|---|---|---|---|---|---|---|---|
arxiv-archive-subcategory | hierarchical classification · science publishing | choice | 200 | 76.8 89% | 69.3 79% | 53.7 30% | 65.4 87% | 70.6 82% | 70.1 78% | JSON |
dbpedia-ontology-hierarchy | hierarchical classification · media and reference | choice | 200 | 92.9 94% | 89.4 89% | 60.6 16% | 89.3 93% | 83.8 87% | 89.0 89% | JSON |
finer-xbrl-fact-tag | structured extraction · finance | choice | 181 | 93.3 92% | 78.4 72% | 53.0 14% | 89.0 85% | 79.6 80% | 79.8 73% | JSON |
trial-outcomes | feature extraction for ml · saas growth | noul | 400 | 78.8 68% | 72.2 70% | 76.4 70% | 69.4 69% | 77.7 70% | 77.8 70% | JSON |
Real-time and agents
| Task | Use case · vertical | Type | Items | Jev | JevFish | NanoJev | OpenJev | System One Scorer | decider-2b | Data |
|---|---|---|---|---|---|---|---|---|---|---|
bfcl-tool-dispatch | agent routing and skill selection · software engineering | choice | 200 | 97.6 97% | 95.7 95% | 61.0 45% · 196/200 | 96.5 96% | 83.2 79% | 92.1 89% | JSON |
k8s-issue-kind-routing | agent routing and skill selection · devops | choice | 200 | 77.5 73% | 73.6 62% | 61.5 34% | 73.8 71% | 73.5 61% | 75.4 66% | JSON |
swe-rebench-run-resolved | run judging · software engineering | noul | 200 | 74.9 60% | 76.2 67% | 0.0 — · 0/200 | 59.2 56% | 73.5 50% | 69.4 52% | JSON |
swe-smith-run-resolved | run judging · software engineering | noul | 200 | 60.2 51% | 81.2 72% | 0.3 0% · 1/200 | 51.1 50% | 71.8 50% | 69.8 46% | JSON |
tau2-airline-run-success | run judging · customer operations | noul | 200 | 68.0 52% | 54.9 47% | 75.5 50% · 198/200 | 55.3 48% | 71.8 49% | 67.3 47% | JSON |
tau2-airline-transcript-success | run judging · customer operations | noul | 199 | 76.7 61% | 54.2 44% | 1.4 67% · 3/199 | 63.7 49% | 74.8 53% | 63.9 32% | JSON |
Workflow control
| Task | Use case · vertical | Type | Items | Jev | JevFish | NanoJev | OpenJev | System One Scorer | decider-2b | Data |
|---|---|---|---|---|---|---|---|---|---|---|
cfpb-complaint-product-queue | support inbox triage · finance | choice | 199 | 83.5 79% | 73.6 64% | 63.4 44% | 78.9 77% | 82.6 78% | 79.7 73% | JSON |
hwu64-tool-select | typed tool dispatch · consumer assistant | choice | 200 | 93.7 88% | 90.2 85% | 67.8 36% | 91.5 87% | 83.4 81% | 91.2 86% | JSON |
ledgar-clause-review-routing | intent and model routing · legal | choice | 199 | 34.5 29% | 44.7 26% | 50.8 14% | 32.8 28% | 48.9 20% | 37.3 25% | JSON |
unfair-tos-publish-gate | confidence gated actions · consumer legal | noul | 198 | 87.1 84% | 79.7 67% | 76.0 63% | 83.9 82% | 83.9 82% | 83.3 76% | JSON |
Safety and quality
| Task | Use case · vertical | Type | Items | Jev | JevFish | NanoJev | OpenJev | System One Scorer | decider-2b | Data |
|---|---|---|---|---|---|---|---|---|---|---|
cuad-checklist | compliance verification · legal | noul | 204 | 91.6 89% | 86.3 81% | 72.7 50% | 91.2 89% | 92.1 93% | 95.1 95% | JSON |
humanevalpack-spec-match | semantic linting · software engineering | noul | 164 | 82.8 89% | 69.4 63% | 68.1 50% · 163/164 | 80.5 88% | 73.2 66% | 71.4 65% | JSON |
moderation-hate-severity | llm guardrails · trust and safety | score | 200 | 93.4 85% | 84.5 65% | 75.0 35% | 90.3 78% | 86.3 74% | 89.6 78% | JSON |
prompt-injection | llm guardrails · security | noul | 194 | 90.5 88% | 81.5 72% | 77.9 64% | 85.4 84% | 82.4 76% | 76.3 70% | JSON |
unfair-tos-clause-type | compliance verification · consumer legal | choice | 200 | 85.9 82% | 90.3 90% | 60.8 26% | 83.9 81% | 79.1 78% | 83.5 80% | JSON |
vitaminc-claim-consistency | self consistency checks · media | choice | 200 | 88.1 83% | 78.1 63% | 60.8 51% | 82.8 77% | 82.3 76% | 82.4 72% | JSON |