Skip to benchmark
Blobfish ResearchDealBench-100 v1.0.0Public release

Can an agent run the deal—not just calculate one valuation cell?

DealBench-100 measures transaction execution across ten synthetic deal worlds. The agent must find the operative sources, reconcile stale and current evidence, calculate the supported answer, compare process options, update the live model and committee deliverable, commit the plan, and leave an honest review-ready handoff.

100analyst workflows
10synthetic deal worlds
36provider-shaped tools
440released source assets
0LLM grading calls
Release gate passed

The tasks are solvable; the qualification controls are not model rows. All 100 oracle episodes reached 100 DealScore, all 100/100 deterministic replays matched, and 500 adversarial episodes produced 0 strict false accepts. A ranked model appears only after a complete version-pinned run on this exact release.

Assets per task
26
Evidence reads
15
Exact criteria
35
Native formats
10
Public design anchors

We used the public APEX product surfaces and Archipelago execution architecture to define what should be inspectable, then authored a separate synthetic corpus and deterministic verifier. Mercor's gated APEX-Agents dataset was not downloaded or scraped.

Measured performance

Leaderboard

DealScore is the mean deterministic outcome score across all 100 tasks. Qualification controls stay below the ranked table so an oracle can never masquerade as a model result.

Ranking

Model leaderboard

Only complete, version-pinned model runs on this exact release receive a rank.

RankAgentCoverageScoreRun
No complete version-pinned model run is published for DealBench-100 yet. The executed controls below establish a 100-point ceiling and five distinct failure floors without receiving ranks.
Qualification evidence

Measured controls

Oracle and adversarial replays prove solvability and verifier discrimination. They are measured, but never ranked as model submissions.

ClassControlCoverageScoreEvidence
REF
Reference oracleDeterministic release qualification
100 tasks
100.0%
29.2 calls avg.Solvability ceiling; not a model submission.
CTRL
No-op controlDeterministic release qualification
100 tasks
4.0%
calls avg.Executed adversarial control; never ranked with models.
CTRL
Answer-only shortcutDeterministic release qualification
100 tasks
46.0%
calls avg.Executed adversarial control; never ranked with models.
CTRL
State-only shortcutDeterministic release qualification
100 tasks
85.0%
calls avg.Executed adversarial control; never ranked with models.
CTRL
Wrong-source controlDeterministic release qualification
100 tasks
83.5%
calls avg.Executed adversarial control; never ranked with models.
CTRL
Wrong-target controlDeterministic release qualification
100 tasks
85.5%
calls avg.Executed adversarial control; never ranked with models.
15%Discovery
25%Model accuracy
15%Decision
20%Committed state
10%Deliverable
10%Readback
5%Containment

100 distinct tasks

Task browser

Every task is an independently authored, high-level analyst request over one frozen transaction world. Open any row to inspect the prompt, 26-file evidence room, operating alternatives, and all 35 checks.

100 shown
dealbench-001Source control and launch readiness

Confirm the operative deal case

**PROJECT-LANTERN · LumaWorks Lighting · Source control and launch readiness** The VP needs the operative case locked before the team changes another number. Establish which forecast, QoE bridge, and approval are current; explain what makes the other versions stale; then commit the supported working case and prepare the deal-team handoff for review. Use the evidence available in the connected systems. Leave unrelated projects and records unchanged, keep outbound communication in review status, and do not claim completion until the durable state and readbacks agree.

Company
PROJECT-LANTERN
As of
2026-01-16T09:00:00Z
Seeded files
26
Reference solution
29 MCP calls

Complete task prompt

35 graded criteria summarized
**PROJECT-LANTERN · LumaWorks Lighting · Source control and launch readiness**

The VP needs the operative case locked before the team changes another number. Establish which forecast, QoE bridge, and approval are current; explain what makes the other versions stale; then commit the supported working case and prepare the deal-team handoff for review.

Use the evidence available in the connected systems. Leave unrelated projects and records unchanged, keep outbound communication in review status, and do not claim completion until the durable state and readbacks agree.
How the employee outcome is evaluated

Reasoning, persisted state, and the answer must agree.

7 semantic milestones

A source-authority and transaction-math chain ending in controlled model, deliverable, plan, and review-only communication state.

Strict success
The exact current sources, calculations, decision, model/deck state, review handoff, readbacks, and containment all agree.
Ordering policy
Exact call order is not graded; required investigations must precede the first controlled write and each write must be read back.
Inspect the task-specific causal milestones
  1. discovery
    discovery

    Discovery contributes 15 DealScore points.

  2. model_accuracy
    model_accuracy

    Model accuracy contributes 25 DealScore points.

  3. decision
    decision

    Decision contributes 15 DealScore points.

  4. committed_state
    committed_state

    Committed state contributes 20 DealScore points.

  5. deliverable
    deliverable

    Deliverable contributes 10 DealScore points.

  6. readback
    readback

    Readback contributes 10 DealScore points.

  7. containment
    containment

    Containment contributes 5 DealScore points.

Decision space (3 grounded options)
  • Supported current-evidence path — selected: Matches current source authority, calculation, approval, and process constraints.
  • Use the headline shortcut: Ignores version authority, risk, funding, or diligence constraints.
  • Stop without analysis: Fails to distinguish a real gate from a resolvable evidence issue.
Inspect all 35 deterministic criteria
  • discovery: Complete the task contract investigation before any controlled write. (1 pts)
  • discovery: Complete the project record investigation before any controlled write. (1 pts)
  • discovery: Complete the data room search investigation before any controlled write. (1 pts)
  • discovery: Complete the current forecast investigation before any controlled write. (1 pts)
  • discovery: Complete the prior forecast investigation before any controlled write. (1 pts)
  • discovery: Complete the version history investigation before any controlled write. (1 pts)
  • discovery: Complete the request mail search investigation before any controlled write. (1 pts)
  • discovery: Complete the request mail investigation before any controlled write. (1 pts)
  • discovery: Complete the team chat search investigation before any controlled write. (1 pts)
  • discovery: Complete the team thread investigation before any controlled write. (1 pts)
  • discovery: Complete the workbook index investigation before any controlled write. (1 pts)
  • discovery: Complete the workbook inputs investigation before any controlled write. (1 pts)
  • discovery: Complete the live model investigation before any controlled write. (1 pts)
  • discovery: Complete the deliverable investigation before any controlled write. (1 pts)
  • discovery: Complete the approval investigation before any controlled write. (1 pts)
  • model_accuracy: Submit the exact source-grounded headline value usd m value (86.09). (5 pts)
  • model_accuracy: Submit the exact source-grounded equity value usd m value (-21.41). (5 pts)
  • model_accuracy: Submit the exact source-grounded per share value usd value (-0.51). (5 pts)
  • model_accuracy: Submit the exact source-grounded primary metric value (19.66). (5 pts)
  • model_accuracy: Submit the exact source-grounded secondary metric value (7.2). (5 pts)
  • decision: Select the supported option lock_current_case. (10 pts)
  • decision: Ground the decision in operative source REV-01-CURRENT. (5 pts)
  • committed_state: Persist the exact model revision, outputs, and source references. (6 pts)
  • committed_state: Persist the exact values to the controlled workbook range. (4 pts)
  • committed_state: Commit the supported plan with status, rationale, model, deliverable, and source refs. (4 pts)
  • committed_state: Persist the task-specific bid, diligence, approval, or source-linked model state required by this workflow. (2 pts)
  • committed_state: Persist the review-only mail and deal-team handoff. (4 pts)
  • deliverable: Update only the controlled deliverable to the expected revision. (5 pts)
  • deliverable: Tie the deliverable values exactly to the live model outputs. (5 pts)
  • readback: Read back the model and workbook after their writes. (3 pts)
  • readback: Read back the plan and deliverable after commit. (3 pts)
  • readback: Read back sent mail and channel history before submission. (2 pts)
  • readback: Read back the durable task submission. (2 pts)
  • containment: Leave every unrelated project and record unchanged. (3 pts)
  • containment: Keep external communications in draft-for-review status. (2 pts)
Open released task dataset

Seeded evidence

Asset room

Current and superseded forecasts, QoE, comps, precedents, debt schedules, bid letters, diligence, approvals, models, decks, mail, chat, permissions, and audit records are all visible.

01-request-email.emlassets/project-lanternCorroboratingEML165 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
02-current-forecast.xlsxassets/project-lanternOperativeXLSX2.9 KB
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
03-prior-forecast.xlsxassets/project-lanternCorroboratingXLSX2.9 KB
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
04-quality-of-earnings.pdfassets/project-lanternCorroboratingPDF798 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
05-management-case.csvassets/project-lanternCorroboratingCSV63 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
06-current-assumptions.jsonassets/project-lanternOperativeJSON149 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
07-retired-assumptions.jsonassets/project-lanternCorroboratingJSON99 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
08-trading-comps.csvassets/project-lanternCorroboratingCSV155 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
09-precedent-transactions.csvassets/project-lanternCorroboratingCSV199 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
10-debt-schedule.xlsxassets/project-lanternCorroboratingXLSX2.6 KB
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
11-bid-letters.pdfassets/project-lanternCorroboratingPDF865 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
12-diligence-log.xlsxassets/project-lanternCorroboratingXLSX2.6 KB
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
13-approval-policy.mdassets/project-lanternCorroboratingMD202 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
14-current-approval.emlassets/project-lanternOperativeEML132 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
15-retired-approval.emlassets/project-lanternCorroboratingEML129 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
16-deal-team-thread.jsonassets/project-lanternCorroboratingJSON249 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
17-model-change-log.csvassets/project-lanternCorroboratingCSV90 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
18-committee-deck-current.pptxassets/project-lanternOperativePPTX2.8 KB
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
19-committee-deck-prior.pptxassets/project-lanternCorroboratingPPTX2.8 KB
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
20-client-update-draft.mdassets/project-lanternCorroboratingMD193 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
21-source-map.yamlassets/project-lanternCorroboratingYAML161 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
22-data-room-permissions.jsonassets/project-lanternCorroboratingJSON190 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
23-audit.logassets/project-lanternCorroboratingLOG101 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
24-deal-timeline.icsassets/project-lanternCorroboratingICS170 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
task-brief.mdassets/tasks/dealbench-001OperativeMD683 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file
starting-snapshot.jsonassets/tasks/dealbench-001CorroboratingJSON265 B
Agent-visible synthetic evidence; current and superseded records are deliberately mixed.Open released file

Runnable world

Environment and tool contract

Seven logical MCP servers expose one isolated SQLite snapshot. Every controlled write is durable, task-scoped, and checked against the complete before/after state.

MCP package pindealbench100@1.0.0
MCP protocol2025-06-18
Catalog SHA-2560032864f237cd9c52f5bebe86f1893f058e819941a476d1b1bcab38882385704
Harbor package SHA-2562dacc4693638ed89d5b5e3822184728a1597646b1c0a8a9bc8682e26b3b47449
Hugging Face commit0c4f25f561b4d5a85687bec83b8b581c7d3f7f1f
Synthetic worldatlas-deal-team-v1
benchmark.get_task
Read onlyIdempotentClosed sandbox

Read the task-scoped outcome contract without revealing the gold answer.

{
  "additionalProperties": false,
  "properties": {
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    }
  },
  "required": [
    "task_id"
  ],
  "type": "object"
}
benchmark.get_submission
Read onlyIdempotentClosed sandbox

Read back the durable submitted answer.

{
  "additionalProperties": false,
  "properties": {
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    }
  },
  "required": [
    "task_id"
  ],
  "type": "object"
}
dealroom.search_files
Read onlyIdempotentClosed sandbox

Search the synthetic data room by project and free-text query.

{
  "additionalProperties": false,
  "properties": {
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    },
    "query": {
      "description": "Search query",
      "type": "string"
    }
  },
  "required": [
    "project_code",
    "query"
  ],
  "type": "object"
}
dealroom.get_file
Read onlyIdempotentClosed sandbox

Read one data-room file record, including authority and contents.

{
  "additionalProperties": false,
  "properties": {
    "file_id": {
      "description": "File identifier",
      "type": "string"
    }
  },
  "required": [
    "file_id"
  ],
  "type": "object"
}
dealroom.get_version_history
Read onlyIdempotentClosed sandbox

List all current and superseded versions for a logical file.

{
  "additionalProperties": false,
  "properties": {
    "logical_name": {
      "description": "Logical document name",
      "type": "string"
    }
  },
  "required": [
    "logical_name"
  ],
  "type": "object"
}
dealroom.get_permissions
Read onlyIdempotentClosed sandbox

Read permissions and controlled-edit status for a data-room file.

{
  "additionalProperties": false,
  "properties": {
    "file_id": {
      "description": "File identifier",
      "type": "string"
    }
  },
  "required": [
    "file_id"
  ],
  "type": "object"
}
mail.search_messages
Read onlyIdempotentClosed sandbox

Search task-scoped mailbox messages.

{
  "additionalProperties": false,
  "properties": {
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    },
    "query": {
      "description": "Search query",
      "type": "string"
    }
  },
  "required": [
    "project_code",
    "query"
  ],
  "type": "object"
}
mail.get_message
Read onlyIdempotentClosed sandbox

Read one mailbox message.

{
  "additionalProperties": false,
  "properties": {
    "message_id": {
      "description": "Message identifier",
      "type": "string"
    }
  },
  "required": [
    "message_id"
  ],
  "type": "object"
}
mail.list_sent
Read onlyIdempotentClosed sandbox

Read back sent draft messages for a task.

{
  "additionalProperties": false,
  "properties": {
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    }
  },
  "required": [
    "project_code",
    "task_id"
  ],
  "type": "object"
}
chat.search_messages
Read onlyIdempotentClosed sandbox

Search a deal-team channel.

{
  "additionalProperties": false,
  "properties": {
    "channel": {
      "description": "Channel",
      "type": "string"
    },
    "query": {
      "description": "Search query",
      "type": "string"
    }
  },
  "required": [
    "channel",
    "query"
  ],
  "type": "object"
}
chat.get_thread
Read onlyIdempotentClosed sandbox

Read one deal-team thread.

{
  "additionalProperties": false,
  "properties": {
    "thread_id": {
      "description": "Thread identifier",
      "type": "string"
    }
  },
  "required": [
    "thread_id"
  ],
  "type": "object"
}
chat.get_channel_history
Read onlyIdempotentClosed sandbox

Read back task posts in a channel.

{
  "additionalProperties": false,
  "properties": {
    "channel": {
      "description": "Channel",
      "type": "string"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    }
  },
  "required": [
    "channel",
    "task_id"
  ],
  "type": "object"
}
sheets.list_workbooks
Read onlyIdempotentClosed sandbox

List controlled workbooks for a project.

{
  "additionalProperties": false,
  "properties": {
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    }
  },
  "required": [
    "project_code"
  ],
  "type": "object"
}
sheets.read_range
Read onlyIdempotentClosed sandbox

Read a range from the synthetic model workbook.

{
  "additionalProperties": false,
  "properties": {
    "range": {
      "description": "A1 range",
      "type": "string"
    },
    "workbook_id": {
      "description": "Workbook identifier",
      "type": "string"
    }
  },
  "required": [
    "workbook_id",
    "range"
  ],
  "type": "object"
}
sheets.get_change_log
Read onlyIdempotentClosed sandbox

Read back controlled workbook writes.

{
  "additionalProperties": false,
  "properties": {
    "workbook_id": {
      "description": "Workbook identifier",
      "type": "string"
    }
  },
  "required": [
    "workbook_id"
  ],
  "type": "object"
}
markets.get_company
Read onlyIdempotentClosed sandbox

Read current buyer or target market data.

{
  "additionalProperties": false,
  "properties": {
    "company_id": {
      "description": "Company identifier",
      "type": "string"
    }
  },
  "required": [
    "company_id"
  ],
  "type": "object"
}
markets.list_comparables
Read onlyIdempotentClosed sandbox

Read the approved trading-comparables set.

{
  "additionalProperties": false,
  "properties": {
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    }
  },
  "required": [
    "project_code"
  ],
  "type": "object"
}
markets.list_transactions
Read onlyIdempotentClosed sandbox

Read the approved precedent-transactions set.

{
  "additionalProperties": false,
  "properties": {
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    }
  },
  "required": [
    "project_code"
  ],
  "type": "object"
}
markets.get_credit_curve
Read onlyIdempotentClosed sandbox

Read the frozen financing curve used by the model.

{
  "additionalProperties": false,
  "properties": {
    "currency": {
      "description": "ISO currency",
      "type": "string"
    }
  },
  "required": [
    "currency"
  ],
  "type": "object"
}
deals.get_project
Read onlyIdempotentClosed sandbox

Read the master deal record and current process status.

{
  "additionalProperties": false,
  "properties": {
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    }
  },
  "required": [
    "project_code"
  ],
  "type": "object"
}
deals.get_model
Read onlyIdempotentClosed sandbox

Read the live model revision and outputs.

{
  "additionalProperties": false,
  "properties": {
    "model_id": {
      "description": "Model identifier",
      "type": "string"
    }
  },
  "required": [
    "model_id"
  ],
  "type": "object"
}
deals.list_bids
Read onlyIdempotentClosed sandbox

Read all current bids with price, certainty and conditions.

{
  "additionalProperties": false,
  "properties": {
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    }
  },
  "required": [
    "project_code"
  ],
  "type": "object"
}
deals.list_diligence_findings
Read onlyIdempotentClosed sandbox

Read current diligence findings.

{
  "additionalProperties": false,
  "properties": {
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    }
  },
  "required": [
    "project_code"
  ],
  "type": "object"
}
deals.get_approval
Read onlyIdempotentClosed sandbox

Read current approval authority and status.

{
  "additionalProperties": false,
  "properties": {
    "approval_id": {
      "description": "Approval identifier",
      "type": "string"
    }
  },
  "required": [
    "approval_id"
  ],
  "type": "object"
}
deals.get_deliverable
Read onlyIdempotentClosed sandbox

Read the controlled committee deliverable.

{
  "additionalProperties": false,
  "properties": {
    "deliverable_id": {
      "description": "Deliverable identifier",
      "type": "string"
    }
  },
  "required": [
    "deliverable_id"
  ],
  "type": "object"
}
deals.get_committed_plan
Read onlyIdempotentClosed sandbox

Read back the committed plan for one task.

{
  "additionalProperties": false,
  "properties": {
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    }
  },
  "required": [
    "project_code",
    "task_id"
  ],
  "type": "object"
}
sheets.write_range
State changingNon-idempotentClosed sandbox

Write one controlled output range; all other cells are preserved.

{
  "additionalProperties": false,
  "properties": {
    "range": {
      "description": "A1 range",
      "type": "string"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    },
    "values": {
      "description": "Two-dimensional values",
      "items": {},
      "type": "array"
    },
    "workbook_id": {
      "description": "Workbook identifier",
      "type": "string"
    }
  },
  "required": [
    "workbook_id",
    "range",
    "values",
    "task_id"
  ],
  "type": "object"
}
deals.update_model
State changingNon-idempotentClosed sandbox

Commit a source-bound model revision and exact outputs.

{
  "additionalProperties": false,
  "properties": {
    "model_id": {
      "description": "Model identifier",
      "type": "string"
    },
    "outputs": {
      "additionalProperties": true,
      "description": "Calculated outputs",
      "type": "object"
    },
    "revision": {
      "description": "New revision",
      "type": "string"
    },
    "source_refs": {
      "items": {
        "type": "string"
      },
      "type": "array"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    }
  },
  "required": [
    "model_id",
    "revision",
    "outputs",
    "source_refs",
    "task_id"
  ],
  "type": "object"
}
deals.update_bid_status
State changingNon-idempotentClosed sandbox

Update only the selected bid's review status.

{
  "additionalProperties": false,
  "properties": {
    "bid_id": {
      "description": "Bid identifier",
      "type": "string"
    },
    "rationale": {
      "description": "Source-grounded rationale",
      "type": "string"
    },
    "status": {
      "description": "Review status",
      "type": "string"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    }
  },
  "required": [
    "bid_id",
    "status",
    "rationale",
    "task_id"
  ],
  "type": "object"
}
deals.update_diligence_finding
State changingNon-idempotentClosed sandbox

Resolve one diligence finding into the current model.

{
  "additionalProperties": false,
  "properties": {
    "finding_id": {
      "description": "Finding identifier",
      "type": "string"
    },
    "resolution": {
      "description": "Resolution",
      "type": "string"
    },
    "status": {
      "description": "Status",
      "type": "string"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    }
  },
  "required": [
    "finding_id",
    "status",
    "resolution",
    "task_id"
  ],
  "type": "object"
}
deals.request_approval
State changingNon-idempotentClosed sandbox

Record an approval-state transition with evidence.

{
  "additionalProperties": false,
  "properties": {
    "approval_id": {
      "description": "Approval identifier",
      "type": "string"
    },
    "evidence_refs": {
      "items": {
        "type": "string"
      },
      "type": "array"
    },
    "status": {
      "description": "Approval status",
      "type": "string"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    }
  },
  "required": [
    "approval_id",
    "status",
    "evidence_refs",
    "task_id"
  ],
  "type": "object"
}
deals.update_deliverable
State changingNon-idempotentClosed sandbox

Update only the controlled output values and revision in the committee deck.

{
  "additionalProperties": false,
  "properties": {
    "deliverable_id": {
      "description": "Deliverable identifier",
      "type": "string"
    },
    "revision": {
      "description": "Revision",
      "type": "string"
    },
    "status": {
      "description": "Review status",
      "type": "string"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    },
    "values": {
      "additionalProperties": true,
      "description": "Exact model-linked values",
      "type": "object"
    }
  },
  "required": [
    "deliverable_id",
    "revision",
    "values",
    "status",
    "task_id"
  ],
  "type": "object"
}
deals.commit_plan
State changingNon-idempotentClosed sandbox

Commit the supported deal plan and its evidence lineage.

{
  "additionalProperties": false,
  "properties": {
    "decision": {
      "description": "Selected option",
      "type": "string"
    },
    "deliverable_id": {
      "description": "Deliverable identifier",
      "type": "string"
    },
    "model_id": {
      "description": "Model identifier",
      "type": "string"
    },
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    },
    "rationale": {
      "description": "Rationale",
      "type": "string"
    },
    "source_refs": {
      "items": {
        "type": "string"
      },
      "type": "array"
    },
    "status": {
      "description": "Decision status",
      "type": "string"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    }
  },
  "required": [
    "project_code",
    "task_id",
    "decision",
    "status",
    "rationale",
    "source_refs",
    "model_id",
    "deliverable_id"
  ],
  "type": "object"
}
mail.send_message
State changingNon-idempotentClosed sandbox

Save a deal-team email as a review-only draft.

{
  "additionalProperties": false,
  "properties": {
    "body": {
      "description": "Message body",
      "type": "string"
    },
    "project_code": {
      "description": "Deal project code",
      "type": "string"
    },
    "review_status": {
      "description": "Must remain draft_for_review",
      "type": "string"
    },
    "subject": {
      "description": "Subject",
      "type": "string"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    },
    "to": {
      "description": "Recipient",
      "type": "string"
    }
  },
  "required": [
    "project_code",
    "to",
    "subject",
    "body",
    "review_status",
    "task_id"
  ],
  "type": "object"
}
chat.post_message
State changingNon-idempotentClosed sandbox

Save a deal-team channel handoff in review status.

{
  "additionalProperties": false,
  "properties": {
    "channel": {
      "description": "Channel",
      "type": "string"
    },
    "review_status": {
      "description": "Must remain draft_for_review",
      "type": "string"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    },
    "text": {
      "description": "Message",
      "type": "string"
    }
  },
  "required": [
    "channel",
    "text",
    "review_status",
    "task_id"
  ],
  "type": "object"
}
benchmark.submit_answer
State changingNon-idempotentClosed sandbox

Persist the structured final answer for deterministic grading.

{
  "additionalProperties": false,
  "properties": {
    "answers": {
      "additionalProperties": true,
      "description": "Task-specific answer",
      "type": "object"
    },
    "task_id": {
      "description": "Task identifier",
      "type": "string"
    }
  },
  "required": [
    "task_id",
    "answers"
  ],
  "type": "object"
}

Architecture comparison

APEX / Archipelago design anchor, deterministic DealBench implementation

The public architecture is preserved where it improves reproducibility; the task corpus and grading implementation are independent.

Inspect Archipelago
LayerPublic APEX / Archipelago patternDealBench-100
WorldData-rich professional project with files and appsTen frozen synthetic deal projects with 26 task-visible files
EnvironmentContainer, MCP gateway, snapshot populationHarbor task, seven MCP servers, task-local SQLite snapshot
TrajectoryMessages, tool calls, artifact edits, final snapshotFull provider-shaped calls, outputs, before/after state and verdict
GradingCriterion-level judge/static/domain verifiers35 executable math, lineage, state, readback and containment checks; no judge model
ReleaseWorld assets, gold outputs, task metadataHF mirror, Harbor dataset, gold contracts, digests and ten public oracle traces

Replay the work

Trajectories

Ten full reference trajectories show messages, provider-shaped calls, results, writes, readbacks, and the deterministic verdict. They prove execution, not model capability.

AgentReference oracleDeterministic release solver
Score100.0%Strict pass
Tool calls29Provider-native MCP trace
Run cost
31 events
  1. Employee request

    **PROJECT-LANTERN · LumaWorks Lighting · Source control and launch readiness** The VP needs the operative case locked before the team changes another number. Establish which forecast, QoE bridge, and approval are current; explain what makes the other versions stale; then commit the supported working case and prepare the deal-team handoff for review. Use the evidence available in the connected systems. Leave unrelated projects and records unchanged, keep outbound communication in review status, and do not claim completion until the durable state and readbacks agree.

  2. 1
    Scopebenchmark.get_taskok
    {
      "task_id": "dealbench-001"
    }

    {"allowed_write_tools": ["deals.update_model", "sheets.write_range", "deals.update_bid_status", "deals.update_diligence_finding", "deals.request_approval", "deals.update_deliverable", "deals.commit_plan", "mail.send_message", "chat.post_message", "benchmark.submit_answer"], "answer_schema": {"additionalProperties": false, "properties": {"decision_status": {"type": "string"}, "deliverable_revision": {"type": "string"}, "equity_value_usd_m": {"multipleOf": 0.01, "type": "number"}, "headline_value_…

  3. 2
    Investigatedeals.get_projectok
    {
      "project_code": "PROJECT-LANTERN"
    }

    {"company": "LumaWorks Lighting", "deal_type": "sell-side auction", "industry": "Commercial lighting", "payload": {"allowed_addbacks": 7.6, "approval_status": "approved", "buyer_net_income": 121.0, "buyer_share_price": 24.5, "buyer_shares": 83.0, "capex_pct": 0.031, "cash": 18.5, "cash_pct": 0.55, "comp_multiple": 9.4, "company": "LumaWorks Lighting", "critical_open": false, "deal_type": "sell-side auction", "debt": 126.0, "disallowed_addbacks": 2.1, "ebitda_margin": 0.184, "entry_multiple": 9.6…

  4. 3
    Investigatedealroom.search_filesok
    {
      "project_code": "PROJECT-LANTERN",
      "query": "source_control"
    }

    {"files": [{"file_id": "PROJECT-LANTERN-APPROVAL-CURRENT", "is_current": 1, "kind": "eml", "logical_name": "PROJECT-LANTERN-APPROVAL", "name": "Current transaction approval", "version": 2}, {"file_id": "PROJECT-LANTERN-DECK-CURRENT", "is_current": 1, "kind": "pptx", "logical_name": "PROJECT-LANTERN-DECK", "name": "Current committee deck", "version": 5}, {"file_id": "PROJECT-LANTERN-FORECAST-CURRENT", "is_current": 1, "kind": "xlsx", "logical_name": "PROJECT-LANTERN-FORECAST", "name": "Current ma…

  5. 4
    Investigatedealroom.get_fileok
    {
      "file_id": "PROJECT-LANTERN-FORECAST-CURRENT"
    }

    {"content": {"approved": true, "growth": 0.072, "revenue": 438.0, "revision": "REV-01-CURRENT"}, "file_id": "PROJECT-LANTERN-FORECAST-CURRENT", "is_current": 1, "kind": "xlsx", "logical_name": "PROJECT-LANTERN-FORECAST", "name": "Current management forecast", "project_code": "PROJECT-LANTERN", "version": 3}

  6. 5
    Investigatedealroom.get_fileok
    {
      "file_id": "PROJECT-LANTERN-FORECAST-PRIOR"
    }

    {"content": {"approved": false, "growth": 0.06, "revenue": 424.86, "revision": "REV-ERN-PRIOR"}, "file_id": "PROJECT-LANTERN-FORECAST-PRIOR", "is_current": 0, "kind": "xlsx", "logical_name": "PROJECT-LANTERN-FORECAST", "name": "Superseded management forecast", "project_code": "PROJECT-LANTERN", "version": 2}

  7. 6
    Investigatedealroom.get_version_historyok
    {
      "logical_name": "PROJECT-LANTERN-FORECAST"
    }

    {"versions": [{"content": {"approved": true, "growth": 0.072, "revenue": 438.0, "revision": "REV-01-CURRENT"}, "file_id": "PROJECT-LANTERN-FORECAST-CURRENT", "is_current": 1, "kind": "xlsx", "name": "Current management forecast", "version": 3}, {"content": {"approved": false, "growth": 0.06, "revenue": 424.86, "revision": "REV-ERN-PRIOR"}, "file_id": "PROJECT-LANTERN-FORECAST-PRIOR", "is_current": 0, "kind": "xlsx", "name": "Superseded management forecast", "version": 2}]}

  8. 7
    Investigatemail.search_messagesok
    {
      "project_code": "PROJECT-LANTERN",
      "query": "dealbench-001"
    }

    {"messages": [{"message_id": "MSG-dealbench-001-REQUEST", "sender": "vp@atlas-sim.example", "sent_at": "2026-01-16T09:00:00Z", "subject": "dealbench-001 operating request", "task_id": "dealbench-001"}, {"message_id": "MSG-dealbench-001-AUTHORITY", "sender": "deal-control@atlas-sim.example", "sent_at": "2026-01-16T09:00:00Z", "subject": "Current source and review authority", "task_id": "dealbench-001"}, {"message_id": "MSG-dealbench-001-DECOY", "sender": "former-analyst@atlas-sim.example", "sent_…

  9. 8
    Investigatemail.get_messageok
    {
      "message_id": "MSG-dealbench-001-REQUEST"
    }

    {"body": "**PROJECT-LANTERN · LumaWorks Lighting · Source control and launch readiness**\n\nThe VP needs the operative case locked before the team changes another number. Establish which forecast, QoE bridge, and approval are current; explain what makes the other versions stale; then commit the supported working case and prepare the deal-team handoff for review.\n\nUse the evidence available in the connected systems. Leave unrelated projects and records unchanged, keep outbound communication in …

  10. 9
    Investigatechat.search_messagesok
    {
      "channel": "deal-project-lantern",
      "query": "dealbench-001"
    }

    {"messages": [{"author": "vp", "channel": "deal-project-lantern", "message_id": "CHAT-dealbench-001-1", "posted_at": "2026-01-16T09:00:00Z", "task_id": "dealbench-001", "text": "Please resolve dealbench-001 against current authority.", "thread_id": "THREAD-dealbench-001"}, {"author": "deal-control", "channel": "deal-project-lantern", "message_id": "CHAT-dealbench-001-2", "posted_at": "2026-01-16T09:00:00Z", "task_id": "dealbench-001", "text": "Current source is REV-01-CURRENT; keep the external …

  11. 10
    Investigatechat.get_threadok
    {
      "thread_id": "THREAD-dealbench-001"
    }

    {"messages": [{"author": "vp", "channel": "deal-project-lantern", "message_id": "CHAT-dealbench-001-1", "posted_at": "2026-01-16T09:00:00Z", "task_id": "dealbench-001", "text": "Please resolve dealbench-001 against current authority.", "thread_id": "THREAD-dealbench-001"}, {"author": "deal-control", "channel": "deal-project-lantern", "message_id": "CHAT-dealbench-001-2", "posted_at": "2026-01-16T09:00:00Z", "task_id": "dealbench-001", "text": "Current source is REV-01-CURRENT; keep the external …

  12. 11
    Investigatesheets.list_workbooksok
    {
      "project_code": "PROJECT-LANTERN"
    }

    {"workbooks": [{"name": "LumaWorks Lighting live deal model", "project_code": "PROJECT-LANTERN", "revision": "R2", "workbook_id": "WB-PROJECT-LANTERN"}]}

  13. 12
    Investigatesheets.read_rangeok
    {
      "range": "Inputs!A1:H20",
      "workbook_id": "WB-PROJECT-LANTERN"
    }

    {"range": "Inputs!A1:H20", "values": {"allowed_addbacks": 7.6, "cash": 18.5, "debt": 126.0, "disallowed_addbacks": 2.1, "ebitda_margin": 0.184, "revenue": 438.0, "terminal_growth": 0.025, "wacc": 0.101}, "workbook_id": "WB-PROJECT-LANTERN"}

  14. 13
    Investigatedeals.get_modelok
    {
      "model_id": "MODEL-PROJECT-LANTERN"
    }

    {"last_task_id": null, "model_id": "MODEL-PROJECT-LANTERN", "outputs": {"headline_value_usd_m": 0}, "project_code": "PROJECT-LANTERN", "revision": "MODEL-PRIOR-R2", "source_refs": ["REV-ERN-PRIOR"], "status": "working"}

  15. 14
    Investigatedeals.get_deliverableok
    {
      "deliverable_id": "DECK-PROJECT-LANTERN"
    }

    {"deliverable_id": "DECK-PROJECT-LANTERN", "last_task_id": null, "project_code": "PROJECT-LANTERN", "revision": "DECK-PRIOR-R1", "status": "working", "unrelated_slides_sha256": "de291458f9d4e4aaefbe809ae643a19d9a694b94e5de0585d83105607aacff08", "values": {"headline_value_usd_m": 0}}

  16. 15
    Investigatedeals.get_approvalok
    {
      "approval_id": "APR-PROJECT-LANTERN"
    }

    {"approval_id": "APR-PROJECT-LANTERN", "authority": "Transaction Committee", "evidence_refs": ["PROJECT-LANTERN-APPROVAL-CURRENT"], "last_task_id": null, "project_code": "PROJECT-LANTERN", "status": "approved"}

  17. 16
    Modeldeals.update_modelok
    {
      "model_id": "MODEL-PROJECT-LANTERN",
      "outputs": {
        "equity_value_usd_m": -21.41,
        "headline_value_usd_m": 86.09,
        "per_share_value_usd": -0.51,
        "primary_metric": 19.66,
        "secondary_metric": 7.2
      },
      "revision": "MODEL-01-SOUR-R3",
      "source_refs": [
        "REV-01-CURRENT",
        "PROJECT-LANTERN-QOE-CURRENT",
        "MSG-dealbench-001-REQUEST",
        "THREAD-dealbench-001"
      ],
      "task_id": "dealbench-001"
    }

    {"last_task_id": "dealbench-001", "model_id": "MODEL-PROJECT-LANTERN", "outputs": {"equity_value_usd_m": -21.41, "headline_value_usd_m": 86.09, "per_share_value_usd": -0.51, "primary_metric": 19.66, "secondary_metric": 7.2}, "project_code": "PROJECT-LANTERN", "revision": "MODEL-01-SOUR-R3", "source_refs": ["REV-01-CURRENT", "PROJECT-LANTERN-QOE-CURRENT", "MSG-dealbench-001-REQUEST", "THREAD-dealbench-001"], "status": "updated_for_review"}

  18. 17
    Modelsheets.write_rangeok
    {
      "range": "Outputs!B2:F3",
      "task_id": "dealbench-001",
      "values": [
        [
          "MODEL-01-SOUR-R3",
          86.09,
          -21.41,
          -0.51,
          19.66,
          7.2
        ]
      ],
      "workbook_id": "WB-PROJECT-LANTERN"
    }

    {"change_id": "CHANGE-dealbench-001", "updatedRange": "Outputs!B2:F3", "updatedRows": 1}

  19. 18
    Deliverdeals.update_deliverableok
    {
      "deliverable_id": "DECK-PROJECT-LANTERN",
      "revision": "DECK-01-SOUR-R2",
      "status": "draft_for_review",
      "task_id": "dealbench-001",
      "values": {
        "equity_value_usd_m": -21.41,
        "headline_value_usd_m": 86.09,
        "per_share_value_usd": -0.51,
        "primary_metric": 19.66,
        "secondary_metric": 7.2
      }
    }

    {"deliverable_id": "DECK-PROJECT-LANTERN", "last_task_id": "dealbench-001", "project_code": "PROJECT-LANTERN", "revision": "DECK-01-SOUR-R2", "status": "draft_for_review", "unrelated_slides_sha256": "de291458f9d4e4aaefbe809ae643a19d9a694b94e5de0585d83105607aacff08", "values": {"equity_value_usd_m": -21.41, "headline_value_usd_m": 86.09, "per_share_value_usd": -0.51, "primary_metric": 19.66, "secondary_metric": 7.2}}

  20. 19
    Decidedeals.commit_planok
    {
      "decision": "lock_current_case",
      "deliverable_id": "DECK-PROJECT-LANTERN",
      "model_id": "MODEL-PROJECT-LANTERN",
      "project_code": "PROJECT-LANTERN",
      "rationale": "Current evidence supports lock_current_case; operative revision REV-01-CURRENT.",
      "source_refs": [
        "REV-01-CURRENT",
        "PROJECT-LANTERN-QOE-CURRENT",
        "MSG-dealbench-001-REQUEST",
        "THREAD-dealbench-001"
      ],
      "status": "ready_for_review",
      "task_id": "dealbench-001"
    }

    {"decision": "lock_current_case", "deliverable_id": "DECK-PROJECT-LANTERN", "model_id": "MODEL-PROJECT-LANTERN", "plan_id": "PLAN-dealbench-001", "project_code": "PROJECT-LANTERN", "rationale": "Current evidence supports lock_current_case; operative revision REV-01-CURRENT.", "source_refs": ["REV-01-CURRENT", "PROJECT-LANTERN-QOE-CURRENT", "MSG-dealbench-001-REQUEST", "THREAD-dealbench-001"], "status": "ready_for_review", "task_id": "dealbench-001"}

  21. 20
    Delivermail.send_messageok
    {
      "body": "PROJECT-LANTERN: recommend lock_current_case. Headline value $86.09m; equity value $-21.41m; model MODEL-01-SOUR-R3; deck DECK-01-SOUR-R2; source REV-01-CURRENT.",
      "project_code": "PROJECT-LANTERN",
      "review_status": "draft_for_review",
      "subject": "dealbench-001 — ready_for_review — review draft",
      "task_id": "dealbench-001",
      "to": "deal-lead-review@atlas-sim.example"
    }

    {"review_status": "draft_for_review", "saved": true, "sent_id": "SENT-dealbench-001"}

  22. 21
    Deliverchat.post_messageok
    {
      "channel": "deal-project-lantern",
      "review_status": "draft_for_review",
      "task_id": "dealbench-001",
      "text": "dealbench-001: ready_for_review / lock_current_case / review draft ready."
    }

    {"post_id": "POST-dealbench-001", "review_status": "draft_for_review", "saved": true}

  23. 22
    Investigatedeals.get_modelok
    {
      "model_id": "MODEL-PROJECT-LANTERN"
    }

    {"last_task_id": "dealbench-001", "model_id": "MODEL-PROJECT-LANTERN", "outputs": {"equity_value_usd_m": -21.41, "headline_value_usd_m": 86.09, "per_share_value_usd": -0.51, "primary_metric": 19.66, "secondary_metric": 7.2}, "project_code": "PROJECT-LANTERN", "revision": "MODEL-01-SOUR-R3", "source_refs": ["REV-01-CURRENT", "PROJECT-LANTERN-QOE-CURRENT", "MSG-dealbench-001-REQUEST", "THREAD-dealbench-001"], "status": "updated_for_review"}

  24. 23
    Verifysheets.get_change_logok
    {
      "workbook_id": "WB-PROJECT-LANTERN"
    }

    {"changes": [{"cell_range": "Outputs!B2:F3", "change_id": "CHANGE-dealbench-001", "task_id": "dealbench-001", "values": [["MODEL-01-SOUR-R3", 86.09, -21.41, -0.51, 19.66, 7.2]], "workbook_id": "WB-PROJECT-LANTERN"}]}

Methodology

How the numbers are made

01

Ten deep project worlds

Each synthetic company has a current and superseded forecast, QoE bridge, comps, precedents, debt schedule, bids, diligence, approvals, committee deck, mailbox, chat and audit trail. Ten workflows reuse each frozen project as a real team would.

02

High-level analyst outcomes

Prompts ask for a business outcome and review-ready handoff. They do not prescribe tool order. The agent must distinguish operative from stale authority, calculate from current evidence, compare options, and decide what may be committed.

03

Stateful cross-application execution

Seven logical MCP servers expose data-room, mail, chat, spreadsheet, market-data, deal-management and benchmark controls over isolated SQLite state. Model, workbook, deck, bid, finding, approval, plan and communication writes are durable.

04

One deterministic DealScore

DealScore allocates 100 executable points to discovery, model accuracy, decision quality, committed state, deliverable consistency, readback and containment. No LLM judge or exact reference call sequence is used.

05

Two-sided qualification

The release passed 100/100 oracle runs and 100/100 exact deterministic replays. Five negative-control families executed 500 attacks with zero false accepts.

06

Clean-room APEX boundary

Public APEX, APEX-Accounting and Archipelago pages informed the product contract. Mercor's gated dataset was not downloaded or scraped. Every company, prompt, file, value, tool, answer and trajectory in DealBench is newly synthetic.

07

Leaderboard honesty

Qualification controls prove solvability and discrimination but are never ranked as models. A model row appears only after a complete version-pinned 100-task run has an inspectable receipt.

Qualification controls — excluded from the leaderboard
  • Reference oracle: 100.0% across 100 tasks — Solvability ceiling; not a model submission.
  • No-op control: 4.0% across 100 tasks — Executed adversarial control; never ranked with models.
  • Answer-only shortcut: 46.0% across 100 tasks — Executed adversarial control; never ranked with models.
  • State-only shortcut: 85.0% across 100 tasks — Executed adversarial control; never ranked with models.
  • Wrong-source control: 83.5% across 100 tasks — Executed adversarial control; never ranked with models.
  • Wrong-target control: 85.5% across 100 tasks — Executed adversarial control; never ranked with models.

Clean-room benchmark

APEX-shaped release discipline. Independently authored deal work.

Public APEX, APEX-Accounting, the APEX paper, and Archipelago informed the inspection contract. The gated APEX-Agents corpus was not downloaded or scraped; every DealBench company, prompt, asset, value, tool, answer, and trajectory is synthetic and new.

Run DealBench on Harbor ↗Inspect the architecture anchor ↗