Benchmark operations / coverage as of 2026-09-16

Every benchmark. An execution plan.

Track the 17 portfolio entries and the benchmarks across 20 product categories. Inspect what is prepared, what has been measured, and what each complete evaluation still requires.

17 portfolio entries20 product categories3 registered full native packages1 dedicated GKE cluster

The complete public portfolio

17 entries, with their real scope.

Nine original benchmarks, six companion worlds, Northstar and HubBench. A companion’s four tasks are distinct from the upstream benchmark’s full task set.

Explore the portfolio ↗
Preparation snapshot · 2026-09-16 · new full-model matrix pending
Benchmark / worldEvaluation scopeTasksExecutionNext action
CounselBench-100Original benchmarkOriginal release identity required100Qualification requiredbf-benchmarks · dedicated run JobPrepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts.
SalesBench-100Original benchmarkOriginal release identity required100Qualification requiredbf-benchmarks · dedicated run JobPrepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts.
DevOpsBench-100Original benchmarkOriginal release identity required100Full package registeredbf-benchmarks · dedicated run JobAll 100 original tasks and immutable runtime images are registered. Complete model attempts and grading remain required; earlier ten-task controls and pilots retain their original scope.
LedgerBench-100Original benchmarkOriginal release identity required100Qualification requiredbf-benchmarks · dedicated run JobPrepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts.
FactoryBench-100Original benchmarkOriginal release identity required100Qualification requiredbf-benchmarks · dedicated run JobPrepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts.
SemiKongBench-100Original benchmarkOriginal release identity required100Qualification requiredbf-benchmarks · dedicated run JobPrepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts.
DealBench-100Original benchmarkOriginal release identity required100Qualification requiredbf-benchmarks · dedicated run JobPrepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts.
ERPBench-100Original benchmarkOriginal release identity required100Qualification requiredbf-benchmarks · dedicated run JobPrepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts.
Arc CRM 6Original benchmarkOriginal release identity required6Qualification requiredbf-benchmarks · dedicated run JobPrepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts.
FilingBench-50Companion world50 upstream source tasks are separate4Qualification requiredbf-benchmarks · dedicated run JobThe four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run.
ServiceBench-375Companion world375 upstream source tasks are separate4Qualification requiredbf-benchmarks · dedicated run JobThe four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run.
DiscoveryBench-102Companion world102 upstream source tasks are separate4Qualification requiredbf-benchmarks · dedicated run JobThe four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run.
DefenseBench-100Companion world100 upstream source tasks are separate4Qualification requiredbf-benchmarks · dedicated run JobThe four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run.
OpsBench-14Companion world14 upstream source tasks are separate4Qualification requiredbf-benchmarks · dedicated run JobThe four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run.
RepoBench-500Companion world500 upstream source tasks are separate4Qualification requiredbf-benchmarks · dedicated run JobThe four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run.
Northstar Composite OpsWorld proofOriginal release identity required100Qualification requiredbf-benchmarks · dedicated run JobFull 100-task native Harbor adapter merged in #4592. All original local controls pass; the prepared image still requires live GKE qualification and a complete measured model run. The separate customer release keeps its own identity.
HubBench13-family suiteOriginal release identity required104Qualification requiredbf-benchmarks · dedicated run JobQualify the complete 104-task release, multiple MCP endpoints, named internal networks and original 1,200-second task budgets.

Existing model results and reference controls keep their original release, task set and provenance. This matrix does not replace those historical results with a new score.

20 product categories

Match the benchmark to the work.

Each benchmark has an execution target and a concrete preparation requirement. Company fit describes candidates for evaluation.

20 / 20 categories

Browser agents · 3 benchmark entries

Successful completion of natural-language browser tasks.

Execution profile: Browser / Harbor. Prepare the complete task roster and benchmark-specific grading; live sites and self-hosted resettable sites require different environments.

Desktop and mobile agents · 4 benchmark entries

Application control, visual grounding and verified changes to application state.

Execution profile: Desktop or emulator. A qualified VM/emulator or grounding adapter, benchmark assets and native evaluator are required; the CPU container lane is insufficient.

Coding agents · 4 benchmark entries

Issue resolution, executable terminal work, regressions, time and cost.

Execution profile: Harbor / CPU Linux. Freeze the complete release, build task and verifier images, and qualify original positive and negative controls.

Customer-service agents · 3 benchmark entries

Correct actions, policy adherence, completion and repeated-conversation consistency.

Execution profile: Tool and conversation simulation. Pin the simulator, policies, domain state and native evaluator; supply required service access and reset behavior.

CRM agents · 2 benchmark entries

Correct record queries and mutations with an explicit schema and tool mapping.

Execution profile: Tool and conversation simulation. Pin the simulator, policies, domain state and native evaluator; supply required service access and reset behavior.

Enterprise workflow agents · 3 benchmark entries

Multi-step completion across applications, documents and business records.

Execution profile: Multi-service world / Harbor. Qualify each world service, MCP endpoint, evidence mount and deterministic reset with its original grader.

Search and research agents · 4 benchmark entries

Difficult retrieval, supported answers, research quality and citation quality.

Execution profile: Browser / Harbor. Prepare the complete task roster and benchmark-specific grading; live sites and self-hosted resettable sites require different environments.

Embeddings and retrieval · 3 benchmark entries

Document relevance and ranking; latency and throughput for complete pipelines.

Execution profile: Retrieval and structured evaluation. Prepare the corpus and relevance labels, embedding/search or answer adapter, and the benchmark-native scoring pipeline.

SQL and business-analysis agents · 3 benchmark entries

SQL execution correctness and multi-step business-analysis accuracy.

Execution profile: Data and artifact evaluation. Prepare databases/workbooks and original grading tools; external warehouses need their benchmark-specific access.

Spreadsheet agents · 2 benchmark entries

Correct workbook artifacts, formulas and graded formatting.

Execution profile: Data and artifact evaluation. Prepare databases/workbooks and original grading tools; external warehouses need their benchmark-specific access.

Document parsing and extraction · 4 benchmark entries

Text, tables, reading order, formulas, fields and supporting evidence.

Execution profile: Document evaluation. Prepare source documents, parser interface and pinned text/table/layout evaluator images.

Legal research and contracts · 4 benchmark entries

Reasoning, evidence retrieval, clause extraction and interpretation.

Execution profile: Retrieval and structured evaluation. Prepare the corpus and relevance labels, embedding/search or answer adapter, and the benchmark-native scoring pipeline.

Financial research · 3 benchmark entries

Evidence-backed answers, numerical reasoning and professional deliverables.

Execution profile: Data and artifact evaluation. Prepare databases/workbooks and original grading tools; external warehouses need their benchmark-specific access.

Speech recognition · 4 benchmark entries

Transcription across languages and domains; diarization scored separately.

Execution profile: Speech input/output. Qualify an ASR or TTS endpoint, source audio and native scoring; reference-speaker and language conditions must be preserved.

Speech synthesis and cloning · 1 benchmark entries

Intelligibility and reference-speaker similarity; separate listening and latency tests.

Execution profile: Speech input/output. Qualify an ASR or TTS endpoint, source audio and native scoring; reference-speaker and language conditions must be preserved.

Conversational voice agents · 2 benchmark entries

Spoken task completion, tool use, interruptions and response latency.

Execution profile: Streaming voice simulation. Qualify streaming audio, interruption timing, tool/user simulation and separate task, ASR and TTS scores.

Image generation · 3 benchmark entries

Prompt adherence, counts, attributes and spatial relationships.

Execution profile: Media generation and scoring. Qualify image/video generation and benchmark-native evaluator images; GPU requirements need a separately bounded execution profile.

Video generation · 1 benchmark entries

Temporal consistency, quality, motion and prompt alignment.

Execution profile: Media generation and scoring. Qualify image/video generation and benchmark-native evaluator images; GPU requirements need a separately bounded execution profile.

Multimodal understanding · 3 benchmark entries

Reasoning and understanding over images and video.

Execution profile: Image and video understanding. Prepare licensed media, modality-capable model input and benchmark-native scoring with the original sampling protocol.

Agent security · 2 benchmark entries

Prompt-injection resistance while preserving legitimate task success.

Execution profile: Tool and conversation simulation. Pin the simulator, policies, domain state and native evaluator; supply required service access and reset behavior.

Dedicated benchmark compute

Isolated tasks. Bounded runs.

Each admitted evaluation uses its own GKE Job, frozen manifest and retained artifacts. Benchmark compute is separate from production world storage.

Inspect harness records ↗

Original harness and graders

Harbor 0.21.0 and HUD 0.6.17 are installed. Benchmark-specific adapters and original positive/negative controls determine readiness.

Dedicated execution

gVisor task pods; a bounded dedicated Job for each admitted evaluation. The task pool scales between 0 and 2 nodes.

Retained outcomes

Task pods are removed after collection; finished run Jobs expire after 600 seconds. Retain benchmark outputs, account caches and all world data.

Runtime limits and remaining adapter work

Target: bf-benchmarks, us-central1-a. System node: e2-standard-2; task nodes: e2-standard-4, maximum 2. Current task limit: 2 CPU, 4096 MiB RAM, 900 seconds. Run Job deadline includes a 180-second finalization window; maximum 10980 seconds.

The current qualified lane does not support arbitrary multi-service networks, evidence volumes, desktop VMs/emulators, GPU evaluators or streaming media protocols. Their assigned profiles are preparation plans, and need implementation and qualification before admission. The published 1,200-second world task budgets must be preserved when extending the current 900-second runner limit.

Registry observation: 2026-09-16T04:23:21.089063+00:00. This page is a dated preparation record, not a live availability monitor.