WebVoyager ↗
Browser agent with live-web access and recorded trajectories.
Assigned compute: bf-benchmarks · Browser / Harbor. This assignment is not an admitted run.
Track the 17 portfolio entries and the benchmarks across 20 product categories. Inspect what is prepared, what has been measured, and what each complete evaluation still requires.
The complete public portfolio
Nine original benchmarks, six companion worlds, Northstar and HubBench. A companion’s four tasks are distinct from the upstream benchmark’s full task set.
| Benchmark / world | Evaluation scope | Tasks | Execution | Next action |
|---|---|---|---|---|
| CounselBench-100 ↗ | Original benchmarkOriginal release identity required | 100 | Qualification requiredbf-benchmarks · dedicated run Job | Prepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts. |
| SalesBench-100 ↗ | Original benchmarkOriginal release identity required | 100 | Qualification requiredbf-benchmarks · dedicated run Job | Prepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts. |
| DevOpsBench-100 ↗ | Original benchmarkOriginal release identity required | 100 | Full package registeredbf-benchmarks · dedicated run Job | All 100 original tasks and immutable runtime images are registered. Complete model attempts and grading remain required; earlier ten-task controls and pilots retain their original scope. |
| LedgerBench-100 ↗ | Original benchmarkOriginal release identity required | 100 | Qualification requiredbf-benchmarks · dedicated run Job | Prepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts. |
| FactoryBench-100 ↗ | Original benchmarkOriginal release identity required | 100 | Qualification requiredbf-benchmarks · dedicated run Job | Prepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts. |
| SemiKongBench-100 ↗ | Original benchmarkOriginal release identity required | 100 | Qualification requiredbf-benchmarks · dedicated run Job | Prepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts. |
| DealBench-100 ↗ | Original benchmarkOriginal release identity required | 100 | Qualification requiredbf-benchmarks · dedicated run Job | Prepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts. |
| ERPBench-100 ↗ | Original benchmarkOriginal release identity required | 100 | Qualification requiredbf-benchmarks · dedicated run Job | Prepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts. |
| Arc CRM 6 ↗ | Original benchmarkOriginal release identity required | 6 | Qualification requiredbf-benchmarks · dedicated run Job | Prepare the exact complete published release, immutable images and original verifier; qualify world services, MCP tools, evidence mounts and native task timeouts. |
| FilingBench-50 ↗ | Companion world50 upstream source tasks are separate | 4 | Qualification requiredbf-benchmarks · dedicated run Job | The four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run. |
| ServiceBench-375 ↗ | Companion world375 upstream source tasks are separate | 4 | Qualification requiredbf-benchmarks · dedicated run Job | The four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run. |
| DiscoveryBench-102 ↗ | Companion world102 upstream source tasks are separate | 4 | Qualification requiredbf-benchmarks · dedicated run Job | The four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run. |
| DefenseBench-100 ↗ | Companion world100 upstream source tasks are separate | 4 | Qualification requiredbf-benchmarks · dedicated run Job | The four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run. |
| OpsBench-14 ↗ | Companion world14 upstream source tasks are separate | 4 | Qualification requiredbf-benchmarks · dedicated run Job | The four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run. |
| RepoBench-500 ↗ | Companion world500 upstream source tasks are separate | 4 | Qualification requiredbf-benchmarks · dedicated run Job | The four-task companion world needs its own prepared GKE package and verifier qualification. Its upstream benchmark remains a separate run. |
| Northstar Composite Ops ↗ | World proofOriginal release identity required | 100 | Qualification requiredbf-benchmarks · dedicated run Job | Full 100-task native Harbor adapter merged in #4592. All original local controls pass; the prepared image still requires live GKE qualification and a complete measured model run. The separate customer release keeps its own identity. |
| HubBench ↗ | 13-family suiteOriginal release identity required | 104 | Qualification requiredbf-benchmarks · dedicated run Job | Qualify the complete 104-task release, multiple MCP endpoints, named internal networks and original 1,200-second task budgets. |
Existing model results and reference controls keep their original release, task set and provenance. This matrix does not replace those historical results with a new score.
20 product categories
Each benchmark has an execution target and a concrete preparation requirement. Company fit describes candidates for evaluation.
20 / 20 categories
Successful completion of natural-language browser tasks.
Execution profile: Browser / Harbor. Prepare the complete task roster and benchmark-specific grading; live sites and self-hosted resettable sites require different environments.
Browser agent with live-web access and recorded trajectories.
Assigned compute: bf-benchmarks · Browser / Harbor. This assignment is not an admitted run.
Live browser, action/screenshot trace and dated task snapshot.
Assigned compute: bf-benchmarks · Browser / Harbor. This assignment is not an admitted run.
Self-hosted sites plus BrowserGym or original harness.
Assigned compute: bf-benchmarks · Browser / Harbor. This assignment is not an admitted run.
Application control, visual grounding and verified changes to application state.
Execution profile: Desktop or emulator. A qualified VM/emulator or grounding adapter, benchmark assets and native evaluator are required; the CPU container lane is insufficient.
Desktop VM, screenshot/input adapter and pinned OS image.
Assigned compute: bf-benchmarks · Desktop or emulator. This assignment is not an admitted run.
Desktop runtime and release-matched gated assets.
Assigned compute: bf-benchmarks · Desktop or emulator. This assignment is not an admitted run.
Model returning screen coordinates for image and instruction.
Assigned compute: bf-benchmarks · Desktop or emulator. This assignment is not an admitted run.
Android emulator, supported apps and action adapter.
Assigned compute: bf-benchmarks · Desktop or emulator. This assignment is not an admitted run.
Issue resolution, executable terminal work, regressions, time and cost.
Execution profile: Harbor / CPU Linux. Freeze the complete release, build task and verifier images, and qualify original positive and negative controls.
Repository-editing agent and isolated executable tests.
Assigned compute: bf-benchmarks · Harbor / CPU Linux. This assignment is not an admitted run.
Repository agent and Pro-specific environments/tests.
Assigned compute: bf-benchmarks · Harbor / CPU Linux. This assignment is not an admitted run.
Containerized command-line agent, Harbor and fixed resource budgets.
Assigned compute: bf-benchmarks · Harbor / CPU Linux. This assignment is not an admitted run.
Harbor terminal agent; some tasks need GPUs or multiple containers.
Assigned compute: bf-benchmarks · Harbor / CPU Linux. This assignment is not an admitted run.
Correct actions, policy adherence, completion and repeated-conversation consistency.
Execution profile: Tool and conversation simulation. Pin the simulator, policies, domain state and native evaluator; supply required service access and reset behavior.
Agent linked to benchmark tools and a fixed user simulator.
Assigned compute: bf-benchmarks · Tool and conversation simulation. This assignment is not an admitted run.
Tool adapter plus user simulation; retrieval setup for knowledge tasks.
Assigned compute: bf-benchmarks · Tool and conversation simulation. This assignment is not an admitted run.
Tool-call model adapter or instrumented agent interface.
Assigned compute: bf-benchmarks · Tool and conversation simulation. This assignment is not an admitted run.
Correct record queries and mutations with an explicit schema and tool mapping.
Execution profile: Tool and conversation simulation. Pin the simulator, policies, domain state and native evaluator; supply required service access and reset behavior.
Salesforce sandbox/API adapter.
Assigned compute: bf-benchmarks · Tool and conversation simulation. This assignment is not an admitted run.
Tool adapter plus user simulation; retrieval setup for knowledge tasks.
Assigned compute: bf-benchmarks · Tool and conversation simulation. This assignment is not an admitted run.
Multi-step completion across applications, documents and business records.
Execution profile: Multi-service world / Harbor. Qualify each world service, MCP endpoint, evidence mount and deterministic reset with its original grader.
ServiceNow instance and BrowserGym integration.
Assigned compute: bf-benchmarks · Multi-service world / Harbor. This assignment is not an admitted run.
AppWorld runtime plus interactive code/tool agent.
Assigned compute: bf-benchmarks · Multi-service world / Harbor. This assignment is not an admitted run.
Multi-service company environment and agent tools.
Assigned compute: bf-benchmarks · Multi-service world / Harbor. This assignment is not an admitted run.
Difficult retrieval, supported answers, research quality and citation quality.
Execution profile: Browser / Harbor. Prepare the complete task roster and benchmark-specific grading; live sites and self-hosted resettable sites require different environments.
Search/browsing agent with fixed budgets.
Assigned compute: bf-benchmarks · Browser / Harbor. This assignment is not an admitted run.
General agent with the required tool and media access.
Assigned compute: bf-benchmarks · Browser / Harbor. This assignment is not an admitted run.
Research system returning reports and citations; fixed judge pipeline.
Assigned compute: bf-benchmarks · Browser / Harbor. This assignment is not an admitted run.
Research agent and the benchmark judge workflow.
Assigned compute: bf-benchmarks · Browser / Harbor. This assignment is not an admitted run.
Document relevance and ranking; latency and throughput for complete pipelines.
Execution profile: Retrieval and structured evaluation. Prepare the corpus and relevance labels, embedding/search or answer adapter, and the benchmark-native scoring pipeline.
Embedding/reranking endpoint through the selected MTEB task collection.
Assigned compute: bf-benchmarks · Retrieval and structured evaluation. This assignment is not an admitted run.
Search/index pipeline with fixed corpus, chunking and embeddings.
Assigned compute: bf-benchmarks · Retrieval and structured evaluation. This assignment is not an admitted run.
Retrieval/reranker or agentic retrieval system.
Assigned compute: bf-benchmarks · Retrieval and structured evaluation. This assignment is not an admitted run.
SQL execution correctness and multi-step business-analysis accuracy.
Execution profile: Data and artifact evaluation. Prepare databases/workbooks and original grading tools; external warehouses need their benchmark-specific access.
SQL generation, database access and query runner.
Assigned compute: bf-benchmarks · Data and artifact evaluation. This assignment is not an admitted run.
Warehouse or local database environment; some tracks require external accounts.
Assigned compute: bf-benchmarks · Data and artifact evaluation. This assignment is not an admitted run.
Data agent with Python/shell and benchmark files.
Assigned compute: bf-benchmarks · Data and artifact evaluation. This assignment is not an admitted run.
Correct workbook artifacts, formulas and graded formatting.
Execution profile: Data and artifact evaluation. Prepare databases/workbooks and original grading tools; external warehouses need their benchmark-specific access.
Agent capable of editing spreadsheet files.
Assigned compute: bf-benchmarks · Data and artifact evaluation. This assignment is not an admitted run.
File and application tools plus the published rubric evaluator.
Assigned compute: bf-benchmarks · Data and artifact evaluation. This assignment is not an admitted run.
Text, tables, reading order, formulas, fields and supporting evidence.
Execution profile: Document evaluation. Prepare source documents, parser interface and pinned text/table/layout evaluator images.
Document-to-structured-text/Markdown adapter.
Assigned compute: bf-benchmarks · Document evaluation. This assignment is not an admitted run.
Document parser producing normalized Markdown.
Assigned compute: bf-benchmarks · Document evaluation. This assignment is not an admitted run.
Document parser with the required output representations.
Assigned compute: bf-benchmarks · Document evaluation. This assignment is not an admitted run.
Structured extraction API returning schema-valid JSON and evidence.
Assigned compute: bf-benchmarks · Document evaluation. This assignment is not an admitted run.
Reasoning, evidence retrieval, clause extraction and interpretation.
Execution profile: Retrieval and structured evaluation. Prepare the corpus and relevance labels, embedding/search or answer adapter, and the benchmark-native scoring pipeline.
Text model or appropriately configured legal assistant.
Assigned compute: bf-benchmarks · Retrieval and structured evaluation. This assignment is not an admitted run.
Retrieval pipeline over the provided legal corpus.
Assigned compute: bf-benchmarks · Retrieval and structured evaluation. This assignment is not an admitted run.
Contract extraction pipeline.
Assigned compute: bf-benchmarks · Retrieval and structured evaluation. This assignment is not an admitted run.
Contract text plus hypothesis; model outputs labels and evidence.
Assigned compute: bf-benchmarks · Retrieval and structured evaluation. This assignment is not an admitted run.
Evidence-backed answers, numerical reasoning and professional deliverables.
Execution profile: Data and artifact evaluation. Prepare databases/workbooks and original grading tools; external warehouses need their benchmark-specific access.
Document retrieval and financial QA pipeline.
Assigned compute: bf-benchmarks · Data and artifact evaluation. This assignment is not an admitted run.
Financial reasoning model with calculator/code tools where allowed.
Assigned compute: bf-benchmarks · Data and artifact evaluation. This assignment is not an admitted run.
File and application tools plus the published rubric evaluator.
Assigned compute: bf-benchmarks · Data and artifact evaluation. This assignment is not an admitted run.
Transcription across languages and domains; diarization scored separately.
Execution profile: Speech input/output. Qualify an ASR or TTS endpoint, source audio and native scoring; reference-speaker and language conditions must be preserved.
Speech-to-text API or local model.
Assigned compute: bf-benchmarks · Speech input/output. This assignment is not an admitted run.
ASR or speech system supporting the evaluated languages.
Assigned compute: bf-benchmarks · Speech input/output. This assignment is not an admitted run.
ASR system handling long audio.
Assigned compute: bf-benchmarks · Speech input/output. This assignment is not an admitted run.
Meeting transcription and speaker-diarization pipeline.
Assigned compute: bf-benchmarks · Speech input/output. This assignment is not an admitted run.
Intelligibility and reference-speaker similarity; separate listening and latency tests.
Execution profile: Speech input/output. Qualify an ASR or TTS endpoint, source audio and native scoring; reference-speaker and language conditions must be preserved.
Reference-conditioned TTS API; generate WAV/audio for the prescribed text.
Assigned compute: bf-benchmarks · Speech input/output. This assignment is not an admitted run.
Spoken task completion, tool use, interruptions and response latency.
Execution profile: Streaming voice simulation. Qualify streaming audio, interruption timing, tool/user simulation and separate task, ASR and TTS scores.
Native real-time audio adapter and benchmark tool/user simulation.
Assigned compute: bf-benchmarks · Streaming voice simulation. This assignment is not an admitted run.
Voice assistant with audio input and benchmark-compatible answers.
Assigned compute: bf-benchmarks · Streaming voice simulation. This assignment is not an admitted run.
Prompt adherence, counts, attributes and spatial relationships.
Execution profile: Media generation and scoring. Qualify image/video generation and benchmark-native evaluator images; GPU requirements need a separately bounded execution profile.
Image generation API with fixed prompts, seeds and sample counts.
Assigned compute: bf-benchmarks · Media generation and scoring. This assignment is not an admitted run.
Image generator and fixed vision-language evaluator.
Assigned compute: bf-benchmarks · Media generation and scoring. This assignment is not an admitted run.
Text-to-image generator and published evaluators.
Assigned compute: bf-benchmarks · Media generation and scoring. This assignment is not an admitted run.
Temporal consistency, quality, motion and prompt alignment.
Execution profile: Media generation and scoring. Qualify image/video generation and benchmark-native evaluator images; GPU requirements need a separately bounded execution profile.
Video model with matched prompts, duration and resolution.
Assigned compute: bf-benchmarks · Media generation and scoring. This assignment is not an admitted run.
Reasoning and understanding over images and video.
Execution profile: Image and video understanding. Prepare licensed media, modality-capable model input and benchmark-native scoring with the original sampling protocol.
Vision-language model.
Assigned compute: bf-benchmarks · Image and video understanding. This assignment is not an admitted run.
Image-and-text reasoning model.
Assigned compute: bf-benchmarks · Image and video understanding. This assignment is not an admitted run.
Video model with declared frame sampling, audio and subtitle access.
Assigned compute: bf-benchmarks · Image and video understanding. This assignment is not an admitted run.
Prompt-injection resistance while preserving legitimate task success.
Execution profile: Tool and conversation simulation. Pin the simulator, policies, domain state and native evaluator; supply required service access and reset behavior.
Agent + defensive product in the same benchmark workflow.
Assigned compute: bf-benchmarks · Tool and conversation simulation. This assignment is not an admitted run.
Tool-integrated model/agent adapter.
Assigned compute: bf-benchmarks · Tool and conversation simulation. This assignment is not an admitted run.
Dedicated benchmark compute
Each admitted evaluation uses its own GKE Job, frozen manifest and retained artifacts. Benchmark compute is separate from production world storage.
Harbor 0.21.0 and HUD 0.6.17 are installed. Benchmark-specific adapters and original positive/negative controls determine readiness.
gVisor task pods; a bounded dedicated Job for each admitted evaluation. The task pool scales between 0 and 2 nodes.
Task pods are removed after collection; finished run Jobs expire after 600 seconds. Retain benchmark outputs, account caches and all world data.
Target: bf-benchmarks, us-central1-a. System node: e2-standard-2; task nodes: e2-standard-4, maximum 2. Current task limit: 2 CPU, 4096 MiB RAM, 900 seconds. Run Job deadline includes a 180-second finalization window; maximum 10980 seconds.
The current qualified lane does not support arbitrary multi-service networks, evidence volumes, desktop VMs/emulators, GPU evaluators or streaming media protocols. Their assigned profiles are preparation plans, and need implementation and qualification before admission. The published 1,200-second world task budgets must be preserved when extending the current 900-second runner limit.
Registry observation: 2026-09-16T04:23:21.089063+00:00. This page is a dated preparation record, not a live availability monitor.