Skip to content
Studio

Factory for agent training

Turn real workflows into agent training gyms.

Describe one hard enterprise workflow. Studio builds the tool environment, realistic data, up to 20 graded tasks, and deterministic graders. Difficulty variants stay near the factory-configured model's edge, while the brief and release proof stay in one workspace.

Drafts save automatically · no setup needed to start

Open finished demoOriginal brief, generated tasks and execution receipts
Create a gym from your workflowup to 20 graded tasks · demo profile

Structured creation brief

Make the task buildable before spending a factory run.Each required field maps directly to the canonical creation contract.
0 / 6

Restoring your saved draft… Editing unlocks when Studio is ready.

Source anchors Optional URLs, policies, or document references
Real source roomOptional · highest realism for document workflows
DOCX · XLSX · PPTX · PDF · EML · text · 512 files / 24 MiB
Original bytes are sealed by SHA-256 and copied unchanged into task source rooms. Only upload material you are authorized to include: anyone you share the released environment with can read the selected originals. The build is explicitly labeled source-backed, not synthetic-only.
Try an example
⌘/Ctrl + Enter to generate
Complete the six-field creation briefGeneration stays locked until the company, operator, workflow, systems, complication, and verifier-visible outcome are explicit.

Studio researches the domain, simulates the systems, mines up to 20 materially distinct graded tasks (only oracle-proven tasks are admitted), and lets you steer the build before release.

Fictional model-lab workflow

One shared room, from expert brief to cluster proof.

See how domain, evaluation, environment, and infrastructure stakeholders hand one environment forward without losing the decisions that make it useful for RL.

4 / 4 reviewers

Capability hypothesisExample targets · remeasure after every train round

Planning baseline
6 / 30 manual probes
Studio admission
≥12 / 20 candidates
Useful boundary
20–80% pass-rate band
WLWorkflow leadDefines business truth
EREval researcherSets reward evidence
EEEnv engineerChecks tool feasibility
PIPlatform engineerSets runtime constraints

Frame the capability boundary

WLWorkflow leadEnterprise sales operations

A useful task cannot be isolated Salesforce CRUD. The agent has to recover conflicting deal context from calls, notes, email, documents, and the forecast sheet—then leave the CRM and approval packet consistent.

Shared creation briefEnterprise deal handoff with contradictory context

The whole team agrees on the systems, the hard exception, the observable state change, and what the grader must prove before generation starts.

Approval recorddemo
CompanyExample model labPlanning baseline6 / 30 manual probesAdmission gate≥12 / 20 candidatesUseful boundary20–80% pass rate
  • 6 named systems stay in scope
  • Cross-system reconciliation is required
  • Success is a state change, not a prose answer

Illustrative plan. Targets are acceptance criteria, not measured customer results. Commands become runnable after a team downloads its world and supplies its own image digests.