Factory for agent training
Turn real workflows into agent training gyms.
Describe one hard enterprise workflow. Studio builds the tool environment, realistic data, up to 20 graded tasks, and deterministic graders. Difficulty variants stay near the factory-configured model's edge, while the brief and release proof stay in one workspace.
Drafts save automatically · no setup needed to start
Open finished demoOriginal brief, generated tasks and execution receiptsCreate a gym from your workflowup to 20 graded tasks · demo profile
Structured creation brief
Make the task buildable before spending a factory run.Each required field maps directly to the canonical creation contract.Studio researches the domain, simulates the systems, mines up to 20 materially distinct graded tasks (only oracle-proven tasks are admitted), and lets you steer the build before release.
Fictional model-lab workflow
One shared room, from expert brief to cluster proof.
See how domain, evaluation, environment, and infrastructure stakeholders hand one environment forward without losing the decisions that make it useful for RL.
Capability hypothesisExample targets · remeasure after every train round
- Planning baseline
- 6 / 30 manual probes
- Studio admission
- ≥12 / 20 candidates
- Useful boundary
- 20–80% pass-rate band
Frame the capability boundary
The whole team agrees on the systems, the hard exception, the observable state change, and what the grader must prove before generation starts.
- 6 named systems stay in scope
- Cross-system reconciliation is required
- Success is a state change, not a prose answer