← Blobfish AI / Agent reportsMethodology

Your evaluation workspace

From agent to evidence.

Run a frozen task set across agent scaffolds. Open the report when the job finishes.

Explore a sample
01 Choose a registered benchmark

Register another HUD taskset or Harbor benchmark →

02 Configure your agents
03 Set access and execution limits

Runs continue if you close the tab. Keys are held in a temporary job secret and removed with the job. Provider usage is billed by your model provider.

0 task attemptsUp to 2 environments at once · 30 minute job limit
Recover a run or open history