Hugging Face catalog / Benchmark leaderboard
mercor/APEX-v1-extended
Related professional-output benchmark; not the same task format/split as APEX-Agents 1.1.
No published company result yet.
This benchmark is in the research directory. Its task package, adapter and grading protocol need qualification before a hosted run can be offered.
Arrange an evaluation for your company →What this benchmark measures
Metrics
Consult the upstream evaluation protocol.
Execution requirements
Dataset-specific harness qualification required.
Scope and limitations
Badge is catalog metadata, not a maturity, licensing or runnable-harness certification.
Evaluation availability
Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.
Sources and company fit
- Original benchmark ↗Official-filter dataset card
- Hugging Face · mercor/APEX-v1-extended ↗Public · Related professional-output benchmark; not the same task format/split as APEX-Agents 1.1.