Hugging Face catalog / Benchmark leaderboard

mercor/APEX-v1-extended

Related professional-output benchmark; not the same task format/split as APEX-Agents 1.1.

No published company result yet.

This benchmark is in the research directory. Its task package, adapter and grading protocol need qualification before a hosted run can be offered.

Arrange an evaluation for your company →

What this benchmark measures

Metrics

Consult the upstream evaluation protocol.

Execution requirements

Dataset-specific harness qualification required.

Scope and limitations

Badge is catalog metadata, not a maturity, licensing or runnable-harness certification.

Evaluation availability

Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.

Sources and company fit