Harbor catalog / Benchmark leaderboard

openthoughts/tasktrove-exp-rle-github-issue-v3

Training/development/subset candidate

No published company result yet.

This benchmark is in the research directory. Its task package, adapter and grading protocol need qualification before a hosted run can be offered.

Arrange an evaluation for your company →

What this benchmark measures

Metrics

Consult the upstream task verifier.

Execution requirements

Dataset-specific harness qualification required.

Scope and limitations

Name/card indicates a training, generated, development or subset collection; inspect held-out protocol before using it as an independent evaluation.

Evaluation availability

Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.

Sources and company fit