General model capability / Benchmark leaderboard
ARC-AGI-2
Novel abstract visual-grid transformations.
No published company result yet.
This benchmark is in the research directory. Its task package, adapter and grading protocol need qualification before a hosted run can be offered.
Arrange an evaluation for your company →What this benchmark measures
Metrics
Exact task accuracy at a reported inference-cost budget.
Execution requirements
Reasoning system using the official task protocol.
Scope and limitations
Public, semi-private and private evaluation sets differ. The Harbor 167-task package is not the full official private evaluation.
Evaluation availability
Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.
Sources and company fit
- Original benchmark ↗Established frontier diagnostic
- Harbor · arcprize/arc-agi-2 ↗167 source tasks · Public · Public detail metadata + family research
Companies whose products may fit
Research recommendations based on product capabilities. These companies have not necessarily run this benchmark or integrated with Blobfish.
Compatibility notes for each company
Alibaba Qwen: Capability-aligned; adapter/access to qualify
Anthropic: Capability-aligned; adapter/access to qualify
Cohere: Direct component API + selected system adapters
DeepSeek: Capability-aligned; adapter/access to qualify
Google DeepMind: Capability-aligned; adapter/access to qualify
Meta: Capability-aligned; adapter/access to qualify
MiniMax: Capability-aligned; adapter/access to qualify
Mistral: Capability-aligned; adapter/access to qualify
OpenAI: Capability-aligned; adapter/access to qualify
Z.ai: Capability-aligned; adapter/access to qualify
xAI: Capability-aligned; adapter/access to qualify