SQL, analytics and data engineering / Benchmark leaderboard
BIRD
Text-to-SQL over realistic databases and domain knowledge.
No published company result yet.
This benchmark is in the research directory. Its task package, adapter and grading protocol need qualification before a hosted run can be offered.
Arrange an evaluation for your company →What this benchmark measures
Metrics
Execution accuracy; efficiency only under the declared official metric.
Execution requirements
SQL generation, database access and query runner.
Scope and limitations
SQL dialect and semantic-layer adaptations must be documented; report full split versus Mini-Dev.
Evaluation availability
Catalog entry. Request a managed evaluation to qualify your agent interface and the benchmark’s native grading requirements.
Sources and company fit
- Original benchmark ↗Established
Companies whose products may fit
Research recommendations based on product capabilities. These companies have not necessarily run this benchmark or integrated with Blobfish.
Compatibility notes for each company
Databricks: Capability-aligned; adapter/access to qualify
Hex: Capability-aligned; adapter/access to qualify
Julius AI: Capability-aligned; adapter/access to qualify
MotherDuck: Capability-aligned; adapter/access to qualify
Sigma: Capability-aligned; adapter/access to qualify
Snowflake: Capability-aligned; adapter/access to qualify
ThoughtSpot: Capability-aligned; adapter/access to qualify