Benchmark Radar
AI BENCHMARK PROFILE

Spider 2.0-AIFunc

General AIKnowledge & ReasoningLeolty

Extends text-to-SQL to AI-native SQL workflows with 465 instances across 125 real-world databases on Snowflake. Tasks require using AI functions like classification and sentiment analysis. Evaluates execution accuracy.

Released
2026-07-07
Readiness
Runnable
Primary field
General AI

Why it matters

Addresses the gap of benchmarks not covering AI-native SQL capabilities that are increasingly available in cloud platforms. Provides a reusable dataset and evaluation harness for a new task type.

Motivation

Major cloud data platforms now expose large language model capabilities as native SQL functions, enabling analysts to perform classification, filtering, sentiment analysis, extraction, similarity search, and aggregation within ordinary SQL queries.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.