Benchmark Radar
AI BENCHMARK PROFILE

HellaSwag

General AIKnowledge & Reasoning

A challenging commonsense natural language inference dataset that uses Adversarial Filtering to create questions trivial for humans (>95% accuracy) but difficult for state-of-the-art models, requiring completion of sentence endings based on physical situations and everyday activities

Released
Unknown
Readiness
Paper only
Primary field
General AI

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.