Benchmark Radar
AI BENCHMARK PROFILE

ESF-Bench

General AIKnowledge & Reasoning

ESF-Bench evaluates slot filling in enterprise contexts, covering 810 multi-turn dialogues and 6,530 slots across 8 domains, with a taxonomy of 57 challenging scenarios.

Released
2026-07-25
Readiness
Paper only
Primary field
General AI

Why it matters

Addresses the gap in evaluating LLMs for slot filling under real-world enterprise constraints and unexpected user behaviors, providing a standard for model comparison in this practical task.

Motivation

The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.