AI BENCHMARK PROFILE
ESF-Bench
ESF-Bench evaluates slot filling in enterprise contexts, covering 810 multi-turn dialogues and 6,530 slots across 8 domains, with a taxonomy of 57 challenging scenarios.
- Released
- 2026-07-25
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Addresses the gap in evaluating LLMs for slot filling under real-world enterprise constraints and unexpected user behaviors, providing a standard for model comparison in this practical task.
Motivation
The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.