AI BENCHMARK PROFILE
SCHEDBench
SCHEDBench evaluates LLM constraint faithfulness in natural-language combinatorial scheduling. It includes 1,132 instances across JSP, RCPSP, nurse rostering, and timetabling, with templated NL variations and solver-verified references.
- Released
- 2026-08-02
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
LLMs must generate constraint-feasible schedules under surface-form variations. SCHEDBench provides a benchmark to test invariance to paraphrasing, revealing faithfulness issues and informing robust model development.
Motivation
This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form variation.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.