Benchmark Radar
AI BENCHMARK PROFILE

SCHEDBench

General AIKnowledge & Reasoning

SCHEDBench evaluates LLM constraint faithfulness in natural-language combinatorial scheduling. It includes 1,132 instances across JSP, RCPSP, nurse rostering, and timetabling, with templated NL variations and solver-verified references.

Released
2026-08-02
Readiness
Paper only
Primary field
General AI

Why it matters

LLMs must generate constraint-feasible schedules under surface-form variations. SCHEDBench provides a benchmark to test invariance to paraphrasing, revealing faithfulness issues and informing robust model development.

Motivation

This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form variation.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.