SCHEDBench: New Benchmark Tests LLM Constraint Faithfulness in Scheduling
A new benchmark called SCHEDBench has been developed by researchers to assess the constraint faithfulness of large language models (LLMs) in combinatorial scheduling tasks. This benchmark, which is outlined in a paper available on arXiv (2608.00991), evaluates whether LLMs produce schedules that consistently follow constraint-feasible behavior across different natural-language expressions. SCHEDBench includes 1,132 instances derived from canonical scheduling scenarios and solver-generated feasibility and optimality, covering job-shop scheduling problems (JSP), resource-constrained project scheduling (RCPSP), nurse scheduling, and curriculum timetabling of various complexities. The instances are transformed into natural language problems using specialized templates and variations, with reference solutions confirmed for feasibility and optimality. The benchmark tests thirteen leading language models to see how well they maintain constraint faithfulness amidst linguistic differences, addressing a significant gap in evaluating LLMs' reasoning skills in scheduling contexts, relevant for applications like workforce management and project planning.
Key facts
- SCHEDBench is a new benchmark for evaluating LLM constraint faithfulness in combinatorial scheduling.
- It includes 1,132 instances across JSP, RCPSP, nurse rostering, and curriculum timetabling.
- Instances are templated into natural language with surface-form variations.
- Reference solutions are verified for feasibility and optimality.
- Thirteen frontier LLMs are evaluated.
- The paper is available on arXiv under ID 2608.00991.
- The benchmark assesses whether LLMs generate constraint-feasible schedules across varied NL surface forms.
- The work is grounded in canonical scheduling instances and solver-derived feasibility.
Entities
Institutions
- arXiv