STAIF: Stage-wise Optimization for Complex Instruction Following in LLMs
A new framework called STAIF has been introduced by researchers to enhance the capacity of large language models in adhering to intricate instructions that involve multiple constraints. This approach separates the alignment of subjective soft constraints from the optimization of hard constraints that can be objectively verified. In Stage 1, preference optimization is utilized with various negative samples to boost sensitivity to soft constraints. Stage 2 implements Reinforcement Learning with Verifiable Rewards (RLVR) to ensure strict adherence to hard constraints. To facilitate this research, the authors developed STAINSTRUCT, a bilingual dataset comprising around 31,000 complex multi-constraint instructions in English and Chinese. The paper critiques existing alignment techniques like DPO, which tend to prioritize overall reward signals and often neglect individual constraint fulfillment, particularly in multi-constraint or out-of-distribution scenarios.
Key facts
- STAIF is a stage-wise optimization framework for instruction following
- Stage 1: preference optimization with multiple negative samples for soft constraints
- Stage 2: RLVR for hard constraints
- STAINSTRUCT dataset: ~31,000 bilingual (English, Chinese) complex instructions
- Addresses limitations of DPO in multi-constraint settings
- Published on arXiv with ID 2607.22649
Entities
Institutions
- arXiv