MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
A new technique called MINT (MIN-selection preference disTillation) has been developed by researchers to align language agents with multiple objectives at once. This method tackles a frequent issue in preference-based training, where combining objectives additively can lead to optimization collapse, prioritizing the simplest goal over others. MINT suggests a straightforward modification to preference distillation: rather than evaluating candidates based on a weighted rewards sum, it assesses them according to their weakest objective, thus identifying the most balanced option instead of the most skewed one. This aligns with the p → -∞ limit of a generalized-mean family, ranging from additive to worst-case selection. Tests on cooperative emotional support and adversarial negotiation tasks showed that min-selection enhances both objectives while significantly minimizing their imbalance. The study can be found on arXiv with the identifier 2608.14828.
Key facts
- MINT is a method for balanced multi-objective alignment in language agents.
- It addresses optimization collapse in additive reward combination.
- The approach ranks candidates by their weakest objective.
- It is the p → -∞ limit of a generalized-mean family.
- Tested on cooperative emotional support and adversarial negotiation.
- Results show improvement in both objectives and reduced imbalance.
- Paper available on arXiv:2608.14828.
- The method uses an unchanged DPO objective.
Entities
Institutions
- arXiv