DecoEvo: Co-Evolving LLM Solver and Rubric Generator in Text Space
A new method called DecoEvo (Decoupled Co-Evolution) has been developed by researchers to enhance large language models (LLMs) in text space without altering the model weights. Unlike traditional text-space methods that maintain a static evaluation, DecoEvo simultaneously evolves a solver skill and a rubric-generator skill with separate objectives. The solver benefits from feedback at the criterion level, while the rubric-generator undergoes revisions through complementary processes. This innovation tackles the limitations of fixed rubrics that fail to capture unmeasured aspects and prevents the solver from exploiting unreliable rubric changes. Notably, gold rubrics are not utilized during the optimization process. The research paper can be found on arXiv.
Key facts
- DecoEvo co-evolves solver and rubric-generator skills in text space
- Text-space optimization edits external natural-language artifacts
- Model weights remain unchanged; model treated as black box
- Existing text-space methods keep evaluation fixed
- Fixed rubrics become a bottleneck on open-ended tasks
- Rubric evolution alone is unreliable when selected by solver's score
- DecoEvo uses decoupled objectives without gold rubrics
- Solver updated via criterion-level feedback
Entities
Institutions
- arXiv