MetaEvolve: Teaching LLMs Self-Evolution Meta-Skills via Reinforcement Learning
MetaEvolve, a newly introduced framework, seeks to enhance essential meta-skills in large language models (LLMs) to promote effective self-evolution. This initiative builds on the achievements of AlphaEvolve, which showcased improvements through iterative self-evolution informed by environmental feedback. MetaEvolve fills a void in conventional post-training methods that overlook vital meta-skills such as self-reflection. The framework incorporates a data synthesis pipeline, evolution-aware reinforcement learning (RL), and evolutionary search during inference. Its foundation lies in coding, where ongoing program execution yields continuous reward signals beyond mere correctness. The training data comprises evolution trajectories, including the current program, fitness score (which integrates correctness and efficiency), and records of previous attempts. Detailed findings are available in a paper on arXiv (2607.21971).
Key facts
- MetaEvolve is a framework for teaching LLMs meta-skills for self-evolution.
- It builds on AlphaEvolve's iterative self-evolution with environment feedback.
- Meta-skills like self-reflection are largely neglected by traditional post-training.
- The framework uses a data synthesis pipeline, evolution-aware RL, and inference-time evolutionary search.
- MetaEvolve is grounded in coding, where program execution provides continuous reward signals.
- Training data includes evolution trajectories with current program, fitness score, and history.
- The paper is available on arXiv with identifier 2607.21971.
- The approach aims to enable effective multi-round refinement in LLMs.
Entities
Institutions
- arXiv