Blind Resampling Outperforms Self-Repair in Small Code Models
A recent study published on arXiv (2607.26117) questions the conventional approach of self-repair in code generation models. Researchers utilized a placebo-controlled framework on MBPP+ across three model sizes (1.5B, 3B, 7B) to evaluate four retry scenarios: blind resampling, a content-free failure notice, authentic execution feedback, and feedback that included verbal self-reflection. The blind resampling method proved to be the most effective for models under 7B and was statistically equivalent at 7B, requiring 2.5-5.5 times fewer tokens. Conditioning on a model's previous failures resulted in a 6.1-point decrease at 1.5B (p=0.006), while execution feedback showed no significant advantage over the placebo. These results indicate that additional attempts, rather than feedback, enhance performance in smaller models.
Key facts
- Study compares self-repair to blind resampling in code models
- Placebo-controlled design on MBPP+ dataset
- Models tested: 1.5B, 3B, 7B parameters
- Four retry conditions: blind resampling, content-free failure notice, execution feedback, feedback with self-reflection
- Blind resampling strongest below 7B, tied at 7B
- Blind resampling uses 2.5-5.5x fewer tokens
- Conditioning on failed attempt costs 6.1 points at 1.5B (p=0.006)
- Execution feedback adds no measurable value over placebo
Entities
Institutions
- arXiv