Frozen-Weight AI Agents Learn from Deployment Feedback via External Memory
A new arXiv paper (2607.22157) demonstrates that AI agents with frozen models can achieve continual learning by pairing them with an external memory that distills each episode into retrievable natural-language rules. Using only the one-bit outcome verdict from deployment feedback, single-trial success on the banking domain of τ-bench rose to 1.6× the static-RAG baseline; learning from corrections achieved 2.6×, converting 22 of 84 tasks the baseline never solved. The method was tested on Mistral Large and spans the deployment spectrum.
Key facts
- arXiv paper 2607.22157
- Frozen-weights agents paired with external memory for continual learning
- Feedback from deployment (outcome verdicts and corrections) used as learning signal
- Tested on τ-bench banking domain
- One-bit outcome verdict lifted success to 1.6× baseline
- Learning from corrections achieved 2.6× baseline
- Converted 22 of 84 tasks previously unsolved
- Tested on Mistral Large model
Entities
Institutions
- arXiv
- Mistral Large