MemOPD: On-Policy Distillation via Memory State Alignment for Long-Horizon Agents
A recent preprint on arXiv (2608.07068v1) presents MemOPD, a technique designed for on-policy distillation (OPD) in agents with long-horizon capabilities. This research tackles the issue of context accumulation during interactions, which negatively affects both performance and stability. By utilizing compact memory, the method compresses and updates historical data between model calls. Typically, the process of determining what to keep depends on proximal policy optimization (PPO) linked to final task rewards, where sparse rewards offer minimal guidance for memory updates. MemOPD introduces on-policy distillation to provide dense supervision from a teacher during student rollouts. For this supervision to be effective, the teacher must assess each action sampled in the same state it was generated. The paper highlights that memory compression can disrupt this alignment, leading to misalignment in scoring actions. The proposed method seeks to synchronize memory states for valid supervision. This preprint is newly released and accessible on arXiv.
Key facts
- arXiv:2608.07068v1
- Announce Type: new
- Introduces MemOPD
- Addresses long-horizon agents
- Uses on-policy distillation (OPD)
- Provides dense teacher supervision
- Identifies issue with context rewriting
- Aims to align memory states
Entities
Institutions
- arXiv