Multi-Agent Protocol Distillation Bridges Distribution Gap in Agentic Search
A new paper on arXiv (2607.24280) introduces Multi-Agent Protocol Distillation (MAPD), a framework combining knowledge distillation and reinforcement learning to improve agentic search in large language models. Agentic search interleaves multi-step reasoning with retrieval for knowledge-intensive tasks, but outcome-based RL provides only sparse supervision. MAPD addresses the heterogeneous distillation problem by using a structured, style-normalized protocol as an intermediate representation, enabling distillation from proprietary models despite hidden logits and mismatched tokenizers. The framework employs an offline multi-agent setup to bridge the distribution gap between proprietary and open-source models, avoiding superficial stylistic artifacts from raw trajectory imitation. The paper proposes MAPD as a joint distillation and RL framework that densifies supervisory signals without requiring logit access.
Key facts
- arXiv paper 2607.24280 introduces Multi-Agent Protocol Distillation (MAPD)
- MAPD combines knowledge distillation and reinforcement learning for agentic search
- Agentic search interleaves multi-step reasoning with retrieval
- Outcome-based RL provides only sparse supervision
- MAPD uses a structured, style-normalized protocol as intermediate representation
- It addresses the heterogeneous distillation problem from proprietary models
- Hidden logits and mismatched tokenizers prevent conventional logit-matching
- The framework employs an offline multi-agent setup
Entities
Institutions
- arXiv