SLMs as Multi-Agent Routers: Progressive SFT and RL Approach
A new research paper on arXiv (2608.00030) proposes training a small language model (SLM) to act as a multi-agent router, selecting specialized retrieval agents and generating structured parameters for downstream tool calls. The approach uses supervised fine-tuning followed by reinforcement learning, with a hierarchical reward function that balances retrieval relevance and query-agent topic alignment. This method addresses limitations of intent-based routing by incorporating feedback from retrieved content, enabling the model to learn task-dependent agent suitability. The paper is authored by researchers (names not provided) and was announced as a cross-type submission. The work is relevant to AI-technology and digital domains, offering a potential improvement in retrieval systems.
Key facts
- Paper on arXiv: 2608.00030
- Trains small language model (SLM) as multi-agent router
- Uses supervised fine-tuning and reinforcement learning
- Jointly performs agent selection and structured parameter generation
- Hierarchical reward function based on retrieval relevance and topic alignment
- Addresses limitations of intent-based routing
- Incorporates signal from retrieved content
- Announced as cross-type submission
Entities
Institutions
- arXiv