LAMaS: A Latency-Aware Orchestration Framework for Multi-Agent Systems
A new framework, Latency-Aware Multi-agent System (LAMaS), has been proposed to address latency in multi-agent systems (MAS) that coordinate multiple LLM-powered agents. The framework, detailed in arXiv paper 2601.10560, tackles the challenge of reducing end-to-end latency without sacrificing task accuracy. LAMaS operates at two levels: at training time, it learns latency-aware orchestration, and at inference time, it optimizes the critical execution path. The paper highlights that existing orchestration methods focus on task performance and inference cost, leaving latency largely unaddressed. LAMaS aims to fill this gap by ensuring that latency optimization does not misassign operator-level credit, which can degrade task accuracy. The framework is designed for learning-based multi-agent systems and is presented as a solution to the high inference latency caused by multi-step execution and repeated model invocations. The research is relevant to the growing field of AI orchestration, where efficiency and accuracy are critical.
Key facts
- LAMaS is a latency-aware orchestration framework for multi-agent systems.
- It addresses high inference latency in MAS caused by multi-step execution and repeated model invocations.
- Existing orchestration methods primarily optimize task performance and inference cost, not latency.
- End-to-end latency in MAS is governed by the critical execution path.
- Naive latency optimization can misassign operator-level credit and degrade task accuracy.
- LAMaS operates at training time and inference time.
- The paper is available on arXiv with identifier 2601.10560.
- The framework is designed for learning-based multi-agent systems.
Entities
Institutions
- arXiv