BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks
A new framework named BALANCE has been proposed to optimize large language model (LLM) inference in wireless edge networks, addressing the latency-memory tradeoff between autoregressive decoding (AD) and speculative decoding (SD). The framework, detailed in a paper on arXiv (2608.05926), allows an edge server to host both a small language model (SLM) and an LLM, assigning each user to either AD or SD and running both modes simultaneously. The goal is to maximize the number of served users by formulating a task throughput maximization problem. This research is relevant to next-generation mobile networks where edge inference is seen as a promising paradigm for providing LLM services. The paper was announced as a cross-type submission on arXiv.
Key facts
- BALANCE is a hybrid autoregressive-speculative inference framework for edge LLM inference.
- It addresses the latency-memory tradeoff between autoregressive decoding (AD) and speculative decoding (SD).
- Edge server hosts both a small language model (SLM) and a large language model (LLM).
- Each user is assigned to either AD or SD, and both modes run simultaneously.
- The framework aims to maximize the number of served users.
- A task throughput maximization problem is formulated.
- The paper is available on arXiv with ID 2608.05926.
- The research targets next-generation mobile networks.
Entities
Institutions
- arXiv