Collaborative Distributed Inference System for Efficient QoS-Aware Autoscaling
A recent research article introduces a collaborative distributed inference framework that integrates dedicated infrastructure with resources provided by users to meet the increasing demand for AI inference. This framework employs dedicated resources to ensure a baseline capacity for quality of service (QoS), while user-contributed resources manage demand surges without necessitating a proportional increase in centralized infrastructure. The authors create a high-dimensional generative Markov model featuring structured temporal factorization to capture the stochastic and dynamic relationships among users, resources, tasks, and policies. This model facilitates simulation and lays the groundwork for optimizing task scheduling and QoS-aware resource allocation. Evaluations demonstrate that distributed scheduling policies can effectively leverage volunteered resources, minimizing the need for centralized infrastructure expansion while preserving QoS. The study can be accessed on arXiv under identifier 2608.11840.
Key facts
- The paper proposes a collaborative distributed inference system combining dedicated and user-contributed resources.
- Dedicated resources provide baseline capacity for QoS, while volunteered resources absorb increasing demand.
- A high-dimensional generative Markov model with structured temporal factorization is developed.
- The model supports simulation and optimization of task scheduling and resource allocation.
- The system is evaluated across user populations, resource capacities, and scheduling policies.
- Simulations show distributed scheduling policies effectively utilize volunteered resources.
- The approach reduces the need for proportional growth in centralized infrastructure.
- The paper is available on arXiv with identifier 2608.11840.
Entities
Institutions
- arXiv