Search papers, labs, and topics across Lattice.
This paper introduces a user-assisted collaborative distributed inference system that leverages both dedicated infrastructure and user-contributed resources to efficiently manage AI inference demand. By employing a high-dimensional generative Markov model with structured temporal factorization, the authors optimize task scheduling and resource allocation to maintain quality of service (QoS) while minimizing the need for centralized infrastructure. The results from simulations indicate that as user populations increase, the proposed system significantly enhances request completion rates and reduces latency, showcasing its potential for infrastructure-efficient autoscaling.
User-assisted collaborative inference can slash dedicated resource usage while boosting performance as demand scales.
Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. We propose a collaborative distributed inference system combining dedicated infrastructure with resources contributed by service users. Dedicated resources provide baseline capacity for maintaining quality of service (QoS), while volunteered resources absorb increasing demand without proportional growth in centralized infrastructure. To capture stochastic and dynamic interactions among users, resources, tasks, and policies, we develop a high-dimensional generative Markov model with structured temporal factorization. The model supports simulation and provides a foundation for task scheduling and QoS-aware resource allocation optimization. We evaluate the system across user populations, resource capacities, and centralized and distributed scheduling policies. Simulations show that distributed scheduling becomes increasingly advantageous as the user population grows, improving request completion and P99 latency while substantially reducing dedicated resource consumption. These results demonstrate the feasibility of user-assisted collaborative inference for infrastructure-efficient autoscaling.