CarletonCASNewcastle UniversityUESTCUMacauUniversity of HyderabadUniversity of ljubljanaApr 19, 2026arXiv:2604.17227

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda

Minxian Xu, Jingfeng Wu, Shengye Song, Satish Narayana Srirama, Bahman Javad, Rajiv Ranjan, Devki Nandan Jha, Sa Wang, Huanle Xu, Zizhao Mo, Shuo Ren, Thomas Kunz, Petar Kochovski, Vlado Stankovski, Chengzhong Xu, Rajkumar Buyya

AI Summary

This paper examines the challenges of deploying and scaling large language models (LLMs) and advocates for the adoption of cloud-native and distributed system architectures. It highlights the limitations of traditional systems in meeting the computational demands of LLMs, particularly during training and inference. The paper proposes a research agenda focused on areas like serverless inference, quantum computing, and federated learning to drive future LLM innovation and scalability.

Key Contribution

LLM scaling bottlenecks demand a shift towards cloud-native architectures and distributed systems, unlocking potential gains from serverless inference and quantum computing.

Abstract

The rapid rise of Large Language Models (LLMs) has revolutionized various artificial intelligence (AI) applications, from natural language processing to code generation. However, the computational demands of these models, particularly in training and inference, present significant challenges. Traditional systems are often unable to meet these requirements, necessitating the integration of cloud-native and distributed architectures. This paper explores the role of cloud platforms and distributed systems in supporting the scalability, efficiency, and optimization of LLMs. We discuss the complexities of LLM deployment, including data management, resource optimization, and the need for microservices, autoscaling, and hybrid cloud-edge solutions. Additionally, we examine emerging research trends, such as serverless inference, quantum computing, and federated learning, and their potential to drive the next phase of LLM innovation. The paper concludes with a roadmap for future developments, emphasizing the need for continued research, standardization, and cross-sector collaboration to sustain the growth of LLMs in both research and enterprise applications.

Distributed Systems & Hardware Scaling Laws & Emergent Abilities Training Efficiency & Optimization

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda

Related Papers