Deploying Large Language Models on Edge Computing Devices: Architectures, Optimization Techniques, Challenges, and Future Directions

30 Jul

Authors: Research Scholar Mr.K.Raju, Assistant Professor Dr.Rahul Kumar Budania

Abstract: By providing sophisticated natural language interpretation, reasoning, and decision-making capabilities, large language models (LLMs) have transformed artificial intelligence (AI). However, many real-time and mission-critical applications cannot profit from the standard cloud-based deployment of LLMs due to concerns such high communication latency, excessive bandwidth utilization, privacy issues, and reliance on constant network connectivity. By moving AI processors closer to data sources and enabling intelligent processing directly on edge devices while lowering reaction times, boosting system autonomy, and improving privacy, integrating LLMs with Edge Intelligence (EI) gets around these restrictions. By looking at the most recent architectural frameworks, deployment techniques, and learning paradigms created for resource-constrained edge contexts, this survey offers a thorough analysis of LLM-based Edge Intelligence. To facilitate the effective deployment of LLMs across different edge infrastructures, it covers cloud-edge, edge-only, cloud-edge-client, federated learning, peer-to-peer learning, and knowledge distillation techniques. Advanced optimization strategies like model compression, quantization, pruning, computation offloading, effective memory management, edge caching, and lightweight Small Language Models (SLMs), which dramatically lower computational overhead while preserving high inference performance on resource-constrained devices, are also covered in the survey. The report highlights numerous real-world applications of LLM-powered Edge Intelligence, including software engineering, robotics, intelligent transportation systems, intelligent healthcare, autonomous driving, the Industrial Internet of Things (IIoT), upcoming 6G communication networks, and smart cities. The survey highlights the significance of creating transparent, moral, and reliable AI systems that adhere to new moral and legal standards. Scalable LLM deployment, effective resource use, energy-conscious computing, intelligent model adaptation, distributed learning, secure edge collaboration, and smooth interaction with future 6G and beyond communication networks are just a few of the subjects covered in the survey's conclusion. It also describes future research directions and highlights existing research shortages. Overall, because it offers a comprehensive analysis of the architectures, optimization techniques, applications, security issues, and potential future advancements of LLM-based Edge Intelligence, the paper is a useful tool for researchers, practitioners, and developers. This enables the creation of effective, safe, scalable, and reliable edge computing systems driven by AI.

DOI: https://doi.org/10.5281/zenodo.21701918