Efficient Deployment of Transformer-Based Large Language Models on Edge Computing Devices: Strategies, Challenges, Optimization Techniques, and Future Research Directions

30 Jul

Authors: Research Scholar Mr.K.Raju, Assistant Professor Dr.Rahul Kumar Budania

Abstract: A fast developing field of artificial intelligence research involves the nexus of edge computing, Transformer architecture, and Large Language Models (LLMs). Because of their high processing and memory needs, LLMs cannot be directly implemented on low-resource edge devices, despite their remarkable capacity to understand and produce natural language. The implementation of Transformer-based LLMs on edge computing platforms is examined in this study, with an emphasis on methods such as system-level approaches, inference optimization, model compression, and architectural improvements. The main issues with edge devices' constrained power, memory, and CPU capabilities are also covered. This work also looks at new compact LLMs and optimization techniques that increase deployment efficiency, lower computational costs, and enhance data security and privacy. A list of upcoming research challenges to enable effective, scalable, and intelligent edge-based AI applications concludes the survey.

DOI: https://doi.org/10.5281/zenodo.21701699