Authors: Shreya Kulkarni
Abstract: Cloud-native architectures have fundamentally redefined modern software engineering by enabling dynamic scalability, elasticity, and rapid deployment through the adoption of microservices architecture, containerization, DevOps practices, and multi-cloud infrastructures. Platforms built on distributed service-oriented principles allow independent deployment and horizontal scaling; however, the inherent decentralization and runtime dynamism introduce substantial challenges in maintaining consistent performance engineering metrics (latency, throughput, resource utilization) and ensuring robust reliability engineering attributes (availability, fault tolerance, resilience). This review critically synthesizes contemporary research on performance optimization, reliability modeling, and resilience engineering within cloud-native systems. Key architectural paradigms—including scalable REST-based microservices, serverless computing models, and automated multi-cloud provisioning—are examined to analyze trade-offs between scalability, latency variability, cold-start overhead, and fault propagation. Comparative insights from microservices versus serverless performance studies highlight workload-sensitive design considerations, while resilience-focused research grounded in well-architected frameworks, redundancy strategies, and disaster recovery planning demonstrates the importance of proactive reliability integration. The review further incorporates probabilistic and analytical reliability modeling techniques, such as Markov chain–based estimation, risk assessment frameworks, and reliability block diagrams, illustrating their applicability in predicting failure states within distributed cloud environments. In addition, cross-disciplinary methodologies derived from structural reliability engineering, manufacturing optimization, and semiconductor performance analysis are discussed to emphasize transferable quantitative approaches for degradation modeling, system robustness evaluation, and performance–reliability co-optimization. Findings indicate that performance degradation often precedes reliability failures in distributed systems, reinforcing the necessity for observability-driven architectures, autonomic management frameworks, adaptive autoscaling mechanisms, and predictive analytics. Emerging solutions leveraging AI-driven anomaly detection, self-healing orchestration, and multi-cloud fault isolation strategies demonstrate promise but remain constrained by challenges such as telemetry data overload, energy-efficient redundancy design, and cross-cloud synchronization complexity. Overall, this review identifies critical research gaps in predictive failure modeling, performance-aware resilience strategies, and sustainable reliability engineering, and outlines future directions toward intelligent, self-adaptive cloud-native infrastructures capable of balancing efficiency, scalability, cost optimization, and operational robustness in highly dynamic distributed ecosystems.
DOI: https://doi.org/10.5281/zenodo.18669308
International Journal of Science, Engineering and Technology