Cloud operations teams face an expanding volume of telemetry, alerts, and service dependencies that exceeds the capacity of manual, reactive practices. This paper reports practical experience applying Artificial Intelligence for IT Operations (AIOps) to improve service reliability across large-scale cloud environments. We organize the relevant practices into a closed-loop framework, named CLAIR, which connects observability, machine-learning-based anomaly detection, event correlation, and automated remediation into a continuous reliability cycle. Drawing on operational deployments, we describe how intelligent automation reduces alert noise, shortens mean time to detection and recovery, and lowers the cognitive load on on-call engineers. We present an illustrative evaluation comparing a reactive baseline against an AIOps-augmented pipeline, summarize lessons learned, and discuss governance and privacy considerations for deploying machine learning on operational data. The results indicate that disciplined, human-supervised automation can materially improve reliability outcomes while keeping engineers in control of consequential actions.
@artical{m11122022ijsea11121073,
Title = "AIOps-Driven Service Reliability: Transforming Cloud Operations Through Intelligent Automation",
Journal ="International Journal of Science and Engineering Applications (IJSEA)",
Volume = "11",
Issue ="12",
Pages ="500 - 506",
Year = "2022",
Authors ="Mourya Chigurupati"}