Modern distributed systems have outpaced traditional SRE tooling. Static thresholds, manual alert correlation, and reactive runbooks were built for a simpler era — not for Kubernetes-native, microservices-heavy production environments generating millions of signals per minute. The result is alert fatigue, unsustainable on-call burden, and SRE teams spending more time firefighting than engineering. This session presents a practitioner’s blueprint for applying AIOps — Artificial Intelligence for IT Operations — directly to the SRE workflow. Drawing from real-world experience managing production reliability at a global financial institution serving 200M+ customers, five US patents in AIOps platform management, and the recently published book SRE with AIOps , the speaker will walk through how AI-driven intelligence transforms three core SRE disciplines:
Intelligent Incident Management — how ML-powered alert correlation collapses noise, accelerates triage, and compresses MTTR without burning out your on-call rotation Predictive Anomaly Detection — replacing static thresholds with dynamic baselines using autoencoders and isolation forests, catching degradation before it breaches SLOs Generative AI in the SRE Loop — architecture and real-world patterns for deploying LLM-powered assistants that execute runbooks, correlate incidents, and draft postmortems through natural language
Attendees will leave with concrete architectural patterns, an ROI framework for justifying AIOps investment internally, and a clear picture of where the SRE role is heading as autonomous remediation and agentic AI move from concept to production reality.
