LLM Inference Optimization at Scale: DevOps Lessons from Production AI Systems

Deploying large language models in production is not a model problem — it is a DevOps problem. In this session, Tejas Patel, Senior Software Development Engineer at Amazon, shares hard-won lessons from building and operating AI inference pipelines at massive scale, including a 1,000 TB data migration and AI personalization systems serving millions of users daily.

Attendees will walk away with practical strategies for reducing LLM inference latency through adaptive token routing and speculative decoding, managing GPU resource contention in multi-tenant environments, scaling distributed AI systems without sacrificing reliability, and integrating AI inference into existing CI/CD and MLOps pipelines.

This session bridges the gap between AI research and production engineering — showing DevOps practitioners exactly what breaks when you scale LLMs, and how to fix it.

Speaker

tejas-pravinbhai-patel

Tejas Pravinbhai Patel

 

Tejas Pravinbhai Patel is an IEEE Award-Winning Researcher, acclaimed Keynote Speaker, and Senior Software Engineer at Amazon, where he architects large-scale, distributed AI systems deployed in

...