← Back to the library
SVS305Level 300Artificial IntelligenceOpen SourceServerless & Containers

vLLM on AWS: testing to production and everything in between

צחי, Solutions Architect (AI Infrastructure), AWS | ואומרי, Senior Open Source Engineer, AWS

AWS Summit Tel Aviv 2026 · AWS for Containers and Platform Engineering · 41 min · 18 pages

This session explores practical architectural patterns for deploying and scaling large language models (LLMs) in production environments. We'll walk through a comprehensive journey from initial testing to production deployment, covering essential steps including model evaluation with vLLM, performance benchmarking, and optimization techniques. Attendees will learn how to implement efficient autoscaling solutions using Ray and vLLM, compare different inference servers like Triton and vLLM, and understand their trade-offs. The session concludes with a deep dive into productionalization using AIBrix, providing attendees with actionable insights for building robust, scalable LLM infrastructures. Suitable for ML engineers and architects working on LLM deployments.

Generated automatically from the recording's transcript. It's meant to replace re-watching, not the recording itself — verify against the video at the timestamp before quoting formally. Slides belong to the speakers and organisers. Produced with recording-to-pdf.

AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog