LLM Inference on Trainium with vLLM & PyTorch Native
Guy Almog, Senior Solutions Architect, AWS · Guy Ben Baruch, Principal Solutions Architect, AWS
AWS Summit Tel Aviv 2026 · AWS Demos · 25 min · 16 pages
Join us for an immersive demo on deploying and optimizing large language models for scalable inference at enterprise scale - powered by the PyTorch Native experience on AWS Neuron. PyTorch Native eliminates custom compilation friction, enabling a familiar PyTorch workflow directly on AWS Trainium hardware. The session demonstrates distributed LLM serving with vLLM on Amazon EKS, showcasing how PyTorch Native on Neuron maximizes throughput and cost efficiency. We'll also dive into AWS Neuron Explorer for profiling — showing how to identify bottlenecks and optimize inference pipelines at scale.
Generated automatically from the recording's transcript. It's meant to replace re-watching, not the recording itself — verify against the video at the timestamp before quoting formally. Slides belong to the speakers and organisers. Produced with recording-to-pdf.