← Back to the library
DEM305Level 300Artificial Intelligence

LLM Inference on Trainium with vLLM & PyTorch Native

Guy Almog, Senior Solutions Architect, AWS · Guy Ben Baruch, Principal Solutions Architect, AWS

AWS Summit Tel Aviv 2026 · AWS Demos · 25 min · 16 pages

Join us for an immersive demo on deploying and optimizing large language models for scalable inference at enterprise scale - powered by the PyTorch Native experience on AWS Neuron. PyTorch Native eliminates custom compilation friction, enabling a familiar PyTorch workflow directly on AWS Trainium hardware. The session demonstrates distributed LLM serving with vLLM on Amazon EKS, showcasing how PyTorch Native on Neuron maximizes throughput and cost efficiency. We'll also dive into AWS Neuron Explorer for profiling — showing how to identify bottlenecks and optimize inference pipelines at scale.

Generated automatically from the recording's transcript. It's meant to replace re-watching, not the recording itself — verify against the video at the timestamp before quoting formally. Slides belong to the speakers and organisers. Produced with recording-to-pdf.

AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog