☕ Social📍 AmsterdamOpen to all

AI Inference in the Real World: From Kubernetes to the Far Edge

WhenTue, Oct 27, 6:00 PMStarts in 25 days📅 Add to calendarWhereAWS AmsterdamMr.Treublaan 7, 1097 DP Amsterdam, Netherlands, Amsterdam🗺 Apple Maps🗺 Google Maps🚕 UberHostAI Native NetherlandsCostNot stated — check with the hostOneJoy doesn't handle payments — settle directly with the host or venue.CapacityOpen — no spot limit

About this event

AI workloads don’t all belong on the same model, hardware, or infrastructure. For our 13th AI Native Netherlands meetup, Christian Melendez from AWS (https://aws.amazon.com/) will look at when smaller models and CPUs make more sense than defaulting to large GPU-backed LLMs. William Rizzo from Mirantis will take that question to the far edge: running LLM inference on immutable, disconnected clusters where connectivity is unreliable and remote access can’t be assumed. Together, the talks explore one practical question: What should run where, on what hardware, and how do you keep it reliable in production? A huge thank you to our friends at AWS for hosting us at their Amsterdam office. Food and drinks will be provided! We'll cover: • When AI workloads actually need GPUs — and when CPUs and smaller models are the better fit. • How to combine Small Language Models and larger LLMs without sending every task to the most expensive model. • What changes when inference moves from the datacenter to factories, vehicles, retail sites, and other edge environments. • How immutable, image-based infrastructure can support updates, rollbacks, and reliability across disconnected fleets. • The real-world trade-offs, failure modes, and production lessons behind both approaches. Speaker 1: Christian Melendez (AWS) (https://www.linkedin.com/in/christianmldz/)https://www.linkedin.com/in/christianmldz/ Christian Melendez is a Principal Specialist Solutions Architect at AWS. He helps the region's largest enterprises build efficient, resilient AI and cloud-native workloads on Kubernetes, with a focus on compute efficiency, cost optimisation, and autoscaling at scale. Author of Kubernetes Autoscaling and creator of Karpenter Blueprints, a best-practices repository that grew Karpenter adoption 16.5x across EMEA, he also built Slemify, an open-source framework for fine-tuning and serving Small Language Models on Kubernetes. A regular speaker at AWS re:Invent, KCDs / CNDs, ContainerDays, and others, Christian focuses on making infrastructure simple, observable, and cost-effective at scale. Talk: The Right Tool for the Right Token: When to Reach for a CPU or a GPU While hard reasoning problems require an LLM and a GPU, other tasks can be solved with simpler means. Classifying, routing, embedding, and reranking are structured, high-frequency jobs that a fine-tuned small model (SLM) can handle on CPU at predictable cost. In this session, we will demo a pipeline where each task runs on the best fitting infrastructure. This will include a real CPU-first agent on Amazon EKS running an SLM, and calling an LLM when needed. We will share our learnings from building this pipeline, including how we measured price/performance, its limitations, and which of our design assumptions we found to be wrong. You will leave with confidence to choose the right tool, and an open-source reference architecture to validate your assumptions. Speaker 2: William Rizzo (Mirantis) (https://www.linkedin.com/in/william-rizzo/) William Rizzo is Global Field CTO at Mirantis, where he helps organisations design, build, and run platform engineering, edge, and AI infrastructure initiatives. His career spans engineering, pre-sales, product ownership, and consulting across high-performance computing, storage, and distributed systems. A CNCF and Linkerd Ambassador and a Kairos maintainer, he's a regular speaker at KubeCon on platform engineering and building resilient internal developer platforms. Talk: Inference at the Far Edge: Running LLMs on Immutable, Disconnected Clusters Most inference architectures assume reliable connectivity, abundant infrastructure, and an operator who can access the system when something goes wrong. At the edge, those assumptions disappear. William will show how LLM inference can run on immutable, read-only edge nodes across factories, retail sites, vehicles, and remote facilities. He’ll cover how the model, runtime, and GPU drivers can be packaged into a

Join this event

Before you join

OneJoy is where people find each other — the host organises the event, not us. Check who is hosting, judge whether it suits you, and take the same care you would meeting anyone new. Under-18s should come with a parent or guardian. Any money changes hands directly with the host; OneJoy never handles payments.

Sign in — Have an account? Sign in and we'll fill this in for you.

Only shared with the host.

More options

Questions & comments

Ask the host anything — replies are visible to everyone.

—

⚑ Report

Report this to the OneJoy team

Tell us what is wrong. We read every report.