Social📍 LondonOpen to all

AWS AI In Practice #6

WhenWed, Aug 26, 6:00 PMStarts in 18 days📅 Add to calendarWhereAutogenAI123 Pentonville Rd, N1 9LG, London🗺 Apple Maps🗺 Google Maps🚕 UberHostAWS AI in PracticeCostNot stated — check with the hostOneJoy doesn't handle payments — settle directly with the host or venue.CapacityOpen — no spot limit

About this event

We’re delighted to welcome Anton Nazaruk (https://www.linkedin.com/in/anton-nazaruk/), CTO, Cloud Combinator, Ivaylo Iliev (https://www.linkedin.com/in/ivaylo-iliev3/), Partner Solutions Architect, AWS and Silvia Lehnis (https://www.linkedin.com/in/silvialehnis/), Chief AI Officer, UBDS Digital Here’s what they’re bringing: Seven ways to buy GPU compute on AWS. One wrong choice and the bill climbs while your training run stalls. Anton builds production-ready data and AI platforms at Cloud Combinator, and he’s bringing the receipts - a pre-recorded walkthrough of a real distributed training run, including a node failure, replacement, and full job recovery. Tonight he’s joined by Ivaylo, giving us the complete blueprint: Capacity Blocks, SageMaker HyperPod, EFA networking, FSx storage, and Slurm or EKS orchestration, with cost control designed in from day one. Silvia and her team at UBDS Digital took an intelligent document processing solution on Amazon Bedrock and tested everything - prompting, model selection, document handling, and validation - to find out what actually moves extraction accuracy on messy, real-world documents. Tonight she’s walking us through the results: which methods earn their keep, which don’t, and the trade-offs between value and implementation effort. A big thank you to our sponsors Cloudscaler (https://rebrand.ly/cloudscaler), Rayo (https://rebrand.ly/rayo-cloud) & The Scale Factory (https://rebrand.ly/scalefactory) for making this event possible. Programme: 18:00: Arrival, registration 18:15: Talks start 20:00: Networking with food and a drink provided by the generosity of our sponsors. Session 1: From Zero to HyperPod: Choosing, Operating, and Cost-Controlling Distributed Model Training Infrastructure on AWS with Anton Nazaruk & Ivaylo Iliev Training large models on AWS is no longer just an ML problem - it is a capacity, infrastructure, reliability, and cost-control problem. This talk gives engineers a practical framework for choosing between EC2 On-Demand, Spot, Savings Plans, Capacity Blocks, SageMaker Training Jobs, SageMaker Training Plans, and SageMaker HyperPod. We’ll look at when each option makes sense, what tradeoffs they introduce, and how to avoid common mistakes around GPU availability, quota planning, interruptions, and runaway cost. We’ll then walk through a repeatable distributed training blueprint: compute fleet, EFA networking, FSx/S3 storage, Slurm or EKS orchestration, observability, checkpointing, and failure recovery. The session includes a demo-style walkthrough of launching a distributed training job and showing how node failure and recovery should be handled in a production-ready setup. The goal is for attendees to leave with a clear mental model of how to run distributed model training on AWS reliably, how to choose the right service or capacity model, and how to make cost and failure recovery part of the architecture from day one. Learning Takeaways • Choose between EC2 On-Demand, Spot, Savings Plans, Capacity Blocks, SageMaker Training Jobs, Training Plans, and HyperPod with a clear framework for when each makes sense. • Build a repeatable distributed training blueprint - compute fleet, EFA networking, FSx/S3 storage, Slurm or EKS orchestration, observability, and checkpointing. • Design cost control and node failure recovery into your training architecture from day one. Anton Anton Nazaruk is CTO at Cloud Combinator, where he works on cloud architecture, AI infrastructure, and distributed systems. He helps teams design production-ready platforms for data and AI workloads on AWS, with a focus on reliability, cost control, and repeatable infrastructure patterns. Ivaylo is a Partner Solutions Architect at AWS, where he helps organisations optimise their cloud solutions and accelerate their digital transformation journeys. With a focus on artificial intelligence and machine learning, he works closely with AWS partners to develop innovative solutions that address real-w

Join this event

Before you join

OneJoy is where people find each other — the host organises the event, not us. Check who is hosting, judge whether it suits you, and take the same care you would meeting anyone new. Under-18s should come with a parent or guardian. Any money changes hands directly with the host; OneJoy never handles payments.

Sign in Have an account? Sign in and we'll fill this in for you.

Only shared with the host.

More options

Questions & comments

Ask the host anything — replies are visible to everyone.

Report

Report this to the OneJoy team

Tell us what is wrong. We read every report.