
vLLM: LLM inference and serving
About this event
This session introduces vLLM, an open-source library for LLM inference and serving. We will build an intuitive systems-level understanding of how vLLM approaches these production challenges, from a token’s journey through prefill and decode to the scheduling and memory-management techniques that make high-throughput serving practical. Start exploring at vllm (https://docs.vllm.ai/en/latest/). Slides for past meetups posted: Github (https://github.com/YanXuHappygela/LLM-reading-group/tree/main) Recordings posted at: YanAITalk (https://www.youtube.com/@yanaitalk/videos) Feel free to reach out if you want to present at upcoming meetups! Note: You must have a Zoom account to login (free account is sufficient). Zoom Link will be posted to the event page one day before the meetup.
This event has ended.
How was it?
Reviews
This event has finished — be the first to review it!
Questions & comments
Ask the host anything — replies are visible to everyone.
—