Learn how to run open-weight AI models and power local agents, like AI coding tools, using vLLM, an open-source and high-performance inference engine. We’ll cover how to optimize models, expose them via OpenAI-compatible APIs, and tune runtime settings to maximize performance on constrained hardware. We’ll also run live benchmarks to inspect token throughput and balance latency trade-offs, which can be done using any AI serving tool.
Cedric Clyburn (@cedricclyburn) is a Senior Developer Advocate at Red Hat and a Kubernetes Community Day organizer. Passionate about open-source software, he contributes to projects like Podman and vLLM and speaks globally at events like Devoxx and Linux Foundation conferences. He also creates tech educational content across written and video formats, reaching over 2 million views online.