Pinned
Introducing llm-d on stage at Red Hat Summit was truly a privilege ...
LLM inference is too slow, too expensive, and too hard to scale.
🚨 Introducing llm-d, a Kubernetes-native distributed inference framework, to change that—using vLLM (@vllm_project), smart scheduling, and disaggregated compute.
Here’s how it works—and how you can use it today:







