Skip to content
@llm-d

llm-d

llm-d enables high performance distributed inference in production on Kubernetes

Welcome to llm-d: a Kubernetes-native high-performance distributed LLM inference framework

GitHub Org's stars Documentation License

Join Slack X (formerly Twitter) Follow Bluesky LinkedIn Reddit YouTube

llm-d is a Kubernetes-native high-performance distributed LLM inference framework that provides the fastest time-to-value and competitive performance per dollar. Built on vLLM, Kubernetes, and Inference Gateway, llm-d offers modular solutions for distributed inference with features like KV-cache aware routing and disaggregated serving.

🚀 Quick Start Guide

New to llm-d? Here's how to get started:

  1. Join our Slack 💬Get your invite and visit llm-d.slack.com
  2. Explore our code 📂GitHub Organization
  3. Join a meeting 📅Add calendar
  4. Pick your area 🎯 → Browse Special Interest Groups.

📚 Key Resources

💬 Communication Channels

🗓️ Regular Meetings

All meetings are open to the public! 🌟

  • 📅 Weekly Standup: Every Wednesday at 12:30pm ET - Project updates and open discussion
  • 🎯 SIG Meetings: Various times throughout the week - See SIG details for schedules

Join to participate, ask questions, or just listen and learn!

🎯 Special Interest Groups (SIGs)

Want to dive deeper into specific areas? Our Special Interest Groups are focused teams working on different aspects of llm-d:

  • Inference Scheduler - Intelligent request routing and load balancing
  • Benchmarking - Performance testing and optimization
  • PD-Disaggregation - Prefill/decode separation patterns
  • KV-Disaggregation - KV caching and distributed storage
  • Installation - Kubernetes integration and deployment
  • Autoscaling - Traffic-aware autoscaling and resource management
  • Observability - Monitoring, logging, and metrics

View more SIG Details →

🤝 How to Contribute

Getting Involved

Contributing Code

  1. Read Guidelines: Review our Code of Conduct and contribution process
  2. Sign Commits: All commits require DCO sign-off (git commit -s)

Ways to Contribute

  • 🐛 Bug fixes and small features - Submit PRs directly to component repos
  • 🚀 New features with APIs - Require project proposals
  • 📚 Documentation - Help improve guides and examples
  • 🧪 Testing & Benchmarking - Contribute to our test coverage
  • 💡 Experimental features - Start in llm-d-incubation org

🔒 Security & Safety

🌐 Connect With Us

Follow llm-d across social platforms for updates, discussions, and community highlights:

❓ Need Help?

Questions? Ideas? Just want to chat? We're here to help! The llm-d community team is friendly and responsive.


License: Apache 2.0

Pinned Loading

  1. llm-d llm-d Public

    Achieve state of the art inference performance with modern accelerators on Kubernetes

    Shell 4.6k 787

  2. llm-d-router llm-d-router Public

    llm-d Router: The intelligent entry point for inference requests

    Go 356 389

  3. llm-d-kv-cache llm-d-kv-cache Public

    Distributed KV cache scheduling & offloading libraries

    Go 180 159

  4. llm-d-benchmark llm-d-benchmark Public

    llm-d benchmark scripts and tooling

    Python 70 144

Repositories

Showing 10 of 22 repositories
  • llm-d-router Public

    llm-d Router: The intelligent entry point for inference requests

    llm-d/llm-d-router's past year of commit activity
    Go 356 Apache-2.0 389 229 (6 issues need help) 109 Updated Sep 23, 2026
  • llm-d Public

    Achieve state of the art inference performance with modern accelerators on Kubernetes

    llm-d/llm-d's past year of commit activity
    Shell 4,635 Apache-2.0 787 127 (11 issues need help) 92 Updated Sep 23, 2026
  • llm-d-benchmark Public

    llm-d benchmark scripts and tooling

    llm-d/llm-d-benchmark's past year of commit activity
    Python 70 Apache-2.0 144 29 14 Updated Sep 23, 2026
  • llm-d-inference-sim Public

    A lightweight, configurable, and real-time simulator designed to mimic the behavior of vLLM without the need for GPUs or running actual heavy models.

    llm-d/llm-d-inference-sim's past year of commit activity
    Go 201 Apache-2.0 153 10 (1 issue needs help) 8 Updated Sep 23, 2026
  • llm-d-batch-gateway Public

    llm-d-batch-gateway is a standalone, backend-agnostic, OpenAI-compatible batch API and processing engine

    llm-d/llm-d-batch-gateway's past year of commit activity
    Go 24 Apache-2.0 52 22 17 Updated Sep 23, 2026
  • llm-d-async Public

    Asynchronous Processor for Inference Gateway. Orchestrator of queues

    llm-d/llm-d-async's past year of commit activity
    Go 14 Apache-2.0 47 17 10 Updated Sep 22, 2026
  • llm-d-kv-cache Public

    Distributed KV cache scheduling & offloading libraries

    llm-d/llm-d-kv-cache's past year of commit activity
    Go 180 Apache-2.0 159 21 26 Updated Sep 22, 2026
  • llm-d-autoscaling Public

    Variant optimization autoscaler for distributed inference workloads

    llm-d/llm-d-autoscaling's past year of commit activity
    Go 57 Apache-2.0 96 18 14 Updated Sep 21, 2026
  • llm-d-prism Public

    Performance analysis for distributed inference systems

    llm-d/llm-d-prism's past year of commit activity
    JavaScript 14 Apache-2.0 25 12 4 Updated Sep 21, 2026
  • llm-d-infra Public

    repo for CI and infrastructure required to maintain llm-d org member repos

    llm-d/llm-d-infra's past year of commit activity
    Shell 2 Apache-2.0 38 26 9 Updated Sep 18, 2026