Skip to content

Pinned Loading

  1. vllm vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 93.3k 23.1k

  2. vllm-omni vllm-omni Public

    A framework for efficient model inference with omni-modality models

    Python 7.1k 1.9k

  3. recipes recipes Public

    Common recipes to run vLLM

    JavaScript 1k 442

  4. llm-compressor llm-compressor Public

    State-of-the-art LLM compression, built for production inference with vLLM

    Python 3.9k 681

  5. speculators speculators Public

    A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

    Python 863 245

  6. semantic-router semantic-router Public

    An open, programmable decision layer for models and compute.

    Go 6k 1k

Repositories

Showing 10 of 50 repositories
  • vllm-metal Public

    Community maintained hardware plugin for vLLM on Apple Silicon

    vllm-project/vllm-metal's past year of commit activity
    Python 1,820 Apache-2.0 289 14 14 Updated Oct 7, 2026
  • semantic-router Public

    An open, programmable decision layer for models and compute.

    vllm-project/semantic-router's past year of commit activity
    Go 6,041 Apache-2.0 1,015 414 (1 issue needs help) 197 Updated Oct 7, 2026
  • vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    vllm-project/vllm's past year of commit activity
    Python 93,312 Apache-2.0 23,075 2,550 (30 issues need help) 5,000+ Updated Oct 7, 2026
  • vllm-ascend Public

    Community maintained hardware plugin for vLLM on Huawei Ascend

    vllm-project/vllm-ascend's past year of commit activity
    Python 2,925 Apache-2.0 2,386 1,569 (77 issues need help) 2,207 Updated Oct 7, 2026
  • vllm-gaudi Public

    Community maintained hardware plugin for vLLM on Intel Gaudi

    vllm-project/vllm-gaudi's past year of commit activity
    Python 61 Apache-2.0 157 14 57 Updated Oct 7, 2026
  • aibrix Public

    Cost-efficient and pluggable Infrastructure components for GenAI inference

    vllm-project/aibrix's past year of commit activity
    Go 5,126 Apache-2.0 718 343 (18 issues need help) 37 Updated Oct 7, 2026
  • tpu-inference Public

    TPU inference for vLLM, with unified JAX and PyTorch support.

    vllm-project/tpu-inference's past year of commit activity
    Python 453 Apache-2.0 331 109 (2 issues need help) 414 Updated Oct 7, 2026
  • guidellm Public

    Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs

    vllm-project/guidellm's past year of commit activity
    Python 1,666 Apache-2.0 257 61 37 Updated Oct 7, 2026
  • humming Public

    Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.

    vllm-project/humming's past year of commit activity
    Python 247 Apache-2.0 47 6 10 Updated Oct 7, 2026
  • ci-infra Public

    This repo hosts code for vLLM CI & Performance Benchmark infrastructure.

    vllm-project/ci-infra's past year of commit activity
    Python 58 Apache-2.0 91 0 82 Updated Oct 7, 2026