There was an error while loading. Please reload this page.
A high-throughput and memory-efficient inference and serving engine for LLMs
Python 89.1k 20.7k
A framework for efficient model inference with omni-modality models
Python 6.1k 1.5k
Common recipes to run vLLM
JavaScript 972 375
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Python 3.7k 615
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
Python 730 183
A programmable Mixture-of-Models router for heterogeneous LLM inference
Go 5.2k 815
vLLM Daily Summarization of Merged PRs
TPU inference for vLLM, with unified JAX and PyTorch support.
Community maintained hardware plugin for vLLM on Ascend
Community maintained hardware plugin for vLLM on Apple Silicon
This repo hosts code for vLLM CI & Performance Benchmark infrastructure.
Community maintained hardware plugin for vLLM on Intel Gaudi
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
Loading…