vLLM
by vLLM Project · docs.vllm.ai ↗
Open-source inference and serving engine for large language models, known for high-throughput PagedAttention.
open-source inference serving llm
- Category
- Platforms & Infrastructure
- Business model
- Open (open weights/source)
- Availability
- Global
- Launched
- 2023-06
- Record updated
- 2026-07-30
- License
- Apache-2.0
- Documentation
- docs.vllm.ai ↗
- Founded
- 2023
- Headquarters
- US
- Canonical URL
- https://globalaiproductindex.com/products/vllm/
Overview
vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.
Key features
- PagedAttention for efficient KV-cache memory management
- High-throughput continuous batching of requests
- OpenAI-compatible API server out of the box
- Broad support for open model architectures and quantization
Use cases
- Serving open LLMs in production at high throughput
- Self-hosting an OpenAI-compatible inference endpoint
- Benchmarking and research on efficient LLM inference
Pricing
vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.
Frequently asked questions
What is vLLM used for?
vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.
Is vLLM free?
Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.
What is PagedAttention?
PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.
Similar products
- Amazon Bedrock — Managed AWS service for building applications with foundation models from multiple providers via one API.
- Amazon Nova — AWS portfolio of Amazon-built foundation models and services for text, multimodal, speech, custom model building, and UI-automating agents.
- Azure AI Foundry — Microsoft's platform for building, evaluating, and deploying AI applications and agents on Azure.
- Baseten — Platform for deploying and serving machine-learning models as scalable APIs.
- Cerebras — AI compute company whose inference cloud serves open language models on wafer-scale hardware.
- Civitai — Community platform for sharing and discovering open image-generation models, LoRAs, and generated art.
All vLLM alternatives → · All Platforms & Infrastructure products →
Sources
This record was last reviewed on 2026-07-30.
Machine-readable record: /api/products/vllm.json · Spot an error? Suggest a correction