vLLM

by vLLM Project · docs.vllm.ai ↗

Platforms & Infrastructure Open (open weights/source) Availability: Global Launched 2023-06

Open-source inference and serving engine for large language models, known for high-throughput PagedAttention.

open-source inference serving llm

Category
Platforms & Infrastructure
Business model
Open (open weights/source)
Availability
Global
Launched
2023-06
Record updated
2026-07-30
License
Apache-2.0
Documentation
docs.vllm.ai ↗
Founded
2023
Headquarters
US
Canonical URL
https://globalaiproductindex.com/products/vllm/

Overview

vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.

Key features

Use cases

Pricing

vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.

Frequently asked questions

What is vLLM used for?

vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.

Is vLLM free?

Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.

What is PagedAttention?

PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.

Similar products

All vLLM alternatives → · All Platforms & Infrastructure products →

Sources

This record was last reviewed on 2026-07-30.

Machine-readable record: /api/products/vllm.json · Spot an error? Suggest a correction