vLLM

来自 vLLM Project · docs.vllm.ai ↗

平台与基础设施 开放(开放权重/源码) 可用性:全球 发布于 2023-06

面向大语言模型的开源推理与服务引擎,以高吞吐的 PagedAttention 而知名。

open-source inference serving llm

分类
平台与基础设施
商业模式
开放(开放权重/源码)
可用性
全球
发布
2023-06
记录更新
2026-07-30
许可证
Apache-2.0
文档
docs.vllm.ai ↗
成立时间
2023
总部所在地
US
规范网址
https://globalaiproductindex.com/zh/products/vllm/

概述

vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.

主要功能

适用场景

定价

vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.

常见问题

What is vLLM used for?

vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.

Is vLLM free?

Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.

What is PagedAttention?

PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.

同类产品

查看vLLM的全部替代品 → · 查看全部平台与基础设施产品 →

来源

本记录最近复核于 2026-07-30。

机器可读记录:/api/products/vllm.json · 发现错误?提交纠错