vLLM

por vLLM Project · docs.vllm.ai ↗

Plataformas e infraestrutura Aberto (pesos/código aberto) Disponibilidade: Global Lançado 2023-06

Motor de inferência e serving de código aberto para grandes modelos de linguagem, conhecido pelo PagedAttention de alto rendimento.

open-source inference serving llm

Categoria
Plataformas e infraestrutura
Modelo de negócio
Aberto (pesos/código aberto)
Disponibilidade
Global
Lançado
2023-06
Registro atualizado
2026-07-30
Licença
Apache-2.0
Documentação
docs.vllm.ai ↗
Fundada
2023
Sede
US
URL canônica
https://globalaiproductindex.com/pt/products/vllm/

Visão geral

vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.

Principais recursos

Casos de uso

Preços

vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.

Perguntas frequentes

What is vLLM used for?

vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.

Is vLLM free?

Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.

What is PagedAttention?

PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.

Produtos similares

Todas as alternativas a vLLM → · Todos os produtos de Plataformas e infraestrutura →

Fontes

Este registro foi revisado pela última vez em 2026-07-30.

Registro legível por máquina: /api/products/vllm.json · Encontrou um erro? Sugerir uma correção