vLLM
por vLLM Project · docs.vllm.ai ↗
Motor de inferência e serving de código aberto para grandes modelos de linguagem, conhecido pelo PagedAttention de alto rendimento.
open-source inference serving llm
- Categoria
- Plataformas e infraestrutura
- Modelo de negócio
- Aberto (pesos/código aberto)
- Disponibilidade
- Global
- Lançado
- 2023-06
- Registro atualizado
- 2026-07-30
- Licença
- Apache-2.0
- Documentação
- docs.vllm.ai ↗
- Fundada
- 2023
- Sede
- US
Visão geral
vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.
Principais recursos
- PagedAttention for efficient KV-cache memory management
- High-throughput continuous batching of requests
- OpenAI-compatible API server out of the box
- Broad support for open model architectures and quantization
Casos de uso
- Serving open LLMs in production at high throughput
- Self-hosting an OpenAI-compatible inference endpoint
- Benchmarking and research on efficient LLM inference
Preços
vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.
Perguntas frequentes
What is vLLM used for?
vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.
Is vLLM free?
Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.
What is PagedAttention?
PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.
Produtos similares
- Amazon Bedrock — Serviço gerenciado da AWS para criar aplicações com modelos de fundação de vários provedores por meio de uma única API.
- Amazon Nova — Portfólio da AWS de modelos de fundação e serviços criados pela Amazon para texto, multimodalidade, fala, construção de modelos personalizados e agentes que automatizam a IU.
- Azure AI Foundry — Plataforma da Microsoft para criar, avaliar e implantar aplicações e agentes de IA no Azure.
- Baseten — Plataforma para implantar e servir modelos de machine learning como APIs escaláveis.
- Cerebras — Empresa de computação de IA cuja nuvem de inferência serve modelos de linguagem abertos em hardware em escala de wafer.
- Civitai — Plataforma comunitária para compartilhar e descobrir modelos abertos de geração de imagens, LoRAs e arte gerada.
Todas as alternativas a vLLM → · Todos os produtos de Plataformas e infraestrutura →
Fontes
Este registro foi revisado pela última vez em 2026-07-30.
Registro legível por máquina: /api/products/vllm.json · Encontrou um erro? Sugerir uma correção