vLLM
de vLLM Project · docs.vllm.ai ↗
Motor de inferencia y servicio de código abierto para grandes modelos de lenguaje, conocido por su PagedAttention de alto rendimiento.
open-source inference serving llm
- Categoría
- Plataformas e infraestructura
- Modelo de negocio
- Abierto (pesos/código abierto)
- Disponibilidad
- Global
- Lanzamiento
- 2023-06
- Registro actualizado
- 2026-07-30
- Licencia
- Apache-2.0
- Documentación
- docs.vllm.ai ↗
- Fundación
- 2023
- Sede
- US
Resumen
vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.
Funciones principales
- PagedAttention for efficient KV-cache memory management
- High-throughput continuous batching of requests
- OpenAI-compatible API server out of the box
- Broad support for open model architectures and quantization
Casos de uso
- Serving open LLMs in production at high throughput
- Self-hosting an OpenAI-compatible inference endpoint
- Benchmarking and research on efficient LLM inference
Precios
vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.
Preguntas frecuentes
What is vLLM used for?
vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.
Is vLLM free?
Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.
What is PagedAttention?
PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.
Productos similares
- Amazon Bedrock — Servicio gestionado de AWS para crear aplicaciones con modelos fundacionales de varios proveedores mediante una sola API.
- Amazon Nova — Cartera de AWS de modelos fundacionales y servicios creados por Amazon para texto, multimodalidad, voz, creación de modelos personalizados y agentes que automatizan la IU.
- Azure AI Foundry — Plataforma de Microsoft para crear, evaluar e implementar aplicaciones y agentes de IA en Azure.
- Baseten — Plataforma para desplegar y servir modelos de aprendizaje automático como API escalables.
- Cerebras — Empresa de computación de IA cuya nube de inferencia sirve modelos de lenguaje abiertos en hardware a escala de oblea.
- Civitai — Plataforma comunitaria para compartir y descubrir modelos abiertos de generación de imágenes, LoRAs y arte generado.
Todas las alternativas a vLLM → · Todos los productos de Plataformas e infraestructura →
Fuentes
Este registro se revisó por última vez el 2026-07-30.
Registro legible por máquina: /api/products/vllm.json · ¿Detectaste un error? Sugerir una corrección