vLLM

de vLLM Project · docs.vllm.ai ↗

Plataformas e infraestructura Abierto (pesos/código abierto) Disponibilidad: Global Lanzado en 2023-06

Motor de inferencia y servicio de código abierto para grandes modelos de lenguaje, conocido por su PagedAttention de alto rendimiento.

open-source inference serving llm

Categoría
Plataformas e infraestructura
Modelo de negocio
Abierto (pesos/código abierto)
Disponibilidad
Global
Lanzamiento
2023-06
Registro actualizado
2026-07-30
Licencia
Apache-2.0
Documentación
docs.vllm.ai ↗
Fundación
2023
Sede
US
URL canónica
https://globalaiproductindex.com/es/products/vllm/

Resumen

vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.

Funciones principales

Casos de uso

Precios

vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.

Preguntas frecuentes

What is vLLM used for?

vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.

Is vLLM free?

Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.

What is PagedAttention?

PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.

Productos similares

Todas las alternativas a vLLM → · Todos los productos de Plataformas e infraestructura →

Fuentes

Este registro se revisó por última vez el 2026-07-30.

Registro legible por máquina: /api/products/vllm.json · ¿Detectaste un error? Sugerir una corrección