vLLM

par vLLM Project · docs.vllm.ai ↗

Plateformes et infrastructure Ouvert (poids/source ouverts) Disponibilité : Mondial Lancé en 2023-06

Moteur d'inférence et de service open source pour grands modèles de langue, connu pour son PagedAttention à haut débit.

open-source inference serving llm

Catégorie
Plateformes et infrastructure
Modèle économique
Ouvert (poids/source ouverts)
Disponibilité
Mondial
Lancement
2023-06
Enregistrement mis à jour
2026-07-30
Licence
Apache-2.0
Documentation
docs.vllm.ai ↗
Création
2023
Siège
US
URL canonique
https://globalaiproductindex.com/fr/products/vllm/

Aperçu

vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.

Fonctionnalités clés

Cas d'usage

Tarifs

vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.

Questions fréquentes

What is vLLM used for?

vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.

Is vLLM free?

Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.

What is PagedAttention?

PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.

Produits similaires

Toutes les alternatives à vLLM → · Tous les produits Plateformes et infrastructure →

Sources

Cet enregistrement a été vérifié pour la dernière fois le 2026-07-30.

Enregistrement lisible par machine : /api/products/vllm.json · Une erreur ? Suggérer une correction