vLLM
par vLLM Project · docs.vllm.ai ↗
Moteur d'inférence et de service open source pour grands modèles de langue, connu pour son PagedAttention à haut débit.
open-source inference serving llm
- Catégorie
- Plateformes et infrastructure
- Modèle économique
- Ouvert (poids/source ouverts)
- Disponibilité
- Mondial
- Lancement
- 2023-06
- Enregistrement mis à jour
- 2026-07-30
- Licence
- Apache-2.0
- Documentation
- docs.vllm.ai ↗
- Création
- 2023
- Siège
- US
- URL canonique
- https://globalaiproductindex.com/fr/products/vllm/
Aperçu
vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.
Fonctionnalités clés
- PagedAttention for efficient KV-cache memory management
- High-throughput continuous batching of requests
- OpenAI-compatible API server out of the box
- Broad support for open model architectures and quantization
Cas d'usage
- Serving open LLMs in production at high throughput
- Self-hosting an OpenAI-compatible inference endpoint
- Benchmarking and research on efficient LLM inference
Tarifs
vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.
Questions fréquentes
What is vLLM used for?
vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.
Is vLLM free?
Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.
What is PagedAttention?
PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.
Produits similaires
- Amazon Bedrock — Service géré d'AWS pour créer des applications avec des modèles de fondation de plusieurs fournisseurs via une seule API.
- Amazon Nova — Portefeuille AWS de modèles de fondation et de services conçus par Amazon pour le texte, le multimodal, la parole, la création de modèles personnalisés et les agents d'automatisation d'interface.
- Azure AI Foundry — Plateforme de Microsoft pour créer, évaluer et déployer des applications et agents d'IA sur Azure.
- Baseten — Plateforme pour déployer et servir des modèles d'apprentissage automatique sous forme d'API évolutives.
- Cerebras — Entreprise de calcul IA dont le cloud d'inférence sert des modèles de langage ouverts sur du matériel à l'échelle du wafer.
- Civitai — Plateforme communautaire pour partager et découvrir des modèles ouverts de génération d'images, des LoRA et des créations générées.
Toutes les alternatives à vLLM → · Tous les produits Plateformes et infrastructure →
Sources
Cet enregistrement a été vérifié pour la dernière fois le 2026-07-30.
Enregistrement lisible par machine : /api/products/vllm.json · Une erreur ? Suggérer une correction