vLLM
von vLLM Project · docs.vllm.ai ↗
Quelloffene Inferenz- und Serving-Engine für große Sprachmodelle, bekannt für das durchsatzstarke PagedAttention.
open-source inference serving llm
- Kategorie
- Plattformen & Infrastruktur
- Geschäftsmodell
- Offen (offene Gewichte/Quellcode)
- Verfügbarkeit
- Weltweit
- Veröffentlicht
- 2023-06
- Eintrag aktualisiert
- 2026-07-30
- Lizenz
- Apache-2.0
- Dokumentation
- docs.vllm.ai ↗
- Gegründet
- 2023
- Hauptsitz
- US
- Kanonische URL
- https://globalaiproductindex.com/de/products/vllm/
Überblick
vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.
Wichtige Funktionen
- PagedAttention for efficient KV-cache memory management
- High-throughput continuous batching of requests
- OpenAI-compatible API server out of the box
- Broad support for open model architectures and quantization
Anwendungsfälle
- Serving open LLMs in production at high throughput
- Self-hosting an OpenAI-compatible inference endpoint
- Benchmarking and research on efficient LLM inference
Preise
vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.
Häufig gestellte Fragen
What is vLLM used for?
vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.
Is vLLM free?
Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.
What is PagedAttention?
PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.
Ähnliche Produkte
- Amazon Bedrock — Verwalteter AWS-Dienst, um Anwendungen mit Foundation-Modellen mehrerer Anbieter über eine einzige API zu erstellen.
- Amazon Nova — AWS-Portfolio amazon-eigener Foundation-Modelle und Dienste für Text, Multimodalität, Sprache, den Bau eigener Modelle und UI-automatisierende Agenten.
- Azure AI Foundry — Microsofts Plattform zum Erstellen, Bewerten und Bereitstellen von KI-Anwendungen und -Agenten auf Azure.
- Baseten — Plattform zum Bereitstellen und Betreiben von Machine-Learning-Modellen als skalierbare APIs.
- Cerebras — KI-Compute-Unternehmen, dessen Inferenz-Cloud offene Sprachmodelle auf Wafer-Scale-Hardware bereitstellt.
- Civitai — Community-Plattform zum Teilen und Entdecken offener Bildgenerierungsmodelle, LoRAs und generierter Kunst.
Alle vLLM-Alternativen → · Alle Produkte in Plattformen & Infrastruktur →
Quellen
Dieser Eintrag wurde zuletzt am 2026-07-30 geprüft.
Maschinenlesbarer Eintrag: /api/products/vllm.json · Fehler entdeckt? Korrektur vorschlagen