vLLM

von vLLM Project · docs.vllm.ai ↗

Plattformen & Infrastruktur Offen (offene Gewichte/Quellcode) Verfügbarkeit: Weltweit Veröffentlicht 2023-06

Quelloffene Inferenz- und Serving-Engine für große Sprachmodelle, bekannt für das durchsatzstarke PagedAttention.

open-source inference serving llm

Kategorie
Plattformen & Infrastruktur
Geschäftsmodell
Offen (offene Gewichte/Quellcode)
Verfügbarkeit
Weltweit
Veröffentlicht
2023-06
Eintrag aktualisiert
2026-07-30
Lizenz
Apache-2.0
Dokumentation
docs.vllm.ai ↗
Gegründet
2023
Hauptsitz
US
Kanonische URL
https://globalaiproductindex.com/de/products/vllm/

Überblick

vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.

Wichtige Funktionen

Anwendungsfälle

Preise

vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.

Häufig gestellte Fragen

What is vLLM used for?

vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.

Is vLLM free?

Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.

What is PagedAttention?

PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.

Ähnliche Produkte

Alle vLLM-Alternativen → · Alle Produkte in Plattformen & Infrastruktur →

Quellen

Dieser Eintrag wurde zuletzt am 2026-07-30 geprüft.

Maschinenlesbarer Eintrag: /api/products/vllm.json · Fehler entdeckt? Korrektur vorschlagen