vLLM

提供: vLLM Project · docs.vllm.ai ↗

プラットフォームとインフラ オープン(オープンウェイト/ソース) 提供状況: グローバル 2023-06 公開

大規模言語モデル向けのオープンソース推論・サービングエンジン。高スループットの PagedAttention で知られる。

open-source inference serving llm

カテゴリ
プラットフォームとインフラ
ビジネスモデル
オープン(オープンウェイト/ソース)
提供状況
グローバル
公開
2023-06
レコード更新
2026-07-30
ライセンス
Apache-2.0
ドキュメント
docs.vllm.ai ↗
設立
2023
本社所在地
US
正規URL
https://globalaiproductindex.com/ja/products/vllm/

概要

vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.

主な機能

ユースケース

料金

vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.

よくある質問

What is vLLM used for?

vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.

Is vLLM free?

Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.

What is PagedAttention?

PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.

類似製品

vLLMの代替製品をすべて見る → · プラットフォームとインフラの製品をすべて見る →

出典

このレコードは 2026-07-30 に最終確認されました。

機械可読レコード: /api/products/vllm.json · 誤りを見つけましたか? 修正を提案