vLLM

제공: vLLM Project · docs.vllm.ai ↗

플랫폼 및 인프라 오픈(오픈 가중치/소스) 이용 가능성: 글로벌 출시 2023-06

대규모 언어 모델을 위한 오픈 소스 추론·서빙 엔진으로, 높은 처리량의 PagedAttention으로 알려져 있다.

open-source inference serving llm

카테고리
플랫폼 및 인프라
비즈니스 모델
오픈(오픈 가중치/소스)
이용 가능성
글로벌
출시
2023-06
레코드 업데이트
2026-07-30
라이선스
Apache-2.0
문서
docs.vllm.ai ↗
설립
2023
본사
US
표준 URL
https://globalaiproductindex.com/ko/products/vllm/

개요

vLLM is an open-source inference and serving engine for large language models, originally developed at UC Berkeley. Its PagedAttention memory management delivers high throughput and efficient GPU utilization, and it exposes an OpenAI-compatible API server. vLLM is free under Apache-2.0 and has become a de facto standard for self-hosted LLM serving.

주요 기능

활용 사례

가격

vLLM is free and open source under Apache-2.0; costs are limited to the GPU infrastructure it runs on.

자주 묻는 질문

What is vLLM used for?

vLLM serves large language models with high throughput, exposing an OpenAI-compatible API for self-hosted deployments.

Is vLLM free?

Yes. vLLM is open source under Apache-2.0; you only pay for the hardware you run it on.

What is PagedAttention?

PagedAttention is vLLM's memory-management technique that pages the KV cache, raising GPU utilization and throughput.

유사 제품

vLLM 대안 모두 보기 → · 플랫폼 및 인프라 제품 모두 보기 →

출처

이 레코드는 2026-07-30에 마지막으로 검토되었습니다.

기계 판독용 레코드: /api/products/vllm.json · 오류를 발견하셨나요? 수정 제안