vLLM
Recommendedby vLLM Project
- Open source
- Local-first
- Privacy-first
- Self-hostable
- Pricing:
- Open source
- License:
- Apache-2.0
High-throughput, memory-efficient inference and serving engine for LLMs. Uses PagedAttention and continuous batching to serve open models at scale behind an OpenAI-compatible API.
Categories
Tags
#python #gpu #inference-server #openai-compatible