Inference servers and platforms for deploying models at scale.
1 tool
vLLM Project
High-throughput, memory-efficient inference and serving engine for LLMs. Uses PagedAttention and continuous batching to serve open models at scale behind an OpenAI-compatible API.