llama.cpp
Recommendedby ggml-org (Georgi Gerganov)
- Open source
- Local-first
- Privacy-first
- Self-hostable
- Pricing:
- Open source
- License:
- MIT
LLM inference engine in C/C++ with minimal dependencies. Runs quantized GGUF models on CPU and GPU across Apple Silicon, x86 and CUDA, and powers many local AI applications.
Categories
Tags
#cpp #gguf #quantization #cpu-inference #llm