AI product Open source

llama.cpp

llama.cpp is a C/C++ implementation for running large language and vision-language model inference locally or in the cloud, including on CPUs, GPUs, laptops, Apple silicon devices, and small computers such as Raspberry Pis. The project is built on the ggml library and uses integer quantization from 1.5-bit through 8-bit to reduce memory use and accelerate inference, with CPU, CUDA, HIP, Metal, Vulkan, SYCL, OpenCL, and other hardware backends. It supports hybrid CPU-GPU inference for models larger than available VRAM and includes command-line tools, a built-in web interface, and an OpenAI-compatible REST API server. The project is distributed as source code, Docker images, and pre-built binaries, with installation instructions at llama.app.

View repository Mentioned in 1 video ↓

What llama.cpp is used for

1 use taken from transcripts — each links to the moment in the video.

  • Runs quantized open-weight language models locally on consumer hardware, including CPUs, GPUs, laptops, and Raspberry Pis, for offline assistants, code assistants, RAG, and AI agents.

Videos mentioning llama.cpp

1 in the library.