mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-10-11 07:20:33 +02:00
b11536
* common : read a GGUF's trained context from its metadata common_get_gguf_n_ctx_train opens only the file's metadata (no_alloc, like common_get_decision_type) and reads <arch>.context_length, so a caller can learn the trained context without loading the model. It accepts both u32 and u64 values and returns 0 when the file is missing, unreadable, invalid, or reports no context length. Assisted-by: pi:zai-org/GLM-5.3-Flash * server : report the trained context in the models listing update_caps already resolves the model file offline to read its modalities, so it now reads the trained context from the same GGUF metadata, and GET /models reports it as context_length when it is known. A router listing then carries the context without any Hub request, which lets the UI sort and filter by it offline. Assisted-by: pi:zai-org/GLM-5.3-Flash * ui : take the trained context from the models listing The router now reports context_length per model, so the option mapping fills contextLength from it and the manager reads it before the Hub record. The Context column, the context sort and the context filter then work with the Hugging Face Hub API turned off. A browser suite guards the sort and the search, the Hub-cache driven context filter and the re-sort when details arrive after the sort was clicked. Assisted-by: pi:zai-org/GLM-5.3-Flash * ui : mark favorite models with a heart A favorited model shows a rose heart in the selector even before its row is hovered, and the crossed heart takes its place on hover, so unfavoriting stays one hover away. The manager table marks its favorited rows with the same heart after the badges and capabilities. Assisted-by: pi:zai-org/GLM-5.3-Flash * fix: UI text nit * fix: UI nits * fix: Favorite models grouping in models table * feat: Remove sorting from Status column in Models Table * server: read the GGUF metadata once per model Read the decision type and the trained context in a single GGUF open, accept only a UINT32 context length like the model loader, and reset n_ctx_train with the other caps so a failed refresh drops it. * fix: Post-review fixes --------- Co-authored-by: Pascal <admin@serveurperso.com>
llama.cpp
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
Quick start
A few options to get llama.cpp installed on your machine:
# curl
curl -LsSf https://llama.app/install.sh | sh
# powershell
irm https://llama.app/install.ps1 | iex
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
Description
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
Supported backends
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
Documentation
Tools
Development
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
Contributing
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
Acknowledgements
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain
Languages
C++
55.8%
C
15.8%
Python
7.3%
Cuda
5.3%
TypeScript
4.3%
Other
11.3%