mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-10-07 13:30:37 +02:00
master
* llama : add GLM5-Next NextN (MTP) graph Build the GLM5-Next multi-token-prediction head as graph_mtp: the NextN block embeds enorm(tok)+hnorm(h) through eh_proj, runs one plain DSA layer and the shared lm_head, reusing the trunk's builders through the no_build tag ctor. llama_memory_recurrent also tolerates a partial seq_rm when the context holds no recurrent layers, which is what the MTP draft context needs. Assisted-by: Claude * llama : glm5-next: skip dead compute in headless NextN forwards A NextN forward with no output rows (the MTP catch-up and the draft-context prefill) persists only through its cache writes, so the headless graph keeps the MLA latent, indexer key|gate and pooled-key writes and drops the query path, the indexer selection, the attention body, the FFN and the LM head. The 4-token catch-up falls from 6.9 ms to 0.33 ms of kernels; the greedy output hashes and the draft acceptance are unchanged. Assisted-by: Claude * llama : glm5-next: fix NextN extraction contracts and shared-tail rollback Three fixes from the architectural review. The headless graph prune now also requires that no unmasked nextn extraction is live, because that mode reads n_tokens hidden rows regardless of the logits flags. Masked extraction publishes the hidden rows gathered by the output ids, so a batch whose output flags are not a prefix exports the right rows. A partial recurrent rollback whose tail cell is shared with another sequence is now rejected instead of silently moving that sequence's tail. Assisted-by: Claude * llama : glm5-next: tidy comments in the MTP changes Assisted-by: Claude * llama : glm5-next: crop the MTP graph to the output rows instead of pruning it Replace the headless NextN prune with the crop pattern the other MTP graphs use: gather the attention output and the block input at the output ids before the position-wise FFN and the shared head. A NextN forward with no output rows (the MTP catch-up and the draft-context prefill) then runs the FFN and the head over zero rows. The 4-token catch-up falls from 6.9 ms to 2.9 ms of kernels; greedy output hashes are unchanged. Assisted-by: Claude * glm5-next: use the nextn crop helpers in the MTP graph Replace the local crop condition and the masked select of t_h_nextn with crop_before_nextn and crop_after_nextn, so the MTP graph narrows its rows the same way as the main graph and the other models. Describe the shared cell and empty filter branches of the recurrent partial rollback. * glm5-next: load MTP-only and trunk-only GGUF files Make the trunk tensors optional when the file only holds the NextN layer, and the NextN tensors optional when the file only holds the trunk, so the split MTP GGUF loads as a draft model. Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Pascal <admin@serveurperso.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (#29852)
llama.cpp
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
Quick start
A few options to get llama.cpp installed on your machine:
# curl
curl -LsSf https://llama.app/install.sh | sh
# powershell
irm https://llama.app/install.ps1 | iex
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
Description
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
Supported backends
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
Documentation
Tools
Development
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
Contributing
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
Acknowledgements
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain
Languages
C++
55.9%
C
16%
Python
7.3%
Cuda
5.3%
TypeScript
4.2%
Other
11.1%