Ed Addario
ebe18bee5a
vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON ( #29912 )
2026-10-05 10:18:55 +03:00
0504396140
imatrix: calculate activation-based statistics for new format (GGUF) imatrices ( #14891 )
...
* Use activations to calculate the stats
* Determine calculation mode
* Compute entropy for activations
* Compute cosine similarity based on activations
* Compute l2 norm
* Add compute_layer_statistics() function
* Update aggregated statistic report layout
* Fix printing l2 norm when calc_mode = 1
* Refactor variable name
* Compute aggregated (per layer) l2 norm
* Update aggregated sum of squared activations per layer
* Make ZD Score two-tailed
* Update report layout
* Reverse conditional logic to match convention
* Rename report heading
* Add --activation-statistics parameter
* Add Euclidean–Cosine Score (ECS)
* Add --activation-statistics logic to avoid doubling the imatrix size by default
* Update stats output sort based on imatrix type
* Process external NextN draft files (-md / --model-draft)
* Refactor to use new llama_batch_ext
Co-authored-by: compilade <git@compilade.net >
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co >
2026-10-04 11:24:53 +02:00
Ed Addario
f3a33dff26
rpc : fix linking when compiling with BUILD_SHARED_LIBS=OFF ( #28492 )
2026-09-12 09:22:57 +03:00
Ed Addario
0596704284
quant : Optimise memory usage by evicting weights after processing each layer ( #22877 )
...
* Evict weights from memory after processing each layer
* Revert changes
* Move unmap to libllama
* Unmap weights offloaded to backend
* Change member's constness
* Remove unmap weights offloaded to backend
2026-08-18 16:22:32 +02:00
Ed Addario and Georgi Gerganov
4951250235
llama : refactor llama_model_quantize_params to expose a pure C interface ( #20346 )
...
* Refactor llama_model_quantize_params to expose a pure C interface
* Restore comment and cleanup struct def
* Code review refactoring
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
* Code review refactoring
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2026-04-01 08:43:00 +03:00
Ed Addario
daf2dd7880
quantize : skip tensor override when in fallback mode ( #14995 )
2025-07-31 21:32:18 +02:00
Ed Addario
e9192bec56
quantize : fix using combined imatrix GGUFs (multiple datasets) ( #14973 )
2025-07-30 21:11:56 +02:00
Ed Addario and Sigbjørn Skjæret
7f97599581
quantize : update README.md ( #14905 )
...
* Update README.md
* Fix trailing whitespace
* Update README.md
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2025-07-27 23:31:11 +02:00
Ed Addario and compilade
d1aa0cc5d1
imatrix: add option to display importance score statistics for a given imatrix file ( #12718 )
...
* Add --show-statistics option
* Add --show-statistics logic
* Add tensor name parsing
* Tidy output format
* Fix typo in title
* Improve tensor influence ranking
* Add better statistics
* Change statistics' sort order
* Add Cosine Similarity
* Add header search path
* Change header search path to private
* Add weighted statistics per layer
* Update report title
* Refactor compute_statistics out of main
* Refactor compute_cossim out of load_imatrix
* Refactor compute_statistics out of load_imatrix
* Move imatrix statistics calculation into its own functions
* Add checks and validations
* Remove unnecessary include directory
* Rename labels
* Add m_stats getter and refactor compute_statistics out of load_imatrix
* Refactor variable names
* Minor cosmetic change
* Retrigger checks (empty commit)
* Rerun checks (empty commit)
* Fix unnecessary type promotion
Co-authored-by: compilade <git@compilade.net >
* Reverting change to improve code readability
* Rerun checks (empty commit)
* Rerun checks (empty commit)
* Rerun checks - third time's the Charm 🤞 (empty commit)
* Minor cosmetic change
* Update README
* Fix typo
* Update README
* Rerun checks (empty commit)
* Re-implement changes on top of #9400
* Update README.md
* Update README
* Update README.md
Co-authored-by: compilade <git@compilade.net >
* Update README.md
Co-authored-by: compilade <git@compilade.net >
* Update README.md
* Remove duplicate option in print_usage()
* Update README.md
* Update README.md
Co-authored-by: compilade <git@compilade.net >
* Update README.md
Co-authored-by: compilade <git@compilade.net >
* Remove input check
* Remove commented out code
---------
Co-authored-by: compilade <git@compilade.net >
2025-07-22 14:33:37 +02:00
Ed Addario
c81f4192f9
gguf-py : dump bpw per layer and model in markdown mode ( #14703 )
2025-07-16 00:04:42 +02:00
Ed Addario
982e347255
quantize : fix minor logic flaw in --tensor-type ( #14572 )
2025-07-13 18:02:17 +02:00
Ed Addario
fa4a9f2a1c
quantize : handle user-defined pruning of whole layers (blocks) ( #13037 )
2025-06-22 23:16:26 +02:00
Ed Addario
30e5b01de2
quantize : change int to unsigned int for KV overrides ( #14197 )
2025-06-15 18:53:45 +02:00
Ed Addario
e5c834f718
quantize : improve tensor-type pattern matching ( #13033 )
2025-05-13 19:12:31 +02:00
Ed Addario
71e90e8813
quantize: Handle user-defined quantization levels for additional tensors ( #12511 )
...
* Add llama_model_quantize_params parameters
* Add new quantize parameters parsing and validation
* Update usage
* Add new parameters defaults
* Add new quantization parameters logic
* Add llama_model_quantize_params parameters
* Add new quantize parameters parsing and validation
* Update usage
* Add new parameters defaults
* Add new quantization parameters logic
* Minor refactoring as per the contributors' coding guidelines
* Update descriptions to match existing style
* Add llama_model_quantize_params parameters
* Add new quantize parameters parsing and validation
* Update usage
* Add new parameters defaults
* Add new quantization parameters logic
* Minor refactoring as per the contributors' guidelines
* Implement general --tensor-type instead of tensor-specific command option
* Fix implied type bug
* Restore missing #includes
* Add regex capability for tensor selection
* Refactor function name and update ALLOWED_TENSOR_TYPE
* Add missing #include
* Handle edge case when tensor name is cls.output
* Minor logging improvement
2025-04-13 21:29:28 +03:00