github ggml-org/llama.cpp b11388

pre-releaseone hour ago
Details

imatrix: calculate activation-based statistics for new format (GGUF) imatrices (#14891)

  • Use activations to calculate the stats
  • Determine calculation mode
  • Compute entropy for activations
  • Compute cosine similarity based on activations
  • Compute l2 norm
  • Add compute_layer_statistics() function
  • Update aggregated statistic report layout
  • Fix printing l2 norm when calc_mode = 1
  • Refactor variable name
  • Compute aggregated (per layer) l2 norm
  • Update aggregated sum of squared activations per layer
  • Make ZD Score two-tailed
  • Update report layout
  • Reverse conditional logic to match convention
  • Rename report heading
  • Add --activation-statistics parameter
  • Add Euclidean–Cosine Score (ECS)
  • Add --activation-statistics logic to avoid doubling the imatrix size by default
  • Update stats output sort based on imatrix type
  • Process external NextN draft files (-md / --model-draft)
  • Refactor to use new llama_batch_ext

Co-authored-by: compilade git@compilade.net
Co-authored-by: Sigbjørn Skjæret sigbjorn.skjaeret@huggingface.co

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.