github ggml-org/llama.cpp b7630

latest releases: b7703, b7700, b7699...
6 days ago
Details

CANN: add operator fusion support for ADD + RMS_NORM (#17512)

This commit implements operator fusion for ADD + RMS_NORM operations
in the CANN backend to reduce memory access overhead and improve
performance. The fusion is controlled by the GGML_CANN_OPERATOR_FUSION
environment variable (default: false).

Changes:

  • Implement ggml_cann_op_add_rms_norm_fused() using ACLNN AddRmsNorm
  • Add ggml_cann_can_fuse() to check fusion eligibility
  • Integrate fusion logic into computation graph evaluation
  • Add test cases for ADD + RMS_NORM fusion
  • Update documentation with new environment variable

The fusion combines ADD and RMS_NORM into a single kernel call,
which is more efficient than executing them separately.

macOS/iOS:

Linux:

Windows:

openEuler:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.