Details
CUDA: Do not mutate cgraph for fused ADDs (#19566)
- Do not mutate cgraph for fused ADDs
- We should try to minimize in-place changes to the incoming
ggml_cgraph where possible (those should happen in graph_optimize) - Modifying in-place leads to an additional, unnecessary graph capture
step as we store the properties before modifying the graph in-place
in the cuda-backend
-
Assert ggml_tensor is trivially copyable
-
Update ggml/src/ggml-cuda/ggml-cuda.cu
Co-authored-by: Aman Gupta amangupta052@gmail.com
Co-authored-by: Aman Gupta amangupta052@gmail.com
macOS/iOS:
Linux:
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows x64 (CUDA 12) - CUDA 12.4 DLLs
- Windows x64 (CUDA 13) - CUDA 13.1 DLLs
- Windows x64 (Vulkan)
- Windows x64 (SYCL)
- Windows x64 (HIP)
openEuler: