github ggml-org/llama.cpp b10920

latest release: b10921
pre-releaseone hour ago
Details

hexagon: support for multi-device model split (aka row-split) (#28589)

  • hex-row-split: add support for multi-device row spliting

Co-authored-by: Max Krasnyansky maxk@qti.qualcomm.com

  • hex-mdev: add work splitting to fused kernels

  • hex-mdev: use mdev_ prefix for all multi-device state

  • hex-mdev: make device configuration more expressive to support device groups

  • hex-mdev: fix mdev session init

  • hex-mdev: fused nx (2x,3x) matmuls must update row counts for each w/o

  • hex-mdev: fix MUL_MAT work partitioning bugs introduced by mdev

  • hex-cont: fix crashes with new tests due to wrong striding

  • hex-mdev: move fences after l2flushes

  • hex-cont: fix work splitting for mnpu -- align chunks to cachelines

  • hex-mdev: fix CPY tests with multi-dev

  • hex-mmid: fix work partitioning with mnpu

  • hex-mm: fix test failures with mdev

  • hex-binary: fix work partitioning for mdev

  • hex-argsort: fix mdev partitioning

  • hex-mdev: fix work partitioning and general updates for all simple ops

  • hex-fa: fix mdev work splitting issues

  • hex-mdev: fixing more failing ops test

  • hex-mdev: update the rest of the ops

  • hex-mdev: refactor all mdev splitting logic to be contained within if (mdev_count > 1) {...}

  • hex-mdev: fix macros

  • hex-mdev: simplify session flush logic

  • hex-sync: fix recursion in session flush

  • hex-mdev: factor out fence buffer and allocator

  • hex-fence: make fence allocation more robust with reserved slots for mdev

  • hex-mdev: keep all mdev state in htp_mdev_group

  • hex-mdev: further cleanup mdev group handling at the host

  • hex-mdev: update group idx in the opbatch before serializing

  • hex-batch: remove separate op_pending and use batch_req/rsp_seq

  • hex-async: workaround another missing tensor_init in ggml-meta

  • hex-fence: cleanup and robustify fences and error handling in multi-device scenarios

  • hex-ar: improve ALLREDUCE error handling

  • hex-async: robust error handling for op_cpy_fence

  • hex-async: use seq0 from allreduce context to allocate fence_seq

  • hex-mdev: fix remaining issues with fence and barrier clearing in CPY_FENCE

  • hex-misc: realign macros and fix misplaces trace events

  • hex-misc: align macros

  • hex-mdev: fix unclone buffer re-entrancy

  • hex-glu: fix mdev partitioning logic

  • hex-mdev: make buffer uncloning/cleanup work with tensor-split scenarios

  • hex-mdev: tighten up the can_split check in act-ops

  • hex-mdev: factor out common bits of the partitioning logic

  • hex-mm: minor realignment of the macros

  • hex-bufs: fix incorrectly placed assert for MAX_BUFS

  • hex-pad: tighten up gating checks for PAD

  • hex-kparams: make sure all kernels properly use kparams->n_threads

  • hex-docs: update user and developer docs with new features and detailed guide for ops development

  • hex-scripts: update run script to properly parse dev groups

  • hex-misc: formatting

  • hex-sess: minor cleanup for session init

  • hex-ar: fix vtcm size calc in allreduce kparams

  • hex-scripts: fix flake8 warnings

  • hex-rope: update ROPE to support mdev work split

  • hex-ops: remove redunant checks and minor reformat

  • hex-dev-guide: update dev-guide to avoid redundant null checks

  • hex-async: improve event_wait, event_sync and fence implementations

  • hex-async: remove synchronous flush from event_sync

  • hex-async: symplify fence recovery protocol and make sync more robust

  • hex-async: futher simplify error recovery for fences

  • hex-err: return status instead of just -1

  • hex-async: print all seq nums in hex

  • hex-async: make sure fences flush dirty ranges

  • hex-async: add dirty ranges merging to reduce fence flushes

  • hex-async: properly sync before freeing the event

  • hex-async: make sure fence owner session is not overriden

  • hex-async: more fence write order more robust

  • hex-async: make sure not to fuse ALLREDUCE+ADD if their dsts overlap

  • hex-fusion: cleanup redundant checks


Co-authored-by: Alexander Lu alexlu@qti.qualcomm.com

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI:

Don't miss a new llama.cpp release

NewReleases is sending notifications on new releases.