CCCL 3.5 Release
The CCCL team is excited to announce the 3.5 release of the CUDA Core Compute Libraries (CCCL). Highlights include public tuning APIs for CUB, reproducible floating-point scans, batched Top-K selection, and reductions that accept problem sizes stored on the GPU.
Public CUB Tuning APIs
CCCL 3.5 exposes public tuning policies for CUB device-wide algorithms. Applications can customize parameters such as threads per block, items per thread, vectorization, and algorithm selection by passing a policy selector through cuda::execution::tune(...) in an execution environment.
Policy selectors choose tuning parameters for a GPU compute capability and return an algorithm-specific policy such as cub::ReducePolicy or cub::ScanPolicy. This provides a supported way to specialize CUB for a workload through the public device-wide APIs.
For example, this example policy below uses 256 threads per block and selects the number of items per thread based on compute capability:
struct ReduceTuning {
__host__ __device__ constexpr cub::ReducePolicy
operator()(cuda::compute_capability cc) const {
auto pass = cub::ReducePassPolicy{
.threads_per_block = 256,
.items_per_thread = cc >= cuda::compute_capability{10, 0} ? 20 : 16,
.vec_size = 4,
.reduce_algorithm = cub::BLOCK_REDUCE_WARP_REDUCTIONS,
.load_modifier = cub::LOAD_LDG
};
return {.multi_tile = pass, .single_tile = pass};
}
};
auto env = cuda::std::execution::env{
cuda::stream_ref{stream}, cuda::execution::tune(ReduceTuning{})
};
auto status = cub::DeviceReduce::Sum(input.begin(), output.begin(), input.size(), env);See the CUB tuning documentation for an example and the CUB environment documentation for more information on how to use environments. Each CUB device-wide algorithm also has a documentation example for how to customize its tuning, see for example how to tune cub::DeviceReduce.
Reproducible Floating Point Scans
cub::DeviceScan now supports run-to-run reproducibility for floating-point summation. Requesting cuda::execution::determinism::run_to_run selects a fixed reduction order, producing repeatable results on the same GPU with the same input and build configuration.
This applies to inclusive and exclusive sums and scans using cuda::std::plus.
See the cub::DeviceScan documentation.
auto env = cuda::std::execution::env{
cuda::stream_ref{stream},
cuda::execution::require(cuda::execution::determinism::run_to_run)
};
auto status = cub::DeviceScan::InclusiveSum(input.begin(), output.begin(), input.size(), env);Batched Top-K Selection
CCCL 3.5 adds cub::DeviceBatchedTopK, which selects the smallest or largest K items independently from many segments. MinKeys, MaxKeys, MinPairs, and MaxPairs support fixed or variable segment sizes and a K value that can vary by segment.
In CCCL 3.5, each segment is processed in one thread block, with a 48 KiB shared-memory limit. With the default policies, the maximum is 8,192 float keys or 4,096 double keys per segment; limits for key-value pairs depend on both types.
Callers must provide a compile-time upper bound on segment size (for example, cuda::args::bounds<1, 8192>() for float keys) and explicitly request non-deterministic, unordered output using determinism::not_guaranteed, tie_break::unspecified, and output_ordering::unsorted through cuda::execution::require(...).
Select the two largest keys per four-element segment from existing device buffers:
constexpr int segment_size = 4, k = 2;
auto segments_in = cuda::make_strided_iterator(
cuda::make_counting_iterator(input.begin()), segment_size);
auto segments_out = cuda::make_strided_iterator(
cuda::make_counting_iterator(output.begin()), k);
auto env = cuda::std::execution::env{
cuda::stream_ref{stream},
cuda::execution::require(
cuda::execution::determinism::not_guaranteed,
cuda::execution::tie_break::unspecified,
cuda::execution::output_ordering::unsorted)
};
auto status = cub::DeviceBatchedTopK::MaxKeys(
segments_in, segments_out,
cuda::args::constant<segment_size>{}, cuda::args::constant<k>{}, num_segments, env);See the CCCL 3.5 cub::DeviceBatchedTopK API for argument annotations, supported types, and complete examples.
CUB Arguments Stored on the GPU
The new <cuda/argument> header provides cuda::args::constant, immediate, deferred, and deferred_sequence, together with argument bounds. These annotations describe values known at compile time, values supplied by the host, and values read from device-accessible memory in stream order.
cub::DeviceReduce::{Reduce, Sum, Min, Max, TransformReduce} now accept a problem size supplied through cuda::args::deferred. A preceding kernel can produce the element count without a round trip to the host. Captured reductions can consume a different count on each CUDA Graph replay without updating or recapturing the graph.
Selection writes its output count to a one-element device buffer. The reduction consumes that count directly on the same stream:
auto selected = cuda::make_device_buffer<float>(stream, device, input.size(), cuda::no_init);
auto count = cuda::make_device_buffer<int>(stream, device, 1, cuda::no_init);
auto sum = cuda::make_device_buffer<float>(stream, device, 1, cuda::no_init);
auto env = cuda::std::execution::env{cuda::stream_ref{stream}};
auto status = cub::DeviceSelect::If(
input.begin(), selected.begin(), count.begin(), input.size(), predicate, env);
status = cub::DeviceReduce::Sum(
selected.begin(), sum.begin(), cuda::args::deferred{count.begin()}, env);cub::DeviceScan::InclusiveScan also accepts an initial value supplied through cuda::args::deferred.
See the DeviceReduce documentation, the InclusiveScan addition, and #8789: support for device-resident problem sizes.
Faster Searches for Sorted Query Values
cub::DeviceFind::LowerBoundSortedValues and UpperBoundSortedValues accelerate batched lower- and upper-bound searches when both the searched range and the query values are sorted under the same comparator. They use a merge-path traversal with O(N + M) total work, where N is the range length and M is the number of queries.
See the cub::DeviceFind documentation.
Additional Library Improvements
- Added
cuda::std::ranges::zip_viewandcuda::std::views::zip, with expanded tuple-like construction and assignment support incuda::std::tupleandcuda::std::pair. - Added
bit_compress,bit_expand,bit_repeat, andbit_reverse, plusshlandshr, to<cuda/std/bit>. - Added
cuda::isclosefor approximate comparisons,cuda::ceil_ilog10, andcuda::device::warp_match_anyfor matching trivially copyable values across warp lanes. - Added
atomic_ref::address()to bothcuda::atomic_refandcuda::std::atomic_ref, andcuda::std::is_virtual_base_ofwhere supported by the compiler. - Added initializer-list overloads for buffer factories, additional memory-pool attribute queries with CUDA 13.3, and updated PTX wrappers for CUDA 13.4.
- Expanded CUB's use of hardware warp-reduction instructions, including floating-point min/max on supported Blackwell targets; added Programmatic Dependent Launch support to
DeviceRadixSortand vectorizedBlockLoad/BlockStoreoperations for contiguous iterators. - Added GDB and LLDB pretty-printers for
cuda::buffer,cuda::std::array, andcuda::std::complexandcuda::complex. - C++23 extended
cuda::std::tupleinterface for tuple like types - Various compile time improvements for type traits
- Added
cuda::std::expected::has_error() - Public macros to detect the host architecture
CCCL_HOST_ARCH
Notable Fixes
- Fixed a race condition in
DeviceTopK. - Fixed out-of-bounds writes in
DeviceHistogramfor large inputs and in streaming reduce-by-key and run-length encoding for single-partition inputs. - Fixed Thrust OpenMP scans processing unused block sums, and improved
cuda::device_bufferiterator compatibility in Thrust andDeviceAdjacentDifference. - Fixed
cuda::std::aligned_allocargument handling,cuda::bufferinitialization for three-byte element types, andto_charswidth calculation for exact powers of ten.
Deprecations and Migration Notes
cub::ChainedPolicy, the CUB dispatch structs and the agent policy types listed below are deprecated. Custom tuning should use policy selectors through the execution environment of the corresponding public Device* API. Both single-call environment APIs and traditional two-phase temporary-storage APIs remain supported. See the tuning migration examples.
All type names in the table are in namespace cub.
| Deprecated customization types | Public tuning policy |
|---|---|
DispatchAdjacentDifference, AgentAdjacentDifferencePolicy
| AdjacentDifferencePolicy
|
AgentBatchMemcpyPolicy
| BatchedCopyPolicy
|
DispatchHistogram, AgentHistogramPolicy
| HistogramPolicy
|
DispatchMergeSort, AgentMergeSortPolicy
| MergeSortPolicy
|
DispatchRadixSort, AgentRadixSortDownsweepPolicy, AgentRadixSortUpsweepPolicy, AgentRadixSortHistogramPolicy, AgentRadixSortExclusiveSumPolicy, AgentRadixSortOnesweepPolicy
| RadixSortPolicy
|
DispatchReduce, DispatchTransformReduce, AgentReducePolicy
| ReducePolicy
|
DispatchReduceByKey, AgentReduceByKeyPolicy
| ReduceByKeyPolicy
|
DeviceRleDispatch, AgentRlePolicy
| RleEncodePolicy or RleNonTrivialRunsPolicy, depending on the operation
|
DispatchScan, AgentScanPolicy
| ScanPolicy
|
DispatchScanByKey, AgentScanByKeyPolicy
| ScanByKeyPolicy
|
DispatchSegmentedRadixSort
| SegmentedRadixSortPolicy
|
DispatchSegmentedReduce, AgentWarpReducePolicy
| SegmentedReducePolicy
|
DispatchSegmentedSort, AgentSubWarpMergeSortPolicy
| SegmentedSortPolicy
|
DispatchSelectIf, AgentSelectIfPolicy
| SelectPolicy or PartitionPolicy, depending on the operation
|
DispatchThreeWayPartitionIf, AgentThreeWayPartitionPolicy
| ThreeWayPartitionPolicy
|
DispatchUniqueByKey, AgentUniqueByKeyPolicy
| UniqueByKeyPolicy
|
cub::DeviceCount,DeviceCountUncached, andDeviceCountCachedValueare deprecated. Usecuda::devices.size().cuda::std::inplace_vector::try_push_backandtry_emplace_backnow returncuda::std::optional<T&>instead ofT*. Update code that stores the result in a pointer or compares it withnullptr.- CUB's umbrella and unsupported device-wide headers now issue explicit diagnostics under NVRTC. Include the specific block-, warp-, or thread-level headers used by runtime-compiled kernels.
Full changelog from v3.4.0 to the 3.5 release branch
What's Changed
🚀 Thrust / CUB
📚 Libcudacxx
- [libcu++][doc] Fix Broken
cuda/functionaldocumentation by @fbusato in #9073 cuda::std::simdComplex by @fbusato in #8475libcudacxx-testSKILL by @fbusato in #9100- Fix
mdspanlayout_strideABI failure in MSVC by @fbusato in #8954 cuda::std::simdload and store functionalities by @fbusato in #8252- Update
libcudacxx-styleSKILL by @fbusato in #9115 std::simdpermute by @fbusato in #8508cuda::std::simdBit by @fbusato in #8704cuda::std::simdCreation by @fbusato in #8653cuda::std::simdAlways use_CCCL_HOST_DEVICE_APIby @fbusato in #9191- Fix/Improve
warp_match_allby @fbusato in #9192 cuda::std::simdAlgorithms by @fbusato in #8659- CCCL works with old DLPack versions by @fbusato in #9346
- Optimize
cuda::std::rotl/rotrby @fbusato in #9352 cuda::device::warp_match_anyby @fbusato in #9243- Introduce public
CCCL_HOST_ARCHby @fbusato in #9494 - Support
__float128forisfinite()andisinf()by @fbusato in #9576 - Don't use
__in,__out,__inoutvariable names by @fbusato in #9599 - Use precise header in
bit_castby @fbusato in #9600 cuda::std::simdMemory permute by @fbusato in #8539- Add
cuda::ceil_ilog10by @fbusato in #9613 cuda::std::simdMath by @fbusato in #8740- fix
shared_memory_mdspandocumentation version by @fbusato in #9734 - Workaround for
cuda::std::simdNVRTC c++17 failure with some math functions by @fbusato in #9754 - Update
__cccl_ptx_isafor CTK 13.4 by @fbusato in #9791 cuda::std::simdOptimize small integer operations by @fbusato in #8873cuda::isclose()by @fbusato in #9577cuda::std::simdF32x2 cleanup/refactoring by @fbusato in #8951- Fix
cuda::ptxshrandbfindlong,long longhandling by @fbusato in #10000 cuda::std::simd: Addshr,shl,bit_reverseby @fbusato in #10058cuda::std::simdOptimize Min/Max by @fbusato in #8949
🔄 Other Changes
- Use signed offset type for DevicePartition by @bernhardmgruber in #8971
- cudax: introduce HLL Policy template parameter by @sleeepyjack in #8857
- [libcu++] Turn memory resource properties docs into a table by @pciolkosz in #8973
- Skip header testing for the NVTX3 header. by @wmaxey in #8986
- docs: fix duplicated words in host_stub_visibility and permutation_iterator comments by @vip892766gma in #8980
- Enable clang-diagnositc clang-tidy checks by @Jacobfaib in #8939
- [libcu++] Fix
__detectably_invalidvalue returned during constant evaluation by @davebayer in #8987 - Use the new tuning API for
detail::radix_sort::dispatchby @bernhardmgruber in #7949 - [c2h] Make catch macros work on device by @davebayer in #8928
- Fix nightly GCC8 build regressions in bit_cast and DeviceReduce by @alliepiper in #8924
- [infra] Add clang-cuda-21 in C++23 job for libcu++ by @davebayer in #8988
- [Tile] avoid returns in switch statements in logarithm by @miscco in #9001
- [Tile] Avoid returns inside loops in
mdspanby @miscco in #9004 - [Tile] avoid return in switch statement in
assume_alignedby @miscco in #9002 - [Tile] Avoid returning in a loop in
envby @miscco in #9007 - [Tile] Avoid return in loop in
charconvby @miscco in #9010 - [Tile] avoid use of CPO in variant by @miscco in #9012
- [Tile] Mark
simdas_CCCL_HOST_DEVICEby @miscco in #9013 - [Tile] Avoid return in loop in
variantby @miscco in #9011 - [Tile] Mark threading support as
_CCCL_HOST_DEVICE_APIby @miscco in #9006 - [Tile] Avoid unused variable warning by @miscco in #9009
- [Tile] Avoid returns in loops in various algorithms by @miscco in #9014
- [libcu++] Fix device fp128 functions test by @davebayer in #9015
- Revert "[libcu++] Fix concurent return value writes during device testing (#8957)" by @miscco in #9003
- [Tile] Make
arch_id_CCCL_HOST_DEVICE_APIby @miscco in #9005 - [Tile] Avoid returns in loops in string functions by @miscco in #9008
- [libcu++] avoid warning about pointless comparison by @miscco in #9016
- [Tile] Avoid unused variable warnings by @miscco in #9017
- [Tile] Mark some algorithms as
_CCCL_HOST_DEVICE_APIby @miscco in #9018 - Fix mdspan support in modernized CUB interfaces by @pauleonix in #9020
- [Tile] Avoid use of
in_placeglobal by @miscco in #9024 - [Tile] Avoid returns in loops in some algorithms by @miscco in #9025
- [Tile] Avoid return statements in loops in
bitsetby @miscco in #9027 - [Tile] Do not access global in tile mode by @miscco in #9026
- [Tile] Avoid internal usage of CPOs by @miscco in #9023
- Use the new tuning API internally for
detail::segmented_sort::dispatchby @bernhardmgruber in #8992 - [libcu++] Cleanup our lit config by @miscco in #8917
- [Tile] Avoid calling a device only function in tile mode by @miscco in #9031
- [Tile] Mark format tests as unsupported by @miscco in #9033
- [Tile] Avoid return in switch statements in
variantby @miscco in #9032 - Use also a SM120 GPU in light CI for CCCL.C by @bernhardmgruber in #8972
- Drop override accum_t from reduce by key by @bernhardmgruber in #8993
- [Tile] Try to appease tile compiler by @miscco in #9035
- [Tile] enable some more tests by @miscco in #9034
- [libcu++] Fix default make_shared_resource construction by @bdice in #9044
- [infra] Update nvhpc to 26.3 by @davebayer in #9048
- Fix segmented radix sort benchmark segment size type by @bernhardmgruber in #9039
- Use the new tuning API internally for
detail::segmented_radix_sort::dispatchby @bernhardmgruber in #8927 - Use the new tuning API internally for
detail::select::dispatchandDeviceSelectby @bernhardmgruber in #8880 - Use
U32inby_keybenchmark by @bernhardmgruber in #9051 - Drop obsolete function by @bernhardmgruber in #9050
- Use the new tuning API internally for
detail::select|three_way_partition::dispatchandDevicePartitionby @bernhardmgruber in #8925 - Add work planning how-to by @jrhemstad in #9042
- Use the new tuning API internally for
detail::reduce_by_key::dispatchby @bernhardmgruber in #8756 - Improve argument names in cub/util_math by @pauleonix in #9053
- Complex tanh (and thus tan) accuracy refinement by @s-oboyle in #9036
- Vectorize contiguous iterators in
cub::BlockLoad/Storeby @bernhardmgruber in #9056 - [STF] Add per-handle exec_place stream resources by @caugonnet in #8905
- [cub] Replace
__NVCOMPILER_CUDA_ARCH__withNV_TARGET_MINIMUM_SM_INTEGERby @davebayer in #9067 - Add
__lazy_call_orby @Jacobfaib in #8767 - [libcu++] Always suppress C++ extensions warnings in prologue by @davebayer in #9019
- [libcu++] Explicitly mark
__halfand__nv_bfloat16as trivially_copyable by @miscco in #9069 - [STF] Move unstable_unique from STF to generic cudax utility by @caugonnet in #8190
- [cudax] Disable cuco hashers
__int128test for nvcc 12.0 + gcc combination by @davebayer in #9076 - [cudax] Remove
CUDAX_MEOWtest macros by @davebayer in #8998 - Env passthrough 4/4 by @gonidelis in #9065
- [cub] Replace
assertwithCHECKorREQUIREin tests by @davebayer in #8999 - Add env DeviceTopK without temp_storage args by @gonidelis in #8982
- Final env-passthrough 2/4 by @gonidelis in #8979
- Use the new tuning API internally for
detail::scan_by_key::dispatchby @bernhardmgruber in #8761 - [STF] Make graph_task and stream_task move-only by @caugonnet in #8915
- [STF] Fix graph capture integration by @caugonnet in #8914
- Use public API for deterministic reduction test by @bernhardmgruber in #9087
- Align offset type in unique_by_key to public API by @bernhardmgruber in #9091
- [Tile] improve documentation by @miscco in #9054
- Small segmented radix sort cleanup by @bernhardmgruber in #9092
- Use public API in unique_by_key benchmark by @bernhardmgruber in #9089
- New complex acos function. by @s-oboyle in #9096
- Use the new tuning API internally for
detail::batch_memcpy::dispatchby @bernhardmgruber in #9093 - Final env-passthrough 3/4 by @gonidelis in #8981
- STF: extend cuda_try output-parameter inference (first/last + ambiguity rejection) by @andralex in #8891
- [Tile] Partially support
__halfand__nv_bfloat16in tile mode by @miscco in #9084 - [Tile] Improve table formatting and merge tables by @miscco in #9104
- [Tile] Improve support for
complex,__halfand__nv_bfloat16by @miscco in #9072 - Use the new tuning API internally for
detail::histogram::dispatchby @bernhardmgruber in #9106 - [libcu++] Do not require
default_initializablefor bit_cast by @miscco in #9105 - [CUB] Fix invalid condition to exclude 128 bit integers by @miscco in #9112
- Fix RAPIDS in CI by @trxcllnt in #9103
- Update
.gitignoreto excludecodegraphfiles by @fbusato in #9114 - Remove dead code by @bernhardmgruber in #9122
- [libcu++] Fix redefinition errors in
cuda::std::simdtest utilities by @shwina in #9125 - Distribute
std::tuple_(size|element)specializations to implementation files by @davebayer in #9123 - use
std::scoped_lockby @charan-003 in #9124 - [libcu++] Remove empty messages in
static_assertby @davebayer in #9131 - [STF] Properly destroy CUDA streams and do not try to initialize CUDA while capturing by @caugonnet in #8919
- Refactor DeviceReduce dispatch logic by @bernhardmgruber in #9088
- [STF] [TRIVIAL] Rename decorated_stream to augmented_stream by @andralex in #9138
- Make more warpspeed scan stage counts tunable by @bernhardmgruber in #9128
- Skip plotting benchmarks without data by @bernhardmgruber in #9085
- Use the new tuning API internally for
detail::batched_topk::dispatchby @bernhardmgruber in #9095 - [STF] Use out parameter for partition mappers by @caugonnet in #9117
- Hardcode
DeviceTransformOffsetTtoint64by @bernhardmgruber in #9127 - [STF] Deduplicate cudaStreamIsCapturing helper by @andralex in #9136
- [libcu++] Implement
std::constant_wrapperby @davebayer in #9046 - [libcu++] Change
inplace_vectorreturn type fortry_meowmethods by @davebayer in #9130 - [libcu++] Fix
ranges::subrangeuse ofpair-likeby @miscco in #9126 - [libcu++] Update
__cccl_ptx_isafor CTK 13.3 by @davebayer in #9164 - Support launching the devcontainer from a git worktree by @alliepiper in #8950
- CUDA runtime unit test by @charan-003 in #8859
- [STF] Add C API context wait helper by @caugonnet in #9161
- Use the public API for the
transform_reducebenchmark by @bernhardmgruber in #9179 - [cudax] Define
_CUDAX_ENABLE_GROUP_FEATURES_IN_LIBCUDACXXglobally in cudax by @davebayer in #9180 - Refactor warpspeed scan 1/2 by @bernhardmgruber in #9168
- [cub] Fix warp reduce of fixed size random access range by @davebayer in #9153
- Only enabled PDL for PTX/SASS supporting it by @bernhardmgruber in #9163
- [cudax] Disable cuco hashers test for nvcc 12.0 by @davebayer in #9185
- Update all the things to CTK 13.3. by @alliepiper in #9143
- [STF] Fix host-thread data races in concurrent task submission by @caugonnet in #9186
- [cudax] Make
lane_maskpart of the mapping result by @davebayer in #9140 - [cudax] Initial
cudax::coop::reduceprototype by @davebayer in #9154 - [libcu++] Fix
cuda::std::constant_wrappertests for nvcc 13.3 by @davebayer in #9198 - Refactor warpspeed scan 2/2 by @bernhardmgruber in #9169
- Use public CUB API to implement thrust::unique_by_key by @bernhardmgruber in #9197
- Fix MSVC returning incorrect values for complex tan/tanh for large inputs. by @s-oboyle in #9189
- [cudax] Implement
binary_partitionmapping for threads within warp by @davebayer in #8894 - [cudax] Remove synchronizers header reintroduced during rebasing by @davebayer in #9204
- CCCL C v2 by @shwina in #8985
- Avoid passing OOB data to scan operator in warpspeed scan by @bernhardmgruber in #9202
- smoke test to verify GPU memory allocation/deallocation by @charan-003 in #9195
- [STF] Add extended C context creation API by @caugonnet in #9162
- [cudax] Implement
cudax::coop::reduceforcudax::this_clusterby @davebayer in #9167 - [libcu++] Implement tuple-like constructors for
tupleby @miscco in #9059 - [libcu++] Replace
cudaStreamPerThreadwithcudaStream{}in PSTL by @davebayer in #9214 - [libcu++] Suppress
-Wattributesin lit tests with nvcc 12.0 and gcc by @davebayer in #9216 - Add coderabbit disclaimer by @shwina in #9221
- Argument annotation framework for segmented algorithms by @pciolkosz in #8875
- [STF] Add re-launchable popped graphs to stackable_ctx by @caugonnet in #9178
- [libcu++] Fix use of exception keywords by @davebayer in #9220
- Marks conditionally-used params in argument annotation framework
[[maybe_unused]]by @elstehle in #9225 - Use the new tuning API internally for
detail::segmented_scan::dispatchby @gonidelis in #9194 - Allow public tuning of
cub::DeviceAdjacentDifferenceby @bernhardmgruber in #9218 - Allow public tuning of
cub::DeviceMergeSortby @bernhardmgruber in #8600 - Fix NVRTC 13.3 warning bug in CCCL and CCCL.C by @NaderAlAwar in #9171
cudax::fill_bytes(mdspan)by @fbusato in #9193- cudax/stf: migrate stackable/ from cuda_safe_call to cuda_try by @andralex in #9165
- run_to_run deterministic device scan by @srinivasyadav18 in #9098
- [cub] Replace cub parameter framework with cuda::argument by @pciolkosz in #9074
- [libcu++] Rename max to highest in arguments framework by @pciolkosz in #9246
- [cudax] Implement
cudax::coop::reduceforcudax::this_gridby @davebayer in #9203 - [libcu++] Use stream's context in PSTL by @davebayer in #9219
- Prevent type-mismatch errors in cuda.compute.reduce_into by @acosmicflamingo in #9206
cudax::copy(mdspan)Optimize shared memory cases by @fbusato in #9137- [libcu++] Fix issues with new tuple constructors by @miscco in #9261
- Use the new tuning API internally for detail::find::dispatch by @gonidelis in #9240
- [cudax] Update lane mask inside mappings only when unit is thread by @davebayer in #9264
- Use uniform type names for init values throughout CUB by @gonidelis in #9267
- Build more RAPIDS libraries in CI by @trxcllnt in #9116
- fix: Replace runtime_static_assert tests with compile tests. by @arnavnagzirkar in #9211
- Rename
block_scan_algorithmtoscan_algorithmby @bernhardmgruber in #9244 - [libcu++] Move argument bounds helpers to bound file by @miscco in #9276
- [CUB] Adds benchmarks for batched indexed top-k (aka batched arg top-k) by @elstehle in #9288
- [cudax] Implement
cudax::invoke_oneby @davebayer in #9230 - [thrust] Fix missing qualifiers for basic_common_reference by @miscco in #9292
- [libcu++] Disable SIMD tests for tile mode by @miscco in #9296
- Deprecate
AgentMergeSortPolicyandAgentAdjacentDifferencePolicyby @bernhardmgruber in #9235 - [CUB] Makes the batched top-k selection direction compile-time only by @elstehle in #9286
- Allow public tuning of
cub::DeviceTransformby @bernhardmgruber in #8745 - [cuda.compute] Enable building cuda-compute against CCCL C V2 by @shwina in #9200
- [STF] Add C bindings for the places layer by @caugonnet in #9232
- cudax/stf: migrate internal/ launch + host_launch_scope from cuda_safe_call to cuda_try by @andralex in #9249
- Restore override key with a comment by @shwina in #9298
- Allow public tuning of
cub::DeviceForby @bernhardmgruber in #9297 - Allow public tuning of
DeviceFindby @bernhardmgruber in #9236 - cudax/stf: migrate internal/ misc files from cuda_safe_call to cuda_try by @andralex in #9241
- cudax/stf: migrate internal/ parallel_for + cuda_kernel scopes from cuda_safe_call to cuda_try by @andralex in #9265
- [cudax] Implement
cudax::coop::reducefor warp groups within a block by @davebayer in #9258 - Allow public tuning of non-ByKey
cub::DeviceScanby @bernhardmgruber in #8853 - Support
cub::DeviceReducewithout initial value by @bernhardmgruber in #9289 - Better asserts for warp primitives without temp storage by @bernhardmgruber in #9294
- Fix the references to permutation_iterator in shuffle_iterator docs by @djns99 in #9307
- [libcu++] Fix
__is_sequencedefinition for arguments framework by @miscco in #9271 - [infra] Add
clang-22CUDA build job to nightly/weekly CI by @davebayer in #8888 - [libcu++] Implement tuple protocol for
integer_sequenceby @davebayer in #9129 - Allow public tuning of
cub::DeviceScan::*ByKeyby @bernhardmgruber in #9215 - [STF] Add C bindings for stackable contexts by @caugonnet in #9233
- Also test Thrust on SM120 in nightly/weekly CI by @bernhardmgruber in #9201
- [cudax] Implement
cuda::coop::reducefor threads within a warp by @davebayer in #9300 - Take environments by
const&inDeviceTransformby @bernhardmgruber in #9336 - Document CUB temp storage alignment by @jrhemstad in #9302
- Deprecate
DispatchScanandAgentScanPolicyby @bernhardmgruber in #9234 - Allow public tuning of the
cub::DeviceReduce::ReduceByKeyandcub::DeviceRunLengthEncode::Encodeby @bernhardmgruber in #9329 - [STF] Migrate __stf/allocators/ from cuda_safe_call to cuda_try by @andralex in #9147
- [Tile] Improve testing of
__tile__and__device__only functions by @miscco in #9313 - Refresh c2h inspect_changes fixture path by @caugonnet in #9356
- [libcu++] Implement C++26 new tuple-assignments by @miscco in #9227
- [libcu++] Adds
stable_sortedoutput_ordering by @elstehle in #9355 - Make thrust distributions compatible with URNG interface by @RAMitchell in #9319
- Update documentation for cuda-cccl to say python 3.10+ and CC 7.5+ by @NaderAlAwar in #9365
- Use sccache-dist build cluster in optional third-party jobs by @trxcllnt in #9118
- Rename non-library uses of "warpspeed" to "lookahead" by @bernhardmgruber in #9327
- cudax/stf: migrate internal/ context + resources from cuda_safe_call to cuda_try by @andralex in #9248
- Update devcontainers missed in #9118 by @trxcllnt in #9368
- docs: add missing doxygen comment for par_nosync_t::on() by @Oxygen56 in #9196
- [libcu++] Make argument namespace and wrappers construction public by @pciolkosz in #9251
- [cudax] Implement
cudax::coop::shufflefor threads within a warp by @davebayer in #9325 - [cudax] Implement
cudax::coop::shuffle_downfor threads within a warp by @davebayer in #9371 - [cudax] Use
coop::shuffle_downincoop::reduceby @davebayer in #9392 - [cudax] Implement
cudax::coop::shuffle_upfor threads within a warp by @davebayer in #9390 - Implement
cuda::std::basic_format_stringby @davebayer in #5569 - Override GPU name by @gevtushenko in #9385
- Adds tests for segment-specific-k-values by @elstehle in #9311
- [cudax][STF] Clarify stream pool capture and teardown comments by @caugonnet in #9395
- [libcu++] Fix the default device pool getter by @pciolkosz in #9351
- [cudax][STF] Auto-free graph allocations on launch by @caugonnet in #9394
- Intentionally do not
dlcloseJIT compiled libraries in v2 (hostJIT) by @shwina in #9402 - Improves segmented top-k test compilation times by @elstehle in #9404
- [libcu++] Adds a
cuda::execution::tie_breakrequirement by @elstehle in #9238 - [PSTL] Use env based overload for
DeviceFindIfby @miscco in #9318 - [libcudacxx] Add cuda::std::__stringof for compile-time values including function names by @andralex in #9299
- [libcu++] harden the preprocessor machinery to avoid user defined tokens by @miscco in #9407
- [cuda.compute]: stop wrapping binary search comparator in python callable by @NaderAlAwar in #9428
- Histogram tuning policy cleanup by @bernhardmgruber in #9361
- Allow public tuning of the
cub::DeviceRunLengthEncode::NonTrivialRunsby @bernhardmgruber in #9347 - Allow public tuning of
cub::DeviceSelect(withoutUniqueByKey) andcub::DevicePartition(without three-way) by @bernhardmgruber in #9316 - Remove unused scan output type local by @fallintoplace in #9410
- Fix tuning docs for DeviceFind by @bernhardmgruber in #9379
- Remove __syncthreads() duplicate from agent_radix_histogram by @gonidelis in #9445
- run to run scan warpspeed impl sm100+ by @srinivasyadav18 in #9263
- [libcu++] Fix device memory pool test by @pciolkosz in #9442
- Smoke test to verify pinned memory by @charan-003 in #9285
- Productize tuning API for
DeviceSegmentedScanby @gonidelis in #9430 - Allow public tuning of three-way
cub::DevicePartition::Ifby @bernhardmgruber in #9324 - Refactor libcudacxx-style skill by moving CCCL wide style guidelines to a specific file by @NaderAlAwar in #9405
- [CUB, docs-only] Adds docs page on the requirements users can express for
DeviceTopKandDeviceBatchedTopKby @elstehle in #9446 - [cuda.compute]: cache np.dtypes properly in stateful ops by @NaderAlAwar in #9469
- [cudax] Implement broadcasted variants of
cudax::coop::reduceby @davebayer in #9360 - Update NVBench by @gonidelis in #9223
- Reorganize cuco implementation headers under detail/ by @PointKernel in #9436
- Allow public tuning of
cub::DeviceHistogramby @bernhardmgruber in #9362 - Allow public tuning for
cub::DeviceSelect:::UniqueByKeyby @gonidelis in #9370 - Add PDL to
cub::DeviceRadixSortby @gonidelis in #9247 - Rename
warp_threadsin tuning policies by @bernhardmgruber in #9485 - Enable bugprone clang-tidy checks by @Jacobfaib in #9467
- CRITICAL: Add PDL guard back missed in #9247 by @gonidelis in #9497
- [libcu++] Implement
cuda::std::vformat_toby @davebayer in #9443 - [libcu++] Indirect
indirect_binary invocableby @miscco in #9417 - [libcu++] Implement
cuda::std::formatted_sizeby @davebayer in #9472 - Radix sort policy name improvements by @bernhardmgruber in #9489
- Allow public tuning of
cub::DeviceMemcpyby @bernhardmgruber in #9359 - Ensure warpspeed scan uses <= 256 threads for NVHPC by @bernhardmgruber in #9490
- bugprone-signed-char-misuse by @Jacobfaib in #9507
- bugprone-sizeof-expression by @Jacobfaib in #9517
- bugprone-multi-level-implicit-pointer-conversion by @Jacobfaib in #9509
- bugprone-unhandled-self-assignment by @Jacobfaib in #9522
- bugprone-empty-catch by @Jacobfaib in #9514
- bugprone-assignment-in-if-condition by @Jacobfaib in #9519
- [Tile] Disable tile mode for NVCC 13.3 by @miscco in #9488
- [Tile] Mark alignment helpers as
_CCCL_HOST_DEVICE_APIby @miscco in #9487 - [CUB] Refactor
DeviceSelect::Flaggedto always take an environment by @miscco in #9455 - Wraps min/max in parantheses to avoid MSVC compilation issues by @elstehle in #9537
- bugprone-suspicious-stringview-data-usage by @Jacobfaib in #9525
- bugprone-inc-dec-in-conditions by @Jacobfaib in #9527
- Fix
BlockTopK+/-0.0 handling by @pauleonix in #9470 - [STF] Make exec_place/data_place singletons thread-safe by @caugonnet in #9541
- [STF] Fix data race on stackable_logical_data across host threads by @caugonnet in #9540
- bugprone-suspicious-include by @Jacobfaib in #9508
- bugprone-forward-declaration-namespace by @Jacobfaib in #9504
- [cccl.c] Split build step into compile + load by @shwina in #8484
- bugprone-unintended-char-ostream-output by @Jacobfaib in #9521
- Add basic communicator concept by @Jacobfaib in #9426
- bugprone-return-const-ref-from-parameter by @Jacobfaib in #9515
- use
__ballot_syncin warpspeed lookahead by @srinivasyadav18 in #9471 - Add a workaround for nvcc bug in constant wrapper and revert CUB tests changes by @pciolkosz in #9382
- Temporarily disable is_device_accessible peer tests by @pciolkosz in #9547
- [HostJit] Properly use
__declspecon windows by @miscco in #9539 - [libcu++] Implement
cuda::std::format_toby @davebayer in #9474 - [libcu++] Skip
__fp_set_expfpclassifytests on denormals by @davebayer in #9536 - [libcu++] Implement
cuda::std::format_to_nby @davebayer in #9482 - Implement
ranges::zip_viewby @Jacobfaib in #8744 - [cuda.compute]: add benchmarks to measure host side overhead by @NaderAlAwar in #9432
- bugprone-undefined-memory-manipulation by @Jacobfaib in #9518
- bugprone-move-forwarding-reference by @Jacobfaib in #9526
- [libcu++] Implement
cuda::std::dynamic_formatby @davebayer in #9483 - Add multi GPU CI job for libcu++ by @pciolkosz in #9435
- bugprone-casting-through-void by @Jacobfaib in #9523
- [libcu++] Guard cuda::args bounds against types without numeric_limits by @edenfunf in #9473
- [libcu++] Consistently waive memory pool tests if unsupported by @pciolkosz in #9479
- [cuda.compute]: add CI job for
minimalcuda-cccl extra by @NaderAlAwar in #9434 - [CUB] Refactor
DeviceAdjacentDifference::SubtractLeftto always take an environment by @miscco in #9418 - [libcu++] Implement
cuda::std::range_formatby @davebayer in #9558 - [libcu++] Implement tuple-like constructors for
pairby @miscco in #9543 - [CUB] Refactor
DeviceAdjacentDifference::SubtractRightto always take an environment by @miscco in #9419 - [STF] Support ctx.wait() on a token by @caugonnet in #9501
- [libcu++] Implement
cuda::std::formattableconcept by @davebayer in #9544 - Add determinism docs by @srinivasyadav18 in #9350
- Fix clang-tidy early return by @Jacobfaib in #9580
- bugprone-unchecked-optional-access by @Jacobfaib in #9520
- [CUB][Bug] Fix DeviceHistogram out-of-bounds write by @fbusato in #9570
- [DOC] Fix sidebar noise from Breathe overload anchors by @gonidelis in #9585
- [libcu++] Fix peer access case in is_pointer_accessible by @pciolkosz in #9478
- [thrust] Use CCCL Runtime in Thrust set operation tests by @pciolkosz in #9551
- bugprone-pointer-arithmetic-on-polymorphic-object by @Jacobfaib in #9528
- [HostJit] Windows support by @miscco in #9502
- bugprone-forwarding-reference-overload by @Jacobfaib in #9512
- bugprone-narrowing-conversions by @Jacobfaib in #9505
- bugprone-use-after-move by @Jacobfaib in #9516
- Refactor
device_adjacent_difference::dispatchto take a tuning by @miscco in #9454 - [CUB] Add DeviceFind lower/upper bound for sorted values via merge-path by @AneeshGidda in #8780
- [CUB] Refactor
cub::DeviceSelect::Ifto always take an environment by @miscco in #9456 - [CUB] Refactor
DeviceSelect::FlaggedIfto always take an environment by @miscco in #9457 - [CUB] Refactor
DeviceSelect::Uniqueto always take an environment by @miscco in #9458 - bugprone-integer-division by @Jacobfaib in #9524
- [cuda.compute]: Relax cache key in histogram by @NaderAlAwar in #9596
- [libcu++] Add initializer_list overloads for make_*_buffer helpers by @pciolkosz in #9586
- Use cuda::buffer in ccclrt algorithm tests by @pciolkosz in #9591
- [libcu++] Optimize integral formatters by @davebayer in #9606
- [libcu++] Implement pair-like assignments for
pairby @miscco in #9579 - [libcu++] Set current context for buffer driver operations by @pciolkosz in #9615
- Move coderabbit disclaimer to PR template and collapse walkthrough by @NaderAlAwar in #9610
- Allow public tuning of
cub::DeviceRadixSortby @bernhardmgruber in #9491 - [libcu++] Optimize
to_charsintegral width calculation by @davebayer in #9601 - [CUB] Refactor
DeviceHistogram::MultiHistogramEvento always take an environment by @miscco in #9552 - [CUB] Refactor
DeviceHistogram::HistogramEvento always take an environment by @miscco in #9553 - [CUB] Refactor
DeviceHistogram::MultiHistogramRangeto always take an environment by @miscco in #9554 - [CUB] Refactor
DeviceHistogram::HistogramRangeto always take an environment by @miscco in #9555 - [libcu++] Add image processing CCCL Runtime / CUB example by @pciolkosz in #8541
- [libcu++] Add tests for cross device APIs and APIs related to a device with a different device set current by @pciolkosz in #9617
- [cccl.c]: Add
serialize()anddeserialize()functions to enable ahead-of-time compilation workflows by @shwina in #9568 - Enforce MergePolicy tuning values by @bernhardmgruber in #9439
- Duplicate radix sort dispatch to deprecate
DispatchRadixSortby @bernhardmgruber in #9530 - [STF] Exempt STF test CMake files from CMake owners by @caugonnet in #9642
- Add NCCL communicator by @Jacobfaib in #9427
- [docs] Add a note about error handling using exceptions to the docs by @pciolkosz in #9632
- [libcu++] Guard tuple and pair against dangling references by @miscco in #9622
- Pacify clang-tidy for ncclCommInitAll() by @Jacobfaib in #9651
- [cuda.compute]: fix v2 issue with some well known ops not being supported by @NaderAlAwar in #9649
- Exposes
DeviceBatchedTopK::{Min,Max}{Keys,Pairs}for non-deterministic, unordered, and small segments-only by @elstehle in #9331 - Unify atomic and two-phase reduction by @bernhardmgruber in #9349
- [CUB] Refactor
DeviceCopyto always take an environment by @miscco in #9416 - [CUB] Cleanup some dispatch arguments by @miscco in #9603
- [cub] Specialize
std::formatterfor enums by @davebayer in #9641 - Move
bits_per_passlast intopk_policyby @bernhardmgruber in #9640 - [cudax] Add
unit_prefix to unit-related data in mapping result by @davebayer in #9625 - Allow public tuning of
cub::DeviceMergeby @bernhardmgruber in #9637 - Allow public tuning of
cub::DeviceCopyby @bernhardmgruber in #9658 - Move CUB test docs into developer docs by @bernhardmgruber in #9662
- Update Node 20 GitHub Actions by @jrhemstad in #9673
- [cudax][cuco] Use the detail config header and proper visibility macros in cuco headers by @PointKernel in #9675
- [STF] Fix slice copy context activation on multi-GPU by @caugonnet in #9677
- [STF] Fix multi-context parallel for for grid places by @caugonnet in #9604
- Improve tuning introduction docs by @bernhardmgruber in #9661
- Unify reduce and rfa policy by @bernhardmgruber in #9639
- [CUB] Refactor
DeviceReduce::Arg{Min, Max}to always take an environment by @miscco in #9403 - Add
DeviceSegmentedScanto device-wide docs by @bernhardmgruber in #9690 - [TRIVIAL] Pin numba below 0.66 for Python CUDA extras by @caugonnet in #9692
- Fix NCCL tests for multi-GPU by @Jacobfaib in #9693
- fix: cache cuda.compute builds for closures over Python scalars by @nethum529 in #9680
- [libcu++] Fix iterator traversal checks for
minmax_elementby @miscco in #9694 - [libcu++] Try and avoid MSVC circular is_constructible chain by @miscco in #9683
- [HostJit] Add all tested device APIs to the PCH cache by @miscco in #9663
- Productize
DeviceSegmentedSorttuning API by @gonidelis in #9681 - Use
cuda::std::is_sufficiently_alignedin CCCL by @davebayer in #9684 - [cudax] Implement
cudax::coop::any_ofalgorithm for <= warp groups by @davebayer in #9665 - [CUB] Refactor
dispatch_streaming_arg_reduceto take a tuning environment by @miscco in #9660 - Allow public tuning of non-ByKey
cub::DeviceReduceby @bernhardmgruber in #8863 - Allow public tuning of
cub::DeviceFind::*BoundSortedValuesby @bernhardmgruber in #9688 - Rename
FindPolicytoFindIfPolicyby @bernhardmgruber in #9689 - Allow public tuning of
cub::DeviceSegmentedRadixSortby @bernhardmgruber in #9685 - Fix
cudafe++< 13.1 with older gcc by @davebayer in #9657 - Fix CUB DeviceReduce env overloads to accept no_init_t by @Jacobfaib in #9676
- Fix use of
reduce_policyby @davebayer in #9704 - Add fixed_capacity_map to cudax by @srinivasyadav18 in #7705
- [libcu++] Additional peer device copy testing by @pciolkosz in #9636
- [places] Align cyclic_shape::size with iteration cardinality by @fallintoplace in #9148
- Allow public tuning of
cub::DeviceSegmentedReduceby @bernhardmgruber in #9686 - Small refactorings for batched memcpy by @bernhardmgruber in #9659
- Improve NV_IF_ELSE_TARGET formatting in segmented reduce dispatch by @bernhardmgruber in #9709
- Control unrolling in
cub::DeviceMerge[Sort]via tuning by @bernhardmgruber in #9181 - Remove
num_from tuning policy members by @bernhardmgruber in #9717 - [cuda.compute] Expose
.serialize()and.deserialize()methods in Python by @shwina in #9644 - bugprone-misplaced-widening-cast by @Jacobfaib in #9506
- Rename
AgentTopKPolicytoagent_topk_policyby @bernhardmgruber in #9711 - Fix doxygen tuning doc headings by @bernhardmgruber in #9713
- Rename policy selector tests by @bernhardmgruber in #9715
- [STF][trivial] Forward declare stf_ctx_handle in C-STF header place section by @caugonnet in #9705
- [docs] Document that streams are created as non-blocking by @pciolkosz in #9707
- Fix cuco clang-tidy error by @Jacobfaib in #9736
- Vectorize output store in ublkcp DeviceTransform kernel by @nanan-nvidia in #9481
- Separate from and deprecate
DispatchScanByKeyby @bernhardmgruber in #9714 - [cudax] Change cuco test target names by @davebayer in #9738
- [cub] Fix clang-tidy narrowing conversion warning by @davebayer in #9741
- Swap
small_segmentandmedium_segmentinSegmentedSortPolicyby @bernhardmgruber in #9716 - Cover __half and __nv_bfloat16 in CUB Reduce/Scan/RadixSort benchmarks by @edenfunf in #9708
- Make stream and memory resource explicit in fixed_capacity_map APIs by @PointKernel in #9719
- [CUB] warpspeed kernel: make scan_resources_t copy by @srinivasyadav18 in #9751
- [cuda.compute]: Enable (AoT) compilation for multiple compute capabilities by @shwina in #9732
- Rename cuco
.hppheaders to.cuhby @PointKernel in #9735 - [STF] Name green context data places by handle and support VMM mem_create by @caugonnet in #9706
- Fix test helper producing inverted interval by @pauleonix in #9753
- [STF] Export blocked_partition_custom into cuda::experimental::stf by @caugonnet in #9758
- Make stream and memory resource explicit in hyperloglog APIs by @PointKernel in #9720
- warpspeed run_to_run deterministic scan for SM90 using atomic global counter by @srinivasyadav18 in #9565
- [CUB]
DeviceReducewith device resident problem size by @NaderAlAwar in #9722 - Symlink agent files instead of telling bots to read other files by @Jacobfaib in #9750
- Deprecate
cub::ChainedPolicyby @bernhardmgruber in #9744 - Deprecate CUB device count functions by @bernhardmgruber in #9743
- Make DeviceTransform a P0 benchmark for QA by @bernhardmgruber in #9774
- Do not vectorize large type reductions by @bernhardmgruber in #9762
- Tuning policy cleanup 1/2 by @bernhardmgruber in #9710
- Strip zero-width or unprintable unicode characters by @Jacobfaib in #9564
- [Tile] tile DeviceTransform port by @nanan-nvidia in #9210
- Create a swapfile in
workflow-run-job-linuxby @trxcllnt in #9739 - Test unaligned temp storage by @bernhardmgruber in #9746
- Run dummy test if unit test would be empty in C++17 by @bernhardmgruber in #9784
- Fully classify the
cub::detail::InputValuetemplate to avoid implicit conversions by @Jacobfaib in #9772 - [cuda.compute][cuda.coop]: Replace all usages of device arrays outside examples with new wrapper hat does not depend on cupy or numba-cuda by @NaderAlAwar in #9653
- Handle a type of size exactly 3 in cuda::buffer by @Jacobfaib in #9776
- Tuning policy cleanup 2/2 by @bernhardmgruber in #9712
- bugprone-branch-clone by @Jacobfaib in #9513
- Add builds for MSVC cccl_c_parallel by @miscco in #9605
- Use __builtin_bswapg in cuda::std::byteswap when available by @Functionhx in #9785
- [infra] Add multi-gpu CI for cudax by @pciolkosz in #9590
- Add MGMN Reduce by @Jacobfaib in #9645
- [libcu++] Use cccl runtime in thrust tests pt 2 by @pciolkosz in #9633
- Remove
_CCCL_GRID_CONSTANTfromDeviceSelectSweepKernelparameters by @nanan-nvidia in #9795 - Fix Thrust contiguous iterator unwraps for cuda::device_buffer by @sleeepyjack in #9756
- [STF] Implement cyclic_partition::get_executor by @caugonnet in #9803
- Add tuning policies to CUB API docs by @bernhardmgruber in #9745
- Nits in
score.pyby @gonidelis in #9809 - Fix scanning out-of-bounds items in OpenMP scan by @bernhardmgruber in #9759
- [docs] Fix saturating overflow arithmetic docs by @davebayer in #9812
- Reduce P0 DeviceTransform benchmarks to babelstream and fill by @bernhardmgruber in #9780
- Pin NVTX to commit before upstream scope refactor by @bernhardmgruber in #9820
- [cub] Fix
bugprone-branch-cloneerror by @davebayer in https://github.com/NVIDIA/cccl/pull/9810 - [cudax][cuco] Port fixed_capacity_map benchmarks from cuCollections by @sleeepyjack in https://github.com/NVIDIA/cccl/pull/9748
- Remove _CCCL_GRID_CONSTANT from merge sort kernel parameters by @nanan-nvidia in https://github.com/NVIDIA/cccl/pull/9829
- [libcu++] Implement P3798R1 The unexpected in
std::expectedby @davebayer in #9733 - Refactor enum to string functions by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9752
- [cudax] Implement
cudax::takemapping by @davebayer in https://github.com/NVIDIA/cccl/pull/9818 - Use plus tunings for lookahead scan more widely by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9813
- Move cccl_get_catch2() into testing block by @codingwithmagga in https://github.com/NVIDIA/cccl/pull/9383
- [clang-tidy] Suppress new warnings emitted by clang-tidy-22 by @davebayer in https://github.com/NVIDIA/cccl/pull/9855
- [cccl.c] Fix scan policy mismatch between host and NVRTC by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9770
- DeviceScan::InclusiveScan() accepts a
cuda::args::deferredby @Jacobfaib in #9826 - Use
cuda::std::numbersin cccl by @davebayer in https://github.com/NVIDIA/cccl/pull/4955 - Remove most
CCCL_GRID_CONSTANTannotations by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9858 - Revert "Pin NVTX to commit before upstream scope refactor" (#9820) by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9850
- Add multi-GPU exclusive_scan by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9796
- Add deferred argument to 2-phase InclusiveScan by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9865
- [CUB] GPU-to-GPU
DeviceReducewith device resident problem size by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9740 - [FEA] Replace not_equal_to_val with cuda::std::not_fn(cuda::equal_to_value{}) by @rkothari3 in https://github.com/NVIDIA/cccl/pull/9833
- Move SASS diff description to a SKILL by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9767
- [Thrust] Split transform test by @miscco in https://github.com/NVIDIA/cccl/pull/9701
- Add pytorch inspired DeviceTransform benchmark by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9764
- Add
cudax::distributedandcudax::returned_tospecifiers for cooperative algorithms by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9862 - Remove unnecessary
is_evenoverloads for complex by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9873 - Replace
bool_constantbyif constexprin agent_sub_warp_merge_sort by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9876 - [cudax] Fix cudax benchmarks prefix by @davebayer in https://github.com/NVIDIA/cccl/pull/9878
- [STF] Add exec places from externally-owned CUDA contexts by @caugonnet in https://github.com/NVIDIA/cccl/pull/9779
- [CUB] Refactor
DeviceSelect::UniqueByKeyto always take an environment by @miscco in https://github.com/NVIDIA/cccl/pull/9459 - [cuda.compute]: make cuda.compute thread safe to enable free threaded wheels by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9475
- [CUB] Refactor
DevicePartition::Flaggedto always take an environment by @miscco in https://github.com/NVIDIA/cccl/pull/9463 - Add multi-GPU inclusive_scan by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9827
- Add
cstdioandcstdarghostlib headers by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9879 - [cudax][cuco] Add HyperLogLog benchmarks by @sleeepyjack in https://github.com/NVIDIA/cccl/pull/9889
- [cuda.cccl] Simplify Python CMake configuration by @caugonnet in https://github.com/NVIDIA/cccl/pull/9883
- [CUB] Refactor
DevicePartition::Ifto always take an environment by @miscco in https://github.com/NVIDIA/cccl/pull/9464 - [places] Add checked execution grid reshaping by @caugonnet in https://github.com/NVIDIA/cccl/pull/9898
- Fix dead 64-bit rotate builtin macros and popcount tile undef in by @temujinkz in https://github.com/NVIDIA/cccl/pull/9874
- [STF] Add structured tensor partitions and placement evaluation to the places layer by @caugonnet in https://github.com/NVIDIA/cccl/pull/9804
- [libcu++] Add
__float128support forcuda::std::fabsby @davebayer in https://github.com/NVIDIA/cccl/pull/9895 - Don't use the thrust size dispatchers anymore in multi-GPU by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9978
- Fix tuning policy links in docs by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9897
- Slightly bump test sizes for mgmn tests by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9980
- Add initial docs for CI + infrastructure systems by @alliepiper in https://github.com/NVIDIA/cccl/pull/9611
- cudax/stf: migrate stream/interfaces/ from cuda_safe_call to cuda_try by @andralex in https://github.com/NVIDIA/cccl/pull/9268
- [STF] Drive parallel_for and data placement from structured partitions by @caugonnet in https://github.com/NVIDIA/cccl/pull/9808
- Use result policies for MGMN algorithms by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9979
- Add unaligned versions of cub::DeviceTransform babelstream and fill benchmarks by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/9881
- Add owning nccl communicator by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9977
- [bench] Enable metatargets for benchmark targets by @davebayer in https://github.com/NVIDIA/cccl/pull/9990
- cudax/stf: cuda_try migration — stream event_types (PR7) by @andralex in https://github.com/NVIDIA/cccl/pull/9303
- cudax/stf: cuda_try migration — graph misc (PR6) by @andralex in https://github.com/NVIDIA/cccl/pull/9301
- [libcu++] Use == and < in cuda::args instead of <= by @pciolkosz in https://github.com/NVIDIA/cccl/pull/9884
- cudax/stf: cuda_try migration — stream_ctx (PR8) by @andralex in https://github.com/NVIDIA/cccl/pull/9306
- [STF] Migrate __stf/utility/ from cuda_safe_call to cuda_try by @andralex in https://github.com/NVIDIA/cccl/pull/9150
- Fix some miscellaneous clang-tidy errors by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9999
- Remove nvcc arch flags from clangd by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9996
- [nvtarget] Add support for
sm_107by @davebayer in https://github.com/NVIDIA/cccl/pull/10008 - [STF] Support compound while-loop conditions in the C API by @caugonnet in https://github.com/NVIDIA/cccl/pull/10006
- Small CUB doc corrections by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10011
- [libcu++] Optimize
cuda::std::rotlandcuda::std::rotrby @davebayer in https://github.com/NVIDIA/cccl/pull/10004 - Update
__cccl_ptx_isafor clang-cuda 22 by @davebayer in https://github.com/NVIDIA/cccl/pull/10007 - Fix unused comparison operators now that cuda::arguments only uses
operator<by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/10021 - [cuda.compute]: Serialize hostjit builds and improve tests by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9998
- [cuda.compute]: Fix segmented sort selector race by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10024
- Use if in _CCCL_TRY_CUDA_API by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9985
- Restructure the RLE encode tuning policy by @nanan-nvidia in https://github.com/NVIDIA/cccl/pull/10028
- Enable
__float128incuda::std::fpclassify,isnormal,ilogbandlogbby @temujinkz in https://github.com/NVIDIA/cccl/pull/10014 - Flatten CUB/Thrust API docs by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10012
- Enable
__float128incuda::std::copysignandcuda::std::signbitby @temujinkz in https://github.com/NVIDIA/cccl/pull/9991 - Docs follow-up nits for flattened CUB/Thrust docs (#10012) by @gonidelis in https://github.com/NVIDIA/cccl/pull/10042
- [cuda.compute]: Add pytest-run-parallel by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9886
cuda_errorfixups by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/10005- [libcu++] Make
DoNotOptimizefunctional on device by @davebayer in https://github.com/NVIDIA/cccl/pull/10037 - Extend ChainedPolicy test for
sm_107by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10047 - [STF] Expose geometry-aware allocation and native partition functions in the C API by @caugonnet in https://github.com/NVIDIA/cccl/pull/10038
- [libcu++] Implement P3793R2 Better Shifting (without SIMD) by @davebayer in #9993
- Exposes
cuda::execution::guaranteeby @elstehle in https://github.com/NVIDIA/cccl/pull/10022 - [cuda.compute]: Relax synchronization around Clang compilation in v2, keeping it only for linking by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10051
- [cuda.compute]: Fix windows race related to get nvrtc type name by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10049
- Remove trailing whitespace from CODEOWNERS by @pauleonix in https://github.com/NVIDIA/cccl/pull/10062
- [cuda.compute]: Fail Windows Python test jobs on native command errors by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10046
- [libcu++] Remove
cuda::__(shl|shr)by @davebayer in https://github.com/NVIDIA/cccl/pull/10050 - Minor QOL improvements to devcontainer launching script by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/10055
- bugprone-exception-escape by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/9510
- Prepare three-way partition tuning policies for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10077
- Prepare non-trivial-runs tuning policy for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10080
- [cuda.compute]: Add smoke tests for benchmarks by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9885
- Add pretty printers for cuda::buffer by @Jacobfaib in #9866
- Prepare scan-by-key tuning policy for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10079
- [cuda.compute]: bind nvJitLink to _12_0 aliases so cu12 wheels load on CTK 12.0 by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10083
- support for p0 explicit benchmarks by @srinivasyadav18 in https://github.com/NVIDIA/cccl/pull/10054
- Prepare reduce-by-key tuning policy for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10078
- Use block-scoped atomics in BlockHistogramAtomic on SM 60+ by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10084
- Take CUB environments by
const&by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10105 - [CUB] Fix DeviceAdjacentDifference for cuda::device_buffer iterators by @Noperi0r in #9861
- Implement prefetching by @gonidelis in https://github.com/NVIDIA/cccl/pull/9723
- [cudax][cuco] Reject undersized hyperloglog_ref storage by @sleeepyjack in https://github.com/NVIDIA/cccl/pull/10220
- Enable
__float128incuda::stdcomparison functions (isgreaterfamily) by @temujinkz in https://github.com/NVIDIA/cccl/pull/10226 - [libcu++] Optimize
cuda::std::saturating_castby @davebayer in https://github.com/NVIDIA/cccl/pull/9724 - [cudax][cuco] Fix const mismatch in cooperative HyperLogLog merge by @sleeepyjack in https://github.com/NVIDIA/cccl/pull/10217
- Split cuopt RAPIDS build by @bdice in https://github.com/NVIDIA/cccl/pull/10215
- Remove cuda.coop._experimental from cuda-cccl by @tpn in https://github.com/NVIDIA/cccl/pull/10104
- [cudax][cuco] Return double from HyperLogLog estimate by @sleeepyjack in https://github.com/NVIDIA/cccl/pull/10218
- [libcu++] Fix
cuda::std::aligned_allocargument order by @davebayer in #9728 - Pin cuda-toolkit wheel to container's CTK major.minor in CI by @leofang in https://github.com/NVIDIA/cccl/pull/8160
- [cudax][cuco] Migrate host and device find APIs for fixed_capacity_map by @PointKernel in https://github.com/NVIDIA/cccl/pull/9868
- [cuda.compute]: Add ci matrix entry for minimal ft testing on windows by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9887
- Improve memory resource documentation by @bdice in https://github.com/NVIDIA/cccl/pull/10228
- [cuda.compute]: add CI job that builds c.parallel with thread sanitizer by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/9986
- [libcudacxx][fp] add fpemu (double-precision emulation) + unit tests by @akolesov-nvidia in https://github.com/NVIDIA/cccl/pull/9777
- Implement
cuda::std::is_virtual_base_ofby @davebayer in #4397 - [libcu++] Avoid compilation issue with tuple_of_iterator_references by @miscco in https://github.com/NVIDIA/cccl/pull/10181
- [cuda.compute]: add back CI for CTK 12.0 by @NaderAlAwar in https://github.com/NVIDIA/cccl/pull/10057
- [libcu++] Use
lit-style tests forfpemu(1/2) by @davebayer in https://github.com/NVIDIA/cccl/pull/10233 - Prepare select and partition tuning policies for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10076
- Prepare batched-copy tuning policy for lookahead by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/10081
- feat(security): onboard pre-commit and Pulse secret scanning by @gmanal in https://github.com/NVIDIA/cccl/pull/10010
- fix(ci): restore actions: read for secret-scan reusable workflow by @gmanal in https://github.com/NVIDIA/cccl/pull/10515
- Implement
P2835R7andP3936R1std::atomic_ref::address()by @charan-003 in #9621 - cudax/stf: cuda_try migration — stream_task (PR9) by @andralex in https://github.com/NVIDIA/cccl/pull/10003
- Prevent OOB write for RBK and RLE in streaming context by @nanan-nvidia in #10504
- Replace deprecated
std::aligned_storage_tin STFsmall_vectorby @caugonnet in https://github.com/NVIDIA/cccl/pull/10507 - Add pretty printers for cuda::std::array by @pieroevcc in #10528
- [libcu++] Avoid use of
__CUDA_ARCH__in fpemu by @davebayer in https://github.com/NVIDIA/cccl/pull/10385 - [libcu++] Fix
_Float64test for fpemu in C++23 by @davebayer in https://github.com/NVIDIA/cccl/pull/10508 - [cuco] Adds support for 1 and 2 byte key and value types in
fixed_capacity_mapby @ryanjspears in https://github.com/NVIDIA/cccl/pull/10025 - cudax/stf: cuda_try migration — graph_task (PR10) by @andralex in https://github.com/NVIDIA/cccl/pull/10523
- [libcu++] Implement internal resizable buffer by @pciolkosz in https://github.com/NVIDIA/cccl/pull/10221
- Add NVRTC compatibility errors to CUB headers by @hzaidi05 in #7079
- Reduce PR NVHPC CI coverage by @jrhemstad in https://github.com/NVIDIA/cccl/pull/10527
- Restore _CCCL_GRID_CONSTANT on streaming_context in DeviceSelect by @nanan-nvidia in https://github.com/NVIDIA/cccl/pull/10537
- [Infra] Split SM75 GPU runners more evenly between RTX2080 and T4 by @miscco in https://github.com/NVIDIA/cccl/pull/10513
- Add environment docs landing page with essential info by @gonidelis in https://github.com/NVIDIA/cccl/pull/10043
- [libcu++] Do not instantiate types for inline variables of constructible traits by @miscco in https://github.com/NVIDIA/cccl/pull/10250
- [libcu++] Implement P3104R6 Bit permutations by @davebayer in #10063
- [Tile] Mark formatters of tuning policies as host_device by @miscco in https://github.com/NVIDIA/cccl/pull/10539
- [Tile] Wrap
char_traits::eqin a functor by @miscco in https://github.com/NVIDIA/cccl/pull/10542 - [cudax] Exempt places CMake files from cmake-codeowners review by @caugonnet in https://github.com/NVIDIA/cccl/pull/9782
- [Tile] Mark fpemu as unsupported in tile mode by @miscco in https://github.com/NVIDIA/cccl/pull/10538
- [libcu++] Fix
cuda::make_tma_descriptor(...)by @davebayer in https://github.com/NVIDIA/cccl/pull/10546 - [Tile] Mark all atomics functions as host device only by @miscco in https://github.com/NVIDIA/cccl/pull/10543
- Fixup _CCCL_HOST_API and inline constexpr variables in memory land by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/10536
- Fix more static constexpr -> inline constexpr by @Jacobfaib in https://github.com/NVIDIA/cccl/pull/10549
- CUB: Add a custom test wrapper with memory classification by @ratnachandugembali-lgtm in https://github.com/NVIDIA/cccl/pull/10529
- [Tile] Fix complex interop with extended floating point types by @miscco in https://github.com/NVIDIA/cccl/pull/10550
- [TILE] Add a workaround to make
invokework in tile mode by @miscco in https://github.com/NVIDIA/cccl/pull/10545 - cudax/stf: add defer_exception and void checks for SCOPE by @andralex in https://github.com/NVIDIA/cccl/pull/10559
- [libcudacxx] Add GDB/LLDB pretty-printers for cuda::std::complex and cuda::complex by @HenrikGharagyozyan in #10563
- Add tunable prefetching to DeviceSelect::Flagged by @anikaj-eng in https://github.com/NVIDIA/cccl/pull/10519
- [cub] Always set smem limit to max for lookahead scan by @davebayer in https://github.com/NVIDIA/cccl/pull/10570
- Use devcontainers from RAPIDS 26.10 by @bdice in https://github.com/NVIDIA/cccl/pull/10526
- [libcu++] Add new memory pool attributes from CUDA 13.3 by @pciolkosz in #9798
- cudax: complete Library Fundamentals TS v3 scope guards by @andralex in https://github.com/NVIDIA/cccl/pull/10565
- Compile time benchmarking tool by @griwes in https://github.com/NVIDIA/cccl/pull/9498
- [Backport branch/3.5.x] [libcu++] Allow no GPUDirect RDMA flush options by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/10633
- [Backport branch/3.5.x] Fixes a race condition in
DeviceTopKby @github-actions[bot] in #10683 - [Backport branch/3.5.x] [Tile] Disable tile support by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/10969
- [Backport branch/3.5.x] [libcu++] Fix the base 10 width of exact powers of ten in to_chars by @github-actions[bot] in #10990
- [Backport branch/3.5.x] [CUB] Fix unqualified calls to libcu++ entities by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11015
- [backport 3.5] Remove fpemu from 3.5 release by @davebayer in https://github.com/NVIDIA/cccl/pull/10613
- [Backport branch/3.5.x] [libcu++] Allow alternate pinned memory type reporting by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/10632
- [backport 3.5.x] Remove
constant_wrapperfrom 3.5 release by @davebayer in https://github.com/NVIDIA/cccl/pull/11065 - [Backport branch/3.5.x] [libcu++] Add
TREAT_WARNINGS_AS_ERRORS.litparser by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11087 - [Backport branch/3.5.x] [libcu++] Disable flaky test by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11114
- [Backport to 3.5] Fix MSVC26 (#11103) by @bernhardmgruber in https://github.com/NVIDIA/cccl/pull/11110
- [Backport 3.5] Update cuda::ptx for CUDA 13.4 by @pciolkosz in #11054
- [Backport branch/3.5.x] Add inputs for controlling which repo and branch are checked out by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11009
- [Backport branch/3.5.x] Deduplicate and extend nightly/weekly workflows. (#11074) by @wmaxey in https://github.com/NVIDIA/cccl/pull/11162
- [Backport branch/3.5.x] [CI] Add fields in CI matrix that determine devcontainer repo and runner labels by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11125
- [Backport 3.5] Backport #10747 and #11111 by @miscco in https://github.com/NVIDIA/cccl/pull/11201
- [Backport branch/3.5.x] Fixup - remove extra parameter in workflow call by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11194
- [Backport branch/3.5.x] [libcu++] Fix
__builtin_bswapgnot being supported by nvcc by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11178 - [Backport 3.5] Backport
tupleandpairimprovements by @miscco in https://github.com/NVIDIA/cccl/pull/11475 - [Backport branch/3.5.x] [libcu++] Fix
arch_traitsfor sm100 by @github-actions[bot] in https://github.com/NVIDIA/cccl/pull/11484
New Contributors
- @vip892766gma made their first contribution in #8980
- @arnavnagzirkar made their first contribution in #9211
- @Oxygen56 made their first contribution in #9196
- @AneeshGidda made their first contribution in #8780
- @nethum529 made their first contribution in #9680
- @Functionhx made their first contribution in #9785
- @codingwithmagga made their first contribution in https://github.com/NVIDIA/cccl/pull/9383
- @rkothari3 made their first contribution in https://github.com/NVIDIA/cccl/pull/9833
- @Noperi0r made their first contribution in #9861
- @gmanal made their first contribution in https://github.com/NVIDIA/cccl/pull/10010
- @pieroevcc made their first contribution in #10528
- @hzaidi05 made their first contribution in #7079
- @anikaj-eng made their first contribution in https://github.com/NVIDIA/cccl/pull/10519
Full Changelog: v3.5.0.dev...v3.5.0