Skip to content

binary_partition

产品支持情况

  • Ascend 950PR/Ascend 950DT:支持
  • Atlas A3 训练系列产品/Atlas A3 推理系列产品:不支持
  • Atlas A2 训练系列产品/Atlas A2 推理系列产品:不支持
  • Atlas 200I/500 A2 推理产品:不支持
  • Atlas 推理系列产品AI Core:不支持
  • Atlas 推理系列产品Vector Core:不支持
  • Atlas 训练系列产品:不支持

功能说明

binary_partition API用于根据一个标签(0或1)将父组划分为两个子组,标签相同的线程会被分配到同一组中。

函数原型

C++
coalesced_group binary_partition(const coalesced_group& g, bool pred)
C++
template <unsigned int Size, typename ParentT>
coalesced_group binary_partition(const thread_block_tile<Size, ParentT>& g, bool pred)

参数说明

表1 参数说明

参数名输入/输出描述
g输入被划分的父组,类型可以是coalesced_groupthread_block_tile
pred输入标签,用于划分子组。

返回值说明

返回划分出的子组coalesced_group对象。

约束说明

  • gthread_block_tile<Size, ParentT>类型时,g必须满足Size小于等于32,否则编译报错。

调用示例

  • SIMT编程场景:

    C++
    using namespace cooperative_groups;
    __global__ void simt_kernel(int *inputArr, ...)
    {
        auto block = this_thread_block();
        auto tile32 = tiled_partition<32>(block);
    
        // inputArr中是随机的整数
        int elem = inputArr[block.thread_rank()];
        // 根据elem&1是否为true将tile32划分为两个子组
        auto subtile = binary_partition(tile32, (elem & 1));
        ...
    }
    
  • SIMD与SIMT混合编程场景:

    C++
    using namespace cooperative_groups;
    __simt_vf__ inline void simt_kernel(__gm__ int *inputArr, ...)
    {
        ...
        auto block = this_thread_block();
        auto tile32 = tiled_partition<32>(block);
    
        // inputArr中是随机的整数
        int elem = inputArr[block.thread_rank()];
        // 根据elem&1是否为true将tile32划分为两个子组
        auto subtile = binary_partition(tile32, (elem & 1));
        ...
    }
    

免责声明:本站内容由 asc-devkit 仓 master 分支自动编译生成,属于持续开发版本,可能存在缺陷,仅供预览与参考。如需稳定及商用资料,请查阅官方 昇腾社区