Skip to content

Use native fp32 atomic min/max when the device supports cl_ext_float_atomics #1466

Description

@pvelesko

PR #1405 lowers atomicrmw fmin and fmax to cmpxchg loops unconditionally, so fp32 atomic min/max never reaches the hardware instruction on devices that have one.

rtdevlib already implements this dispatch for atomic add. appendRuntimeObjects() in src/backend/OpenCL/CHIPBackendOpenCL.cc picks atomicAddFloat_native or atomicAddFloat_emulation from hasFP32AtomicAdd(), but atomicMinMaxFloat_emulation is appended unconditionally and has no _native peer in bitcode/.

Measured on Aurora (Intel(R) Data Center GPU Max 1550), atomic min over 4M work items, native atomic_fetch_min_explicit vs a cmpxchg loop, best of 5:

contention                   fp32 native   fp32 cmpxchg    ratio
1 addr                          5.596 ms       11.192 ms    2.00x
64 addrs                        2.608 ms       15.954 ms    6.12x
1K addrs                        0.279 ms        1.808 ms    6.48x
64K addrs                       0.026 ms        0.067 ms    2.57x
4M addrs (uncontended)          0.026 ms        0.067 ms    2.57x

The peak is at moderate contention, which is where reduction trees operate.

This is fp32 only. fp64 gains nothing and does not need a native path. PVC reports CL_DEVICE_DOUBLE_FP_ATOMIC_CAPABILITIES_EXT = 0x70007, so the min/max bits are set, but IGC emulates the operation rather than issuing an instruction, and native measures 1.00x against a cmpxchg loop at every contention level. From an ocloc ISA dump for -device pvc:

//.kernel n32
        atomic_fmin.ugm.d32.a64 (32|M0)  null:0 [r2:4] r6:2

//.kernel n64
        atomic_or.ugm.d64.a64   (32|M0)  r10:4  [r2:4] r6:4
        sel (16|M0) (lt)f0.0  r30.0<1>:df  r38.0<1;1,0>:df ...
        atomic_icas.ugm.d64.a64 (32|M0)  r22:4  [r18:4] r26:8

Sketch of the change:

  1. Add bitcode/atomicMinMaxFloat_native.cl using atomic_fetch_min_explicit and atomic_fetch_max_explicit, following atomicAddFloat_native.cl.
  2. Add hasFP32AtomicMinMax() in src/backend/OpenCL/CHIPBackendOpenCL.hh reading the CL_DEVICE_GLOBAL_FP_ATOMIC_MIN_MAX_EXT and CL_DEVICE_LOCAL_FP_ATOMIC_MIN_MAX_EXT bits of Fp32AtomicAddCapabilities_, which already holds the full CL_DEVICE_SINGLE_FP_ATOMIC_CAPABILITIES_EXT bitfield, so no new clGetDeviceInfo call is needed.
  3. Switch on it in appendRuntimeObjects().
  4. Have the pass from llvm_passes: lower atomicrmw fmin and fmax to cmpxchg loops #1405 emit a call to the existing __chip_atomic_min_f32 import instead of an inline cmpxchg loop. spirvNeedsRtdevlib() in src/Utils.hh already matches on __chip_atomic_min.

One caveat. rtdevlib is resolved by clLinkProgram, which fails with -59 on Mali-G52, so routing fmin and fmax through it would regress Mali unless the inline lowering is kept for devices where the link step is unavailable.

Measurements come from a debug queue job on Aurora, single tile, ZE_AFFINITY_MASK=0.0, OpenCL backend.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions