You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
PR #1405 lowers atomicrmw fmin and fmax to cmpxchg loops unconditionally, so fp32 atomic min/max never reaches the hardware instruction on devices that have one.
rtdevlib already implements this dispatch for atomic add. appendRuntimeObjects() in src/backend/OpenCL/CHIPBackendOpenCL.cc picks atomicAddFloat_native or atomicAddFloat_emulation from hasFP32AtomicAdd(), but atomicMinMaxFloat_emulation is appended unconditionally and has no _native peer in bitcode/.
Measured on Aurora (Intel(R) Data Center GPU Max 1550), atomic min over 4M work items, native atomic_fetch_min_explicit vs a cmpxchg loop, best of 5:
contention fp32 native fp32 cmpxchg ratio
1 addr 5.596 ms 11.192 ms 2.00x
64 addrs 2.608 ms 15.954 ms 6.12x
1K addrs 0.279 ms 1.808 ms 6.48x
64K addrs 0.026 ms 0.067 ms 2.57x
4M addrs (uncontended) 0.026 ms 0.067 ms 2.57x
The peak is at moderate contention, which is where reduction trees operate.
This is fp32 only. fp64 gains nothing and does not need a native path. PVC reports CL_DEVICE_DOUBLE_FP_ATOMIC_CAPABILITIES_EXT = 0x70007, so the min/max bits are set, but IGC emulates the operation rather than issuing an instruction, and native measures 1.00x against a cmpxchg loop at every contention level. From an ocloc ISA dump for -device pvc:
Add bitcode/atomicMinMaxFloat_native.cl using atomic_fetch_min_explicit and atomic_fetch_max_explicit, following atomicAddFloat_native.cl.
Add hasFP32AtomicMinMax() in src/backend/OpenCL/CHIPBackendOpenCL.hh reading the CL_DEVICE_GLOBAL_FP_ATOMIC_MIN_MAX_EXT and CL_DEVICE_LOCAL_FP_ATOMIC_MIN_MAX_EXT bits of Fp32AtomicAddCapabilities_, which already holds the full CL_DEVICE_SINGLE_FP_ATOMIC_CAPABILITIES_EXT bitfield, so no new clGetDeviceInfo call is needed.
One caveat. rtdevlib is resolved by clLinkProgram, which fails with -59 on Mali-G52, so routing fmin and fmax through it would regress Mali unless the inline lowering is kept for devices where the link step is unavailable.
Measurements come from a debug queue job on Aurora, single tile, ZE_AFFINITY_MASK=0.0, OpenCL backend.
PR #1405 lowers
atomicrmw fminandfmaxto cmpxchg loops unconditionally, so fp32 atomic min/max never reaches the hardware instruction on devices that have one.rtdevlib already implements this dispatch for atomic add.
appendRuntimeObjects()insrc/backend/OpenCL/CHIPBackendOpenCL.ccpicksatomicAddFloat_nativeoratomicAddFloat_emulationfromhasFP32AtomicAdd(), butatomicMinMaxFloat_emulationis appended unconditionally and has no_nativepeer inbitcode/.Measured on Aurora (
Intel(R) Data Center GPU Max 1550), atomic min over 4M work items, nativeatomic_fetch_min_explicitvs a cmpxchg loop, best of 5:The peak is at moderate contention, which is where reduction trees operate.
This is fp32 only. fp64 gains nothing and does not need a native path. PVC reports
CL_DEVICE_DOUBLE_FP_ATOMIC_CAPABILITIES_EXT = 0x70007, so the min/max bits are set, but IGC emulates the operation rather than issuing an instruction, and native measures 1.00x against a cmpxchg loop at every contention level. From an ocloc ISA dump for-device pvc:Sketch of the change:
bitcode/atomicMinMaxFloat_native.clusingatomic_fetch_min_explicitandatomic_fetch_max_explicit, followingatomicAddFloat_native.cl.hasFP32AtomicMinMax()insrc/backend/OpenCL/CHIPBackendOpenCL.hhreading theCL_DEVICE_GLOBAL_FP_ATOMIC_MIN_MAX_EXTandCL_DEVICE_LOCAL_FP_ATOMIC_MIN_MAX_EXTbits ofFp32AtomicAddCapabilities_, which already holds the fullCL_DEVICE_SINGLE_FP_ATOMIC_CAPABILITIES_EXTbitfield, so no newclGetDeviceInfocall is needed.appendRuntimeObjects().__chip_atomic_min_f32import instead of an inline cmpxchg loop.spirvNeedsRtdevlib()insrc/Utils.hhalready matches on__chip_atomic_min.One caveat. rtdevlib is resolved by
clLinkProgram, which fails with-59on Mali-G52, so routingfminandfmaxthrough it would regress Mali unless the inline lowering is kept for devices where the link step is unavailable.Measurements come from a
debugqueue job on Aurora, single tile,ZE_AFFINITY_MASK=0.0, OpenCL backend.