Skip to content

Commit 07e646c

Browse files
committed
GPU: give Metal its own spellings for two math helpers
GPUCA_CHOICE routes Metal down the OpenCL arm, and three of those spellings do not exist in MSL. One of the three needs no Metal case: MSL has no nan() in any form, but it does define NAN, and so does OpenCL -- a constant expression of type float representing a quiet NaN -- so the shared arm uses that instead of nan(0u). remainder() does not exist either, and __builtin_remainderf is not a way around it: that compiles, then fails to link against an undefined remainderf. MSL does have fmod, which is exact, so Remainderf reduces with that and nudges the result into [-|y|/2, |y|/2], rounding the quotient of a tie to even. It agrees with remainderf bit for bit over six million samples, from the TwoPI wrapping its only caller does up to |x| of 1e30, ties swept exhaustively for a power-of-two divisor. Reducing as x - y * rint(x / y) would have been shorter, but x / y rounds in float, so rint picks the wrong multiple outside a narrow range around the divisor and the result is wrong almost everywhere else. MSL's sincos returns the sine and takes the cosine by thread reference rather than by pointer, and cannot write through the generic reference SinCos is given, so the result goes via a local. Host, CUDA, HIP and cling keep the GPUCA_CHOICE arms they had; OpenCL swaps nan(0u) for NAN.
1 parent c0d5eb7 commit 07e646c

1 file changed

Lines changed: 21 additions & 2 deletions

File tree

‎GPU/Common/GPUCommonMath.h‎

Lines changed: 21 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -105,7 +105,7 @@ class GPUCommonMath
105105
GPUd() constexpr static bool Finite(float x);
106106
GPUd() constexpr static bool IsNaN(float x);
107107
#ifndef __FAST_MATH__
108-
GPUd() constexpr static float QuietNaN() { return GPUCA_CHOICE(std::numeric_limits<float>::quiet_NaN(), __builtin_nanf(""), nan(0u)); }
108+
GPUd() constexpr static float QuietNaN() { return GPUCA_CHOICE(std::numeric_limits<float>::quiet_NaN(), __builtin_nanf(""), NAN); }
109109
#endif
110110
GPUd() constexpr static uint32_t Clz(uint32_t val);
111111
GPUd() constexpr static uint32_t Ctz(uint32_t val);
@@ -245,7 +245,21 @@ GPUdi() float2 GPUCommonMath::MakeFloat2(float x, float y)
245245
}
246246

247247
GPUdi() constexpr float GPUCommonMath::Modf(float x, float y) { return GPUCA_CHOICE(fmodf(x, y), fmodf(x, y), fmod(x, y)); }
248-
GPUhdi() float GPUCommonMath::Remainderf(float x, float y) { return GPUCA_CHOICE(std::remainderf(x, y), remainderf(x, y), remainder(x, y)); }
248+
GPUhdi() float GPUCommonMath::Remainderf(float x, float y)
249+
{
250+
#ifdef __METAL__ // MSL has no remainder(), so reduce with fmod, which is exact
251+
const float r = fmod(x, y), h = 0.5f * fabs(y), a = fabs(r);
252+
if (a > h) {
253+
return r - copysign(y, r);
254+
}
255+
if (a == h && fmod((x - r) / y, 2.f) != 0.f) { // the quotient of a tie is rounded to even
256+
return r - copysign(y, r);
257+
}
258+
return r;
259+
#else
260+
return GPUCA_CHOICE(std::remainderf(x, y), remainderf(x, y), remainder(x, y));
261+
#endif
262+
}
249263

250264
GPUdi() uint32_t GPUCommonMath::Float2UIntReint(const float& x)
251265
{
@@ -302,6 +316,11 @@ GPUhdi() void GPUCommonMath::SinCos(float x, float& s, float& c)
302316
__sincosf(x, &s, &c);
303317
#elif !defined(GPUCA_GPUCODE_DEVICE) && (defined(__GNU_SOURCE__) || defined(_GNU_SOURCE) || defined(GPUCA_GPUCODE))
304318
sincosf(x, &s, &c);
319+
#elif defined(__METAL__) // MSL's sincos returns sin and takes cos by thread
320+
// reference, so it cannot write through a generic one
321+
float metalCos;
322+
s = sincos(x, metalCos);
323+
c = metalCos;
305324
#else
306325
GPUCA_CHOICE((void)((s = sinf(x)) + (c = cosf(x))), sincosf(x, &s, &c), s = sincos(x, &c));
307326
#endif

0 commit comments

Comments
 (0)