| Age | Commit message (Collapse) | Author | Lines |
|
the new code uses that for all real |x| in [0x1p-999,0x1p999]
y = (float)x
is the same as
r = (double)x
t = x - r
if (t!=0 && r.bits%2==0)
r.bits += (r<0)==(t<0) ? 1 : -1
y = (float)r
in all rounding modes, with the same fenv effects.
this can be interpreted as a round to odd adjustment[1].
in fmaf t is computed with a fast2sum variant.
- optimized common case.
- implicit uflow handling instead of fenv calls.
- no fegetround check.
- no api calls on hf targets.
- fixed missed uflow on targets that signal it before rounding:
fmaf(-0x1p-100f, 0x1p-100f, 0x1p-126f)
- removed freebsd code references and comments.
- fmaf.o code size vs before the halfway subnormal fix:
x86_64: 514 -> 210
armhf: 304 -> 156 (v7 thumb, no vfma op)
arm: 400 -> 348 (soft float, no fenv)
round to odd paper (suggested by Sergey Davidoff):
[1] S. Boldo et al., Emulation of a FMA and correctly-rounded sums:
proved algorithms using rounding to odd, 2008
|
|
inexact halfway cases were not handled correctly for subnormals
fmaf(0x20201p-92f, 0x1fe01p-92f, 0x1p-130f)
= (float)(0x1p-130 + 0x1.000000004p-150)
was rounded to 0x1.00001p-130 first in double precision, then to
0x1p-130 in the float subnormal range instead of 0x1.00002p-130.
this is a minimal fix of the halfway check.
Reported-by: Sergey Davidoff <shnatsel@gmail.com>
|
|
|
|
if double precision r=x*y+z is not a half way case between two single
precision floats or it is an exact result then fmaf returns (float)r.
however the exactness check was wrong when |x*y| < |z| and could cause
incorrectly rounded result in nearest rounding mode when r is a half
way case.
fmaf(-0x1.26524ep-54, -0x1.cb7868p+11, 0x1.d10f5ep-29)
was incorrectly rounded up to 0x1.d117ap-29 instead of 0x1.d1179ep-29.
(exact result is 0x1.d1179efffffffecp-29, r is 0x1.d1179fp-29)
|
|
the issue is described in commits 1e5eb73545ca6cfe8b918798835aaf6e07af5beb
and ffd8ac2dd50f99c3c83d7d9d845df9874ec3e7d5
|
|
The underflow exception is not raised correctly in some
cornercases (see previous fma commit), added comments
with examples for fmaf, fmal and non-x86 fma.
In fmaf store the result before returning so it has the
correct precision when FLT_EVAL_METHOD!=0
|
|
|
|
this is necessary to support archs where fenv is incomplete or
unavailable (presently arm). fma, fmal, and the lrint family should
work perfectly fine with this change; fmaf is slightly broken with
respect to rounding as it depends on non-default rounding modes to do
its work.
|
|
thanks to the hard work of Szabolcs Nagy (nsz), identifying the best
(from correctness and license standpoint) implementations from freebsd
and openbsd and cleaning them up! musl should now fully support c99
float and long double math functions, and has near-complete complex
math support. tgmath should also work (fully on gcc-compatible
compilers, and mostly on any c99 compiler).
based largely on commit 0376d44a890fea261506f1fc63833e7a686dca19 from
nsz's libm git repo, with some additions (dummy versions of a few
missing long double complex functions, etc.) by me.
various cleanups still need to be made, including re-adding (if
they're correct) some asm functions that were dropped.
|