Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c > #402657
| From | Terje Mathisen <terje.mathisen@tmsw.no> |
|---|---|
| Newsgroups | comp.lang.c, comp.arch |
| Subject | Re: Taking hypot(3) To The Edge |
| Date | 2026-10-03 20:22 +0200 |
| Organization | A noiseless patient Spider |
| Message-ID | <119rh5s$38e29$1@dont-email.me> (permalink) |
| References | (2 earlier) <119l4vl$t831$2@dont-email.me> <20261001141212.00002f20@yahoo.com> <119lh8n$14irr$2@dont-email.me> <119otie$2c89q$1@dont-email.me> <20261003210613.00000b2c@yahoo.com> |
Cross-posted to 2 groups.
Michael S wrote: > On Fri, 2 Oct 2026 20:35:56 +0200 > Terje Mathisen <terje.mathisen@tmsw.no> wrote: > >> David Brown wrote: >>> On 01/10/2026 13:12, Michael S wrote: >>>> On Thu, 1 Oct 2026 10:17:57 +0200 >>>> David Brown <david.brown@hesbynett.no> wrote: >>>> >>>>> On 01/10/2026 03:33, bart wrote: >>>>>> On 01/10/2026 01:47, Lawrence D’Oliveiro wrote: >>>>>>> A couple of things stand out immediately: one is that the >>>>>>> “long double†type doesn’t seem to make use of all the >>>>>>> 128 bits it occupies. I was expecting a mantissa length closer >>>>>>> to 100 bits, but it’s nowhere near that. >>>>>> >>>>>> Probably it uses Intel's 80-bit x87 FPU format. That uses a >>>>>> 64-bit mantissa (with explicit top bit). >>>>> >>>>> Yes, that's the default for "long double" in the standard x86-64 >>>>> ABI. >>>>> >>>>> On 32-bit x86, these 10-byte doubles were often stored in 12-byte >>>>> (96-bit) containers for better alignment, but I don't know the ABI >>>>> standards here. >>>>> >>>>> On 64-bit x86, for better alignment they are stored in 16-byte >>>>> containers. The rest of the space will be padding. >>>>> >>>>> gcc supports "-mlong-double-64", "-mlong-double-80" and >>>>> "-mlong-double-128" flags. >>>> >>>> The latter appears to be a default on ARM64 Linux. >>> >>> That makes sense. Although software 128-bit floating point is >>> going to be very slow compared to hardware 64-bit, the reason you >>> would use "long double" is to get more than "double". >> >> Significantly slower, yes. >> >> For Mill we figured out that if the HW provided a small number of >> helper functions (easy to do in HW, much harder in SW), then you can >> do most ops in maybe 5x the HW cycle count. >> >> Without that help, you need to manually extract >> sign/exp/mantissa(inserting leading bit unless subnormal), then for >> add/sub you must normalize the smaller number (including sticky bit), >> do the add, normalize again, then merge with exponent and round. >> >> For FMUL you don't pre-normalize (so no extra cycles for subnormal), >> but you have to handle the potential for up to 111 bits of >> post-normalization. >> >> Michael S is my current benchmark source here, I'd guess FADD128 in >> less than 40 cycles, about the same for FMUL128, while FMAC128 has to >> be a little bit harder with a _very_ wide intermediate result. >> >> Terje >> > > Of course, an actual cycle count depends on what you measure, latency > or throughput, on your CPU, on how many corners you are willing to cut > w.r.t. support for rounding modes and for FP exceptions and on the ABI. > Current x86-64 SYSV ABI is quite problematic and rather far from > well-thought. > Windows currently has no official FP128 ABI at all, the closest to > official is the ABI implemented by gcc under msys2. It is even more > problematic than SysV. > As far as I am concerned, the most problematic part of both ABIs is > program status word that is shared with FP32/64. But registers choice > is also bad. > According to what I hear from Thomas Koenig, in Fortran quite a few > corners can be cut without violating language assumptions. For C it > would be harder. > With most corners cut, i.e. without non-default rounding modes and > without support for Inexact exception, and for throughput rather than > latency Zen3/Linux runs at ~50 clocks per FMUL+FADD. I never tried to > separate between the two. Would think that they are about the same. > > Pay attention, that nearly all cost of implementing non-default > rounding mode is because of need to read rounding mode from HW register. > The same applies to implementing Inexact exception - the main cost is > update of HW register. So what you are saying is the my guesstimate was in the right ballpark, but that it would be far better for fp128 to be a totally separate environment, with all flags and modes maintained in the library instead of sharing anything with the HW? Regarding rounding modes, in my own Mill work a single 64-bit const contain all the rounding rules for the four non-truncate alternatives, so this didn't really cost much. Terje -- - <Terje.Mathisen at tmsw.no> "almost all programming can be viewed as an exercise in caching"
Back to comp.lang.c | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-01 00:47 +0000
Re: Taking hypot(3) To The Edge "Steven G. Kargl" <sgk@REMOVEtroutmask.apl.washington.edu> - 2026-10-01 01:02 +0000
Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-01 01:55 +0000
Re: Taking hypot(3) To The Edge "Steven G. Kargl" <sgk@REMOVEtroutmask.apl.washington.edu> - 2026-10-01 03:44 +0000
Re: Taking hypot(3) To The Edge bart <bc@freeuk.com> - 2026-10-01 02:33 +0100
Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-01 02:38 +0000
Re: Taking hypot(3) To The Edge David Brown <david.brown@hesbynett.no> - 2026-10-01 10:17 +0200
Re: Taking hypot(3) To The Edge Michael S <already5chosen@yahoo.com> - 2026-10-01 14:12 +0300
Re: Taking hypot(3) To The Edge David Brown <david.brown@hesbynett.no> - 2026-10-01 13:47 +0200
Re: Taking hypot(3) To The Edge Terje Mathisen <terje.mathisen@tmsw.no> - 2026-10-02 20:35 +0200
Re: Taking hypot(3) To The Edge bart <bc@freeuk.com> - 2026-10-02 21:17 +0100
Re: Taking hypot(3) To The Edge BGB <cr88192@gmail.com> - 2026-10-02 17:13 -0500
Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-03 02:17 +0000
Re: Taking hypot(3) To The Edge BGB <cr88192@gmail.com> - 2026-10-03 15:20 -0500
Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-03 21:46 +0000
Re: Taking hypot(3) To The Edge BGB <cr88192@gmail.com> - 2026-10-03 20:35 -0500
Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-04 23:54 +0000
Re: Taking hypot(3) To The Edge Michael S <already5chosen@yahoo.com> - 2026-10-03 21:06 +0300
Re: Taking hypot(3) To The Edge Terje Mathisen <terje.mathisen@tmsw.no> - 2026-10-03 20:22 +0200
Re: Taking hypot(3) To The Edge Michael S <already5chosen@yahoo.com> - 2026-10-03 22:01 +0300
Re: Taking hypot(3) To The Edge Terje Mathisen <terje.mathisen@tmsw.no> - 2026-10-03 23:20 +0200
Re: Taking hypot(3) To The Edge MitchAlsup <user5857@newsgrouper.org.invalid> - 2026-10-04 17:49 +0000
Re: Taking hypot(3) To The Edge Thomas Koenig <tkoenig@netcologne.de> - 2026-10-03 19:47 +0000
Re: Taking hypot(3) To The Edge bart <bc@freeuk.com> - 2026-10-01 23:05 +0100
Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-02 02:34 +0000
Re: Taking hypot(3) To The Edge Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-10-01 21:14 -0700
Re: Taking hypot(3) To The Edge bart <bc@freeuk.com> - 2026-10-02 11:39 +0100
csiph-web