Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c > #402657

Re: Taking hypot(3) To The Edge

From Terje Mathisen <terje.mathisen@tmsw.no>
Newsgroups comp.lang.c, comp.arch
Subject Re: Taking hypot(3) To The Edge
Date 2026-10-03 20:22 +0200
Organization A noiseless patient Spider
Message-ID <119rh5s$38e29$1@dont-email.me> (permalink)
References (2 earlier) <119l4vl$t831$2@dont-email.me> <20261001141212.00002f20@yahoo.com> <119lh8n$14irr$2@dont-email.me> <119otie$2c89q$1@dont-email.me> <20261003210613.00000b2c@yahoo.com>

Cross-posted to 2 groups.

Show all headers | View raw


Michael S wrote:
> On Fri, 2 Oct 2026 20:35:56 +0200
> Terje Mathisen <terje.mathisen@tmsw.no> wrote:
> 
>> David Brown wrote:
>>> On 01/10/2026 13:12, Michael S wrote:
>>>> On Thu, 1 Oct 2026 10:17:57 +0200
>>>> David Brown <david.brown@hesbynett.no> wrote:
>>>>   
>>>>> On 01/10/2026 03:33, bart wrote:
>>>>>> On 01/10/2026 01:47, Lawrence D’Oliveiro wrote:
>>>>>>> A couple of things stand out immediately: one is that the
>>>>>>> “long double” type doesn’t seem to make use of all the
>>>>>>> 128 bits it occupies. I was expecting a mantissa length closer
>>>>>>> to 100 bits, but it’s nowhere near that.
>>>>>>
>>>>>> Probably it uses Intel's 80-bit x87 FPU format. That uses a
>>>>>> 64-bit mantissa (with explicit top bit).
>>>>>
>>>>> Yes, that's the default for "long double" in the standard x86-64
>>>>> ABI.
>>>>>
>>>>> On 32-bit x86, these 10-byte doubles were often stored in 12-byte
>>>>> (96-bit) containers for better alignment, but I don't know the ABI
>>>>> standards here.
>>>>>
>>>>> On 64-bit x86, for better alignment they are stored in 16-byte
>>>>> containers.  The rest of the space will be padding.
>>>>>
>>>>> gcc supports "-mlong-double-64", "-mlong-double-80" and
>>>>> "-mlong-double-128" flags.
>>>>
>>>> The latter appears to be a default on ARM64 Linux.
>>>
>>> That makes sense.  Although software 128-bit floating point is
>>> going to be very slow compared to hardware 64-bit, the reason you
>>> would use "long double" is to get more than "double".
>>
>> Significantly slower, yes.
>>
>> For Mill we figured out that if the HW provided a small number of
>> helper functions (easy to do in HW, much harder in SW), then you can
>> do most ops in maybe 5x the HW cycle count.
>>
>> Without that help, you need to manually extract
>> sign/exp/mantissa(inserting leading bit unless subnormal), then for
>> add/sub you must normalize the smaller number (including sticky bit),
>> do the add, normalize again, then merge with exponent and round.
>>
>> For FMUL you don't pre-normalize (so no extra cycles for subnormal),
>> but you have to handle the potential for up to 111 bits of
>> post-normalization.
>>
>> Michael S is my current benchmark source here, I'd guess FADD128 in
>> less than 40 cycles, about the same for FMUL128, while FMAC128 has to
>> be a little bit harder with a _very_ wide intermediate result.
>>
>> Terje
>>
> 
> Of course, an actual cycle count depends on what you measure, latency
> or throughput, on your CPU, on how many corners you are willing to cut
> w.r.t. support for rounding modes and for FP exceptions and on the ABI.
> Current x86-64 SYSV ABI is quite problematic and rather far from
> well-thought.
> Windows currently has no official FP128 ABI at all, the closest to
> official is the ABI implemented by gcc under msys2. It is even more
> problematic than SysV.
> As far as I am concerned, the most problematic part of both ABIs is
> program status word that is shared with FP32/64. But registers choice
> is also bad.
> According to what I hear from Thomas Koenig, in Fortran quite a few
> corners can be cut without violating language assumptions. For C it
> would be harder.
> With most corners cut, i.e. without non-default rounding modes and
> without support for Inexact exception, and for throughput rather than
> latency Zen3/Linux runs at ~50 clocks per FMUL+FADD. I never tried to
> separate between the two. Would think that they are about the same.
> 
> Pay attention, that nearly all cost of implementing non-default
> rounding mode is because of need to read rounding mode from HW register.
> The same applies to implementing Inexact exception - the main cost is
> update of HW register.

So what you are saying is the my guesstimate was in the right ballpark, 
but that it would be far better for fp128 to be a totally separate 
environment, with all flags and modes maintained in the library instead 
of sharing anything with the HW?

Regarding rounding modes, in my own Mill work a single 64-bit const 
contain all the rounding rules for the four non-truncate alternatives, 
so this didn't really cost much.

Terje


-- 
- <Terje.Mathisen at tmsw.no>
"almost all programming can be viewed as an exercise in caching"

Back to comp.lang.c | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread


Thread

Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-01 00:47 +0000
  Re: Taking hypot(3) To The Edge "Steven G. Kargl" <sgk@REMOVEtroutmask.apl.washington.edu> - 2026-10-01 01:02 +0000
    Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-01 01:55 +0000
      Re: Taking hypot(3) To The Edge "Steven G. Kargl" <sgk@REMOVEtroutmask.apl.washington.edu> - 2026-10-01 03:44 +0000
  Re: Taking hypot(3) To The Edge bart <bc@freeuk.com> - 2026-10-01 02:33 +0100
    Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-01 02:38 +0000
    Re: Taking hypot(3) To The Edge David Brown <david.brown@hesbynett.no> - 2026-10-01 10:17 +0200
      Re: Taking hypot(3) To The Edge Michael S <already5chosen@yahoo.com> - 2026-10-01 14:12 +0300
        Re: Taking hypot(3) To The Edge David Brown <david.brown@hesbynett.no> - 2026-10-01 13:47 +0200
          Re: Taking hypot(3) To The Edge Terje Mathisen <terje.mathisen@tmsw.no> - 2026-10-02 20:35 +0200
            Re: Taking hypot(3) To The Edge bart <bc@freeuk.com> - 2026-10-02 21:17 +0100
            Re: Taking hypot(3) To The Edge BGB <cr88192@gmail.com> - 2026-10-02 17:13 -0500
              Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-03 02:17 +0000
                Re: Taking hypot(3) To The Edge BGB <cr88192@gmail.com> - 2026-10-03 15:20 -0500
                Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-03 21:46 +0000
                Re: Taking hypot(3) To The Edge BGB <cr88192@gmail.com> - 2026-10-03 20:35 -0500
                Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-04 23:54 +0000
            Re: Taking hypot(3) To The Edge Michael S <already5chosen@yahoo.com> - 2026-10-03 21:06 +0300
              Re: Taking hypot(3) To The Edge Terje Mathisen <terje.mathisen@tmsw.no> - 2026-10-03 20:22 +0200
                Re: Taking hypot(3) To The Edge Michael S <already5chosen@yahoo.com> - 2026-10-03 22:01 +0300
                Re: Taking hypot(3) To The Edge Terje Mathisen <terje.mathisen@tmsw.no> - 2026-10-03 23:20 +0200
                Re: Taking hypot(3) To The Edge MitchAlsup <user5857@newsgrouper.org.invalid> - 2026-10-04 17:49 +0000
                Re: Taking hypot(3) To The Edge Thomas Koenig <tkoenig@netcologne.de> - 2026-10-03 19:47 +0000
  Re: Taking hypot(3) To The Edge bart <bc@freeuk.com> - 2026-10-01 23:05 +0100
    Re: Taking hypot(3) To The Edge Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-10-02 02:34 +0000
      Re: Taking hypot(3) To The Edge Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-10-01 21:14 -0700
      Re: Taking hypot(3) To The Edge bart <bc@freeuk.com> - 2026-10-02 11:39 +0100

csiph-web