Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c++ > #84173 > unrolled thread

Constexpr evaluation of heavy stuff

Started byAndrey Tarasevich <andreytarasevich@hotmail.com>
First post2022-05-18 16:17 -0700
Last post2022-05-19 18:33 +0200
Articles 19 on this page of 39 — 14 participants

Back to article view | Back to comp.lang.c++


Contents

  Constexpr evaluation of heavy stuff Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-05-18 16:17 -0700
    Re: Constexpr evaluation of heavy stuff "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-05-19 09:19 +0200
      Re: Constexpr evaluation of heavy stuff David Brown <david.brown@hesbynett.no> - 2022-05-19 11:59 +0200
        Re: Constexpr evaluation of heavy stuff "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-05-19 12:37 +0200
          Re: Constexpr evaluation of heavy stuff David Brown <david.brown@hesbynett.no> - 2022-05-19 14:49 +0200
            Re: Constexpr evaluation of heavy stuff Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-05-19 06:19 -0700
              Re: Constexpr evaluation of heavy stuff Paavo Helde <eesnimi@osa.pri.ee> - 2022-05-19 20:32 +0300
                Re: Constexpr evaluation of heavy stuff Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-05-20 01:12 -0700
                  Re: Constexpr evaluation of heavy stuff Juha Nieminen <nospam@thanks.invalid> - 2022-05-20 08:37 +0000
                    Re: Constexpr evaluation of heavy stuff Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-05-20 02:21 -0700
                      Re: Constexpr evaluation of heavy stuff Ben <ben.usenet@bsb.me.uk> - 2022-05-20 11:25 +0100
                        Re: Constexpr evaluation of heavy stuff David Brown <david.brown@hesbynett.no> - 2022-05-20 12:47 +0200
                          Re: Constexpr evaluation of heavy stuff Ben <ben.usenet@bsb.me.uk> - 2022-05-20 12:13 +0100
                            Re: Constexpr evaluation of heavy stuff David Brown <david.brown@hesbynett.no> - 2022-05-20 14:27 +0200
                            Re: Constexpr evaluation of heavy stuff Öö Tiib <ootiib@hot.ee> - 2022-05-20 06:33 -0700
                              Re: Constexpr evaluation of heavy stuff Paavo Helde <eesnimi@osa.pri.ee> - 2022-05-20 17:15 +0300
                                Re: Constexpr evaluation of heavy stuff Christian Gollwitzer <auriocus@gmx.de> - 2022-05-20 18:57 +0200
                                  Re: Constexpr evaluation of heavy stuff Paavo Helde <eesnimi@osa.pri.ee> - 2022-05-20 22:31 +0300
                                    Re: Constexpr evaluation of heavy stuff Manfred <noname@add.invalid> - 2022-05-21 21:56 +0200
                                      Re: Constexpr evaluation of heavy stuff Öö Tiib <ootiib@hot.ee> - 2022-05-21 23:24 -0700
                                        Re: Constexpr evaluation of heavy stuff Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-05-22 04:56 -0700
                                          Re: Constexpr evaluation of heavy stuff Öö Tiib <ootiib@hot.ee> - 2022-05-22 06:15 -0700
                                        Re: Constexpr evaluation of heavy stuff Manfred <noname@add.invalid> - 2022-05-22 21:46 +0200
                                          Re: Constexpr evaluation of heavy stuff "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-05-22 23:23 +0200
                                            Re: Constexpr evaluation of heavy stuff Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-05-22 16:26 -0700
                                            Re: Constexpr evaluation of heavy stuff Manfred <noname@add.invalid> - 2022-05-23 17:24 +0200
                                            Re: Constexpr evaluation of heavy stuff Richard Damon <Richard@Damon-Family.org> - 2022-05-23 19:41 -0400
                                              Re: Constexpr evaluation of heavy stuff scott@slp53.sl.home (Scott Lurndal) - 2022-05-24 16:15 +0000
                                            Re: Constexpr evaluation of heavy stuff Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-05-24 05:37 -0700
                        Re: Constexpr evaluation of heavy stuff Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-05-20 04:19 -0700
                          Re: Constexpr evaluation of heavy stuff Ben <ben.usenet@bsb.me.uk> - 2022-05-20 12:32 +0100
                      Re: Constexpr evaluation of heavy stuff Juha Nieminen <nospam@thanks.invalid> - 2022-05-20 10:40 +0000
                    Re: Constexpr evaluation of heavy stuff David Brown <david.brown@hesbynett.no> - 2022-05-20 12:31 +0200
                      Re: Constexpr evaluation of heavy stuff Juha Nieminen <nospam@thanks.invalid> - 2022-05-20 10:43 +0000
                        Re: Constexpr evaluation of heavy stuff David Brown <david.brown@hesbynett.no> - 2022-05-20 14:52 +0200
            Re: Constexpr evaluation of heavy stuff "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-05-19 15:34 +0200
      Re: Constexpr evaluation of heavy stuff "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-05-19 17:03 +0200
    Re: Constexpr evaluation of heavy stuff Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-05-19 08:53 -0700
      Re: Constexpr evaluation of heavy stuff Marcel Mueller <news.5.maazl@spamgourmet.org> - 2022-05-19 18:33 +0200

Page 2 of 2 — ← Prev page 1 [2]


#84236

FromMalcolm McLean <malcolm.arthur.mclean@gmail.com>
Date2022-05-22 04:56 -0700
Message-ID<1d635773-de4d-451e-ae65-acecddb7839en@googlegroups.com>
In reply to#84235
On Sunday, 22 May 2022 at 07:24:40 UTC+1, Öö Tiib wrote:
> On Saturday, 21 May 2022 at 22:56:22 UTC+3, Manfred wrote: 
> > On 5/20/2022 9:31 PM, Paavo Helde wrote: 
> > > 20.05.2022 19:57 Christian Gollwitzer kirjutas: 
> > >> Am 20.05.22 um 16:15 schrieb Paavo Helde: 
> > >>> 20.05.2022 16:33 Öö Tiib kirjutas: 
> > >>> 
> > >>>> Modular arithmetic around likes of 42949672956 however is typically 
> > >>>> utterly useless. 
> > >>> I agree in general, but I have found one use case. Suppose we have a 
> > >>> string s which might or might not contain a separator like ':' and we 
> > >>> want to either return the part after separator, or the whole string 
> > >>> if there is no separator. This is the code: 
> > >>> 
> > >>> return s.substr(s.find(':')+1); 
> > >>> 
> > >> And where exactly do you need mod "ridiculous number" here? WHy is 
> > >> that better than having a signed index variable, where negative values 
> > >> indicate invalid indices? 
> > > 
> > > You are right, if std::string::find() returned a signed integer -1 for 
> > > no-find, this code would work just as fine. The only thing is that it 
> > > doesn't work that way, it returns size_t(-1) which is typically 
> > > 18446744073709551615ULL nowadays. 
> > But your example is still valid: size_t(-1) + 1 yields the correct value 
> > because of unsigned wrapping, which is defined behaviour. For the rest, 
> > using an unsigned number as an index inside a string makes sense, 
> > especially in the standard library - a signed type would waste half its 
> > range for this use.
> The history shows that either half of that range is enough or way too short. 
> So where 7 bit ASCII is not enough there usage of 16 bit UCS-2 can become 
> obsolete decade later.
>
Sizes don't have a linear distribution. Whilst it would be hard to work out exactly
what the distribution of the sizes of arrays in programs is, it would likely be
a log normal distribution or something similar. So whilst losing a bit halves
your range, it doesn't halve the number of arrays you can index.

[toc] | [prev] | [next] | [standalone]


#84237

FromÖö Tiib <ootiib@hot.ee>
Date2022-05-22 06:15 -0700
Message-ID<2de7ce52-9117-4ec4-898a-0c3e72e88fb1n@googlegroups.com>
In reply to#84236
On Sunday, 22 May 2022 at 14:56:55 UTC+3, Malcolm McLean wrote:
> On Sunday, 22 May 2022 at 07:24:40 UTC+1, Öö Tiib wrote: 
> > On Saturday, 21 May 2022 at 22:56:22 UTC+3, Manfred wrote: 
> > > On 5/20/2022 9:31 PM, Paavo Helde wrote: 
> > > > 20.05.2022 19:57 Christian Gollwitzer kirjutas: 
> > > >> Am 20.05.22 um 16:15 schrieb Paavo Helde: 
> > > >>> 20.05.2022 16:33 Öö Tiib kirjutas: 
> > > >>> 
> > > >>>> Modular arithmetic around likes of 42949672956 however is typically 
> > > >>>> utterly useless. 
> > > >>> I agree in general, but I have found one use case. Suppose we have a 
> > > >>> string s which might or might not contain a separator like ':' and we 
> > > >>> want to either return the part after separator, or the whole string 
> > > >>> if there is no separator. This is the code: 
> > > >>> 
> > > >>> return s.substr(s.find(':')+1); 
> > > >>> 
> > > >> And where exactly do you need mod "ridiculous number" here? WHy is 
> > > >> that better than having a signed index variable, where negative values 
> > > >> indicate invalid indices? 
> > > > 
> > > > You are right, if std::string::find() returned a signed integer -1 for 
> > > > no-find, this code would work just as fine. The only thing is that it 
> > > > doesn't work that way, it returns size_t(-1) which is typically 
> > > > 18446744073709551615ULL nowadays. 
> > > But your example is still valid: size_t(-1) + 1 yields the correct value 
> > > because of unsigned wrapping, which is defined behaviour. For the rest, 
> > > using an unsigned number as an index inside a string makes sense, 
> > > especially in the standard library - a signed type would waste half its 
> > > range for this use. 
> > The history shows that either half of that range is enough or way too short. 
> > So where 7 bit ASCII is not enough there usage of 16 bit UCS-2 can become 
> > obsolete decade later. 
> >
> Sizes don't have a linear distribution. Whilst it would be hard to work out exactly 
> what the distribution of the sizes of arrays in programs is, it would likely be 
> a log normal distribution or something similar. So whilst losing a bit halves 
> your range, it doesn't halve the number of arrays you can index.

Yes. Quite common are arrays that grow never to thousands of elements.
However as our generic dynamic array  (std::vector) is pointer and two sizes
we can't reduce most of that actual waste. We can't get any benefit from
going lower than one 64 bit pointer and 2 32 bit sizes. Compiler would just
add unused padding to end of it because of alignment requirement of pointer.
So the bits "wasted" as signs to sizes ... are already idle anyway.

[toc] | [prev] | [next] | [standalone]


#84239

FromManfred <noname@add.invalid>
Date2022-05-22 21:46 +0200
Message-ID<t6e3u2$d0l$1@gioia.aioe.org>
In reply to#84235
On 5/22/2022 8:24 AM, Öö Tiib wrote:
> On Saturday, 21 May 2022 at 22:56:22 UTC+3, Manfred wrote:
>> On 5/20/2022 9:31 PM, Paavo Helde wrote:
>>> 20.05.2022 19:57 Christian Gollwitzer kirjutas:
>>>> Am 20.05.22 um 16:15 schrieb Paavo Helde:
>>>>> 20.05.2022 16:33 Öö Tiib kirjutas:
>>>>>
>>>>>> Modular arithmetic around likes of 42949672956 however is typically
>>>>>> utterly useless.
>>>>> I agree in general, but I have found one use case. Suppose we have a
>>>>> string s which might or might not contain a separator like ':' and we
>>>>> want to either return the part after separator, or the whole string
>>>>> if there is no separator. This is the code:
>>>>>
>>>>> return s.substr(s.find(':')+1);
>>>>>
>>>> And where exactly do you need mod "ridiculous number" here? WHy is
>>>> that better than having a signed index variable, where negative values
>>>> indicate invalid indices?
>>>
>>> You are right, if std::string::find() returned a signed integer -1 for
>>> no-find, this code would work just as fine. The only thing is that it
>>> doesn't work that way, it returns size_t(-1) which is typically
>>> 18446744073709551615ULL nowadays.
>> But your example is still valid: size_t(-1) + 1 yields the correct value
>> because of unsigned wrapping, which is defined behaviour. For the rest,
>> using an unsigned number as an index inside a string makes sense,
>> especially in the standard library - a signed type would waste half its
>> range for this use.
> 
> The history shows that either half of that range is enough or way too short.
> So where 7 bit ASCII is not enough there usage of 16 bit UCS-2 can become
> obsolete decade later.

Your example about character sets is wrong, however I may agree that 
your argument about size distribution is often valid for application 
development - or rather most of it.
But, as I wrote, this is part of the standard library, where efficiency 
is a must.

So, no, gratuitous waste of half the range is not acceptable in the 
standard library.

Especially in a case like this, where there is a trivial solution to 
that, which by the way is the one adopted by the standard.

[toc] | [prev] | [next] | [standalone]


#84240

From"Alf P. Steinbach" <alf.p.steinbach@gmail.com>
Date2022-05-22 23:23 +0200
Message-ID<t6e9ki$n4a$1@dont-email.me>
In reply to#84239
On 22 May 2022 21:46, Manfred wrote:
> On 5/22/2022 8:24 AM, Öö Tiib wrote:
>> On Saturday, 21 May 2022 at 22:56:22 UTC+3, Manfred wrote:
>>> On 5/20/2022 9:31 PM, Paavo Helde wrote:
>>>> 20.05.2022 19:57 Christian Gollwitzer kirjutas:
>>>>> Am 20.05.22 um 16:15 schrieb Paavo Helde:
>>>>>> 20.05.2022 16:33 Öö Tiib kirjutas:
>>>>>>
>>>>>>> Modular arithmetic around likes of 42949672956 however is typically
>>>>>>> utterly useless.
>>>>>> I agree in general, but I have found one use case. Suppose we have a
>>>>>> string s which might or might not contain a separator like ':' and we
>>>>>> want to either return the part after separator, or the whole string
>>>>>> if there is no separator. This is the code:
>>>>>>
>>>>>> return s.substr(s.find(':')+1);
>>>>>>
>>>>> And where exactly do you need mod "ridiculous number" here? WHy is
>>>>> that better than having a signed index variable, where negative values
>>>>> indicate invalid indices?
>>>>
>>>> You are right, if std::string::find() returned a signed integer -1 for
>>>> no-find, this code would work just as fine. The only thing is that it
>>>> doesn't work that way, it returns size_t(-1) which is typically
>>>> 18446744073709551615ULL nowadays.
>>> But your example is still valid: size_t(-1) + 1 yields the correct value
>>> because of unsigned wrapping, which is defined behaviour. For the rest,
>>> using an unsigned number as an index inside a string makes sense,
>>> especially in the standard library - a signed type would waste half its
>>> range for this use.
>>
>> The history shows that either half of that range is enough or way too 
>> short.
>> So where 7 bit ASCII is not enough there usage of 16 bit UCS-2 can become
>> obsolete decade later.
> 
> Your example about character sets is wrong, however I may agree that 
> your argument about size distribution is often valid for application 
> development - or rather most of it.
> But, as I wrote, this is part of the standard library, where efficiency 
> is a must.
> 
> So, no, gratuitous waste of half the range is not acceptable in the 
> standard library.
> 
> Especially in a case like this, where there is a trivial solution to 
> that, which by the way is the one adopted by the standard.

That was (and for embedded possibly still is) in support of 16-bit systems.

And that support is half-baked with the pointer difference type being 
signed, and therefore required to be at least 17 bits (not joking here).

When you have an array of bytes that fills up more than half the address 
range (not very common) and you feel the urgent need to treat it 
code-wise as just any other array, nothing special, no ma!, well, dumbness.

Of course it's OK to not use half of the possible size range for the 
direct size representation (it comes into play in other contexts). As a 
concrete example, it's decidedly OK on a 64-bit system. And as another 
concrete example, it was OK in 32-bit Windows, where half the range was 
all you had available anyway without very special configuration, which I 
believe was quite risky. It can be more difficult to argue in favor of 
unsigned sizes for the cases of 32-bit Linux and 16-bit embedded. But 
the former is a hopefully soon extinct beast, and the latter requires 
special considerations anyway, not just unthinking use of vanilla C++.

I gather that that reality is very much part of why the language 
designer Bjarne Stroustrup says it's OK with that apparent/alleged range 
waste, and a mistake to use unsigned for sizes. For example, Bjarne is 
one of the three main authors of the C++ Core Guidelines that recommends 
signed types for numbers. He's even written a paper titled "Subscripts 
and sizes should be signed", available at <url: 
https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p1428r0.pdf>; 
please read it.



Cheers,

- Alf

[toc] | [prev] | [next] | [standalone]


#84241

FromMalcolm McLean <malcolm.arthur.mclean@gmail.com>
Date2022-05-22 16:26 -0700
Message-ID<4980e2c4-5202-4b17-ab5b-9a82f522f226n@googlegroups.com>
In reply to#84240
On Sunday, 22 May 2022 at 22:23:47 UTC+1, alf.p.s...@gmail.com wrote:
> 
> I gather that that reality is very much part of why the language 
> designer Bjarne Stroustrup says it's OK with that apparent/alleged range 
> waste, and a mistake to use unsigned for sizes. For example, Bjarne is 
> one of the three main authors of the C++ Core Guidelines that recommends 
> signed types for numbers. He's even written a paper titled "Subscripts 
> and sizes should be signed", available at <url: 
> https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p1428r0.pdf>; 
> please read it. 
> 
I've been saying the same thing for a long time. The root of the mistake was that 
sizeof returned a size_t. So you could support sizeof(massive_array). The 
cost of that was to inflict mixed mode integer arithemetic  on the whole
of the rest of the language. It's worse for C++ than for C because of the way
that C++ uses types for symbol resolution.

[toc] | [prev] | [next] | [standalone]


#84242

FromManfred <noname@add.invalid>
Date2022-05-23 17:24 +0200
Message-ID<t6g901$5ed$1@gioia.aioe.org>
In reply to#84240
On 5/22/2022 11:23 PM, Alf P. Steinbach wrote:
> On 22 May 2022 21:46, Manfred wrote:
>> On 5/22/2022 8:24 AM, Öö Tiib wrote:
>>> On Saturday, 21 May 2022 at 22:56:22 UTC+3, Manfred wrote:
>>>> On 5/20/2022 9:31 PM, Paavo Helde wrote:
>>>>> 20.05.2022 19:57 Christian Gollwitzer kirjutas:
>>>>>> Am 20.05.22 um 16:15 schrieb Paavo Helde:
>>>>>>> 20.05.2022 16:33 Öö Tiib kirjutas:
>>>>>>>
>>>>>>>> Modular arithmetic around likes of 42949672956 however is typically
>>>>>>>> utterly useless.
>>>>>>> I agree in general, but I have found one use case. Suppose we have a
>>>>>>> string s which might or might not contain a separator like ':' 
>>>>>>> and we
>>>>>>> want to either return the part after separator, or the whole string
>>>>>>> if there is no separator. This is the code:
>>>>>>>
>>>>>>> return s.substr(s.find(':')+1);
>>>>>>>
>>>>>> And where exactly do you need mod "ridiculous number" here? WHy is
>>>>>> that better than having a signed index variable, where negative 
>>>>>> values
>>>>>> indicate invalid indices?
>>>>>
>>>>> You are right, if std::string::find() returned a signed integer -1 for
>>>>> no-find, this code would work just as fine. The only thing is that it
>>>>> doesn't work that way, it returns size_t(-1) which is typically
>>>>> 18446744073709551615ULL nowadays.
>>>> But your example is still valid: size_t(-1) + 1 yields the correct 
>>>> value
>>>> because of unsigned wrapping, which is defined behaviour. For the rest,
>>>> using an unsigned number as an index inside a string makes sense,
>>>> especially in the standard library - a signed type would waste half its
>>>> range for this use.
>>>
>>> The history shows that either half of that range is enough or way too 
>>> short.
>>> So where 7 bit ASCII is not enough there usage of 16 bit UCS-2 can 
>>> become
>>> obsolete decade later.
>>
>> Your example about character sets is wrong, however I may agree that 
>> your argument about size distribution is often valid for application 
>> development - or rather most of it.
>> But, as I wrote, this is part of the standard library, where 
>> efficiency is a must.
>>
>> So, no, gratuitous waste of half the range is not acceptable in the 
>> standard library.
>>
>> Especially in a case like this, where there is a trivial solution to 
>> that, which by the way is the one adopted by the standard.
> 
> That was (and for embedded possibly still is) in support of 16-bit systems.
> 
> And that support is half-baked with the pointer difference type being 
> signed, and therefore required to be at least 17 bits (not joking here).
> 

True, pointer differences are one of the problems here.

> When you have an array of bytes that fills up more than half the address 
> range (not very common) and you feel the urgent need to treat it 
> code-wise as just any other array, nothing special, no ma!, well, dumbness.

Well, obviously it wouldn't be like "just any other array". However, I 
wouldn't call "dumbness" the possibility of using /some/ standard 
library features (like find()) with it.

> 
> Of course it's OK to not use half of the possible size range for the 
> direct size representation (it comes into play in other contexts). As a 
> concrete example, it's decidedly OK on a 64-bit system. And as another 
> concrete example, it was OK in 32-bit Windows, where half the range was 
> all you had available anyway without very special configuration, which I 
> believe was quite risky. It can be more difficult to argue in favor of 
> unsigned sizes for the cases of 32-bit Linux and 16-bit embedded. But 
> the former is a hopefully soon extinct beast, and the latter requires 
> special considerations anyway, not just unthinking use of vanilla C++.
> 

"Unthinking" is obviously out of the picture here - if you are 
programming a system where most of the available memory is dedicated to 
one text buffer, clearly you have to think about how you are using it.
"Special considerations", however, do not necessarily mean that you 
happily discard all of the standard library features with its use.

Just to be clear, having such a system is a rare event. However, since 
you mention embedded, having tight hardware constraints on these 
architectures is not that rare an event.

(Incidentally, I still happily use a 32-bit Linux laptop)

> I gather that that reality is very much part of why the language 
> designer Bjarne Stroustrup says it's OK with that apparent/alleged range 
> waste, and a mistake to use unsigned for sizes. For example, Bjarne is 
> one of the three main authors of the C++ Core Guidelines that recommends 
> signed types for numbers. He's even written a paper titled "Subscripts 
> and sizes should be signed", available at <url: 
> https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p1428r0.pdf>; 
> please read it.
> 

Read it, thanks for sharing.
I find that the most convincing point is about signed/unsigned mixed 
arithmetic. Which is a good reason to introduce ssize_t.
Still, throwing away size_t entirely? Hmm, still dubious.

Again, we are talking about the standard library here, which is supposed 
to run /everywhere/, and serve a very broad spectrum of goals. It is 
inherently different from developing any specific application.

> 
> 
> Cheers,
> 
> - Alf

[toc] | [prev] | [next] | [standalone]


#84244

FromRichard Damon <Richard@Damon-Family.org>
Date2022-05-23 19:41 -0400
Message-ID<IYUiK.1428$x7oc.412@fx01.iad>
In reply to#84240
On 5/22/22 5:23 PM, Alf P. Steinbach wrote:
> On 22 May 2022 21:46, Manfred wrote:
>> On 5/22/2022 8:24 AM, Öö Tiib wrote:
>>> On Saturday, 21 May 2022 at 22:56:22 UTC+3, Manfred wrote:
>>>> On 5/20/2022 9:31 PM, Paavo Helde wrote:
>>>>> 20.05.2022 19:57 Christian Gollwitzer kirjutas:
>>>>>> Am 20.05.22 um 16:15 schrieb Paavo Helde:
>>>>>>> 20.05.2022 16:33 Öö Tiib kirjutas:
>>>>>>>
>>>>>>>> Modular arithmetic around likes of 42949672956 however is typically
>>>>>>>> utterly useless.
>>>>>>> I agree in general, but I have found one use case. Suppose we have a
>>>>>>> string s which might or might not contain a separator like ':' 
>>>>>>> and we
>>>>>>> want to either return the part after separator, or the whole string
>>>>>>> if there is no separator. This is the code:
>>>>>>>
>>>>>>> return s.substr(s.find(':')+1);
>>>>>>>
>>>>>> And where exactly do you need mod "ridiculous number" here? WHy is
>>>>>> that better than having a signed index variable, where negative 
>>>>>> values
>>>>>> indicate invalid indices?
>>>>>
>>>>> You are right, if std::string::find() returned a signed integer -1 for
>>>>> no-find, this code would work just as fine. The only thing is that it
>>>>> doesn't work that way, it returns size_t(-1) which is typically
>>>>> 18446744073709551615ULL nowadays.
>>>> But your example is still valid: size_t(-1) + 1 yields the correct 
>>>> value
>>>> because of unsigned wrapping, which is defined behaviour. For the rest,
>>>> using an unsigned number as an index inside a string makes sense,
>>>> especially in the standard library - a signed type would waste half its
>>>> range for this use.
>>>
>>> The history shows that either half of that range is enough or way too 
>>> short.
>>> So where 7 bit ASCII is not enough there usage of 16 bit UCS-2 can 
>>> become
>>> obsolete decade later.
>>
>> Your example about character sets is wrong, however I may agree that 
>> your argument about size distribution is often valid for application 
>> development - or rather most of it.
>> But, as I wrote, this is part of the standard library, where 
>> efficiency is a must.
>>
>> So, no, gratuitous waste of half the range is not acceptable in the 
>> standard library.
>>
>> Especially in a case like this, where there is a trivial solution to 
>> that, which by the way is the one adopted by the standard.
> 
> That was (and for embedded possibly still is) in support of 16-bit systems.
> 
> And that support is half-baked with the pointer difference type being 
> signed, and therefore required to be at least 17 bits (not joking here).
> 
> When you have an array of bytes that fills up more than half the address 
> range (not very common) and you feel the urgent need to treat it 
> code-wise as just any other array, nothing special, no ma!, well, dumbness.
> 
> Of course it's OK to not use half of the possible size range for the 
> direct size representation (it comes into play in other contexts). As a 
> concrete example, it's decidedly OK on a 64-bit system. And as another 
> concrete example, it was OK in 32-bit Windows, where half the range was 
> all you had available anyway without very special configuration, which I 
> believe was quite risky. It can be more difficult to argue in favor of 
> unsigned sizes for the cases of 32-bit Linux and 16-bit embedded. But 
> the former is a hopefully soon extinct beast, and the latter requires 
> special considerations anyway, not just unthinking use of vanilla C++.
> 
> I gather that that reality is very much part of why the language 
> designer Bjarne Stroustrup says it's OK with that apparent/alleged range 
> waste, and a mistake to use unsigned for sizes. For example, Bjarne is 
> one of the three main authors of the C++ Core Guidelines that recommends 
> signed types for numbers. He's even written a paper titled "Subscripts 
> and sizes should be signed", available at <url: 
> https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p1428r0.pdf>; 
> please read it.
> 
> 
> 
> Cheers,
> 
> - Alf

Actually, during the formative years of the inital versions of the 
standard, the segmented archicture of the x86 system was prevalent, 
which could well have 32 bit pointers that only could span a 64k range 
of addresses at a time, made a 16 bit size_t not that impractical even 
if a 16 bit pointer couldn't handle all of memory.

A lot of the rules in the standard have features designed to support 
this sort of architecture.

[toc] | [prev] | [next] | [standalone]


#84255

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-05-24 16:15 +0000
Message-ID<rw7jK.34871$GTEb.20506@fx48.iad>
In reply to#84244
Richard Damon <Richard@Damon-Family.org> writes:
>On 5/22/22 5:23 PM, Alf P. Steinbach wrote:

>Actually, during the formative years of the inital versions of the 
>standard, the segmented archicture of the x86 system was prevalent, 

And it (x86) was running in 32-bit flat mode with paging enabled
for the majority of systems (unix) where C mattered at the time.  Full
32-bit pointers and 32-bit int.

> A lot of the rules in the standard have features designed to support 
> this sort of architecture.

While this is true, it had little to do with x86, but rather the
hundred or so other architectures used at the time in the micro,  mini
and mainframe worlds.

[toc] | [prev] | [next] | [standalone]


#84252

FromTim Rentsch <tr.17687@z991.linuxsc.com>
Date2022-05-24 05:37 -0700
Message-ID<86pmk314z1.fsf@linuxsc.com>
In reply to#84240
"Alf P. Steinbach" <alf.p.steinbach@gmail.com> writes:

[...]

> I gather that that reality is very much part of why the language
> designer Bjarne Stroustrup says it's OK with that apparent/alleged
> range waste, and a mistake to use unsigned for sizes.  For example,
> Bjarne is one of the three main authors of the C++ Core Guidelines
> that recommends signed types for numbers.  He's even written a paper
> titled "Subscripts and sizes should be signed", available at <url:
> https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p1428r0.pdf>;
> please read it.

I saw this essay several years ago (probably because someone
posted a link to it in the newsgroup here).  After seeing this
comment I got a fresh copy and re-read it, perhaps a bit more
carefully than my first reading.  Even with a fresh second look,
I still don't find it convincing.

[toc] | [prev] | [next] | [standalone]


#84216

FromMalcolm McLean <malcolm.arthur.mclean@gmail.com>
Date2022-05-20 04:19 -0700
Message-ID<6caff9be-a81b-4a8e-b302-add0b3207e99n@googlegroups.com>
In reply to#84210
On Friday, 20 May 2022 at 11:26:07 UTC+1, Ben wrote:
> Malcolm McLean <malcolm.ar...@gmail.com> writes: 
> 
> > On Friday, 20 May 2022 at 09:37:19 UTC+1, Juha Nieminen wrote: 
> >> Malcolm McLean <malcolm.ar...@gmail.com> wrote: 
> >> > The snag is that to use it, you have to be able to put the compiler into a 
> >> > mode where it traps on overflow instead of silently wrapping. Most 
> >> > compilers don't have such a mode. 
> >> I'm not entirely sure what the point of that would be. What exactly should 
> >> happen in such a program if an unsigned arithmetic operation overflows or 
> >> underflows? Should the program end? Should the code throw an exception? 
> >> (What would you do if such an exception happens?) How should the program 
> >> behave in such a situation? 
> >> 
> >> Also, should bit-shifting to the left trap if a 1-bit gets shifted out? 
> >> Ostensibly multiplying by a power of 2 should trap if it overflows. 
> >> Should bit-shifting to the left also do so (given that it's effectively 
> >> the exact same thing)? 
> >> 
> >> Maybe it should behave like a few other programming languages which switch 
> >> from hardware-register-sized integers to software multiprecision integers 
> >> on any overflow? (But even then, what about underflow?) 
> >> 
> >> If you want integers that grow as needed, just use GMP? 
> >> 
> > Say we've got this code. 
> > 
> > int width = getinteger(): 
> > int height = getinteger(); 
> > int Npixels = width* height; 
> > 
> > There's a potential vulnerability there if Npixels wraps. An attacker 
> > might be able to construct an input file that causes the program to 
> > access memory out of bounds in such a way as to cause arbitrary code 
> > to be executed.
> I found the suggestion (not from you) that UB is an advantage a very odd 
> one. I would consider it too unpredictable to be any practical help. 
> 
> But here, why not just make all the sizes unsigned? You have to do some 
> bounds checking since the data come from a file which can be maliciously 
> constructed, so provided you do that check, a wrapped Npixels will 
> presumably just produce a junk result. 
> 
> And, if you decide to add an overflow check, it's no harder using 
> unsigned. 
> 
You frequently generate co-ordinates that go outside of image bounds when 
working with images. It's easier to deal with this if the values to the top and
left go negative instead of wrapping to high positive values.

As a general rule, it's better if all integers are signed.

If you make a mistake, and despite your best efforts to catch overflow before it
occurs, an overflow happens, then usually it is better for the program to
terminate than for it to continue. In theory, you can set up a C++ compiler to
achieve this. In reality, as you say, it's too unpredictable.

[toc] | [prev] | [next] | [standalone]


#84217

FromBen <ben.usenet@bsb.me.uk>
Date2022-05-20 12:32 +0100
Message-ID<871qwo5tib.fsf@bsb.me.uk>
In reply to#84216
Malcolm McLean <malcolm.arthur.mclean@gmail.com> writes:

> On Friday, 20 May 2022 at 11:26:07 UTC+1, Ben wrote:
>> Malcolm McLean <malcolm.ar...@gmail.com> writes: 
>> 
>> > On Friday, 20 May 2022 at 09:37:19 UTC+1, Juha Nieminen wrote: 
>> >> Malcolm McLean <malcolm.ar...@gmail.com> wrote: 
>> >> > The snag is that to use it, you have to be able to put the compiler into a 
>> >> > mode where it traps on overflow instead of silently wrapping. Most 
>> >> > compilers don't have such a mode. 
>> >> I'm not entirely sure what the point of that would be. What exactly should 
>> >> happen in such a program if an unsigned arithmetic operation overflows or 
>> >> underflows? Should the program end? Should the code throw an exception? 
>> >> (What would you do if such an exception happens?) How should the program 
>> >> behave in such a situation? 
>> >> 
>> >> Also, should bit-shifting to the left trap if a 1-bit gets shifted out? 
>> >> Ostensibly multiplying by a power of 2 should trap if it overflows. 
>> >> Should bit-shifting to the left also do so (given that it's effectively 
>> >> the exact same thing)? 
>> >> 
>> >> Maybe it should behave like a few other programming languages which switch 
>> >> from hardware-register-sized integers to software multiprecision integers 
>> >> on any overflow? (But even then, what about underflow?) 
>> >> 
>> >> If you want integers that grow as needed, just use GMP? 
>> >> 
>> > Say we've got this code. 
>> > 
>> > int width = getinteger(): 
>> > int height = getinteger(); 
>> > int Npixels = width* height; 
>> > 
>> > There's a potential vulnerability there if Npixels wraps. An attacker 
>> > might be able to construct an input file that causes the program to 
>> > access memory out of bounds in such a way as to cause arbitrary code 
>> > to be executed.
>> I found the suggestion (not from you) that UB is an advantage a very odd 
>> one. I would consider it too unpredictable to be any practical help. 
>> 
>> But here, why not just make all the sizes unsigned? You have to do some 
>> bounds checking since the data come from a file which can be maliciously 
>> constructed, so provided you do that check, a wrapped Npixels will 
>> presumably just produce a junk result. 
>> 
>> And, if you decide to add an overflow check, it's no harder using 
>> unsigned. 
>> 
> You frequently generate co-ordinates that go outside of image bounds
> when working with images. It's easier to deal with this if the values
> to the top and left go negative instead of wrapping to high positive
> values.

But the bounds of the image won't correspond to the signed wrapping so
this has to be coded and checked for anyway.

> As a general rule, it's better if all integers are signed.

As a general rule, general rules are too general to be considered rules.

-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#84212

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-05-20 10:40 +0000
Message-ID<t67r62$16g0$1@gioia.aioe.org>
In reply to#84209
Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote:
> Say we've got this code.
> 
> int width = getinteger():
> int height = getinteger();
> int Npixels = width* height;
> 
> There's a potential vulnerability there if Npixels wraps. An attacker might be able to construct
> an input file that causes the program to access memory out of bounds in such a way as
> to cause arbitrary code to be executed.
> 
> If the program terminates with an error message, that vulnerability is removed.

Being able to crash a program (that's not supposed to ever end) can be
a vulnerability in itself.

I don't think you can safeguard against every single exploit with language
features.

> Usually, terminating will be better. The program won't process the image correctly, so ignoring
> the problem won't help.

As with all such situations, the programmer ought to validate the input
explicitly. It's quite easy to see if the dimensions of an image will be too
large.

[toc] | [prev] | [next] | [standalone]


#84211

FromDavid Brown <david.brown@hesbynett.no>
Date2022-05-20 12:31 +0200
Message-ID<t67ql7$e7h$1@dont-email.me>
In reply to#84208
On 20/05/2022 10:37, Juha Nieminen wrote:
> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote:
>> The snag is that to use it, you have to be able to put the compiler into a
>> mode where it traps on overflow instead of silently wrapping. Most
>> compilers don't have such a mode.
> 
> I'm not entirely sure what the point of that would be. What exactly should
> happen in such a program if an unsigned arithmetic operation overflows or
> underflows? Should the program end? Should the code throw an exception?
> (What would you do if such an exception happens?) How should the program
> behave in such a situation?
> 

That's the big question in regard to overflow (signed or unsigned). 
/Sometimes/ you actually want wrapping behaviour, but it's rare.

Possible desirable behaviour include:

1. Halting the program with an error message.

2. Returning an error value or other special value (like a NaN or 
infinity in floating point).

3. Returning a default value.

4. Saturating.

5. Wrapping.

6. Logging the error in some way, and continuing (with the wrapped 
value, default, value, or whatever).

7. Setting errno and continuing.

8. Throwing a C++ exception.

9. Calling an error-handler function.

10. Returning an unspecified value.  (i.e., the result is a valid value 
of the type, but you have no idea what value it is).

11. Undefined behaviour, so the compiler can assume it won't happen if 
it aids optimisation.

12. Automatically growing the size of the types.


I'm sure there are more possibilities.  And you might want different 
behaviour for different types, or at different points in the code, or 
with different operations - such as having "x << 1" be defined as 
wrapping while "x * 2" be defined as trapping with an error message on 
overflow.

And then there is the question of how to combine this all with 
optimisation, re-arrangements, and intermediary calculations.


Usually, but not always, an overflow (signed or unsigned) is a bug in 
the code.  So having compiler tools that help catch that bug is useful 
during debugging and testing.  (And it doesn't matter whether "most 
compilers don't have such a mode" - it only matters that the compiler 
and tools you use during development have such a mode.)

Having options aimed at minimising the impact of run-time errors can be 
useful for deployed code - though the compiler needs guidance for how to 
do that.  (Halting with an error is perhaps the safest choice for a 
desktop program, avoiding the risk of security holes - but it's not a 
great choice for an aeroplane engine controller.)


I guess the only answer is to make your own C++ integer classes that 
have the overflow behaviour /you/ want in the code at the time.  The 
only thing you can be really sure of, is that there is /no/ single 
"correct" answer.


> Also, should bit-shifting to the left trap if a 1-bit gets shifted out?
> Ostensibly multiplying by a power of 2 should trap if it overflows.
> Should bit-shifting to the left also do so (given that it's effectively
> the exact same thing)?
> 
> Maybe it should behave like a few other programming languages which switch
> from hardware-register-sized integers to software multiprecision integers
> on any overflow? (But even then, what about underflow?)
> 
> If you want integers that grow as needed, just use GMP?

[toc] | [prev] | [next] | [standalone]


#84213

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-05-20 10:43 +0000
Message-ID<t67rc6$190e$1@gioia.aioe.org>
In reply to#84211
David Brown <david.brown@hesbynett.no> wrote:
> That's the big question in regard to overflow (signed or unsigned). 
> /Sometimes/ you actually want wrapping behaviour, but it's rare.
> 
> Possible desirable behaviour include:

From the perspective of C++ (and C) in particular, the problem is that
unless the CPU itself supports catching overflow (and underflow), the
compiler would need to generate checking code for every single
arithmetic operation made in the program. If your program is 100%
valid you would be "paying for something you don't use".

[toc] | [prev] | [next] | [standalone]


#84219

FromDavid Brown <david.brown@hesbynett.no>
Date2022-05-20 14:52 +0200
Message-ID<t682v1$cq9$1@dont-email.me>
In reply to#84213
On 20/05/2022 12:43, Juha Nieminen wrote:
> David Brown <david.brown@hesbynett.no> wrote:
>> That's the big question in regard to overflow (signed or unsigned).
>> /Sometimes/ you actually want wrapping behaviour, but it's rare.
>>
>> Possible desirable behaviour include:
> 
>  From the perspective of C++ (and C) in particular, the problem is that
> unless the CPU itself supports catching overflow (and underflow), the
> compiler would need to generate checking code for every single
> arithmetic operation made in the program. If your program is 100%
> valid you would be "paying for something you don't use".

That is absolutely a relevant issue, for some code.

For a lot of other code, speed is not important.  For Malcolm's example, 
the speed of the code that calculates the sizes and number of pixels is 
going to be irrelevant compared to the library code that creates or 
manipulates the image.

Many of the possible behaviours I listed are only relevant during 
testing and debugging, where efficiency is usually less important in 
comparison to finding issues with the code.

(Even if a process supports overflow trapping instructions, these are 
rarely free - they can mess with branch prediction, pipelining, or 
superscaling.  And they often mess with the compiler's abilities to 
re-arrange code for optimisation.)

[toc] | [prev] | [next] | [standalone]


#84179

From"Alf P. Steinbach" <alf.p.steinbach@gmail.com>
Date2022-05-19 15:34 +0200
Message-ID<t65h0i$9li$1@dont-email.me>
In reply to#84177
On 19 May 2022 14:49, David Brown wrote:
> On 19/05/2022 12:37, Alf P. Steinbach wrote:
> > [snipped]
> 
> [snipped general argumentative part]
>
> In this particular situation, all the values are non-negative, and well 
> below the range of int.  So int, unsigned int, size_t, or many other 
> types would be fine both for the indexing and for the calculation 
> values.  So no, it does not matter for correctness of code.

That's a circular argument, starting with observing that this code is 
correct, and that the choice of signed or unsigned as type for all 
values in it doesn't matter technically, and concluding that that means 
that the choice couldn't matter for the original correctness.

But of course it could, and would: that's the reason for the C++ core 
guidelines' recommendation.

The circular argument hasn't disproved anything; it's just a fallacy.


> "int" is 
> often a convenient default choice, but there are no direct advantages to 
> it here - and a possible direct disadvantage for the calculations in 
> that division is sometimes slower with signed types.

Again, the burden of proof for the efficiency claim, is as I see it on you.

A proof would involve MEASURING this code, and/or other code with 
remainder operations on signed type representation of non-negative 
numbers in the same range, compared to ditto measurement of same code 
but with unsigned type representation.

It could be that you're right in that the relative efficiency for e.g. 
unsigned division by constant, spills over into general division of 
non-negative numbers on general computers. If so then we (or at least I) 
have learned something. If not, then we (both) have learned something. ;-)


> [snipped argumentative /ad hominem/ fallacy]
>>
>> I criticized the herd opinions with those words.
>>
>> I criticized the other person's code by showing an alternative way to 
>> express it.
>>
>> And that is not what you wrote (quoted).
>>
> 
> You told the OP to "consider disregarding silly context independent herd 
> opinion".  This is not a court of law - we do not need to argue over 
> tiny details of wording.

Misrepresentation in the direction of claiming that I'd characterized a 
person with derogatory wording, is not a detail.


> [snipped irrelevant]
> 
> If you prefer to stick to x86 cpus, following the herd opinion that 
> these are the only processors that matter, you can look at the timings 
> of signed and unsigned integer division instructions for large numbers 
> of processors in Agner Fog's timing tables:
> 
> <https://www.agner.org/optimize/instruction_tables.pdf>
> 
> A quick look at a few cases suggests it is common for signed integer 
> division to be a little slower than unsigned integer division, even on 
> quite modern devices.
> 
> Of course this is unlikely to make a measurable difference in practice - 
> if code is performance-critical, you'd try to avoid integer division 
> anyway (floating point division is usually much faster on modern "big" 
> processors).

Not that I'm necessarily disagreeing about the last paragraph, but 
counting cycles of /assumed/ machine code instructions (like IDIV), and 
comparing the /ranges/ of number of clock cycles, is not a reliable way 
to determine efficiency of particular C++ code on a general computer.

So, this just adds some claims.

I would tend to agree with the last added claim, but not with the 
original claim about the alleged possible speed of unsigned division 
being a reason to choose unsigned types for code like this. If it turns 
out that there is a significant efficiency gain, then a simple `1u*` 
would optimize the remainder operation (the type conversion is a NOP 
since C++20). Much preferable to wholesale adoption of unsigned type.


>  The point is, changing the unsigned type to a signed type 
> is an unnecessary pessimisation

On the contrary.

Here's a slightly fallacious appeal to authority, since we're into 
fallacious reasoning anyway: that's why my choice is the same as the C++ 
core guidelines, and why your choice is in disagreement. :)


>  [snip]
> The guidelines are not focused on code efficiency, but on more important 
> issues of code correctness, readability, and maintainability (while 
> avoiding serious inefficiencies).  I think it would be very rare to have 
> the choice of signed or unsigned types determined mainly by the 
> efficiency of integer division - but a prime sieve might conceivably be 
> one of these rare cases where it is relevant.  However, I would be 
> surprised to see a measurable difference in real code.

Agreed.


Cheers,

- Alf

[toc] | [prev] | [next] | [standalone]


#84182

From"Alf P. Steinbach" <alf.p.steinbach@gmail.com>
Date2022-05-19 17:03 +0200
Message-ID<t65m7f$m3t$1@dont-email.me>
In reply to#84174
On 19 May 2022 09:19, Alf P. Steinbach wrote:
> 
> Code features different from original:
> 
> * Self-describing, including the alias with more reasonable parameter 
> order (like "5 int", not "int 5") for std library thing.
> * Signed types like `int` for numbers, not unsigned types like `size_t`.
> * `goto` where appropriate (would be nice with a labeled `continue`!).

I discovered that g++ diagnoses the `goto` as invalid in a `constexpr` 
function. Earlier I only compiled the code with MSVC, which accepted it.

I.e. `goto` is apparently not appropriate in this context.

Mea culpa. :(


- Alf

[toc] | [prev] | [next] | [standalone]


#84186

FromTim Rentsch <tr.17687@z991.linuxsc.com>
Date2022-05-19 08:53 -0700
Message-ID<86tu9l1pul.fsf@linuxsc.com>
In reply to#84173
Andrey Tarasevich <andreytarasevich@hotmail.com> writes:

> Hello
>
> Just for the sake of experiment I tried this
>
>   #include <array>
>
>   template <std::size_t N> constexpr auto collect_primes()
>   {
>     std::array<unsigned, N> primes = { 2, 3 };
>     std::size_t n_primes = 2;
>
>     for (unsigned i = 0; n_primes < N; ++i)
>     {
>       unsigned value = 6u * (i / 2u + 1u) + i % 2u * 2u - 1;
>
>       std::size_t i_prime = 0;
>       for (; i_prime < n_primes; ++i_prime)
>         if (value % primes[i_prime] == 0)
>           break;
>
>       if (i_prime == n_primes)
>         primes[n_primes++] = value;
>     }
>
>     return primes;
>   }
>
>   int main()
>   {
>     constexpr auto primes = collect_primes<78498>();
>     ...
>   }
>
> Given the above code GCC will [attempt to] build the table of primes
> at compile time.  Command line options like
>
>   -fconstexpr-ops-limit=10000000000
>
> are needed to prevent it from aborting this overly heavy compile-time
> evaluation prematurely.  And this compile-time evaluation takes hours
> and hours and hours and... well... honestly, I haven't been able to
> see it to completion.
>
> If I get rid of the `constexpr`, i.e. switch to regular run-time
> evaluation, the above table will be built in about 15 seconds on the
> same machine.
>
> Reducing the table size to 10000 results in 8 minute `constexpr`
> evaluation vs. 0.25 second run-time table generation.
>
> Granted, it is naive to expect the same efficiency from compile-time
> constexpr` evaluation as we expect from `-O3` run-time code.  They
> can't just paste the `constexpr` code into a temp file, compile it
> with `-O3` and run it to obtain the result (can they?).  The compiler
> will apparently have to "pseudo-run" the code inside an "abstract
> virtual C++ machine" of sorts... But the actual difference in
> performance is still rather startling.  What could be the primary
> reason behind `constexpr` evaluation being so much slower?  Is it just
> a quality of implementation issue?  Or is there a bunch of fundamental
> reasons it will always be slow?  I see that GCC consumes huge amounts
> of memory wile compiling this code... Is it trying to fully unwrap the
> cycle/logic, which could be the reason for massive memory consumption
> and subsequent slowdown once swapping begins?

I don't have answers for your questions, but I did play around
with this a little bit, enough to offer some ideas for what they
might be worth.  Disclaimer: my version of gcc is almost
certainly less advanced than yours, so that may have affected my
observations.

One:  I think -O0 is a better basis than -O3 for comparing a
constexpr version to a non-constexpr version.  Probably gcc
evaluates constexpr expressions using a very straightforward
interpretation than an optimized interpretation.

Two:  In evaluating constexpr code, std::array is much more
expensive than just a plain array (eg, unsigned primes[N];).
I measured the time ratio between them at about 6.7.

Three:  Apparently the amount of memory required for std::array
is also significantly more than a plain array.  I didn't try to
measure this but there were other indications in the respective
compilation behaviors.

Four:  Changing the source so a plain array could be used instead
of a std::array, and also a somewhat different method of testing
for primality, I saw a time ratio (between a compile-time version
and run-time only version) of about 125-150 to 1.  That is on a
much smaller interval than the original above because trying to
run the full interval made the compiler go belly up.

Five:  Related to the last point, I suspect there is a knee in
the curve for compile-time evaluation, perhaps because the array
of primes is a local variable, which implies that a very large
stack frame may be needed for the full interval.

Just some ideas for possible directions to explore...

[toc] | [prev] | [next] | [standalone]


#84188

FromMarcel Mueller <news.5.maazl@spamgourmet.org>
Date2022-05-19 18:33 +0200
Message-ID<t65rh0$3edh$1@gwaiyur.mb-net.net>
In reply to#84186
Am 19.05.22 um 17:53 schrieb Tim Rentsch:
> One:  I think -O0 is a better basis than -O3 for comparing a
> constexpr version to a non-constexpr version.  Probably gcc
> evaluates constexpr expressions using a very straightforward
> interpretation than an optimized interpretation.

I am in doubt whether the constexpr evaluator will ever enter the code 
generation stage. Probably it is just an interpreter which operates 
directly on the internal execution trees. This will run some orders of 
magnitude slower.
In most scenarios this is probably even faster than generating temporary 
code in the local platform which might not even be supported if you 
invoked a cross compiler.

> Two:  In evaluating constexpr code, std::array is much more
> expensive than just a plain array (eg, unsigned primes[N];).
> I measured the time ratio between them at about 6.7.

I would not expect anything else. All the trivial function calls cannot 
be eliminated in this mode.

> Three:  Apparently the amount of memory required for std::array
> is also significantly more than a plain array.

The execution tree is more similar to a DOM tree versus HTML. Each 
operator is a dynamically allocated, most likely polymorphic object. 
Even the buckets where the data is stored are probably generic value 
objects. Again the native binary representation may not fit to the 
target platform.

> Five:  Related to the last point, I suspect there is a knee in
> the curve for compile-time evaluation, perhaps because the array
> of primes is a local variable, which implies that a very large
> stack frame may be needed for the full interval.

I think the assumption that compile time evaluation is in any kind 
similar to runtime evaluation is wrong in any way.
Maybe future compiler versions have internal optimizations to improve 
the situation but the principal problem, that the compilation platform 
is not the same than the target platform will restrict this or at least 
make it very complex. The best which might happen is something like a 
JIT with hotspot detection.

But there is one thing that prevents large optimization of compilers. 
While it is a primary goal that compiled code runs as fast as possible, 
it is the primary goal of a compiler to avoid any difficult to find 
bugs. And optimizations are likely subject to introduce new bugs.


Marcel

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | comp.lang.c++


csiph-web