Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c++ > #85614 > unrolled thread

Differences between C and C++

Started byJuha Nieminen <nospam@thanks.invalid>
First post2022-07-25 13:14 +0000
Last post2022-07-28 06:16 -0700
Articles 20 on this page of 84 — 18 participants

Back to article view | Back to comp.lang.c++


Contents

  Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-25 13:14 +0000
    Re: Differences between C and C++ "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-07-25 16:18 +0200
    Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-25 18:47 +0300
      Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-25 19:49 +0300
        Re: Differences between C and C++ Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-07-25 19:38 +0100
    Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-25 08:51 -0700
    Re: Differences between C and C++ Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-07-25 17:14 +0100
      Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-25 16:19 +0000
        Re: Differences between C and C++ Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-07-25 19:28 +0100
        Re: Differences between C and C++ Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-07-25 12:21 -0700
          Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-26 07:54 +0000
            Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-26 08:00 +0000
              Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-26 08:09 +0000
                Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-26 02:42 -0700
                  Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-26 15:29 +0000
                    Re: Differences between C and C++ scott@slp53.sl.home (Scott Lurndal) - 2022-07-26 16:30 +0000
                    Re: Differences between C and C++ Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-07-26 11:59 -0700
                      Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-27 07:52 +0000
                        Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-27 01:56 -0700
                          Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-27 14:51 +0000
                            Re: Differences between C and C++ Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-07-27 11:11 -0700
      Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-26 06:09 +0000
        Re: Differences between C and C++ Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-07-27 17:33 +0100
    Re: Differences between C and C++ Chris Vine <chris@cvine--nospam--.freeserve.co.uk> - 2022-07-25 17:45 +0100
    Re: Differences between C and C++ Paul N <gw7rib@aol.com> - 2022-07-27 10:09 -0700
      Re: Differences between C and C++ Bo Persson <bo@bo-persson.se> - 2022-07-27 19:28 +0200
        Re: Differences between C and C++ Paul N <gw7rib@aol.com> - 2022-07-27 11:43 -0700
          Re: Differences between C and C++ David Brown <david.brown@hesbynett.no> - 2022-07-29 18:04 +0200
            Re: Differences between C and C++ Richard Damon <Richard@Damon-Family.org> - 2022-07-29 18:47 -0400
              Re: Differences between C and C++ David Brown <david.brown@hesbynett.no> - 2022-07-30 14:08 +0200
        Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-28 08:03 +0000
          Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-28 01:57 -0700
            Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-28 09:04 +0000
      Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-28 06:53 +0000
        Re: Differences between C and C++ Richard Damon <Richard@Damon-Family.org> - 2022-07-28 07:33 -0400
          Re: Differences between C and C++ Bo Persson <bo@bo-persson.se> - 2022-07-28 15:04 +0200
            Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-28 15:11 +0000
              Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-28 20:03 +0300
                Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-29 07:56 +0000
                  Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-29 12:03 +0300
                    Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-29 09:28 +0000
                    Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-29 12:27 +0000
                      Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-29 05:43 -0700
                        Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-29 14:14 +0000
                          Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-29 08:25 -0700
                            Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-29 15:48 +0000
                              Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-29 13:12 -0700
                                Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-30 09:28 +0000
                                  Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-30 03:39 -0700
                                    Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-30 14:23 +0000
                                      Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-30 19:07 -0700
                                        Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-31 07:22 +0000
                                          Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-31 06:07 -0700
                                            Re: Differences between C and C++ muttley@dastardlyhq.com - 2022-07-31 16:36 +0000
                                              Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-08-01 01:12 -0700
                        Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-29 22:47 +0000
                          Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-30 03:48 -0700
                  Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-29 05:16 -0700
                  Re: Differences between C and C++ Manfred <noname@add.invalid> - 2022-07-30 02:14 +0200
                    Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-31 14:39 +0000
                      Re: Differences between C and C++ "Fred. Zwarts" <F.Zwarts@KVI.nl> - 2022-07-31 18:49 +0200
                        Re: Differences between C and C++ Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-07-31 14:49 -0700
                          Re: Differences between C and C++ Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-07-31 17:03 -0700
                            Re: Differences between C and C++ Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-08-01 00:16 -0700
                              Re: Differences between C and C++ Ike Naar <ike@sdf.org> - 2022-08-01 07:27 +0000
                              Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-08-01 10:47 +0300
                          Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-08-01 10:34 +0300
                          Re: Differences between C and C++ Manfred <noname@add.invalid> - 2022-08-03 01:20 +0200
                            Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-08-03 07:59 +0000
                              Re: Differences between C and C++ Bo Persson <bo@bo-persson.se> - 2022-08-03 12:45 +0200
                                Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-08-03 10:53 +0000
                                  Re: Differences between C and C++ Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-08-03 03:58 -0700
                                  Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-08-03 14:49 +0300
                                    Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-08-03 23:31 -0700
                        Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-08-01 11:29 +0000
                          Re: Differences between C and C++ Bo Persson <bo@bo-persson.se> - 2022-08-01 14:39 +0200
                            Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-08-02 05:43 +0000
                      Re: Differences between C and C++ Manfred <noname@add.invalid> - 2022-08-03 01:13 +0200
                        Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-08-03 09:34 +0300
                Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-29 09:22 +0000
                  Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-29 13:53 +0300
                    Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-29 14:13 +0300
                    Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-29 14:11 +0000
          Re: Differences between C and C++ Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-07-28 06:16 -0700

Page 4 of 5 — ← Prev page 1 2 3 [4] 5  Next page →


#85766

From"Fred. Zwarts" <F.Zwarts@KVI.nl>
Date2022-07-31 18:49 +0200
Message-ID<tc6bq1$1r17$1@gioia.aioe.org>
In reply to#85764
Op 31.jul..2022 om 16:39 schreef Juha Nieminen:
> Manfred <noname@add.invalid> wrote:
>> That, and below, is generally true for a generic dynamic container.
>> However, std::string is not a generic container. It is a very specific
>> and very optimized container of chars.
>> So, if you use a decent implementation of std the performance of
>> std::string and char* is usually pretty close.
> 
> I don't know if you are referring to it, but many people seem to think that
> "short string optimization" makes std::string pretty much as efficient
> as an array of char (for strings that are short enough).
> 
> They fail to take into consideration that short string optimization
> requires conditionals in almost all member functons that access the
> string data. Conditionals are not free.


I wonder if that is true. The string object has a pointer to its buffer 
and a length. I assume that if the short string optimization is used, 
this pointer points to the internal buffer. Only member functions that 
want to increase the size of the buffer need those conditionals. I 
wonder whether "almost all member functions" modify the buffer size.

[toc] | [prev] | [next] | [standalone]


#85767

FromMalcolm McLean <malcolm.arthur.mclean@gmail.com>
Date2022-07-31 14:49 -0700
Message-ID<8699b402-e8e2-4e85-8de6-d67249cbeca9n@googlegroups.com>
In reply to#85766
On Sunday, 31 July 2022 at 17:49:22 UTC+1, F.Zwarts wrote:
> Op 31.jul..2022 om 16:39 schreef Juha Nieminen:
> > Manfred <non...@add.invalid> wrote: 
> >> That, and below, is generally true for a generic dynamic container. 
> >> However, std::string is not a generic container. It is a very specific 
> >> and very optimized container of chars. 
> >> So, if you use a decent implementation of std the performance of 
> >> std::string and char* is usually pretty close. 
> > 
> > I don't know if you are referring to it, but many people seem to think that 
> > "short string optimization" makes std::string pretty much as efficient 
> > as an array of char (for strings that are short enough). 
> > 
> > They fail to take into consideration that short string optimization 
> > requires conditionals in almost all member functons that access the 
> > string data. Conditionals are not free.
> I wonder if that is true. The string object has a pointer to its buffer 
> and a length. I assume that if the short string optimization is used, 
> this pointer points to the internal buffer. Only member functions that 
> want to increase the size of the buffer need those conditionals. I 
> wonder whether "almost all member functions" modify the buffer size.
>
The std::string is (in C)
struct string
{
char *buff;
size_t len;
}
so 16 bytes on a 64-bit machine.
However in reality not all 64 bits of a pointer are wired to memory addresses. 
It's likely that the most significant bit has to be clear in a valid address.
So we can exploit this by (in C)
struct shortstring
{
   bool flag;  // set when the shortstring is valid
   int pad:7;  // Maybe use for length or encoding or other stuff
   char data[15]; // UP to 14 characters of ASCII string data.
};

union
{
   struct string s;
   struct shortstring ss;
} std_string;

Now string::size is implemented as
if(s->ss.flag)
    return strlen(s->ss.data);
 else
   return s->s.len;


The other string member functions are implemented similarly.

[toc] | [prev] | [next] | [standalone]


#85768

FromKeith Thompson <Keith.S.Thompson+u@gmail.com>
Date2022-07-31 17:03 -0700
Message-ID<871qu0x21j.fsf@nosuchdomain.example.com>
In reply to#85767
Malcolm McLean <malcolm.arthur.mclean@gmail.com> writes:
> On Sunday, 31 July 2022 at 17:49:22 UTC+1, F.Zwarts wrote:
>> Op 31.jul..2022 om 16:39 schreef Juha Nieminen:
>> > Manfred <non...@add.invalid> wrote: 
>> >> That, and below, is generally true for a generic dynamic container. 
>> >> However, std::string is not a generic container. It is a very specific 
>> >> and very optimized container of chars. 
>> >> So, if you use a decent implementation of std the performance of 
>> >> std::string and char* is usually pretty close. 
>> > 
>> > I don't know if you are referring to it, but many people seem to think that 
>> > "short string optimization" makes std::string pretty much as efficient 
>> > as an array of char (for strings that are short enough). 
>> > 
>> > They fail to take into consideration that short string optimization 
>> > requires conditionals in almost all member functons that access the 
>> > string data. Conditionals are not free.
>> I wonder if that is true. The string object has a pointer to its buffer 
>> and a length. I assume that if the short string optimization is used, 
>> this pointer points to the internal buffer. Only member functions that 
>> want to increase the size of the buffer need those conditionals. I 
>> wonder whether "almost all member functions" modify the buffer size.
>>
> The std::string is (in C)
> struct string
> {
> char *buff;
> size_t len;
> }
> so 16 bytes on a 64-bit machine.

That's one possible implementation.

> However in reality not all 64 bits of a pointer are wired to memory addresses. 
> It's likely that the most significant bit has to be clear in a valid address.
> So we can exploit this by (in C)
> struct shortstring
> {
>    bool flag;  // set when the shortstring is valid
>    int pad:7;  // Maybe use for length or encoding or other stuff
>    char data[15]; // UP to 14 characters of ASCII string data.
> };
>
> union
> {
>    struct string s;
>    struct shortstring ss;
> } std_string;
>
> Now string::size is implemented as
> if(s->ss.flag)
>     return strlen(s->ss.data);
>  else
>    return s->s.len;
>
>
> The other string member functions are implemented similarly.

That won't work without extra code to ensure that the short string
optimization isn't used in all cases.  The length of a std::string
is not determined by a null character.  This program must print 3 :

#include <iostream>
#include <string>

int main() {
    std::string s = "a";
    s += '\0';
    s += 'b';
    std::cout << s.size() << '\n';
}

-- 
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
Working, but not speaking, for Philips
void Void(void) { Void(); } /* The recursive call of the void */

[toc] | [prev] | [next] | [standalone]


#85769

FromMalcolm McLean <malcolm.arthur.mclean@gmail.com>
Date2022-08-01 00:16 -0700
Message-ID<96324449-0709-4099-a262-d44bc65d27e9n@googlegroups.com>
In reply to#85768
On Monday, 1 August 2022 at 01:04:09 UTC+1, Keith Thompson wrote:
> Malcolm McLean <malcolm.ar...@gmail.com> writes: 
> > On Sunday, 31 July 2022 at 17:49:22 UTC+1, F.Zwarts wrote: 
> >> Op 31.jul..2022 om 16:39 schreef Juha Nieminen: 
> >> > Manfred <non...@add.invalid> wrote: 
> >> >> That, and below, is generally true for a generic dynamic container. 
> >> >> However, std::string is not a generic container. It is a very specific 
> >> >> and very optimized container of chars. 
> >> >> So, if you use a decent implementation of std the performance of 
> >> >> std::string and char* is usually pretty close. 
> >> > 
> >> > I don't know if you are referring to it, but many people seem to think that 
> >> > "short string optimization" makes std::string pretty much as efficient 
> >> > as an array of char (for strings that are short enough). 
> >> > 
> >> > They fail to take into consideration that short string optimization 
> >> > requires conditionals in almost all member functons that access the 
> >> > string data. Conditionals are not free. 
> >> I wonder if that is true. The string object has a pointer to its buffer 
> >> and a length. I assume that if the short string optimization is used, 
> >> this pointer points to the internal buffer. Only member functions that 
> >> want to increase the size of the buffer need those conditionals. I 
> >> wonder whether "almost all member functions" modify the buffer size. 
> >> 
> > The std::string is (in C) 
> > struct string 
> > { 
> > char *buff; 
> > size_t len; 
> > } 
> > so 16 bytes on a 64-bit machine.
> That's one possible implementation.
> > However in reality not all 64 bits of a pointer are wired to memory addresses. 
> > It's likely that the most significant bit has to be clear in a valid address. 
> > So we can exploit this by (in C) 
> > struct shortstring 
> > { 
> > bool flag; // set when the shortstring is valid 
> > int pad:7; // Maybe use for length or encoding or other stuff 
> > char data[15]; // UP to 14 characters of ASCII string data. 
> > }; 
> > 
> > union 
> > { 
> > struct string s; 
> > struct shortstring ss; 
> > } std_string; 
> > 
> > Now string::size is implemented as 
> > if(s->ss.flag) 
> > return strlen(s->ss.data); 
> > else 
> > return s->s.len; 
> > 
> > 
> > The other string member functions are implemented similarly.
> That won't work without extra code to ensure that the short string 
> optimization isn't used in all cases. The length of a std::string 
> is not determined by a null character. This program must print 3 : 
> 
> #include <iostream> 
> #include <string> 
> 
> int main() { 
> std::string s = "a"; 
> s += '\0'; 
> s += 'b'; 
> std::cout << s.size() << '\n'; 
> } 
> 
You're right. std::size includes embedded nuls. So in fact you need to use four bits
for the "pad" field to store the length. Or nul-pad the string and work backwards
until you hit a set byte. 

[toc] | [prev] | [next] | [standalone]


#85770

FromIke Naar <ike@sdf.org>
Date2022-08-01 07:27 +0000
Message-ID<slrntef039.g3m.ike@sdf.org>
In reply to#85769
On 2022-08-01, Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote:
> You're right. std::size includes embedded nuls. So in fact you need to use four bits
> for the "pad" field to store the length. Or nul-pad the string and work backwards
> until you hit a set byte. 

Working backwards will give the wrong result if there are embedded nuls
at the end of the string.

[toc] | [prev] | [next] | [standalone]


#85772

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-08-01 10:47 +0300
Message-ID<tc80ep$peih$1@dont-email.me>
In reply to#85769
01.08.2022 10:16 Malcolm McLean kirjutas:
> On Monday, 1 August 2022 at 01:04:09 UTC+1, Keith Thompson wrote:
>> Malcolm McLean <malcolm.ar...@gmail.com> writes:
>>> On Sunday, 31 July 2022 at 17:49:22 UTC+1, F.Zwarts wrote:
>>>> Op 31.jul..2022 om 16:39 schreef Juha Nieminen:
>>>>> Manfred <non...@add.invalid> wrote:
>>>>>> That, and below, is generally true for a generic dynamic container.
>>>>>> However, std::string is not a generic container. It is a very specific
>>>>>> and very optimized container of chars.
>>>>>> So, if you use a decent implementation of std the performance of
>>>>>> std::string and char* is usually pretty close.
>>>>>
>>>>> I don't know if you are referring to it, but many people seem to think that
>>>>> "short string optimization" makes std::string pretty much as efficient
>>>>> as an array of char (for strings that are short enough).
>>>>>
>>>>> They fail to take into consideration that short string optimization
>>>>> requires conditionals in almost all member functons that access the
>>>>> string data. Conditionals are not free.
>>>> I wonder if that is true. The string object has a pointer to its buffer
>>>> and a length. I assume that if the short string optimization is used,
>>>> this pointer points to the internal buffer. Only member functions that
>>>> want to increase the size of the buffer need those conditionals. I
>>>> wonder whether "almost all member functions" modify the buffer size.
>>>>
>>> The std::string is (in C)
>>> struct string
>>> {
>>> char *buff;
>>> size_t len;
>>> }
>>> so 16 bytes on a 64-bit machine.
>> That's one possible implementation.
>>> However in reality not all 64 bits of a pointer are wired to memory addresses.
>>> It's likely that the most significant bit has to be clear in a valid address.
>>> So we can exploit this by (in C)
>>> struct shortstring
>>> {
>>> bool flag; // set when the shortstring is valid
>>> int pad:7; // Maybe use for length or encoding or other stuff
>>> char data[15]; // UP to 14 characters of ASCII string data.
>>> };
>>>
>>> union
>>> {
>>> struct string s;
>>> struct shortstring ss;
>>> } std_string;
>>>
>>> Now string::size is implemented as
>>> if(s->ss.flag)
>>> return strlen(s->ss.data);
>>> else
>>> return s->s.len;
>>>
>>>
>>> The other string member functions are implemented similarly.
>> That won't work without extra code to ensure that the short string
>> optimization isn't used in all cases. The length of a std::string
>> is not determined by a null character. This program must print 3 :
>>
>> #include <iostream>
>> #include <string>
>>
>> int main() {
>> std::string s = "a";
>> s += '\0';
>> s += 'b';
>> std::cout << s.size() << '\n';
>> }
>>
> You're right. std::size includes embedded nuls. So in fact you need to use four bits
> for the "pad" field to store the length. Or nul-pad the string and work backwards
> until you hit a set byte.

Or check for the presence of embedded nuls and switch off SSO for such 
strings. Alas, that would mean non-zero overhead.

[toc] | [prev] | [next] | [standalone]


#85771

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-08-01 10:34 +0300
Message-ID<tc7vm4$p7sr$1@dont-email.me>
In reply to#85767
01.08.2022 00:49 Malcolm McLean kirjutas:

> The std::string is (in C)
> struct string
> {
> char *buff;
> size_t len;
> }
> so 16 bytes on a 64-bit machine.

In reality some popular implementations have sizeof(std::string)==32 
(e.g. MSVC2019 on Windows, gcc 8.3 on Linux). When trying to study the 
internals I see lots of attention to allocators. Indeed, all string 
constructors also take an allocator argument which must be stored somewhere.

With 32 bytes, there would be a lot more room for SSO. Unfortunately it 
seems the implementations are not very eager to make use of that space, 
I think only 16 bytes gets typically used as SSO. Presumably they have 
the same char* buff and size_t len in an SSO string, it's just that the 
char* buff pointer just points 16 bytes ahead, into the internal SSO 
buffer of 16 bytes. This ensures there would be zero overhead for SSO 
and no conditionals involved for non-mutable access.

[toc] | [prev] | [next] | [standalone]


#85785

FromManfred <noname@add.invalid>
Date2022-08-03 01:20 +0200
Message-ID<tccbff$aah$1@gioia.aioe.org>
In reply to#85767
On 7/31/2022 11:49 PM, Malcolm McLean wrote:
> On Sunday, 31 July 2022 at 17:49:22 UTC+1, F.Zwarts wrote:
>> Op 31.jul..2022 om 16:39 schreef Juha Nieminen:
>>> Manfred <non...@add.invalid> wrote:
>>>> That, and below, is generally true for a generic dynamic container.
>>>> However, std::string is not a generic container. It is a very specific
>>>> and very optimized container of chars.
>>>> So, if you use a decent implementation of std the performance of
>>>> std::string and char* is usually pretty close.
>>>
>>> I don't know if you are referring to it, but many people seem to think that
>>> "short string optimization" makes std::string pretty much as efficient
>>> as an array of char (for strings that are short enough).
>>>
>>> They fail to take into consideration that short string optimization
>>> requires conditionals in almost all member functons that access the
>>> string data. Conditionals are not free.
>> I wonder if that is true. The string object has a pointer to its buffer
>> and a length. I assume that if the short string optimization is used,
>> this pointer points to the internal buffer. Only member functions that
>> want to increase the size of the buffer need those conditionals. I
>> wonder whether "almost all member functions" modify the buffer size.
>>
> The std::string is (in C)
> struct string
> {
> char *buff;
> size_t len;
> }
> so 16 bytes on a 64-bit machine.
> However in reality not all 64 bits of a pointer are wired to memory addresses.
> It's likely that the most significant bit has to be clear in a valid address.

On most modern architectures, it's even more likely that the /least/ 
significant bit is zero, which is in fact what your implementation does 
on a LE architecture.

> So we can exploit this by (in C)
> struct shortstring
> {
>     bool flag;  // set when the shortstring is valid
>     int pad:7;  // Maybe use for length or encoding or other stuff
>     char data[15]; // UP to 14 characters of ASCII string data.
> };
> 
> union
> {
>     struct string s;
>     struct shortstring ss;
> } std_string;
> 
> Now string::size is implemented as
> if(s->ss.flag)
>      return strlen(s->ss.data);
>   else
>     return s->s.len;
> 
> 
> The other string member functions are implemented similarly.

[toc] | [prev] | [next] | [standalone]


#85788

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-08-03 07:59 +0000
Message-ID<tcd9te$19s4$1@gioia.aioe.org>
In reply to#85785
Manfred <noname@add.invalid> wrote:
>> However in reality not all 64 bits of a pointer are wired to memory addresses.
>> It's likely that the most significant bit has to be clear in a valid address.
> 
> On most modern architectures, it's even more likely that the /least/ 
> significant bit is zero, which is in fact what your implementation does 
> on a LE architecture.

That would mean that a pointer can only point to even addresses. I don't
think that's the case. Rather obviously a pointer must be able to point
to any address (else you would be able to eg. traverse a string with
a pointer).

IIRC in Linux the kernel reserves the upper half of the memory range to the
kernel space and the lower half to user space (and IIRC user code has no
business using any kernel space pointer for any reason in any situation)
so it may be that at least in Linux the most-significant bit of pointers
in user code is indeed always zero. However, I would never write code
that makes this assumptions because it's not really an assumption you
can make at any level (I don't think even the Linux kernel makes that
absolute guarantee).

[toc] | [prev] | [next] | [standalone]


#85790

FromBo Persson <bo@bo-persson.se>
Date2022-08-03 12:45 +0200
Message-ID<jkv1ucFqa8tU1@mid.individual.net>
In reply to#85788
On 2022-08-03 at 09:59, Juha Nieminen wrote:
> Manfred <noname@add.invalid> wrote:
>>> However in reality not all 64 bits of a pointer are wired to memory addresses.
>>> It's likely that the most significant bit has to be clear in a valid address.
>>
>> On most modern architectures, it's even more likely that the /least/
>> significant bit is zero, which is in fact what your implementation does
>> on a LE architecture.
> 
> That would mean that a pointer can only point to even addresses. I don't
> think that's the case. Rather obviously a pointer must be able to point
> to any address (else you would be able to eg. traverse a string with
> a pointer).
> 

The observation is about *heap memory* used by the containers, where you 
are likley not allowed to alloocate a single byte, but will get a 
pointer properly aligned at 8 or 16 bytes. And therefore has some zeros 
at the lower end.

[toc] | [prev] | [next] | [standalone]


#85792

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-08-03 10:53 +0000
Message-ID<tcdk2t$1dc1$2@gioia.aioe.org>
In reply to#85790
Bo Persson <bo@bo-persson.se> wrote:
> On 2022-08-03 at 09:59, Juha Nieminen wrote:
>> Manfred <noname@add.invalid> wrote:
>>>> However in reality not all 64 bits of a pointer are wired to memory addresses.
>>>> It's likely that the most significant bit has to be clear in a valid address.
>>>
>>> On most modern architectures, it's even more likely that the /least/
>>> significant bit is zero, which is in fact what your implementation does
>>> on a LE architecture.
>> 
>> That would mean that a pointer can only point to even addresses. I don't
>> think that's the case. Rather obviously a pointer must be able to point
>> to any address (else you would be able to eg. traverse a string with
>> a pointer).
>> 
> 
> The observation is about *heap memory* used by the containers, where you 
> are likley not allowed to alloocate a single byte, but will get a 
> pointer properly aligned at 8 or 16 bytes. And therefore has some zeros 
> at the lower end.

The topic is short-string-optimization of a std::string-like class.
If the internal pointer used in that class is pointing to either allocated
memory or an internal buffer, you can't be sure that the internal buffer
will be aligned to a 16-byte, 8-byte or even a 2-byte boundary. That's
because the class instance itself may not be allocated on the heap.

[toc] | [prev] | [next] | [standalone]


#85793

FromMalcolm McLean <malcolm.arthur.mclean@gmail.com>
Date2022-08-03 03:58 -0700
Message-ID<4b8578a8-cc4d-434c-9ec9-3730ece082bcn@googlegroups.com>
In reply to#85792
On Wednesday, 3 August 2022 at 11:53:34 UTC+1, Juha Nieminen wrote:
> Bo Persson <b...@bo-persson.se> wrote: 
> > On 2022-08-03 at 09:59, Juha Nieminen wrote: 
> >> Manfred <non...@add.invalid> wrote: 
> >>>> However in reality not all 64 bits of a pointer are wired to memory addresses. 
> >>>> It's likely that the most significant bit has to be clear in a valid address. 
> >>> 
> >>> On most modern architectures, it's even more likely that the /least/ 
> >>> significant bit is zero, which is in fact what your implementation does 
> >>> on a LE architecture. 
> >> 
> >> That would mean that a pointer can only point to even addresses. I don't 
> >> think that's the case. Rather obviously a pointer must be able to point 
> >> to any address (else you would be able to eg. traverse a string with 
> >> a pointer). 
> >> 
> > 
> > The observation is about *heap memory* used by the containers, where you 
> > are likley not allowed to alloocate a single byte, but will get a 
> > pointer properly aligned at 8 or 16 bytes. And therefore has some zeros 
> > at the lower end.
> The topic is short-string-optimization of a std::string-like class. 
> If the internal pointer used in that class is pointing to either allocated 
> memory or an internal buffer, you can't be sure that the internal buffer 
> will be aligned to a 16-byte, 8-byte or even a 2-byte boundary. That's 
> because the class instance itself may not be allocated on the heap.
>
But only if the architecture doesn't require alignment. If a structure has a
member requiring 4 byte alignment, then the whole structure has to be 4-byte
aligned when created on the stack.

[toc] | [prev] | [next] | [standalone]


#85796

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-08-03 14:49 +0300
Message-ID<tcdnd4$272cv$1@dont-email.me>
In reply to#85792
03.08.2022 13:53 Juha Nieminen kirjutas:
> Bo Persson <bo@bo-persson.se> wrote:
>> On 2022-08-03 at 09:59, Juha Nieminen wrote:
>>> Manfred <noname@add.invalid> wrote:
>>>>> However in reality not all 64 bits of a pointer are wired to memory addresses.
>>>>> It's likely that the most significant bit has to be clear in a valid address.
>>>>
>>>> On most modern architectures, it's even more likely that the /least/
>>>> significant bit is zero, which is in fact what your implementation does
>>>> on a LE architecture.
>>>
>>> That would mean that a pointer can only point to even addresses. I don't
>>> think that's the case. Rather obviously a pointer must be able to point
>>> to any address (else you would be able to eg. traverse a string with
>>> a pointer).
>>>
>>
>> The observation is about *heap memory* used by the containers, where you
>> are likley not allowed to alloocate a single byte, but will get a
>> pointer properly aligned at 8 or 16 bytes. And therefore has some zeros
>> at the lower end.
> 
> The topic is short-string-optimization of a std::string-like class.
> If the internal pointer used in that class is pointing to either allocated
> memory or an internal buffer, you can't be sure that the internal buffer
> will be aligned to a 16-byte, 8-byte or even a 2-byte boundary. That's
> because the class instance itself may not be allocated on the heap.

Even when allocated on stack, it will be aligned properly as needed for 
the pointer or size_t members, enforcing at least 8-byte alignment in 
64-bit.

That's not to say SSO ought to use such bit-fiddling, the current SSO 
implementations apparently do no such thing (presumably because it would 
cause overhead even for non-mutable access).


[toc] | [prev] | [next] | [standalone]


#85799

FromÖö Tiib <ootiib@hot.ee>
Date2022-08-03 23:31 -0700
Message-ID<fe4f0340-2291-41e0-be00-7ec98affe779n@googlegroups.com>
In reply to#85796
On Wednesday, 3 August 2022 at 14:50:10 UTC+3, Paavo Helde wrote:
> 03.08.2022 13:53 Juha Nieminen kirjutas: 
> > Bo Persson <b...@bo-persson.se> wrote: 
> >> On 2022-08-03 at 09:59, Juha Nieminen wrote: 
> >>> Manfred <non...@add.invalid> wrote: 
> >>>>> However in reality not all 64 bits of a pointer are wired to memory addresses. 
> >>>>> It's likely that the most significant bit has to be clear in a valid address. 
> >>>> 
> >>>> On most modern architectures, it's even more likely that the /least/ 
> >>>> significant bit is zero, which is in fact what your implementation does 
> >>>> on a LE architecture. 
> >>> 
> >>> That would mean that a pointer can only point to even addresses. I don't 
> >>> think that's the case. Rather obviously a pointer must be able to point 
> >>> to any address (else you would be able to eg. traverse a string with 
> >>> a pointer). 
> >>> 
> >> 
> >> The observation is about *heap memory* used by the containers, where you 
> >> are likley not allowed to alloocate a single byte, but will get a 
> >> pointer properly aligned at 8 or 16 bytes. And therefore has some zeros 
> >> at the lower end. 
> > 
> > The topic is short-string-optimization of a std::string-like class. 
> > If the internal pointer used in that class is pointing to either allocated 
> > memory or an internal buffer, you can't be sure that the internal buffer 
> > will be aligned to a 16-byte, 8-byte or even a 2-byte boundary. That's 
> > because the class instance itself may not be allocated on the heap.
> Even when allocated on stack, it will be aligned properly as needed for 
> the pointer or size_t members, enforcing at least 8-byte alignment in 
> 64-bit. 
> 
> That's not to say SSO ought to use such bit-fiddling, the current SSO 
> implementations apparently do no such thing (presumably because it would 
> cause overhead even for non-mutable access).

Yes it could be possible to fit 30 character SSO into 32 byte std::string
but most current implementations have 15 character SSO. 
That means as it is performance optimisation they deliver performance
not storage.

[toc] | [prev] | [next] | [standalone]


#85774

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-08-01 11:29 +0000
Message-ID<tc8dfd$713$1@gioia.aioe.org>
In reply to#85766
Fred. Zwarts <F.Zwarts@kvi.nl> wrote:
>> I don't know if you are referring to it, but many people seem to think that
>> "short string optimization" makes std::string pretty much as efficient
>> as an array of char (for strings that are short enough).
>> 
>> They fail to take into consideration that short string optimization
>> requires conditionals in almost all member functons that access the
>> string data. Conditionals are not free.
> 
> 
> I wonder if that is true. The string object has a pointer to its buffer 
> and a length. I assume that if the short string optimization is used, 
> this pointer points to the internal buffer. Only member functions that 
> want to increase the size of the buffer need those conditionals. I 
> wonder whether "almost all member functions" modify the buffer size.

I wonder why they don't just outright add "static" versions of all the
data containers. As in, you specify the maximum size of the container
as a template parameter, and the side of the object itself will be
that much. Just like std::array, but containing all the member
functions of the data container. So for example you could have a
"static" std::string of a maximum length that you specify as a
template parameter, which does no dynamic memory allocations.

(Of course nothing stops me from implementing such a thing
myself, but...)

[toc] | [prev] | [next] | [standalone]


#85777

FromBo Persson <bo@bo-persson.se>
Date2022-08-01 14:39 +0200
Message-ID<jkpvrkF1d0oU1@mid.individual.net>
In reply to#85774
On 2022-08-01 at 13:29, Juha Nieminen wrote:
> Fred. Zwarts <F.Zwarts@kvi.nl> wrote:
>>> I don't know if you are referring to it, but many people seem to think that
>>> "short string optimization" makes std::string pretty much as efficient
>>> as an array of char (for strings that are short enough).
>>>
>>> They fail to take into consideration that short string optimization
>>> requires conditionals in almost all member functons that access the
>>> string data. Conditionals are not free.
>>
>>
>> I wonder if that is true. The string object has a pointer to its buffer
>> and a length. I assume that if the short string optimization is used,
>> this pointer points to the internal buffer. Only member functions that
>> want to increase the size of the buffer need those conditionals. I
>> wonder whether "almost all member functions" modify the buffer size.
> 
> I wonder why they don't just outright add "static" versions of all the
> data containers. As in, you specify the maximum size of the container
> as a template parameter, and the side of the object itself will be
> that much. Just like std::array, but containing all the member
> functions of the data container. So for example you could have a
> "static" std::string of a maximum length that you specify as a
> template parameter, which does no dynamic memory allocations.
> 
> (Of course nothing stops me from implementing such a thing
> myself, but...)


One inconvenience would be that each different length would be a 
separate type. That would make many string ops really awkward, if it 
changes type halfway through.

You can already reserve() the space for a std::string, at the cost of a 
single dynamic allocation.

[toc] | [prev] | [next] | [standalone]


#85782

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-08-02 05:43 +0000
Message-ID<tcadi4$3g0$1@gioia.aioe.org>
In reply to#85777
Bo Persson <bo@bo-persson.se> wrote:
> You can already reserve() the space for a std::string, at the cost of a 
> single dynamic allocation.

The main idea would be that if it's eg. a member variable of a class,
its maximum length is something like 20 characters, and the class is
instantiated millions of times, it will save a lot of memory, cause
less memory fragmentation and be more efficient.

But yeah, in such situations it would probably not be a huge amount
of trouble to create a custom implementation.

[toc] | [prev] | [next] | [standalone]


#85784

FromManfred <noname@add.invalid>
Date2022-08-03 01:13 +0200
Message-ID<tccb2c$62b$1@gioia.aioe.org>
In reply to#85764
On 7/31/2022 4:39 PM, Juha Nieminen wrote:
> Manfred <noname@add.invalid> wrote:
>> That, and below, is generally true for a generic dynamic container.
>> However, std::string is not a generic container. It is a very specific
>> and very optimized container of chars.
>> So, if you use a decent implementation of std the performance of
>> std::string and char* is usually pretty close.
> 
> I don't know if you are referring to it, but many people seem to think that
> "short string optimization" makes std::string pretty much as efficient
> as an array of char (for strings that are short enough).
> 
> They fail to take into consideration that short string optimization
> requires conditionals in almost all member functons that access the
> string data. Conditionals are not free.

It is more general than plain SSO.

My point is that std::string is different from a generic container when 
it comes to performance, because implementors know two things:
1) std::string is made to contain char's (and, more rarely, wchar_t's) - 
it's not made to contain other stuff (well, maybe technically it can, 
but then it can happily perform poorly)
2) users are more eager to get performance out of std::string than out 
of say std::vector.

So, SSO is /one/ of the technologies that implementors can use to 
improve performance, but they (expecially the good ones) are known to be 
quite creative when it comes to performance. SSO is a good example to 
explain my point, though: while it makes sense for strings, I guess very 
few, if anyone, would be willing to pay for some sort of SVO (Short 
Vector Optimization)

[toc] | [prev] | [next] | [standalone]


#85787

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-08-03 09:34 +0300
Message-ID<tcd4uk$22gbl$1@dont-email.me>
In reply to#85784
03.08.2022 02:13 Manfred kirjutas:
> On 7/31/2022 4:39 PM, Juha Nieminen wrote:
>> Manfred <noname@add.invalid> wrote:
>>> That, and below, is generally true for a generic dynamic container.
>>> However, std::string is not a generic container. It is a very specific
>>> and very optimized container of chars.
>>> So, if you use a decent implementation of std the performance of
>>> std::string and char* is usually pretty close.
>>
>> I don't know if you are referring to it, but many people seem to think 
>> that
>> "short string optimization" makes std::string pretty much as efficient
>> as an array of char (for strings that are short enough).
>>
>> They fail to take into consideration that short string optimization
>> requires conditionals in almost all member functons that access the
>> string data. Conditionals are not free.
> 
> It is more general than plain SSO.
> 
> My point is that std::string is different from a generic container when 
> it comes to performance, because implementors know two things:
> 1) std::string is made to contain char's (and, more rarely, wchar_t's) - 
> it's not made to contain other stuff (well, maybe technically it can, 
> but then it can happily perform poorly)
> 2) users are more eager to get performance out of std::string than out 
> of say std::vector.
> 
> So, SSO is /one/ of the technologies that implementors can use to 
> improve performance, but they (expecially the good ones) are known to be 
> quite creative when it comes to performance. SSO is a good example to 
> explain my point, though: while it makes sense for strings, I guess very 
> few, if anyone, would be willing to pay for some sort of SVO (Short 
> Vector Optimization)

Right, SSO makes more sense for strings, mostly because the character 
size is small, compared to a typical vector element size (as long as one 
is using ASCII or UTF-8 and not something like std::u32string).

 From what I see in common C++ implementations, sizeof(std::string) is 
32, from which 16 bytes gets reused as the SSO buffer, meaning that SSO 
applies to strings up to length 15. That's something.

Sizeof(std::vector) seems to be something like 24, from which 8 bytes 
could be used as "SVO" buffer, meaning that this optimization would 
apply for vectors of max 1-2 elements only. Not so useful.

[toc] | [prev] | [next] | [standalone]


#85733

FromMuttley@dastardlyhq.com
Date2022-07-29 09:22 +0000
Message-ID<tc08s1$1u7$1@gioia.aioe.org>
In reply to#85723
On Thu, 28 Jul 2022 20:03:41 +0300
Paavo Helde <eesnimi@osa.pri.ee> wrote:
>28.07.2022 18:11 Muttley@dastardlyhq.com kirjutas:
>> On Thu, 28 Jul 2022 15:04:12 +0200
>> Bo Persson <bo@bo-persson.se> wrote:
>>
>>> std::string full_name = first_name + last_name;
>>>
>>> I wouldn't expect a floating point addition.
>> 
>> Unfortunately even in C++ 2020 that won't compile unless first_name is a
>> string and won't work if its char* due to precendence rules.
>
>Why on earth are you using char* in a C++ program? The only thing they 

Is that supposed to be a serious question?

>bring is problems and slowdowns (because string length needs to be 
>re-calculated all the time).

Depends on what you're doing with it. You do realise char* don't always
just point to a text string?

>#include <string>
>using namespace std::string_literals;
>
>int main() {
>     auto first_name = "Bernhard"s, last_name = "Shaw"s;
>     std::string full_name = first_name + last_name;

Hmm, I wonder whats more efficient. Counting the length of 2 short strings or
creating 2 objects...

[toc] | [prev] | [next] | [standalone]


Page 4 of 5 — ← Prev page 1 2 3 [4] 5  Next page →

Back to top | Article view | comp.lang.c++


csiph-web