Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #85614 > unrolled thread
| Started by | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| First post | 2022-07-25 13:14 +0000 |
| Last post | 2022-07-28 06:16 -0700 |
| Articles | 20 on this page of 84 — 18 participants |
Back to article view | Back to comp.lang.c++
Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-25 13:14 +0000
Re: Differences between C and C++ "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-07-25 16:18 +0200
Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-25 18:47 +0300
Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-25 19:49 +0300
Re: Differences between C and C++ Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-07-25 19:38 +0100
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-25 08:51 -0700
Re: Differences between C and C++ Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-07-25 17:14 +0100
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-25 16:19 +0000
Re: Differences between C and C++ Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-07-25 19:28 +0100
Re: Differences between C and C++ Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-07-25 12:21 -0700
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-26 07:54 +0000
Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-26 08:00 +0000
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-26 08:09 +0000
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-26 02:42 -0700
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-26 15:29 +0000
Re: Differences between C and C++ scott@slp53.sl.home (Scott Lurndal) - 2022-07-26 16:30 +0000
Re: Differences between C and C++ Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-07-26 11:59 -0700
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-27 07:52 +0000
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-27 01:56 -0700
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-27 14:51 +0000
Re: Differences between C and C++ Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-07-27 11:11 -0700
Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-26 06:09 +0000
Re: Differences between C and C++ Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-07-27 17:33 +0100
Re: Differences between C and C++ Chris Vine <chris@cvine--nospam--.freeserve.co.uk> - 2022-07-25 17:45 +0100
Re: Differences between C and C++ Paul N <gw7rib@aol.com> - 2022-07-27 10:09 -0700
Re: Differences between C and C++ Bo Persson <bo@bo-persson.se> - 2022-07-27 19:28 +0200
Re: Differences between C and C++ Paul N <gw7rib@aol.com> - 2022-07-27 11:43 -0700
Re: Differences between C and C++ David Brown <david.brown@hesbynett.no> - 2022-07-29 18:04 +0200
Re: Differences between C and C++ Richard Damon <Richard@Damon-Family.org> - 2022-07-29 18:47 -0400
Re: Differences between C and C++ David Brown <david.brown@hesbynett.no> - 2022-07-30 14:08 +0200
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-28 08:03 +0000
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-28 01:57 -0700
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-28 09:04 +0000
Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-28 06:53 +0000
Re: Differences between C and C++ Richard Damon <Richard@Damon-Family.org> - 2022-07-28 07:33 -0400
Re: Differences between C and C++ Bo Persson <bo@bo-persson.se> - 2022-07-28 15:04 +0200
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-28 15:11 +0000
Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-28 20:03 +0300
Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-29 07:56 +0000
Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-29 12:03 +0300
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-29 09:28 +0000
Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-29 12:27 +0000
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-29 05:43 -0700
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-29 14:14 +0000
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-29 08:25 -0700
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-29 15:48 +0000
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-29 13:12 -0700
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-30 09:28 +0000
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-30 03:39 -0700
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-30 14:23 +0000
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-30 19:07 -0700
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-31 07:22 +0000
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-31 06:07 -0700
Re: Differences between C and C++ muttley@dastardlyhq.com - 2022-07-31 16:36 +0000
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-08-01 01:12 -0700
Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-29 22:47 +0000
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-30 03:48 -0700
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-07-29 05:16 -0700
Re: Differences between C and C++ Manfred <noname@add.invalid> - 2022-07-30 02:14 +0200
Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-07-31 14:39 +0000
Re: Differences between C and C++ "Fred. Zwarts" <F.Zwarts@KVI.nl> - 2022-07-31 18:49 +0200
Re: Differences between C and C++ Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-07-31 14:49 -0700
Re: Differences between C and C++ Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-07-31 17:03 -0700
Re: Differences between C and C++ Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-08-01 00:16 -0700
Re: Differences between C and C++ Ike Naar <ike@sdf.org> - 2022-08-01 07:27 +0000
Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-08-01 10:47 +0300
Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-08-01 10:34 +0300
Re: Differences between C and C++ Manfred <noname@add.invalid> - 2022-08-03 01:20 +0200
Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-08-03 07:59 +0000
Re: Differences between C and C++ Bo Persson <bo@bo-persson.se> - 2022-08-03 12:45 +0200
Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-08-03 10:53 +0000
Re: Differences between C and C++ Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-08-03 03:58 -0700
Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-08-03 14:49 +0300
Re: Differences between C and C++ Öö Tiib <ootiib@hot.ee> - 2022-08-03 23:31 -0700
Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-08-01 11:29 +0000
Re: Differences between C and C++ Bo Persson <bo@bo-persson.se> - 2022-08-01 14:39 +0200
Re: Differences between C and C++ Juha Nieminen <nospam@thanks.invalid> - 2022-08-02 05:43 +0000
Re: Differences between C and C++ Manfred <noname@add.invalid> - 2022-08-03 01:13 +0200
Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-08-03 09:34 +0300
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-29 09:22 +0000
Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-29 13:53 +0300
Re: Differences between C and C++ Paavo Helde <eesnimi@osa.pri.ee> - 2022-07-29 14:13 +0300
Re: Differences between C and C++ Muttley@dastardlyhq.com - 2022-07-29 14:11 +0000
Re: Differences between C and C++ Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-07-28 06:16 -0700
Page 4 of 5 — ← Prev page 1 2 3 [4] 5 Next page →
| From | "Fred. Zwarts" <F.Zwarts@KVI.nl> |
|---|---|
| Date | 2022-07-31 18:49 +0200 |
| Message-ID | <tc6bq1$1r17$1@gioia.aioe.org> |
| In reply to | #85764 |
Op 31.jul..2022 om 16:39 schreef Juha Nieminen: > Manfred <noname@add.invalid> wrote: >> That, and below, is generally true for a generic dynamic container. >> However, std::string is not a generic container. It is a very specific >> and very optimized container of chars. >> So, if you use a decent implementation of std the performance of >> std::string and char* is usually pretty close. > > I don't know if you are referring to it, but many people seem to think that > "short string optimization" makes std::string pretty much as efficient > as an array of char (for strings that are short enough). > > They fail to take into consideration that short string optimization > requires conditionals in almost all member functons that access the > string data. Conditionals are not free. I wonder if that is true. The string object has a pointer to its buffer and a length. I assume that if the short string optimization is used, this pointer points to the internal buffer. Only member functions that want to increase the size of the buffer need those conditionals. I wonder whether "almost all member functions" modify the buffer size.
[toc] | [prev] | [next] | [standalone]
| From | Malcolm McLean <malcolm.arthur.mclean@gmail.com> |
|---|---|
| Date | 2022-07-31 14:49 -0700 |
| Message-ID | <8699b402-e8e2-4e85-8de6-d67249cbeca9n@googlegroups.com> |
| In reply to | #85766 |
On Sunday, 31 July 2022 at 17:49:22 UTC+1, F.Zwarts wrote:
> Op 31.jul..2022 om 16:39 schreef Juha Nieminen:
> > Manfred <non...@add.invalid> wrote:
> >> That, and below, is generally true for a generic dynamic container.
> >> However, std::string is not a generic container. It is a very specific
> >> and very optimized container of chars.
> >> So, if you use a decent implementation of std the performance of
> >> std::string and char* is usually pretty close.
> >
> > I don't know if you are referring to it, but many people seem to think that
> > "short string optimization" makes std::string pretty much as efficient
> > as an array of char (for strings that are short enough).
> >
> > They fail to take into consideration that short string optimization
> > requires conditionals in almost all member functons that access the
> > string data. Conditionals are not free.
> I wonder if that is true. The string object has a pointer to its buffer
> and a length. I assume that if the short string optimization is used,
> this pointer points to the internal buffer. Only member functions that
> want to increase the size of the buffer need those conditionals. I
> wonder whether "almost all member functions" modify the buffer size.
>
The std::string is (in C)
struct string
{
char *buff;
size_t len;
}
so 16 bytes on a 64-bit machine.
However in reality not all 64 bits of a pointer are wired to memory addresses.
It's likely that the most significant bit has to be clear in a valid address.
So we can exploit this by (in C)
struct shortstring
{
bool flag; // set when the shortstring is valid
int pad:7; // Maybe use for length or encoding or other stuff
char data[15]; // UP to 14 characters of ASCII string data.
};
union
{
struct string s;
struct shortstring ss;
} std_string;
Now string::size is implemented as
if(s->ss.flag)
return strlen(s->ss.data);
else
return s->s.len;
The other string member functions are implemented similarly.
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2022-07-31 17:03 -0700 |
| Message-ID | <871qu0x21j.fsf@nosuchdomain.example.com> |
| In reply to | #85767 |
Malcolm McLean <malcolm.arthur.mclean@gmail.com> writes:
> On Sunday, 31 July 2022 at 17:49:22 UTC+1, F.Zwarts wrote:
>> Op 31.jul..2022 om 16:39 schreef Juha Nieminen:
>> > Manfred <non...@add.invalid> wrote:
>> >> That, and below, is generally true for a generic dynamic container.
>> >> However, std::string is not a generic container. It is a very specific
>> >> and very optimized container of chars.
>> >> So, if you use a decent implementation of std the performance of
>> >> std::string and char* is usually pretty close.
>> >
>> > I don't know if you are referring to it, but many people seem to think that
>> > "short string optimization" makes std::string pretty much as efficient
>> > as an array of char (for strings that are short enough).
>> >
>> > They fail to take into consideration that short string optimization
>> > requires conditionals in almost all member functons that access the
>> > string data. Conditionals are not free.
>> I wonder if that is true. The string object has a pointer to its buffer
>> and a length. I assume that if the short string optimization is used,
>> this pointer points to the internal buffer. Only member functions that
>> want to increase the size of the buffer need those conditionals. I
>> wonder whether "almost all member functions" modify the buffer size.
>>
> The std::string is (in C)
> struct string
> {
> char *buff;
> size_t len;
> }
> so 16 bytes on a 64-bit machine.
That's one possible implementation.
> However in reality not all 64 bits of a pointer are wired to memory addresses.
> It's likely that the most significant bit has to be clear in a valid address.
> So we can exploit this by (in C)
> struct shortstring
> {
> bool flag; // set when the shortstring is valid
> int pad:7; // Maybe use for length or encoding or other stuff
> char data[15]; // UP to 14 characters of ASCII string data.
> };
>
> union
> {
> struct string s;
> struct shortstring ss;
> } std_string;
>
> Now string::size is implemented as
> if(s->ss.flag)
> return strlen(s->ss.data);
> else
> return s->s.len;
>
>
> The other string member functions are implemented similarly.
That won't work without extra code to ensure that the short string
optimization isn't used in all cases. The length of a std::string
is not determined by a null character. This program must print 3 :
#include <iostream>
#include <string>
int main() {
std::string s = "a";
s += '\0';
s += 'b';
std::cout << s.size() << '\n';
}
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
Working, but not speaking, for Philips
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | Malcolm McLean <malcolm.arthur.mclean@gmail.com> |
|---|---|
| Date | 2022-08-01 00:16 -0700 |
| Message-ID | <96324449-0709-4099-a262-d44bc65d27e9n@googlegroups.com> |
| In reply to | #85768 |
On Monday, 1 August 2022 at 01:04:09 UTC+1, Keith Thompson wrote:
> Malcolm McLean <malcolm.ar...@gmail.com> writes:
> > On Sunday, 31 July 2022 at 17:49:22 UTC+1, F.Zwarts wrote:
> >> Op 31.jul..2022 om 16:39 schreef Juha Nieminen:
> >> > Manfred <non...@add.invalid> wrote:
> >> >> That, and below, is generally true for a generic dynamic container.
> >> >> However, std::string is not a generic container. It is a very specific
> >> >> and very optimized container of chars.
> >> >> So, if you use a decent implementation of std the performance of
> >> >> std::string and char* is usually pretty close.
> >> >
> >> > I don't know if you are referring to it, but many people seem to think that
> >> > "short string optimization" makes std::string pretty much as efficient
> >> > as an array of char (for strings that are short enough).
> >> >
> >> > They fail to take into consideration that short string optimization
> >> > requires conditionals in almost all member functons that access the
> >> > string data. Conditionals are not free.
> >> I wonder if that is true. The string object has a pointer to its buffer
> >> and a length. I assume that if the short string optimization is used,
> >> this pointer points to the internal buffer. Only member functions that
> >> want to increase the size of the buffer need those conditionals. I
> >> wonder whether "almost all member functions" modify the buffer size.
> >>
> > The std::string is (in C)
> > struct string
> > {
> > char *buff;
> > size_t len;
> > }
> > so 16 bytes on a 64-bit machine.
> That's one possible implementation.
> > However in reality not all 64 bits of a pointer are wired to memory addresses.
> > It's likely that the most significant bit has to be clear in a valid address.
> > So we can exploit this by (in C)
> > struct shortstring
> > {
> > bool flag; // set when the shortstring is valid
> > int pad:7; // Maybe use for length or encoding or other stuff
> > char data[15]; // UP to 14 characters of ASCII string data.
> > };
> >
> > union
> > {
> > struct string s;
> > struct shortstring ss;
> > } std_string;
> >
> > Now string::size is implemented as
> > if(s->ss.flag)
> > return strlen(s->ss.data);
> > else
> > return s->s.len;
> >
> >
> > The other string member functions are implemented similarly.
> That won't work without extra code to ensure that the short string
> optimization isn't used in all cases. The length of a std::string
> is not determined by a null character. This program must print 3 :
>
> #include <iostream>
> #include <string>
>
> int main() {
> std::string s = "a";
> s += '\0';
> s += 'b';
> std::cout << s.size() << '\n';
> }
>
You're right. std::size includes embedded nuls. So in fact you need to use four bits
for the "pad" field to store the length. Or nul-pad the string and work backwards
until you hit a set byte.
[toc] | [prev] | [next] | [standalone]
| From | Ike Naar <ike@sdf.org> |
|---|---|
| Date | 2022-08-01 07:27 +0000 |
| Message-ID | <slrntef039.g3m.ike@sdf.org> |
| In reply to | #85769 |
On 2022-08-01, Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote: > You're right. std::size includes embedded nuls. So in fact you need to use four bits > for the "pad" field to store the length. Or nul-pad the string and work backwards > until you hit a set byte. Working backwards will give the wrong result if there are embedded nuls at the end of the string.
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-08-01 10:47 +0300 |
| Message-ID | <tc80ep$peih$1@dont-email.me> |
| In reply to | #85769 |
01.08.2022 10:16 Malcolm McLean kirjutas:
> On Monday, 1 August 2022 at 01:04:09 UTC+1, Keith Thompson wrote:
>> Malcolm McLean <malcolm.ar...@gmail.com> writes:
>>> On Sunday, 31 July 2022 at 17:49:22 UTC+1, F.Zwarts wrote:
>>>> Op 31.jul..2022 om 16:39 schreef Juha Nieminen:
>>>>> Manfred <non...@add.invalid> wrote:
>>>>>> That, and below, is generally true for a generic dynamic container.
>>>>>> However, std::string is not a generic container. It is a very specific
>>>>>> and very optimized container of chars.
>>>>>> So, if you use a decent implementation of std the performance of
>>>>>> std::string and char* is usually pretty close.
>>>>>
>>>>> I don't know if you are referring to it, but many people seem to think that
>>>>> "short string optimization" makes std::string pretty much as efficient
>>>>> as an array of char (for strings that are short enough).
>>>>>
>>>>> They fail to take into consideration that short string optimization
>>>>> requires conditionals in almost all member functons that access the
>>>>> string data. Conditionals are not free.
>>>> I wonder if that is true. The string object has a pointer to its buffer
>>>> and a length. I assume that if the short string optimization is used,
>>>> this pointer points to the internal buffer. Only member functions that
>>>> want to increase the size of the buffer need those conditionals. I
>>>> wonder whether "almost all member functions" modify the buffer size.
>>>>
>>> The std::string is (in C)
>>> struct string
>>> {
>>> char *buff;
>>> size_t len;
>>> }
>>> so 16 bytes on a 64-bit machine.
>> That's one possible implementation.
>>> However in reality not all 64 bits of a pointer are wired to memory addresses.
>>> It's likely that the most significant bit has to be clear in a valid address.
>>> So we can exploit this by (in C)
>>> struct shortstring
>>> {
>>> bool flag; // set when the shortstring is valid
>>> int pad:7; // Maybe use for length or encoding or other stuff
>>> char data[15]; // UP to 14 characters of ASCII string data.
>>> };
>>>
>>> union
>>> {
>>> struct string s;
>>> struct shortstring ss;
>>> } std_string;
>>>
>>> Now string::size is implemented as
>>> if(s->ss.flag)
>>> return strlen(s->ss.data);
>>> else
>>> return s->s.len;
>>>
>>>
>>> The other string member functions are implemented similarly.
>> That won't work without extra code to ensure that the short string
>> optimization isn't used in all cases. The length of a std::string
>> is not determined by a null character. This program must print 3 :
>>
>> #include <iostream>
>> #include <string>
>>
>> int main() {
>> std::string s = "a";
>> s += '\0';
>> s += 'b';
>> std::cout << s.size() << '\n';
>> }
>>
> You're right. std::size includes embedded nuls. So in fact you need to use four bits
> for the "pad" field to store the length. Or nul-pad the string and work backwards
> until you hit a set byte.
Or check for the presence of embedded nuls and switch off SSO for such
strings. Alas, that would mean non-zero overhead.
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-08-01 10:34 +0300 |
| Message-ID | <tc7vm4$p7sr$1@dont-email.me> |
| In reply to | #85767 |
01.08.2022 00:49 Malcolm McLean kirjutas:
> The std::string is (in C)
> struct string
> {
> char *buff;
> size_t len;
> }
> so 16 bytes on a 64-bit machine.
In reality some popular implementations have sizeof(std::string)==32
(e.g. MSVC2019 on Windows, gcc 8.3 on Linux). When trying to study the
internals I see lots of attention to allocators. Indeed, all string
constructors also take an allocator argument which must be stored somewhere.
With 32 bytes, there would be a lot more room for SSO. Unfortunately it
seems the implementations are not very eager to make use of that space,
I think only 16 bytes gets typically used as SSO. Presumably they have
the same char* buff and size_t len in an SSO string, it's just that the
char* buff pointer just points 16 bytes ahead, into the internal SSO
buffer of 16 bytes. This ensures there would be zero overhead for SSO
and no conditionals involved for non-mutable access.
[toc] | [prev] | [next] | [standalone]
| From | Manfred <noname@add.invalid> |
|---|---|
| Date | 2022-08-03 01:20 +0200 |
| Message-ID | <tccbff$aah$1@gioia.aioe.org> |
| In reply to | #85767 |
On 7/31/2022 11:49 PM, Malcolm McLean wrote:
> On Sunday, 31 July 2022 at 17:49:22 UTC+1, F.Zwarts wrote:
>> Op 31.jul..2022 om 16:39 schreef Juha Nieminen:
>>> Manfred <non...@add.invalid> wrote:
>>>> That, and below, is generally true for a generic dynamic container.
>>>> However, std::string is not a generic container. It is a very specific
>>>> and very optimized container of chars.
>>>> So, if you use a decent implementation of std the performance of
>>>> std::string and char* is usually pretty close.
>>>
>>> I don't know if you are referring to it, but many people seem to think that
>>> "short string optimization" makes std::string pretty much as efficient
>>> as an array of char (for strings that are short enough).
>>>
>>> They fail to take into consideration that short string optimization
>>> requires conditionals in almost all member functons that access the
>>> string data. Conditionals are not free.
>> I wonder if that is true. The string object has a pointer to its buffer
>> and a length. I assume that if the short string optimization is used,
>> this pointer points to the internal buffer. Only member functions that
>> want to increase the size of the buffer need those conditionals. I
>> wonder whether "almost all member functions" modify the buffer size.
>>
> The std::string is (in C)
> struct string
> {
> char *buff;
> size_t len;
> }
> so 16 bytes on a 64-bit machine.
> However in reality not all 64 bits of a pointer are wired to memory addresses.
> It's likely that the most significant bit has to be clear in a valid address.
On most modern architectures, it's even more likely that the /least/
significant bit is zero, which is in fact what your implementation does
on a LE architecture.
> So we can exploit this by (in C)
> struct shortstring
> {
> bool flag; // set when the shortstring is valid
> int pad:7; // Maybe use for length or encoding or other stuff
> char data[15]; // UP to 14 characters of ASCII string data.
> };
>
> union
> {
> struct string s;
> struct shortstring ss;
> } std_string;
>
> Now string::size is implemented as
> if(s->ss.flag)
> return strlen(s->ss.data);
> else
> return s->s.len;
>
>
> The other string member functions are implemented similarly.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2022-08-03 07:59 +0000 |
| Message-ID | <tcd9te$19s4$1@gioia.aioe.org> |
| In reply to | #85785 |
Manfred <noname@add.invalid> wrote: >> However in reality not all 64 bits of a pointer are wired to memory addresses. >> It's likely that the most significant bit has to be clear in a valid address. > > On most modern architectures, it's even more likely that the /least/ > significant bit is zero, which is in fact what your implementation does > on a LE architecture. That would mean that a pointer can only point to even addresses. I don't think that's the case. Rather obviously a pointer must be able to point to any address (else you would be able to eg. traverse a string with a pointer). IIRC in Linux the kernel reserves the upper half of the memory range to the kernel space and the lower half to user space (and IIRC user code has no business using any kernel space pointer for any reason in any situation) so it may be that at least in Linux the most-significant bit of pointers in user code is indeed always zero. However, I would never write code that makes this assumptions because it's not really an assumption you can make at any level (I don't think even the Linux kernel makes that absolute guarantee).
[toc] | [prev] | [next] | [standalone]
| From | Bo Persson <bo@bo-persson.se> |
|---|---|
| Date | 2022-08-03 12:45 +0200 |
| Message-ID | <jkv1ucFqa8tU1@mid.individual.net> |
| In reply to | #85788 |
On 2022-08-03 at 09:59, Juha Nieminen wrote: > Manfred <noname@add.invalid> wrote: >>> However in reality not all 64 bits of a pointer are wired to memory addresses. >>> It's likely that the most significant bit has to be clear in a valid address. >> >> On most modern architectures, it's even more likely that the /least/ >> significant bit is zero, which is in fact what your implementation does >> on a LE architecture. > > That would mean that a pointer can only point to even addresses. I don't > think that's the case. Rather obviously a pointer must be able to point > to any address (else you would be able to eg. traverse a string with > a pointer). > The observation is about *heap memory* used by the containers, where you are likley not allowed to alloocate a single byte, but will get a pointer properly aligned at 8 or 16 bytes. And therefore has some zeros at the lower end.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2022-08-03 10:53 +0000 |
| Message-ID | <tcdk2t$1dc1$2@gioia.aioe.org> |
| In reply to | #85790 |
Bo Persson <bo@bo-persson.se> wrote: > On 2022-08-03 at 09:59, Juha Nieminen wrote: >> Manfred <noname@add.invalid> wrote: >>>> However in reality not all 64 bits of a pointer are wired to memory addresses. >>>> It's likely that the most significant bit has to be clear in a valid address. >>> >>> On most modern architectures, it's even more likely that the /least/ >>> significant bit is zero, which is in fact what your implementation does >>> on a LE architecture. >> >> That would mean that a pointer can only point to even addresses. I don't >> think that's the case. Rather obviously a pointer must be able to point >> to any address (else you would be able to eg. traverse a string with >> a pointer). >> > > The observation is about *heap memory* used by the containers, where you > are likley not allowed to alloocate a single byte, but will get a > pointer properly aligned at 8 or 16 bytes. And therefore has some zeros > at the lower end. The topic is short-string-optimization of a std::string-like class. If the internal pointer used in that class is pointing to either allocated memory or an internal buffer, you can't be sure that the internal buffer will be aligned to a 16-byte, 8-byte or even a 2-byte boundary. That's because the class instance itself may not be allocated on the heap.
[toc] | [prev] | [next] | [standalone]
| From | Malcolm McLean <malcolm.arthur.mclean@gmail.com> |
|---|---|
| Date | 2022-08-03 03:58 -0700 |
| Message-ID | <4b8578a8-cc4d-434c-9ec9-3730ece082bcn@googlegroups.com> |
| In reply to | #85792 |
On Wednesday, 3 August 2022 at 11:53:34 UTC+1, Juha Nieminen wrote: > Bo Persson <b...@bo-persson.se> wrote: > > On 2022-08-03 at 09:59, Juha Nieminen wrote: > >> Manfred <non...@add.invalid> wrote: > >>>> However in reality not all 64 bits of a pointer are wired to memory addresses. > >>>> It's likely that the most significant bit has to be clear in a valid address. > >>> > >>> On most modern architectures, it's even more likely that the /least/ > >>> significant bit is zero, which is in fact what your implementation does > >>> on a LE architecture. > >> > >> That would mean that a pointer can only point to even addresses. I don't > >> think that's the case. Rather obviously a pointer must be able to point > >> to any address (else you would be able to eg. traverse a string with > >> a pointer). > >> > > > > The observation is about *heap memory* used by the containers, where you > > are likley not allowed to alloocate a single byte, but will get a > > pointer properly aligned at 8 or 16 bytes. And therefore has some zeros > > at the lower end. > The topic is short-string-optimization of a std::string-like class. > If the internal pointer used in that class is pointing to either allocated > memory or an internal buffer, you can't be sure that the internal buffer > will be aligned to a 16-byte, 8-byte or even a 2-byte boundary. That's > because the class instance itself may not be allocated on the heap. > But only if the architecture doesn't require alignment. If a structure has a member requiring 4 byte alignment, then the whole structure has to be 4-byte aligned when created on the stack.
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-08-03 14:49 +0300 |
| Message-ID | <tcdnd4$272cv$1@dont-email.me> |
| In reply to | #85792 |
03.08.2022 13:53 Juha Nieminen kirjutas: > Bo Persson <bo@bo-persson.se> wrote: >> On 2022-08-03 at 09:59, Juha Nieminen wrote: >>> Manfred <noname@add.invalid> wrote: >>>>> However in reality not all 64 bits of a pointer are wired to memory addresses. >>>>> It's likely that the most significant bit has to be clear in a valid address. >>>> >>>> On most modern architectures, it's even more likely that the /least/ >>>> significant bit is zero, which is in fact what your implementation does >>>> on a LE architecture. >>> >>> That would mean that a pointer can only point to even addresses. I don't >>> think that's the case. Rather obviously a pointer must be able to point >>> to any address (else you would be able to eg. traverse a string with >>> a pointer). >>> >> >> The observation is about *heap memory* used by the containers, where you >> are likley not allowed to alloocate a single byte, but will get a >> pointer properly aligned at 8 or 16 bytes. And therefore has some zeros >> at the lower end. > > The topic is short-string-optimization of a std::string-like class. > If the internal pointer used in that class is pointing to either allocated > memory or an internal buffer, you can't be sure that the internal buffer > will be aligned to a 16-byte, 8-byte or even a 2-byte boundary. That's > because the class instance itself may not be allocated on the heap. Even when allocated on stack, it will be aligned properly as needed for the pointer or size_t members, enforcing at least 8-byte alignment in 64-bit. That's not to say SSO ought to use such bit-fiddling, the current SSO implementations apparently do no such thing (presumably because it would cause overhead even for non-mutable access).
[toc] | [prev] | [next] | [standalone]
| From | Öö Tiib <ootiib@hot.ee> |
|---|---|
| Date | 2022-08-03 23:31 -0700 |
| Message-ID | <fe4f0340-2291-41e0-be00-7ec98affe779n@googlegroups.com> |
| In reply to | #85796 |
On Wednesday, 3 August 2022 at 14:50:10 UTC+3, Paavo Helde wrote: > 03.08.2022 13:53 Juha Nieminen kirjutas: > > Bo Persson <b...@bo-persson.se> wrote: > >> On 2022-08-03 at 09:59, Juha Nieminen wrote: > >>> Manfred <non...@add.invalid> wrote: > >>>>> However in reality not all 64 bits of a pointer are wired to memory addresses. > >>>>> It's likely that the most significant bit has to be clear in a valid address. > >>>> > >>>> On most modern architectures, it's even more likely that the /least/ > >>>> significant bit is zero, which is in fact what your implementation does > >>>> on a LE architecture. > >>> > >>> That would mean that a pointer can only point to even addresses. I don't > >>> think that's the case. Rather obviously a pointer must be able to point > >>> to any address (else you would be able to eg. traverse a string with > >>> a pointer). > >>> > >> > >> The observation is about *heap memory* used by the containers, where you > >> are likley not allowed to alloocate a single byte, but will get a > >> pointer properly aligned at 8 or 16 bytes. And therefore has some zeros > >> at the lower end. > > > > The topic is short-string-optimization of a std::string-like class. > > If the internal pointer used in that class is pointing to either allocated > > memory or an internal buffer, you can't be sure that the internal buffer > > will be aligned to a 16-byte, 8-byte or even a 2-byte boundary. That's > > because the class instance itself may not be allocated on the heap. > Even when allocated on stack, it will be aligned properly as needed for > the pointer or size_t members, enforcing at least 8-byte alignment in > 64-bit. > > That's not to say SSO ought to use such bit-fiddling, the current SSO > implementations apparently do no such thing (presumably because it would > cause overhead even for non-mutable access). Yes it could be possible to fit 30 character SSO into 32 byte std::string but most current implementations have 15 character SSO. That means as it is performance optimisation they deliver performance not storage.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2022-08-01 11:29 +0000 |
| Message-ID | <tc8dfd$713$1@gioia.aioe.org> |
| In reply to | #85766 |
Fred. Zwarts <F.Zwarts@kvi.nl> wrote: >> I don't know if you are referring to it, but many people seem to think that >> "short string optimization" makes std::string pretty much as efficient >> as an array of char (for strings that are short enough). >> >> They fail to take into consideration that short string optimization >> requires conditionals in almost all member functons that access the >> string data. Conditionals are not free. > > > I wonder if that is true. The string object has a pointer to its buffer > and a length. I assume that if the short string optimization is used, > this pointer points to the internal buffer. Only member functions that > want to increase the size of the buffer need those conditionals. I > wonder whether "almost all member functions" modify the buffer size. I wonder why they don't just outright add "static" versions of all the data containers. As in, you specify the maximum size of the container as a template parameter, and the side of the object itself will be that much. Just like std::array, but containing all the member functions of the data container. So for example you could have a "static" std::string of a maximum length that you specify as a template parameter, which does no dynamic memory allocations. (Of course nothing stops me from implementing such a thing myself, but...)
[toc] | [prev] | [next] | [standalone]
| From | Bo Persson <bo@bo-persson.se> |
|---|---|
| Date | 2022-08-01 14:39 +0200 |
| Message-ID | <jkpvrkF1d0oU1@mid.individual.net> |
| In reply to | #85774 |
On 2022-08-01 at 13:29, Juha Nieminen wrote: > Fred. Zwarts <F.Zwarts@kvi.nl> wrote: >>> I don't know if you are referring to it, but many people seem to think that >>> "short string optimization" makes std::string pretty much as efficient >>> as an array of char (for strings that are short enough). >>> >>> They fail to take into consideration that short string optimization >>> requires conditionals in almost all member functons that access the >>> string data. Conditionals are not free. >> >> >> I wonder if that is true. The string object has a pointer to its buffer >> and a length. I assume that if the short string optimization is used, >> this pointer points to the internal buffer. Only member functions that >> want to increase the size of the buffer need those conditionals. I >> wonder whether "almost all member functions" modify the buffer size. > > I wonder why they don't just outright add "static" versions of all the > data containers. As in, you specify the maximum size of the container > as a template parameter, and the side of the object itself will be > that much. Just like std::array, but containing all the member > functions of the data container. So for example you could have a > "static" std::string of a maximum length that you specify as a > template parameter, which does no dynamic memory allocations. > > (Of course nothing stops me from implementing such a thing > myself, but...) One inconvenience would be that each different length would be a separate type. That would make many string ops really awkward, if it changes type halfway through. You can already reserve() the space for a std::string, at the cost of a single dynamic allocation.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2022-08-02 05:43 +0000 |
| Message-ID | <tcadi4$3g0$1@gioia.aioe.org> |
| In reply to | #85777 |
Bo Persson <bo@bo-persson.se> wrote: > You can already reserve() the space for a std::string, at the cost of a > single dynamic allocation. The main idea would be that if it's eg. a member variable of a class, its maximum length is something like 20 characters, and the class is instantiated millions of times, it will save a lot of memory, cause less memory fragmentation and be more efficient. But yeah, in such situations it would probably not be a huge amount of trouble to create a custom implementation.
[toc] | [prev] | [next] | [standalone]
| From | Manfred <noname@add.invalid> |
|---|---|
| Date | 2022-08-03 01:13 +0200 |
| Message-ID | <tccb2c$62b$1@gioia.aioe.org> |
| In reply to | #85764 |
On 7/31/2022 4:39 PM, Juha Nieminen wrote: > Manfred <noname@add.invalid> wrote: >> That, and below, is generally true for a generic dynamic container. >> However, std::string is not a generic container. It is a very specific >> and very optimized container of chars. >> So, if you use a decent implementation of std the performance of >> std::string and char* is usually pretty close. > > I don't know if you are referring to it, but many people seem to think that > "short string optimization" makes std::string pretty much as efficient > as an array of char (for strings that are short enough). > > They fail to take into consideration that short string optimization > requires conditionals in almost all member functons that access the > string data. Conditionals are not free. It is more general than plain SSO. My point is that std::string is different from a generic container when it comes to performance, because implementors know two things: 1) std::string is made to contain char's (and, more rarely, wchar_t's) - it's not made to contain other stuff (well, maybe technically it can, but then it can happily perform poorly) 2) users are more eager to get performance out of std::string than out of say std::vector. So, SSO is /one/ of the technologies that implementors can use to improve performance, but they (expecially the good ones) are known to be quite creative when it comes to performance. SSO is a good example to explain my point, though: while it makes sense for strings, I guess very few, if anyone, would be willing to pay for some sort of SVO (Short Vector Optimization)
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-08-03 09:34 +0300 |
| Message-ID | <tcd4uk$22gbl$1@dont-email.me> |
| In reply to | #85784 |
03.08.2022 02:13 Manfred kirjutas: > On 7/31/2022 4:39 PM, Juha Nieminen wrote: >> Manfred <noname@add.invalid> wrote: >>> That, and below, is generally true for a generic dynamic container. >>> However, std::string is not a generic container. It is a very specific >>> and very optimized container of chars. >>> So, if you use a decent implementation of std the performance of >>> std::string and char* is usually pretty close. >> >> I don't know if you are referring to it, but many people seem to think >> that >> "short string optimization" makes std::string pretty much as efficient >> as an array of char (for strings that are short enough). >> >> They fail to take into consideration that short string optimization >> requires conditionals in almost all member functons that access the >> string data. Conditionals are not free. > > It is more general than plain SSO. > > My point is that std::string is different from a generic container when > it comes to performance, because implementors know two things: > 1) std::string is made to contain char's (and, more rarely, wchar_t's) - > it's not made to contain other stuff (well, maybe technically it can, > but then it can happily perform poorly) > 2) users are more eager to get performance out of std::string than out > of say std::vector. > > So, SSO is /one/ of the technologies that implementors can use to > improve performance, but they (expecially the good ones) are known to be > quite creative when it comes to performance. SSO is a good example to > explain my point, though: while it makes sense for strings, I guess very > few, if anyone, would be willing to pay for some sort of SVO (Short > Vector Optimization) Right, SSO makes more sense for strings, mostly because the character size is small, compared to a typical vector element size (as long as one is using ASCII or UTF-8 and not something like std::u32string). From what I see in common C++ implementations, sizeof(std::string) is 32, from which 16 bytes gets reused as the SSO buffer, meaning that SSO applies to strings up to length 15. That's something. Sizeof(std::vector) seems to be something like 24, from which 8 bytes could be used as "SVO" buffer, meaning that this optimization would apply for vectors of max 1-2 elements only. Not so useful.
[toc] | [prev] | [next] | [standalone]
| From | Muttley@dastardlyhq.com |
|---|---|
| Date | 2022-07-29 09:22 +0000 |
| Message-ID | <tc08s1$1u7$1@gioia.aioe.org> |
| In reply to | #85723 |
On Thu, 28 Jul 2022 20:03:41 +0300
Paavo Helde <eesnimi@osa.pri.ee> wrote:
>28.07.2022 18:11 Muttley@dastardlyhq.com kirjutas:
>> On Thu, 28 Jul 2022 15:04:12 +0200
>> Bo Persson <bo@bo-persson.se> wrote:
>>
>>> std::string full_name = first_name + last_name;
>>>
>>> I wouldn't expect a floating point addition.
>>
>> Unfortunately even in C++ 2020 that won't compile unless first_name is a
>> string and won't work if its char* due to precendence rules.
>
>Why on earth are you using char* in a C++ program? The only thing they
Is that supposed to be a serious question?
>bring is problems and slowdowns (because string length needs to be
>re-calculated all the time).
Depends on what you're doing with it. You do realise char* don't always
just point to a text string?
>#include <string>
>using namespace std::string_literals;
>
>int main() {
> auto first_name = "Bernhard"s, last_name = "Shaw"s;
> std::string full_name = first_name + last_name;
Hmm, I wonder whats more efficient. Counting the length of 2 short strings or
creating 2 objects...
[toc] | [prev] | [next] | [standalone]
Page 4 of 5 — ← Prev page 1 2 3 [4] 5 Next page →
Back to top | Article view | comp.lang.c++
csiph-web