Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c++ > #87510 > unrolled thread

An argument *against* (the liberal use of) references

Started byJuha Nieminen <nospam@thanks.invalid>
First post2022-11-22 09:52 +0000
Last post2022-11-23 20:08 +0100
Articles 20 on this page of 58 — 17 participants

Back to article view | Back to comp.lang.c++


Contents

  An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-11-22 09:52 +0000
    Re: An argument *against* (the liberal use of) references "Fred. Zwarts" <F.Zwarts@KVI.nl> - 2022-11-22 11:23 +0100
    Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-11-22 13:27 +0200
      Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-11-23 06:52 +0000
        Re: An argument *against* (the liberal use of) references "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-11-23 13:51 +0100
          Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-11-23 13:46 +0000
            Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-11-30 13:59 +0200
              Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-01 06:45 +0000
                Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 09:19 +0200
                  Re: An argument *against* (the liberal use of) references Stuart Redmann <DerTopper@web.de> - 2022-12-01 14:06 +0100
                    Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 16:45 +0200
                      Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-01 16:02 +0000
                        Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 20:21 +0200
                          Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-01 19:09 +0000
                            Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-02 05:53 -0800
                              Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-02 14:36 +0000
                                Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 12:29 -0800
                                  Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-02 20:55 +0000
                                    Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 12:58 -0800
                                      Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-02 21:18 +0000
                                        Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 14:16 -0800
                                    Re: An argument *against* (the liberal use of) references Branimir Maksimovic <branimir.maksimovic@icloud.com> - 2022-12-02 21:09 +0000
                                      Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 13:22 -0800
                                        Re: An argument *against* (the liberal use of) references Branimir Maksimovic <branimir.maksimovic@icloud.com> - 2022-12-02 21:36 +0000
                                          Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 14:13 -0800
                              Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-02 18:33 +0200
                          Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-01 11:59 -0800
                            Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 23:53 +0200
                              Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-02 04:11 -0800
                    Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-01 12:21 -0800
                  Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-06 11:36 +0000
                    Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-06 12:26 -0800
                      Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-06 20:39 +0000
                        Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-06 23:53 +0200
                          Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-06 13:59 -0800
                            Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-06 22:04 +0000
                              Re: An argument *against* (the liberal use of) references Öö Tiib <ootiib@hot.ee> - 2022-12-06 21:38 -0800
                                Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-07 15:22 +0000
                                  Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-07 18:37 +0200
                          Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-06 14:01 -0800
                      Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-07 09:05 +0000
                        Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-07 12:48 -0800
                          Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-07 12:50 -0800
                          Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-08 07:52 +0000
                            Re: An argument *against* (the liberal use of) references Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-12-08 05:50 -0800
                              Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-08 13:30 -0800
                                Re: An argument *against* (the liberal use of) references Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-12-09 12:50 -0800
                                  Re: An argument *against* (the liberal use of) references David Brown <david.brown@hesbynett.no> - 2022-12-11 12:18 +0100
                          Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-08 13:38 +0200
                            Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-08 16:45 -0800
                    Re: An argument *against* (the liberal use of) references Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-12-09 04:35 -0800
    Re: An argument *against* (the liberal use of) references Bonita Montero <Bonita.Montero@gmail.com> - 2022-11-22 15:36 +0100
    Re: An argument *against* (the liberal use of) references Chris Vine <chris@cvine--nospam--.freeserve.co.uk> - 2022-11-22 15:21 +0000
    Re: An argument *against* (the liberal use of) references Richard Damon <Richard@Damon-Family.org> - 2022-11-22 10:48 -0500
      Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-22 13:18 -0800
    Re: An argument *against* (the liberal use of) references El Jo <giorgio.zoppi@gmail.com> - 2022-11-22 13:41 -0800
      Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-22 13:51 -0800
    Re: An argument *against* (the liberal use of) references Bonita Montero <Bonita.Montero@gmail.com> - 2022-11-23 20:08 +0100

Page 2 of 3 — ← Prev page 1 [2] 3  Next page →


#87703

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-12-02 14:16 -0800
Message-ID<tmdtg0$358pk$6@dont-email.me>
In reply to#87699
On 12/2/2022 1:18 PM, Scott Lurndal wrote:
> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>> On 12/2/2022 12:55 PM, Scott Lurndal wrote:
>>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>>> On 12/2/2022 6:36 AM, Scott Lurndal wrote:
>>>>> Michael S <already5chosen@yahoo.com> writes:
>>>>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:
>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas:
>>>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas:
>>>>>>>>>
>>>>>
>>>>>>>
>>>>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the
>>>>>>> systemwide lock is the only possibility [*].
>>>>>>>
>>>>>>
>>>>>> Huh?
>>>>>> There is absolute no relationship between what prefix is used (instruction
>>>>>> encoding issue) and implementation.
>>>>>
>>>>> The only way to specify an atomic access in those chips was
>>>>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic
>>>>> add, et alia).
>>>>>
>>>>
>>>> Iirc, XCHG has an implicit LOCK prefix?
>>>
>>> Not sure about XCHG, but CMPXCHG requires the prefix
>>> when used in a multiprocessor system; on a uniprocessor
>>> it will be atomic because interrupts are always taken
>>> between instructions (unlike, the VAX, for instance,
>>> where certain instructions (MOVC3/5) can be interrupted
>>> and restarted).    I suspect that XCHG has similar
>>> characteristics.
>>>
>>>
>>
>> Iirc, XCHG is the _only_ atomic RMW instruction that has an _implicit_
>> LOCK prefix. CMPXCHG _needs_ the programmer to put in a LOCK prefix.
> 
> Yes, that is the case
> 
> XCHG:
>    "If a memory operand is referenced, the processor's locking protocol is automatically
>     implemented for the duration of the exchange operation, regardless of the presence
>     or absence of the LOCK prefix or of the value of the IOPL. (See the LOCK prefix
>     description in this chapter for more information on the locking protocol.)
> 

Indeed. When I remembered this I kind of doubted myself for a moment. 
Then, I said well, I know its true, and if I am wrong then I must of 
fried my brain a bit.

[toc] | [prev] | [next] | [standalone]


#87698

FromBranimir Maksimovic <branimir.maksimovic@icloud.com>
Date2022-12-02 21:09 +0000
Message-ID<jQtiL.1308$_Y84.549@fx46.iad>
In reply to#87696
On 2022-12-02, Scott Lurndal <scott@slp53.sl.home> wrote:
> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>On 12/2/2022 6:36 AM, Scott Lurndal wrote:
>>> Michael S <already5chosen@yahoo.com> writes:
>>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:
>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas:
>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas:
>>>>>>>
>>> 
>>>>>
>>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the
>>>>> systemwide lock is the only possibility [*].
>>>>>
>>>>
>>>> Huh?
>>>> There is absolute no relationship between what prefix is used (instruction
>>>> encoding issue) and implementation.
>>> 
>>> The only way to specify an atomic access in those chips was
>>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic
>>> add, et alia).
>>> 
>>
>>Iirc, XCHG has an implicit LOCK prefix?
>
> Not sure about XCHG, but CMPXCHG requires the prefix
> when used in a multiprocessor system; on a uniprocessor
> it will be atomic because interrupts are always taken
> between instructions (unlike, the VAX, for instance,
> where certain instructions (MOVC3/5) can be interrupted
> and restarted).    I suspect that XCHG has similar
> characteristics.
>
>
  "lock" prefix causes the processor's bus-lock signal to be asserted during
execution of the accompanying instruction. In a multiprocessor environment,
the bus-lock signal insures that the processor has exclusive use of any shared
memory while the signal is asserted. The "lock" prefix can be prepended only
to the following instructions and only to those forms of the instructions
where the destination operand is a memory operand: "add", "adc", "and", "btc",
"btr", "bts", "cmpxchg", "cmpxchg8b", "dec", "inc", "neg", "not", "or", "sbb",
"sub", "xor", "xadd" and "xchg". If the "lock" prefix is used with one of
these instructions and the source operand is a memory operand, an undefined
opcode exception may be generated. An undefined opcode exception will also be
generated if the "lock" prefix is used with any instruction not in the above
list. The "xchg" instruction always asserts the bus-lock signal regardless of
the presence or absence of the "lock" prefix.


-- 

7-77-777
Evil Sinner!
with software, you repeat same experiment, expecting different results...

[toc] | [prev] | [next] | [standalone]


#87700

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-12-02 13:22 -0800
Message-ID<tmdqa9$358pk$1@dont-email.me>
In reply to#87698
On 12/2/2022 1:09 PM, Branimir Maksimovic wrote:
> On 2022-12-02, Scott Lurndal <scott@slp53.sl.home> wrote:
>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>> On 12/2/2022 6:36 AM, Scott Lurndal wrote:
>>>> Michael S <already5chosen@yahoo.com> writes:
>>>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:
>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas:
>>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas:
>>>>>>>>
>>>>
>>>>>>
>>>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the
>>>>>> systemwide lock is the only possibility [*].
>>>>>>
>>>>>
>>>>> Huh?
>>>>> There is absolute no relationship between what prefix is used (instruction
>>>>> encoding issue) and implementation.
>>>>
>>>> The only way to specify an atomic access in those chips was
>>>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic
>>>> add, et alia).
>>>>
>>>
>>> Iirc, XCHG has an implicit LOCK prefix?
>>
>> Not sure about XCHG, but CMPXCHG requires the prefix
>> when used in a multiprocessor system; on a uniprocessor
>> it will be atomic because interrupts are always taken
>> between instructions (unlike, the VAX, for instance,
>> where certain instructions (MOVC3/5) can be interrupted
>> and restarted).    I suspect that XCHG has similar
>> characteristics.
>>
>>
>    "lock" prefix causes the processor's bus-lock signal to be asserted during
> execution of the accompanying instruction. In a multiprocessor environment,
> the bus-lock signal insures that the processor has exclusive use of any shared
> memory while the signal is asserted. The "lock" prefix can be prepended only
> to the following instructions and only to those forms of the instructions
> where the destination operand is a memory operand: "add", "adc", "and", "btc",
> "btr", "bts", "cmpxchg", "cmpxchg8b", "dec", "inc", "neg", "not", "or", "sbb",
> "sub", "xor", "xadd" and "xchg". If the "lock" prefix is used with one of
> these instructions and the source operand is a memory operand, an undefined
> opcode exception may be generated. An undefined opcode exception will also be
> generated if the "lock" prefix is used with any instruction not in the above
> list. 


> The "xchg" instruction always asserts the bus-lock signal regardless of
> the presence or absence of the "lock" prefix.
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

BINGO! I had a strong feeling I was right.

https://youtu.be/TnZrWWUFl8I

[toc] | [prev] | [next] | [standalone]


#87701

FromBranimir Maksimovic <branimir.maksimovic@icloud.com>
Date2022-12-02 21:36 +0000
Message-ID<6duiL.41721$f9D6.225@fx09.iad>
In reply to#87700
On 2022-12-02, Chris M. Thomasson <chris.m.thomasson.1@gmail.com> wrote:
> On 12/2/2022 1:09 PM, Branimir Maksimovic wrote:
>> On 2022-12-02, Scott Lurndal <scott@slp53.sl.home> wrote:
>>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>>> On 12/2/2022 6:36 AM, Scott Lurndal wrote:
>>>>> Michael S <already5chosen@yahoo.com> writes:
>>>>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:
>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas:
>>>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas:
>>>>>>>>>
>>>>>
>>>>>>>
>>>>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the
>>>>>>> systemwide lock is the only possibility [*].
>>>>>>>
>>>>>>
>>>>>> Huh?
>>>>>> There is absolute no relationship between what prefix is used (instruction
>>>>>> encoding issue) and implementation.
>>>>>
>>>>> The only way to specify an atomic access in those chips was
>>>>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic
>>>>> add, et alia).
>>>>>
>>>>
>>>> Iirc, XCHG has an implicit LOCK prefix?
>>>
>>> Not sure about XCHG, but CMPXCHG requires the prefix
>>> when used in a multiprocessor system; on a uniprocessor
>>> it will be atomic because interrupts are always taken
>>> between instructions (unlike, the VAX, for instance,
>>> where certain instructions (MOVC3/5) can be interrupted
>>> and restarted).    I suspect that XCHG has similar
>>> characteristics.
>>>
>>>
>>    "lock" prefix causes the processor's bus-lock signal to be asserted during
>> execution of the accompanying instruction. In a multiprocessor environment,
>> the bus-lock signal insures that the processor has exclusive use of any shared
>> memory while the signal is asserted. The "lock" prefix can be prepended only
>> to the following instructions and only to those forms of the instructions
>> where the destination operand is a memory operand: "add", "adc", "and", "btc",
>> "btr", "bts", "cmpxchg", "cmpxchg8b", "dec", "inc", "neg", "not", "or", "sbb",
>> "sub", "xor", "xadd" and "xchg". If the "lock" prefix is used with one of
>> these instructions and the source operand is a memory operand, an undefined
>> opcode exception may be generated. An undefined opcode exception will also be
>> generated if the "lock" prefix is used with any instruction not in the above
>> list. 
>
>
>> The "xchg" instruction always asserts the bus-lock signal regardless of
>> the presence or absence of the "lock" prefix.
> ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
>
> BINGO! I had a strong feeling I was right.
>
> https://youtu.be/TnZrWWUFl8I
Nice music :p

-- 

7-77-777
Evil Sinner!
with software, you repeat same experiment, expecting different results...

[toc] | [prev] | [next] | [standalone]


#87702

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-12-02 14:13 -0800
Message-ID<tmdtao$358pk$5@dont-email.me>
In reply to#87701
On 12/2/2022 1:36 PM, Branimir Maksimovic wrote:
> On 2022-12-02, Chris M. Thomasson <chris.m.thomasson.1@gmail.com> wrote:
>> On 12/2/2022 1:09 PM, Branimir Maksimovic wrote:
>>> On 2022-12-02, Scott Lurndal <scott@slp53.sl.home> wrote:
>>>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>>>> On 12/2/2022 6:36 AM, Scott Lurndal wrote:
>>>>>> Michael S <already5chosen@yahoo.com> writes:
>>>>>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:
>>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas:
>>>>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas:
>>>>>>>>>>
>>>>>>
>>>>>>>>
>>>>>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the
>>>>>>>> systemwide lock is the only possibility [*].
>>>>>>>>
>>>>>>>
>>>>>>> Huh?
>>>>>>> There is absolute no relationship between what prefix is used (instruction
>>>>>>> encoding issue) and implementation.
>>>>>>
>>>>>> The only way to specify an atomic access in those chips was
>>>>>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic
>>>>>> add, et alia).
>>>>>>
>>>>>
>>>>> Iirc, XCHG has an implicit LOCK prefix?
>>>>
>>>> Not sure about XCHG, but CMPXCHG requires the prefix
>>>> when used in a multiprocessor system; on a uniprocessor
>>>> it will be atomic because interrupts are always taken
>>>> between instructions (unlike, the VAX, for instance,
>>>> where certain instructions (MOVC3/5) can be interrupted
>>>> and restarted).    I suspect that XCHG has similar
>>>> characteristics.
>>>>
>>>>
>>>     "lock" prefix causes the processor's bus-lock signal to be asserted during
>>> execution of the accompanying instruction. In a multiprocessor environment,
>>> the bus-lock signal insures that the processor has exclusive use of any shared
>>> memory while the signal is asserted. The "lock" prefix can be prepended only
>>> to the following instructions and only to those forms of the instructions
>>> where the destination operand is a memory operand: "add", "adc", "and", "btc",
>>> "btr", "bts", "cmpxchg", "cmpxchg8b", "dec", "inc", "neg", "not", "or", "sbb",
>>> "sub", "xor", "xadd" and "xchg". If the "lock" prefix is used with one of
>>> these instructions and the source operand is a memory operand, an undefined
>>> opcode exception may be generated. An undefined opcode exception will also be
>>> generated if the "lock" prefix is used with any instruction not in the above
>>> list.
>>
>>
>>> The "xchg" instruction always asserts the bus-lock signal regardless of
>>> the presence or absence of the "lock" prefix.
>> ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
>>
>> BINGO! I had a strong feeling I was right.
>>
>> https://youtu.be/TnZrWWUFl8I
> Nice music :p
> 


Thanks again, Branimir, for the quote of the docs.

https://youtu.be/UZ2-FfXZlAU
(super mario music in a live big band format? Nice... :^)

My hat is off to you.

[toc] | [prev] | [next] | [standalone]


#87688

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-12-02 18:33 +0200
Message-ID<tmd9d8$341b0$1@dont-email.me>
In reply to#87684
02.12.2022 15:53 Michael S kirjutas:
> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:

> BTW, it means that your claim in post above "no overhead at all unless ..."
> is incorrect in the absolute sense. There is an overhead even without "unless".
> But the overhead in uncontended case is small - order of dozen or two of CPU
> clocks. So, undetectable in Paavo's case of only 1000 updates per second.
> For 1M updates per second impact would me detectable with precise time
>   measurements and for 100M per second there would be big slowdown.

Just a minor note, earlier I posted numbers for only a single 
smartpointer, whereas in the full program there are probably tens or 
hundreds of thousands of them. I just measured it and the total rate of 
refcount changes is something like 5M per second.

With this 5M/s rate I cannot see any slowdown (of using std::atomic<int> 
instead of int) in my measurements (with no contention), variations 
caused by other uncontrollable factors seem to be much larger. Maybe I 
should rerun this on a Linux box where the things are more stable.

[toc] | [prev] | [next] | [standalone]


#87675

FromMichael S <already5chosen@yahoo.com>
Date2022-12-01 11:59 -0800
Message-ID<b54ae238-40c3-4003-a49a-8e44dc7d9eb7n@googlegroups.com>
In reply to#87671
On Thursday, December 1, 2022 at 8:22:12 PM UTC+2, Paavo Helde wrote:
> 01.12.2022 18:02 Scott Lurndal kirjutas: 
> > Paavo Helde <ees...@osa.pri.ee> writes: 
> >> 01.12.2022 15:06 Stuart Redmann kirjutas: 
> > 
> > <snip> 
> > 
> >>> Another thought: if thread-safety is too costly, you could use two smart 
> >>> pointer classes: thread-safe pointers and forwarding non-thread-safe smart 
> >>> pointers. The forwarding smart pointers have their own thread-UNsafe 
> >>> refcount and the thread-safe smart pointer as member. 
> >> 
> >> I tried to measure the impact of std::atomic<int> refcounters and in 
> >> first tests it seems the overhead on x86_64 is zero (with no 
> >> contention). 
> > 
> > Which follows naturally from the fact that the core doing 
> > the atomic access has exclusive access to the cache line 
> > containing the refcounter. No overhead at all, unless
> Thanks for the clarifications!
> > the ref counter isn't aligned and crosses a cache-line 
> > boundary (or the access is to an uncached memory 
> > range or caching is disabled), in which case the processor 
> > will take a system-wide 
> > lock to perform the operation, which is catastrophic 
> > on systems with large processor counts.
> It is clear that having a misaligned cross-border atomic would be very 
> bad. But what about normal uncached memory ranges, wouldn't these be 
> just loaded into the cache, without disturbing other processors, and 
> without any "catastrophic" consequences?

"Unchached memory range" is a misnormer.
A proper name is uncacheable range (region).
Unfortunately "uncached" in the meaning of "uncacheable" is used 
quite often. Even Intel's official manuals suffer from such
inconsistent vocabulary.

[toc] | [prev] | [next] | [standalone]


#87678

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-12-01 23:53 +0200
Message-ID<tmb7p9$2sgre$1@dont-email.me>
In reply to#87675
01.12.2022 21:59 Michael S kirjutas:

> "Unchached memory range" is a misnormer.
> A proper name is uncacheable range (region).
> Unfortunately "uncached" in the meaning of "uncacheable" is used
> quite often. Even Intel's official manuals suffer from such
> inconsistent vocabulary.

Thanks, I had to look up what is "uncacheable memory". I guess the 80186 
processor where I learned my basics did not have such a thing.

[toc] | [prev] | [next] | [standalone]


#87683

FromMichael S <already5chosen@yahoo.com>
Date2022-12-02 04:11 -0800
Message-ID<a1c3687f-6b3e-4b60-a0dc-6264b3306026n@googlegroups.com>
In reply to#87678
On Thursday, December 1, 2022 at 11:54:02 PM UTC+2, Paavo Helde wrote:
> 01.12.2022 21:59 Michael S kirjutas: 
> 
> > "Unchached memory range" is a misnormer. 
> > A proper name is uncacheable range (region). 
> > Unfortunately "uncached" in the meaning of "uncacheable" is used 
> > quite often. Even Intel's official manuals suffer from such 
> > inconsistent vocabulary.
> Thanks, I had to look up what is "uncacheable memory". I guess the 80186 
> processor where I learned my basics did not have such a thing.

80186 was "embedded" microprocessor similar at core to 8086.
It was typically used with no cache so didn't need a concept of uncacheable 
regions.
80286 and especially i386 was used with (external) cache quite often, but
according to my understanding their caches were what we call today 
"memory-side caches" associated with main memory. 
From system (both CPU and other bus masters) perspective such caches 
are totally transparent (except for entering/leaving deep sleep states, but 
back then they didn't do it) so  there still was no need for uncacheable regions.

System-side caches and associated problem first appear in x86 world in i486.
Still, the in original i486 the system cache had strict write-through policy, so the
problems were minor.
Then came Pentium with 8 KB of write-back Data cache and a little later came
new models of i486 with even bigger write-back cache and problems became
quite real, especially because approximately in the same time PCI took over
I/O bus  role and suddenly multiple bus masters that were high-end curiosity
before then, became common in consumer PC hardware. 
But even then x86 architecture lacked adequate answer to a new challenge.

The first reasonable answer (MTRR registers) came only in PPro but it still
had problems of scalability - too few regions. 
Later (P-III) they invented PAT which from theoretical point of view is inferior
to MTRRs because in PAT scheme cachability is an attribute of virtual address
rather than of physical address. But PAT is ultimately scalable and is one
solution that, as long as OS does a proper plumbing, is one solution that can
rule over all aspects of cachabilty.  So it won.

[toc] | [prev] | [next] | [standalone]


#87676

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-12-01 12:21 -0800
Message-ID<tmb2co$2s3ov$3@dont-email.me>
In reply to#87666
On 12/1/2022 5:06 AM, Stuart Redmann wrote:
> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>> 01.12.2022 08:45 Juha Nieminen kirjutas:
>>> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>>>> Actually once located, the bug was simple. The object which was copied
>>>> was a single-threaded refcounted smartpointer, and by copying it the
>>>> refcounter got incremented (and later decremented). Alas, this was
>>>> accidentally done from parallel threads at the same time, without any
>>>> synchronization, so eventually the refcounter got messed up.
>>>
>>> I think that if the reference count is declared atomic, it can be safely
>>> directly incremented. When decrementing you would need to use the
>>> fetch_sub() function to see if the object needs to be destroyed.
>>>
>>> While modifying an atomic might not be equally fast as a non-atomic,
>>> it shouldn't be all that much slower either, at least if the target
>>> architecture supports atomic operations.
>>
>> I have pondered this myself. Maybe I should measure the actual slowdown
>> after temporarily making the refcounters atomic. But this seems overkill
>> because these smartpointers would still point to single-threaded objects
>> which are meant to be primarily used in single-thread regime, so in most
>> cases making the smartpointers atomic does not buy anything.
>>
>> When tracking down this bug, I monitored all refcounter changes for a
>> particular single smartpointer during the program run (ca 10 min). There
>> were 591848 increments and decrements, from which 1526 came from the
>> problematic (parallelized) part. It looks like a pessimization to slow
>> down 99.75% of accesses when only 0.25% would actually benefit from this.
>>
> 
> 600k changes in reference counts look suspicious to me. When you pass a
> ref-counted object to a worker thread, there should only be a single change
> in the refcount. This would be because inside the worker thread some object
> takes (shared) ownership of shared object. If the shared object needs to be
> passed to sub-routines, you should pass them as references or plain
> pointers (if the subroutine must be able to cope with non-existing
> objects). It should be rare occurrence that another object in the worker
> thread needs to take ownership of the shared object.
> 
> Another thought: if thread-safety is too costly, you could use two smart
> pointer classes: thread-safe pointers and forwarding non-thread-safe smart
> pointers. The forwarding smart pointers have their own thread-UNsafe
> refcount and the thread-safe smart pointer as member.

Check this out:

https://github.com/jseigh/atomic-ptr-plus/blob/master/atomic-ptr/atomic_ptr.h

It is a truly atomic reference counted pointer. A thread can take a 
reference without owning a prior reference.

Here is a patent:

https://patents.justia.com/patent/5295262

[toc] | [prev] | [next] | [standalone]


#87706

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-12-06 11:36 +0000
Message-ID<tmn9fl$1kpl$1@gioia.aioe.org>
In reply to#87665
Paavo Helde <eesnimi@osa.pri.ee> wrote:
> When tracking down this bug, I monitored all refcounter changes for a 
> particular single smartpointer during the program run (ca 10 min). There 
> were 591848 increments and decrements, from which 1526 came from the 
> problematic (parallelized) part. It looks like a pessimization to slow 
> down 99.75% of accesses when only 0.25% would actually benefit from this.

There's place for micro-optimization and there's place to do the
Right Thing (TM) instead.

In the vast, vast majority of situations micro-optimization will have
little to no effect on the program. It's only when you have number
crunching code that does something billions of times per second that
micro-optimization may start having some discernible effect. Those
situations tend to be very rare and far-in-between. And when you do
have such situations you can make faster versions of things for that
alone.

Micro-optimization is extra useless if it's surrounded by, and thus
swamped by code that's a lot slower than it. Even when micro-optimizing
you should start with the worst offenders, not the smallest things.

(By "micro-optimization" I'm referring to things that do not change
the computational complexity of something and only makes that something
some clock cycles faster.)

So unless your smart pointer is being copied and assigned around
millions of times per second in tight number-crunching loops, you can
safely ignore any lost clock cycles by making the reference counter
thread-safe.

(If you actually need to copy and assign smart pointers around
millions of times per second in a tight number-crunching inner
loop, perhaps create a specialized version of the pointer for
that particular purpose...)

[toc] | [prev] | [next] | [standalone]


#87707

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-12-06 12:26 -0800
Message-ID<tmo8ho$ahd1$1@dont-email.me>
In reply to#87706
On 12/6/2022 3:36 AM, Juha Nieminen wrote:
> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>> When tracking down this bug, I monitored all refcounter changes for a
>> particular single smartpointer during the program run (ca 10 min). There
>> were 591848 increments and decrements, from which 1526 came from the
>> problematic (parallelized) part. It looks like a pessimization to slow
>> down 99.75% of accesses when only 0.25% would actually benefit from this.
[...]
> So unless your smart pointer is being copied and assigned around
> millions of times per second in tight number-crunching loops, you can
> safely ignore any lost clock cycles by making the reference counter
> thread-safe.
[...]
A thread-safe reference counted pointer can heavily damage performance 
in certain usage scenarios. Blasting the system with memory barriers and 
atomic RMW ops all over the place. Now, there is a work around called 
proxy reference counting. Are you familiar with it?

[toc] | [prev] | [next] | [standalone]


#87708

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-12-06 20:39 +0000
Message-ID<sMNjL.66$ZhSc.56@fx38.iad>
In reply to#87707
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>On 12/6/2022 3:36 AM, Juha Nieminen wrote:
>> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>>> When tracking down this bug, I monitored all refcounter changes for a
>>> particular single smartpointer during the program run (ca 10 min). There
>>> were 591848 increments and decrements, from which 1526 came from the
>>> problematic (parallelized) part. It looks like a pessimization to slow
>>> down 99.75% of accesses when only 0.25% would actually benefit from this.
>[...]
>> So unless your smart pointer is being copied and assigned around
>> millions of times per second in tight number-crunching loops, you can
>> safely ignore any lost clock cycles by making the reference counter
>> thread-safe.
>[...]
>A thread-safe reference counted pointer can heavily damage performance 
>in certain usage scenarios. Blasting the system with memory barriers and 
>atomic RMW ops all over the place. Now, there is a work around called 
>proxy reference counting. Are you familiar with it?

The fact that smart pointers do allocation/deallocation has made
them useless for high-performance threaded code, IMO.

[toc] | [prev] | [next] | [standalone]


#87709

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-12-06 23:53 +0200
Message-ID<tmodko$avn3$1@dont-email.me>
In reply to#87708
06.12.2022 22:39 Scott Lurndal kirjutas:

> 
> The fact that smart pointers do allocation/deallocation has made
> them useless for high-performance threaded code, IMO.

You have got it backwards. Smartpointers are taken into use for coping 
with the fact that objects need to by dynamically allocated and 
deallocated, by the program logic.

And this allocation/deallocation would happen relatively rarely. If the 
object lifetimes were short, then typically they could be controlled 
much better, and there would be no need for refcounted smartpointers, or 
maybe even no need for dynamic allocation of objects in the first place.


[toc] | [prev] | [next] | [standalone]


#87710

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-12-06 13:59 -0800
Message-ID<tmodvl$auhg$2@dont-email.me>
In reply to#87709
On 12/6/2022 1:53 PM, Paavo Helde wrote:
> 06.12.2022 22:39 Scott Lurndal kirjutas:
> 
>>
>> The fact that smart pointers do allocation/deallocation has made
>> them useless for high-performance threaded code, IMO.
> 
> You have got it backwards. Smartpointers are taken into use for coping 
> with the fact that objects need to by dynamically allocated and 
> deallocated, by the program logic.
> 
> And this allocation/deallocation would happen relatively rarely. If the 
> object lifetimes were short, then typically they could be controlled 
> much better, and there would be no need for refcounted smartpointers, or 
> maybe even no need for dynamic allocation of objects in the first place.
> 
> 
> 

Imvvho, a smart pointer should not need to allocate anything under the 
covers. Also, are you familiar with proxy reference counting? It has the 
ability to amortize a single reference over n objects.

[toc] | [prev] | [next] | [standalone]


#87712

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-12-06 22:04 +0000
Message-ID<A%OjL.2208$iS99.2139@fx16.iad>
In reply to#87710
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>On 12/6/2022 1:53 PM, Paavo Helde wrote:
>> 06.12.2022 22:39 Scott Lurndal kirjutas:
>> 
>>>
>>> The fact that smart pointers do allocation/deallocation has made
>>> them useless for high-performance threaded code, IMO.
>> 
>> You have got it backwards. Smartpointers are taken into use for coping 
>> with the fact that objects need to by dynamically allocated and 
>> deallocated, by the program logic.
>> 
>> And this allocation/deallocation would happen relatively rarely. 

Assumption not in evidence.  I've personnally had to rip smart pointers
out of code because the allocation/deallocation happened very
frequently.   One if the applications was simulating a processor pipeline,
another was handling network packets both were written by well-educated
people familiar with C++.

Granted, one can specify a more efficient allocator, but

  1) most C++ programmers don't bother or don't know how
  2) Even then there is unnecessary overhead unless the allocator is pool based.

KISS applies, always.

[toc] | [prev] | [next] | [standalone]


#87713

FromÖö Tiib <ootiib@hot.ee>
Date2022-12-06 21:38 -0800
Message-ID<4f4fff56-27ce-4181-910d-fb2fa0307b67n@googlegroups.com>
In reply to#87712
On Wednesday, 7 December 2022 at 00:04:34 UTC+2, Scott Lurndal wrote:
> "Chris M. Thomasson" <chris.m.t...@gmail.com> writes:
> >On 12/6/2022 1:53 PM, Paavo Helde wrote: 
> >> 06.12.2022 22:39 Scott Lurndal kirjutas: 
> >> 
> >>> 
> >>> The fact that smart pointers do allocation/deallocation has made 
> >>> them useless for high-performance threaded code, IMO. 
> >> 
> >> You have got it backwards. Smartpointers are taken into use for coping 
> >> with the fact that objects need to by dynamically allocated and 
> >> deallocated, by the program logic. 
> >> 
> >> And this allocation/deallocation would happen relatively rarely.
> Assumption not in evidence. I've personnally had to rip smart pointers 
> out of code because the allocation/deallocation happened very 
> frequently. One if the applications was simulating a processor pipeline, 
> another was handling network packets both were written by well-educated 
> people familiar with C++. 
> 
> Granted, one can specify a more efficient allocator, but 
> 
> 1) most C++ programmers don't bother or don't know how 
> 2) Even then there is unnecessary overhead unless the allocator is pool based. 
> 
> KISS applies, always.

No one argues with that. Just that keeping it simple is far from simple. 
For example it is tricky to keep dynamic allocations minimal.  That is
not fault of smart pointers.

The std::unique_ptr helps at places where dynamic allocations are 
needed greatly, especially when there can be exceptions. It has next to
no overhead. 

Yes, the std::shared_ptr is loaded. Usage of std::make_shared
helps a bit but thinking about how to make it simpler and to get rid of 
shared ownership or even dynamic allocations is hard and not always
fruitful.
 

[toc] | [prev] | [next] | [standalone]


#87725

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-12-07 15:22 +0000
Message-ID<Ic2kL.1981$jiuc.1610@fx44.iad>
In reply to#87713
=?UTF-8?B?w5bDtiBUaWli?= <ootiib@hot.ee> writes:
>On Wednesday, 7 December 2022 at 00:04:34 UTC+2, Scott Lurndal wrote:
>> "Chris M. Thomasson" <chris.m.t...@gmail.com> writes:
>> >On 12/6/2022 1:53 PM, Paavo Helde wrote: 
>> >> 06.12.2022 22:39 Scott Lurndal kirjutas: 
>> >> 
>> >>> 
>> >>> The fact that smart pointers do allocation/deallocation has made 
>> >>> them useless for high-performance threaded code, IMO. 
>> >> 
>> >> You have got it backwards. Smartpointers are taken into use for coping 
>> >> with the fact that objects need to by dynamically allocated and 
>> >> deallocated, by the program logic. 
>> >> 
>> >> And this allocation/deallocation would happen relatively rarely.
>> Assumption not in evidence. I've personnally had to rip smart pointers 
>> out of code because the allocation/deallocation happened very 
>> frequently. One if the applications was simulating a processor pipeline, 
>> another was handling network packets both were written by well-educated 
>> people familiar with C++. 
>> 
>> Granted, one can specify a more efficient allocator, but 
>> 
>> 1) most C++ programmers don't bother or don't know how 
>> 2) Even then there is unnecessary overhead unless the allocator is pool based. 
>> 
>> KISS applies, always.
>
>No one argues with that. Just that keeping it simple is far from simple. 
>For example it is tricky to keep dynamic allocations minimal.  That is
>not fault of smart pointers.

In my experience, it has been generally sufficent to pre-allocate the
data structures and store them in a table or look-aside list,
as the maximum number is bounded.

For example, an application handling network packets on a processor
with 64 cores, may only need 128 jumbo packet buffers if the packet processing
thread count matches the core count.  These can be preallocated
and then passed as regular pointers throughout the flow.\

(Specialized DPUs have a custom hardware block (network pool allocator)
that allocates hardware buffers to packets on ingress and those buffers
are passed by hardware to the other blocks in the flow, such as
blocks to identify the flow, fragment/defragment a packet,
apply encryption/decryption algorithms, all controlled by a
hardware scheduler block, etc.).


Likewise for a simulation of an internal processor interconnect
such as ring or mesh structure, there are a fixed maximum number
of flits than can be active at any point in time.   Preallocating
them into a lookaside list eliminates allocation and deallocation
overhead on every flit.

When simulating a full SoC, the maximum inflight objects is
likewise bounded and for the most part, can be preallocated.

[toc] | [prev] | [next] | [standalone]


#87726

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-12-07 18:37 +0200
Message-ID<tmqffe$jl7n$1@dont-email.me>
In reply to#87725
07.12.2022 17:22 Scott Lurndal kirjutas:
> =?UTF-8?B?w5bDtiBUaWli?= <ootiib@hot.ee> writes:
>>
>> No one argues with that. Just that keeping it simple is far from simple.
>> For example it is tricky to keep dynamic allocations minimal.  That is
>> not fault of smart pointers.
> 
> In my experience, it has been generally sufficent to pre-allocate the
> data structures and store them in a table or look-aside list,
> as the maximum number is bounded.

My experience is more that the user wants to read in unknown number of 
tiff files containing unknown number of image frames of unknown sizes, 
then start to process them by script-driven flexible algorithms, 
producing an unknown number of intermediate and final results of unknown 
size. And this processing ought to be as fast as possible, as nobody 
wants to wait for hours (although with large data sets and complex 
processing it inevitable gets into hours). And this processing ought 
better to make use of all the cpu cores and should not run out of 
computer memory while doing that.

So it seems preallocating a fixed number of data structures of fixed 
size would not really work in my case.

[toc] | [prev] | [next] | [standalone]


#87711

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-12-06 14:01 -0800
Message-ID<tmoe3m$auhg$3@dont-email.me>
In reply to#87709
On 12/6/2022 1:53 PM, Paavo Helde wrote:
> 06.12.2022 22:39 Scott Lurndal kirjutas:
> 
>>
>> The fact that smart pointers do allocation/deallocation has made
>> them useless for high-performance threaded code, IMO.
> 
> You have got it backwards. Smartpointers are taken into use for coping 
> with the fact that objects need to by dynamically allocated and 
> deallocated, by the program logic.
> 
> And this allocation/deallocation would happen relatively rarely. If the 
> object lifetimes were short, then typically they could be controlled 
> much better, and there would be no need for refcounted smartpointers, or 
> maybe even no need for dynamic allocation of objects in the first place.
> 
> 
> 

Fwiw, RCU is a form of proxy collection. Fwiw, I wrote an experimental 
one using pure C++.

https://pastebin.com/raw/f71480694
(goes to pure text page, no ads and shit like that...)

[toc] | [prev] | [next] | [standalone]


Page 2 of 3 — ← Prev page 1 [2] 3  Next page →

Back to top | Article view | comp.lang.c++


csiph-web