Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #87510 > unrolled thread
| Started by | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| First post | 2022-11-22 09:52 +0000 |
| Last post | 2022-11-23 20:08 +0100 |
| Articles | 20 on this page of 58 — 17 participants |
Back to article view | Back to comp.lang.c++
An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-11-22 09:52 +0000
Re: An argument *against* (the liberal use of) references "Fred. Zwarts" <F.Zwarts@KVI.nl> - 2022-11-22 11:23 +0100
Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-11-22 13:27 +0200
Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-11-23 06:52 +0000
Re: An argument *against* (the liberal use of) references "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-11-23 13:51 +0100
Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-11-23 13:46 +0000
Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-11-30 13:59 +0200
Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-01 06:45 +0000
Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 09:19 +0200
Re: An argument *against* (the liberal use of) references Stuart Redmann <DerTopper@web.de> - 2022-12-01 14:06 +0100
Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 16:45 +0200
Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-01 16:02 +0000
Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 20:21 +0200
Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-01 19:09 +0000
Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-02 05:53 -0800
Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-02 14:36 +0000
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 12:29 -0800
Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-02 20:55 +0000
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 12:58 -0800
Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-02 21:18 +0000
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 14:16 -0800
Re: An argument *against* (the liberal use of) references Branimir Maksimovic <branimir.maksimovic@icloud.com> - 2022-12-02 21:09 +0000
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 13:22 -0800
Re: An argument *against* (the liberal use of) references Branimir Maksimovic <branimir.maksimovic@icloud.com> - 2022-12-02 21:36 +0000
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 14:13 -0800
Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-02 18:33 +0200
Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-01 11:59 -0800
Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 23:53 +0200
Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-02 04:11 -0800
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-01 12:21 -0800
Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-06 11:36 +0000
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-06 12:26 -0800
Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-06 20:39 +0000
Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-06 23:53 +0200
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-06 13:59 -0800
Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-06 22:04 +0000
Re: An argument *against* (the liberal use of) references Öö Tiib <ootiib@hot.ee> - 2022-12-06 21:38 -0800
Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-07 15:22 +0000
Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-07 18:37 +0200
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-06 14:01 -0800
Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-07 09:05 +0000
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-07 12:48 -0800
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-07 12:50 -0800
Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-08 07:52 +0000
Re: An argument *against* (the liberal use of) references Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-12-08 05:50 -0800
Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-08 13:30 -0800
Re: An argument *against* (the liberal use of) references Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-12-09 12:50 -0800
Re: An argument *against* (the liberal use of) references David Brown <david.brown@hesbynett.no> - 2022-12-11 12:18 +0100
Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-08 13:38 +0200
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-08 16:45 -0800
Re: An argument *against* (the liberal use of) references Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-12-09 04:35 -0800
Re: An argument *against* (the liberal use of) references Bonita Montero <Bonita.Montero@gmail.com> - 2022-11-22 15:36 +0100
Re: An argument *against* (the liberal use of) references Chris Vine <chris@cvine--nospam--.freeserve.co.uk> - 2022-11-22 15:21 +0000
Re: An argument *against* (the liberal use of) references Richard Damon <Richard@Damon-Family.org> - 2022-11-22 10:48 -0500
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-22 13:18 -0800
Re: An argument *against* (the liberal use of) references El Jo <giorgio.zoppi@gmail.com> - 2022-11-22 13:41 -0800
Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-22 13:51 -0800
Re: An argument *against* (the liberal use of) references Bonita Montero <Bonita.Montero@gmail.com> - 2022-11-23 20:08 +0100
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-12-02 14:16 -0800 |
| Message-ID | <tmdtg0$358pk$6@dont-email.me> |
| In reply to | #87699 |
On 12/2/2022 1:18 PM, Scott Lurndal wrote: > "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >> On 12/2/2022 12:55 PM, Scott Lurndal wrote: >>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >>>> On 12/2/2022 6:36 AM, Scott Lurndal wrote: >>>>> Michael S <already5chosen@yahoo.com> writes: >>>>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote: >>>>>>> Paavo Helde <ees...@osa.pri.ee> writes: >>>>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas: >>>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes: >>>>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas: >>>>>>>>> >>>>> >>>>>>> >>>>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the >>>>>>> systemwide lock is the only possibility [*]. >>>>>>> >>>>>> >>>>>> Huh? >>>>>> There is absolute no relationship between what prefix is used (instruction >>>>>> encoding issue) and implementation. >>>>> >>>>> The only way to specify an atomic access in those chips was >>>>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic >>>>> add, et alia). >>>>> >>>> >>>> Iirc, XCHG has an implicit LOCK prefix? >>> >>> Not sure about XCHG, but CMPXCHG requires the prefix >>> when used in a multiprocessor system; on a uniprocessor >>> it will be atomic because interrupts are always taken >>> between instructions (unlike, the VAX, for instance, >>> where certain instructions (MOVC3/5) can be interrupted >>> and restarted). I suspect that XCHG has similar >>> characteristics. >>> >>> >> >> Iirc, XCHG is the _only_ atomic RMW instruction that has an _implicit_ >> LOCK prefix. CMPXCHG _needs_ the programmer to put in a LOCK prefix. > > Yes, that is the case > > XCHG: > "If a memory operand is referenced, the processor's locking protocol is automatically > implemented for the duration of the exchange operation, regardless of the presence > or absence of the LOCK prefix or of the value of the IOPL. (See the LOCK prefix > description in this chapter for more information on the locking protocol.) > Indeed. When I remembered this I kind of doubted myself for a moment. Then, I said well, I know its true, and if I am wrong then I must of fried my brain a bit.
[toc] | [prev] | [next] | [standalone]
| From | Branimir Maksimovic <branimir.maksimovic@icloud.com> |
|---|---|
| Date | 2022-12-02 21:09 +0000 |
| Message-ID | <jQtiL.1308$_Y84.549@fx46.iad> |
| In reply to | #87696 |
On 2022-12-02, Scott Lurndal <scott@slp53.sl.home> wrote: > "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >>On 12/2/2022 6:36 AM, Scott Lurndal wrote: >>> Michael S <already5chosen@yahoo.com> writes: >>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote: >>>>> Paavo Helde <ees...@osa.pri.ee> writes: >>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas: >>>>>>> Paavo Helde <ees...@osa.pri.ee> writes: >>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas: >>>>>>> >>> >>>>> >>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the >>>>> systemwide lock is the only possibility [*]. >>>>> >>>> >>>> Huh? >>>> There is absolute no relationship between what prefix is used (instruction >>>> encoding issue) and implementation. >>> >>> The only way to specify an atomic access in those chips was >>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic >>> add, et alia). >>> >> >>Iirc, XCHG has an implicit LOCK prefix? > > Not sure about XCHG, but CMPXCHG requires the prefix > when used in a multiprocessor system; on a uniprocessor > it will be atomic because interrupts are always taken > between instructions (unlike, the VAX, for instance, > where certain instructions (MOVC3/5) can be interrupted > and restarted). I suspect that XCHG has similar > characteristics. > > "lock" prefix causes the processor's bus-lock signal to be asserted during execution of the accompanying instruction. In a multiprocessor environment, the bus-lock signal insures that the processor has exclusive use of any shared memory while the signal is asserted. The "lock" prefix can be prepended only to the following instructions and only to those forms of the instructions where the destination operand is a memory operand: "add", "adc", "and", "btc", "btr", "bts", "cmpxchg", "cmpxchg8b", "dec", "inc", "neg", "not", "or", "sbb", "sub", "xor", "xadd" and "xchg". If the "lock" prefix is used with one of these instructions and the source operand is a memory operand, an undefined opcode exception may be generated. An undefined opcode exception will also be generated if the "lock" prefix is used with any instruction not in the above list. The "xchg" instruction always asserts the bus-lock signal regardless of the presence or absence of the "lock" prefix. -- 7-77-777 Evil Sinner! with software, you repeat same experiment, expecting different results...
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-12-02 13:22 -0800 |
| Message-ID | <tmdqa9$358pk$1@dont-email.me> |
| In reply to | #87698 |
On 12/2/2022 1:09 PM, Branimir Maksimovic wrote: > On 2022-12-02, Scott Lurndal <scott@slp53.sl.home> wrote: >> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >>> On 12/2/2022 6:36 AM, Scott Lurndal wrote: >>>> Michael S <already5chosen@yahoo.com> writes: >>>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote: >>>>>> Paavo Helde <ees...@osa.pri.ee> writes: >>>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas: >>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes: >>>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas: >>>>>>>> >>>> >>>>>> >>>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the >>>>>> systemwide lock is the only possibility [*]. >>>>>> >>>>> >>>>> Huh? >>>>> There is absolute no relationship between what prefix is used (instruction >>>>> encoding issue) and implementation. >>>> >>>> The only way to specify an atomic access in those chips was >>>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic >>>> add, et alia). >>>> >>> >>> Iirc, XCHG has an implicit LOCK prefix? >> >> Not sure about XCHG, but CMPXCHG requires the prefix >> when used in a multiprocessor system; on a uniprocessor >> it will be atomic because interrupts are always taken >> between instructions (unlike, the VAX, for instance, >> where certain instructions (MOVC3/5) can be interrupted >> and restarted). I suspect that XCHG has similar >> characteristics. >> >> > "lock" prefix causes the processor's bus-lock signal to be asserted during > execution of the accompanying instruction. In a multiprocessor environment, > the bus-lock signal insures that the processor has exclusive use of any shared > memory while the signal is asserted. The "lock" prefix can be prepended only > to the following instructions and only to those forms of the instructions > where the destination operand is a memory operand: "add", "adc", "and", "btc", > "btr", "bts", "cmpxchg", "cmpxchg8b", "dec", "inc", "neg", "not", "or", "sbb", > "sub", "xor", "xadd" and "xchg". If the "lock" prefix is used with one of > these instructions and the source operand is a memory operand, an undefined > opcode exception may be generated. An undefined opcode exception will also be > generated if the "lock" prefix is used with any instruction not in the above > list. > The "xchg" instruction always asserts the bus-lock signal regardless of > the presence or absence of the "lock" prefix. ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ BINGO! I had a strong feeling I was right. https://youtu.be/TnZrWWUFl8I
[toc] | [prev] | [next] | [standalone]
| From | Branimir Maksimovic <branimir.maksimovic@icloud.com> |
|---|---|
| Date | 2022-12-02 21:36 +0000 |
| Message-ID | <6duiL.41721$f9D6.225@fx09.iad> |
| In reply to | #87700 |
On 2022-12-02, Chris M. Thomasson <chris.m.thomasson.1@gmail.com> wrote: > On 12/2/2022 1:09 PM, Branimir Maksimovic wrote: >> On 2022-12-02, Scott Lurndal <scott@slp53.sl.home> wrote: >>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >>>> On 12/2/2022 6:36 AM, Scott Lurndal wrote: >>>>> Michael S <already5chosen@yahoo.com> writes: >>>>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote: >>>>>>> Paavo Helde <ees...@osa.pri.ee> writes: >>>>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas: >>>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes: >>>>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas: >>>>>>>>> >>>>> >>>>>>> >>>>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the >>>>>>> systemwide lock is the only possibility [*]. >>>>>>> >>>>>> >>>>>> Huh? >>>>>> There is absolute no relationship between what prefix is used (instruction >>>>>> encoding issue) and implementation. >>>>> >>>>> The only way to specify an atomic access in those chips was >>>>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic >>>>> add, et alia). >>>>> >>>> >>>> Iirc, XCHG has an implicit LOCK prefix? >>> >>> Not sure about XCHG, but CMPXCHG requires the prefix >>> when used in a multiprocessor system; on a uniprocessor >>> it will be atomic because interrupts are always taken >>> between instructions (unlike, the VAX, for instance, >>> where certain instructions (MOVC3/5) can be interrupted >>> and restarted). I suspect that XCHG has similar >>> characteristics. >>> >>> >> "lock" prefix causes the processor's bus-lock signal to be asserted during >> execution of the accompanying instruction. In a multiprocessor environment, >> the bus-lock signal insures that the processor has exclusive use of any shared >> memory while the signal is asserted. The "lock" prefix can be prepended only >> to the following instructions and only to those forms of the instructions >> where the destination operand is a memory operand: "add", "adc", "and", "btc", >> "btr", "bts", "cmpxchg", "cmpxchg8b", "dec", "inc", "neg", "not", "or", "sbb", >> "sub", "xor", "xadd" and "xchg". If the "lock" prefix is used with one of >> these instructions and the source operand is a memory operand, an undefined >> opcode exception may be generated. An undefined opcode exception will also be >> generated if the "lock" prefix is used with any instruction not in the above >> list. > > >> The "xchg" instruction always asserts the bus-lock signal regardless of >> the presence or absence of the "lock" prefix. > ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ > > BINGO! I had a strong feeling I was right. > > https://youtu.be/TnZrWWUFl8I Nice music :p -- 7-77-777 Evil Sinner! with software, you repeat same experiment, expecting different results...
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-12-02 14:13 -0800 |
| Message-ID | <tmdtao$358pk$5@dont-email.me> |
| In reply to | #87701 |
On 12/2/2022 1:36 PM, Branimir Maksimovic wrote: > On 2022-12-02, Chris M. Thomasson <chris.m.thomasson.1@gmail.com> wrote: >> On 12/2/2022 1:09 PM, Branimir Maksimovic wrote: >>> On 2022-12-02, Scott Lurndal <scott@slp53.sl.home> wrote: >>>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >>>>> On 12/2/2022 6:36 AM, Scott Lurndal wrote: >>>>>> Michael S <already5chosen@yahoo.com> writes: >>>>>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote: >>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes: >>>>>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas: >>>>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes: >>>>>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas: >>>>>>>>>> >>>>>> >>>>>>>> >>>>>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the >>>>>>>> systemwide lock is the only possibility [*]. >>>>>>>> >>>>>>> >>>>>>> Huh? >>>>>>> There is absolute no relationship between what prefix is used (instruction >>>>>>> encoding issue) and implementation. >>>>>> >>>>>> The only way to specify an atomic access in those chips was >>>>>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic >>>>>> add, et alia). >>>>>> >>>>> >>>>> Iirc, XCHG has an implicit LOCK prefix? >>>> >>>> Not sure about XCHG, but CMPXCHG requires the prefix >>>> when used in a multiprocessor system; on a uniprocessor >>>> it will be atomic because interrupts are always taken >>>> between instructions (unlike, the VAX, for instance, >>>> where certain instructions (MOVC3/5) can be interrupted >>>> and restarted). I suspect that XCHG has similar >>>> characteristics. >>>> >>>> >>> "lock" prefix causes the processor's bus-lock signal to be asserted during >>> execution of the accompanying instruction. In a multiprocessor environment, >>> the bus-lock signal insures that the processor has exclusive use of any shared >>> memory while the signal is asserted. The "lock" prefix can be prepended only >>> to the following instructions and only to those forms of the instructions >>> where the destination operand is a memory operand: "add", "adc", "and", "btc", >>> "btr", "bts", "cmpxchg", "cmpxchg8b", "dec", "inc", "neg", "not", "or", "sbb", >>> "sub", "xor", "xadd" and "xchg". If the "lock" prefix is used with one of >>> these instructions and the source operand is a memory operand, an undefined >>> opcode exception may be generated. An undefined opcode exception will also be >>> generated if the "lock" prefix is used with any instruction not in the above >>> list. >> >> >>> The "xchg" instruction always asserts the bus-lock signal regardless of >>> the presence or absence of the "lock" prefix. >> ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ >> >> BINGO! I had a strong feeling I was right. >> >> https://youtu.be/TnZrWWUFl8I > Nice music :p > Thanks again, Branimir, for the quote of the docs. https://youtu.be/UZ2-FfXZlAU (super mario music in a live big band format? Nice... :^) My hat is off to you.
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-12-02 18:33 +0200 |
| Message-ID | <tmd9d8$341b0$1@dont-email.me> |
| In reply to | #87684 |
02.12.2022 15:53 Michael S kirjutas: > On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote: > BTW, it means that your claim in post above "no overhead at all unless ..." > is incorrect in the absolute sense. There is an overhead even without "unless". > But the overhead in uncontended case is small - order of dozen or two of CPU > clocks. So, undetectable in Paavo's case of only 1000 updates per second. > For 1M updates per second impact would me detectable with precise time > measurements and for 100M per second there would be big slowdown. Just a minor note, earlier I posted numbers for only a single smartpointer, whereas in the full program there are probably tens or hundreds of thousands of them. I just measured it and the total rate of refcount changes is something like 5M per second. With this 5M/s rate I cannot see any slowdown (of using std::atomic<int> instead of int) in my measurements (with no contention), variations caused by other uncontrollable factors seem to be much larger. Maybe I should rerun this on a Linux box where the things are more stable.
[toc] | [prev] | [next] | [standalone]
| From | Michael S <already5chosen@yahoo.com> |
|---|---|
| Date | 2022-12-01 11:59 -0800 |
| Message-ID | <b54ae238-40c3-4003-a49a-8e44dc7d9eb7n@googlegroups.com> |
| In reply to | #87671 |
On Thursday, December 1, 2022 at 8:22:12 PM UTC+2, Paavo Helde wrote: > 01.12.2022 18:02 Scott Lurndal kirjutas: > > Paavo Helde <ees...@osa.pri.ee> writes: > >> 01.12.2022 15:06 Stuart Redmann kirjutas: > > > > <snip> > > > >>> Another thought: if thread-safety is too costly, you could use two smart > >>> pointer classes: thread-safe pointers and forwarding non-thread-safe smart > >>> pointers. The forwarding smart pointers have their own thread-UNsafe > >>> refcount and the thread-safe smart pointer as member. > >> > >> I tried to measure the impact of std::atomic<int> refcounters and in > >> first tests it seems the overhead on x86_64 is zero (with no > >> contention). > > > > Which follows naturally from the fact that the core doing > > the atomic access has exclusive access to the cache line > > containing the refcounter. No overhead at all, unless > Thanks for the clarifications! > > the ref counter isn't aligned and crosses a cache-line > > boundary (or the access is to an uncached memory > > range or caching is disabled), in which case the processor > > will take a system-wide > > lock to perform the operation, which is catastrophic > > on systems with large processor counts. > It is clear that having a misaligned cross-border atomic would be very > bad. But what about normal uncached memory ranges, wouldn't these be > just loaded into the cache, without disturbing other processors, and > without any "catastrophic" consequences? "Unchached memory range" is a misnormer. A proper name is uncacheable range (region). Unfortunately "uncached" in the meaning of "uncacheable" is used quite often. Even Intel's official manuals suffer from such inconsistent vocabulary.
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-12-01 23:53 +0200 |
| Message-ID | <tmb7p9$2sgre$1@dont-email.me> |
| In reply to | #87675 |
01.12.2022 21:59 Michael S kirjutas: > "Unchached memory range" is a misnormer. > A proper name is uncacheable range (region). > Unfortunately "uncached" in the meaning of "uncacheable" is used > quite often. Even Intel's official manuals suffer from such > inconsistent vocabulary. Thanks, I had to look up what is "uncacheable memory". I guess the 80186 processor where I learned my basics did not have such a thing.
[toc] | [prev] | [next] | [standalone]
| From | Michael S <already5chosen@yahoo.com> |
|---|---|
| Date | 2022-12-02 04:11 -0800 |
| Message-ID | <a1c3687f-6b3e-4b60-a0dc-6264b3306026n@googlegroups.com> |
| In reply to | #87678 |
On Thursday, December 1, 2022 at 11:54:02 PM UTC+2, Paavo Helde wrote: > 01.12.2022 21:59 Michael S kirjutas: > > > "Unchached memory range" is a misnormer. > > A proper name is uncacheable range (region). > > Unfortunately "uncached" in the meaning of "uncacheable" is used > > quite often. Even Intel's official manuals suffer from such > > inconsistent vocabulary. > Thanks, I had to look up what is "uncacheable memory". I guess the 80186 > processor where I learned my basics did not have such a thing. 80186 was "embedded" microprocessor similar at core to 8086. It was typically used with no cache so didn't need a concept of uncacheable regions. 80286 and especially i386 was used with (external) cache quite often, but according to my understanding their caches were what we call today "memory-side caches" associated with main memory. From system (both CPU and other bus masters) perspective such caches are totally transparent (except for entering/leaving deep sleep states, but back then they didn't do it) so there still was no need for uncacheable regions. System-side caches and associated problem first appear in x86 world in i486. Still, the in original i486 the system cache had strict write-through policy, so the problems were minor. Then came Pentium with 8 KB of write-back Data cache and a little later came new models of i486 with even bigger write-back cache and problems became quite real, especially because approximately in the same time PCI took over I/O bus role and suddenly multiple bus masters that were high-end curiosity before then, became common in consumer PC hardware. But even then x86 architecture lacked adequate answer to a new challenge. The first reasonable answer (MTRR registers) came only in PPro but it still had problems of scalability - too few regions. Later (P-III) they invented PAT which from theoretical point of view is inferior to MTRRs because in PAT scheme cachability is an attribute of virtual address rather than of physical address. But PAT is ultimately scalable and is one solution that, as long as OS does a proper plumbing, is one solution that can rule over all aspects of cachabilty. So it won.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-12-01 12:21 -0800 |
| Message-ID | <tmb2co$2s3ov$3@dont-email.me> |
| In reply to | #87666 |
On 12/1/2022 5:06 AM, Stuart Redmann wrote: > Paavo Helde <eesnimi@osa.pri.ee> wrote: >> 01.12.2022 08:45 Juha Nieminen kirjutas: >>> Paavo Helde <eesnimi@osa.pri.ee> wrote: >>>> Actually once located, the bug was simple. The object which was copied >>>> was a single-threaded refcounted smartpointer, and by copying it the >>>> refcounter got incremented (and later decremented). Alas, this was >>>> accidentally done from parallel threads at the same time, without any >>>> synchronization, so eventually the refcounter got messed up. >>> >>> I think that if the reference count is declared atomic, it can be safely >>> directly incremented. When decrementing you would need to use the >>> fetch_sub() function to see if the object needs to be destroyed. >>> >>> While modifying an atomic might not be equally fast as a non-atomic, >>> it shouldn't be all that much slower either, at least if the target >>> architecture supports atomic operations. >> >> I have pondered this myself. Maybe I should measure the actual slowdown >> after temporarily making the refcounters atomic. But this seems overkill >> because these smartpointers would still point to single-threaded objects >> which are meant to be primarily used in single-thread regime, so in most >> cases making the smartpointers atomic does not buy anything. >> >> When tracking down this bug, I monitored all refcounter changes for a >> particular single smartpointer during the program run (ca 10 min). There >> were 591848 increments and decrements, from which 1526 came from the >> problematic (parallelized) part. It looks like a pessimization to slow >> down 99.75% of accesses when only 0.25% would actually benefit from this. >> > > 600k changes in reference counts look suspicious to me. When you pass a > ref-counted object to a worker thread, there should only be a single change > in the refcount. This would be because inside the worker thread some object > takes (shared) ownership of shared object. If the shared object needs to be > passed to sub-routines, you should pass them as references or plain > pointers (if the subroutine must be able to cope with non-existing > objects). It should be rare occurrence that another object in the worker > thread needs to take ownership of the shared object. > > Another thought: if thread-safety is too costly, you could use two smart > pointer classes: thread-safe pointers and forwarding non-thread-safe smart > pointers. The forwarding smart pointers have their own thread-UNsafe > refcount and the thread-safe smart pointer as member. Check this out: https://github.com/jseigh/atomic-ptr-plus/blob/master/atomic-ptr/atomic_ptr.h It is a truly atomic reference counted pointer. A thread can take a reference without owning a prior reference. Here is a patent: https://patents.justia.com/patent/5295262
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2022-12-06 11:36 +0000 |
| Message-ID | <tmn9fl$1kpl$1@gioia.aioe.org> |
| In reply to | #87665 |
Paavo Helde <eesnimi@osa.pri.ee> wrote: > When tracking down this bug, I monitored all refcounter changes for a > particular single smartpointer during the program run (ca 10 min). There > were 591848 increments and decrements, from which 1526 came from the > problematic (parallelized) part. It looks like a pessimization to slow > down 99.75% of accesses when only 0.25% would actually benefit from this. There's place for micro-optimization and there's place to do the Right Thing (TM) instead. In the vast, vast majority of situations micro-optimization will have little to no effect on the program. It's only when you have number crunching code that does something billions of times per second that micro-optimization may start having some discernible effect. Those situations tend to be very rare and far-in-between. And when you do have such situations you can make faster versions of things for that alone. Micro-optimization is extra useless if it's surrounded by, and thus swamped by code that's a lot slower than it. Even when micro-optimizing you should start with the worst offenders, not the smallest things. (By "micro-optimization" I'm referring to things that do not change the computational complexity of something and only makes that something some clock cycles faster.) So unless your smart pointer is being copied and assigned around millions of times per second in tight number-crunching loops, you can safely ignore any lost clock cycles by making the reference counter thread-safe. (If you actually need to copy and assign smart pointers around millions of times per second in a tight number-crunching inner loop, perhaps create a specialized version of the pointer for that particular purpose...)
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-12-06 12:26 -0800 |
| Message-ID | <tmo8ho$ahd1$1@dont-email.me> |
| In reply to | #87706 |
On 12/6/2022 3:36 AM, Juha Nieminen wrote: > Paavo Helde <eesnimi@osa.pri.ee> wrote: >> When tracking down this bug, I monitored all refcounter changes for a >> particular single smartpointer during the program run (ca 10 min). There >> were 591848 increments and decrements, from which 1526 came from the >> problematic (parallelized) part. It looks like a pessimization to slow >> down 99.75% of accesses when only 0.25% would actually benefit from this. [...] > So unless your smart pointer is being copied and assigned around > millions of times per second in tight number-crunching loops, you can > safely ignore any lost clock cycles by making the reference counter > thread-safe. [...] A thread-safe reference counted pointer can heavily damage performance in certain usage scenarios. Blasting the system with memory barriers and atomic RMW ops all over the place. Now, there is a work around called proxy reference counting. Are you familiar with it?
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-12-06 20:39 +0000 |
| Message-ID | <sMNjL.66$ZhSc.56@fx38.iad> |
| In reply to | #87707 |
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >On 12/6/2022 3:36 AM, Juha Nieminen wrote: >> Paavo Helde <eesnimi@osa.pri.ee> wrote: >>> When tracking down this bug, I monitored all refcounter changes for a >>> particular single smartpointer during the program run (ca 10 min). There >>> were 591848 increments and decrements, from which 1526 came from the >>> problematic (parallelized) part. It looks like a pessimization to slow >>> down 99.75% of accesses when only 0.25% would actually benefit from this. >[...] >> So unless your smart pointer is being copied and assigned around >> millions of times per second in tight number-crunching loops, you can >> safely ignore any lost clock cycles by making the reference counter >> thread-safe. >[...] >A thread-safe reference counted pointer can heavily damage performance >in certain usage scenarios. Blasting the system with memory barriers and >atomic RMW ops all over the place. Now, there is a work around called >proxy reference counting. Are you familiar with it? The fact that smart pointers do allocation/deallocation has made them useless for high-performance threaded code, IMO.
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-12-06 23:53 +0200 |
| Message-ID | <tmodko$avn3$1@dont-email.me> |
| In reply to | #87708 |
06.12.2022 22:39 Scott Lurndal kirjutas: > > The fact that smart pointers do allocation/deallocation has made > them useless for high-performance threaded code, IMO. You have got it backwards. Smartpointers are taken into use for coping with the fact that objects need to by dynamically allocated and deallocated, by the program logic. And this allocation/deallocation would happen relatively rarely. If the object lifetimes were short, then typically they could be controlled much better, and there would be no need for refcounted smartpointers, or maybe even no need for dynamic allocation of objects in the first place.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-12-06 13:59 -0800 |
| Message-ID | <tmodvl$auhg$2@dont-email.me> |
| In reply to | #87709 |
On 12/6/2022 1:53 PM, Paavo Helde wrote: > 06.12.2022 22:39 Scott Lurndal kirjutas: > >> >> The fact that smart pointers do allocation/deallocation has made >> them useless for high-performance threaded code, IMO. > > You have got it backwards. Smartpointers are taken into use for coping > with the fact that objects need to by dynamically allocated and > deallocated, by the program logic. > > And this allocation/deallocation would happen relatively rarely. If the > object lifetimes were short, then typically they could be controlled > much better, and there would be no need for refcounted smartpointers, or > maybe even no need for dynamic allocation of objects in the first place. > > > Imvvho, a smart pointer should not need to allocate anything under the covers. Also, are you familiar with proxy reference counting? It has the ability to amortize a single reference over n objects.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-12-06 22:04 +0000 |
| Message-ID | <A%OjL.2208$iS99.2139@fx16.iad> |
| In reply to | #87710 |
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >On 12/6/2022 1:53 PM, Paavo Helde wrote: >> 06.12.2022 22:39 Scott Lurndal kirjutas: >> >>> >>> The fact that smart pointers do allocation/deallocation has made >>> them useless for high-performance threaded code, IMO. >> >> You have got it backwards. Smartpointers are taken into use for coping >> with the fact that objects need to by dynamically allocated and >> deallocated, by the program logic. >> >> And this allocation/deallocation would happen relatively rarely. Assumption not in evidence. I've personnally had to rip smart pointers out of code because the allocation/deallocation happened very frequently. One if the applications was simulating a processor pipeline, another was handling network packets both were written by well-educated people familiar with C++. Granted, one can specify a more efficient allocator, but 1) most C++ programmers don't bother or don't know how 2) Even then there is unnecessary overhead unless the allocator is pool based. KISS applies, always.
[toc] | [prev] | [next] | [standalone]
| From | Öö Tiib <ootiib@hot.ee> |
|---|---|
| Date | 2022-12-06 21:38 -0800 |
| Message-ID | <4f4fff56-27ce-4181-910d-fb2fa0307b67n@googlegroups.com> |
| In reply to | #87712 |
On Wednesday, 7 December 2022 at 00:04:34 UTC+2, Scott Lurndal wrote: > "Chris M. Thomasson" <chris.m.t...@gmail.com> writes: > >On 12/6/2022 1:53 PM, Paavo Helde wrote: > >> 06.12.2022 22:39 Scott Lurndal kirjutas: > >> > >>> > >>> The fact that smart pointers do allocation/deallocation has made > >>> them useless for high-performance threaded code, IMO. > >> > >> You have got it backwards. Smartpointers are taken into use for coping > >> with the fact that objects need to by dynamically allocated and > >> deallocated, by the program logic. > >> > >> And this allocation/deallocation would happen relatively rarely. > Assumption not in evidence. I've personnally had to rip smart pointers > out of code because the allocation/deallocation happened very > frequently. One if the applications was simulating a processor pipeline, > another was handling network packets both were written by well-educated > people familiar with C++. > > Granted, one can specify a more efficient allocator, but > > 1) most C++ programmers don't bother or don't know how > 2) Even then there is unnecessary overhead unless the allocator is pool based. > > KISS applies, always. No one argues with that. Just that keeping it simple is far from simple. For example it is tricky to keep dynamic allocations minimal. That is not fault of smart pointers. The std::unique_ptr helps at places where dynamic allocations are needed greatly, especially when there can be exceptions. It has next to no overhead. Yes, the std::shared_ptr is loaded. Usage of std::make_shared helps a bit but thinking about how to make it simpler and to get rid of shared ownership or even dynamic allocations is hard and not always fruitful.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-12-07 15:22 +0000 |
| Message-ID | <Ic2kL.1981$jiuc.1610@fx44.iad> |
| In reply to | #87713 |
=?UTF-8?B?w5bDtiBUaWli?= <ootiib@hot.ee> writes: >On Wednesday, 7 December 2022 at 00:04:34 UTC+2, Scott Lurndal wrote: >> "Chris M. Thomasson" <chris.m.t...@gmail.com> writes: >> >On 12/6/2022 1:53 PM, Paavo Helde wrote: >> >> 06.12.2022 22:39 Scott Lurndal kirjutas: >> >> >> >>> >> >>> The fact that smart pointers do allocation/deallocation has made >> >>> them useless for high-performance threaded code, IMO. >> >> >> >> You have got it backwards. Smartpointers are taken into use for coping >> >> with the fact that objects need to by dynamically allocated and >> >> deallocated, by the program logic. >> >> >> >> And this allocation/deallocation would happen relatively rarely. >> Assumption not in evidence. I've personnally had to rip smart pointers >> out of code because the allocation/deallocation happened very >> frequently. One if the applications was simulating a processor pipeline, >> another was handling network packets both were written by well-educated >> people familiar with C++. >> >> Granted, one can specify a more efficient allocator, but >> >> 1) most C++ programmers don't bother or don't know how >> 2) Even then there is unnecessary overhead unless the allocator is pool based. >> >> KISS applies, always. > >No one argues with that. Just that keeping it simple is far from simple. >For example it is tricky to keep dynamic allocations minimal. That is >not fault of smart pointers. In my experience, it has been generally sufficent to pre-allocate the data structures and store them in a table or look-aside list, as the maximum number is bounded. For example, an application handling network packets on a processor with 64 cores, may only need 128 jumbo packet buffers if the packet processing thread count matches the core count. These can be preallocated and then passed as regular pointers throughout the flow.\ (Specialized DPUs have a custom hardware block (network pool allocator) that allocates hardware buffers to packets on ingress and those buffers are passed by hardware to the other blocks in the flow, such as blocks to identify the flow, fragment/defragment a packet, apply encryption/decryption algorithms, all controlled by a hardware scheduler block, etc.). Likewise for a simulation of an internal processor interconnect such as ring or mesh structure, there are a fixed maximum number of flits than can be active at any point in time. Preallocating them into a lookaside list eliminates allocation and deallocation overhead on every flit. When simulating a full SoC, the maximum inflight objects is likewise bounded and for the most part, can be preallocated.
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-12-07 18:37 +0200 |
| Message-ID | <tmqffe$jl7n$1@dont-email.me> |
| In reply to | #87725 |
07.12.2022 17:22 Scott Lurndal kirjutas: > =?UTF-8?B?w5bDtiBUaWli?= <ootiib@hot.ee> writes: >> >> No one argues with that. Just that keeping it simple is far from simple. >> For example it is tricky to keep dynamic allocations minimal. That is >> not fault of smart pointers. > > In my experience, it has been generally sufficent to pre-allocate the > data structures and store them in a table or look-aside list, > as the maximum number is bounded. My experience is more that the user wants to read in unknown number of tiff files containing unknown number of image frames of unknown sizes, then start to process them by script-driven flexible algorithms, producing an unknown number of intermediate and final results of unknown size. And this processing ought to be as fast as possible, as nobody wants to wait for hours (although with large data sets and complex processing it inevitable gets into hours). And this processing ought better to make use of all the cpu cores and should not run out of computer memory while doing that. So it seems preallocating a fixed number of data structures of fixed size would not really work in my case.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-12-06 14:01 -0800 |
| Message-ID | <tmoe3m$auhg$3@dont-email.me> |
| In reply to | #87709 |
On 12/6/2022 1:53 PM, Paavo Helde wrote: > 06.12.2022 22:39 Scott Lurndal kirjutas: > >> >> The fact that smart pointers do allocation/deallocation has made >> them useless for high-performance threaded code, IMO. > > You have got it backwards. Smartpointers are taken into use for coping > with the fact that objects need to by dynamically allocated and > deallocated, by the program logic. > > And this allocation/deallocation would happen relatively rarely. If the > object lifetimes were short, then typically they could be controlled > much better, and there would be no need for refcounted smartpointers, or > maybe even no need for dynamic allocation of objects in the first place. > > > Fwiw, RCU is a form of proxy collection. Fwiw, I wrote an experimental one using pure C++. https://pastebin.com/raw/f71480694 (goes to pure text page, no ads and shit like that...)
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | comp.lang.c++
csiph-web