Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #87211 > unrolled thread
| Started by | Lynn McGuire <lynnmcguire5@gmail.com> |
|---|---|
| First post | 2022-11-03 14:30 -0500 |
| Last post | 2022-11-09 07:59 -0500 |
| Articles | 17 on this page of 37 — 13 participants |
Back to article view | Back to comp.lang.c++
“The pool of talented C++ developers is running dry” Lynn McGuire <lynnmcguire5@gmail.com> - 2022-11-03 14:30 -0500
Re: “The pool of talented C++ developers is running dry” "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-03 13:03 -0700
Re: “The pool of talented C++ developers is running dry” "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-03 13:04 -0700
Re: “The pool of talented C++ developers is running dry” Lynn McGuire <lynnmcguire5@gmail.com> - 2022-11-03 16:07 -0500
Re: “The pool of talented C++ developers is running dry” Lynn McGuire <lynnmcguire5@gmail.com> - 2022-11-03 16:05 -0500
Re: “The pool of talented C++ developers is running dry” "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-03 14:25 -0700
Re: “The pool of talented C++ developers is running dry” Lynn McGuire <lynnmcguire5@gmail.com> - 2022-11-05 14:12 -0500
Re: “The pool of talented C++ developers is running dry” "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-05 14:20 -0700
Re: “The pool of talented C++ developers is running dry” Bonita Montero <Bonita.Montero@gmail.com> - 2022-11-04 09:24 +0100
Re: “The pool of talented C++ developers is running dry” Vir Campestris <vir.campestris@invalid.invalid> - 2022-11-05 21:42 +0000
Re: “The pool of talented C++ developers is running dry” "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-06 11:55 -0800
Re: “The pool of talented C++ developers is running dry” Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-11-08 02:44 -0800
Re: ???The pool of talented C++ developers is running dry??? Juha Nieminen <nospam@thanks.invalid> - 2022-11-08 11:29 +0000
Re: ???The pool of talented C++ developers is running dry??? Öö Tiib <ootiib@hot.ee> - 2022-11-08 03:54 -0800
Re: ???The pool of talented C++ developers is running dry??? Stuart Redmann <DerTopper@web.de> - 2022-11-08 14:26 +0100
Re: ???The pool of talented C++ developers is running dry??? Öö Tiib <ootiib@hot.ee> - 2022-11-09 00:22 -0800
Re: ???The pool of talented C++ developers is running dry??? Juha Nieminen <nospam@thanks.invalid> - 2022-11-08 15:16 +0000
Re: ???The pool of talented C++ developers is running dry??? Bonita Montero <Bonita.Montero@gmail.com> - 2022-11-08 14:52 +0100
Re: ???The pool of talented C++ developers is running dry??? Jorgen Grahn <grahn+nntp@snipabacken.se> - 2022-11-08 22:48 +0000
Re: ???The pool of talented C++ developers is running dry??? "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-08 15:39 -0800
Re: ???The pool of talented C++ developers is running dry??? David Brown <david.brown@hesbynett.no> - 2022-11-09 09:29 +0100
Re: ???The pool of talented C++ developers is running dry??? scott@slp53.sl.home (Scott Lurndal) - 2022-11-09 14:39 +0000
Re: ???The pool of talented C++ developers is running dry??? David Brown <david.brown@hesbynett.no> - 2022-11-09 17:05 +0100
Re: ???The pool of talented C++ developers is running dry??? scott@slp53.sl.home (Scott Lurndal) - 2022-11-09 17:58 +0000
Re: ???The pool of talented C++ developers is running dry??? "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-09 11:47 -0800
Re: ???The pool of talented C++ developers is running dry??? scott@slp53.sl.home (Scott Lurndal) - 2022-11-09 20:43 +0000
Re: ???The pool of talented C++ developers is running dry??? David Brown <david.brown@hesbynett.no> - 2022-11-10 09:52 +0100
Re: ???The pool of talented C++ developers is running dry??? Michael S <already5chosen@yahoo.com> - 2022-11-10 03:07 -0800
Re: ???The pool of talented C++ developers is running dry??? "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-10 11:53 -0800
Re: ???The pool of talented C++ developers is running dry??? David Brown <david.brown@hesbynett.no> - 2022-11-10 21:44 +0100
Re: ???The pool of talented C++ developers is running dry??? "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-10 12:53 -0800
Re: ???The pool of talented C++ developers is running dry??? scott@slp53.sl.home (Scott Lurndal) - 2022-11-10 21:20 +0000
Re: ???The pool of talented C++ developers is running dry??? "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-10 13:23 -0800
Re: ???The pool of talented C++ developers is running dry??? David Brown <david.brown@hesbynett.no> - 2022-11-11 08:10 +0100
Re: ???The pool of talented C++ developers is running dry??? Öö Tiib <ootiib@hot.ee> - 2022-11-11 01:45 -0800
Re: ???The pool of talented C++ developers is running dry??? Michael S <already5chosen@yahoo.com> - 2022-11-10 03:14 -0800
Re: ???The pool of talented C++ developers is running dry??? Sam <sam@email-scan.com> - 2022-11-09 07:59 -0500
Page 2 of 2 — ← Prev page 1 [2]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2022-11-09 09:29 +0100 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <tkfoch$695m$1@dont-email.me> |
| In reply to | #87294 |
On 09/11/2022 00:39, Chris M. Thomasson wrote: > On 11/8/2022 3:29 AM, Juha Nieminen wrote: >> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote: >>> You keep on adding features to the language which have unintuitive >>> syntax and odd rules, and >>> don't do much to increase the number of programs you can write >>> quickly. So what happens? >>> Theres not much motivation to learn these features until forced to do >>> so. So codebases tend >>> to be mainly legacy, and C++ programmers' skills fall behind. >> >> For the longest time I quite strongly disagreed with the claim that >> C++ is >> becoming too big and too complicated. >> >> However, C++20 has eroded this conviction of mine somewhat. C++23 is >> eroding >> it even more. >> >> C++11 felt like a big bunch of features that the language was in dire >> need >> of, and genuinely made programming easier. C++14 and C++17 fixed and >> patched >> many of the minor problems and defects that turned out to exist in C++11, >> so C++17 felt like "what C++11 should have been in the first place". I agree with that. A challenge for C++ is that even when a new and better feature is added, the older and clumsier methods still have to be supported. This also means that syntax can be awkward because it can't conflict with existing syntax, and the details get more complex all the time. > > I was really excited and happy when C++ finally made atomics and membars > part of the actual standard, C++11 iirc. Before that, I would have to > code these things up in assembly language. > Standard atomics would be great if they worked for my targets. The gcc implementations (and I haven't seen any others) for "advanced" use (read-modify-write, or sizes larger than a standard register) is completely broken for single-core systems, and even on multi-core systems it is limited if you use thread priorities. The trouble with them is that no one has addressed the elephant in the room - in general, you need OS support and locks to implement large atomics.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-11-09 14:39 +0000 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <CYOaL.12305$eyq6.10545@fx03.iad> |
| In reply to | #87296 |
David Brown <david.brown@hesbynett.no> writes: >On 09/11/2022 00:39, Chris M. Thomasson wrote: >> On 11/8/2022 3:29 AM, Juha Nieminen wrote: > >> I was really excited and happy when C++ finally made atomics and membars >> part of the actual standard, C++11 iirc. Before that, I would have to >> code these things up in assembly language. >> > >Standard atomics would be great if they worked for my targets. The gcc >implementations (and I haven't seen any others) for "advanced" use >(read-modify-write, or sizes larger than a standard register) is >completely broken for single-core systems, and even on multi-core >systems it is limited if you use thread priorities. > The trouble with >them is that no one has addressed the elephant in the room - in general, >you need OS support and locks to implement large atomics. IFF the target architecture doesn't have a comprehensive set of atomic access instructions, perhaps. ARMv8 LSE, for example, has individual instructions for most of the gcc atomic intrinsics (e.g. __sync_fetch_and_add will generate a single LDADD atomic instruction). The instructions support the common arithmetic operations (add, or, etc). Before LSE, the ARMv8 implementations were built using the arm LL/SC equivalent (load exclusive/store exclusive) instructions.
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2022-11-09 17:05 +0100 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <tkgj53$8nkh$1@dont-email.me> |
| In reply to | #87299 |
On 09/11/2022 15:39, Scott Lurndal wrote: > David Brown <david.brown@hesbynett.no> writes: >> On 09/11/2022 00:39, Chris M. Thomasson wrote: >>> On 11/8/2022 3:29 AM, Juha Nieminen wrote: >> >>> I was really excited and happy when C++ finally made atomics and membars >>> part of the actual standard, C++11 iirc. Before that, I would have to >>> code these things up in assembly language. >>> >> >> Standard atomics would be great if they worked for my targets. The gcc >> implementations (and I haven't seen any others) for "advanced" use >> (read-modify-write, or sizes larger than a standard register) is >> completely broken for single-core systems, and even on multi-core >> systems it is limited if you use thread priorities. > >> The trouble with >> them is that no one has addressed the elephant in the room - in general, >> you need OS support and locks to implement large atomics. > > IFF the target architecture doesn't have a comprehensive set of > atomic access instructions, perhaps. > > ARMv8 LSE, for example, has individual instructions for most of the > gcc atomic intrinsics (e.g. __sync_fetch_and_add will generate a single > LDADD atomic instruction). The instructions support the common > arithmetic operations (add, or, etc). > > Before LSE, the ARMv8 implementations were built using the arm > LL/SC equivalent (load exclusive/store exclusive) instructions. You are more familiar with the details of these things than most people, so I hope you (or someone else) will correct me if my logic below is wrong. There's no problem when the target has a single unbreakable instruction for the action. And LL/SC are fine for atomic loads or stores of different sizes. But LL/SC is not sufficient for read-modify-write sequences of a size larger than can be handled by a single atomic instruction. Imagine you have a processor that can atomically read or write an unsigned integer type "uint". Your sequence for "uint_inc" will be : retry: load link x = *p x++ if (store conditional *p = x fails) goto retry If two processes try this, they can interleave and be started or stopped without trouble - the result will be an atomic increment. Now consider a double-sized type containing two "uint" fields: retry: load link x_lo = *p x_hi = *(p + 1) x_lo++ if (!x_lo) x_hi++ if (store conditional *p = x_lo fails) goto retry *(p + 1) = x_hi If the process executing this is stopped after the first write, and a second process is run that calls a similar function, then the new process will see a half-changed value for the object resulting in a corrupted object. Resumption of the first process will half-change the value again. Different combinations of using "store_conditional" on the two stores will result in similar problems. The only way to make a multi-unit RMW operation work is if other processes are /blocked/ from breaking in during the actual write sequence. Reads and the calculation can be re-retried, but not the writes - they must be made an unbreakable sequence. And that, in general, means a lock and OS support to ensure that the locking process gets to finish. The gcc implementation of atomic operations (larger than can be handled with a single instruction) uses simple user-space spin locks (the lock can be accessed atomically - with an LL/SC sequence, for the ARM). If one process tries to access the atomic while another process has the lock, it will spin - running a busy wait loop. As long as these processes are running on different cores, there's no problem with one core running a few rounds of a tight loop while another core does a quick load or store. Given that contention is rare and cores are often plentiful, this results in a very efficient atomic operation. But it can deadlock - a process could take the spin lock and then get descheduled by the OS, and other threads wanting the lock could be activated. If these fill up the cores (maybe you have multiple threads all using the same supposedly lock-free atomic structure), you are screwed. And if you have only one core (like almost all microcontrollers), and the thread that has the lock is interrupted by an interrupt routine that wants to access the same atomic variable, you are /really/ screwed. This can happen with such simple code as a 64-bit atomic counter in an interrupt routine that is also accessed atomically from a background task. It's very unlikely that you'll hit a problem, but it is possible. To me, that is useless - atomics need guaranteed forward progress. That means the std::atomic<> stuff needs to use OS-level locks for advanced cases that can't be handled directly by instructions or LL/SC sequences, or for a microcontroller you'd want to disable interrupts around the access. The alternative is to refuse to compile the operations and only support atomics that are smaller or simpler.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-11-09 17:58 +0000 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <CTRaL.31519$NeJ8.1285@fx09.iad> |
| In reply to | #87300 |
David Brown <david.brown@hesbynett.no> writes:
>On 09/11/2022 15:39, Scott Lurndal wrote:
>> David Brown <david.brown@hesbynett.no> writes:
>>> On 09/11/2022 00:39, Chris M. Thomasson wrote:
>>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote:
>>>
>>>> I was really excited and happy when C++ finally made atomics and membars
>>>> part of the actual standard, C++11 iirc. Before that, I would have to
>>>> code these things up in assembly language.
>>>>
>>>
>>> Standard atomics would be great if they worked for my targets. The gcc
>>> implementations (and I haven't seen any others) for "advanced" use
>>> (read-modify-write, or sizes larger than a standard register) is
>>> completely broken for single-core systems, and even on multi-core
>>> systems it is limited if you use thread priorities.
>>
>>> The trouble with
>>> them is that no one has addressed the elephant in the room - in general,
>>> you need OS support and locks to implement large atomics.
>>
>> IFF the target architecture doesn't have a comprehensive set of
>> atomic access instructions, perhaps.
>>
>> ARMv8 LSE, for example, has individual instructions for most of the
>> gcc atomic intrinsics (e.g. __sync_fetch_and_add will generate a single
>> LDADD atomic instruction). The instructions support the common
>> arithmetic operations (add, or, etc).
>>
>> Before LSE, the ARMv8 implementations were built using the arm
>> LL/SC equivalent (load exclusive/store exclusive) instructions.
>
>
>You are more familiar with the details of these things than most people,
>so I hope you (or someone else) will correct me if my logic below is wrong.
>
>
>There's no problem when the target has a single unbreakable instruction
>for the action. And LL/SC are fine for atomic loads or stores of
>different sizes.
Here's the code generated by GCC for
q = __sync_fetch_and_add(&q, 1u);
Without LSE (atomics) support:
401034: 885ffc60 ldaxr w0, [x3]
401038: 11000401 add w1, w0, #0x1
40103c: 8804fc61 stlxr w4, w1, [x3]
c01040: 35ffffa4 cbnz w4, 401034 <main+0x34>
With LSE (atomics) support:
12c: b8e10001 ldaddal w1, w1, [x0]
>
>But LL/SC is not sufficient for read-modify-write sequences of a size
>larger than can be handled by a single atomic instruction.
>
>Imagine you have a processor that can atomically read or write an
>unsigned integer type "uint". Your sequence for "uint_inc" will be :
>
>retry:
> load link x = *p
> x++
> if (store conditional *p = x fails) goto retry
>
>
>If two processes try this, they can interleave and be started or stopped
>without trouble - the result will be an atomic increment.
>
>Now consider a double-sized type containing two "uint" fields:
>
>retry:
> load link x_lo = *p
> x_hi = *(p + 1)
> x_lo++
> if (!x_lo) x_hi++
> if (store conditional *p = x_lo fails) goto retry
> *(p + 1) = x_hi
For such sequences, one uses the LL/SC as a spinlock;
acquire the spinlock, perform the non-atomic operation
and release the spinlock. On uniprocessor systems,
alternate mechanisms like disabling interrupts are the
common solution.
Although in this case, using a wider type if available is a
better option.
>
>If the process executing this is stopped after the first write, and a
>second process is run that calls a similar function, then the new
>process will see a half-changed value for the object resulting in a
>corrupted object. Resumption of the first process will half-change the
>value again. Different combinations of using "store_conditional" on the
>two stores will result in similar problems.
>
>The only way to make a multi-unit RMW operation work is if other
>processes are /blocked/ from breaking in during the actual write
>sequence. Reads and the calculation can be re-retried, but not the
>writes - they must be made an unbreakable sequence. And that, in
>general, means a lock and OS support to ensure that the locking process
>gets to finish.
>
>
>The gcc implementation of atomic operations (larger than can be handled
>with a single instruction) uses simple user-space spin locks (the lock
>can be accessed atomically - with an LL/SC sequence, for the ARM).
>
>If one process tries to access the atomic while another process has the
>lock, it will spin - running a busy wait loop. As long as these
>processes are running on different cores, there's no problem with one
>core running a few rounds of a tight loop while another core does a
>quick load or store. Given that contention is rare and cores are often
>plentiful, this results in a very efficient atomic operation. But it
>can deadlock - a process could take the spin lock and then get
>descheduled by the OS, and other threads wanting the lock could be
>activated. If these fill up the cores (maybe you have multiple threads
>all using the same supposedly lock-free atomic structure), you are screwed.
This is a typical priority inheritance problem.
>
>And if you have only one core (like almost all microcontrollers), and
>the thread that has the lock is interrupted by an interrupt routine that
>wants to access the same atomic variable, you are /really/ screwed.
To be fair, the programmer should be aware of these issues and not
use mechanisms subject to deadlock. As noted above, the typical
solution is to disable interrupts during a critical section.
>
>It's very unlikely that you'll hit a problem, but it is possible.
Famous last words, indeed.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-11-09 11:47 -0800 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <tkh03q$a45v$3@dont-email.me> |
| In reply to | #87296 |
On 11/9/2022 12:29 AM, David Brown wrote: > On 09/11/2022 00:39, Chris M. Thomasson wrote: >> On 11/8/2022 3:29 AM, Juha Nieminen wrote: >>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote: >>>> You keep on adding features to the language which have unintuitive >>>> syntax and odd rules, and >>>> don't do much to increase the number of programs you can write >>>> quickly. So what happens? >>>> Theres not much motivation to learn these features until forced to >>>> do so. So codebases tend >>>> to be mainly legacy, and C++ programmers' skills fall behind. >>> >>> For the longest time I quite strongly disagreed with the claim that >>> C++ is >>> becoming too big and too complicated. >>> >>> However, C++20 has eroded this conviction of mine somewhat. C++23 is >>> eroding >>> it even more. >>> >>> C++11 felt like a big bunch of features that the language was in dire >>> need >>> of, and genuinely made programming easier. C++14 and C++17 fixed and >>> patched >>> many of the minor problems and defects that turned out to exist in >>> C++11, >>> so C++17 felt like "what C++11 should have been in the first place". > > I agree with that. > > A challenge for C++ is that even when a new and better feature is added, > the older and clumsier methods still have to be supported. This also > means that syntax can be awkward because it can't conflict with existing > syntax, and the details get more complex all the time. > >> >> I was really excited and happy when C++ finally made atomics and >> membars part of the actual standard, C++11 iirc. Before that, I would >> have to code these things up in assembly language. >> > > Standard atomics would be great if they worked for my targets. The gcc > implementations (and I haven't seen any others) for "advanced" use > (read-modify-write, or sizes larger than a standard register) is > completely broken for single-core systems, and even on multi-core > systems it is limited if you use thread priorities. The trouble with > them is that no one has addressed the elephant in the room - in general, > you need OS support and locks to implement large atomics. Are you referring to double-width compare-and-swap (DWCAS)? C++ should be able to handle it directly using the processors instruction set. Say C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a double word. Double word in the sense that they are two _contiguous_ words. In other words, a lock-free CAS of a double word on a 64 bit x64 should use CMPXCHG16B.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-11-09 20:43 +0000 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <diUaL.5684$BaF9.4221@fx39.iad> |
| In reply to | #87303 |
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >On 11/9/2022 12:29 AM, David Brown wrote: >> On 09/11/2022 00:39, Chris M. Thomasson wrote: >>> On 11/8/2022 3:29 AM, Juha Nieminen wrote: >>>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote: >>>>> You keep on adding features to the language which have unintuitive >>>>> syntax and odd rules, and >>>>> don't do much to increase the number of programs you can write >>>>> quickly. So what happens? >>>>> Theres not much motivation to learn these features until forced to >>>>> do so. So codebases tend >>>>> to be mainly legacy, and C++ programmers' skills fall behind. >>>> >>>> For the longest time I quite strongly disagreed with the claim that >>>> C++ is >>>> becoming too big and too complicated. >>>> >>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is >>>> eroding >>>> it even more. >>>> >>>> C++11 felt like a big bunch of features that the language was in dire >>>> need >>>> of, and genuinely made programming easier. C++14 and C++17 fixed and >>>> patched >>>> many of the minor problems and defects that turned out to exist in >>>> C++11, >>>> so C++17 felt like "what C++11 should have been in the first place". >> >> I agree with that. >> >> A challenge for C++ is that even when a new and better feature is added, >> the older and clumsier methods still have to be supported. This also >> means that syntax can be awkward because it can't conflict with existing >> syntax, and the details get more complex all the time. >> >>> >>> I was really excited and happy when C++ finally made atomics and >>> membars part of the actual standard, C++11 iirc. Before that, I would >>> have to code these things up in assembly language. >>> >> >> Standard atomics would be great if they worked for my targets. The gcc >> implementations (and I haven't seen any others) for "advanced" use >> (read-modify-write, or sizes larger than a standard register) is >> completely broken for single-core systems, and even on multi-core >> systems it is limited if you use thread priorities. The trouble with >> them is that no one has addressed the elephant in the room - in general, >> you need OS support and locks to implement large atomics. > >Are you referring to double-width compare-and-swap (DWCAS)? C++ should >be able to handle it directly using the processors instruction set. Say >C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a >double word. Double word in the sense that they are two _contiguous_ >words. In other words, a lock-free CAS of a double word on a 64 bit x64 >should use CMPXCHG16B. David works with low-end embedded processors, as I understand it, with limited and/or restricted instruction sets.
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2022-11-10 09:52 +0100 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <tkie4g$h5am$1@dont-email.me> |
| In reply to | #87304 |
On 09/11/2022 21:43, Scott Lurndal wrote: > "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >> On 11/9/2022 12:29 AM, David Brown wrote: >>> On 09/11/2022 00:39, Chris M. Thomasson wrote: >>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote: >>>>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote: >>>>>> You keep on adding features to the language which have unintuitive >>>>>> syntax and odd rules, and >>>>>> don't do much to increase the number of programs you can write >>>>>> quickly. So what happens? >>>>>> Theres not much motivation to learn these features until forced to >>>>>> do so. So codebases tend >>>>>> to be mainly legacy, and C++ programmers' skills fall behind. >>>>> >>>>> For the longest time I quite strongly disagreed with the claim that >>>>> C++ is >>>>> becoming too big and too complicated. >>>>> >>>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is >>>>> eroding >>>>> it even more. >>>>> >>>>> C++11 felt like a big bunch of features that the language was in dire >>>>> need >>>>> of, and genuinely made programming easier. C++14 and C++17 fixed and >>>>> patched >>>>> many of the minor problems and defects that turned out to exist in >>>>> C++11, >>>>> so C++17 felt like "what C++11 should have been in the first place". >>> >>> I agree with that. >>> >>> A challenge for C++ is that even when a new and better feature is added, >>> the older and clumsier methods still have to be supported. This also >>> means that syntax can be awkward because it can't conflict with existing >>> syntax, and the details get more complex all the time. >>> >>>> >>>> I was really excited and happy when C++ finally made atomics and >>>> membars part of the actual standard, C++11 iirc. Before that, I would >>>> have to code these things up in assembly language. >>>> >>> >>> Standard atomics would be great if they worked for my targets. The gcc >>> implementations (and I haven't seen any others) for "advanced" use >>> (read-modify-write, or sizes larger than a standard register) is >>> completely broken for single-core systems, and even on multi-core >>> systems it is limited if you use thread priorities. The trouble with >>> them is that no one has addressed the elephant in the room - in general, >>> you need OS support and locks to implement large atomics. >> >> Are you referring to double-width compare-and-swap (DWCAS)? C++ should >> be able to handle it directly using the processors instruction set. Say >> C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a >> double word. Double word in the sense that they are two _contiguous_ >> words. In other words, a lock-free CAS of a double word on a 64 bit x64 >> should use CMPXCHG16B. > > David works with low-end embedded processors, as I understand it, with > limited and/or restricted instruction sets. > Yes. But the principle is the same on bigger systems too. If your processor can do a single-instruction 64-bit write, you see the problems for atomics bigger than 64-bit. If it can handle 128-bit writes, you see the problems for atomics bigger than 128-bit. Obviously the need for big atomics is much lower than the need for smaller ones. Once you have a DCAS, or LL/SC, you have covered most needs. However, these alone will not give you read-modify-write operations on anything bigger than you can handle with a single read (or more importantly, with a single unbreakable write operation). Anything where the implementation is "use small atomics to get a spin lock, then do the work" is /broken/. It has a small but non-zero chance of failing in general use on big multi-core systems. On small single-core systems, it is guaranteed broken from the outset. The C++ (and C) language, standard library, common toolchains and library implementations give the programmer the impression that they can make atomics as they like. You can write : std::atomic<std::array<int, 32>> xs; and it looks like you have a big atomic object. But it will not work - you cannot rely on it. It will /seem/ to work in all your testing, because the chance of hitting a problem is small - but it can fail at any time. The atomics that the programmer can use should either be absolutely correct, guaranteed by design in all circumstances, or they should not be compile-time errors when you try to use atomics that are too big, or where the operations are too complex, for the implementation to guarantee. It would be even better for the implementation to handle these correctly. That means OS support for /real/ locks, not fake sort-of-works userland spin locks, but futexes or something like that for big systems, and interrupt disabling for single-core microcontrollers. (Dual-core microcontrollers are an extra complication.) Common library implementations could rely on an extra library or code for their "lock" and "unlock" calls - if they are not provided, you at least have a link error.
[toc] | [prev] | [next] | [standalone]
| From | Michael S <already5chosen@yahoo.com> |
|---|---|
| Date | 2022-11-10 03:07 -0800 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <1c041b9f-9424-41a2-b536-5f21a32addd8n@googlegroups.com> |
| In reply to | #87307 |
On Thursday, November 10, 2022 at 10:52:49 AM UTC+2, David Brown wrote: > On 09/11/2022 21:43, Scott Lurndal wrote: > > "Chris M. Thomasson" <chris.m.t...@gmail.com> writes: > >> On 11/9/2022 12:29 AM, David Brown wrote: > >>> On 09/11/2022 00:39, Chris M. Thomasson wrote: > >>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote: > >>>>> Malcolm McLean <malcolm.ar...@gmail.com> wrote: > >>>>>> You keep on adding features to the language which have unintuitive > >>>>>> syntax and odd rules, and > >>>>>> don't do much to increase the number of programs you can write > >>>>>> quickly. So what happens? > >>>>>> Theres not much motivation to learn these features until forced to > >>>>>> do so. So codebases tend > >>>>>> to be mainly legacy, and C++ programmers' skills fall behind. > >>>>> > >>>>> For the longest time I quite strongly disagreed with the claim that > >>>>> C++ is > >>>>> becoming too big and too complicated. > >>>>> > >>>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is > >>>>> eroding > >>>>> it even more. > >>>>> > >>>>> C++11 felt like a big bunch of features that the language was in dire > >>>>> need > >>>>> of, and genuinely made programming easier. C++14 and C++17 fixed and > >>>>> patched > >>>>> many of the minor problems and defects that turned out to exist in > >>>>> C++11, > >>>>> so C++17 felt like "what C++11 should have been in the first place". > >>> > >>> I agree with that. > >>> > >>> A challenge for C++ is that even when a new and better feature is added, > >>> the older and clumsier methods still have to be supported. This also > >>> means that syntax can be awkward because it can't conflict with existing > >>> syntax, and the details get more complex all the time. > >>> > >>>> > >>>> I was really excited and happy when C++ finally made atomics and > >>>> membars part of the actual standard, C++11 iirc. Before that, I would > >>>> have to code these things up in assembly language. > >>>> > >>> > >>> Standard atomics would be great if they worked for my targets. The gcc > >>> implementations (and I haven't seen any others) for "advanced" use > >>> (read-modify-write, or sizes larger than a standard register) is > >>> completely broken for single-core systems, and even on multi-core > >>> systems it is limited if you use thread priorities. The trouble with > >>> them is that no one has addressed the elephant in the room - in general, > >>> you need OS support and locks to implement large atomics. > >> > >> Are you referring to double-width compare-and-swap (DWCAS)? C++ should > >> be able to handle it directly using the processors instruction set. Say > >> C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a > >> double word. Double word in the sense that they are two _contiguous_ > >> words. In other words, a lock-free CAS of a double word on a 64 bit x64 > >> should use CMPXCHG16B. > > > > David works with low-end embedded processors, as I understand it, with > > limited and/or restricted instruction sets. > > > Yes. > > But the principle is the same on bigger systems too. If your processor > can do a single-instruction 64-bit write, you see the problems for > atomics bigger than 64-bit. If it can handle 128-bit writes, you see > the problems for atomics bigger than 128-bit. > > Obviously the need for big atomics is much lower than the need for > smaller ones. Once you have a DCAS, or LL/SC, you have covered most needs. > > However, these alone will not give you read-modify-write operations on > anything bigger than you can handle with a single read (or more > importantly, with a single unbreakable write operation). Anything where > the implementation is "use small atomics to get a spin lock, then do the > work" is /broken/. It has a small but non-zero chance of failing in > general use on big multi-core systems. On small single-core systems, it > is guaranteed broken from the outset. > > The C++ (and C) language, standard library, common toolchains and > library implementations give the programmer the impression that they can > make atomics as they like. You can write : > > std::atomic<std::array<int, 32>> xs; > > and it looks like you have a big atomic object. But it will not work - > you cannot rely on it. It will /seem/ to work in all your testing, > because the chance of hitting a problem is small - but it can fail at > any time. > > The atomics that the programmer can use should either be absolutely > correct, guaranteed by design in all circumstances, or they should not > be compile-time errors when you try to use atomics that are too big, or > where the operations are too complex, for the implementation to guarantee. > Yes, failing in compile time is the most reasonable. > It would be even better for the implementation to handle these > correctly. That means OS support for /real/ locks, not fake > sort-of-works userland spin locks, but futexes or something like that > for big systems, and interrupt disabling for single-core > microcontrollers. (Dual-core microcontrollers are an extra > complication.) Common library implementations could rely on an extra > library or code for their "lock" and "unlock" calls - if they are not > provided, you at least have a link error. It certainly would be against the spirit of 'C'. I'm not sure about about relationship to spirit of C++. Also I'm not sure that C++ has spirit.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-11-10 11:53 -0800 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <tkjkqu$k10j$3@dont-email.me> |
| In reply to | #87307 |
On 11/10/2022 12:52 AM, David Brown wrote: > On 09/11/2022 21:43, Scott Lurndal wrote: >> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >>> On 11/9/2022 12:29 AM, David Brown wrote: >>>> On 09/11/2022 00:39, Chris M. Thomasson wrote: >>>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote: >>>>>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote: >>>>>>> You keep on adding features to the language which have unintuitive >>>>>>> syntax and odd rules, and >>>>>>> don't do much to increase the number of programs you can write >>>>>>> quickly. So what happens? >>>>>>> Theres not much motivation to learn these features until forced to >>>>>>> do so. So codebases tend >>>>>>> to be mainly legacy, and C++ programmers' skills fall behind. >>>>>> >>>>>> For the longest time I quite strongly disagreed with the claim that >>>>>> C++ is >>>>>> becoming too big and too complicated. >>>>>> >>>>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is >>>>>> eroding >>>>>> it even more. >>>>>> >>>>>> C++11 felt like a big bunch of features that the language was in dire >>>>>> need >>>>>> of, and genuinely made programming easier. C++14 and C++17 fixed and >>>>>> patched >>>>>> many of the minor problems and defects that turned out to exist in >>>>>> C++11, >>>>>> so C++17 felt like "what C++11 should have been in the first place". >>>> >>>> I agree with that. >>>> >>>> A challenge for C++ is that even when a new and better feature is >>>> added, >>>> the older and clumsier methods still have to be supported. This also >>>> means that syntax can be awkward because it can't conflict with >>>> existing >>>> syntax, and the details get more complex all the time. >>>> >>>>> >>>>> I was really excited and happy when C++ finally made atomics and >>>>> membars part of the actual standard, C++11 iirc. Before that, I would >>>>> have to code these things up in assembly language. >>>>> >>>> >>>> Standard atomics would be great if they worked for my targets. The gcc >>>> implementations (and I haven't seen any others) for "advanced" use >>>> (read-modify-write, or sizes larger than a standard register) is >>>> completely broken for single-core systems, and even on multi-core >>>> systems it is limited if you use thread priorities. The trouble with >>>> them is that no one has addressed the elephant in the room - in >>>> general, >>>> you need OS support and locks to implement large atomics. >>> >>> Are you referring to double-width compare-and-swap (DWCAS)? C++ should >>> be able to handle it directly using the processors instruction set. Say >>> C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a >>> double word. Double word in the sense that they are two _contiguous_ >>> words. In other words, a lock-free CAS of a double word on a 64 bit x64 >>> should use CMPXCHG16B. >> >> David works with low-end embedded processors, as I understand it, with >> limited and/or restricted instruction sets. >> > > Yes. > > But the principle is the same on bigger systems too. If your processor > can do a single-instruction 64-bit write, you see the problems for > atomics bigger than 64-bit. If it can handle 128-bit writes, you see > the problems for atomics bigger than 128-bit. > > Obviously the need for big atomics is much lower than the need for > smaller ones. Once you have a DCAS, or LL/SC, you have covered most needs. > > However, these alone will not give you read-modify-write operations on > anything bigger than you can handle with a single read (or more > importantly, with a single unbreakable write operation). Anything where > the implementation is "use small atomics to get a spin lock, then do the > work" is /broken/. It has a small but non-zero chance of failing in > general use on big multi-core systems. On small single-core systems, it > is guaranteed broken from the outset. > > The C++ (and C) language, standard library, common toolchains and > library implementations give the programmer the impression that they can > make atomics as they like. You can write : > > std::atomic<std::array<int, 32>> xs; > > and it looks like you have a big atomic object. But it will not work - > you cannot rely on it. It will /seem/ to work in all your testing, > because the chance of hitting a problem is small - but it can fail at > any time. A rule of thumb... Imvho, always check the result of: https://en.cppreference.com/w/cpp/atomic/atomic/is_lock_free Just to be, sure... ;^) Fwiw, DWCAS is very different than DCAS. The latter can work on two non-contiguous words. The former only works with contiguous words. A main reason for DWCAS to exist in the first place is to be able to handle a lock-free stack. A pointer and an version count to combat the ABA problem. Although, there are many other interesting uses for DWCAS... https://groups.google.com/g/comp.lang.c++/c/nUDtke-H1io/m/g87spoMUCgAJ > The atomics that the programmer can use should either be absolutely > correct, guaranteed by design in all circumstances, or they should not > be compile-time errors when you try to use atomics that are too big, or > where the operations are too complex, for the implementation to guarantee. > > It would be even better for the implementation to handle these > correctly. That means OS support for /real/ locks, not fake > sort-of-works userland spin locks, but futexes or something like that > for big systems, and interrupt disabling for single-core > microcontrollers. (Dual-core microcontrollers are an extra > complication.) Common library implementations could rely on an extra > library or code for their "lock" and "unlock" calls - if they are not > provided, you at least have a link error. > If the result of is_lock_free is not true, then you should really think about digging into how the locking is actually implemented. Hash based address locking is one simple way to do it. Fwiw, I created one called multi-mutex: https://groups.google.com/g/comp.lang.c++/c/sV4WC_cBb9Q/m/Ti8LFyH4CgAJ
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2022-11-10 21:44 +0100 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <tkjnqt$k9i0$1@dont-email.me> |
| In reply to | #87317 |
On 10/11/2022 20:53, Chris M. Thomasson wrote: > On 11/10/2022 12:52 AM, David Brown wrote: >> On 09/11/2022 21:43, Scott Lurndal wrote: >>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >>>> On 11/9/2022 12:29 AM, David Brown wrote: >>>>> On 09/11/2022 00:39, Chris M. Thomasson wrote: >>>>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote: >>>>>>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote: >>>>>>>> You keep on adding features to the language which have unintuitive >>>>>>>> syntax and odd rules, and >>>>>>>> don't do much to increase the number of programs you can write >>>>>>>> quickly. So what happens? >>>>>>>> Theres not much motivation to learn these features until forced to >>>>>>>> do so. So codebases tend >>>>>>>> to be mainly legacy, and C++ programmers' skills fall behind. >>>>>>> >>>>>>> For the longest time I quite strongly disagreed with the claim that >>>>>>> C++ is >>>>>>> becoming too big and too complicated. >>>>>>> >>>>>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is >>>>>>> eroding >>>>>>> it even more. >>>>>>> >>>>>>> C++11 felt like a big bunch of features that the language was in >>>>>>> dire >>>>>>> need >>>>>>> of, and genuinely made programming easier. C++14 and C++17 fixed and >>>>>>> patched >>>>>>> many of the minor problems and defects that turned out to exist in >>>>>>> C++11, >>>>>>> so C++17 felt like "what C++11 should have been in the first place". >>>>> >>>>> I agree with that. >>>>> >>>>> A challenge for C++ is that even when a new and better feature is >>>>> added, >>>>> the older and clumsier methods still have to be supported. This also >>>>> means that syntax can be awkward because it can't conflict with >>>>> existing >>>>> syntax, and the details get more complex all the time. >>>>> >>>>>> >>>>>> I was really excited and happy when C++ finally made atomics and >>>>>> membars part of the actual standard, C++11 iirc. Before that, I would >>>>>> have to code these things up in assembly language. >>>>>> >>>>> >>>>> Standard atomics would be great if they worked for my targets. The >>>>> gcc >>>>> implementations (and I haven't seen any others) for "advanced" use >>>>> (read-modify-write, or sizes larger than a standard register) is >>>>> completely broken for single-core systems, and even on multi-core >>>>> systems it is limited if you use thread priorities. The trouble with >>>>> them is that no one has addressed the elephant in the room - in >>>>> general, >>>>> you need OS support and locks to implement large atomics. >>>> >>>> Are you referring to double-width compare-and-swap (DWCAS)? C++ should >>>> be able to handle it directly using the processors instruction set. Say >>>> C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a >>>> double word. Double word in the sense that they are two _contiguous_ >>>> words. In other words, a lock-free CAS of a double word on a 64 bit x64 >>>> should use CMPXCHG16B. >>> >>> David works with low-end embedded processors, as I understand it, with >>> limited and/or restricted instruction sets. >>> >> >> Yes. >> >> But the principle is the same on bigger systems too. If your >> processor can do a single-instruction 64-bit write, you see the >> problems for atomics bigger than 64-bit. If it can handle 128-bit >> writes, you see the problems for atomics bigger than 128-bit. >> >> Obviously the need for big atomics is much lower than the need for >> smaller ones. Once you have a DCAS, or LL/SC, you have covered most >> needs. >> >> However, these alone will not give you read-modify-write operations on >> anything bigger than you can handle with a single read (or more >> importantly, with a single unbreakable write operation). Anything >> where the implementation is "use small atomics to get a spin lock, >> then do the work" is /broken/. It has a small but non-zero chance of >> failing in general use on big multi-core systems. On small >> single-core systems, it is guaranteed broken from the outset. >> >> The C++ (and C) language, standard library, common toolchains and >> library implementations give the programmer the impression that they >> can make atomics as they like. You can write : >> >> std::atomic<std::array<int, 32>> xs; >> >> and it looks like you have a big atomic object. But it will not work >> - you cannot rely on it. It will /seem/ to work in all your testing, >> because the chance of hitting a problem is small - but it can fail at >> any time. > > A rule of thumb... Imvho, always check the result of: > > https://en.cppreference.com/w/cpp/atomic/atomic/is_lock_free And it if is not, what is the point in allowing it if the locks don't work? > > Just to be, sure... ;^) Fwiw, DWCAS is very different than DCAS. The > latter can work on two non-contiguous words. The former only works with > contiguous words. A main reason for DWCAS to exist in the first place is > to be able to handle a lock-free stack. A pointer and an version count > to combat the ABA problem. Although, there are many other interesting > uses for DWCAS... > > https://groups.google.com/g/comp.lang.c++/c/nUDtke-H1io/m/g87spoMUCgAJ > > >> The atomics that the programmer can use should either be absolutely >> correct, guaranteed by design in all circumstances, or they should not >> be compile-time errors when you try to use atomics that are too big, >> or where the operations are too complex, for the implementation to >> guarantee. >> >> It would be even better for the implementation to handle these >> correctly. That means OS support for /real/ locks, not fake >> sort-of-works userland spin locks, but futexes or something like that >> for big systems, and interrupt disabling for single-core >> microcontrollers. (Dual-core microcontrollers are an extra >> complication.) Common library implementations could rely on an extra >> library or code for their "lock" and "unlock" calls - if they are not >> provided, you at least have a link error. >> > > If the result of is_lock_free is not true, then you should really think > about digging into how the locking is actually implemented. Hash based > address locking is one simple way to do it. Fwiw, I created one called > multi-mutex: > > https://groups.google.com/g/comp.lang.c++/c/sV4WC_cBb9Q/m/Ti8LFyH4CgAJ Hash-based arrays of locks are as bad as a single lock for all atomics, in that it does not work unless it is a proper OS lock. The larger your array of locks, the lower your chances of problems, but it all comes down to one thing - are your locks safe or not?
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-11-10 12:53 -0800 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <tkjocq$kaut$1@dont-email.me> |
| In reply to | #87318 |
On 11/10/2022 12:44 PM, David Brown wrote: > On 10/11/2022 20:53, Chris M. Thomasson wrote: >> On 11/10/2022 12:52 AM, David Brown wrote: >>> On 09/11/2022 21:43, Scott Lurndal wrote: >>>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >>>>> On 11/9/2022 12:29 AM, David Brown wrote: >>>>>> On 09/11/2022 00:39, Chris M. Thomasson wrote: >>>>>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote: >>>>>>>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote: >>>>>>>>> You keep on adding features to the language which have unintuitive >>>>>>>>> syntax and odd rules, and >>>>>>>>> don't do much to increase the number of programs you can write >>>>>>>>> quickly. So what happens? >>>>>>>>> Theres not much motivation to learn these features until forced to >>>>>>>>> do so. So codebases tend >>>>>>>>> to be mainly legacy, and C++ programmers' skills fall behind. >>>>>>>> >>>>>>>> For the longest time I quite strongly disagreed with the claim that >>>>>>>> C++ is >>>>>>>> becoming too big and too complicated. >>>>>>>> >>>>>>>> However, C++20 has eroded this conviction of mine somewhat. >>>>>>>> C++23 is >>>>>>>> eroding >>>>>>>> it even more. >>>>>>>> >>>>>>>> C++11 felt like a big bunch of features that the language was in >>>>>>>> dire >>>>>>>> need >>>>>>>> of, and genuinely made programming easier. C++14 and C++17 fixed >>>>>>>> and >>>>>>>> patched >>>>>>>> many of the minor problems and defects that turned out to exist in >>>>>>>> C++11, >>>>>>>> so C++17 felt like "what C++11 should have been in the first >>>>>>>> place". >>>>>> >>>>>> I agree with that. >>>>>> >>>>>> A challenge for C++ is that even when a new and better feature is >>>>>> added, >>>>>> the older and clumsier methods still have to be supported. This also >>>>>> means that syntax can be awkward because it can't conflict with >>>>>> existing >>>>>> syntax, and the details get more complex all the time. >>>>>> >>>>>>> >>>>>>> I was really excited and happy when C++ finally made atomics and >>>>>>> membars part of the actual standard, C++11 iirc. Before that, I >>>>>>> would >>>>>>> have to code these things up in assembly language. >>>>>>> >>>>>> >>>>>> Standard atomics would be great if they worked for my targets. >>>>>> The gcc >>>>>> implementations (and I haven't seen any others) for "advanced" use >>>>>> (read-modify-write, or sizes larger than a standard register) is >>>>>> completely broken for single-core systems, and even on multi-core >>>>>> systems it is limited if you use thread priorities. The trouble with >>>>>> them is that no one has addressed the elephant in the room - in >>>>>> general, >>>>>> you need OS support and locks to implement large atomics. >>>>> >>>>> Are you referring to double-width compare-and-swap (DWCAS)? C++ should >>>>> be able to handle it directly using the processors instruction set. >>>>> Say >>>>> C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used >>>>> for a >>>>> double word. Double word in the sense that they are two _contiguous_ >>>>> words. In other words, a lock-free CAS of a double word on a 64 bit >>>>> x64 >>>>> should use CMPXCHG16B. >>>> >>>> David works with low-end embedded processors, as I understand it, with >>>> limited and/or restricted instruction sets. >>>> >>> >>> Yes. >>> >>> But the principle is the same on bigger systems too. If your >>> processor can do a single-instruction 64-bit write, you see the >>> problems for atomics bigger than 64-bit. If it can handle 128-bit >>> writes, you see the problems for atomics bigger than 128-bit. >>> >>> Obviously the need for big atomics is much lower than the need for >>> smaller ones. Once you have a DCAS, or LL/SC, you have covered most >>> needs. >>> >>> However, these alone will not give you read-modify-write operations >>> on anything bigger than you can handle with a single read (or more >>> importantly, with a single unbreakable write operation). Anything >>> where the implementation is "use small atomics to get a spin lock, >>> then do the work" is /broken/. It has a small but non-zero chance of >>> failing in general use on big multi-core systems. On small >>> single-core systems, it is guaranteed broken from the outset. >>> >>> The C++ (and C) language, standard library, common toolchains and >>> library implementations give the programmer the impression that they >>> can make atomics as they like. You can write : >>> >>> std::atomic<std::array<int, 32>> xs; >>> >>> and it looks like you have a big atomic object. But it will not work >>> - you cannot rely on it. It will /seem/ to work in all your testing, >>> because the chance of hitting a problem is small - but it can fail at >>> any time. >> >> A rule of thumb... Imvho, always check the result of: >> >> https://en.cppreference.com/w/cpp/atomic/atomic/is_lock_free > > And it if is not, what is the point in allowing it if the locks don't work? > >> >> Just to be, sure... ;^) Fwiw, DWCAS is very different than DCAS. The >> latter can work on two non-contiguous words. The former only works >> with contiguous words. A main reason for DWCAS to exist in the first >> place is to be able to handle a lock-free stack. A pointer and an >> version count to combat the ABA problem. Although, there are many >> other interesting uses for DWCAS... >> >> https://groups.google.com/g/comp.lang.c++/c/nUDtke-H1io/m/g87spoMUCgAJ >> >> >>> The atomics that the programmer can use should either be absolutely >>> correct, guaranteed by design in all circumstances, or they should >>> not be compile-time errors when you try to use atomics that are too >>> big, or where the operations are too complex, for the implementation >>> to guarantee. >>> >>> It would be even better for the implementation to handle these >>> correctly. That means OS support for /real/ locks, not fake >>> sort-of-works userland spin locks, but futexes or something like that >>> for big systems, and interrupt disabling for single-core >>> microcontrollers. (Dual-core microcontrollers are an extra >>> complication.) Common library implementations could rely on an extra >>> library or code for their "lock" and "unlock" calls - if they are not >>> provided, you at least have a link error. >>> >> >> If the result of is_lock_free is not true, then you should really >> think about digging into how the locking is actually implemented. Hash >> based address locking is one simple way to do it. Fwiw, I created one >> called multi-mutex: >> >> https://groups.google.com/g/comp.lang.c++/c/sV4WC_cBb9Q/m/Ti8LFyH4CgAJ > > Hash-based arrays of locks are as bad as a single lock for all atomics, > in that it does not work unless it is a proper OS lock. The larger your > array of locks, the lower your chances of problems, but it all comes > down to one thing - are your locks safe or not? Well, my multi-mutex uses std::mutex as elements of its vector of locks. It hashes an address into said table. So, it's only as good as the implementation of std::mutex... std::vector<std::mutex> m_locks; Fair enough?
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-11-10 21:20 +0000 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <AWdbL.84469$2Rs3.11964@fx12.iad> |
| In reply to | #87319 |
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >On 11/10/2022 12:44 PM, David Brown wrote: >> Hash-based arrays of locks are as bad as a single lock for all atomics, >> in that it does not work unless it is a proper OS lock. The larger your >> array of locks, the lower your chances of problems, but it all comes >> down to one thing - are your locks safe or not? > >Well, my multi-mutex uses std::mutex as elements of its vector of locks. >It hashes an address into said table. So, it's only as good as the >implementation of std::mutex... > >std::vector<std::mutex> m_locks; > >Fair enough? > I'd point out that using a vector is sub-optimal if the various elements of the vector are accessed from different cores due to false sharing....
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-11-10 13:23 -0800 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <tkjq48$kflb$1@dont-email.me> |
| In reply to | #87320 |
On 11/10/2022 1:20 PM, Scott Lurndal wrote: > "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >> On 11/10/2022 12:44 PM, David Brown wrote: > >>> Hash-based arrays of locks are as bad as a single lock for all atomics, >>> in that it does not work unless it is a proper OS lock. The larger your >>> array of locks, the lower your chances of problems, but it all comes >>> down to one thing - are your locks safe or not? >> >> Well, my multi-mutex uses std::mutex as elements of its vector of locks. >> It hashes an address into said table. So, it's only as good as the >> implementation of std::mutex... >> >> std::vector<std::mutex> m_locks; >> >> Fair enough? >> > > I'd point out that using a vector is sub-optimal if the various > elements of the vector are accessed from different cores due to > false sharing.... Well, the mutex state should really be isolated on their own cache lines. Padding and alignment on a l2 cacheline boundary. Iirc, this can be done in modern c++. Even aligning elements of a vector.
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2022-11-11 08:10 +0100 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <tkksi3$pom6$1@dont-email.me> |
| In reply to | #87319 |
On 10/11/2022 21:53, Chris M. Thomasson wrote: > On 11/10/2022 12:44 PM, David Brown wrote: >> On 10/11/2022 20:53, Chris M. Thomasson wrote: >>> On 11/10/2022 12:52 AM, David Brown wrote: >>> If the result of is_lock_free is not true, then you should really >>> think about digging into how the locking is actually implemented. >>> Hash based address locking is one simple way to do it. Fwiw, I >>> created one called multi-mutex: >>> >>> https://groups.google.com/g/comp.lang.c++/c/sV4WC_cBb9Q/m/Ti8LFyH4CgAJ >> >> Hash-based arrays of locks are as bad as a single lock for all >> atomics, in that it does not work unless it is a proper OS lock. The >> larger your array of locks, the lower your chances of problems, but it >> all comes down to one thing - are your locks safe or not? > > Well, my multi-mutex uses std::mutex as elements of its vector of locks. > It hashes an address into said table. So, it's only as good as the > implementation of std::mutex... > > std::vector<std::mutex> m_locks; > > Fair enough? > As long as std::mutex is a wrapper for a real OS mutex, that will be fine (after padding for cache line sizes). The problem with the atomics library in gcc is not the array of locks or the hashing on address, but that it doesn't use /real/ locks.
[toc] | [prev] | [next] | [standalone]
| From | Öö Tiib <ootiib@hot.ee> |
|---|---|
| Date | 2022-11-11 01:45 -0800 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <4990e23a-9da7-4934-8e7c-40a0255b115dn@googlegroups.com> |
| In reply to | #87322 |
On Friday, 11 November 2022 at 09:11:17 UTC+2, David Brown wrote: > On 10/11/2022 21:53, Chris M. Thomasson wrote: > > On 11/10/2022 12:44 PM, David Brown wrote: > >> On 10/11/2022 20:53, Chris M. Thomasson wrote: > >>> On 11/10/2022 12:52 AM, David Brown wrote: > > >>> If the result of is_lock_free is not true, then you should really > >>> think about digging into how the locking is actually implemented. > >>> Hash based address locking is one simple way to do it. Fwiw, I > >>> created one called multi-mutex: > >>> > >>> https://groups.google.com/g/comp.lang.c++/c/sV4WC_cBb9Q/m/Ti8LFyH4CgAJ > >> > >> Hash-based arrays of locks are as bad as a single lock for all > >> atomics, in that it does not work unless it is a proper OS lock. The > >> larger your array of locks, the lower your chances of problems, but it > >> all comes down to one thing - are your locks safe or not? > > > > Well, my multi-mutex uses std::mutex as elements of its vector of locks. > > It hashes an address into said table. So, it's only as good as the > > implementation of std::mutex... > > > > std::vector<std::mutex> m_locks; > > > > Fair enough? > > > As long as std::mutex is a wrapper for a real OS mutex, that will be > fine (after padding for cache line sizes). > > The problem with the atomics library in gcc is not the array of locks or > the hashing on address, but that it doesn't use /real/ locks. Yes. Libraries look often like sabotaged and so who needs quality and portability has to use conditional compiling for what they use std::atomic<Something> and for what manually mutex-protected Something instance. Purpose defeated.
[toc] | [prev] | [next] | [standalone]
| From | Michael S <already5chosen@yahoo.com> |
|---|---|
| Date | 2022-11-10 03:14 -0800 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <eac1413e-eec0-4530-9ea4-8cb5f97ecd6fn@googlegroups.com> |
| In reply to | #87304 |
On Wednesday, November 9, 2022 at 10:44:10 PM UTC+2, Scott Lurndal wrote: > "Chris M. Thomasson" <chris.m.t...@gmail.com> writes: > >On 11/9/2022 12:29 AM, David Brown wrote: > >> On 09/11/2022 00:39, Chris M. Thomasson wrote: > >>> On 11/8/2022 3:29 AM, Juha Nieminen wrote: > >>>> Malcolm McLean <malcolm.ar...@gmail.com> wrote: > >>>>> You keep on adding features to the language which have unintuitive > >>>>> syntax and odd rules, and > >>>>> don't do much to increase the number of programs you can write > >>>>> quickly. So what happens? > >>>>> Theres not much motivation to learn these features until forced to > >>>>> do so. So codebases tend > >>>>> to be mainly legacy, and C++ programmers' skills fall behind. > >>>> > >>>> For the longest time I quite strongly disagreed with the claim that > >>>> C++ is > >>>> becoming too big and too complicated. > >>>> > >>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is > >>>> eroding > >>>> it even more. > >>>> > >>>> C++11 felt like a big bunch of features that the language was in dire > >>>> need > >>>> of, and genuinely made programming easier. C++14 and C++17 fixed and > >>>> patched > >>>> many of the minor problems and defects that turned out to exist in > >>>> C++11, > >>>> so C++17 felt like "what C++11 should have been in the first place". > >> > >> I agree with that. > >> > >> A challenge for C++ is that even when a new and better feature is added, > >> the older and clumsier methods still have to be supported. This also > >> means that syntax can be awkward because it can't conflict with existing > >> syntax, and the details get more complex all the time. > >> > >>> > >>> I was really excited and happy when C++ finally made atomics and > >>> membars part of the actual standard, C++11 iirc. Before that, I would > >>> have to code these things up in assembly language. > >>> > >> > >> Standard atomics would be great if they worked for my targets. The gcc > >> implementations (and I haven't seen any others) for "advanced" use > >> (read-modify-write, or sizes larger than a standard register) is > >> completely broken for single-core systems, and even on multi-core > >> systems it is limited if you use thread priorities. The trouble with > >> them is that no one has addressed the elephant in the room - in general, > >> you need OS support and locks to implement large atomics. > > > >Are you referring to double-width compare-and-swap (DWCAS)? C++ should > >be able to handle it directly using the processors instruction set. Say > >C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a > >double word. Double word in the sense that they are two _contiguous_ > >words. In other words, a lock-free CAS of a double word on a 64 bit x64 > >should use CMPXCHG16B. > David works with low-end embedded processors, as I understand it, with > limited and/or restricted instruction sets. In my book Cortex-M (except M0) is not a low end. M7 in particular is too big and too complicated not even for proverbial 99%, but for solid 100% of my microcontroller needs.
[toc] | [prev] | [next] | [standalone]
| From | Sam <sam@email-scan.com> |
|---|---|
| Date | 2022-11-09 07:59 -0500 |
| Subject | Re: ???The pool of talented C++ developers is running dry??? |
| Message-ID | <cone.1667998785.254454.147087.1004@monster.email-scan.com> |
| In reply to | #87283 |
Juha Nieminen writes: > C++20, however, doesn't feel like this anymore. It has a few new features > that genuinely help in programming, but most of it feels like just adding > features for the sake of adding them. C++23 even moreso. Or the features were specifically added to make sucky operating systems suck a little less. Specifically: co-routines. Microsoft hijacked the standardization process to push through co-routines, because real multiple execution threads on MS-Windows blows chunks, and the OS can only implement co-routines in a passable manner.
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | comp.lang.c++
csiph-web