Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c++ > #87211 > unrolled thread

“The pool of talented C++ developers is running dry”

Started byLynn McGuire <lynnmcguire5@gmail.com>
First post2022-11-03 14:30 -0500
Last post2022-11-09 07:59 -0500
Articles 17 on this page of 37 — 13 participants

Back to article view | Back to comp.lang.c++


Contents

  “The pool of talented C++ developers is running dry” Lynn McGuire <lynnmcguire5@gmail.com> - 2022-11-03 14:30 -0500
    Re: “The pool of talented C++ developers is running dry” "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-03 13:03 -0700
      Re: “The pool of talented C++ developers is running dry” "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-03 13:04 -0700
        Re: “The pool of talented C++ developers is running dry” Lynn McGuire <lynnmcguire5@gmail.com> - 2022-11-03 16:07 -0500
      Re: “The pool of talented C++ developers is running dry” Lynn McGuire <lynnmcguire5@gmail.com> - 2022-11-03 16:05 -0500
        Re: “The pool of talented C++ developers is running dry” "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-03 14:25 -0700
          Re: “The pool of talented C++ developers is running dry” Lynn McGuire <lynnmcguire5@gmail.com> - 2022-11-05 14:12 -0500
            Re: “The pool of talented C++ developers is running dry” "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-05 14:20 -0700
        Re: “The pool of talented C++ developers is running dry” Bonita Montero <Bonita.Montero@gmail.com> - 2022-11-04 09:24 +0100
    Re: “The pool of talented C++ developers is running dry” Vir Campestris <vir.campestris@invalid.invalid> - 2022-11-05 21:42 +0000
    Re: “The pool of talented C++ developers is running dry” "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-06 11:55 -0800
    Re: “The pool of talented C++ developers is running dry” Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-11-08 02:44 -0800
      Re: ???The pool of talented C++ developers is running dry??? Juha Nieminen <nospam@thanks.invalid> - 2022-11-08 11:29 +0000
        Re: ???The pool of talented C++ developers is running dry??? Öö Tiib <ootiib@hot.ee> - 2022-11-08 03:54 -0800
          Re: ???The pool of talented C++ developers is running dry??? Stuart Redmann <DerTopper@web.de> - 2022-11-08 14:26 +0100
            Re: ???The pool of talented C++ developers is running dry??? Öö Tiib <ootiib@hot.ee> - 2022-11-09 00:22 -0800
          Re: ???The pool of talented C++ developers is running dry??? Juha Nieminen <nospam@thanks.invalid> - 2022-11-08 15:16 +0000
        Re: ???The pool of talented C++ developers is running dry??? Bonita Montero <Bonita.Montero@gmail.com> - 2022-11-08 14:52 +0100
        Re: ???The pool of talented C++ developers is running dry??? Jorgen Grahn <grahn+nntp@snipabacken.se> - 2022-11-08 22:48 +0000
        Re: ???The pool of talented C++ developers is running dry??? "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-08 15:39 -0800
          Re: ???The pool of talented C++ developers is running dry??? David Brown <david.brown@hesbynett.no> - 2022-11-09 09:29 +0100
            Re: ???The pool of talented C++ developers is running dry??? scott@slp53.sl.home (Scott Lurndal) - 2022-11-09 14:39 +0000
              Re: ???The pool of talented C++ developers is running dry??? David Brown <david.brown@hesbynett.no> - 2022-11-09 17:05 +0100
                Re: ???The pool of talented C++ developers is running dry??? scott@slp53.sl.home (Scott Lurndal) - 2022-11-09 17:58 +0000
            Re: ???The pool of talented C++ developers is running dry??? "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-09 11:47 -0800
              Re: ???The pool of talented C++ developers is running dry??? scott@slp53.sl.home (Scott Lurndal) - 2022-11-09 20:43 +0000
                Re: ???The pool of talented C++ developers is running dry??? David Brown <david.brown@hesbynett.no> - 2022-11-10 09:52 +0100
                  Re: ???The pool of talented C++ developers is running dry??? Michael S <already5chosen@yahoo.com> - 2022-11-10 03:07 -0800
                  Re: ???The pool of talented C++ developers is running dry??? "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-10 11:53 -0800
                    Re: ???The pool of talented C++ developers is running dry??? David Brown <david.brown@hesbynett.no> - 2022-11-10 21:44 +0100
                      Re: ???The pool of talented C++ developers is running dry??? "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-10 12:53 -0800
                        Re: ???The pool of talented C++ developers is running dry??? scott@slp53.sl.home (Scott Lurndal) - 2022-11-10 21:20 +0000
                          Re: ???The pool of talented C++ developers is running dry??? "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-10 13:23 -0800
                        Re: ???The pool of talented C++ developers is running dry??? David Brown <david.brown@hesbynett.no> - 2022-11-11 08:10 +0100
                          Re: ???The pool of talented C++ developers is running dry??? Öö Tiib <ootiib@hot.ee> - 2022-11-11 01:45 -0800
                Re: ???The pool of talented C++ developers is running dry??? Michael S <already5chosen@yahoo.com> - 2022-11-10 03:14 -0800
        Re: ???The pool of talented C++ developers is running         dry??? Sam <sam@email-scan.com> - 2022-11-09 07:59 -0500

Page 2 of 2 — ← Prev page 1 [2]


#87296 — Re: ???The pool of talented C++ developers is running dry???

FromDavid Brown <david.brown@hesbynett.no>
Date2022-11-09 09:29 +0100
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<tkfoch$695m$1@dont-email.me>
In reply to#87294
On 09/11/2022 00:39, Chris M. Thomasson wrote:
> On 11/8/2022 3:29 AM, Juha Nieminen wrote:
>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote:
>>> You keep on adding features to the language which have unintuitive 
>>> syntax and odd rules, and
>>> don't do much to increase the number of programs you can write 
>>> quickly. So what happens?
>>> Theres not much motivation to learn these features until forced to do 
>>> so. So codebases tend
>>> to be mainly legacy, and C++ programmers' skills fall behind.
>>
>> For the longest time I quite strongly disagreed with the claim that 
>> C++ is
>> becoming too big and too complicated.
>>
>> However, C++20 has eroded this conviction of mine somewhat. C++23 is 
>> eroding
>> it even more.
>>
>> C++11 felt like a big bunch of features that the language was in dire 
>> need
>> of, and genuinely made programming easier. C++14 and C++17 fixed and 
>> patched
>> many of the minor problems and defects that turned out to exist in C++11,
>> so C++17 felt like "what C++11 should have been in the first place".

I agree with that.

A challenge for C++ is that even when a new and better feature is added, 
the older and clumsier methods still have to be supported.  This also 
means that syntax can be awkward because it can't conflict with existing 
syntax, and the details get more complex all the time.

> 
> I was really excited and happy when C++ finally made atomics and membars 
> part of the actual standard, C++11 iirc. Before that, I would have to 
> code these things up in assembly language.
> 

Standard atomics would be great if they worked for my targets.  The gcc 
implementations (and I haven't seen any others) for "advanced" use 
(read-modify-write, or sizes larger than a standard register) is 
completely broken for single-core systems, and even on multi-core 
systems it is limited if you use thread priorities.  The trouble with 
them is that no one has addressed the elephant in the room - in general, 
you need OS support and locks to implement large atomics.


[toc] | [prev] | [next] | [standalone]


#87299 — Re: ???The pool of talented C++ developers is running dry???

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-11-09 14:39 +0000
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<CYOaL.12305$eyq6.10545@fx03.iad>
In reply to#87296
David Brown <david.brown@hesbynett.no> writes:
>On 09/11/2022 00:39, Chris M. Thomasson wrote:
>> On 11/8/2022 3:29 AM, Juha Nieminen wrote:
>
>> I was really excited and happy when C++ finally made atomics and membars 
>> part of the actual standard, C++11 iirc. Before that, I would have to 
>> code these things up in assembly language.
>> 
>
>Standard atomics would be great if they worked for my targets.  The gcc 
>implementations (and I haven't seen any others) for "advanced" use 
>(read-modify-write, or sizes larger than a standard register) is 
>completely broken for single-core systems, and even on multi-core 
>systems it is limited if you use thread priorities. 

> The trouble with 
>them is that no one has addressed the elephant in the room - in general, 
>you need OS support and locks to implement large atomics.

IFF the target architecture doesn't have a comprehensive set of
atomic access instructions, perhaps.

ARMv8 LSE, for example, has individual instructions for most of the
gcc atomic intrinsics (e.g. __sync_fetch_and_add will generate a single
LDADD atomic instruction).   The instructions support the common
arithmetic operations (add, or, etc).

Before LSE, the ARMv8 implementations were built using the arm
LL/SC equivalent (load exclusive/store exclusive) instructions.

[toc] | [prev] | [next] | [standalone]


#87300 — Re: ???The pool of talented C++ developers is running dry???

FromDavid Brown <david.brown@hesbynett.no>
Date2022-11-09 17:05 +0100
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<tkgj53$8nkh$1@dont-email.me>
In reply to#87299
On 09/11/2022 15:39, Scott Lurndal wrote:
> David Brown <david.brown@hesbynett.no> writes:
>> On 09/11/2022 00:39, Chris M. Thomasson wrote:
>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote:
>>
>>> I was really excited and happy when C++ finally made atomics and membars
>>> part of the actual standard, C++11 iirc. Before that, I would have to
>>> code these things up in assembly language.
>>>
>>
>> Standard atomics would be great if they worked for my targets.  The gcc
>> implementations (and I haven't seen any others) for "advanced" use
>> (read-modify-write, or sizes larger than a standard register) is
>> completely broken for single-core systems, and even on multi-core
>> systems it is limited if you use thread priorities.
> 
>> The trouble with
>> them is that no one has addressed the elephant in the room - in general,
>> you need OS support and locks to implement large atomics.
> 
> IFF the target architecture doesn't have a comprehensive set of
> atomic access instructions, perhaps.
> 
> ARMv8 LSE, for example, has individual instructions for most of the
> gcc atomic intrinsics (e.g. __sync_fetch_and_add will generate a single
> LDADD atomic instruction).   The instructions support the common
> arithmetic operations (add, or, etc).
> 
> Before LSE, the ARMv8 implementations were built using the arm
> LL/SC equivalent (load exclusive/store exclusive) instructions.


You are more familiar with the details of these things than most people, 
so I hope you (or someone else) will correct me if my logic below is wrong.


There's no problem when the target has a single unbreakable instruction 
for the action.  And LL/SC are fine for atomic loads or stores of 
different sizes.

But LL/SC is not sufficient for read-modify-write sequences of a size 
larger than can be handled by a single atomic instruction.

Imagine you have a processor that can atomically read or write an 
unsigned integer type "uint".  Your sequence for "uint_inc" will be :

retry:
	load link x = *p
	x++
	if (store conditional *p = x fails) goto retry


If two processes try this, they can interleave and be started or stopped 
without trouble - the result will be an atomic increment.

Now consider a double-sized type containing two "uint" fields:

retry:
	load link x_lo = *p
	x_hi = *(p + 1)
	x_lo++
	if (!x_lo) x_hi++
	if (store conditional *p = x_lo fails) goto retry
	*(p + 1) = x_hi

If the process executing this is stopped after the first write, and a 
second process is run that calls a similar function, then the new 
process will see a half-changed value for the object resulting in a 
corrupted object.  Resumption of the first process will half-change the 
value again.  Different combinations of using "store_conditional" on the 
two stores will result in similar problems.

The only way to make a multi-unit RMW operation work is if other 
processes are /blocked/ from breaking in during the actual write 
sequence.  Reads and the calculation can be re-retried, but not the 
writes - they must be made an unbreakable sequence.  And that, in 
general, means a lock and OS support to ensure that the locking process 
gets to finish.


The gcc implementation of atomic operations (larger than can be handled 
with a single instruction) uses simple user-space spin locks (the lock 
can be accessed atomically - with an LL/SC sequence, for the ARM).

If one process tries to access the atomic while another process has the 
lock, it will spin - running a busy wait loop.  As long as these 
processes are running on different cores, there's no problem with one 
core running a few rounds of a tight loop while another core does a 
quick load or store.  Given that contention is rare and cores are often 
plentiful, this results in a very efficient atomic operation.  But it 
can deadlock - a process could take the spin lock and then get 
descheduled by the OS, and other threads wanting the lock could be 
activated.  If these fill up the cores (maybe you have multiple threads 
all using the same supposedly lock-free atomic structure), you are screwed.

And if you have only one core (like almost all microcontrollers), and 
the thread that has the lock is interrupted by an interrupt routine that 
wants to access the same atomic variable, you are /really/ screwed. 
This can happen with such simple code as a 64-bit atomic counter in an 
interrupt routine that is also accessed atomically from a background task.


It's very unlikely that you'll hit a problem, but it is possible.  To 
me, that is useless - atomics need guaranteed forward progress.  That 
means the std::atomic<> stuff needs to use OS-level locks for advanced 
cases that can't be handled directly by instructions or LL/SC sequences, 
or for a microcontroller you'd want to disable interrupts around the 
access.  The alternative is to refuse to compile the operations and only 
support atomics that are smaller or simpler.





	

[toc] | [prev] | [next] | [standalone]


#87301 — Re: ???The pool of talented C++ developers is running dry???

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-11-09 17:58 +0000
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<CTRaL.31519$NeJ8.1285@fx09.iad>
In reply to#87300
David Brown <david.brown@hesbynett.no> writes:
>On 09/11/2022 15:39, Scott Lurndal wrote:
>> David Brown <david.brown@hesbynett.no> writes:
>>> On 09/11/2022 00:39, Chris M. Thomasson wrote:
>>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote:
>>>
>>>> I was really excited and happy when C++ finally made atomics and membars
>>>> part of the actual standard, C++11 iirc. Before that, I would have to
>>>> code these things up in assembly language.
>>>>
>>>
>>> Standard atomics would be great if they worked for my targets.  The gcc
>>> implementations (and I haven't seen any others) for "advanced" use
>>> (read-modify-write, or sizes larger than a standard register) is
>>> completely broken for single-core systems, and even on multi-core
>>> systems it is limited if you use thread priorities.
>> 
>>> The trouble with
>>> them is that no one has addressed the elephant in the room - in general,
>>> you need OS support and locks to implement large atomics.
>> 
>> IFF the target architecture doesn't have a comprehensive set of
>> atomic access instructions, perhaps.
>> 
>> ARMv8 LSE, for example, has individual instructions for most of the
>> gcc atomic intrinsics (e.g. __sync_fetch_and_add will generate a single
>> LDADD atomic instruction).   The instructions support the common
>> arithmetic operations (add, or, etc).
>> 
>> Before LSE, the ARMv8 implementations were built using the arm
>> LL/SC equivalent (load exclusive/store exclusive) instructions.
>
>
>You are more familiar with the details of these things than most people, 
>so I hope you (or someone else) will correct me if my logic below is wrong.
>
>
>There's no problem when the target has a single unbreakable instruction 
>for the action.  And LL/SC are fine for atomic loads or stores of 
>different sizes.

Here's the code generated by GCC for

  q = __sync_fetch_and_add(&q, 1u);


Without LSE (atomics) support:

  401034:       885ffc60        ldaxr   w0, [x3]
  401038:       11000401        add     w1, w0, #0x1
  40103c:       8804fc61        stlxr   w4, w1, [x3]
  c01040:       35ffffa4        cbnz    w4, 401034 <main+0x34>


With LSE (atomics) support:

     12c:       b8e10001        ldaddal w1, w1, [x0]

>
>But LL/SC is not sufficient for read-modify-write sequences of a size 
>larger than can be handled by a single atomic instruction.

>
>Imagine you have a processor that can atomically read or write an 
>unsigned integer type "uint".  Your sequence for "uint_inc" will be :
>
>retry:
>	load link x = *p
>	x++
>	if (store conditional *p = x fails) goto retry
>
>
>If two processes try this, they can interleave and be started or stopped 
>without trouble - the result will be an atomic increment.
>
>Now consider a double-sized type containing two "uint" fields:
>
>retry:
>	load link x_lo = *p
>	x_hi = *(p + 1)
>	x_lo++
>	if (!x_lo) x_hi++
>	if (store conditional *p = x_lo fails) goto retry
>	*(p + 1) = x_hi

For such sequences, one uses the LL/SC as a spinlock;
acquire the spinlock, perform the non-atomic operation
and release the spinlock. On uniprocessor systems,
alternate mechanisms like disabling interrupts are the
common solution.

Although in this case, using a wider type if available is a
better option.

>
>If the process executing this is stopped after the first write, and a 
>second process is run that calls a similar function, then the new 
>process will see a half-changed value for the object resulting in a 
>corrupted object.  Resumption of the first process will half-change the 
>value again.  Different combinations of using "store_conditional" on the 
>two stores will result in similar problems.
>
>The only way to make a multi-unit RMW operation work is if other 
>processes are /blocked/ from breaking in during the actual write 
>sequence.  Reads and the calculation can be re-retried, but not the 
>writes - they must be made an unbreakable sequence.  And that, in 
>general, means a lock and OS support to ensure that the locking process 
>gets to finish.
>
>
>The gcc implementation of atomic operations (larger than can be handled 
>with a single instruction) uses simple user-space spin locks (the lock 
>can be accessed atomically - with an LL/SC sequence, for the ARM).
>
>If one process tries to access the atomic while another process has the 
>lock, it will spin - running a busy wait loop.  As long as these 
>processes are running on different cores, there's no problem with one 
>core running a few rounds of a tight loop while another core does a 
>quick load or store.  Given that contention is rare and cores are often 
>plentiful, this results in a very efficient atomic operation.  But it 
>can deadlock - a process could take the spin lock and then get 
>descheduled by the OS, and other threads wanting the lock could be 
>activated.  If these fill up the cores (maybe you have multiple threads 
>all using the same supposedly lock-free atomic structure), you are screwed.

This is a typical priority inheritance problem.

>
>And if you have only one core (like almost all microcontrollers), and 
>the thread that has the lock is interrupted by an interrupt routine that 
>wants to access the same atomic variable, you are /really/ screwed. 

To be fair, the programmer should be aware of these issues and not
use mechanisms subject to deadlock.   As noted above, the typical
solution is to disable interrupts during a critical section.


>
>It's very unlikely that you'll hit a problem, but it is possible. 

Famous last words, indeed.

[toc] | [prev] | [next] | [standalone]


#87303 — Re: ???The pool of talented C++ developers is running dry???

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-11-09 11:47 -0800
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<tkh03q$a45v$3@dont-email.me>
In reply to#87296
On 11/9/2022 12:29 AM, David Brown wrote:
> On 09/11/2022 00:39, Chris M. Thomasson wrote:
>> On 11/8/2022 3:29 AM, Juha Nieminen wrote:
>>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote:
>>>> You keep on adding features to the language which have unintuitive 
>>>> syntax and odd rules, and
>>>> don't do much to increase the number of programs you can write 
>>>> quickly. So what happens?
>>>> Theres not much motivation to learn these features until forced to 
>>>> do so. So codebases tend
>>>> to be mainly legacy, and C++ programmers' skills fall behind.
>>>
>>> For the longest time I quite strongly disagreed with the claim that 
>>> C++ is
>>> becoming too big and too complicated.
>>>
>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is 
>>> eroding
>>> it even more.
>>>
>>> C++11 felt like a big bunch of features that the language was in dire 
>>> need
>>> of, and genuinely made programming easier. C++14 and C++17 fixed and 
>>> patched
>>> many of the minor problems and defects that turned out to exist in 
>>> C++11,
>>> so C++17 felt like "what C++11 should have been in the first place".
> 
> I agree with that.
> 
> A challenge for C++ is that even when a new and better feature is added, 
> the older and clumsier methods still have to be supported.  This also 
> means that syntax can be awkward because it can't conflict with existing 
> syntax, and the details get more complex all the time.
> 
>>
>> I was really excited and happy when C++ finally made atomics and 
>> membars part of the actual standard, C++11 iirc. Before that, I would 
>> have to code these things up in assembly language.
>>
> 
> Standard atomics would be great if they worked for my targets.  The gcc 
> implementations (and I haven't seen any others) for "advanced" use 
> (read-modify-write, or sizes larger than a standard register) is 
> completely broken for single-core systems, and even on multi-core 
> systems it is limited if you use thread priorities.  The trouble with 
> them is that no one has addressed the elephant in the room - in general, 
> you need OS support and locks to implement large atomics.

Are you referring to double-width compare-and-swap (DWCAS)? C++ should 
be able to handle it directly using the processors instruction set. Say 
C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a 
double word. Double word in the sense that they are two _contiguous_ 
words. In other words, a lock-free CAS of a double word on a 64 bit x64 
should use CMPXCHG16B.

[toc] | [prev] | [next] | [standalone]


#87304 — Re: ???The pool of talented C++ developers is running dry???

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-11-09 20:43 +0000
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<diUaL.5684$BaF9.4221@fx39.iad>
In reply to#87303
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>On 11/9/2022 12:29 AM, David Brown wrote:
>> On 09/11/2022 00:39, Chris M. Thomasson wrote:
>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote:
>>>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote:
>>>>> You keep on adding features to the language which have unintuitive 
>>>>> syntax and odd rules, and
>>>>> don't do much to increase the number of programs you can write 
>>>>> quickly. So what happens?
>>>>> Theres not much motivation to learn these features until forced to 
>>>>> do so. So codebases tend
>>>>> to be mainly legacy, and C++ programmers' skills fall behind.
>>>>
>>>> For the longest time I quite strongly disagreed with the claim that 
>>>> C++ is
>>>> becoming too big and too complicated.
>>>>
>>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is 
>>>> eroding
>>>> it even more.
>>>>
>>>> C++11 felt like a big bunch of features that the language was in dire 
>>>> need
>>>> of, and genuinely made programming easier. C++14 and C++17 fixed and 
>>>> patched
>>>> many of the minor problems and defects that turned out to exist in 
>>>> C++11,
>>>> so C++17 felt like "what C++11 should have been in the first place".
>> 
>> I agree with that.
>> 
>> A challenge for C++ is that even when a new and better feature is added, 
>> the older and clumsier methods still have to be supported.  This also 
>> means that syntax can be awkward because it can't conflict with existing 
>> syntax, and the details get more complex all the time.
>> 
>>>
>>> I was really excited and happy when C++ finally made atomics and 
>>> membars part of the actual standard, C++11 iirc. Before that, I would 
>>> have to code these things up in assembly language.
>>>
>> 
>> Standard atomics would be great if they worked for my targets.  The gcc 
>> implementations (and I haven't seen any others) for "advanced" use 
>> (read-modify-write, or sizes larger than a standard register) is 
>> completely broken for single-core systems, and even on multi-core 
>> systems it is limited if you use thread priorities.  The trouble with 
>> them is that no one has addressed the elephant in the room - in general, 
>> you need OS support and locks to implement large atomics.
>
>Are you referring to double-width compare-and-swap (DWCAS)? C++ should 
>be able to handle it directly using the processors instruction set. Say 
>C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a 
>double word. Double word in the sense that they are two _contiguous_ 
>words. In other words, a lock-free CAS of a double word on a 64 bit x64 
>should use CMPXCHG16B.

David works with low-end embedded processors, as I understand it, with
limited and/or restricted instruction sets.

[toc] | [prev] | [next] | [standalone]


#87307 — Re: ???The pool of talented C++ developers is running dry???

FromDavid Brown <david.brown@hesbynett.no>
Date2022-11-10 09:52 +0100
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<tkie4g$h5am$1@dont-email.me>
In reply to#87304
On 09/11/2022 21:43, Scott Lurndal wrote:
> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>> On 11/9/2022 12:29 AM, David Brown wrote:
>>> On 09/11/2022 00:39, Chris M. Thomasson wrote:
>>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote:
>>>>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote:
>>>>>> You keep on adding features to the language which have unintuitive
>>>>>> syntax and odd rules, and
>>>>>> don't do much to increase the number of programs you can write
>>>>>> quickly. So what happens?
>>>>>> Theres not much motivation to learn these features until forced to
>>>>>> do so. So codebases tend
>>>>>> to be mainly legacy, and C++ programmers' skills fall behind.
>>>>>
>>>>> For the longest time I quite strongly disagreed with the claim that
>>>>> C++ is
>>>>> becoming too big and too complicated.
>>>>>
>>>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is
>>>>> eroding
>>>>> it even more.
>>>>>
>>>>> C++11 felt like a big bunch of features that the language was in dire
>>>>> need
>>>>> of, and genuinely made programming easier. C++14 and C++17 fixed and
>>>>> patched
>>>>> many of the minor problems and defects that turned out to exist in
>>>>> C++11,
>>>>> so C++17 felt like "what C++11 should have been in the first place".
>>>
>>> I agree with that.
>>>
>>> A challenge for C++ is that even when a new and better feature is added,
>>> the older and clumsier methods still have to be supported.  This also
>>> means that syntax can be awkward because it can't conflict with existing
>>> syntax, and the details get more complex all the time.
>>>
>>>>
>>>> I was really excited and happy when C++ finally made atomics and
>>>> membars part of the actual standard, C++11 iirc. Before that, I would
>>>> have to code these things up in assembly language.
>>>>
>>>
>>> Standard atomics would be great if they worked for my targets.  The gcc
>>> implementations (and I haven't seen any others) for "advanced" use
>>> (read-modify-write, or sizes larger than a standard register) is
>>> completely broken for single-core systems, and even on multi-core
>>> systems it is limited if you use thread priorities.  The trouble with
>>> them is that no one has addressed the elephant in the room - in general,
>>> you need OS support and locks to implement large atomics.
>>
>> Are you referring to double-width compare-and-swap (DWCAS)? C++ should
>> be able to handle it directly using the processors instruction set. Say
>> C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a
>> double word. Double word in the sense that they are two _contiguous_
>> words. In other words, a lock-free CAS of a double word on a 64 bit x64
>> should use CMPXCHG16B.
> 
> David works with low-end embedded processors, as I understand it, with
> limited and/or restricted instruction sets.
> 

Yes.

But the principle is the same on bigger systems too.  If your processor 
can do a single-instruction 64-bit write, you see the problems for 
atomics bigger than 64-bit.  If it can handle 128-bit writes, you see 
the problems for atomics bigger than 128-bit.

Obviously the need for big atomics is much lower than the need for 
smaller ones.  Once you have a DCAS, or LL/SC, you have covered most needs.

However, these alone will not give you read-modify-write operations on 
anything bigger than you can handle with a single read (or more 
importantly, with a single unbreakable write operation).  Anything where 
the implementation is "use small atomics to get a spin lock, then do the 
work" is /broken/.  It has a small but non-zero chance of failing in 
general use on big multi-core systems.  On small single-core systems, it 
is guaranteed broken from the outset.

The C++ (and C) language, standard library, common toolchains and 
library implementations give the programmer the impression that they can 
make atomics as they like.  You can write :

	std::atomic<std::array<int, 32>> xs;

and it looks like you have a big atomic object.  But it will not work - 
you cannot rely on it.  It will /seem/ to work in all your testing, 
because the chance of hitting a problem is small - but it can fail at 
any time.

The atomics that the programmer can use should either be absolutely 
correct, guaranteed by design in all circumstances, or they should not 
be compile-time errors when you try to use atomics that are too big, or 
where the operations are too complex, for the implementation to guarantee.

It would be even better for the implementation to handle these 
correctly.  That means OS support for /real/ locks, not fake 
sort-of-works userland spin locks, but futexes or something like that 
for big systems, and interrupt disabling for single-core 
microcontrollers.  (Dual-core microcontrollers are an extra 
complication.)  Common library implementations could rely on an extra 
library or code for their "lock" and "unlock" calls - if they are not 
provided, you at least have a link error.

[toc] | [prev] | [next] | [standalone]


#87309 — Re: ???The pool of talented C++ developers is running dry???

FromMichael S <already5chosen@yahoo.com>
Date2022-11-10 03:07 -0800
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<1c041b9f-9424-41a2-b536-5f21a32addd8n@googlegroups.com>
In reply to#87307
On Thursday, November 10, 2022 at 10:52:49 AM UTC+2, David Brown wrote:
> On 09/11/2022 21:43, Scott Lurndal wrote: 
> > "Chris M. Thomasson" <chris.m.t...@gmail.com> writes: 
> >> On 11/9/2022 12:29 AM, David Brown wrote: 
> >>> On 09/11/2022 00:39, Chris M. Thomasson wrote: 
> >>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote: 
> >>>>> Malcolm McLean <malcolm.ar...@gmail.com> wrote: 
> >>>>>> You keep on adding features to the language which have unintuitive 
> >>>>>> syntax and odd rules, and 
> >>>>>> don't do much to increase the number of programs you can write 
> >>>>>> quickly. So what happens? 
> >>>>>> Theres not much motivation to learn these features until forced to 
> >>>>>> do so. So codebases tend 
> >>>>>> to be mainly legacy, and C++ programmers' skills fall behind. 
> >>>>> 
> >>>>> For the longest time I quite strongly disagreed with the claim that 
> >>>>> C++ is 
> >>>>> becoming too big and too complicated. 
> >>>>> 
> >>>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is 
> >>>>> eroding 
> >>>>> it even more. 
> >>>>> 
> >>>>> C++11 felt like a big bunch of features that the language was in dire 
> >>>>> need 
> >>>>> of, and genuinely made programming easier. C++14 and C++17 fixed and 
> >>>>> patched 
> >>>>> many of the minor problems and defects that turned out to exist in 
> >>>>> C++11, 
> >>>>> so C++17 felt like "what C++11 should have been in the first place". 
> >>> 
> >>> I agree with that. 
> >>> 
> >>> A challenge for C++ is that even when a new and better feature is added, 
> >>> the older and clumsier methods still have to be supported.  This also 
> >>> means that syntax can be awkward because it can't conflict with existing 
> >>> syntax, and the details get more complex all the time. 
> >>> 
> >>>> 
> >>>> I was really excited and happy when C++ finally made atomics and 
> >>>> membars part of the actual standard, C++11 iirc. Before that, I would 
> >>>> have to code these things up in assembly language. 
> >>>> 
> >>> 
> >>> Standard atomics would be great if they worked for my targets.  The gcc 
> >>> implementations (and I haven't seen any others) for "advanced" use 
> >>> (read-modify-write, or sizes larger than a standard register) is 
> >>> completely broken for single-core systems, and even on multi-core 
> >>> systems it is limited if you use thread priorities.  The trouble with 
> >>> them is that no one has addressed the elephant in the room - in general, 
> >>> you need OS support and locks to implement large atomics. 
> >> 
> >> Are you referring to double-width compare-and-swap (DWCAS)? C++ should 
> >> be able to handle it directly using the processors instruction set. Say 
> >> C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a 
> >> double word. Double word in the sense that they are two _contiguous_ 
> >> words. In other words, a lock-free CAS of a double word on a 64 bit x64 
> >> should use CMPXCHG16B. 
> > 
> > David works with low-end embedded processors, as I understand it, with 
> > limited and/or restricted instruction sets. 
> >
> Yes. 
> 
> But the principle is the same on bigger systems too. If your processor 
> can do a single-instruction 64-bit write, you see the problems for 
> atomics bigger than 64-bit. If it can handle 128-bit writes, you see 
> the problems for atomics bigger than 128-bit. 
> 
> Obviously the need for big atomics is much lower than the need for 
> smaller ones. Once you have a DCAS, or LL/SC, you have covered most needs. 
> 
> However, these alone will not give you read-modify-write operations on 
> anything bigger than you can handle with a single read (or more 
> importantly, with a single unbreakable write operation). Anything where 
> the implementation is "use small atomics to get a spin lock, then do the 
> work" is /broken/. It has a small but non-zero chance of failing in 
> general use on big multi-core systems. On small single-core systems, it 
> is guaranteed broken from the outset. 
> 
> The C++ (and C) language, standard library, common toolchains and 
> library implementations give the programmer the impression that they can 
> make atomics as they like. You can write : 
> 
> std::atomic<std::array<int, 32>> xs; 
> 
> and it looks like you have a big atomic object. But it will not work - 
> you cannot rely on it. It will /seem/ to work in all your testing, 
> because the chance of hitting a problem is small - but it can fail at 
> any time. 
> 
> The atomics that the programmer can use should either be absolutely 
> correct, guaranteed by design in all circumstances, or they should not 
> be compile-time errors when you try to use atomics that are too big, or 
> where the operations are too complex, for the implementation to guarantee. 
> 

Yes, failing in compile time is the most reasonable.

> It would be even better for the implementation to handle these 
> correctly. That means OS support for /real/ locks, not fake 
> sort-of-works userland spin locks, but futexes or something like that 
> for big systems, and interrupt disabling for single-core 
> microcontrollers. (Dual-core microcontrollers are an extra 
> complication.) Common library implementations could rely on an extra 
> library or code for their "lock" and "unlock" calls - if they are not 
> provided, you at least have a link error.

It certainly would be against the spirit of 'C'.
I'm not sure about about relationship to spirit of C++.
Also I'm not sure that C++ has spirit.

[toc] | [prev] | [next] | [standalone]


#87317 — Re: ???The pool of talented C++ developers is running dry???

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-11-10 11:53 -0800
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<tkjkqu$k10j$3@dont-email.me>
In reply to#87307
On 11/10/2022 12:52 AM, David Brown wrote:
> On 09/11/2022 21:43, Scott Lurndal wrote:
>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>> On 11/9/2022 12:29 AM, David Brown wrote:
>>>> On 09/11/2022 00:39, Chris M. Thomasson wrote:
>>>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote:
>>>>>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote:
>>>>>>> You keep on adding features to the language which have unintuitive
>>>>>>> syntax and odd rules, and
>>>>>>> don't do much to increase the number of programs you can write
>>>>>>> quickly. So what happens?
>>>>>>> Theres not much motivation to learn these features until forced to
>>>>>>> do so. So codebases tend
>>>>>>> to be mainly legacy, and C++ programmers' skills fall behind.
>>>>>>
>>>>>> For the longest time I quite strongly disagreed with the claim that
>>>>>> C++ is
>>>>>> becoming too big and too complicated.
>>>>>>
>>>>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is
>>>>>> eroding
>>>>>> it even more.
>>>>>>
>>>>>> C++11 felt like a big bunch of features that the language was in dire
>>>>>> need
>>>>>> of, and genuinely made programming easier. C++14 and C++17 fixed and
>>>>>> patched
>>>>>> many of the minor problems and defects that turned out to exist in
>>>>>> C++11,
>>>>>> so C++17 felt like "what C++11 should have been in the first place".
>>>>
>>>> I agree with that.
>>>>
>>>> A challenge for C++ is that even when a new and better feature is 
>>>> added,
>>>> the older and clumsier methods still have to be supported.  This also
>>>> means that syntax can be awkward because it can't conflict with 
>>>> existing
>>>> syntax, and the details get more complex all the time.
>>>>
>>>>>
>>>>> I was really excited and happy when C++ finally made atomics and
>>>>> membars part of the actual standard, C++11 iirc. Before that, I would
>>>>> have to code these things up in assembly language.
>>>>>
>>>>
>>>> Standard atomics would be great if they worked for my targets.  The gcc
>>>> implementations (and I haven't seen any others) for "advanced" use
>>>> (read-modify-write, or sizes larger than a standard register) is
>>>> completely broken for single-core systems, and even on multi-core
>>>> systems it is limited if you use thread priorities.  The trouble with
>>>> them is that no one has addressed the elephant in the room - in 
>>>> general,
>>>> you need OS support and locks to implement large atomics.
>>>
>>> Are you referring to double-width compare-and-swap (DWCAS)? C++ should
>>> be able to handle it directly using the processors instruction set. Say
>>> C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a
>>> double word. Double word in the sense that they are two _contiguous_
>>> words. In other words, a lock-free CAS of a double word on a 64 bit x64
>>> should use CMPXCHG16B.
>>
>> David works with low-end embedded processors, as I understand it, with
>> limited and/or restricted instruction sets.
>>
> 
> Yes.
> 
> But the principle is the same on bigger systems too.  If your processor 
> can do a single-instruction 64-bit write, you see the problems for 
> atomics bigger than 64-bit.  If it can handle 128-bit writes, you see 
> the problems for atomics bigger than 128-bit.
> 
> Obviously the need for big atomics is much lower than the need for 
> smaller ones.  Once you have a DCAS, or LL/SC, you have covered most needs.
> 
> However, these alone will not give you read-modify-write operations on 
> anything bigger than you can handle with a single read (or more 
> importantly, with a single unbreakable write operation).  Anything where 
> the implementation is "use small atomics to get a spin lock, then do the 
> work" is /broken/.  It has a small but non-zero chance of failing in 
> general use on big multi-core systems.  On small single-core systems, it 
> is guaranteed broken from the outset.
> 
> The C++ (and C) language, standard library, common toolchains and 
> library implementations give the programmer the impression that they can 
> make atomics as they like.  You can write :
> 
>      std::atomic<std::array<int, 32>> xs;
> 
> and it looks like you have a big atomic object.  But it will not work - 
> you cannot rely on it.  It will /seem/ to work in all your testing, 
> because the chance of hitting a problem is small - but it can fail at 
> any time.

A rule of thumb... Imvho, always check the result of:

https://en.cppreference.com/w/cpp/atomic/atomic/is_lock_free

Just to be, sure... ;^) Fwiw, DWCAS is very different than DCAS. The 
latter can work on two non-contiguous words. The former only works with 
contiguous words. A main reason for DWCAS to exist in the first place is 
to be able to handle a lock-free stack. A pointer and an version count 
to combat the ABA problem. Although, there are many other interesting 
uses for DWCAS...

https://groups.google.com/g/comp.lang.c++/c/nUDtke-H1io/m/g87spoMUCgAJ


> The atomics that the programmer can use should either be absolutely 
> correct, guaranteed by design in all circumstances, or they should not 
> be compile-time errors when you try to use atomics that are too big, or 
> where the operations are too complex, for the implementation to guarantee.
> 
> It would be even better for the implementation to handle these 
> correctly.  That means OS support for /real/ locks, not fake 
> sort-of-works userland spin locks, but futexes or something like that 
> for big systems, and interrupt disabling for single-core 
> microcontrollers.  (Dual-core microcontrollers are an extra 
> complication.)  Common library implementations could rely on an extra 
> library or code for their "lock" and "unlock" calls - if they are not 
> provided, you at least have a link error.
> 

If the result of is_lock_free is not true, then you should really think 
about digging into how the locking is actually implemented. Hash based 
address locking is one simple way to do it. Fwiw, I created one called 
multi-mutex:

https://groups.google.com/g/comp.lang.c++/c/sV4WC_cBb9Q/m/Ti8LFyH4CgAJ

[toc] | [prev] | [next] | [standalone]


#87318 — Re: ???The pool of talented C++ developers is running dry???

FromDavid Brown <david.brown@hesbynett.no>
Date2022-11-10 21:44 +0100
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<tkjnqt$k9i0$1@dont-email.me>
In reply to#87317
On 10/11/2022 20:53, Chris M. Thomasson wrote:
> On 11/10/2022 12:52 AM, David Brown wrote:
>> On 09/11/2022 21:43, Scott Lurndal wrote:
>>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>>> On 11/9/2022 12:29 AM, David Brown wrote:
>>>>> On 09/11/2022 00:39, Chris M. Thomasson wrote:
>>>>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote:
>>>>>>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote:
>>>>>>>> You keep on adding features to the language which have unintuitive
>>>>>>>> syntax and odd rules, and
>>>>>>>> don't do much to increase the number of programs you can write
>>>>>>>> quickly. So what happens?
>>>>>>>> Theres not much motivation to learn these features until forced to
>>>>>>>> do so. So codebases tend
>>>>>>>> to be mainly legacy, and C++ programmers' skills fall behind.
>>>>>>>
>>>>>>> For the longest time I quite strongly disagreed with the claim that
>>>>>>> C++ is
>>>>>>> becoming too big and too complicated.
>>>>>>>
>>>>>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is
>>>>>>> eroding
>>>>>>> it even more.
>>>>>>>
>>>>>>> C++11 felt like a big bunch of features that the language was in 
>>>>>>> dire
>>>>>>> need
>>>>>>> of, and genuinely made programming easier. C++14 and C++17 fixed and
>>>>>>> patched
>>>>>>> many of the minor problems and defects that turned out to exist in
>>>>>>> C++11,
>>>>>>> so C++17 felt like "what C++11 should have been in the first place".
>>>>>
>>>>> I agree with that.
>>>>>
>>>>> A challenge for C++ is that even when a new and better feature is 
>>>>> added,
>>>>> the older and clumsier methods still have to be supported.  This also
>>>>> means that syntax can be awkward because it can't conflict with 
>>>>> existing
>>>>> syntax, and the details get more complex all the time.
>>>>>
>>>>>>
>>>>>> I was really excited and happy when C++ finally made atomics and
>>>>>> membars part of the actual standard, C++11 iirc. Before that, I would
>>>>>> have to code these things up in assembly language.
>>>>>>
>>>>>
>>>>> Standard atomics would be great if they worked for my targets.  The 
>>>>> gcc
>>>>> implementations (and I haven't seen any others) for "advanced" use
>>>>> (read-modify-write, or sizes larger than a standard register) is
>>>>> completely broken for single-core systems, and even on multi-core
>>>>> systems it is limited if you use thread priorities.  The trouble with
>>>>> them is that no one has addressed the elephant in the room - in 
>>>>> general,
>>>>> you need OS support and locks to implement large atomics.
>>>>
>>>> Are you referring to double-width compare-and-swap (DWCAS)? C++ should
>>>> be able to handle it directly using the processors instruction set. Say
>>>> C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a
>>>> double word. Double word in the sense that they are two _contiguous_
>>>> words. In other words, a lock-free CAS of a double word on a 64 bit x64
>>>> should use CMPXCHG16B.
>>>
>>> David works with low-end embedded processors, as I understand it, with
>>> limited and/or restricted instruction sets.
>>>
>>
>> Yes.
>>
>> But the principle is the same on bigger systems too.  If your 
>> processor can do a single-instruction 64-bit write, you see the 
>> problems for atomics bigger than 64-bit.  If it can handle 128-bit 
>> writes, you see the problems for atomics bigger than 128-bit.
>>
>> Obviously the need for big atomics is much lower than the need for 
>> smaller ones.  Once you have a DCAS, or LL/SC, you have covered most 
>> needs.
>>
>> However, these alone will not give you read-modify-write operations on 
>> anything bigger than you can handle with a single read (or more 
>> importantly, with a single unbreakable write operation).  Anything 
>> where the implementation is "use small atomics to get a spin lock, 
>> then do the work" is /broken/.  It has a small but non-zero chance of 
>> failing in general use on big multi-core systems.  On small 
>> single-core systems, it is guaranteed broken from the outset.
>>
>> The C++ (and C) language, standard library, common toolchains and 
>> library implementations give the programmer the impression that they 
>> can make atomics as they like.  You can write :
>>
>>      std::atomic<std::array<int, 32>> xs;
>>
>> and it looks like you have a big atomic object.  But it will not work 
>> - you cannot rely on it.  It will /seem/ to work in all your testing, 
>> because the chance of hitting a problem is small - but it can fail at 
>> any time.
> 
> A rule of thumb... Imvho, always check the result of:
> 
> https://en.cppreference.com/w/cpp/atomic/atomic/is_lock_free

And it if is not, what is the point in allowing it if the locks don't work?

> 
> Just to be, sure... ;^) Fwiw, DWCAS is very different than DCAS. The 
> latter can work on two non-contiguous words. The former only works with 
> contiguous words. A main reason for DWCAS to exist in the first place is 
> to be able to handle a lock-free stack. A pointer and an version count 
> to combat the ABA problem. Although, there are many other interesting 
> uses for DWCAS...
> 
> https://groups.google.com/g/comp.lang.c++/c/nUDtke-H1io/m/g87spoMUCgAJ
> 
> 
>> The atomics that the programmer can use should either be absolutely 
>> correct, guaranteed by design in all circumstances, or they should not 
>> be compile-time errors when you try to use atomics that are too big, 
>> or where the operations are too complex, for the implementation to 
>> guarantee.
>>
>> It would be even better for the implementation to handle these 
>> correctly.  That means OS support for /real/ locks, not fake 
>> sort-of-works userland spin locks, but futexes or something like that 
>> for big systems, and interrupt disabling for single-core 
>> microcontrollers.  (Dual-core microcontrollers are an extra 
>> complication.)  Common library implementations could rely on an extra 
>> library or code for their "lock" and "unlock" calls - if they are not 
>> provided, you at least have a link error.
>>
> 
> If the result of is_lock_free is not true, then you should really think 
> about digging into how the locking is actually implemented. Hash based 
> address locking is one simple way to do it. Fwiw, I created one called 
> multi-mutex:
> 
> https://groups.google.com/g/comp.lang.c++/c/sV4WC_cBb9Q/m/Ti8LFyH4CgAJ

Hash-based arrays of locks are as bad as a single lock for all atomics, 
in that it does not work unless it is a proper OS lock.  The larger your 
array of locks, the lower your chances of problems, but it all comes 
down to one thing - are your locks safe or not?


[toc] | [prev] | [next] | [standalone]


#87319 — Re: ???The pool of talented C++ developers is running dry???

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-11-10 12:53 -0800
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<tkjocq$kaut$1@dont-email.me>
In reply to#87318
On 11/10/2022 12:44 PM, David Brown wrote:
> On 10/11/2022 20:53, Chris M. Thomasson wrote:
>> On 11/10/2022 12:52 AM, David Brown wrote:
>>> On 09/11/2022 21:43, Scott Lurndal wrote:
>>>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>>>> On 11/9/2022 12:29 AM, David Brown wrote:
>>>>>> On 09/11/2022 00:39, Chris M. Thomasson wrote:
>>>>>>> On 11/8/2022 3:29 AM, Juha Nieminen wrote:
>>>>>>>> Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote:
>>>>>>>>> You keep on adding features to the language which have unintuitive
>>>>>>>>> syntax and odd rules, and
>>>>>>>>> don't do much to increase the number of programs you can write
>>>>>>>>> quickly. So what happens?
>>>>>>>>> Theres not much motivation to learn these features until forced to
>>>>>>>>> do so. So codebases tend
>>>>>>>>> to be mainly legacy, and C++ programmers' skills fall behind.
>>>>>>>>
>>>>>>>> For the longest time I quite strongly disagreed with the claim that
>>>>>>>> C++ is
>>>>>>>> becoming too big and too complicated.
>>>>>>>>
>>>>>>>> However, C++20 has eroded this conviction of mine somewhat. 
>>>>>>>> C++23 is
>>>>>>>> eroding
>>>>>>>> it even more.
>>>>>>>>
>>>>>>>> C++11 felt like a big bunch of features that the language was in 
>>>>>>>> dire
>>>>>>>> need
>>>>>>>> of, and genuinely made programming easier. C++14 and C++17 fixed 
>>>>>>>> and
>>>>>>>> patched
>>>>>>>> many of the minor problems and defects that turned out to exist in
>>>>>>>> C++11,
>>>>>>>> so C++17 felt like "what C++11 should have been in the first 
>>>>>>>> place".
>>>>>>
>>>>>> I agree with that.
>>>>>>
>>>>>> A challenge for C++ is that even when a new and better feature is 
>>>>>> added,
>>>>>> the older and clumsier methods still have to be supported.  This also
>>>>>> means that syntax can be awkward because it can't conflict with 
>>>>>> existing
>>>>>> syntax, and the details get more complex all the time.
>>>>>>
>>>>>>>
>>>>>>> I was really excited and happy when C++ finally made atomics and
>>>>>>> membars part of the actual standard, C++11 iirc. Before that, I 
>>>>>>> would
>>>>>>> have to code these things up in assembly language.
>>>>>>>
>>>>>>
>>>>>> Standard atomics would be great if they worked for my targets.  
>>>>>> The gcc
>>>>>> implementations (and I haven't seen any others) for "advanced" use
>>>>>> (read-modify-write, or sizes larger than a standard register) is
>>>>>> completely broken for single-core systems, and even on multi-core
>>>>>> systems it is limited if you use thread priorities.  The trouble with
>>>>>> them is that no one has addressed the elephant in the room - in 
>>>>>> general,
>>>>>> you need OS support and locks to implement large atomics.
>>>>>
>>>>> Are you referring to double-width compare-and-swap (DWCAS)? C++ should
>>>>> be able to handle it directly using the processors instruction set. 
>>>>> Say
>>>>> C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used 
>>>>> for a
>>>>> double word. Double word in the sense that they are two _contiguous_
>>>>> words. In other words, a lock-free CAS of a double word on a 64 bit 
>>>>> x64
>>>>> should use CMPXCHG16B.
>>>>
>>>> David works with low-end embedded processors, as I understand it, with
>>>> limited and/or restricted instruction sets.
>>>>
>>>
>>> Yes.
>>>
>>> But the principle is the same on bigger systems too.  If your 
>>> processor can do a single-instruction 64-bit write, you see the 
>>> problems for atomics bigger than 64-bit.  If it can handle 128-bit 
>>> writes, you see the problems for atomics bigger than 128-bit.
>>>
>>> Obviously the need for big atomics is much lower than the need for 
>>> smaller ones.  Once you have a DCAS, or LL/SC, you have covered most 
>>> needs.
>>>
>>> However, these alone will not give you read-modify-write operations 
>>> on anything bigger than you can handle with a single read (or more 
>>> importantly, with a single unbreakable write operation).  Anything 
>>> where the implementation is "use small atomics to get a spin lock, 
>>> then do the work" is /broken/.  It has a small but non-zero chance of 
>>> failing in general use on big multi-core systems.  On small 
>>> single-core systems, it is guaranteed broken from the outset.
>>>
>>> The C++ (and C) language, standard library, common toolchains and 
>>> library implementations give the programmer the impression that they 
>>> can make atomics as they like.  You can write :
>>>
>>>      std::atomic<std::array<int, 32>> xs;
>>>
>>> and it looks like you have a big atomic object.  But it will not work 
>>> - you cannot rely on it.  It will /seem/ to work in all your testing, 
>>> because the chance of hitting a problem is small - but it can fail at 
>>> any time.
>>
>> A rule of thumb... Imvho, always check the result of:
>>
>> https://en.cppreference.com/w/cpp/atomic/atomic/is_lock_free
> 
> And it if is not, what is the point in allowing it if the locks don't work?
> 
>>
>> Just to be, sure... ;^) Fwiw, DWCAS is very different than DCAS. The 
>> latter can work on two non-contiguous words. The former only works 
>> with contiguous words. A main reason for DWCAS to exist in the first 
>> place is to be able to handle a lock-free stack. A pointer and an 
>> version count to combat the ABA problem. Although, there are many 
>> other interesting uses for DWCAS...
>>
>> https://groups.google.com/g/comp.lang.c++/c/nUDtke-H1io/m/g87spoMUCgAJ
>>
>>
>>> The atomics that the programmer can use should either be absolutely 
>>> correct, guaranteed by design in all circumstances, or they should 
>>> not be compile-time errors when you try to use atomics that are too 
>>> big, or where the operations are too complex, for the implementation 
>>> to guarantee.
>>>
>>> It would be even better for the implementation to handle these 
>>> correctly.  That means OS support for /real/ locks, not fake 
>>> sort-of-works userland spin locks, but futexes or something like that 
>>> for big systems, and interrupt disabling for single-core 
>>> microcontrollers.  (Dual-core microcontrollers are an extra 
>>> complication.)  Common library implementations could rely on an extra 
>>> library or code for their "lock" and "unlock" calls - if they are not 
>>> provided, you at least have a link error.
>>>
>>
>> If the result of is_lock_free is not true, then you should really 
>> think about digging into how the locking is actually implemented. Hash 
>> based address locking is one simple way to do it. Fwiw, I created one 
>> called multi-mutex:
>>
>> https://groups.google.com/g/comp.lang.c++/c/sV4WC_cBb9Q/m/Ti8LFyH4CgAJ
> 
> Hash-based arrays of locks are as bad as a single lock for all atomics, 
> in that it does not work unless it is a proper OS lock.  The larger your 
> array of locks, the lower your chances of problems, but it all comes 
> down to one thing - are your locks safe or not?

Well, my multi-mutex uses std::mutex as elements of its vector of locks. 
It hashes an address into said table. So, it's only as good as the 
implementation of std::mutex...

std::vector<std::mutex> m_locks;

Fair enough?

[toc] | [prev] | [next] | [standalone]


#87320 — Re: ???The pool of talented C++ developers is running dry???

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-11-10 21:20 +0000
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<AWdbL.84469$2Rs3.11964@fx12.iad>
In reply to#87319
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>On 11/10/2022 12:44 PM, David Brown wrote:

>> Hash-based arrays of locks are as bad as a single lock for all atomics, 
>> in that it does not work unless it is a proper OS lock.  The larger your 
>> array of locks, the lower your chances of problems, but it all comes 
>> down to one thing - are your locks safe or not?
>
>Well, my multi-mutex uses std::mutex as elements of its vector of locks. 
>It hashes an address into said table. So, it's only as good as the 
>implementation of std::mutex...
>
>std::vector<std::mutex> m_locks;
>
>Fair enough?
>

I'd point out that using a vector is sub-optimal if the various
elements of the vector are accessed from different cores due to
false sharing....

[toc] | [prev] | [next] | [standalone]


#87321 — Re: ???The pool of talented C++ developers is running dry???

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-11-10 13:23 -0800
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<tkjq48$kflb$1@dont-email.me>
In reply to#87320
On 11/10/2022 1:20 PM, Scott Lurndal wrote:
> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>> On 11/10/2022 12:44 PM, David Brown wrote:
> 
>>> Hash-based arrays of locks are as bad as a single lock for all atomics,
>>> in that it does not work unless it is a proper OS lock.  The larger your
>>> array of locks, the lower your chances of problems, but it all comes
>>> down to one thing - are your locks safe or not?
>>
>> Well, my multi-mutex uses std::mutex as elements of its vector of locks.
>> It hashes an address into said table. So, it's only as good as the
>> implementation of std::mutex...
>>
>> std::vector<std::mutex> m_locks;
>>
>> Fair enough?
>>
> 
> I'd point out that using a vector is sub-optimal if the various
> elements of the vector are accessed from different cores due to
> false sharing....

Well, the mutex state should really be isolated on their own cache 
lines. Padding and alignment on a l2 cacheline boundary. Iirc, this can 
be done in modern c++. Even aligning elements of a vector.

[toc] | [prev] | [next] | [standalone]


#87322 — Re: ???The pool of talented C++ developers is running dry???

FromDavid Brown <david.brown@hesbynett.no>
Date2022-11-11 08:10 +0100
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<tkksi3$pom6$1@dont-email.me>
In reply to#87319
On 10/11/2022 21:53, Chris M. Thomasson wrote:
> On 11/10/2022 12:44 PM, David Brown wrote:
>> On 10/11/2022 20:53, Chris M. Thomasson wrote:
>>> On 11/10/2022 12:52 AM, David Brown wrote:

>>> If the result of is_lock_free is not true, then you should really 
>>> think about digging into how the locking is actually implemented. 
>>> Hash based address locking is one simple way to do it. Fwiw, I 
>>> created one called multi-mutex:
>>>
>>> https://groups.google.com/g/comp.lang.c++/c/sV4WC_cBb9Q/m/Ti8LFyH4CgAJ
>>
>> Hash-based arrays of locks are as bad as a single lock for all 
>> atomics, in that it does not work unless it is a proper OS lock.  The 
>> larger your array of locks, the lower your chances of problems, but it 
>> all comes down to one thing - are your locks safe or not?
> 
> Well, my multi-mutex uses std::mutex as elements of its vector of locks. 
> It hashes an address into said table. So, it's only as good as the 
> implementation of std::mutex...
> 
> std::vector<std::mutex> m_locks;
> 
> Fair enough?
> 

As long as std::mutex is a wrapper for a real OS mutex, that will be 
fine (after padding for cache line sizes).

The problem with the atomics library in gcc is not the array of locks or 
the hashing on address, but that it doesn't use /real/ locks.

[toc] | [prev] | [next] | [standalone]


#87323 — Re: ???The pool of talented C++ developers is running dry???

FromÖö Tiib <ootiib@hot.ee>
Date2022-11-11 01:45 -0800
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<4990e23a-9da7-4934-8e7c-40a0255b115dn@googlegroups.com>
In reply to#87322
On Friday, 11 November 2022 at 09:11:17 UTC+2, David Brown wrote:
> On 10/11/2022 21:53, Chris M. Thomasson wrote: 
> > On 11/10/2022 12:44 PM, David Brown wrote: 
> >> On 10/11/2022 20:53, Chris M. Thomasson wrote: 
> >>> On 11/10/2022 12:52 AM, David Brown wrote: 
> 
> >>> If the result of is_lock_free is not true, then you should really 
> >>> think about digging into how the locking is actually implemented. 
> >>> Hash based address locking is one simple way to do it. Fwiw, I 
> >>> created one called multi-mutex: 
> >>> 
> >>> https://groups.google.com/g/comp.lang.c++/c/sV4WC_cBb9Q/m/Ti8LFyH4CgAJ 
> >> 
> >> Hash-based arrays of locks are as bad as a single lock for all 
> >> atomics, in that it does not work unless it is a proper OS lock.  The 
> >> larger your array of locks, the lower your chances of problems, but it 
> >> all comes down to one thing - are your locks safe or not? 
> > 
> > Well, my multi-mutex uses std::mutex as elements of its vector of locks. 
> > It hashes an address into said table. So, it's only as good as the 
> > implementation of std::mutex... 
> > 
> > std::vector<std::mutex> m_locks; 
> > 
> > Fair enough? 
> >
> As long as std::mutex is a wrapper for a real OS mutex, that will be 
> fine (after padding for cache line sizes). 
> 
> The problem with the atomics library in gcc is not the array of locks or 
> the hashing on address, but that it doesn't use /real/ locks.

Yes.  Libraries look often like sabotaged and so who needs quality
and portability has to use conditional compiling for what they
use std::atomic<Something> and for what manually mutex-protected
Something instance. Purpose defeated.

[toc] | [prev] | [next] | [standalone]


#87310 — Re: ???The pool of talented C++ developers is running dry???

FromMichael S <already5chosen@yahoo.com>
Date2022-11-10 03:14 -0800
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<eac1413e-eec0-4530-9ea4-8cb5f97ecd6fn@googlegroups.com>
In reply to#87304
On Wednesday, November 9, 2022 at 10:44:10 PM UTC+2, Scott Lurndal wrote:
> "Chris M. Thomasson" <chris.m.t...@gmail.com> writes: 
> >On 11/9/2022 12:29 AM, David Brown wrote: 
> >> On 09/11/2022 00:39, Chris M. Thomasson wrote: 
> >>> On 11/8/2022 3:29 AM, Juha Nieminen wrote: 
> >>>> Malcolm McLean <malcolm.ar...@gmail.com> wrote: 
> >>>>> You keep on adding features to the language which have unintuitive 
> >>>>> syntax and odd rules, and 
> >>>>> don't do much to increase the number of programs you can write 
> >>>>> quickly. So what happens? 
> >>>>> Theres not much motivation to learn these features until forced to 
> >>>>> do so. So codebases tend 
> >>>>> to be mainly legacy, and C++ programmers' skills fall behind. 
> >>>> 
> >>>> For the longest time I quite strongly disagreed with the claim that 
> >>>> C++ is 
> >>>> becoming too big and too complicated. 
> >>>> 
> >>>> However, C++20 has eroded this conviction of mine somewhat. C++23 is 
> >>>> eroding 
> >>>> it even more. 
> >>>> 
> >>>> C++11 felt like a big bunch of features that the language was in dire 
> >>>> need 
> >>>> of, and genuinely made programming easier. C++14 and C++17 fixed and 
> >>>> patched 
> >>>> many of the minor problems and defects that turned out to exist in 
> >>>> C++11, 
> >>>> so C++17 felt like "what C++11 should have been in the first place". 
> >> 
> >> I agree with that. 
> >> 
> >> A challenge for C++ is that even when a new and better feature is added, 
> >> the older and clumsier methods still have to be supported.  This also 
> >> means that syntax can be awkward because it can't conflict with existing 
> >> syntax, and the details get more complex all the time. 
> >> 
> >>> 
> >>> I was really excited and happy when C++ finally made atomics and 
> >>> membars part of the actual standard, C++11 iirc. Before that, I would 
> >>> have to code these things up in assembly language. 
> >>> 
> >> 
> >> Standard atomics would be great if they worked for my targets.  The gcc 
> >> implementations (and I haven't seen any others) for "advanced" use 
> >> (read-modify-write, or sizes larger than a standard register) is 
> >> completely broken for single-core systems, and even on multi-core 
> >> systems it is limited if you use thread priorities.  The trouble with 
> >> them is that no one has addressed the elephant in the room - in general, 
> >> you need OS support and locks to implement large atomics. 
> > 
> >Are you referring to double-width compare-and-swap (DWCAS)? C++ should 
> >be able to handle it directly using the processors instruction set. Say 
> >C++ on a modern 64 bit x64 system, well CMPXCHG16B should be used for a 
> >double word. Double word in the sense that they are two _contiguous_ 
> >words. In other words, a lock-free CAS of a double word on a 64 bit x64 
> >should use CMPXCHG16B.
> David works with low-end embedded processors, as I understand it, with 
> limited and/or restricted instruction sets.

In my book Cortex-M (except M0) is not a low end.
M7 in particular is too big and too complicated not even for 
proverbial 99%, but for solid 100% of my microcontroller needs.

[toc] | [prev] | [next] | [standalone]


#87297 — Re: ???The pool of talented C++ developers is running dry???

FromSam <sam@email-scan.com>
Date2022-11-09 07:59 -0500
SubjectRe: ???The pool of talented C++ developers is running dry???
Message-ID<cone.1667998785.254454.147087.1004@monster.email-scan.com>
In reply to#87283
Juha Nieminen writes:

> C++20, however, doesn't feel like this anymore. It has a few new features
> that genuinely help in programming, but most of it feels like just adding
> features for the sake of adding them. C++23 even moreso.

Or the features were specifically added to make sucky operating systems suck  
a little less. Specifically: co-routines. Microsoft hijacked the  
standardization process to push through co-routines, because real multiple  
execution threads on MS-Windows blows chunks, and the OS can only implement  
co-routines in a passable manner.

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | comp.lang.c++


csiph-web