Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #86454 > unrolled thread
| Started by | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| First post | 2022-09-21 10:04 +0000 |
| Last post | 2022-09-22 21:03 -0500 |
| Articles | 20 on this page of 90 — 16 participants |
Back to article view | Back to comp.lang.c++
Never use strncpy! Juha Nieminen <nospam@thanks.invalid> - 2022-09-21 10:04 +0000
Re: Never use strncpy! "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-09-21 15:06 +0200
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-21 15:24 +0200
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-21 15:30 +0200
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-21 19:20 +0200
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-21 19:33 +0200
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-21 23:14 +0200
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-22 04:40 +0200
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-22 08:56 +0200
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-22 12:02 +0200
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-23 13:23 +0200
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-23 13:49 +0200
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-23 15:25 +0200
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-23 16:03 +0200
Re: Never use strncpy! scott@slp53.sl.home (Scott Lurndal) - 2022-09-23 14:34 +0000
Re: Never use strncpy! Michael S <already5chosen@yahoo.com> - 2022-09-23 08:19 -0700
Re: Never use strncpy! scott@slp53.sl.home (Scott Lurndal) - 2022-09-23 17:57 +0000
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-23 18:36 +0200
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-24 16:22 +0200
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-24 16:25 +0200
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-24 17:09 +0200
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-24 17:14 +0200
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-09-26 17:34 -0700
Re: Never use strncpy! Juha Nieminen <nospam@thanks.invalid> - 2022-09-26 07:53 +0000
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-26 10:21 +0200
Re: Never use strncpy! Juha Nieminen <nospam@thanks.invalid> - 2022-09-26 08:34 +0000
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-26 13:39 +0200
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-09-26 17:47 -0700
Re: Never use strncpy! scott@slp53.sl.home (Scott Lurndal) - 2022-09-27 14:17 +0000
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-27 18:07 +0200
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-09-27 12:42 -0700
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-28 09:32 +0200
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-09-28 13:34 -0700
Re: Never use strncpy! scott@slp53.sl.home (Scott Lurndal) - 2022-09-28 21:11 +0000
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-10-01 06:26 +0200
Re: Never use strncpy! scott@slp53.sl.home (Scott Lurndal) - 2022-10-01 17:01 +0000
Re: Never use strncpy! Michael S <already5chosen@yahoo.com> - 2022-10-02 15:32 -0700
Re: Never use strncpy! scott@slp53.sl.home (Scott Lurndal) - 2022-10-03 14:01 +0000
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-10-03 13:49 -0700
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-10-02 12:46 -0700
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-10-03 04:14 +0200
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-10-02 19:43 -0700
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-10-02 19:48 -0700
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-10-03 04:53 +0200
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-10-02 19:58 -0700
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-10-03 08:27 +0200
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-10-03 13:49 -0700
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-10-04 10:53 -0700
Re: Never use strncpy! "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-10-02 12:44 -0700
Re: Never use strncpy! Juha Nieminen <nospam@thanks.invalid> - 2022-09-22 05:59 +0000
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-22 09:32 +0200
Re: Never use strncpy! Muttley@dastardlyhq.com - 2022-09-21 15:42 +0000
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-21 18:10 +0200
Re: Never use strncpy! Muttley@dastardlyhq.com - 2022-09-21 16:19 +0000
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-21 18:27 +0200
Re: Never use strncpy! Muttley@dastardlyhq.com - 2022-09-23 14:47 +0000
Re: Never use strncpy! scott@slp53.sl.home (Scott Lurndal) - 2022-09-21 16:12 +0000
Re: Never use strncpy! Philipp Klaus Krause <pkk@spth.de> - 2022-09-23 19:52 +0200
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-24 16:25 +0200
Re: Never use strncpy! scott@slp53.sl.home (Scott Lurndal) - 2022-09-21 14:00 +0000
Re: Never use strncpy! Juha Nieminen <nospam@thanks.invalid> - 2022-09-22 06:03 +0000
Re: Never use strncpy! Philipp Klaus Krause <pkk@spth.de> - 2022-09-23 19:50 +0200
Re: Never use strncpy! Juha Nieminen <nospam@thanks.invalid> - 2022-09-26 07:54 +0000
Re: Never use strncpy! Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-09-21 15:56 +0100
Re: Never use strncpy! Muttley@dastardlyhq.com - 2022-09-21 15:36 +0000
Re: Never use strncpy! Frederick Virchanza Gotham <cauldwell.thomas@gmail.com> - 2022-09-21 08:59 -0700
Re: Never use strncpy! Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-09-21 13:23 -0700
Re: Never use strncpy! Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-09-21 21:49 +0100
Re: Never use strncpy! Richard Damon <Richard@Damon-Family.org> - 2022-09-21 19:18 -0400
Re: Never use strncpy! Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-09-22 00:39 +0100
Re: Never use strncpy! Richard Damon <Richard@Damon-Family.org> - 2022-09-21 20:45 -0400
Re: Never use strncpy! Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-09-21 18:00 -0700
Re: Never use strncpy! Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-09-21 17:55 -0700
Re: Never use strncpy! Juha Nieminen <nospam@thanks.invalid> - 2022-09-22 06:09 +0000
Re: Never use strncpy! Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-09-22 11:44 -0700
Re: Never use strncpy! Richard Damon <Richard@Damon-Family.org> - 2022-09-22 23:33 -0400
Re: Never use strncpy! Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-09-22 19:00 -0700
Re: Never use strncpy! Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-09-22 19:04 -0700
Re: Never use strncpy! Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-09-22 22:49 -0700
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-24 11:38 +0200
Re: Never use strncpy! Muttley@dastardlyhq.com - 2022-09-24 09:43 +0000
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-24 11:59 +0200
Re: Never use strncpy! Richard Damon <Richard@Damon-Family.org> - 2022-09-24 06:27 -0400
Re: Never use strncpy! Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-09-24 11:53 -0700
Re: Never use strncpy! Bonita Montero <Bonita.Montero@gmail.com> - 2022-09-24 20:58 +0200
Re: Never use strncpy! Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-09-21 16:01 -0700
Re: Never use strncpy! David Brown <david.brown@hesbynett.no> - 2022-09-22 09:37 +0200
Re: Never use strncpy! Manfred <noname@add.invalid> - 2022-10-01 01:24 +0200
Re: Never use strncpy! Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-09-30 16:43 -0700
Re: Never use strncpy! Lynn McGuire <lynnmcguire5@gmail.com> - 2022-09-22 21:03 -0500
Page 2 of 5 — ← Prev page 1 [2] 3 4 5 Next page →
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2022-09-24 17:09 +0200 |
| Message-ID | <tgn6k6$30jog$1@dont-email.me> |
| In reply to | #86557 |
On 24/09/2022 16:25, Bonita Montero wrote: > Am 24.09.2022 um 16:22 schrieb David Brown: > >> There are a number of ways to implement atomics that are larger than >> a single bus operation can handle, or that involve multiple bus >> operations. ... > > It's always done by STM, and that's slow. No, it is not. You really have no idea about these things - repeating yourself does not make it any less myopic. I think I'm done trying to explain them to you.
[toc] | [prev] | [next] | [standalone]
| From | Bonita Montero <Bonita.Montero@gmail.com> |
|---|---|
| Date | 2022-09-24 17:14 +0200 |
| Message-ID | <tgn6s2$30ka3$1@dont-email.me> |
| In reply to | #86560 |
Am 24.09.2022 um 17:09 schrieb David Brown: > On 24/09/2022 16:25, Bonita Montero wrote: >> Am 24.09.2022 um 16:22 schrieb David Brown: >> >>> There are a number of ways to implement atomics that are larger than >>> a single bus operation can handle, or that involve multiple bus >>> operations. ... >> >> It's always done by STM, and that's slow. > > No, it is not. ... I checked the disassembly of atomics beyond native types with MSVC, clang Windows / Linux and g++ - you don't. If you don't use STM you use the kernel and the atomic operation never fails. Try that out yourself: atomics beyond native types can fail, so you have STM.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-09-26 17:34 -0700 |
| Message-ID | <tgtgeb$3rt9o$1@dont-email.me> |
| In reply to | #86561 |
On 9/24/2022 8:14 AM, Bonita Montero wrote: > Am 24.09.2022 um 17:09 schrieb David Brown: >> On 24/09/2022 16:25, Bonita Montero wrote: >>> Am 24.09.2022 um 16:22 schrieb David Brown: >>> >>>> There are a number of ways to implement atomics that are larger than >>>> a single bus operation can handle, or that involve multiple bus >>>> operations. ... >>> >>> It's always done by STM, and that's slow. >> >> No, it is not. ... > > I checked the disassembly of atomics beyond native types with > MSVC, clang Windows / Linux and g++ - you don't. > If you don't use STM you use the kernel and the atomic operation > never fails. Try that out yourself: atomics beyond native types > can fail, so you have STM. > > Huh?
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2022-09-26 07:53 +0000 |
| Message-ID | <tgrlov$1ctg$1@gioia.aioe.org> |
| In reply to | #86554 |
David Brown <david.brown@hesbynett.no> wrote: > There are a number of ways to implement atomics that are larger than a > single bus operation can handle, or that involve multiple bus > operations. Different processor types have different solutions, and > some are optimised or limited to particular setups (such as single > writer, single processor, etc.). Some processors can handle atomic > accesses for sizes that are bigger than native C/C++ types (such as > 128-bit accesses). Some cannot handle atomic writes for the bigger > native types. I think there's a bit of confusion here about what the term "atomic" means. You seem to be talking about a concept of "atomic" with some kind of meaning like "mutual exclusion supported by the CPU itself". That's not what "atomic" means in general, when talking about multithreaded programming. In general "atomic" merely means that the resource in question can only be accessed by one thread at a time (in other words, it implements some sort of mutual exclusion). As a concrete example: POSIX requires that fwrite() be atomic (for a particular FILE object). This means that no two threads can write to the same FILE object with a singular fwrite() call at the same time. In other words, fwrite() implements (at some level) some kind of (per FILE object) mutex. "Atomic" is actually a stronger guarantee than merely "thread-safe". If fwrite() were merely guaranteed to be "thread-safe", it would just mean that it won't break (eg. corrupt its internal state, or any other data anywhere else) if two threads call it at the same time, but it wouldn't guarantee that the data written by those two threads won't be interleaved somehow. However, since fwrite() is "atomic", not just "thread-safe" (if it conforms to POSIX), then it implements a mutex for the entire function call (for that particular FILE object).
[toc] | [prev] | [next] | [standalone]
| From | Bonita Montero <Bonita.Montero@gmail.com> |
|---|---|
| Date | 2022-09-26 10:21 +0200 |
| Message-ID | <tgrndt$3mo17$1@dont-email.me> |
| In reply to | #86595 |
Am 26.09.2022 um 09:53 schrieb Juha Nieminen: > David Brown <david.brown@hesbynett.no> wrote: >> There are a number of ways to implement atomics that are larger than a >> single bus operation can handle, or that involve multiple bus >> operations. Different processor types have different solutions, and >> some are optimised or limited to particular setups (such as single >> writer, single processor, etc.). Some processors can handle atomic >> accesses for sizes that are bigger than native C/C++ types (such as >> 128-bit accesses). Some cannot handle atomic writes for the bigger >> native types. > > I think there's a bit of confusion here about what the term "atomic" > means. You seem to be talking about a concept of "atomic" with some > kind of meaning like "mutual exclusion supported by the CPU itself". You're confused, David not.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2022-09-26 08:34 +0000 |
| Message-ID | <tgro5s$bj0$3@gioia.aioe.org> |
| In reply to | #86602 |
Bonita Montero <Bonita.Montero@gmail.com> wrote: > You're confused, David not. Just go away, asshole.
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2022-09-26 13:39 +0200 |
| Message-ID | <tgs31b$3nogl$1@dont-email.me> |
| In reply to | #86595 |
On 26/09/2022 09:53, Juha Nieminen wrote: > David Brown <david.brown@hesbynett.no> wrote: >> There are a number of ways to implement atomics that are larger than a >> single bus operation can handle, or that involve multiple bus >> operations. Different processor types have different solutions, and >> some are optimised or limited to particular setups (such as single >> writer, single processor, etc.). Some processors can handle atomic >> accesses for sizes that are bigger than native C/C++ types (such as >> 128-bit accesses). Some cannot handle atomic writes for the bigger >> native types. > > I think there's a bit of confusion here about what the term "atomic" > means. You seem to be talking about a concept of "atomic" with some kind > of meaning like "mutual exclusion supported by the CPU itself". > No, that's not what I am saying. > That's not what "atomic" means in general, when talking about > multithreaded programming. In general "atomic" merely means that the > resource in question can only be accessed by one thread at a time > (in other words, it implements some sort of mutual exclusion). And that's not quite right either. "Atomic" means that accesses are indivisible. As many threads as you want can read or write to the data at the same time - the defining feature is that there is no possibility of a partial access succeeding. We've mostly mentioned reads and writes - but more complex transactions can be atomic too, such as increments. The term can also apply to collections of accesses, well-known from the database world. Such atomic transactions need to be built on top of low-level atomic accesses with locks, lock-free algorithms, or more advanced protocols such as software transactional memory. Atomic accesses do not have to be purely hardware implementations, though that is the most efficient - and anything software-based is going to depend on smaller hardware-based atomic accesses. By far the most convenient accesses are when you can read or write the memory with normal memory access instructions, or at most by using things such as a "bus lock prefix" available on some processors. On RISC processors, anything beyond a single read or write of a size handled directly by hardware typically involves load-store-exclusive sequences. When you have to use code sequences for access, then it's common that you end up with mutual exclusion - one thread at a time has access. But it doesn't have to be that way, and different software sequences can be used to optimise different usage patterns. All that matters is that if a read sequence exits happily saying "I've read the data", then the data it read matches exactly the data that some thread wrote at some point. > > As a concrete example: POSIX requires that fwrite() be atomic (for a > particular FILE object). This means that no two threads can write > to the same FILE object with a singular fwrite() call at the same > time. In other words, fwrite() implements (at some level) some kind > of (per FILE object) mutex. > That's at a much higher level than has been under discussion here - but yes, that is applying the same term and guarantees for different purposes. (The "atomic" requirement does not force a mutex, but fwrite() has other guarantees beyond mere atomicity.) > "Atomic" is actually a stronger guarantee than merely "thread-safe". > > If fwrite() were merely guaranteed to be "thread-safe", it would just > mean that it won't break (eg. corrupt its internal state, or any other > data anywhere else) if two threads call it at the same time, but it > wouldn't guarantee that the data written by those two threads won't > be interleaved somehow. > "Thread safe" is not as well-defined a term as "atomic", as far as I see it. > However, since fwrite() is "atomic", not just "thread-safe" (if it > conforms to POSIX), then it implements a mutex for the entire function > call (for that particular FILE object). "Atomic" is not really enough to describe the behaviour of a function like "fwrite", since the function does not act on a single "state". If you have two threads trying to write A and B to the same object simultaneously, atomicity means that a third thread reading the object will see A or B, and never a mixture. It's fine if this is implemented by a write of A then a write of B, a write of B then a write of A, a write of A alone, a write of B alone, a lock blocking the thread then a mix of A, B, C and D that gets sorted into one of A or B before the lock is released, or any other combination. Clearly that is not the behaviour you want from fwrite() - here there should be either A then B, or B then A.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-09-26 17:47 -0700 |
| Message-ID | <tgth7s$3rvll$1@dont-email.me> |
| In reply to | #86611 |
On 9/26/2022 4:39 AM, David Brown wrote:
> On 26/09/2022 09:53, Juha Nieminen wrote:
>> David Brown <david.brown@hesbynett.no> wrote:
>>> There are a number of ways to implement atomics that are larger than a
>>> single bus operation can handle, or that involve multiple bus
>>> operations. Different processor types have different solutions, and
>>> some are optimised or limited to particular setups (such as single
>>> writer, single processor, etc.). Some processors can handle atomic
>>> accesses for sizes that are bigger than native C/C++ types (such as
>>> 128-bit accesses). Some cannot handle atomic writes for the bigger
>>> native types.
>>
>> I think there's a bit of confusion here about what the term "atomic"
>> means. You seem to be talking about a concept of "atomic" with some kind
>> of meaning like "mutual exclusion supported by the CPU itself".
>>
>
> No, that's not what I am saying.
>
>> That's not what "atomic" means in general, when talking about
>> multithreaded programming. In general "atomic" merely means that the
>> resource in question can only be accessed by one thread at a time
>> (in other words, it implements some sort of mutual exclusion).
>
> And that's not quite right either.
>
> "Atomic" means that accesses are indivisible.
Exactly.
> As many threads as you
> want can read or write to the data at the same time - the defining
> feature is that there is no possibility of a partial access succeeding.
[...]
Atomic to me, say a RMW sequence:
<pseudo-code>
________________________
int g_value = 0;
// A RMW operation
int
fetch_add_busted(
int& origin,
) {
int result = origin;
origin = result + 1;
return result;
}
________________________
Well, there are problems with this. Its not atomic, and can give garbage
in multi-threaded environments... So, we can lock it up, using a hashed
based mutex algorithm (hashing on address into a table of mutexes). I
posted one in the past:
<pseudo-code>
________________________
int g_value = 0;
// A RMW operation
int
fetch_add(
int& origin,
) {
hash_lock(&origin);
int result = origin;
origin = result + 1;
hash_unlock(&origin);
return result;
}
________________________
Okay, we are atomic. There are other ways to get this done.
https://en.cppreference.com/w/cpp/atomic/atomic/fetch_add
;^)
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-09-27 14:17 +0000 |
| Message-ID | <SBDYK.231093$51Rb.96992@fx45.iad> |
| In reply to | #86630 |
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>On 9/26/2022 4:39 AM, David Brown wrote:
>// A RMW operation
>int
>fetch_add(
> int& origin,
>) {
> hash_lock(&origin);
> int result = origin;
> origin = result + 1;
> hash_unlock(&origin);
> return result;
>}
>________________________
>
>Okay, we are atomic. There are other ways to get this done.
>
>https://en.cppreference.com/w/cpp/atomic/atomic/fetch_add
GCC has had built-ins to generate atomic accesses (e.g. __sync_fetch_and_add)
for many years now.
On intel/amd these generate lock prefixes, on other architectures
with atomic support (e.g. ARMv8 LDADD, et alia) those instructions
will be generated.
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2022-09-27 18:07 +0200 |
| Message-ID | <tgv74b$3ilt$1@dont-email.me> |
| In reply to | #86648 |
On 27/09/2022 16:17, Scott Lurndal wrote:
> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>> On 9/26/2022 4:39 AM, David Brown wrote:
>
>> // A RMW operation
>> int
>> fetch_add(
>> int& origin,
>> ) {
>> hash_lock(&origin);
>> int result = origin;
>> origin = result + 1;
>> hash_unlock(&origin);
>> return result;
>> }
>> ________________________
>>
>> Okay, we are atomic. There are other ways to get this done.
>>
>> https://en.cppreference.com/w/cpp/atomic/atomic/fetch_add
>
> GCC has had built-ins to generate atomic accesses (e.g. __sync_fetch_and_add)
> for many years now.
>
> On intel/amd these generate lock prefixes, on other architectures
> with atomic support (e.g. ARMv8 LDADD, et alia) those instructions
> will be generated.
For more demanding cases - sizes larger than the hardware supports
directly, or read-write-modify on RISC - gcc uses a library that does
much what Chris has shown here. The locks are simple busy-wait
user-space spin locks on an atomic flag, which are very efficient on
most systems (especially in the common case of no contention).
Unfortunately, this solution is worse than useless in some cases, such
as real-time systems and single-core systems.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-09-27 12:42 -0700 |
| Message-ID | <tgvjmu$4j86$1@dont-email.me> |
| In reply to | #86654 |
On 9/27/2022 9:07 AM, David Brown wrote:
> On 27/09/2022 16:17, Scott Lurndal wrote:
>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>> On 9/26/2022 4:39 AM, David Brown wrote:
>>
>>> // A RMW operation
>>> int
>>> fetch_add(
>>> int& origin,
>>> ) {
>>> hash_lock(&origin);
>>> int result = origin;
>>> origin = result + 1;
>>> hash_unlock(&origin);
>>> return result;
>>> }
>>> ________________________
>>>
>>> Okay, we are atomic. There are other ways to get this done.
>>>
>>> https://en.cppreference.com/w/cpp/atomic/atomic/fetch_add
>>
>> GCC has had built-ins to generate atomic accesses (e.g.
>> __sync_fetch_and_add)
>> for many years now.
>>
>> On intel/amd these generate lock prefixes, on other architectures
>> with atomic support (e.g. ARMv8 LDADD, et alia) those instructions
>> will be generated.
>
> For more demanding cases - sizes larger than the hardware supports
> directly, or read-write-modify on RISC - gcc uses a library that does
> much what Chris has shown here. The locks are simple busy-wait
> user-space spin locks on an atomic flag, which are very efficient on
> most systems (especially in the common case of no contention).
> Unfortunately, this solution is worse than useless in some cases, such
> as real-time systems and single-core systems.
>
Yes. Using an address based hashed locking scheme works just in case the
arch does not support the direct CPU instruction(s) (think CAS vs LL/SC)
for an atomic RMW operation. However, the locking emulation is most
definitely, not ideal. Not lock-free, indeed. When the arch supports it,
the compiler should be using lock-free operations wrt:
https://en.cppreference.com/w/cpp/atomic/atomic/is_lock_free
Using a hashed locking scheme for the atomic fetch-add impl would return
false wrt is_lock_free... Also, I forgot to add the rest of the
fetch-add, wrt the god damn dangling comma. Notice the original pseudo
code I posted upthread?
________________________
// A RMW operation
int
fetch_add_busted(
int& origin,
) {
int result = origin;
origin = result + 1;
return result;
}
________________________
Here as well:
________________________
// A RMW operation
int
fetch_add(
int& origin,
) {
hash_lock(&origin);
int result = origin;
origin = result + 1;
hash_unlock(&origin);
return result;
}
________________________
Humm... WTF? Let me correct them:
________________________
// A RMW operation
int
fetch_add_busted(
int& origin,
int addend
) {
int result = origin;
origin = result + addend;
return result;
}
________________________
Corrected here as well:
________________________
// A RMW operation
int
fetch_add(
int& origin,
int addend
) {
hash_lock(&origin);
int result = origin;
origin = result + addend;
hash_unlock(&origin);
return result;
}
________________________
Sorry about that non-sense David: Wrt the dangling comma. Forgot to
introduce the addend for the fetch-add RMW operation.
Shit happens. :^)
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2022-09-28 09:32 +0200 |
| Message-ID | <th0t9s$alqp$1@dont-email.me> |
| In reply to | #86658 |
On 27/09/2022 21:42, Chris M. Thomasson wrote:
> On 9/27/2022 9:07 AM, David Brown wrote:
>> On 27/09/2022 16:17, Scott Lurndal wrote:
>>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>>> On 9/26/2022 4:39 AM, David Brown wrote:
>>>
>>>> // A RMW operation
>>>> int
>>>> fetch_add(
>>>> int& origin,
>>>> ) {
>>>> hash_lock(&origin);
>>>> int result = origin;
>>>> origin = result + 1;
>>>> hash_unlock(&origin);
>>>> return result;
>>>> }
>>>> ________________________
>>>>
>>>> Okay, we are atomic. There are other ways to get this done.
>>>>
>>>> https://en.cppreference.com/w/cpp/atomic/atomic/fetch_add
>>>
>>> GCC has had built-ins to generate atomic accesses (e.g.
>>> __sync_fetch_and_add)
>>> for many years now.
>>>
>>> On intel/amd these generate lock prefixes, on other architectures
>>> with atomic support (e.g. ARMv8 LDADD, et alia) those instructions
>>> will be generated.
>>
>> For more demanding cases - sizes larger than the hardware supports
>> directly, or read-write-modify on RISC - gcc uses a library that does
>> much what Chris has shown here. The locks are simple busy-wait
>> user-space spin locks on an atomic flag, which are very efficient on
>> most systems (especially in the common case of no contention).
>> Unfortunately, this solution is worse than useless in some cases, such
>> as real-time systems and single-core systems.
>>
>
> Yes. Using an address based hashed locking scheme works just in case the
> arch does not support the direct CPU instruction(s) (think CAS vs LL/SC)
> for an atomic RMW operation.
LL/SC /is/ a locking scheme - using a hardware lock. And neither CAS
nor LL/SC work for RMW or even plain write operations that are bigger
than the processor can handle in a single write action.
> However, the locking emulation is most
> definitely, not ideal. Not lock-free, indeed.
Processors can generally handle lock-free atomic access of a single
object of limited size - usually the natural width for the processor.
Some processors have instructions for double-width atomic accesses (such
as a double compare-and-swap). And sometimes instruction sequences,
such as LL/SC with loops, are needed - especially for RMW.
Lock-free algorithms beyond that are for specific data structures. You
can't make lock-free atomic access to a 32 byte object. You either have
to use locks (as will be done with a std::atomic<> for the type, or
using the C11 _Atomic qualifier). If you want lock-free access, you
have to wrap it all up in a more advanced structure, using something
like a lock-free atomic pointer to the "current" version of the data
allocated on a heap.
>
> Sorry about that non-sense David: Wrt the dangling comma. Forgot to
> introduce the addend for the fetch-add RMW operation.
>
> Shit happens. :^)
>
That's just minor detail, so not a problem at all.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-09-28 13:34 -0700 |
| Message-ID | <th2b5h$eg6k$1@dont-email.me> |
| In reply to | #86665 |
On 9/28/2022 12:32 AM, David Brown wrote:
> On 27/09/2022 21:42, Chris M. Thomasson wrote:
>> On 9/27/2022 9:07 AM, David Brown wrote:
>>> On 27/09/2022 16:17, Scott Lurndal wrote:
>>>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>>>> On 9/26/2022 4:39 AM, David Brown wrote:
>>>>
>>>>> // A RMW operation
>>>>> int
>>>>> fetch_add(
>>>>> int& origin,
>>>>> ) {
>>>>> hash_lock(&origin);
>>>>> int result = origin;
>>>>> origin = result + 1;
>>>>> hash_unlock(&origin);
>>>>> return result;
>>>>> }
>>>>> ________________________
>>>>>
>>>>> Okay, we are atomic. There are other ways to get this done.
>>>>>
>>>>> https://en.cppreference.com/w/cpp/atomic/atomic/fetch_add
>>>>
>>>> GCC has had built-ins to generate atomic accesses (e.g.
>>>> __sync_fetch_and_add)
>>>> for many years now.
>>>>
>>>> On intel/amd these generate lock prefixes, on other architectures
>>>> with atomic support (e.g. ARMv8 LDADD, et alia) those instructions
>>>> will be generated.
>>>
>>> For more demanding cases - sizes larger than the hardware supports
>>> directly, or read-write-modify on RISC - gcc uses a library that does
>>> much what Chris has shown here. The locks are simple busy-wait
>>> user-space spin locks on an atomic flag, which are very efficient on
>>> most systems (especially in the common case of no contention).
>>> Unfortunately, this solution is worse than useless in some cases,
>>> such as real-time systems and single-core systems.
>>>
>>
>> Yes. Using an address based hashed locking scheme works just in case
>> the arch does not support the direct CPU instruction(s) (think CAS vs
>> LL/SC) for an atomic RMW operation.
>
> LL/SC /is/ a locking scheme - using a hardware lock. And neither CAS
> nor LL/SC work for RMW or even plain write operations that are bigger
> than the processor can handle in a single write action.
Correct. Imvho, the hardware itself is a _lot_ more efficient at these
types of things... Agreed in a sense? I actually prefer pessimistic CAS
over optimistic primitives like LL/SC. Iirc, a LL/SC can fail just by
reading from the reservation granule. Let alone writing to it... PPC had
a special section in its docs that explain the possible issue of a live
lock. Iirc, even CAS has some special logic in the processor that can
actually assert a bus lock.
>> However, the locking emulation is most definitely, not ideal. Not
>> lock-free, indeed.
>
> Processors can generally handle lock-free atomic access of a single
> object of limited size - usually the natural width for the processor.
> Some processors have instructions for double-width atomic accesses (such
> as a double compare-and-swap). And sometimes instruction sequences,
> such as LL/SC with loops, are needed - especially for RMW.
Afaict, DWCAS is there to help get around the ABA problem ala IBM sysv
appendix, oh shit, I forgot the appendix number. I used to know it,
decades ago. I will try to find it.
> Lock-free algorithms beyond that are for specific data structures. You
> can't make lock-free atomic access to a 32 byte object. You either have
> to use locks (as will be done with a std::atomic<> for the type, or
> using the C11 _Atomic qualifier). If you want lock-free access, you
> have to wrap it all up in a more advanced structure, using something
> like a lock-free atomic pointer to the "current" version of the data
> allocated on a heap.
Agreed. Although, I have created lock-free allocators that never used
dynamic memory, believe it or not. Everything exists on threads stacks.
And memory from thread A could be "freed" by another thread. I remember
a project I had to do for a Quadros based system. Completely based on
stacks. Wow, what a time.
>> Sorry about that non-sense David: Wrt the dangling comma. Forgot to
>> introduce the addend for the fetch-add RMW operation.
>>
>> Shit happens. :^)
>>
>
> That's just minor detail, so not a problem at all.
Thanks. :^)
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-09-28 21:11 +0000 |
| Message-ID | <IL2ZK.114890$6gz7.16864@fx37.iad> |
| In reply to | #86688 |
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes: >On 9/28/2022 12:32 AM, David Brown wrote: >> On 27/09/2022 21:42, Chris M. Thomasson wrote: >>> Yes. Using an address based hashed locking scheme works just in case >>> the arch does not support the direct CPU instruction(s) (think CAS vs >>> LL/SC) for an atomic RMW operation. >> >> LL/SC /is/ a locking scheme - using a hardware lock. And neither CAS >> nor LL/SC work for RMW or even plain write operations that are bigger >> than the processor can handle in a single write action. > >Correct. Imvho, the hardware itself is a _lot_ more efficient at these >types of things... Agreed in a sense? I actually prefer pessimistic CAS >over optimistic primitives like LL/SC. Iirc, a LL/SC can fail just by >reading from the reservation granule. Let alone writing to it... PPC had >a special section in its docs that explain the possible issue of a live >lock. Iirc, even CAS has some special logic in the processor that can >actually assert a bus lock. When ARM was designing their 64-bit architecture (ARMv8) circa 2011/2, they only provided a LL/SC equivalent (load-exclusive/store-exclusive). Their architecture partners at the time quickly requested support for real RMW atomics, which were added as part of the LSE (Large System ISA Extensions). LDADD, LDCLR (and with complement), LDSET (or), LDEOR (xor), LDSMAX (signed maximum), LDUMAX (unsigned maximum), LDSMIN, LDUMIN. The processor fabric forwards the operation to the point of coherency (e.g. the L2/LLC) for cachable memory locations and to the endpoint for uncachable memory locations (e.g. a PCIexpress or CXL endpoint).
[toc] | [prev] | [next] | [standalone]
| From | Bonita Montero <Bonita.Montero@gmail.com> |
|---|---|
| Date | 2022-10-01 06:26 +0200 |
| Message-ID | <th8fih$18v7q$1@dont-email.me> |
| In reply to | #86689 |
Am 28.09.2022 um 23:11 schrieb Scott Lurndal: > When ARM was designing their 64-bit architecture (ARMv8) circa 2011/2, they > only provided a LL/SC equivalent (load-exclusive/store-exclusive). Their > architecture partners at the time quickly requested support for > real RMW atomics, which were added as part of the LSE (Large System ISA > Extensions). LDADD, LDCLR (and with complement), LDSET (or), LDEOR (xor), > LDSMAX (signed maximum), LDUMAX (unsigned maximum), LDSMIN, LDUMIN. Eh, RMW can be emulated with LL/SC but not vice versa. A CAS emulated by LL/SC isn't slower than a native CAS. But atomic increments, decrements, ands, ors or whatever ebulated with LL/SC is sometimes slower.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-10-01 17:01 +0000 |
| Message-ID | <Yn_ZK.97606$chF5.85448@fx08.iad> |
| In reply to | #86747 |
Bonita Montero <Bonita.Montero@gmail.com> writes: >Am 28.09.2022 um 23:11 schrieb Scott Lurndal: > >> When ARM was designing their 64-bit architecture (ARMv8) circa 2011/2, they >> only provided a LL/SC equivalent (load-exclusive/store-exclusive). Their >> architecture partners at the time quickly requested support for >> real RMW atomics, which were added as part of the LSE (Large System ISA >> Extensions). LDADD, LDCLR (and with complement), LDSET (or), LDEOR (xor), >> LDSMAX (signed maximum), LDUMAX (unsigned maximum), LDSMIN, LDUMIN. > >Eh, RMW can be emulated with LL/SC but not vice versa. >A CAS emulated by LL/SC isn't slower than a native CAS. >But atomic increments, decrements, ands, ors or whatever >ebulated with LL/SC is sometimes slower. > Who said anything about CAS[*]? [*] For your edification, CAS on modern archtitectures isn't handled by the CPU, but rather by the point of coherency (LLC or PCI-Express/CXL endpoint). Something you can't do with LL/SC at all.
[toc] | [prev] | [next] | [standalone]
| From | Michael S <already5chosen@yahoo.com> |
|---|---|
| Date | 2022-10-02 15:32 -0700 |
| Message-ID | <837042c7-ced4-4fd6-92c4-8952d2d4ed07n@googlegroups.com> |
| In reply to | #86754 |
On Saturday, October 1, 2022 at 8:02:05 PM UTC+3, Scott Lurndal wrote: > Bonita Montero <Bonita....@gmail.com> writes: > >Am 28.09.2022 um 23:11 schrieb Scott Lurndal: > > > >> When ARM was designing their 64-bit architecture (ARMv8) circa 2011/2, they > >> only provided a LL/SC equivalent (load-exclusive/store-exclusive). Their > >> architecture partners at the time quickly requested support for > >> real RMW atomics, which were added as part of the LSE (Large System ISA > >> Extensions). LDADD, LDCLR (and with complement), LDSET (or), LDEOR (xor), > >> LDSMAX (signed maximum), LDUMAX (unsigned maximum), LDSMIN, LDUMIN. > > > >Eh, RMW can be emulated with LL/SC but not vice versa. > >A CAS emulated by LL/SC isn't slower than a native CAS. > >But atomic increments, decrements, ands, ors or whatever > >ebulated with LL/SC is sometimes slower. > > > Who said anything about CAS[*]? > > > [*] For your edification, CAS on modern archtitectures isn't > handled by the CPU, but rather by the point of coherency (LLC > or PCI-Express/CXL endpoint). I don't think so. IMHO, a typical implementation is that CPU acquires the ownership of location and refuses all attempts by other agents to take it back until both parts of CAS are completed and committed to L1$. What you say is an idea that floats widely but never implemented on general-purpose CPUs. Mostly, because for workloads that run on general-purpose CPUs, it's a very bad idea. May be, on some network processor it works the way, you suggest, but I wouldn't call architectures of these processors "modern". > Something you can't do with LL/SC at all.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-10-03 14:01 +0000 |
| Message-ID | <XWB_K.233427$51Rb.232244@fx45.iad> |
| In reply to | #86772 |
Michael S <already5chosen@yahoo.com> writes: >On Saturday, October 1, 2022 at 8:02:05 PM UTC+3, Scott Lurndal wrote: >> Bonita Montero <Bonita....@gmail.com> writes: >> >Am 28.09.2022 um 23:11 schrieb Scott Lurndal: >> > >> >> When ARM was designing their 64-bit architecture (ARMv8) circa 2011/2, they >> >> only provided a LL/SC equivalent (load-exclusive/store-exclusive). Their >> >> architecture partners at the time quickly requested support for >> >> real RMW atomics, which were added as part of the LSE (Large System ISA >> >> Extensions). LDADD, LDCLR (and with complement), LDSET (or), LDEOR (xor), >> >> LDSMAX (signed maximum), LDUMAX (unsigned maximum), LDSMIN, LDUMIN. >> > >> >Eh, RMW can be emulated with LL/SC but not vice versa. >> >A CAS emulated by LL/SC isn't slower than a native CAS. >> >But atomic increments, decrements, ands, ors or whatever >> >ebulated with LL/SC is sometimes slower. >> > >> Who said anything about CAS[*]? >> >> >> [*] For your edification, CAS on modern archtitectures isn't >> handled by the CPU, but rather by the point of coherency (LLC >> or PCI-Express/CXL endpoint). > >I don't think so. The processors that we build implement it exactly as I say. >IMHO, a typical implementation is that CPU acquires the ownership >of location and refuses all attempts by other agents to take it back >until both parts of CAS are completed and committed to L1$. That hasn't been my experience (in the past with Intel/AMD x86_64 where we extended the coherency protocol from HT/QPI over a high speed fabric (IB/10Ge), and currently with custom high-end ARM64 processors that my CPOE sells). Note that I said at the "point of coherency". That may very well be the L1 cache for some processors, the L2/LLC for others (depending on cache inclusivity etc.).
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-10-03 13:49 -0700 |
| Message-ID | <thfhtg$2ahuv$2@dont-email.me> |
| In reply to | #86783 |
On 10/3/2022 7:01 AM, Scott Lurndal wrote: > Michael S <already5chosen@yahoo.com> writes: >> On Saturday, October 1, 2022 at 8:02:05 PM UTC+3, Scott Lurndal wrote: >>> Bonita Montero <Bonita....@gmail.com> writes: >>>> Am 28.09.2022 um 23:11 schrieb Scott Lurndal: >>>> >>>>> When ARM was designing their 64-bit architecture (ARMv8) circa 2011/2, they >>>>> only provided a LL/SC equivalent (load-exclusive/store-exclusive). Their >>>>> architecture partners at the time quickly requested support for >>>>> real RMW atomics, which were added as part of the LSE (Large System ISA >>>>> Extensions). LDADD, LDCLR (and with complement), LDSET (or), LDEOR (xor), >>>>> LDSMAX (signed maximum), LDUMAX (unsigned maximum), LDSMIN, LDUMIN. >>>> >>>> Eh, RMW can be emulated with LL/SC but not vice versa. >>>> A CAS emulated by LL/SC isn't slower than a native CAS. >>>> But atomic increments, decrements, ands, ors or whatever >>>> ebulated with LL/SC is sometimes slower. >>>> >>> Who said anything about CAS[*]? >>> >>> >>> [*] For your edification, CAS on modern archtitectures isn't >>> handled by the CPU, but rather by the point of coherency (LLC >>> or PCI-Express/CXL endpoint). >> >> I don't think so. > > The processors that we build implement it exactly as I say. > >> IMHO, a typical implementation is that CPU acquires the ownership >> of location and refuses all attempts by other agents to take it back >> until both parts of CAS are completed and committed to L1$. > > That hasn't been my experience (in the past with Intel/AMD x86_64 > where we extended the coherency protocol from HT/QPI over a high > speed fabric (IB/10Ge), and currently with custom high-end ARM64 > processors that my CPOE sells). > > Note that I said at the "point of coherency". That may very > well be the L1 cache for some processors, the L2/LLC for others > (depending on cache inclusivity etc.). Correct.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2022-10-02 12:46 -0700 |
| Message-ID | <thcpqr$1q2fa$3@dont-email.me> |
| In reply to | #86747 |
On 9/30/2022 9:26 PM, Bonita Montero wrote: > Am 28.09.2022 um 23:11 schrieb Scott Lurndal: > >> When ARM was designing their 64-bit architecture (ARMv8) circa 2011/2, >> they >> only provided a LL/SC equivalent (load-exclusive/store-exclusive). >> Their >> architecture partners at the time quickly requested support for >> real RMW atomics, which were added as part of the LSE (Large System ISA >> Extensions). LDADD, LDCLR (and with complement), LDSET (or), LDEOR >> (xor), >> LDSMAX (signed maximum), LDUMAX (unsigned maximum), LDSMIN, LDUMIN. > > Eh, RMW can be emulated with LL/SC but not vice versa. > A CAS emulated by LL/SC isn't slower than a native CAS. Using LL/SC can be tricky. You really need to isolate the reservation granule... > But atomic increments, decrements, ands, ors or whatever > ebulated with LL/SC is sometimes slower. How many times do you spin on a SC failure before you get, pissed off?
[toc] | [prev] | [next] | [standalone]
Page 2 of 5 — ← Prev page 1 [2] 3 4 5 Next page →
Back to top | Article view | comp.lang.c++
csiph-web