Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c++ > #87510 > unrolled thread

An argument *against* (the liberal use of) references

Started byJuha Nieminen <nospam@thanks.invalid>
First post2022-11-22 09:52 +0000
Last post2022-11-23 20:08 +0100
Articles 20 on this page of 58 — 17 participants

Back to article view | Back to comp.lang.c++


Contents

  An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-11-22 09:52 +0000
    Re: An argument *against* (the liberal use of) references "Fred. Zwarts" <F.Zwarts@KVI.nl> - 2022-11-22 11:23 +0100
    Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-11-22 13:27 +0200
      Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-11-23 06:52 +0000
        Re: An argument *against* (the liberal use of) references "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-11-23 13:51 +0100
          Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-11-23 13:46 +0000
            Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-11-30 13:59 +0200
              Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-01 06:45 +0000
                Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 09:19 +0200
                  Re: An argument *against* (the liberal use of) references Stuart Redmann <DerTopper@web.de> - 2022-12-01 14:06 +0100
                    Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 16:45 +0200
                      Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-01 16:02 +0000
                        Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 20:21 +0200
                          Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-01 19:09 +0000
                            Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-02 05:53 -0800
                              Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-02 14:36 +0000
                                Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 12:29 -0800
                                  Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-02 20:55 +0000
                                    Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 12:58 -0800
                                      Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-02 21:18 +0000
                                        Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 14:16 -0800
                                    Re: An argument *against* (the liberal use of) references Branimir Maksimovic <branimir.maksimovic@icloud.com> - 2022-12-02 21:09 +0000
                                      Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 13:22 -0800
                                        Re: An argument *against* (the liberal use of) references Branimir Maksimovic <branimir.maksimovic@icloud.com> - 2022-12-02 21:36 +0000
                                          Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-02 14:13 -0800
                              Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-02 18:33 +0200
                          Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-01 11:59 -0800
                            Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-01 23:53 +0200
                              Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-02 04:11 -0800
                    Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-01 12:21 -0800
                  Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-06 11:36 +0000
                    Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-06 12:26 -0800
                      Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-06 20:39 +0000
                        Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-06 23:53 +0200
                          Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-06 13:59 -0800
                            Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-06 22:04 +0000
                              Re: An argument *against* (the liberal use of) references Öö Tiib <ootiib@hot.ee> - 2022-12-06 21:38 -0800
                                Re: An argument *against* (the liberal use of) references scott@slp53.sl.home (Scott Lurndal) - 2022-12-07 15:22 +0000
                                  Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-07 18:37 +0200
                          Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-06 14:01 -0800
                      Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-07 09:05 +0000
                        Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-07 12:48 -0800
                          Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-07 12:50 -0800
                          Re: An argument *against* (the liberal use of) references Juha Nieminen <nospam@thanks.invalid> - 2022-12-08 07:52 +0000
                            Re: An argument *against* (the liberal use of) references Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-12-08 05:50 -0800
                              Re: An argument *against* (the liberal use of) references Michael S <already5chosen@yahoo.com> - 2022-12-08 13:30 -0800
                                Re: An argument *against* (the liberal use of) references Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-12-09 12:50 -0800
                                  Re: An argument *against* (the liberal use of) references David Brown <david.brown@hesbynett.no> - 2022-12-11 12:18 +0100
                          Re: An argument *against* (the liberal use of) references Paavo Helde <eesnimi@osa.pri.ee> - 2022-12-08 13:38 +0200
                            Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-12-08 16:45 -0800
                    Re: An argument *against* (the liberal use of) references Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-12-09 04:35 -0800
    Re: An argument *against* (the liberal use of) references Bonita Montero <Bonita.Montero@gmail.com> - 2022-11-22 15:36 +0100
    Re: An argument *against* (the liberal use of) references Chris Vine <chris@cvine--nospam--.freeserve.co.uk> - 2022-11-22 15:21 +0000
    Re: An argument *against* (the liberal use of) references Richard Damon <Richard@Damon-Family.org> - 2022-11-22 10:48 -0500
      Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-22 13:18 -0800
    Re: An argument *against* (the liberal use of) references El Jo <giorgio.zoppi@gmail.com> - 2022-11-22 13:41 -0800
      Re: An argument *against* (the liberal use of) references "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-11-22 13:51 -0800
    Re: An argument *against* (the liberal use of) references Bonita Montero <Bonita.Montero@gmail.com> - 2022-11-23 20:08 +0100

Page 1 of 3  [1] 2 3  Next page →


#87510 — An argument *against* (the liberal use of) references

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-11-22 09:52 +0000
SubjectAn argument *against* (the liberal use of) references
Message-ID<tli64v$1k4l$1@gioia.aioe.org>
Recently I made a post about references, about how I think many C++
programmers think of them in the wrong way (ie. they think of them
as being effectively an alternative syntax for pointers, a "more
limited pointer syntax", or "a safer pointer syntax", when in fact
references shouldn't be semantically thought of as pointers at all,
but as aliases for the objects they are referring to). In that thread
some criticism was presented about the use of references, and arguments
for their use.

For the sake of fairness and balance, here's an argument *against*
(the liberal use of) references.

Most C++ programmers have been conditioned to always take larger objects
(and even not so large objects) by const reference parameters in functions,
because that's more efficient. (The heavier an object is to deep-copy,
the more efficient a reference to it becomes, obviously.)

However, not many of them consider the *thread-safety* problem that taking
a parameter by reference introduces. In single-threaded programs this is
rather irrelevant, but multithreaded programming is becoming more and more
common every day. Also, if you are writing a library to be used in programs
out there, you have to consider its thread-safety even if the library itself
doesn't use threads.

What is this thread-safety problem introduced by references (especially
when we are writing a function that takes a parameter by reference)? The
fact that in theory another thread could modify the object being referred
to, at the same time that this function is trying to read it.

And there's nothing this function can do to defend against that. (Unless
this function cooperates with the calling code to make it thread-safe, eg.
by using a mutex offered by the calling code.) Even if the function is
internally thread-safe, it can't help the fact that it's using an external
resource without mutual exclusion (unless provided by the calling code).

How many times have you thought about the fact that taking a parameter
by reference makes the function automatically not-thread-safe? I know
I haven't. Like ever.

And, as has been pointed out, in the calling code itself it's not obvious
that a function is taking a parameter by reference, and thus might need
mutual exclusion.

[toc] | [next] | [standalone]


#87511

From"Fred. Zwarts" <F.Zwarts@KVI.nl>
Date2022-11-22 11:23 +0100
Message-ID<tli7uo$gc3$1@gioia.aioe.org>
In reply to#87510
Op 22.von..2022 om 10:52 schreef Juha Nieminen:
> Recently I made a post about references, about how I think many C++
> programmers think of them in the wrong way (ie. they think of them
> as being effectively an alternative syntax for pointers, a "more
> limited pointer syntax", or "a safer pointer syntax", when in fact
> references shouldn't be semantically thought of as pointers at all,
> but as aliases for the objects they are referring to). In that thread
> some criticism was presented about the use of references, and arguments
> for their use.
> 
> For the sake of fairness and balance, here's an argument *against*
> (the liberal use of) references.
> 
> Most C++ programmers have been conditioned to always take larger objects
> (and even not so large objects) by const reference parameters in functions,
> because that's more efficient. (The heavier an object is to deep-copy,
> the more efficient a reference to it becomes, obviously.)
> 
> However, not many of them consider the *thread-safety* problem that taking
> a parameter by reference introduces. In single-threaded programs this is
> rather irrelevant, but multithreaded programming is becoming more and more
> common every day. Also, if you are writing a library to be used in programs
> out there, you have to consider its thread-safety even if the library itself
> doesn't use threads.
> 
> What is this thread-safety problem introduced by references (especially
> when we are writing a function that takes a parameter by reference)? The
> fact that in theory another thread could modify the object being referred
> to, at the same time that this function is trying to read it.
> 
> And there's nothing this function can do to defend against that. (Unless
> this function cooperates with the calling code to make it thread-safe, eg.
> by using a mutex offered by the calling code.) Even if the function is
> internally thread-safe, it can't help the fact that it's using an external
> resource without mutual exclusion (unless provided by the calling code).
> 
> How many times have you thought about the fact that taking a parameter
> by reference makes the function automatically not-thread-safe? I know
> I haven't. Like ever.
> 
> And, as has been pointed out, in the calling code itself it's not obvious
> that a function is taking a parameter by reference, and thus might need
> mutual exclusion.

I have seen such problems often. When possible, I try to make the class 
itself thread-safe, such that it can only be accessed with thread-safe 
member functions. But sometimes one has to use classes that are not 
thread-safe themselves. Then this is something to think of. Using a copy 
instead of a reference is not always the solution, because making a copy 
is not always thread-safe either.

[toc] | [prev] | [next] | [standalone]


#87512

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-11-22 13:27 +0200
Message-ID<tlibnk$3298$1@dont-email.me>
In reply to#87510
22.11.2022 11:52 Juha Nieminen kirjutas:
> Recently I made a post about references, about how I think many C++
> programmers think of them in the wrong way (ie. they think of them
> as being effectively an alternative syntax for pointers, a "more
> limited pointer syntax", or "a safer pointer syntax", when in fact
> references shouldn't be semantically thought of as pointers at all,
> but as aliases for the objects they are referring to). In that thread
> some criticism was presented about the use of references, and arguments
> for their use.
> 
> For the sake of fairness and balance, here's an argument *against*
> (the liberal use of) references.
> 
> Most C++ programmers have been conditioned to always take larger objects
> (and even not so large objects) by const reference parameters in functions,
> because that's more efficient. (The heavier an object is to deep-copy,
> the more efficient a reference to it becomes, obviously.)
> 
> However, not many of them consider the *thread-safety* problem that taking
> a parameter by reference introduces. In single-threaded programs this is
> rather irrelevant, but multithreaded programming is becoming more and more
> common every day. Also, if you are writing a library to be used in programs
> out there, you have to consider its thread-safety even if the library itself
> doesn't use threads.
> 
> What is this thread-safety problem introduced by references (especially
> when we are writing a function that takes a parameter by reference)? The
> fact that in theory another thread could modify the object being referred
> to, at the same time that this function is trying to read it.
> 
> And there's nothing this function can do to defend against that. (Unless
> this function cooperates with the calling code to make it thread-safe, eg.
> by using a mutex offered by the calling code.) Even if the function is
> internally thread-safe, it can't help the fact that it's using an external
> resource without mutual exclusion (unless provided by the calling code).

If the code passing a reference is not thread-safe, then just changing 
it to pass by value will not magically make it thread-safe. Without 
proper synchronization, the object may be changed by another thread at 
any moment, for example in the middle of the copy operation, and thus 
the copy might become internally inconsistent.

In multithreaded programs typically there are 3 types of objects:

1. Non-mutable shared objects which can be accessed by multiple threads 
without locking. Can be passed by reference.

2. Shared objects which need locking when accessed. The locking will be 
best placed inside the objects methods, so that the callers do not need 
to worry about that. Can be passed by reference. Also lock-free data 
structures would belong here.

3. Single-threaded objects which are accessed and modified in a single 
thread only. Do not need locking, but require deep copying when passed 
to another thread. The copying must happen before the copy will become 
accessible in the other thread. This copying can be indeed done by 
passing the object to a function by value - that's what is done e.g. by 
the std::thread constructor.

In their own thread such objects can be accessed without locking and can 
be passed via references, no problems.

In short, just casual pass-by-value is neither sufficient nor needed for 
multi-thread safety. Copying is needed when passing over single-threaded 
objects to other threads, which ought better to happen in clearly 
defined points in the program.

[toc] | [prev] | [next] | [standalone]


#87522

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-11-23 06:52 +0000
Message-ID<tlkfv5$rss$1@gioia.aioe.org>
In reply to#87512
Paavo Helde <eesnimi@osa.pri.ee> wrote:
> If the code passing a reference is not thread-safe, then just changing 
> it to pass by value will not magically make it thread-safe. Without 
> proper synchronization, the object may be changed by another thread at 
> any moment, for example in the middle of the copy operation, and thus 
> the copy might become internally inconsistent.

I think you have a point there. I was thinking that since the function is
getting a local copy of the object then mutual exclusion problems go away
because that local copy is completely independent of the original.

I didn't think that *copying* the object for the function in itself isn't
automatically thread-safe. Yet still, for a few seconds, I had the strong
instinct that there has to be something "more thread-safe" about making
a copy of the value for the function than have the function take a
reference to it... but the more I try to figure out how, I fail. Copying
merely moves the mutual exclusion problem to a slightly different place,
but it doesn't solve it. The calling code *still* needs to solve the
mutual exclusion problem regardless of which way the function takes the
parameter.

[toc] | [prev] | [next] | [standalone]


#87526

From"Alf P. Steinbach" <alf.p.steinbach@gmail.com>
Date2022-11-23 13:51 +0100
Message-ID<tll502$ce5l$1@dont-email.me>
In reply to#87522
On 23 Nov 2022 07:52, Juha Nieminen wrote:
> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>> If the code passing a reference is not thread-safe, then just changing
>> it to pass by value will not magically make it thread-safe. Without
>> proper synchronization, the object may be changed by another thread at
>> any moment, for example in the middle of the copy operation, and thus
>> the copy might become internally inconsistent.
> 
> I think you have a point there. I was thinking that since the function is
> getting a local copy of the object then mutual exclusion problems go away
> because that local copy is completely independent of the original.
> 
> I didn't think that *copying* the object for the function in itself isn't
> automatically thread-safe. Yet still, for a few seconds, I had the strong
> instinct that there has to be something "more thread-safe" about making
> a copy of the value for the function than have the function take a
> reference to it... but the more I try to figure out how, I fail. Copying
> merely moves the mutual exclusion problem to a slightly different place,
> but it doesn't solve it. The calling code *still* needs to solve the
> mutual exclusion problem regardless of which way the function takes the
> parameter.

Depends on when the copying is done.

Copying to the tread instantiation is safe.


- Alf

[toc] | [prev] | [next] | [standalone]


#87527

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-11-23 13:46 +0000
Message-ID<tll88a$13vs$1@gioia.aioe.org>
In reply to#87526
Alf P. Steinbach <alf.p.steinbach@gmail.com> wrote:
> On 23 Nov 2022 07:52, Juha Nieminen wrote:
>> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>>> If the code passing a reference is not thread-safe, then just changing
>>> it to pass by value will not magically make it thread-safe. Without
>>> proper synchronization, the object may be changed by another thread at
>>> any moment, for example in the middle of the copy operation, and thus
>>> the copy might become internally inconsistent.
>> 
>> I think you have a point there. I was thinking that since the function is
>> getting a local copy of the object then mutual exclusion problems go away
>> because that local copy is completely independent of the original.
>> 
>> I didn't think that *copying* the object for the function in itself isn't
>> automatically thread-safe. Yet still, for a few seconds, I had the strong
>> instinct that there has to be something "more thread-safe" about making
>> a copy of the value for the function than have the function take a
>> reference to it... but the more I try to figure out how, I fail. Copying
>> merely moves the mutual exclusion problem to a slightly different place,
>> but it doesn't solve it. The calling code *still* needs to solve the
>> mutual exclusion problem regardless of which way the function takes the
>> parameter.
> 
> Depends on when the copying is done.
> 
> Copying to the tread instantiation is safe.

If you are calling a function and giving as parameter a variable that may
be modified by another thread, you need to take care of the mutual
exclusion problem regardless of whether that function takes the parameter
by value or by reference.

I originally didn't think about the fact that the function taking it by
value doesn't solve the problem because now copying the value needs the
mutual exclusion. The only thing that happens is that the point of
potential conflict has been moved to a slightly different place.

You could make a local copy of the value before passing it to the
function (which could be quite efficient if the variable is atomic),
but even then it doesn't really matter if the function takes it by
value or by reference. (OTOH this may be much more efficient if
the function takes a significant amount of time because you don't
need to keep the mutex locked for the duration of the function.)

[toc] | [prev] | [next] | [standalone]


#87645

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-11-30 13:59 +0200
Message-ID<tm7gjh$2hau4$1@dont-email.me>
In reply to#87527
23.11.2022 15:46 Juha Nieminen kirjutas:
> Alf P. Steinbach <alf.p.steinbach@gmail.com> wrote:
>> On 23 Nov 2022 07:52, Juha Nieminen wrote:
>>> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>>>> If the code passing a reference is not thread-safe, then just changing
>>>> it to pass by value will not magically make it thread-safe. Without
>>>> proper synchronization, the object may be changed by another thread at
>>>> any moment, for example in the middle of the copy operation, and thus
>>>> the copy might become internally inconsistent.
>>>
>>> I think you have a point there. I was thinking that since the function is
>>> getting a local copy of the object then mutual exclusion problems go away
>>> because that local copy is completely independent of the original.
>>>
>>> I didn't think that *copying* the object for the function in itself isn't
>>> automatically thread-safe. Yet still, for a few seconds, I had the strong
>>> instinct that there has to be something "more thread-safe" about making
>>> a copy of the value for the function than have the function take a
>>> reference to it... but the more I try to figure out how, I fail. Copying
>>> merely moves the mutual exclusion problem to a slightly different place,
>>> but it doesn't solve it. The calling code *still* needs to solve the
>>> mutual exclusion problem regardless of which way the function takes the
>>> parameter.
>>
>> Depends on when the copying is done.
>>
>> Copying to the tread instantiation is safe.
> 
> If you are calling a function and giving as parameter a variable that may
> be modified by another thread, you need to take care of the mutual
> exclusion problem regardless of whether that function takes the parameter
> by value or by reference.
> 
> I originally didn't think about the fact that the function taking it by
> value doesn't solve the problem because now copying the value needs the
> mutual exclusion. The only thing that happens is that the point of
> potential conflict has been moved to a slightly different place.
> 
> You could make a local copy of the value before passing it to the
> function (which could be quite efficient if the variable is atomic),
> but even then it doesn't really matter if the function takes it by
> value or by reference. (OTOH this may be much more efficient if
> the function takes a significant amount of time because you don't
> need to keep the mutex locked for the duration of the function.)

Curiously enough, I just spent 3 days for tracking down a random race 
condition bug in a large application. It appeared that for fixing it I 
had to add a single ampersand character, i.e. instead of making a copy 
of the object I had to just take a reference to it. Note this is the 
exact opposite of the general suggestion you advocated earlier ;-)

Actually once located, the bug was simple. The object which was copied 
was a single-threaded refcounted smartpointer, and by copying it the 
refcounter got incremented (and later decremented). Alas, this was 
accidentally done from parallel threads at the same time, without any 
synchronization, so eventually the refcounter got messed up.

After fixing it by taking a reference to the smartpointer instead of 
copying it, the refcounter now remains constant all the time throughout 
the parallel regime (and all other access is read-only as well), so 
everything now works fine.

In principle one could make copies of the pointed objects before the 
parallel regime, or in this particular case it would have been enough to 
use thread-safe smartpointers, but both these approaches would affect 
the performance, and we are always struggling with the performance.

[toc] | [prev] | [next] | [standalone]


#87661

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-12-01 06:45 +0000
Message-ID<tm9iig$1can$1@gioia.aioe.org>
In reply to#87645
Paavo Helde <eesnimi@osa.pri.ee> wrote:
> Actually once located, the bug was simple. The object which was copied 
> was a single-threaded refcounted smartpointer, and by copying it the 
> refcounter got incremented (and later decremented). Alas, this was 
> accidentally done from parallel threads at the same time, without any 
> synchronization, so eventually the refcounter got messed up.

I think that if the reference count is declared atomic, it can be safely
directly incremented. When decrementing you would need to use the
fetch_sub() function to see if the object needs to be destroyed.

While modifying an atomic might not be equally fast as a non-atomic,
it shouldn't be all that much slower either, at least if the target
architecture supports atomic operations.

[toc] | [prev] | [next] | [standalone]


#87665

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-12-01 09:19 +0200
Message-ID<tm9khr$2oi11$3@dont-email.me>
In reply to#87661
01.12.2022 08:45 Juha Nieminen kirjutas:
> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>> Actually once located, the bug was simple. The object which was copied
>> was a single-threaded refcounted smartpointer, and by copying it the
>> refcounter got incremented (and later decremented). Alas, this was
>> accidentally done from parallel threads at the same time, without any
>> synchronization, so eventually the refcounter got messed up.
> 
> I think that if the reference count is declared atomic, it can be safely
> directly incremented. When decrementing you would need to use the
> fetch_sub() function to see if the object needs to be destroyed.
> 
> While modifying an atomic might not be equally fast as a non-atomic,
> it shouldn't be all that much slower either, at least if the target
> architecture supports atomic operations.

I have pondered this myself. Maybe I should measure the actual slowdown 
after temporarily making the refcounters atomic. But this seems overkill 
because these smartpointers would still point to single-threaded objects 
which are meant to be primarily used in single-thread regime, so in most 
cases making the smartpointers atomic does not buy anything.

When tracking down this bug, I monitored all refcounter changes for a 
particular single smartpointer during the program run (ca 10 min). There 
were 591848 increments and decrements, from which 1526 came from the 
problematic (parallelized) part. It looks like a pessimization to slow 
down 99.75% of accesses when only 0.25% would actually benefit from this.


[toc] | [prev] | [next] | [standalone]


#87666

FromStuart Redmann <DerTopper@web.de>
Date2022-12-01 14:06 +0100
Message-ID<526054851.691592086.700998.DerTopper-web.de@news.eternal-september.org>
In reply to#87665
Paavo Helde <eesnimi@osa.pri.ee> wrote:
> 01.12.2022 08:45 Juha Nieminen kirjutas:
>> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>>> Actually once located, the bug was simple. The object which was copied
>>> was a single-threaded refcounted smartpointer, and by copying it the
>>> refcounter got incremented (and later decremented). Alas, this was
>>> accidentally done from parallel threads at the same time, without any
>>> synchronization, so eventually the refcounter got messed up.
>> 
>> I think that if the reference count is declared atomic, it can be safely
>> directly incremented. When decrementing you would need to use the
>> fetch_sub() function to see if the object needs to be destroyed.
>> 
>> While modifying an atomic might not be equally fast as a non-atomic,
>> it shouldn't be all that much slower either, at least if the target
>> architecture supports atomic operations.
> 
> I have pondered this myself. Maybe I should measure the actual slowdown 
> after temporarily making the refcounters atomic. But this seems overkill 
> because these smartpointers would still point to single-threaded objects 
> which are meant to be primarily used in single-thread regime, so in most 
> cases making the smartpointers atomic does not buy anything.
> 
> When tracking down this bug, I monitored all refcounter changes for a 
> particular single smartpointer during the program run (ca 10 min). There 
> were 591848 increments and decrements, from which 1526 came from the 
> problematic (parallelized) part. It looks like a pessimization to slow 
> down 99.75% of accesses when only 0.25% would actually benefit from this.
> 

600k changes in reference counts look suspicious to me. When you pass a
ref-counted object to a worker thread, there should only be a single change
in the refcount. This would be because inside the worker thread some object
takes (shared) ownership of shared object. If the shared object needs to be
passed to sub-routines, you should pass them as references or plain
pointers (if the subroutine must be able to cope with non-existing
objects). It should be rare occurrence that another object in the worker
thread needs to take ownership of the shared object.

Another thought: if thread-safety is too costly, you could use two smart
pointer classes: thread-safe pointers and forwarding non-thread-safe smart
pointers. The forwarding smart pointers have their own thread-UNsafe
refcount and the thread-safe smart pointer as member.

Regards,
Stuart

[toc] | [prev] | [next] | [standalone]


#87668

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-12-01 16:45 +0200
Message-ID<tmaeli$2qj9s$1@dont-email.me>
In reply to#87666
01.12.2022 15:06 Stuart Redmann kirjutas:
> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>> 01.12.2022 08:45 Juha Nieminen kirjutas:
>>> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>>>> Actually once located, the bug was simple. The object which was copied
>>>> was a single-threaded refcounted smartpointer, and by copying it the
>>>> refcounter got incremented (and later decremented). Alas, this was
>>>> accidentally done from parallel threads at the same time, without any
>>>> synchronization, so eventually the refcounter got messed up.
>>>
>>> I think that if the reference count is declared atomic, it can be safely
>>> directly incremented. When decrementing you would need to use the
>>> fetch_sub() function to see if the object needs to be destroyed.
>>>
>>> While modifying an atomic might not be equally fast as a non-atomic,
>>> it shouldn't be all that much slower either, at least if the target
>>> architecture supports atomic operations.
>>
>> I have pondered this myself. Maybe I should measure the actual slowdown
>> after temporarily making the refcounters atomic. But this seems overkill
>> because these smartpointers would still point to single-threaded objects
>> which are meant to be primarily used in single-thread regime, so in most
>> cases making the smartpointers atomic does not buy anything.
>>
>> When tracking down this bug, I monitored all refcounter changes for a
>> particular single smartpointer during the program run (ca 10 min). There
>> were 591848 increments and decrements, from which 1526 came from the
>> problematic (parallelized) part. It looks like a pessimization to slow
>> down 99.75% of accesses when only 0.25% would actually benefit from this.
>>
> 
> 600k changes in reference counts look suspicious to me. When you pass a
> ref-counted object to a worker thread, there should only be a single change
> in the refcount. This would be because inside the worker thread some object
> takes (shared) ownership of shared object. If the shared object needs to be
> passed to sub-routines, you should pass them as references or plain
> pointers (if the subroutine must be able to cope with non-existing
> objects). It should be rare occurrence that another object in the worker
> thread needs to take ownership of the shared object.

You are right, that's how I fixed the bug (by using a reference). There 
are now 1526 less changes in refcounts ;-)

As for others ~600k changes, these seem legitimate. This is a scripting 
language engine (think something like Python) with complex data 
structures built up via refcounted smartpointers. This particular object 
is apparently used as some default column in data tables. I think it was 
used for 2000 columns in some 4000-column table. And there were many 
tables like that. If you insert the same refcounted vector as a new 
column in a table 2000 times, via a member function having a 
smartpointer parameter, then you already get something like at least 
6000 refcount changes.

> 
> Another thought: if thread-safety is too costly, you could use two smart
> pointer classes: thread-safe pointers and forwarding non-thread-safe smart
> pointers. The forwarding smart pointers have their own thread-UNsafe
> refcount and the thread-safe smart pointer as member.

I tried to measure the impact of std::atomic<int> refcounters and in 
first tests it seems the overhead on x86_64 is zero (with no 
contention). So it seems I could use them without drawbacks, but this 
would not save much because the pointed objects would still be not 
thread-safe. I guess it might work out if I could ensure that all 
objects are physically immutable after some initialization. Hmm, time 
for thoughts.

[toc] | [prev] | [next] | [standalone]


#87670

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-12-01 16:02 +0000
Message-ID<ge4iL.2764$lzK9.854@fx35.iad>
In reply to#87668
Paavo Helde <eesnimi@osa.pri.ee> writes:
>01.12.2022 15:06 Stuart Redmann kirjutas:

 <snip>

>> Another thought: if thread-safety is too costly, you could use two smart
>> pointer classes: thread-safe pointers and forwarding non-thread-safe smart
>> pointers. The forwarding smart pointers have their own thread-UNsafe
>> refcount and the thread-safe smart pointer as member.
>
>I tried to measure the impact of std::atomic<int> refcounters and in 
>first tests it seems the overhead on x86_64 is zero (with no 
>contention). 

Which follows naturally from the fact that the core doing
the atomic access has exclusive access to the cache line
containing the refcounter.   No overhead at all, unless
the ref counter isn't aligned and crosses a cache-line
boundary (or the access is to an uncached memory
range or caching is disabled), in which case the processor
will take a system-wide
lock to perform the operation, which is catastrophic
on systems with large processor counts.

(Note that both Intel and AMD processors will fall back to
the system-wide lock if a cache line is highly contended
after some time period has elapsed in order to make forward
progress.)

If the atomic access is to memory on a CXL.memory device, the
operation will not benefit from local cache line latencies
and the atomicity will be guaranteed by the CXL.memory device
exporting the memory to the host in some implementation defined
manner.

[toc] | [prev] | [next] | [standalone]


#87671

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-12-01 20:21 +0200
Message-ID<tmarc2$2rgl0$1@dont-email.me>
In reply to#87670
01.12.2022 18:02 Scott Lurndal kirjutas:
> Paavo Helde <eesnimi@osa.pri.ee> writes:
>> 01.12.2022 15:06 Stuart Redmann kirjutas:
> 
>   <snip>
> 
>>> Another thought: if thread-safety is too costly, you could use two smart
>>> pointer classes: thread-safe pointers and forwarding non-thread-safe smart
>>> pointers. The forwarding smart pointers have their own thread-UNsafe
>>> refcount and the thread-safe smart pointer as member.
>>
>> I tried to measure the impact of std::atomic<int> refcounters and in
>> first tests it seems the overhead on x86_64 is zero (with no
>> contention).
> 
> Which follows naturally from the fact that the core doing
> the atomic access has exclusive access to the cache line
> containing the refcounter.   No overhead at all, unless

Thanks for the clarifications!

> the ref counter isn't aligned and crosses a cache-line
> boundary (or the access is to an uncached memory
> range or caching is disabled), in which case the processor
> will take a system-wide
> lock to perform the operation, which is catastrophic
> on systems with large processor counts.

It is clear that having a misaligned cross-border atomic would be very 
bad. But what about normal uncached memory ranges, wouldn't these be 
just loaded into the cache, without disturbing other processors, and 
without any "catastrophic" consequences?

[toc] | [prev] | [next] | [standalone]


#87674

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-12-01 19:09 +0000
Message-ID<DZ6iL.6667$jXi9.3798@fx34.iad>
In reply to#87671
Paavo Helde <eesnimi@osa.pri.ee> writes:
>01.12.2022 18:02 Scott Lurndal kirjutas:
>> Paavo Helde <eesnimi@osa.pri.ee> writes:
>>> 01.12.2022 15:06 Stuart Redmann kirjutas:
>> 
>>   <snip>
>> 
>>>> Another thought: if thread-safety is too costly, you could use two smart
>>>> pointer classes: thread-safe pointers and forwarding non-thread-safe smart
>>>> pointers. The forwarding smart pointers have their own thread-UNsafe
>>>> refcount and the thread-safe smart pointer as member.
>>>
>>> I tried to measure the impact of std::atomic<int> refcounters and in
>>> first tests it seems the overhead on x86_64 is zero (with no
>>> contention).
>> 
>> Which follows naturally from the fact that the core doing
>> the atomic access has exclusive access to the cache line
>> containing the refcounter.   No overhead at all, unless
>
>Thanks for the clarifications!
>
>> the ref counter isn't aligned and crosses a cache-line
>> boundary (or the access is to an uncached memory
>> range or caching is disabled), in which case the processor
>> will take a system-wide
>> lock to perform the operation, which is catastrophic
>> on systems with large processor counts.
>
>It is clear that having a misaligned cross-border atomic would be very 
>bad. But what about normal uncached memory ranges, wouldn't these be 
>just loaded into the cache, without disturbing other processors, and 
>without any "catastrophic" consequences?

Generally uncached means that the processor fetches
directly from memory bypassing the cache and never evicting any
lines.   This is an important characteristic for MMIO space
where a read access has a side effect (e.g. reading a UART
Data Register).

It also depends on how the processor atomic instructions are implemented.

In legacy Intel/AMD systems, where the LOCK prefix is being used, the
systemwide lock is the only possibility [*].

For ARM64 with the Large System Extensions (LSE) atomic instructions,
the processor can send the atomic operation to the point of coherency
(either the cache subsystem, or if caching is disabled, to the DRAM
controller or PCI-Express device (PCIe supports atomics) and the
synchronization happens at the "endpoint".

Without support all the way to the memory controller or endpoint,
there is no other way to sychronize all agents accessing the controller
or endpoint without acquiring a global mutex of some sort.

[*] It's been a decade since I worked directly with those processors
    and they may have added support for atomic operations to the
    internal ring or now mesh structures used to communicate between
    the processing elements and the memory controllers and PCI root port
    bridges, in which case, like ARM64, they can push the atomic op
    all the way out to the endpoint/controller.

[toc] | [prev] | [next] | [standalone]


#87684

FromMichael S <already5chosen@yahoo.com>
Date2022-12-02 05:53 -0800
Message-ID<a3cd87b9-578b-4ec9-8111-2f70eb5e985cn@googlegroups.com>
In reply to#87674
On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:
> Paavo Helde <ees...@osa.pri.ee> writes: 
> >01.12.2022 18:02 Scott Lurndal kirjutas: 
> >> Paavo Helde <ees...@osa.pri.ee> writes: 
> >>> 01.12.2022 15:06 Stuart Redmann kirjutas: 
> >> 
> >> <snip> 
> >> 
> >>>> Another thought: if thread-safety is too costly, you could use two smart 
> >>>> pointer classes: thread-safe pointers and forwarding non-thread-safe smart 
> >>>> pointers. The forwarding smart pointers have their own thread-UNsafe 
> >>>> refcount and the thread-safe smart pointer as member. 
> >>> 
> >>> I tried to measure the impact of std::atomic<int> refcounters and in 
> >>> first tests it seems the overhead on x86_64 is zero (with no 
> >>> contention). 
> >> 
> >> Which follows naturally from the fact that the core doing 
> >> the atomic access has exclusive access to the cache line 
> >> containing the refcounter. No overhead at all, unless 
> > 
> >Thanks for the clarifications! 
> > 
> >> the ref counter isn't aligned and crosses a cache-line 
> >> boundary (or the access is to an uncached memory 
> >> range or caching is disabled), in which case the processor 
> >> will take a system-wide 
> >> lock to perform the operation, which is catastrophic 
> >> on systems with large processor counts. 
> > 
> >It is clear that having a misaligned cross-border atomic would be very 
> >bad. But what about normal uncached memory ranges, wouldn't these be 
> >just loaded into the cache, without disturbing other processors, and 
> >without any "catastrophic" consequences?
> Generally uncached means that the processor fetches 
> directly from memory bypassing the cache and never evicting any 
> lines. This is an important characteristic for MMIO space 
> where a read access has a side effect (e.g. reading a UART 
> Data Register). 
> 
> It also depends on how the processor atomic instructions are implemented. 
> 
> In legacy Intel/AMD systems, where the LOCK prefix is being used, the 
> systemwide lock is the only possibility [*]. 
> 

Huh?
There is absolute no relationship between what prefix is used (instruction
encoding issue) and implementation.

> For ARM64 with the Large System Extensions (LSE) atomic instructions, 
> the processor can send the atomic operation to the point of coherency 
> (either the cache subsystem, or if caching is disabled, to the DRAM 
> controller or PCI-Express device (PCIe supports atomics) and the 
> synchronization happens at the "endpoint". 
> 
> Without support all the way to the memory controller or endpoint, 
> there is no other way to sychronize all agents accessing the controller 
> or endpoint without acquiring a global mutex of some sort. 
> 
> [*] It's been a decade since I worked directly with those processors 
> and they may have added support for atomic operations to the 
> internal ring or now mesh structures used to communicate between 
> the processing elements and the memory controllers and PCI root port 
> bridges, in which case, like ARM64, they can push the atomic op 
> all the way out to the endpoint/controller.

x86 atomic operations have global order, which is somewhat stronger 
than "total order" of the rest of normal* x86 stores. The difference is that
unlike "total order" global order makes no exceptions for store-to-load 
forwarding from core's local store queue.

BTW, it means that your claim in post above "no overhead at all unless ..."
is incorrect in the absolute sense. There is an overhead even without "unless".
But the overhead in uncontended case is small - order of dozen or two of CPU
clocks. So, undetectable in Paavo's case of only 1000 updates per second.
For 1M updates per second impact would me detectable with precise time
 measurements and for 100M per second there would be big slowdown.

Last September Intel published this manual:
https://cdrdv2-public.intel.com/671368/architecture-instruction-set-extensions-programming-reference.pdf
Manual contains new atomic instruction AADD/AAND/AOR/AXOR that provide
weaker (WC) ordering in WB memory regions. The manual does not say when
this instructions are going to be implemented nor if they will be implemented at all.
It also does not explain in which situation they are expected to be useful.
However one thing is clear: they will *not* be useful in typical user-mode code
that deals with reference counting.
May be, usable in userland  in extreme fire-and-foget situations like counting
events where counter is updated not too often, but often enough to matter,
typically not from the same core as the last one and read approximately never.

My guess is that this instructions were invented to help Optane DIMMs.
So today, with Optane DIMMs officially dead, it would be logical for Intel to
never implement this strange instructions that could easily lead to 
programmer's mistakes.

[*] normal in this case means WB or UC. For WC stores it's more relaxed.

[toc] | [prev] | [next] | [standalone]


#87686

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-12-02 14:36 +0000
Message-ID<e4oiL.36061$f9D6.33011@fx09.iad>
In reply to#87684
Michael S <already5chosen@yahoo.com> writes:
>On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:
>> Paavo Helde <ees...@osa.pri.ee> writes: 
>> >01.12.2022 18:02 Scott Lurndal kirjutas: 
>> >> Paavo Helde <ees...@osa.pri.ee> writes: 
>> >>> 01.12.2022 15:06 Stuart Redmann kirjutas: 
>> >> 

>> 
>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the 
>> systemwide lock is the only possibility [*]. 
>> 
>
>Huh?
>There is absolute no relationship between what prefix is used (instruction
>encoding issue) and implementation.

The only way to specify an atomic access in those chips was
to use the LOCK prefix (e.g. LOCK ADD generates an atomic
add, et alia).

[toc] | [prev] | [next] | [standalone]


#87695

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-12-02 12:29 -0800
Message-ID<tmdn7j$352d1$1@dont-email.me>
In reply to#87686
On 12/2/2022 6:36 AM, Scott Lurndal wrote:
> Michael S <already5chosen@yahoo.com> writes:
>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:
>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>> 01.12.2022 18:02 Scott Lurndal kirjutas:
>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas:
>>>>>
> 
>>>
>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the
>>> systemwide lock is the only possibility [*].
>>>
>>
>> Huh?
>> There is absolute no relationship between what prefix is used (instruction
>> encoding issue) and implementation.
> 
> The only way to specify an atomic access in those chips was
> to use the LOCK prefix (e.g. LOCK ADD generates an atomic
> add, et alia).
> 

Iirc, XCHG has an implicit LOCK prefix?

[toc] | [prev] | [next] | [standalone]


#87696

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-12-02 20:55 +0000
Message-ID<iDtiL.87934$Q0m1.45700@fx18.iad>
In reply to#87695
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>On 12/2/2022 6:36 AM, Scott Lurndal wrote:
>> Michael S <already5chosen@yahoo.com> writes:
>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:
>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas:
>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas:
>>>>>>
>> 
>>>>
>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the
>>>> systemwide lock is the only possibility [*].
>>>>
>>>
>>> Huh?
>>> There is absolute no relationship between what prefix is used (instruction
>>> encoding issue) and implementation.
>> 
>> The only way to specify an atomic access in those chips was
>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic
>> add, et alia).
>> 
>
>Iirc, XCHG has an implicit LOCK prefix?

Not sure about XCHG, but CMPXCHG requires the prefix
when used in a multiprocessor system; on a uniprocessor
it will be atomic because interrupts are always taken
between instructions (unlike, the VAX, for instance,
where certain instructions (MOVC3/5) can be interrupted
and restarted).    I suspect that XCHG has similar
characteristics.

[toc] | [prev] | [next] | [standalone]


#87697

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2022-12-02 12:58 -0800
Message-ID<tmdott$352d1$2@dont-email.me>
In reply to#87696
On 12/2/2022 12:55 PM, Scott Lurndal wrote:
> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>> On 12/2/2022 6:36 AM, Scott Lurndal wrote:
>>> Michael S <already5chosen@yahoo.com> writes:
>>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:
>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas:
>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas:
>>>>>>>
>>>
>>>>>
>>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the
>>>>> systemwide lock is the only possibility [*].
>>>>>
>>>>
>>>> Huh?
>>>> There is absolute no relationship between what prefix is used (instruction
>>>> encoding issue) and implementation.
>>>
>>> The only way to specify an atomic access in those chips was
>>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic
>>> add, et alia).
>>>
>>
>> Iirc, XCHG has an implicit LOCK prefix?
> 
> Not sure about XCHG, but CMPXCHG requires the prefix
> when used in a multiprocessor system; on a uniprocessor
> it will be atomic because interrupts are always taken
> between instructions (unlike, the VAX, for instance,
> where certain instructions (MOVC3/5) can be interrupted
> and restarted).    I suspect that XCHG has similar
> characteristics.
> 
> 

Iirc, XCHG is the _only_ atomic RMW instruction that has an _implicit_ 
LOCK prefix. CMPXCHG _needs_ the programmer to put in a LOCK prefix.

[toc] | [prev] | [next] | [standalone]


#87699

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-12-02 21:18 +0000
Message-ID<YYtiL.88375$Q0m1.65033@fx18.iad>
In reply to#87697
"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>On 12/2/2022 12:55 PM, Scott Lurndal wrote:
>> "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> writes:
>>> On 12/2/2022 6:36 AM, Scott Lurndal wrote:
>>>> Michael S <already5chosen@yahoo.com> writes:
>>>>> On Thursday, December 1, 2022 at 9:09:39 PM UTC+2, Scott Lurndal wrote:
>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>> 01.12.2022 18:02 Scott Lurndal kirjutas:
>>>>>>>> Paavo Helde <ees...@osa.pri.ee> writes:
>>>>>>>>> 01.12.2022 15:06 Stuart Redmann kirjutas:
>>>>>>>>
>>>>
>>>>>>
>>>>>> In legacy Intel/AMD systems, where the LOCK prefix is being used, the
>>>>>> systemwide lock is the only possibility [*].
>>>>>>
>>>>>
>>>>> Huh?
>>>>> There is absolute no relationship between what prefix is used (instruction
>>>>> encoding issue) and implementation.
>>>>
>>>> The only way to specify an atomic access in those chips was
>>>> to use the LOCK prefix (e.g. LOCK ADD generates an atomic
>>>> add, et alia).
>>>>
>>>
>>> Iirc, XCHG has an implicit LOCK prefix?
>> 
>> Not sure about XCHG, but CMPXCHG requires the prefix
>> when used in a multiprocessor system; on a uniprocessor
>> it will be atomic because interrupts are always taken
>> between instructions (unlike, the VAX, for instance,
>> where certain instructions (MOVC3/5) can be interrupted
>> and restarted).    I suspect that XCHG has similar
>> characteristics.
>> 
>> 
>
>Iirc, XCHG is the _only_ atomic RMW instruction that has an _implicit_ 
>LOCK prefix. CMPXCHG _needs_ the programmer to put in a LOCK prefix.

Yes, that is the case

XCHG:
  "If a memory operand is referenced, the processor's locking protocol is automatically
   implemented for the duration of the exchange operation, regardless of the presence
   or absence of the LOCK prefix or of the value of the IOPL. (See the LOCK prefix
   description in this chapter for more information on the locking protocol.)

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | comp.lang.c++


csiph-web