Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.programming.threads > #4042 > unrolled thread

Re: Detecting mutexes owned by dead processes

Started bygee.akyol@gmail.com
First post2018-02-02 13:16 -0800
Last post2018-02-06 00:23 +0000
Articles 3 — 3 participants

Back to article view | Back to comp.programming.threads

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: Detecting mutexes owned by dead processes gee.akyol@gmail.com - 2018-02-02 13:16 -0800
    Re: Detecting mutexes owned by dead processes Kaz Kylheku <217-679-0842@kylheku.com> - 2018-02-02 23:47 +0000
    Re: Detecting mutexes owned by dead processes Steve Watt <steve.removethis@Watt.COM> - 2018-02-06 00:23 +0000

#4042 — Re: Detecting mutexes owned by dead processes

Fromgee.akyol@gmail.com
Date2018-02-02 13:16 -0800
SubjectRe: Detecting mutexes owned by dead processes
Message-ID<8aa2d575-75b9-4bbd-aa75-4dc9605931d2@googlegroups.com>
I have my own lock solution which solves this problem but with some extra
work and ONLY for write locking.  This is so since my structure stores
only the last writer pid & thread in the structure.  Unfortunately, since read
locks are granted to multiple threads and since I do not store all the
threads which have so far acquired the read lock, the mechanism protects
only against write locks.  If one was to also add code to maintain the
entire list of reader threads, it would also work for reader crashes too
but that would make the lock object quite complicated.

I have a mutex and a bunch of counters & thread ids which are in
my own private lock structure.

When a read or write lock is requested, the code checks whether there
is currently a write lock already granted.  If so, AND the pid/thread
which currently has the write lock is NOT the same as the requesting 
pid/thread, there and then a quick check is made using the pid/thread id
of the current writer to see if it is alive.  If so, no problem, the 
write lock is not granted.  But if the pid/thread has already died, then
the lock is granted to the new requester.

Basically, what is being done is detecting whether a pid/thread is alive
only at the instance when another pid/thread requests the write lock.
This limits the liveness checking to be done only when the locks are
required and is not important any other times.

As I said at the beginning, since write lock is granted only to one thread
at a time, it is easy to do.  If it is required that "dead" threads be also
detected when readers are involved, a list of readers have to be maintained.

So make sure you crash only when u have the write lock & not the read lock :-)

[toc] | [next] | [standalone]


#4043

FromKaz Kylheku <217-679-0842@kylheku.com>
Date2018-02-02 23:47 +0000
Message-ID<20180202152326.514@kylheku.com>
In reply to#4042
On 2018-02-02, gee.akyol@gmail.com <gee.akyol@gmail.com> wrote:
>
> I have my own lock solution which solves this problem but with some extra
> work and ONLY for write locking.

I think even POSIX doesn't solve this problem for read-write locks;
the whole "EOWNERDEAD" error check is only specified for mutexes.

This is as good a reason to avoid read-write locks as any, and implement
some sort of process-distributed RCU-like mechanism instead
(in which any necessary locks are process-shared robust mutexes).

The problem then probably reduces to just detecting when there are
processes that died while in a RCU read-side critical section, which
seems simpler due to being batched inside the "synchronize_rcu"
operation rather than tied to a lock. You need just one global table of
processes that are in a read-side or something like that.

RCU means that you need data structures that can be traversed by readers
even while updated; not an easy retrofit for all read-write lock
situations.

[toc] | [prev] | [next] | [standalone]


#4045

FromSteve Watt <steve.removethis@Watt.COM>
Date2018-02-06 00:23 +0000
Message-ID<p5asi1$cfn$1@wattres.Watt.COM>
In reply to#4042
In article <8aa2d575-75b9-4bbd-aa75-4dc9605931d2@googlegroups.com>,
 <gee.akyol@gmail.com> wrote:
[ ... ]
>As I said at the beginning, since write lock is granted only to one thread
>at a time, it is easy to do.  If it is required that "dead" threads be also
>detected when readers are involved, a list of readers have to be maintained.
>
>So make sure you crash only when u have the write lock & not the read lock :-)

And make sure you crash only when you're leaving the data structure in a
state that won't nuke other accessors.  But the owning thread crashed,
so it's a huge gamble about whether the protected data structure is
safe.

Having a shared lock silently succeed when the previous write-locking
owner dies is a terrible idea.  EOWNERDEAD exists to allow code that can
do sanity checks or reconstruction on data structures.

Basically what will happen with your scheme is that some thread will
take the write lock, start manipulating the data structure, and explode
while pointers are still in an inconsistent state.  Then the next (more
innocent) thread will come along, attempt to modify the data structure
in some other way, and also crash.  Lather, rinse, repeat, until all
processes that might update the shared structure have crashed.

Good luck debugging the mess that ensues.
-- 
Steve Watt KD6GGD  PP-ASEL-IA          ICBM: 121W 56' 57.5" / 37N 20' 15.3"
 Internet: steve @ Watt.COM                      Whois: SW32-ARIN
   Free time?  There's no such thing.  It just comes in varying prices...

[toc] | [prev] | [standalone]


Back to top | Article view | comp.programming.threads


csiph-web