Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.programming.threads > #2375 > unrolled thread

Paper about RCU (Read-Copy Update)

Started byaminer <aminer@toto.net>
First post2014-05-26 16:49 -0700
Last post2014-05-29 18:11 -0700
Articles 4 — 2 participants

Back to article view | Back to comp.programming.threads


Contents

  Paper about RCU (Read-Copy Update) aminer <aminer@toto.net> - 2014-05-26 16:49 -0700
    Re: Paper about RCU (Read-Copy Update) aminer <aminer@toto.net> - 2014-05-26 17:14 -0700
    Re: Paper about RCU (Read-Copy Update) "Chris M. Thomasson" <no@spam.invalid> - 2014-05-29 00:41 -0700
      Re: Paper about RCU (Read-Copy Update) "Chris M. Thomasson" <no@spam.invalid> - 2014-05-29 18:11 -0700

#2375 — Paper about RCU (Read-Copy Update)

Fromaminer <aminer@toto.net>
Date2014-05-26 16:49 -0700
SubjectPaper about RCU (Read-Copy Update)
Message-ID<lm09ba$89v$1@news.albasani.net>
Hello,

I have read the following paper about RCU (Read-Copy Update), and as you 
will notice they are testing this RCU implementation  against the 
pthread reader-writer lock, and as you have noticed the pthread 
reader-writer lock doesn't scale well cause the reader side of the 
pthread reader-writer lock is expensive...

Here is the paper:

https://www.efficios.com/pub/rcu/urcu-main.pdf

But as you will notice that  the quiescent-state based reclamation 
(QSBR) and RCU scales very well cause there reader side functions scale 
very well, but don't worry , you don't need the RCU, cause my scalable 
RWLocks are also scaling very well on read-mostly scenarios, why my 
scalable RWLocks are scaling very well ? Cause read the following about 
my LW_RWLock algorithm:

"Notice carefully that my RWLock is scalable cause each element of the 
FCount1^ array resides in a  seperate cache line , hence when i am 
incrementing  with  LockedExchangeAdd(FCount1^[myid].fcount1,1) it's 
scaling, notice also that i am using the following: 
myid:=GetCurrentProcessorNumber so i am puting the processor number 
inside myid variable."

Read this:

http://pages.videotron.com/aminer/rwlock1.html

So as you have noticed in my algorithm, the threads that have the same 
"myid" will have and will increment the "FCount1^[myid].fcount1" in the 
same local cache, so there will be no cache-lines transfers between the 
cores on the reader side of my scalable RWlocks algorithms, this is why 
my RWLock algorithms are scaling very well on read-mostly scenarios.

So hope that you will be happy with my RWLocks algorithms...

You can download my scalable RWLocks from:

https://sites.google.com/site/aminer68/scalable-rwlock



Thank you,
Amine Moulay Ramdane.

















[toc] | [next] | [standalone]


#2376

Fromaminer <aminer@toto.net>
Date2014-05-26 17:14 -0700
Message-ID<lm0aqc$bdf$1@news.albasani.net>
In reply to#2375
Hello,


As you have noticed i have invented four variants of my scalable RWLock, 
the ones that ends with an X in there names are starvation-free the 
others are not..

Why i have decided to come with scalable and starvation-free RWLocks? 
cause you can have frequent reads and infrequent writes but from time to 
time you can have frequent writes, so i think it is important to have a 
scalable and starvation-free RWLock , this is why i have also come up 
with two variants of scalable and starvation-free RWLocks.

My lightweight variant and starvation-free RWLock called LW_RWLockX and 
also my LW_RWLOCK both scales even at 3% of writes, that's also very 
interresting to know...

And you have to know also that even if we are in a scenario with many 
more writes than 3% of writes and my scalable LW_RWLockX don't scale 
globally in the timing you have to know that LW_RWLockX will still be 
useful , cause even if it doen't scale globally in the timing ,  you 
have to understand that if from the time t1 to the time t2 there is 
writes and reads and from the time t2 to time t3 there is only reads, so 
even if the time from t1 to t2 will add more to the overall time cause 
there is writes threads, the perceived throughput from time t2 to t3 
will be higher and the waiting time from t2 to t3 will be lower cause 
all my RWLocks will be scalable from t2 to t3 and this will make all my 
scalable RWLocks useful cause for example the database clients from 
internet or intranet waiting from t2 to t3 will be served more quickly 
and this is still useful and this will make my scalable RWLocks still 
useful even if it doesn't scale above 3% of writes.

I think that my scalable RWLocks algorithms are scalable
as both quiescent-state based reclamation (QSBR) and RCU (Read-Copy Update)


You can download my scalable RWLocks from:

https://sites.google.com/site/aminer68/scalable-rwlock



Thank you,
Amine Moulay Ramdane.

[toc] | [prev] | [next] | [standalone]


#2383

From"Chris M. Thomasson" <no@spam.invalid>
Date2014-05-29 00:41 -0700
Message-ID<lm6ocd$9gc$1@speranza.aioe.org>
In reply to#2375
If you can create a rw-mutex that is not
asymmetrical, aka prior art, that beats a
virtually zero overhead read region...

Well, then I am all ears and eyes!

:^D

[toc] | [prev] | [next] | [standalone]


#2384

From"Chris M. Thomasson" <no@spam.invalid>
Date2014-05-29 18:11 -0700
Message-ID<lm8lso$7e9$1@speranza.aioe.org>
In reply to#2383
> "Chris M. Thomasson"  wrote in message 
> news:lm6ocd$9gc$1@speranza.aioe.org...

> If you can create a rw-mutex that is not
> asymmetrical, aka prior art, that beats a
> virtually zero overhead read region...

> Well, then I am all ears and eyes!

I HIGHLY doubt that using one of your rw-mutexs
will beat a RCU implementation. Have you actually
benchmarked anything? 

[toc] | [prev] | [standalone]


Back to top | Article view | comp.programming.threads


csiph-web