Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.programming.threads > #2515 > unrolled thread

pthread reader-writer mutex...

Started byRamine <ramine@1.1>
First post2014-06-14 18:38 -0700
Last post2014-06-14 18:54 -0700
Articles 2 — 1 participant

Back to article view | Back to comp.programming.threads


Contents

  pthread reader-writer mutex... Ramine <ramine@1.1> - 2014-06-14 18:38 -0700
    Re: pthread reader-writer mutex... Ramine <ramine@1.1> - 2014-06-14 18:54 -0700

#2515 — pthread reader-writer mutex...

FromRamine <ramine@1.1>
Date2014-06-14 18:38 -0700
Subjectpthread reader-writer mutex...
Message-ID<lnitfc$6og$1@dont-email.me>
Hello,


Question:


If you ask me a question such us why do we have to use
such us your scalable RWLocks and not simply use
pthread reader-writer lock ?


Answer:

The Pthread reader-writer lock uses an expensive atomic operation
on the reader side, so when a thread on the reader side executes this 
atomic operation it will generate a cache-line transfer between one core 
to the other, and this cache-line transfer is expensive, so each threads 
on the reader side of the Pthread reader-writer lock
generates a cache-line transfer, so this will make the serial part
of the Amdahl equation much bigger, so if you need to scale your
Pthread reader-writer lock, the parallel part of the reader side
must be much bigger than the time that it takes to transfer a cache-line 
between cores, so since Pthread reader-writer lock don't scale for 
reader sides that are for example smaller than the time that it takes to 
transfer a cache-line between one core to the other , this will make 
Pthread reader-writer not scalable, that not the same for my scalable 
RWLock, my scalable RWLocks scale much better than Pthread reader-writer 
lock even if the time under the reader side is equal or smaller than the 
time that it takes to tranfer a cache-line from one core to the other, 
and for bigger parallel part io have said that
as you have noticed with me that to scale my scalable RWLocks
you have to use a distributed memory configuration like in
NUMA systems, and on the reader and writer part of my scalable RWLocks 
you have to use more expensive processing, that means you have
to transfer data in parallel from the distributed memory to the CPUs,
cause this memory transfers from distributed memory to
the cores is more expensive than the cache-line transfers from one core 
to the other that i am generating inside my scalable RWLocks algorithms, 
so this will make the parallel part much bigger than the serial part so 
this will make my scalable RWLocks algorithms much more scalable, and 
this is good, that's the same for hardisks, you have to distribute your 
hardisks , and use multiple hardisks an access them in parallel,this 
also will make my scalable RWLocks  more scalable, and this is good.

So as you have noticed my scalable RWLocks are still useful
and good.



Thank you,
Amine Moulay Ramdane.

[toc] | [next] | [standalone]


#2516

FromRamine <ramine@1.1>
Date2014-06-14 18:54 -0700
Message-ID<lniuc4$afe$1@dont-email.me>
In reply to#2515
On 6/14/2014 6:38 PM, Ramine wrote:
>
> Hello,
>
>
> Question:
>
>
> If you ask me a question such us why do we have to use
> such us your scalable RWLocks and not simply use
> pthread reader-writer lock ?
>


I correct my english typos:

If you ask me a question such as why do we have to use
such as your scalable RWLocks and not simply use
pthread reader-writer lock ?



>
> Answer:
>
> The Pthread reader-writer lock uses an expensive atomic operation
> on the reader side, so when a thread on the reader side executes this
> atomic operation it will generate a cache-line transfer between one core
> to the other, and this cache-line transfer is expensive, so each threads
> on the reader side of the Pthread reader-writer lock
> generates a cache-line transfer, so this will make the serial part
> of the Amdahl equation much bigger, so if you need to scale your
> Pthread reader-writer lock, the parallel part of the reader side
> must be much bigger than the time that it takes to transfer a cache-line
> between cores, so since Pthread reader-writer lock don't scale for
> reader sides that are for example smaller than the time that it takes to
> transfer a cache-line between one core to the other , this will make
> Pthread reader-writer not scalable, that not the same for my scalable
> RWLock, my scalable RWLocks scale much better than Pthread reader-writer
> lock even if the time under the reader side is equal or smaller than the
> time that it takes to tranfer a cache-line from one core to the other,
> and for bigger parallel part io have said that
> as you have noticed with me that to scale my scalable RWLocks
> you have to use a distributed memory configuration like in
> NUMA systems, and on the reader and writer part of my scalable RWLocks
> you have to use more expensive processing, that means you have
> to transfer data in parallel from the distributed memory to the CPUs,
> cause this memory transfers from distributed memory to
> the cores is more expensive than the cache-line transfers from one core
> to the other that i am generating inside my scalable RWLocks algorithms,
> so this will make the parallel part much bigger than the serial part so
> this will make my scalable RWLocks algorithms much more scalable, and
> this is good, that's the same for hardisks, you have to distribute your
> hardisks , and use multiple hardisks an access them in parallel,this
> also will make my scalable RWLocks  more scalable, and this is good.
>
> So as you have noticed my scalable RWLocks are still useful
> and good.
>
>
>
> Thank you,
> Amine Moulay Ramdane.
>

[toc] | [prev] | [standalone]


Back to top | Article view | comp.programming.threads


csiph-web