Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.arch.embedded > #30874

Re: Multithreaded disk access

From Don Y <blockedofcourse@foo.invalid>
Newsgroups comp.arch.embedded
Subject Re: Multithreaded disk access
Date 2021-10-25 22:29 -0700
Organization A noiseless patient Spider
Message-ID <sl83op$per$1@dont-email.me> (permalink)
References (4 earlier) <skk2fr$ooe$1@dont-email.me> <skk753$tmp$1@dont-email.me> <skkmg9$hl3$1@dont-email.me> <sl63sc$olb$1@dont-email.me> <2UHdJ.2$452.1@fx22.iad>

Show all headers | View raw


On 10/25/2021 5:23 PM, Richard Damon wrote:
> On 10/25/21 7:19 AM, Don Y wrote:
>> On 10/18/2021 1:46 PM, Don Y wrote:
>>>> I think this will give you plenty of an idea how to go about it.
>>>> Once you know the limit you can run at some reasonable figure
>>>> below it and be happy. Getting more precise figures about all
>>>> that is neither easy nor will it buy you anything.
>>>
>>> I suspect "1" is going to end up as the "best compromise".  So,
>>> I'm treating this as an exercise in *validating* that assumption.
>>> I'll see if I can salvage some of the performance monitoring code
>>> from the sanitizer to give me details from which I might be able
>>> to ferret out "opportunities".  If I start by restricting my
>>> observations to non-destructive synthetic loads, then I can
>>> pull a drive and see how it fares in a different host while
>>> running the same code, etc.
>>
>> Actually, '2' turns out to be marginally better than '1'.
>> Beyond that, its hard to generalize without controlling some
>> of the other variables.
>>
>> '2' wins because there is always the potential to make the
>> disk busy, again, just after it satisfies the access for
>> the 1st thread (which is now busy using the data, etc.)
>>
>> But, if the first thread finishes up before the second thread's
>> request has been satisfied, then the presence of a THIRD thread
>> would just be clutter.  (i.e., the work performed correlates
>> with the number of threads that can have value)
> 
> Actually, 2 might be slower than 1, because the new request from the second 
> thread is apt to need a seek, while a single thread making all the calls is 
> more apt to sequentially read much more of the disk.

Second thread is another instance of first thread; same code, same data.
So, it is just "looking ahead" -- i.e., it WILL do what the first thread
WOULD do if the second thread hadn't gotten to it, first!  The strategy
is predicated on tightly coupled actors so they each can see what
the other has done/is doing.

If the layout of the object that thread 1 is processing calls for
a seek, then thread 2 will perform that seek as if thread 1 were to
do it "when done chewing on the past data".

If the disk caches data, then thread 2 will reap the benefits of
that cache AS IF it was thread 1 acting later.

Etc.

A second thread just hides the cost of the data processing.

> The controller, if not given a new chained command, might choose to 
> automatically start reading the next sector of the cylinder, which could be 
> likely the next one asked for.
> 
> The real optimum is likely a single process doing asynchronous requests queuing 
> up a series of requests, and then distributing the data as it comes in to 
> processing threads to do what ever crunching needs to be done.
> 
> These threads then send the data to a single thread that does asynchronous 
> writes of the results.
> 
> You can easily tell if the input or output processes are I/O bound or not and 
> use that to adjust the number of crunching threads in the middle.
> 

Back to comp.arch.embedded | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread


Thread

Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 07:08 -0700
  Re: Multithreaded disk access Richard Damon <Richard@Damon-Family.org> - 2021-10-15 11:38 -0400
    Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 09:00 -0700
      Re: Multithreaded disk access Richard Damon <Richard@Damon-Family.org> - 2021-10-15 12:48 -0400
        Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 11:40 -0700
  Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-15 18:46 +0300
    Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 09:19 -0700
      Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-15 19:28 +0300
  Re: Multithreaded disk access Brett <ggtgp@yahoo.com> - 2021-10-17 20:27 +0000
    Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-17 14:49 -0700
      Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-17 15:01 -0700
      Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-18 01:09 +0300
        Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-18 08:05 -0700
          Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-18 19:25 +0300
            Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-18 13:46 -0700
              Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-19 00:46 +0300
                Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-18 16:17 -0700
                Re: Multithreaded disk access antispam@math.uni.wroc.pl - 2021-10-31 19:21 +0000
                Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-31 21:55 +0200
              Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-25 04:19 -0700
                Re: Multithreaded disk access Richard Damon <Richard@Damon-Family.org> - 2021-10-25 20:23 -0400
                Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-25 22:29 -0700
      Re: Multithreaded disk access Clifford Heath <no.spam@please.net> - 2021-10-18 11:58 +1100

csiph-web