Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.arch.embedded > #30816

Re: Multithreaded disk access

From Dimiter_Popoff <dp@tgi-sci.com>
Newsgroups comp.arch.embedded
Subject Re: Multithreaded disk access
Date 2021-10-18 19:25 +0300
Organization TGI
Message-ID <skk753$tmp$1@dont-email.me> (permalink)
References <skc214$tm7$1@dont-email.me> <ski0v5$7iu$1@dont-email.me> <ski5q2$77e$1@dont-email.me> <ski6vd$hf1$1@dont-email.me> <skk2fr$ooe$1@dont-email.me>

Show all headers | View raw


On 10/18/2021 18:05, Don Y wrote:
> On 10/17/2021 3:09 PM, Dimiter_Popoff wrote:
>>> You're assuming files are laid out contiguously -- that no seeks are 
>>> needed
>>> "between sectors".
>>
>> This is the typical case anyway, most files are contiguously allocated.
> 
> I'm not sure that is the case for files that have been modified on a 
> medium.

It is not the case for files which are appended to, say logs etc.
But files like that do not make such a high percentage. I looked at a log
file which logs some activities several times a day, it has grown to
52 megabytes (ascii text) for something like 15 months (first entry
is June 1 2020). It is spread over 32 segments, as one would expect.

Then I looked at the directory where I archive emails, 311 files
(one of them being Don.txt :-). Just one or two were two segmented,
the rest were all contiguous.

The devil is not that black (literally translating a Bulgarian saying)
as you see. Worst fit allocation is of course crucial to get to
such figures, the mainstream OS-s don't do it and things there
must be much worse.

 > ....
>> Even on popular filesystems which have long forgotten how to do worst
>> fit allocation and have to defragment their disks not so infrequently.
>> But I think they have to access at least 3 locations to get to a file;
>> the directory entry, some kind of FAT-like thing, then the file.
>> Unlike dps, where 2 accesses are enough. And of course dps does
>> worst fit allocation so defragmentating is just unnecessary.
> 
> I think directories are cached.  And, possibly entire drive structures
> (depending on how much physical RAM you have available).

Well of course they must be caching them, especially since there are
gigabytes of RAM available. I know what dps does: it caches longnamed
directories which coexist with the old 8.4 ones in the same filesystem
and work faster than the 8.4 ones which typically don't get cached
(these were done to work well even on floppies, a directory entry
update writes back only the sector(s , if crossing) it occupies
etc. Then in dps the CAT (cluster allocation tables) are cached all
the time (do that for a 500G partition and enjoy reading all the
4 megabytes each time the CAT is needed to allocate new space...
it can be done, in fact the caches are enabled upon boot explicitly
on a per LUN/partition basis).

 > ...
> 
> E.g., my disk sanitizer times each (fixed size) access to profile the
> drive's performance as well as looking for trouble spots on the media.
> But, things like recal cycles or remapping bad sectors introduce
> completely unpredictable blips in the throughput.  So much so that
> I've had to implement a fair bit of logic to identify whether a
> "delay" was part of normal operation *or* a sign of an exceptional
> event.
> 
> [But, the sanitizer has a very predictable access pattern so
> there's no filesystem/content -specific issues involved; just
> process sectors as fast as possible.  (also, there is no
> need to have multiple threads per spindle; just a thread *per*
> spindle -- plus some overhead threads)
> 
> And, the sanitizer isn't as concerned with throughput as the
> human operator is the bottleneck (I can crank out a 500+GB drive
> every few minutes).]

I did something similar many years ago, wneh the largest drive
a nukeman had was 200 (230 IIRC) megabytes, i.e. in prior to
magnetoresistive heads came to the world. It did develop
bad sectors and did not do much internally about it (1993).
So I wrote the "lockout" command (still available, I see I have
recompiled it for power, last change 2016 - can't remember if it
did anything useful, nor why I did that). It accessed sector by
sector the LUN it was told to and built the lockout CAT on
its filesystem (LCAT being ORed to CAT prior to the LUN being
usable for the OS). Took quite some time on that drive but
did the job back then.

> 
> I'll mock up some synthetic loads and try various thread-spawning
> strategies to see the sorts of performance I *might* be able
> to get -- with different "preexisting" media (to minimize my
> impact on that).
> 
> I'm sure I can round up a dozen or more platforms to try -- just
> from stuff I have lying around here!  :>

I think this will give you plenty of an idea how to go about it.
Once you know the limit you can run at some reasonable figure
below it and be happy. Getting more precise figures about all
that is neither easy nor will it buy you anything.


======================================================
Dimiter Popoff, TGI             http://www.tgi-sci.com
======================================================
http://www.flickr.com/photos/didi_tgi/

Back to comp.arch.embedded | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread


Thread

Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 07:08 -0700
  Re: Multithreaded disk access Richard Damon <Richard@Damon-Family.org> - 2021-10-15 11:38 -0400
    Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 09:00 -0700
      Re: Multithreaded disk access Richard Damon <Richard@Damon-Family.org> - 2021-10-15 12:48 -0400
        Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 11:40 -0700
  Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-15 18:46 +0300
    Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 09:19 -0700
      Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-15 19:28 +0300
  Re: Multithreaded disk access Brett <ggtgp@yahoo.com> - 2021-10-17 20:27 +0000
    Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-17 14:49 -0700
      Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-17 15:01 -0700
      Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-18 01:09 +0300
        Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-18 08:05 -0700
          Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-18 19:25 +0300
            Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-18 13:46 -0700
              Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-19 00:46 +0300
                Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-18 16:17 -0700
                Re: Multithreaded disk access antispam@math.uni.wroc.pl - 2021-10-31 19:21 +0000
                Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-31 21:55 +0200
              Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-25 04:19 -0700
                Re: Multithreaded disk access Richard Damon <Richard@Damon-Family.org> - 2021-10-25 20:23 -0400
                Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-25 22:29 -0700
      Re: Multithreaded disk access Clifford Heath <no.spam@please.net> - 2021-10-18 11:58 +1100

csiph-web