Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.arch.embedded > #30816
| From | Dimiter_Popoff <dp@tgi-sci.com> |
|---|---|
| Newsgroups | comp.arch.embedded |
| Subject | Re: Multithreaded disk access |
| Date | 2021-10-18 19:25 +0300 |
| Organization | TGI |
| Message-ID | <skk753$tmp$1@dont-email.me> (permalink) |
| References | <skc214$tm7$1@dont-email.me> <ski0v5$7iu$1@dont-email.me> <ski5q2$77e$1@dont-email.me> <ski6vd$hf1$1@dont-email.me> <skk2fr$ooe$1@dont-email.me> |
On 10/18/2021 18:05, Don Y wrote: > On 10/17/2021 3:09 PM, Dimiter_Popoff wrote: >>> You're assuming files are laid out contiguously -- that no seeks are >>> needed >>> "between sectors". >> >> This is the typical case anyway, most files are contiguously allocated. > > I'm not sure that is the case for files that have been modified on a > medium. It is not the case for files which are appended to, say logs etc. But files like that do not make such a high percentage. I looked at a log file which logs some activities several times a day, it has grown to 52 megabytes (ascii text) for something like 15 months (first entry is June 1 2020). It is spread over 32 segments, as one would expect. Then I looked at the directory where I archive emails, 311 files (one of them being Don.txt :-). Just one or two were two segmented, the rest were all contiguous. The devil is not that black (literally translating a Bulgarian saying) as you see. Worst fit allocation is of course crucial to get to such figures, the mainstream OS-s don't do it and things there must be much worse. > .... >> Even on popular filesystems which have long forgotten how to do worst >> fit allocation and have to defragment their disks not so infrequently. >> But I think they have to access at least 3 locations to get to a file; >> the directory entry, some kind of FAT-like thing, then the file. >> Unlike dps, where 2 accesses are enough. And of course dps does >> worst fit allocation so defragmentating is just unnecessary. > > I think directories are cached. And, possibly entire drive structures > (depending on how much physical RAM you have available). Well of course they must be caching them, especially since there are gigabytes of RAM available. I know what dps does: it caches longnamed directories which coexist with the old 8.4 ones in the same filesystem and work faster than the 8.4 ones which typically don't get cached (these were done to work well even on floppies, a directory entry update writes back only the sector(s , if crossing) it occupies etc. Then in dps the CAT (cluster allocation tables) are cached all the time (do that for a 500G partition and enjoy reading all the 4 megabytes each time the CAT is needed to allocate new space... it can be done, in fact the caches are enabled upon boot explicitly on a per LUN/partition basis). > ... > > E.g., my disk sanitizer times each (fixed size) access to profile the > drive's performance as well as looking for trouble spots on the media. > But, things like recal cycles or remapping bad sectors introduce > completely unpredictable blips in the throughput. So much so that > I've had to implement a fair bit of logic to identify whether a > "delay" was part of normal operation *or* a sign of an exceptional > event. > > [But, the sanitizer has a very predictable access pattern so > there's no filesystem/content -specific issues involved; just > process sectors as fast as possible. (also, there is no > need to have multiple threads per spindle; just a thread *per* > spindle -- plus some overhead threads) > > And, the sanitizer isn't as concerned with throughput as the > human operator is the bottleneck (I can crank out a 500+GB drive > every few minutes).] I did something similar many years ago, wneh the largest drive a nukeman had was 200 (230 IIRC) megabytes, i.e. in prior to magnetoresistive heads came to the world. It did develop bad sectors and did not do much internally about it (1993). So I wrote the "lockout" command (still available, I see I have recompiled it for power, last change 2016 - can't remember if it did anything useful, nor why I did that). It accessed sector by sector the LUN it was told to and built the lockout CAT on its filesystem (LCAT being ORed to CAT prior to the LUN being usable for the OS). Took quite some time on that drive but did the job back then. > > I'll mock up some synthetic loads and try various thread-spawning > strategies to see the sorts of performance I *might* be able > to get -- with different "preexisting" media (to minimize my > impact on that). > > I'm sure I can round up a dozen or more platforms to try -- just > from stuff I have lying around here! :> I think this will give you plenty of an idea how to go about it. Once you know the limit you can run at some reasonable figure below it and be happy. Getting more precise figures about all that is neither easy nor will it buy you anything. ====================================================== Dimiter Popoff, TGI http://www.tgi-sci.com ====================================================== http://www.flickr.com/photos/didi_tgi/
Back to comp.arch.embedded | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 07:08 -0700
Re: Multithreaded disk access Richard Damon <Richard@Damon-Family.org> - 2021-10-15 11:38 -0400
Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 09:00 -0700
Re: Multithreaded disk access Richard Damon <Richard@Damon-Family.org> - 2021-10-15 12:48 -0400
Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 11:40 -0700
Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-15 18:46 +0300
Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-15 09:19 -0700
Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-15 19:28 +0300
Re: Multithreaded disk access Brett <ggtgp@yahoo.com> - 2021-10-17 20:27 +0000
Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-17 14:49 -0700
Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-17 15:01 -0700
Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-18 01:09 +0300
Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-18 08:05 -0700
Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-18 19:25 +0300
Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-18 13:46 -0700
Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-19 00:46 +0300
Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-18 16:17 -0700
Re: Multithreaded disk access antispam@math.uni.wroc.pl - 2021-10-31 19:21 +0000
Re: Multithreaded disk access Dimiter_Popoff <dp@tgi-sci.com> - 2021-10-31 21:55 +0200
Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-25 04:19 -0700
Re: Multithreaded disk access Richard Damon <Richard@Damon-Family.org> - 2021-10-25 20:23 -0400
Re: Multithreaded disk access Don Y <blockedofcourse@foo.invalid> - 2021-10-25 22:29 -0700
Re: Multithreaded disk access Clifford Heath <no.spam@please.net> - 2021-10-18 11:58 +1100
csiph-web