Path: csiph.com!eternal-september.org!feeder.eternal-september.org!mx02.eternal-september.org!.POSTED!not-for-mail From: Matthew Carter Newsgroups: comp.lang.php Subject: Re: Fast PHP way to find a file given the leftmost characters of the file name Date: Sat, 05 Dec 2015 14:45:58 -0500 Organization: Ahungry (http://ahungry.com) Lines: 50 Message-ID: <87egf0k7bd.fsf@ahungry.com> References: Mime-Version: 1.0 Content-Type: text/plain Injection-Info: mx02.eternal-september.org; posting-host="7c986cd4736462de309a749b207746fe"; logging-data="19613"; mail-complaints-to="abuse@eternal-september.org"; posting-account="U2FsdGVkX1+APRdAgPVVOAkGLPYO2fiz" User-Agent: Gnus/5.13 (Gnus v5.13) Emacs/24.5 (gnu/linux) Cancel-Lock: sha1:5qkQifcWJjFdiWjTPE9LOIvM/oM= sha1:sAFZH4dUOiWFyNSMPvcxCRXFLHY= Xref: csiph.com comp.lang.php:15913 Markus Heinz writes: > Hello. > > On 2015-12-05 at 18:39 James Harris wrote: >> Basic query: What is the best way in PHP to find a file given just the >> leftmost characters of its name? I am looking for something that will >> scale so that it can be expected to execute quickly even if there are >> thousands of files in a given directory. > [...] >> Any comments? Even if you think the idea is a bad one I would appreciate >> the feedback. > > The glob function might be helpful to accomplish your goal: > > > Another alternative might be to store the full filenames in a database > table and then do a SQL query like the following: > > SELECT filename FROM files WHERE filename LIKE 'prefix%' > > In this query "prefix" is the prefix which is being searched for and > the query will return all complete filenames matching this prefix. > > Which solution scales better should be examined in a setup like the > target environment and is influenced by parameters such as number of > files, filesystem type, available RAM, CPU speed etc. > >> James > > Regards > > Markus > FWIW, I just tested in a directory with 50,000 files on a system with 1.5G RAM (and non-SD disk) and was able to get 190 matches when specifying the first 2 letters in 0.103 seconds using glob($letters.'*'), as well as similar results when specifying all the way to a single unique name. So, I would stick with glob vs attempting to over-engineer it (if you are working with real files and not just database content), as the only reason to micro-optimize would be if you had extremely high traffic (in which case I think you could afford the $20 or less a month to just get an SD Linode, where the cost of disk I/O is almost non-existent). -- Matthew Carter (m@ahungry.com) http://ahungry.com