Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.os.linux.advocacy > #349365 > unrolled thread
| Started by | DFS <nospam@dfs.com> |
|---|---|
| First post | 2016-04-11 11:23 -0400 |
| Last post | 2016-04-12 12:31 -0700 |
| Articles | 20 on this page of 51 — 13 participants |
Back to article view | Back to comp.os.linux.advocacy
Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 11:23 -0400
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-11 17:09 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-11 19:25 +0000
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-11 20:58 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 17:38 -0400
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 18:11 -0400
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-11 22:25 +0000
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-12 06:16 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-11 21:46 +0000
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-12 07:18 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-12 08:15 +0000
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-12 10:34 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 13:32 -0400
Re: Algorithm to find data range Steve Carroll <fretwizzer@gmail.com> - 2016-04-12 10:39 -0700
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 13:58 -0400
Re: Algorithm to find data range Steve Carroll <fretwizzer@gmail.com> - 2016-04-12 11:59 -0700
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 23:23 -0400
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-12 07:14 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-12 10:44 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-12 15:17 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-13 17:15 -0400
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-13 22:23 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-13 18:32 -0400
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-14 04:33 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-14 05:51 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-14 20:15 -0400
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-15 01:18 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 13:25 -0400
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-12 19:08 +0000
Re: Algorithm to find data range vallor <vallor@cultnix.org> - 2016-04-12 04:40 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-11 18:06 +0000
Re: Algorithm to find data range 7 <7@enemygadgets.com> - 2016-04-11 22:46 +0000
Re: Algorithm to find data range Omar <omarsayeed@linuxmail.org> - 2016-04-11 18:52 -0400
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 13:28 -0400
Re: Algorithm to find data range Omar <omarsayeed@linuxmail.org> - 2016-04-12 13:46 -0400
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 22:13 -0400
Re: Algorithm to find data range Fabian Russell <fb@zen.info> - 2016-04-11 23:22 +0000
Re: Algorithm to find data range vallor <vallor@cultnix.org> - 2016-04-12 00:06 +0000
Re: Algorithm to find data range Omar <omarsayeed@linuxmail.org> - 2016-04-11 20:26 -0400
Re: Algorithm to find data range Chris Ahlstrom <OFeem1987@teleworm.us> - 2016-04-12 05:30 -0400
Re: Algorithm to find data range chrisv <chrisv@nospam.invalid> - 2016-04-12 06:50 -0500
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 22:34 -0400
Re: Algorithm to find data range Fabian Russell <fb@zen.info> - 2016-04-12 09:25 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 13:32 -0400
Is it your own personal NNTP/Usenet server ? Jeff-Relf.Me <@.> - 2016-04-11 16:32 -0700
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 12:47 -0400
Re: Algorithm to find data range Peter Köhlmann <peter-koehlmann@t-online.de> - 2016-04-12 20:46 +0200
Re: Algorithm to find data range chrisv <chrisv@nospam.invalid> - 2016-04-12 13:59 -0500
Re: Algorithm to find data range Silver Slimer <linux@sucks.balls> - 2016-04-12 17:08 -0400
My "newsReader" (X.ZIP) is also a console. Jeff-Relf.Me <@.> - 2016-04-12 12:27 -0700
My "newsReader" (X.ZIP) is also a console. Jeff-Relf.Me <@.> - 2016-04-12 12:31 -0700
Page 1 of 3 [1] 2 3 Next page →
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-11 11:23 -0400 |
| Subject | Algorithm to find data range |
| Message-ID | <negfcb$4ee$1@dont-email.me> |
Take a list of data:
ID Date
1 2004-05-04
2 2004-05-04
3 2004-05-05
4 2004-05-06
5 2004-05-08
...
500000 2014-05-03
I want to find the best approach to determine the unknown range of ID
numbers that correspond to a known range of dates.
That is:
* you know the date range you're interested in (say all of Sep 2008)
* you want to know the corresponding ID range (say 124385 to 125008).
The problem is you can't query or retrieve the data by date, only by ID
number. This restriction is what makes the whole thing an ordeal.
Other constraints:
* data is pulled off a busy server. You can't be hitting it
all day long.
* you can retrieve and examine max of 100 rows at a time. This is a
judgement call. I'll try 100 and see what works best, and
increase or decrease it based on performance.
* the numbers and dates are ordered low to high, but not continuous -
there are gaps in both numbers and dates.
* can't load the data into a SQL table and do min() or max(), so it's
a code-only solution (python here).
The brute-force way is to download all 500K rows, put them in a db table
and query them. That's not an option (well, it could technically be
done but then there's no thinking and creativity and challenge involved
- and what else is geek life for?).
I was thinking about three approaches that I call:
---------------------------------------------------------------------
1. Half-Height Elimination:
examine 1 row at a time (starting with the middle row), cutting the data
to be examined in half on each iteration. This will require some kind
of recursive coding methodology. With a domain of 500K rows, this
approach could require as many as 19 iterations (and each iteration
would involve a small request from the server) in its simplest form:
Iter Rows Remaining
500000
1 250000
2 125000
3 62500
4 31250
5 15625
6 7813
7 3906
8 1953
9 977
10 488
11 244
12 122
13 61
14 31
15 15
16 8
17 4
18 2
19 1 (found your search date)
This pic will help to see how it works:
http://i.imgur.com/AFfIjrI.png
---------------------------------------------------------------------
2. Decile Seek:
create N class objects and have each one scan a short range of data
simultaneous with the other objects, then on to the next range of data
an so on. Each object would be responsible for scanning up to 50K rows,
low to high. Eventually one will hit the beginning of the known date
range, and one will hit the end of the known date range, and then I'll
know the corresponding number range. This would probably be the slowest
and most "wasteful" approach I've come up with, but it could be
interesting to develop.
pic: http://i.imgur.com/iKjhbOM.jpg
---------------------------------------------------------------------
3. Date Distribution:
if you assume the ID numbers and the dates are evenly distributed across
time (ID 1 to 500000, 500K rows over 10 years, 50K per year, 4167 per
month, 962 per week, 137 per day, etc), then it's a simple calculation
to determine at which ID number your date range starts. It won't be
exact because the data isn't actually evenly distributed (gaps as
mentioned before), but it will be a good start.
Calculation:
A = first date of data
B = first date of search
C = last date of search
D = last date of data
firstID = earliest ID number
lastID = latest ID number
Start search ID at: ((B-A)/(D-A)) * (lastID - firstID)
End search ID at: ((C-A)/(D-A)) * (lastID - firstID)
Note: for most searches, A < B < C < D.
pic http://i.imgur.com/SQcq9Hz.jpg
---------------------------------------------------------------------
I haven't coded any of these approaches, but the Date Distribution
approach will probably be fastest and most efficient by far (as measured
by both server hits and processing time).
If you read this far you probably have an opinion or thought. Let's
hear it!
[toc] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2016-04-11 17:09 +0000 |
| Message-ID | <sandman-2ada1150bc27e3d11b868223a447a75e@individual.net> |
| In reply to | #349365 |
In article <negfcb$4ee$1@dont-email.me>, DFS wrote: > Take a list of data: > ID Date > 1 2004-05-04 > 2 2004-05-04 > 3 2004-05-05 > 4 2004-05-06 > 5 2004-05-08 > ... > 500000 2014-05-03 > I want to find the best approach to determine the unknown range of > ID numbers that correspond to a known range of dates. > That is: * you know the date range you're interested in (say all of > Sep 2008) * you want to know the corresponding ID range (say 124385 > to 125008). > The problem is you can't query or retrieve the data by date, only by > ID number. This restriction is what makes the whole thing an ordeal. > Other constraints: > * data is pulled off a busy server. You can't be hitting it > all day long. > * you can retrieve and examine max of 100 rows at a time. This is a > judgement call. I'll try 100 and see what works best, and increase > or decrease it based on performance. > * the numbers and dates are ordered low to high, but not continuous > - there are gaps in both numbers and dates. > * can't load the data into a SQL table and do min() or max(), so > it's a code-only solution (python here). That's some strange limitations. > I was thinking about three approaches that I call: > --------------------------------------------------------------------- > 1. Half-Height Elimination: examine 1 row at a time (starting with > the middle row), cutting the data to be examined in half on each > iteration. This will require some kind of recursive coding > methodology. With a domain of 500K rows, this approach could require > as many as 19 iterations (and each iteration would involve a small > request from the server) in its simplest form: > This pic will help to see how it works: > http://i.imgur.com/AFfIjrI.png That supposes that for each cut, the date range is in the cut. I.e. it could just as easily be twice as many cuts. Also, given restriction 3 above, there is no way for you to know what ID is at the middle of a sample. If you have 500k rows of data, you'd think that ID 250,000 would be in the middle, but since it was stipulated that ID could contain caps and not be continuous, you may have cut two thirds up in the series. In fact, ID 250,000 may be the very last ID according to rule #3. Or the first, for that matter. > --------------------------------------------------------------------- > 2. Decile Seek: create N class objects and have each one scan a > short range of data simultaneous with the other objects, then on to > the next range of data an so on. Each object would be responsible > for scanning up to 50K rows, low to high. Eventually one will hit > the beginning of the known date range, and one will hit the end of > the known date range, and then I'll know the corresponding number > range. This would probably be the slowest and most "wasteful" > approach I've come up with, but it could be interesting to develop. > pic: http://i.imgur.com/iKjhbOM.jpg Also, would seem to be in violation of rule #1 and #2. > --------------------------------------------------------------------- > 3. Date Distribution: if you assume the ID numbers and the dates are > evenly distributed across time (ID 1 to 500000, 500K rows over 10 > years, 50K per year, 4167 per month, 962 per week, 137 per day, > etc), then it's a simple calculation to determine at which ID number > your date range starts. It won't be exact because the data isn't > actually evenly distributed (gaps as mentioned before), but it will > be a good start. > pic http://i.imgur.com/SQcq9Hz.jpg It still relies on assumptions that could be 100% false. I.e. according to your rules, the first ID could be 800,000 for all we know, in spite of your example table. I.e. you'd make a query for ID 1, which returns nothing, so you step to 2, which returns nothing, etc etc. Suddenly you've wasted 799,999 queries that result in nothing. > I haven't coded any of these approaches, but the Date Distribution > approach will probably be fastest and most efficient by far (as > measured by both server hits and processing time). It would, if those assumptions can be trusted. In the end, traversing 500k lines of data is just as quick with todays CPU's. -- Sandman
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-11 19:25 +0000 |
| Message-ID | <hgjdke902.ata34p@rooftop.invalid> |
| In reply to | #349378 |
Sandman <mr@sandman.net> wrote: > In article <negfcb$4ee$1@dont-email.me>, DFS wrote: > >> Take a list of data: > >> ID Date >> 1 2004-05-04 >> 2 2004-05-04 >> 3 2004-05-05 >> 4 2004-05-06 >> 5 2004-05-08 >> ... >> 500000 2014-05-03 > >> I want to find the best approach to determine the unknown range of >> ID numbers that correspond to a known range of dates. > >> That is: * you know the date range you're interested in (say all of >> Sep 2008) * you want to know the corresponding ID range (say 124385 >> to 125008). > >> The problem is you can't query or retrieve the data by date, only by >> ID number. This restriction is what makes the whole thing an ordeal. > >> Other constraints: > >> * data is pulled off a busy server. You can't be hitting it >> all day long. > >> * you can retrieve and examine max of 100 rows at a time. This is a >> judgement call. I'll try 100 and see what works best, and increase >> or decrease it based on performance. > >> * the numbers and dates are ordered low to high, but not continuous >> - there are gaps in both numbers and dates. > >> * can't load the data into a SQL table and do min() or max(), so >> it's a code-only solution (python here). > > That's some strange limitations. > >> I was thinking about three approaches that I call: >> --------------------------------------------------------------------- > >> 1. Half-Height Elimination: examine 1 row at a time (starting with >> the middle row), cutting the data to be examined in half on each >> iteration. This will require some kind of recursive coding >> methodology. With a domain of 500K rows, this approach could require >> as many as 19 iterations (and each iteration would involve a small >> request from the server) in its simplest form: > >> This pic will help to see how it works: >> http://i.imgur.com/AFfIjrI.png > > That supposes that for each cut, the date range is in the cut. I.e. it could > just as easily be twice as many cuts. > It would be left or right, unless that cut just happened to hit somewhere in the middle of the range. If it's a hit, then you test how far left and right that date it extends. If this is usenet, I don't think that for a single group you're going to need to do too many more cuts once you've landed on a date match. (Even a Snit-based cola probably had fewer than 1000 posts in a single day). > Also, given restriction 3 above, there is no way for you to know what ID is > at the middle of a sample. If you have 500k rows of data, you'd think that ID > 250,000 would be in the middle, but since it was stipulated that ID could > contain caps and not be continuous, you may have cut two thirds up in the > series. In fact, ID 250,000 may be the very last ID according to rule #3. Or > the first, for that matter. > I took it to mean 500,000 items, not necessarily ID 1 - ID 500,000.
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2016-04-11 20:58 +0000 |
| Message-ID | <sandman-b3b70fc4cbc0697fbcd2cb88cf7e38a7@individual.net> |
| In reply to | #349396 |
In article <hgjdke902.ata34p@rooftop.invalid>, owl wrote: > > Sandman: > > Also, given restriction 3 above, there is no way for you to know > > what ID is at the middle of a sample. If you have 500k rows of > > data, you'd think that ID 250,000 would be in the middle, but > > since it was stipulated that ID could contain caps and not be > > continuous, you may have cut two thirds up in the series. In fact, > > ID 250,000 may be the very last ID according to rule #3. Or the > > first, for that matter. > > I took it to mean 500,000 items, not necessarily ID 1 - ID 500,000. Exactly, meaning that the ID could be anything, and since we were supposed to fetch posts only by ID, it's impossible to fetch anything, right? For all we know, the ID's could be 1,982,827 to 3,462,987 with enough gaps in the ID's to amount ot 500,000 posts. The only known variables are the number of posts, and that the posts have dates ranging from 2004-05-03 to 2014-05-03, right? So, basically, the data could be: 2,765,982 2004-03-05 2,765,983 2004-03-05 2,765,984 2004-03-05 2,765,987 2004-03-05 .. 4,756,234 2014-03-05 At this point it would be impossible to efficiently figure out what ID numbers correlate to what dates if we can't query what the first and last ID is. I currently have 172,946 posts in my usenet table, and if I were interested in posts made between two dates, I would go: : select * from usenet where date >= '2014-08-01' and date < '2014-09-01'; And I could get two posts back or several thousands, depending on traffic. And since the date field is indexed, the query takes 15ms (and returned 3,513 posts). -- Sandman
[toc] | [prev] | [next] | [standalone]
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-11 17:38 -0400 |
| Message-ID | <neh5ai$vrp$1@dont-email.me> |
| In reply to | #349401 |
On 4/11/2016 4:58 PM, Sandman wrote: > In article <hgjdke902.ata34p@rooftop.invalid>, owl wrote: > >>> Sandman: >>> Also, given restriction 3 above, there is no way for you to know >>> what ID is at the middle of a sample. If you have 500k rows of >>> data, you'd think that ID 250,000 would be in the middle, but >>> since it was stipulated that ID could contain caps and not be >>> continuous, you may have cut two thirds up in the series. In fact, >>> ID 250,000 may be the very last ID according to rule #3. Or the >>> first, for that matter. >> >> I took it to mean 500,000 items, not necessarily ID 1 - ID 500,000. In my example, they were the same. In real life, probably not. I think the gaps occur when people cancel messages. > Exactly, meaning that the ID could be anything, and since we were supposed to > fetch posts only by ID, it's impossible to fetch anything, right? For all we > know, the ID's could be 1,982,827 to 3,462,987 with enough gaps in the ID's to > amount ot 500,000 posts. > > The only known variables are the number of posts, and that the posts have dates > ranging from 2004-05-03 to 2014-05-03, right? > > So, basically, the data could be: > > 2,765,982 2004-03-05 > 2,765,983 2004-03-05 > 2,765,984 2004-03-05 > 2,765,987 2004-03-05 > .. > 4,756,234 2014-03-05 > > At this point it would be impossible to efficiently figure out what ID numbers > correlate to what dates if we can't query what the first and last ID is. I do know what they are. The news server gives you a GROUP response like this: 211 367904 1 578023 comp.os.linux.advocacy Which is: 211 = GROUP command completed 367904 = # of articles on the server 1 = first ID of article on the server 578023 = last ID of article on the server The IDs are unique to that server. > I currently have 172,946 posts in my usenet table, and if I were interested in > posts made between two dates, I would go: > > : select * from usenet where date >= '2014-08-01' and date < '2014-09-01'; > > And I could get two posts back or several thousands, depending on traffic. And > since the date field is indexed, the query takes 15ms (and returned 3,513 > posts).
[toc] | [prev] | [next] | [standalone]
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-11 18:11 -0400 |
| Message-ID | <neh77o$a07$1@dont-email.me> |
| In reply to | #349402 |
On 4/11/2016 5:57 PM, owl wrote: > Here's a query for Nov 2014 from telnet'ing to my news server and > using xpat: > > Trying ::1... > Trying 127.0.0.1... > Connected to localhost. > Escape character is '^]'. > 200 Check out http://redacted/ for info about NNTP access (posting ok). > 381 PASS required > 281 Ok > 211 81864 2152569 2234432 comp.os.linux.advocacy > 221 date matches follow (NOV) > 2152768 Sun, 23 Nov 2014 19:34:33 -0500 > 2152769 Sun, 23 Nov 2014 16:37:39 -0800 (PST) > 2152770 Sun, 23 Nov 2014 19:47:52 -0500 > 2152771 Sun, 23 Nov 2014 16:50:34 -0800 (PST) > 2152772 Sun, 23 Nov 2014 19:51:26 -0500 <snip> Is (NOV) your only xpat spec? Why didn't it return anything from NOV 2013?
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-11 22:25 +0000 |
| Message-ID | <f9ela.angia4@rooftop.invalid> |
| In reply to | #349404 |
DFS <nospam@dfs.com> wrote: > On 4/11/2016 5:57 PM, owl wrote: > >> Here's a query for Nov 2014 from telnet'ing to my news server and >> using xpat: >> >> Trying ::1... >> Trying 127.0.0.1... >> Connected to localhost. >> Escape character is '^]'. >> 200 Check out http://redacted/ for info about NNTP access (posting ok). >> 381 PASS required >> 281 Ok >> 211 81864 2152569 2234432 comp.os.linux.advocacy >> 221 date matches follow (NOV) >> 2152768 Sun, 23 Nov 2014 19:34:33 -0500 >> 2152769 Sun, 23 Nov 2014 16:37:39 -0800 (PST) >> 2152770 Sun, 23 Nov 2014 19:47:52 -0500 >> 2152771 Sun, 23 Nov 2014 16:50:34 -0800 (PST) >> 2152772 Sun, 23 Nov 2014 19:51:26 -0500 > > <snip> > > > Is (NOV) your only xpat spec? Why didn't it return anything from NOV 2013? > > That's just a coincidence. (NOV) means "news overview". My spec was: xpat date 2152569-2234432 *Nov 2014* Nov 23, 2014 is the earliest retention on my server. (The advertised 2152569 ID number is wrong). xpat date 2152767 * 221 date matches follow (NOV) . xpat date 2152768 * 221 date matches follow (NOV) 2152768 Sun, 23 Nov 2014 19:34:33 -0500 .
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2016-04-12 06:16 +0000 |
| Message-ID | <sandman-2577555869e1835f16de7c621f1158bf@individual.net> |
| In reply to | #349402 |
In article <neh5ai$vrp$1@dont-email.me>, DFS wrote: > > > > Sandman: > > > > Also, given restriction 3 above, there is no way for > > > > you to know what ID is at the middle of a sample. If you have > > > > 500k rows of data, you'd think that ID 250,000 would be in the > > > > middle, but since it was stipulated that ID could contain caps > > > > and not be continuous, you may have cut two thirds up in the > > > > series. In fact, ID 250,000 may be the very last ID according > > > > to rule #3. Or the first, for that matter. > > > > > > owl: > > > I took it to mean 500,000 items, not necessarily ID 1 - ID > > > 500,000. > > > In my example, they were the same. > In real life, probably not. I think the gaps occur when people > cancel messages. > > Sandman: > > Exactly, meaning that the ID could be anything, and since we were > > supposed to fetch posts only by ID, it's impossible to fetch > > anything, right? For all we know, the ID's could be 1,982,827 to > > 3,462,987 with enough gaps in the ID's to amount ot 500,000 posts. > > > The only known variables are the number of posts, and that the > > posts have dates ranging from 2004-05-03 to 2014-05-03, right? > > > So, basically, the data could be: > > > 2,765,982 2004-03-05 > > 2,765,983 2004-03-05 > > 2,765,984 2004-03-05 > > 2,765,987 2004-03-05 > > .. > > 4,756,234 2014-03-05 > > > At this point it would be impossible to efficiently figure out > > what ID numbers correlate to what dates if we can't query what the > > first and last ID is. > > I do know what they are. The news server gives you a GROUP response > like this: > 211 367904 1 578023 comp.os.linux.advocacy > 211 = GROUP command completed > 367904 = # of articles on the server > 1 = first ID of article on the server > 578023 = last ID of article on the server > The IDs are unique to that server. You should have mentioned form the start that you were querying a NNTP server. XPAT is the command you're looking for. -> GROUP comp.sys.mac.advocacy 211 18294 1736615 1754937 comp.sys.mac.advocacy Like you said, this shows us the first and last ID. So when can search that entire range using XPAT: -> XPAT Date 1736615-1754937 *Sep 2014* 221 Date matches follow. 1740170 1 Sep 2014 06:51:59 GMT 1740171 1 Sep 2014 07:27:03 GMT 1740172 1 Sep 2014 07:27:36 GMT 1740173 Mon, 01 Sep 2014 13:42:21 +0100 1740174 Mon, 1 Sep 2014 08:51:13 -0400 1740175 Mon, 01 Sep 2014 07:58:05 -0500 1740176 Mon, 1 Sep 2014 09:41:37 -0400 ... 1741098 Tue, 30 Sep 2014 19:45:19 -0700 (PDT) And then you have a full list of posts that match your search criteria. Done, two requests who are fully indexed (Date should be in the overview.fmt) -- Sandman
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-11 21:46 +0000 |
| Message-ID | <hgjei0a3.fauo@rooftop.invalid> |
| In reply to | #349401 |
Sandman <mr@sandman.net> wrote: > In article <hgjdke902.ata34p@rooftop.invalid>, owl wrote: > >> > Sandman: >> > Also, given restriction 3 above, there is no way for you to know >> > what ID is at the middle of a sample. If you have 500k rows of >> > data, you'd think that ID 250,000 would be in the middle, but >> > since it was stipulated that ID could contain caps and not be >> > continuous, you may have cut two thirds up in the series. In fact, >> > ID 250,000 may be the very last ID according to rule #3. Or the >> > first, for that matter. >> >> I took it to mean 500,000 items, not necessarily ID 1 - ID 500,000. > > Exactly, meaning that the ID could be anything, and since we were supposed to > fetch posts only by ID, it's impossible to fetch anything, right? For all we > know, the ID's could be 1,982,827 to 3,462,987 with enough gaps in the ID's to > amount ot 500,000 posts. But the assumption is that he knows the first and last ID number and the total count. Also, assuming this is usenet posts, if there's a gap at all, it's not going to be large. > > The only known variables are the number of posts, and that the posts have dates > ranging from 2004-05-03 to 2014-05-03, right? > > So, basically, the data could be: > > 2,765,982 2004-03-05 > 2,765,983 2004-03-05 > 2,765,984 2004-03-05 > 2,765,987 2004-03-05 > .. > 4,756,234 2014-03-05 > > At this point it would be impossible to efficiently figure out what ID numbers > correlate to what dates if we can't query what the first and last ID is. > > I currently have 172,946 posts in my usenet table, and if I were interested in > posts made between two dates, I would go: > > : select * from usenet where date >= '2014-08-01' and date < '2014-09-01'; > He said he's not able to query by date. He has to use the article ID and test for date. > And I could get two posts back or several thousands, depending on traffic. And > since the date field is indexed, the query takes 15ms (and returned 3,513 > posts). > I believe he's querying a news server, not an SQL database.
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2016-04-12 07:18 +0000 |
| Message-ID | <sandman-57484cfe7bf5206e6be66023a452081f@individual.net> |
| In reply to | #349403 |
In article <hgjei0a3.fauo@rooftop.invalid>, owl wrote: > > Sandman: > > Exactly, meaning that the ID could be anything, and since we were > > supposed to fetch posts only by ID, it's impossible to fetch > > anything, right? For all we know, the ID's could be 1,982,827 to > > 3,462,987 with enough gaps in the ID's to amount ot 500,000 posts. > > But the assumption is that he knows the first and last ID number and > the total count. Also, assuming this is usenet posts, if there's a > gap at all, it's not going to be large. I have since figured out that this question pertained to usenet posts queried from a NNTP server, and also given DFS the method to solve his query problem super efficiently. But at the time, I didn't know it was and just went by the restrictions and info we were given :) > > Sandman: > > The only known variables are the number of posts, and that the > > posts have dates ranging from 2004-05-03 to 2014-05-03, right? > > > So, basically, the data could be: > > > 2,765,982 2004-03-05 > > 2,765,983 2004-03-05 > > 2,765,984 2004-03-05 > > 2,765,987 2004-03-05 > > .. > > 4,756,234 2014-03-05 > > > At this point it would be impossible to efficiently figure out > > what ID numbers correlate to what dates if we can't query what the > > first and last ID is. > > > I currently have 172,946 posts in my usenet table, and if I were > > interested in posts made between two dates, I would go: > > > : select * from usenet where date >= '2014-08-01' and date < > > '2014-09-01'; > > He said he's not able to query by date. He has to use the article ID > and test for date. Indeed, I just said how I do it :) > > Sandman: > > And I could get two posts back or several thousands, depending on > > traffic. And since the date field is indexed, the query takes 15ms > > (and returned 3,513 posts). > > I believe he's querying a news server, not an SQL database. I have realized that as well, now... :) And given him the most efficient method to solve his problem :) -- Sandman
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-12 08:15 +0000 |
| Message-ID | <ghjd30a.atp348@rooftop.invalid> |
| In reply to | #349473 |
Sandman <mr@sandman.net> wrote: > In article <hgjei0a3.fauo@rooftop.invalid>, owl wrote: > >> > Sandman: >> > Exactly, meaning that the ID could be anything, and since we were >> > supposed to fetch posts only by ID, it's impossible to fetch >> > anything, right? For all we know, the ID's could be 1,982,827 to >> > 3,462,987 with enough gaps in the ID's to amount ot 500,000 posts. >> >> But the assumption is that he knows the first and last ID number and >> the total count. Also, assuming this is usenet posts, if there's a >> gap at all, it's not going to be large. > > I have since figured out that this question pertained to usenet posts queried > from a NNTP server, and also given DFS the method to solve his query problem > super efficiently. But at the time, I didn't know it was and just went by the > restrictions and info we were given :) > >> > Sandman: >> > The only known variables are the number of posts, and that the >> > posts have dates ranging from 2004-05-03 to 2014-05-03, right? >> >> > So, basically, the data could be: >> >> > 2,765,982 2004-03-05 >> > 2,765,983 2004-03-05 >> > 2,765,984 2004-03-05 >> > 2,765,987 2004-03-05 >> > .. >> > 4,756,234 2014-03-05 >> >> > At this point it would be impossible to efficiently figure out >> > what ID numbers correlate to what dates if we can't query what the >> > first and last ID is. >> >> > I currently have 172,946 posts in my usenet table, and if I were >> > interested in posts made between two dates, I would go: >> >> > : select * from usenet where date >= '2014-08-01' and date < >> > '2014-09-01'; >> >> He said he's not able to query by date. He has to use the article ID >> and test for date. > > Indeed, I just said how I do it :) > >> > Sandman: >> > And I could get two posts back or several thousands, depending on >> > traffic. And since the date field is indexed, the query takes 15ms >> > (and returned 3,513 posts). >> >> I believe he's querying a news server, not an SQL database. > > I have realized that as well, now... :) > > And given him the most efficient method to solve his problem :) > I told him about xpat yesterday. We've already discussed it.
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2016-04-12 10:34 +0000 |
| Message-ID | <sandman-bbdc0255e0c3eee713dc78c4782ee760@individual.net> |
| In reply to | #349477 |
In article <ghjd30a.atp348@rooftop.invalid>, owl wrote: > > Sandman: > > I have realized that as well, now... :) > > > And given him the most efficient method to solve his problem :) > > I told him about xpat yesterday. We've already discussed it. I had not seen that when I wrote about it, though, read it later. :) -- Sandman
[toc] | [prev] | [next] | [standalone]
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-12 13:32 -0400 |
| Message-ID | <nejb8b$r1o$3@dont-email.me> |
| In reply to | #349473 |
On 4/12/2016 3:18 AM, Sandman wrote: > And given him the most efficient method to solve his problem :) It does /look/ efficient, but I'm sorry hoss, owl beat you to it. Somehow he guessed I was reading Usenet data, and told me about xpat. The helpful old fella even made a video of him using it. I'm touched - he gets a new greeting card this year, rather than the recycled one I sent last year. Unfortunately, it's not as easy as just saying 'use xpat'. My program is python (2.7), and for some reason xpat isn't part of the nntplib for 2.7 (or 3.5) - even though it was considered a "common NNTP extension" 15 years ago. Until I figure that out, I'm progressing with at least one of my methods. I did some tests, and downloading 15,000 article IDs and dates took <1 second! Kewl! But it's a serious violation of my commitment to low impact computing...
[toc] | [prev] | [next] | [standalone]
| From | Steve Carroll <fretwizzer@gmail.com> |
|---|---|
| Date | 2016-04-12 10:39 -0700 |
| Message-ID | <6a893a70-5c54-452a-8f32-eb07ce3452e9@googlegroups.com> |
| In reply to | #349545 |
On Tuesday, April 12, 2016 at 11:32:05 AM UTC-6, DFS wrote: > On 4/12/2016 3:18 AM, Sandman wrote: > > > > And given him the most efficient method to solve his problem :) > > > It does /look/ efficient, but I'm sorry hoss, owl beat you to it. > Somehow he guessed I was reading Usenet data, and told me about xpat. > The helpful old fella even made a video of him using it. I'm touched - > he gets a new greeting card this year, rather than the recycled one I > sent last year. > > Unfortunately, it's not as easy as just saying 'use xpat'. My program > is python (2.7), and for some reason xpat isn't part of the nntplib for > 2.7 (or 3.5) - even though it was considered a "common NNTP extension" > 15 years ago. > > Until I figure that out, I'm progressing with at least one of my > methods. I did some tests, and downloading 15,000 article IDs and dates > took <1 second! Kewl! > > But it's a serious violation of my commitment to low impact computing... Not at all like your COLA stats ;)
[toc] | [prev] | [next] | [standalone]
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-12 13:58 -0400 |
| Message-ID | <nejcp7$cod$1@dont-email.me> |
| In reply to | #349549 |
On 4/12/2016 1:39 PM, Steve Carroll wrote: > On Tuesday, April 12, 2016 at 11:32:05 AM UTC-6, DFS wrote: >> But it's a serious violation of my commitment to low impact computing... > > Not at all like your COLA stats ;) ha! But... it would be interesting to see who blows the most hot air. I don't think it's me.
[toc] | [prev] | [next] | [standalone]
| From | Steve Carroll <fretwizzer@gmail.com> |
|---|---|
| Date | 2016-04-12 11:59 -0700 |
| Message-ID | <6ea915c1-0741-4515-b322-9c9d26acabfd@googlegroups.com> |
| In reply to | #349556 |
On Tuesday, April 12, 2016 at 11:58:09 AM UTC-6, DFS wrote: > On 4/12/2016 1:39 PM, Steve Carroll wrote: > > On Tuesday, April 12, 2016 at 11:32:05 AM UTC-6, DFS wrote: > > > >> But it's a serious violation of my commitment to low impact computing... > > > > Not at all like your COLA stats ;) > > > ha! > > But... it would be interesting to see who blows the most hot air. I > don't think it's me. I'm just giving you crap cuz I've seen you and owl go back and forth on stats. To me blowing hot air is spewing crap you that know is crap, I don't see you doing that; to the contrary, you seem to really believe the garbage you spew. See? I just can't help myself ;)
[toc] | [prev] | [next] | [standalone]
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-11 23:23 -0400 |
| Message-ID | <nehpge$pok$1@dont-email.me> |
| In reply to | #349378 |
On 4/11/2016 1:09 PM, Sandman wrote:
> In article <negfcb$4ee$1@dont-email.me>, DFS wrote:
>
>> Take a list of data:
>
>> ID Date
>> 1 2004-05-04
>> 2 2004-05-04
>> 3 2004-05-05
>> 4 2004-05-06
>> 5 2004-05-08
>> ...
>> 500000 2014-05-03
>
>> I want to find the best approach to determine the unknown range of
>> ID numbers that correspond to a known range of dates.
>
>> That is: * you know the date range you're interested in (say all of
>> Sep 2008) * you want to know the corresponding ID range (say 124385
>> to 125008).
>
>> The problem is you can't query or retrieve the data by date, only by
>> ID number. This restriction is what makes the whole thing an ordeal.
>
>> Other constraints:
>
>> * data is pulled off a busy server. You can't be hitting it
>> all day long.
>
>> * you can retrieve and examine max of 100 rows at a time. This is a
>> judgement call. I'll try 100 and see what works best, and increase
>> or decrease it based on performance.
>
>> * the numbers and dates are ordered low to high, but not continuous
>> - there are gaps in both numbers and dates.
>
>> * can't load the data into a SQL table and do min() or max(), so
>> it's a code-only solution (python here).
>
> That's some strange limitations.
Yes. A couple are self-imposed to make it interesting. But the data
really isn't continuous. And it may or may not start with ID 1 - the
solutions I'm developing never make that assumption.
>> I was thinking about three approaches that I call:
>> ---------------------------------------------------------------------
>
>> 1. Half-Height Elimination: examine 1 row at a time (starting with
>> the middle row), cutting the data to be examined in half on each
>> iteration. This will require some kind of recursive coding
>> methodology. With a domain of 500K rows, this approach could require
>> as many as 19 iterations (and each iteration would involve a small
>> request from the server) in its simplest form:
>
>> This pic will help to see how it works:
>> http://i.imgur.com/AFfIjrI.png
>
> That supposes that for each cut, the date range is in the cut. I.e. it could
> just as easily be twice as many cuts.
> Also, given restriction 3 above, there is no way for you to know what ID is
> at the middle of a sample. If you have 500k rows of data, you'd think that ID
> 250,000 would be in the middle, but since it was stipulated that ID could
> contain caps and not be continuous, you may have cut two thirds up in the
> series. In fact, ID 250,000 may be the very last ID according to rule #3. Or
> the first, for that matter.
I knew this one would be hard to explain, but I'm not examining the
entire cut of data. Only the exact ID that cuts the remaining data in half.
It's easy to show and understand (and program) when the IDs are 1 -
500000 continuous, but the ID range might be 38 to 1427 with gaps.
Start: 0.5 * (1427 - 38) = 1389 rows to examine
Cut 1389 rows at most
1 695
2 347
3 174
4 87
5 43
6 22
7 11
8 5
9 3
10 1
Cut 1. retrieve date where ID = 695. If no row, recurse down until a
row is retrieved (this could be 'expensive' as far as server hits
are concerned).
If you keep chopping the data down, you'll eventually hit your search
date (or the closest date to it - the data gaps might mean your search
date isn't actually in the data set. Another issue to consider!).
>> ---------------------------------------------------------------------
>
>> 2. Decile Seek: create N class objects and have each one scan a
>> short range of data simultaneous with the other objects, then on to
>> the next range of data an so on. Each object would be responsible
>> for scanning up to 50K rows, low to high. Eventually one will hit
>> the beginning of the known date range, and one will hit the end of
>> the known date range, and then I'll know the corresponding number
>> range. This would probably be the slowest and most "wasteful"
>> approach I've come up with, but it could be interesting to develop.
>
>> pic: http://i.imgur.com/iKjhbOM.jpg
>
> Also, would seem to be in violation of rule #1 and #2.
Most likely.
>> ---------------------------------------------------------------------
>
>> 3. Date Distribution: if you assume the ID numbers and the dates are
>> evenly distributed across time (ID 1 to 500000, 500K rows over 10
>> years, 50K per year, 4167 per month, 962 per week, 137 per day,
>> etc), then it's a simple calculation to determine at which ID number
>> your date range starts. It won't be exact because the data isn't
>> actually evenly distributed (gaps as mentioned before), but it will
>> be a good start.
>
>> pic http://i.imgur.com/SQcq9Hz.jpg
>
> It still relies on assumptions that could be 100% false. I.e. according to
> your rules, the first ID could be 800,000 for all we know, in spite of your
> example table.
>
> I.e. you'd make a query for ID 1, which returns nothing, so you step to 2,
> which returns nothing, etc etc. Suddenly you've wasted 799,999 queries that
> result in nothing.
I see I didn't make it clear, but you will know the first and last ID
numbers in the data range (and the dates that go along with them).
You'll also know the number of rows in the data range. That will help.
For example, right now
--------------------------------------------------------
Server: news.eternal-september.org
Group: comp.os.linux.advocacy (c.o.l.a)
Contains 367,961 articles
ID range: 1 to 578082
First post: Sat, 17 Nov 2007 11:28:52 +0000
Last post: Mon, 11 Apr 2016 22:52:23 -0400
--------------------------------------------------------
>> I haven't coded any of these approaches, but the Date Distribution
>> approach will probably be fastest and most efficient by far (as
>> measured by both server hits and processing time).
>
> It would, if those assumptions can be trusted.
It's newsgroup articles on an NNTP server - any newsgroup - so I think
it's fairly evenly distributed.
> In the end, traversing 500k lines of data is just as quick with todays CPU's.
Sure, but where's the fun in that? I could download it once, put it in
a db table and query it right away. Even a Linux advocate could do that.
But downloading and storing 500K rows in a table or in memory is
unacceptable if I want a portable system.
What if I can write code that gets the answers quickly, doesn't require
a db, and allows a user to just type:
-stats sci.math 30 days ending 20080930
-stats comp.os.linux.advocacy 7 days beginning 20160101
That's the ticket.
Thanks for your input, Sandman! It brought up some issues.
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2016-04-12 07:14 +0000 |
| Message-ID | <sandman-6927da386733528d120e976ccb2bcd24@individual.net> |
| In reply to | #349460 |
In article <nehpge$pok$1@dont-email.me>, DFS wrote: > > Sandman: > > In the end, traversing 500k lines of data is just as quick with > > todays CPU's. > > Sure, but where's the fun in that? I could download it once, put it > in a db table and query it right away. Even a Linux advocate could > do that. That's what I'm doing - using suck to download and put all posts from my newsgroups in a DB to make them fully indexed and searchable. :) > But downloading and storing 500K rows in a table or in memory is > unacceptable if I want a portable system. > What if I can write code that gets the answers quickly, doesn't > require a db, and allows a user to just type: > -stats sci.math 30 days ending 20080930 > -stats comp.os.linux.advocacy 7 days beginning 20160101 I gave you the answer to that in an earlier post, but the answer is: group sci.math 211 69183 1478617 1547851 sci.mah XPAT Date 1478617- *Sep 2014* <list of ID's and matched headers> All your program have to do is figure out the name of the months and years for the date range of the command and fetch all months and then sort after the fact. I.e. "stats comp.os.linux.advocacy 7 days beginning 20160101" Would execute: group comp.os.linux.advocacy 211 103059 2128492 2231645 comp.os.linux.advocacy XPAT Date 2128492- *Dec 2015* <list of ID's and full date> XPAT Date 2128492- *Jan 2016* <list of ID's and full date> And since the full date is in the result, you can from that result determine what ID's you want to fetch data from. > Thanks for your input, Sandman! It brought up some issues. :) -- Sandman
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-12 10:44 +0000 |
| Message-ID | <hjgi30ara.kgi4r3@rooftop.invalid> |
| In reply to | #349460 |
DFS <nospam@dfs.com> wrote: > On 4/11/2016 1:09 PM, Sandman wrote: >> ... >> In the end, traversing 500k lines of data is just as quick with todays CPU's. > > Sure, but where's the fun in that? I could download it once, put it in > a db table and query it right away. Even a Linux advocate could do that. > > But downloading and storing 500K rows in a table or in memory is > unacceptable if I want a portable system. > > What if I can write code that gets the answers quickly, doesn't require > a db, and allows a user to just type: > > -stats sci.math 30 days ending 20080930 > -stats comp.os.linux.advocacy 7 days beginning 20160101 > > > That's the ticket. > Like this? anon@lowtide:~$ ./blah.sh localhost cola 20150615 +6days Earliest date searched: 15 Jun 2015 Latest date searched : 21 Jun 2015 First ID in range: 2183608 Mon, 15 Jun 2015 00:43:10 +0200 Last ID in range : 2185782 Sun, 21 Jun 2015 22:42:47 -0700 anon@lowtide:~$ anon@lowtide:~$ ./blah.sh localhost cola 20150615 -2months Earliest date searched: 15 Apr 2015 Latest date searched : 15 Jun 2015 First ID in range: 2173942 Wed, 15 Apr 2015 01:13:48 +0200 Last ID in range : 2183869 Mon, 15 Jun 2015 21:36:04 -0700 (PDT) anon@lowtide:~$ anon@lowtide:~$ ./blah.sh localhost alt.test 20151231 -1year Earliest date searched: 31 Dec 2014 Latest date searched : 31 Dec 2015 First ID in range: 4645896 Wed, 31 Dec 2014 00:08:41 +0000 (UTC) Last ID in range : 4747061 Thu, 31 Dec 2015 18:40:06 -0500 anon@lowtide:~$ I'll post the code later after I add some logic to handle leading zeros on the day portion of the date string. I have them stripped in this version and add a leading space to the date query so as not to have "1 Dec" find "11 Dec", "21 Dec", etc. Unfortunately some date headers have the leading zero and some don't, so with this version it can end up with a null result on either end. It's slow. Takes about 20 some odd seconds to complete. It will be slower still with the leading zero code.
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-12 15:17 +0000 |
| Message-ID | <fhjg0a9a.aer3@rooftop.invalid> |
| In reply to | #349494 |
owl <owl@rooftop.invalid> wrote:
> DFS <nospam@dfs.com> wrote:
>> On 4/11/2016 1:09 PM, Sandman wrote:
>>>
> ...
>>> In the end, traversing 500k lines of data is just as quick with todays CPU's.
>>
>> Sure, but where's the fun in that? I could download it once, put it in
>> a db table and query it right away. Even a Linux advocate could do that.
>>
>> But downloading and storing 500K rows in a table or in memory is
>> unacceptable if I want a portable system.
>>
>> What if I can write code that gets the answers quickly, doesn't require
>> a db, and allows a user to just type:
>>
>> -stats sci.math 30 days ending 20080930
>> -stats comp.os.linux.advocacy 7 days beginning 20160101
>>
>>
>> That's the ticket.
>>
>
> Like this?
>
> anon@lowtide:~$ ./blah.sh localhost cola 20150615 +6days
> Earliest date searched: 15 Jun 2015
> Latest date searched : 21 Jun 2015
> First ID in range: 2183608 Mon, 15 Jun 2015 00:43:10 +0200
> Last ID in range : 2185782 Sun, 21 Jun 2015 22:42:47 -0700
> anon@lowtide:~$
>
> anon@lowtide:~$ ./blah.sh localhost cola 20150615 -2months
> Earliest date searched: 15 Apr 2015
> Latest date searched : 15 Jun 2015
> First ID in range: 2173942 Wed, 15 Apr 2015 01:13:48 +0200
> Last ID in range : 2183869 Mon, 15 Jun 2015 21:36:04 -0700 (PDT)
> anon@lowtide:~$
>
> anon@lowtide:~$ ./blah.sh localhost alt.test 20151231 -1year
> Earliest date searched: 31 Dec 2014
> Latest date searched : 31 Dec 2015
> First ID in range: 4645896 Wed, 31 Dec 2014 00:08:41 +0000 (UTC)
> Last ID in range : 4747061 Thu, 31 Dec 2015 18:40:06 -0500
> anon@lowtide:~$
>
> I'll post the code later after I add some logic to handle leading
> zeros on the day portion of the date string. I have them stripped
> in this version and add a leading space to the date query so as not
> to have "1 Dec" find "11 Dec", "21 Dec", etc. Unfortunately some
> date headers have the leading zero and some don't, so with this version
> it can end up with a null result on either end.
>
> It's slow. Takes about 20 some odd seconds to complete. It will
> be slower still with the leading zero code.
>
OK. The code's below. Not doing every conceivable sanity check,
but probably enough unless you're trying to break it.
anon@lowtide:~$ ./blah.sh
usage: ./blah.sh <server> <newsgroup> <YYYYMMDD> [date offset]
anon@lowtide:~$
The optional date offset should be something like +1day, -2weeks
It depends on `expect`, so you may have to install that.
It gets server login credentials from .newsauth file (lines in this form):
serverA password username
serverB password username
serverC password username
Performance is horrible, but it seems to work OK.
You'll think it's hung, but it's not.
Takes anywhere from 20-50 seconds to return.
Oh yeah, almost forgot: Linux FTW. ;)
---------------------------------------------------------
#!/bin/bash
if [ ${#} -lt 3 ]; then
echo "usage: ${0} <server> <newsgroup> <YYYYMMDD> [date offset]"
exit
fi
if [ ${2} = "cola" ];then
GROUP="comp.os.linux.advocacy"
else
GROUP=${2}
fi
SERVER=${1}
LOGIN=$(grep ${1} ~/.newsauth |awk '{print $3}')
PASSWORD=$(grep ${1} ~/.newsauth |awk '{print $2}' |sed -e 's/\$/\\$/g')
END_ONE=$(date -d "${3}" "+ %-d %b %Y")
END_ONE_LEADING=$(date -d "${3}" "+%d %b %Y")
END_ONE_SECS=$(date -d "${3}" "+%s")
END_TWO=$(date -d "${3} ${4}" "+ %-d %b %Y")
END_TWO_LEADING=$(date -d "${3} ${4}" "+%d %b %Y")
END_TWO_SECS=$(date -d "${3} ${4}" "+%s")
if [ ${END_ONE_SECS} -lt ${END_TWO_SECS} ];then
EARLIEST=${END_ONE}
EARLIEST_LEADING=${END_ONE_LEADING}
LATEST=${END_TWO}
LATEST_LEADING=${END_TWO_LEADING}
else
EARLIEST=${END_TWO}
EARLIEST_LEADING=${END_TWO_LEADING}
LATEST=${END_ONE}
LATEST_LEADING=${END_ONE_LEADING}
fi
echo "Earliest date searched: ${EARLIEST}"
echo "Latest date searched : ${LATEST}"
GROUPINFO=$(tempfile)
START_DATE_LIST=$(tempfile)
START_DATE_LIST_LEADING=$(tempfile)
if [ ${#} -ne 4 ];then
END_DATE_LIST=${START_DATE_LIST}
END_DATE_LIST_LEADING=${START_DATE_LIST_LEADING}
else
END_DATE_LIST=$(tempfile)
END_DATE_LIST_LEADING=$(tempfile)
fi
expect -c "
spawn telnet ${SERVER} 119
expect \"200\"
send \"authinfo user ${LOGIN}\r\"
expect \"381 PASS required\"
send \"authinfo pass ${PASSWORD}\r\"
expect \"281 Ok\"
send \"group ${GROUP}\r\"
expect \"211\"
send \"quit\r\"
" > ${GROUPINFO}
LOWER=$(tail -n 1 ${GROUPINFO} | awk '{print $3}')
UPPER=$(tail -n 1 ${GROUPINFO} | awk '{print $4}')
expect -c "
spawn telnet ${SERVER} 119
expect \"200\"
send \"authinfo user ${LOGIN}\r\"
expect \"381 PASS required\"
send \"authinfo pass ${PASSWORD}\r\"
expect \"281 Ok\"
send \"group ${GROUP}\r\"
expect \"211\"
send \"xpat date ${LOWER}-${UPPER} *${EARLIEST}*\r\"
expect \"^205 .\"
send \"quit\r\"
" > ${START_DATE_LIST}
expect -c "
spawn telnet ${SERVER} 119
expect \"200\"
send \"authinfo user ${LOGIN}\r\"
expect \"381 PASS required\"
send \"authinfo pass ${PASSWORD}\r\"
expect \"281 Ok\"
send \"group ${GROUP}\r\"
expect \"211\"
send \"xpat date ${LOWER}-${UPPER} *${EARLIEST_LEADING}*\r\"
expect \"^205 .\"
send \"quit\r\"
" > ${START_DATE_LIST_LEADING}
if [ ${#} -eq 4 ]; then
expect -c "
spawn telnet ${SERVER} 119
expect \"200\"
send \"authinfo user ${LOGIN}\r\"
expect \"381 PASS required\"
send \"authinfo pass ${PASSWORD}\r\"
expect \"281 Ok\"
send \"group ${GROUP}\r\"
expect \"211\"
send \"xpat date ${LOWER}-${UPPER} *${LATEST}*\r\"
expect \"^205 .\"
send \"quit\r\"
" > ${END_DATE_LIST}
expect -c "
spawn telnet ${SERVER} 119
expect \"200\"
send \"authinfo user ${LOGIN}\r\"
expect \"381 PASS required\"
send \"authinfo pass ${PASSWORD}\r\"
expect \"281 Ok\"
send \"group ${GROUP}\r\"
expect \"211\"
send \"xpat date ${LOWER}-${UPPER} *${LATEST_LEADING}*\r\"
expect \"^205 .\"
send \"quit\r\"
" > ${END_DATE_LIST_LEADING}
fi
FIRST_IN_RANGE=$(grep "${EARLIEST}" ${START_DATE_LIST} |grep -v xpat | head -n 1)
FIRST_IN_RANGE_VAL=$(echo ${FIRST_IN_RANGE} | cut -f1 -d' ')
FIRST_IN_RANGE_LEADING=$(grep "${EARLIEST_LEADING}" ${START_DATE_LIST_LEADING} |grep -v xpat | head -n 1)
FIRST_IN_RANGE_LEADING_VAL=$(echo ${FIRST_IN_RANGE_LEADING} | cut -f1 -d' ')
LAST_IN_RANGE=$(grep "${LATEST}" ${END_DATE_LIST} |grep -v xpat |tail -n 1)
LAST_IN_RANGE_VAL=$(echo ${LAST_IN_RANGE} |cut -f1 -d' ')
LAST_IN_RANGE_LEADING=$(grep "${LATEST_LEADING}" ${END_DATE_LIST_LEADING} |grep -v xpat |tail -n 1)
LAST_IN_RANGE_LEADING_VAL=$(echo ${LAST_IN_RANGE_LEADING} |cut -f1 -d' ')
if [ ${FIRST_IN_RANGE_LEADING_VAL} -lt ${FIRST_IN_RANGE_VAL} ]; then
RANGE_BEGIN=${FIRST_IN_RANGE_LEADING}
else
RANGE_BEGIN=${FIRST_IN_RANGE}
fi
if [ ${LAST_IN_RANGE_LEADING_VAL} -gt ${LAST_IN_RANGE_VAL} ];then
RANGE_END=${LAST_IN_RANGE_LEADING}
else
RANGE_END=${LAST_IN_RANGE}
fi
echo "First ID in range: ${RANGE_BEGIN}"
echo "Last ID in range : ${RANGE_END}"
rm $GROUPINFO
rm $START_DATE_LIST
rm $START_DATE_LIST_LEADING
if [ ${#} -eq 4 ];then
rm $END_DATE_LIST
rm $END_DATE_LIST_LEADING
fi
---------------------------------------------------------
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | comp.os.linux.advocacy
csiph-web