Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.os.linux.advocacy > #349365 > unrolled thread
| Started by | DFS <nospam@dfs.com> |
|---|---|
| First post | 2016-04-11 11:23 -0400 |
| Last post | 2016-06-14 00:48 +0000 |
| Articles | 20 on this page of 52 — 14 participants |
Back to article view | Back to comp.os.linux.advocacy
Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 11:23 -0400
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-11 17:09 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-11 19:25 +0000
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-11 20:58 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 17:38 -0400
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 18:11 -0400
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-11 22:25 +0000
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-12 06:16 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-11 21:46 +0000
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-12 07:18 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-12 08:15 +0000
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-12 10:34 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 13:32 -0400
Re: Algorithm to find data range Steve Carroll <fretwizzer@gmail.com> - 2016-04-12 10:39 -0700
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 13:58 -0400
Re: Algorithm to find data range Steve Carroll <fretwizzer@gmail.com> - 2016-04-12 11:59 -0700
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 23:23 -0400
Re: Algorithm to find data range Sandman <mr@sandman.net> - 2016-04-12 07:14 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-12 10:44 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-12 15:17 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-13 17:15 -0400
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-13 22:23 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-13 18:32 -0400
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-14 04:33 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-14 05:51 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-14 20:15 -0400
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-15 01:18 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 13:25 -0400
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-12 19:08 +0000
Re: Algorithm to find data range vallor <vallor@cultnix.org> - 2016-04-12 04:40 +0000
Re: Algorithm to find data range owl <owl@rooftop.invalid> - 2016-04-11 18:06 +0000
Re: Algorithm to find data range 7 <7@enemygadgets.com> - 2016-04-11 22:46 +0000
Re: Algorithm to find data range Omar <omarsayeed@linuxmail.org> - 2016-04-11 18:52 -0400
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 13:28 -0400
Re: Algorithm to find data range Omar <omarsayeed@linuxmail.org> - 2016-04-12 13:46 -0400
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 22:13 -0400
Re: Algorithm to find data range Fabian Russell <fb@zen.info> - 2016-04-11 23:22 +0000
Re: Algorithm to find data range vallor <vallor@cultnix.org> - 2016-04-12 00:06 +0000
Re: Algorithm to find data range Omar <omarsayeed@linuxmail.org> - 2016-04-11 20:26 -0400
Re: Algorithm to find data range Chris Ahlstrom <OFeem1987@teleworm.us> - 2016-04-12 05:30 -0400
Re: Algorithm to find data range chrisv <chrisv@nospam.invalid> - 2016-04-12 06:50 -0500
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-11 22:34 -0400
Re: Algorithm to find data range Fabian Russell <fb@zen.info> - 2016-04-12 09:25 +0000
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 13:32 -0400
Is it your own personal NNTP/Usenet server ? Jeff-Relf.Me <@.> - 2016-04-11 16:32 -0700
Re: Algorithm to find data range DFS <nospam@dfs.com> - 2016-04-12 12:47 -0400
Re: Algorithm to find data range Peter Köhlmann <peter-koehlmann@t-online.de> - 2016-04-12 20:46 +0200
Re: Algorithm to find data range chrisv <chrisv@nospam.invalid> - 2016-04-12 13:59 -0500
Re: Algorithm to find data range Silver Slimer <linux@sucks.balls> - 2016-04-12 17:08 -0400
My "newsReader" (X.ZIP) is also a console. Jeff-Relf.Me <@.> - 2016-04-12 12:27 -0700
My "newsReader" (X.ZIP) is also a console. Jeff-Relf.Me <@.> - 2016-04-12 12:31 -0700
Re: My "newsReader" (X.ZIP) is also a console. Unknown <dog@gmail.com> - 2016-06-14 00:48 +0000
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-13 17:15 -0400 |
| Message-ID | <nemcmq$b8c$4@dont-email.me> |
| In reply to | #349521 |
On 4/12/2016 11:17 AM, owl wrote: > Performance is horrible, but it seems to work OK. > You'll think it's hung, but it's not. > Takes anywhere from 20-50 seconds to return. Sweet! I like a nice, relaxed program that finishes when it's damn good and ready. > Oh yeah, almost forgot: Linux FTW. ;) No doubt. Windows could never be that slow. > #!/bin/bash <snip> bash produces such beautiful code! You sent, you expected, you spawned, and you echoed... are you a dolphin?
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-13 22:23 +0000 |
| Message-ID | <hjgi403.uut83gko@rooftop.invalid> |
| In reply to | #349709 |
DFS <nospam@dfs.com> wrote: > On 4/12/2016 11:17 AM, owl wrote: > > >> Performance is horrible, but it seems to work OK. >> You'll think it's hung, but it's not. >> Takes anywhere from 20-50 seconds to return. > > Sweet! I like a nice, relaxed program that finishes when it's damn good > and ready. > > >> Oh yeah, almost forgot: Linux FTW. ;) > > No doubt. Windows could never be that slow. > > > >> #!/bin/bash > <snip> > > > bash produces such beautiful code! > > You sent, you expected, you spawned, and you echoed... are you a dolphin? > If fixed it. It's a lot faster now. Also added option to set timeout in case the server is slow to return. (Default timeout is set to 1 sec otherwise). anon@lowtide:~/code/usenet$ time ./blah.sh localhost cola 20150515 +180days Earliest date searched: 15 May 2015 Latest date searched : 11 Nov 2015 First ID in range: 2178717 Fri, 15 May 2015 02:19:47 +0100 (BST) Last ID in range : 2217475 Wed, 11 Nov 2015 21:33:03 -0600 real 0m4.276s user 0m0.028s sys 0m0.032s anon@lowtide:~/code/usenet$ time ./blah.sh localhost cola 20150515 Earliest date searched: 15 May 2015 Latest date searched : 15 May 2015 First ID in range: 2178717 Fri, 15 May 2015 02:19:47 +0100 (BST) Last ID in range : 2178859 Fri, 15 May 2015 22:41:01 -0700 real 0m2.775s user 0m0.032s sys 0m0.028s anon@lowtide:~/code/usenet$
[toc] | [prev] | [next] | [standalone]
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-13 18:32 -0400 |
| Message-ID | <nemh84$vf4$1@dont-email.me> |
| In reply to | #349718 |
On 4/13/2016 6:23 PM, owl wrote: > DFS <nospam@dfs.com> wrote: >> On 4/12/2016 11:17 AM, owl wrote: >> >> >>> Performance is horrible, but it seems to work OK. >>> You'll think it's hung, but it's not. >>> Takes anywhere from 20-50 seconds to return. >> >> Sweet! I like a nice, relaxed program that finishes when it's damn good >> and ready. >> >> >>> Oh yeah, almost forgot: Linux FTW. ;) >> >> No doubt. Windows could never be that slow. >> >> >> >>> #!/bin/bash >> <snip> >> >> >> bash produces such beautiful code! >> >> You sent, you expected, you spawned, and you echoed... are you a dolphin? >> > > If fixed it. It's a lot faster now. > Also added option to set timeout in case the server is slow to return. > (Default timeout is set to 1 sec otherwise). > > anon@lowtide:~/code/usenet$ time ./blah.sh localhost cola 20150515 +180days > Earliest date searched: 15 May 2015 > Latest date searched : 11 Nov 2015 > First ID in range: 2178717 Fri, 15 May 2015 02:19:47 +0100 (BST) > Last ID in range : 2217475 Wed, 11 Nov 2015 21:33:03 -0600 > > real 0m4.276s > user 0m0.028s > sys 0m0.032s > anon@lowtide:~/code/usenet$ time ./blah.sh localhost cola 20150515 > Earliest date searched: 15 May 2015 > Latest date searched : 15 May 2015 > First ID in range: 2178717 Fri, 15 May 2015 02:19:47 +0100 (BST) > Last ID in range : 2178859 Fri, 15 May 2015 22:41:01 -0700 > > real 0m2.775s > user 0m0.032s > sys 0m0.028s > anon@lowtide:~/code/usenet$ Good job. 2.77 secs, that's something to shoot for. I can't (or can't figure out how to) use xpat with python, so I'll try different methods to get an ID range given a date range. Discovered as of last night on eternal-september: - article ID range is 1 to 578254 - article count is 368130 - huge gap: nothing from ID 1 (Nov 2007) to ID 201860 (Oct 2010)
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-14 04:33 +0000 |
| Message-ID | <nmbjd83.klguey@rooftop.invalid> |
| In reply to | #349719 |
DFS <nospam@dfs.com> wrote:
> On 4/13/2016 6:23 PM, owl wrote:
>> DFS <nospam@dfs.com> wrote:
>>> On 4/12/2016 11:17 AM, owl wrote:
>>>
>>>
>>>> Performance is horrible, but it seems to work OK.
>>>> You'll think it's hung, but it's not.
>>>> Takes anywhere from 20-50 seconds to return.
>>>
>>> Sweet! I like a nice, relaxed program that finishes when it's damn good
>>> and ready.
>>>
>>>
>>>> Oh yeah, almost forgot: Linux FTW. ;)
>>>
>>> No doubt. Windows could never be that slow.
>>>
>>>
>>>
>>>> #!/bin/bash
>>> <snip>
>>>
>>>
>>> bash produces such beautiful code!
>>>
>>> You sent, you expected, you spawned, and you echoed... are you a dolphin?
>>>
>>
>> If fixed it. It's a lot faster now.
>> Also added option to set timeout in case the server is slow to return.
>> (Default timeout is set to 1 sec otherwise).
>>
>> anon@lowtide:~/code/usenet$ time ./blah.sh localhost cola 20150515 +180days
>> Earliest date searched: 15 May 2015
>> Latest date searched : 11 Nov 2015
>> First ID in range: 2178717 Fri, 15 May 2015 02:19:47 +0100 (BST)
>> Last ID in range : 2217475 Wed, 11 Nov 2015 21:33:03 -0600
>>
>> real 0m4.276s
>> user 0m0.028s
>> sys 0m0.032s
>> anon@lowtide:~/code/usenet$ time ./blah.sh localhost cola 20150515
>> Earliest date searched: 15 May 2015
>> Latest date searched : 15 May 2015
>> First ID in range: 2178717 Fri, 15 May 2015 02:19:47 +0100 (BST)
>> Last ID in range : 2178859 Fri, 15 May 2015 22:41:01 -0700
>>
>> real 0m2.775s
>> user 0m0.032s
>> sys 0m0.028s
>> anon@lowtide:~/code/usenet$
>
>
> Good job. 2.77 secs, that's something to shoot for.
>
> I can't (or can't figure out how to) use xpat with python, so I'll try
> different methods to get an ID range given a date range.
>
Does it not give you some way to write arbitrary string to the socket?
The server doesn't know what's connecting. It sends the same reponses to
whatever connects (telnet, netcat, your app, whatever). All you need is
a way to send commands and read its responses. If the python nntplib lets
you get the available range of IDs, then you have that range to build your
command string and send it with whatever write() function your language
provides.
Like the way I do it with telnet/expect:
send \"group ${GROUP}\r\"
expect \"211 \"
send \"${XPAT_EARLIEST}\r\"
The "${XPAT_EARLIEST}" is just a string like:
"xpat date 1111111-2222222 * [^0-3]2 May 2015"
(for day number < 10)
or
"xpat date 1111111-2222222 * 12 May 2015"
(for day number >= 10)
> Discovered as of last night on eternal-september:
> - article ID range is 1 to 578254
> - article count is 368130
> - huge gap: nothing from ID 1 (Nov 2007) to ID 201860 (Oct 2010)
>
5.5 years is still a nice batch available.
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-14 05:51 +0000 |
| Message-ID | <hjgo30a.rt4aef@rooftop.invalid> |
| In reply to | #349775 |
owl <owl@rooftop.invalid> wrote:
> DFS <nospam@dfs.com> wrote:
>> On 4/13/2016 6:23 PM, owl wrote:
>>> DFS <nospam@dfs.com> wrote:
>>>> On 4/12/2016 11:17 AM, owl wrote:
>>>>
>>>>
>>>>> Performance is horrible, but it seems to work OK.
>>>>> You'll think it's hung, but it's not.
>>>>> Takes anywhere from 20-50 seconds to return.
>>>>
>>>> Sweet! I like a nice, relaxed program that finishes when it's damn good
>>>> and ready.
>>>>
>>>>
>>>>> Oh yeah, almost forgot: Linux FTW. ;)
>>>>
>>>> No doubt. Windows could never be that slow.
>>>>
>>>>
>>>>
>>>>> #!/bin/bash
>>>> <snip>
>>>>
>>>>
>>>> bash produces such beautiful code!
>>>>
>>>> You sent, you expected, you spawned, and you echoed... are you a dolphin?
>>>>
>>>
>>> If fixed it. It's a lot faster now.
>>> Also added option to set timeout in case the server is slow to return.
>>> (Default timeout is set to 1 sec otherwise).
>>>
>>> anon@lowtide:~/code/usenet$ time ./blah.sh localhost cola 20150515 +180days
>>> Earliest date searched: 15 May 2015
>>> Latest date searched : 11 Nov 2015
>>> First ID in range: 2178717 Fri, 15 May 2015 02:19:47 +0100 (BST)
>>> Last ID in range : 2217475 Wed, 11 Nov 2015 21:33:03 -0600
>>>
>>> real 0m4.276s
>>> user 0m0.028s
>>> sys 0m0.032s
>>> anon@lowtide:~/code/usenet$ time ./blah.sh localhost cola 20150515
>>> Earliest date searched: 15 May 2015
>>> Latest date searched : 15 May 2015
>>> First ID in range: 2178717 Fri, 15 May 2015 02:19:47 +0100 (BST)
>>> Last ID in range : 2178859 Fri, 15 May 2015 22:41:01 -0700
>>>
>>> real 0m2.775s
>>> user 0m0.032s
>>> sys 0m0.028s
>>> anon@lowtide:~/code/usenet$
>>
>>
>> Good job. 2.77 secs, that's something to shoot for.
>>
>> I can't (or can't figure out how to) use xpat with python, so I'll try
>> different methods to get an ID range given a date range.
>>
>
> Does it not give you some way to write arbitrary string to the socket?
> The server doesn't know what's connecting. It sends the same reponses to
> whatever connects (telnet, netcat, your app, whatever). All you need is
> a way to send commands and read its responses. If the python nntplib lets
> you get the available range of IDs, then you have that range to build your
> command string and send it with whatever write() function your language
> provides.
>
I don't do python at all, but the following worked for passing the xpat.
I had already run once with just group command to get the range.
Of course, you'll need to supply the correct credentials and server.
#!/usr/bin/python
import socket
s=socket.socket(socket.AF_INET,socket.SOCK_STREAM)
s.connect(("localhost",119))
print s.recv(8192)
s.send("authinfo user USERNAME\n")
print s.recv(8192)
s.send("authinfo pass PASSWORD\n")
print s.recv(8192)
s.send("group comp.os.linux.advocacy\n")
print s.recv(8192)
s.send("xpat date 2152569-2234805 * 13 Apr 2016*\n")
print s.recv(8192)
print s.recv(8192)
s.close()
[toc] | [prev] | [next] | [standalone]
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-14 20:15 -0400 |
| Message-ID | <nepblm$66a$1@dont-email.me> |
| In reply to | #349778 |
On 4/14/2016 1:51 AM, owl wrote:
> I don't do python at all,
But you did. How did you know to write that code?
> but the following worked for passing the xpat.
> I had already run once with just group command to get the range.
> Of course, you'll need to supply the correct credentials and server.
> import socket
>
> s=socket.socket(socket.AF_INET,socket.SOCK_STREAM)
> s.connect(("localhost",119))
> print s.recv(8192)
> s.send("authinfo user USERNAME\n")
> print s.recv(8192)
> s.send("authinfo pass PASSWORD\n")
> print s.recv(8192)
> s.send("group comp.os.linux.advocacy\n")
> print s.recv(8192)
> s.send("xpat date 2152569-2234805 * 13 Apr 2016*\n")
> print s.recv(8192)
> print s.recv(8192)
> s.close()
Nice! Gave me 4 responses:
* 200 mx02.eternal-september.org InterNetNews NNRP server INN 2.7.0
(20151029 snapshot) ready (posting ok)
* 381 Enter password
* 281 Authentication succeeded
* 211 368279 1 578403 comp.os.linux.advocacy
Then I submitted
s.send("xpat date 578000-578400 *Apr 2016*\n")
print s.recv(8192)
print s.recv(8192)
and it gave dozens of hits, but the last one was cut off
578197 Tue, 12 Apr 2016 20:43:51 +0200
578198 Tue, 12 Apr 201
I increased the recv buffer to 16384 but same thing.
Some setting I need to make?
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-15 01:18 +0000 |
| Message-ID | <nbvmcjuf83.jilgda@rooftop.invalid> |
| In reply to | #349934 |
DFS <nospam@dfs.com> wrote:
> On 4/14/2016 1:51 AM, owl wrote:
>
>> I don't do python at all,
>
> But you did. How did you know to write that code?
>
>
>
>> but the following worked for passing the xpat.
>> I had already run once with just group command to get the range.
>> Of course, you'll need to supply the correct credentials and server.
>
>> import socket
>>
>> s=socket.socket(socket.AF_INET,socket.SOCK_STREAM)
>> s.connect(("localhost",119))
>> print s.recv(8192)
>> s.send("authinfo user USERNAME\n")
>> print s.recv(8192)
>> s.send("authinfo pass PASSWORD\n")
>> print s.recv(8192)
>> s.send("group comp.os.linux.advocacy\n")
>> print s.recv(8192)
>> s.send("xpat date 2152569-2234805 * 13 Apr 2016*\n")
>> print s.recv(8192)
>> print s.recv(8192)
>> s.close()
>
>
> Nice! Gave me 4 responses:
>
> * 200 mx02.eternal-september.org InterNetNews NNRP server INN 2.7.0
> (20151029 snapshot) ready (posting ok)
> * 381 Enter password
> * 281 Authentication succeeded
> * 211 368279 1 578403 comp.os.linux.advocacy
>
>
> Then I submitted
>
> s.send("xpat date 578000-578400 *Apr 2016*\n")
> print s.recv(8192)
> print s.recv(8192)
>
>
> and it gave dozens of hits, but the last one was cut off
>
> 578197 Tue, 12 Apr 2016 20:43:51 +0200
> 578198 Tue, 12 Apr 201
>
> I increased the recv buffer to 16384 but same thing.
>
> Some setting I need to make?
>
Probably really need to call recv() in a loop and buffer the data.
Take a look at this and other similar solutions:
http://code.activestate.com/recipes/408859-socketrecv-three-ways-to-turn-it-into-recvall/
The way I was doing it with bash/expect sometimes required that I set
the timeout higher to get the data. Fortunately, I would get null return
if expect did not get a match before the timeout, so it was obvious that
I needed to increase.
As basic and predicatable as this data is, you may can get by with
just multiple recv calls like below. Take note of the wildpat [^1-3]
for capturing day values less than 10. That is necessary or else it
won't get them all (i.e., "04" will not get "4" and "4" will not get
"04" values). In real code, of course, you'll need to calculate date
differences, grab the earliest ID from the initial date file, the latest
date from the offset date file, etc.
---------------------------
#!/usr/bin/python
import os
import socket
from tempfile import mkstemp
fd1, temp_path1 = mkstemp()
fd2, temp_path2 = mkstemp()
print "orig date IDs in here " + temp_path1
print "offset date IDs in here " + temp_path2
s=socket.socket(socket.AF_INET,socket.SOCK_STREAM)
s.connect(("SERVER",119))
print s.recv(8192)
s.send("authinfo user USERNAME\n")
print s.recv(8192)
s.send("authinfo pass PASSWORD\n")
print s.recv(8192)
s.send("group comp.os.linux.advocacy\n")
print s.recv(8192)
s.send("xpat date 2152569-2234805 *[^1-3]2 Jul 2015*\n")
os.write(fd1,s.recv(32768))
os.write(fd1,s.recv(32768))
os.write(fd1,s.recv(32768))
os.close(fd1)
s.send("xpat date 2152569-2234805 *[^1-3]4 Mar 2016*\n")
os.write(fd2,s.recv(32768))
os.write(fd2,s.recv(32768))
os.write(fd2,s.recv(32768))
os.close(fd2)
s.close()
---------------------------
anon@lowtide:~/code/usenet$ time ./flah.py
orig date IDs in here /tmp/tmpjIe6Fq
offset date IDs in here /tmp/tmpC79Jjs
200 Check out http://www.altopia.com/ for info about NNTP access (posting ok).
381 PASS required
281 Ok
211 81986 2152993 2234978 comp.os.linux.advocacy
real 0m2.568s
user 0m0.044s
sys 0m0.016s
anon@lowtide:~/code/usenet$
anon@lowtide:~/code/usenet$ head /tmp/tmpjIe6Fq
221 date matches follow (NOV)
2187893 Thu, 02 Jul 2015 00:41:19 +0200
2187900 Thu, 02 Jul 2015 01:01:17 +0200
2187901 Thu, 02 Jul 2015 01:03:01 +0200
2187922 Thu, 2 Jul 2015 06:42:17 +0200
2187923 Thu, 2 Jul 2015 06:53:11 +0200
2187926 Thu, 2 Jul 2015 07:29:27 +0200
2187928 Thu, 2 Jul 2015 01:31:27 -0700 (PDT)
2187929 Thu, 2 Jul 2015 11:09:51 +0200
2187930 Thu, 2 Jul 2015 05:52:49 -0400
anon@lowtide:~/code/usenet$
anon@lowtide:~/code/usenet$ tail /tmp/tmpC79Jjs
2230279 Fri, 04 Mar 2016 17:35:51 -0800 (Seattle)
2230282 Fri, 4 Mar 2016 20:52:04 -0500
2230284 Fri, 4 Mar 2016 19:50:29 -0600
2230285 Fri, 4 Mar 2016 20:59:46 -0500
2230288 Fri, 4 Mar 2016 18:18:51 -0800 (PST)
2230292 Fri, 4 Mar 2016 21:44:11 -0500
2230293 Fri, 4 Mar 2016 21:46:07 -0500
2230294 Fri, 4 Mar 2016 21:50:04 -0500
2230296 Fri, 4 Mar 2016 21:10:53 -0600
2230299 Fri, 04 Mar 2016 19:41:13 -0800 (Seattle)
anon@lowtide:~/code/usenet$
Here you would want to parse out the 2187893 from /tmp/tmpjIe6Fq and
2230299 from /tmp/tmpC79Jjs.
[toc] | [prev] | [next] | [standalone]
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-12 13:25 -0400 |
| Message-ID | <nejarm$p7n$2@dont-email.me> |
| In reply to | #349494 |
On 4/12/2016 6:44 AM, owl wrote: > DFS <nospam@dfs.com> wrote: >> On 4/11/2016 1:09 PM, Sandman wrote: >>> > ... >>> In the end, traversing 500k lines of data is just as quick with todays CPU's. >> >> Sure, but where's the fun in that? I could download it once, put it in >> a db table and query it right away. Even a Linux advocate could do that. >> >> But downloading and storing 500K rows in a table or in memory is >> unacceptable if I want a portable system. >> >> What if I can write code that gets the answers quickly, doesn't require >> a db, and allows a user to just type: >> >> -stats sci.math 30 days ending 20080930 >> -stats comp.os.linux.advocacy 7 days beginning 20160101 >> >> >> That's the ticket. >> > > Like this? Much much much more than that. It looks like you're just finding the IDs corresponding to the date range entered (which is what this thread is about). When I/we find a good algorithm, I expect that part to take no more than 3 seconds no matter how many days are covered. My -stats will be a version of the weekly stats we've seen here for years. I'll release the code and anyone can use it to summarize any group over any time period (where any is some arbitrary limit you choose - typically up to one month. You don't want to hog server cycles. And apparently some servers limit your data downloads.) > anon@lowtide:~$ ./blah.sh localhost cola 20150615 +6days > Earliest date searched: 15 Jun 2015 > Latest date searched : 21 Jun 2015 > First ID in range: 2183608 Mon, 15 Jun 2015 00:43:10 +0200 > Last ID in range : 2185782 Sun, 21 Jun 2015 22:42:47 -0700 > anon@lowtide:~$ > > anon@lowtide:~$ ./blah.sh localhost cola 20150615 -2months > Earliest date searched: 15 Apr 2015 > Latest date searched : 15 Jun 2015 > First ID in range: 2173942 Wed, 15 Apr 2015 01:13:48 +0200 > Last ID in range : 2183869 Mon, 15 Jun 2015 21:36:04 -0700 (PDT) > anon@lowtide:~$ > > anon@lowtide:~$ ./blah.sh localhost alt.test 20151231 -1year > Earliest date searched: 31 Dec 2014 > Latest date searched : 31 Dec 2015 > First ID in range: 4645896 Wed, 31 Dec 2014 00:08:41 +0000 (UTC) > Last ID in range : 4747061 Thu, 31 Dec 2015 18:40:06 -0500 > anon@lowtide:~$ > > I'll post the code later after I add some logic to handle leading > zeros on the day portion of the date string. I have them stripped > in this version and add a leading space to the date query so as not > to have "1 Dec" find "11 Dec", "21 Dec", etc. Unfortunately some > date headers have the leading zero and some don't, so with this version > it can end up with a null result on either end. > > It's slow. Takes about 20 some odd seconds to complete. It will > be slower still with the leading zero code. 20 seconds seems way slow. It's probably not in your code. xpat? If xpat is this slow, it won't be the ticket! I have a big python module (uses nntplib, xover, xhdr mostly) that does tons of processing and even when run against 5000 posts it runs in 2-3 seconds.
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-12 19:08 +0000 |
| Message-ID | <ghjdi03a.f3@rooftop.invalid> |
| In reply to | #349539 |
DFS <nospam@dfs.com> wrote: > On 4/12/2016 6:44 AM, owl wrote: >> DFS <nospam@dfs.com> wrote: >>> On 4/11/2016 1:09 PM, Sandman wrote: >>>> >> ... >>>> In the end, traversing 500k lines of data is just as quick with todays CPU's. >>> >>> Sure, but where's the fun in that? I could download it once, put it in >>> a db table and query it right away. Even a Linux advocate could do that. >>> >>> But downloading and storing 500K rows in a table or in memory is >>> unacceptable if I want a portable system. >>> >>> What if I can write code that gets the answers quickly, doesn't require >>> a db, and allows a user to just type: >>> >>> -stats sci.math 30 days ending 20080930 >>> -stats comp.os.linux.advocacy 7 days beginning 20160101 >>> >>> >>> That's the ticket. >>> >> >> Like this? > > Much much much more than that. It looks like you're just finding the > IDs corresponding to the date range entered (which is what this thread > is about). When I/we find a good algorithm, I expect that part to take > no more than 3 seconds no matter how many days are covered. > > My -stats will be a version of the weekly stats we've seen here for > years. I'll release the code and anyone can use it to summarize any > group over any time period (where any is some arbitrary limit you choose > - typically up to one month. You don't want to hog server cycles. And > apparently some servers limit your data downloads.) > > >> anon@lowtide:~$ ./blah.sh localhost cola 20150615 +6days >> Earliest date searched: 15 Jun 2015 >> Latest date searched : 21 Jun 2015 >> First ID in range: 2183608 Mon, 15 Jun 2015 00:43:10 +0200 >> Last ID in range : 2185782 Sun, 21 Jun 2015 22:42:47 -0700 >> anon@lowtide:~$ >> >> anon@lowtide:~$ ./blah.sh localhost cola 20150615 -2months >> Earliest date searched: 15 Apr 2015 >> Latest date searched : 15 Jun 2015 >> First ID in range: 2173942 Wed, 15 Apr 2015 01:13:48 +0200 >> Last ID in range : 2183869 Mon, 15 Jun 2015 21:36:04 -0700 (PDT) >> anon@lowtide:~$ >> >> anon@lowtide:~$ ./blah.sh localhost alt.test 20151231 -1year >> Earliest date searched: 31 Dec 2014 >> Latest date searched : 31 Dec 2015 >> First ID in range: 4645896 Wed, 31 Dec 2014 00:08:41 +0000 (UTC) >> Last ID in range : 4747061 Thu, 31 Dec 2015 18:40:06 -0500 >> anon@lowtide:~$ >> >> I'll post the code later after I add some logic to handle leading >> zeros on the day portion of the date string. I have them stripped >> in this version and add a leading space to the date query so as not >> to have "1 Dec" find "11 Dec", "21 Dec", etc. Unfortunately some >> date headers have the leading zero and some don't, so with this version >> it can end up with a null result on either end. >> >> It's slow. Takes about 20 some odd seconds to complete. It will >> be slower still with the leading zero code. > > > 20 seconds seems way slow. It's probably not in your code. xpat? If > xpat is this slow, it won't be the ticket! > Nah, it's `expect` that's slowing it. Didn't realize it has a default 10 sec timeout. I set it to 1 and speed goes from ~45 sec on the offset search to about 9 sec; and from 20 some to about 5 sec on a single day search. I played with the code a bit more, and for the life of me I cannot get it to accept a [^1-3]N pattern in the handoff to an `expect` send. When it's echo'ed it shows exactly as I would type it at the console, but for some reason it wants to treat the pattern as a command. No amount of escaping seems to help. It's unfortunate, because this would cut out the two extra `expect` runs that I have to do to handle the leading zeros issue. > I have a big python module (uses nntplib, xover, xhdr mostly) that does > tons of processing and even when run against 5000 posts it runs in 2-3 > seconds. > From telnet, I get similar. Hell, I could probably manually go through the steps at the console as fast as I can automate it with expect.
[toc] | [prev] | [next] | [standalone]
| From | vallor <vallor@cultnix.org> |
|---|---|
| Date | 2016-04-12 04:40 +0000 |
| Message-ID | <dn3ci9Fq114U2@mid.individual.net> |
| In reply to | #349378 |
On Mon, 11 Apr 2016 17:09:19 +0000, Sandman wrote:
> In article <negfcb$4ee$1@dont-email.me>, DFS wrote:
>
>> Take a list of data:
>
>> ID Date 1 2004-05-04 2 2004-05-04 3 2004-05-05 4 2004-05-06 5
>> 2004-05-08 ...
>> 500000 2014-05-03
>
>> I want to find the best approach to determine the unknown range of ID
>> numbers that correspond to a known range of dates.
>
>> That is: * you know the date range you're interested in (say all of Sep
>> 2008) * you want to know the corresponding ID range (say 124385 to
>> 125008).
>
>> The problem is you can't query or retrieve the data by date, only by ID
>> number. This restriction is what makes the whole thing an ordeal.
>
>> Other constraints:
>
>> * data is pulled off a busy server. You can't be hitting it all day
>> long.
>
>> * you can retrieve and examine max of 100 rows at a time. This is a
>> judgement call. I'll try 100 and see what works best, and increase or
>> decrease it based on performance.
>
>> * the numbers and dates are ordered low to high, but not continuous -
>> there are gaps in both numbers and dates.
>
>> * can't load the data into a SQL table and do min() or max(), so it's a
>> code-only solution (python here).
>
> That's some strange limitations.
>
>> I was thinking about three approaches that I call:
>> ---------------------------------------------------------------------
>
>> 1. Half-Height Elimination: examine 1 row at a time (starting with the
>> middle row), cutting the data to be examined in half on each iteration.
>> This will require some kind of recursive coding methodology. With a
>> domain of 500K rows, this approach could require as many as 19
>> iterations (and each iteration would involve a small request from the
>> server) in its simplest form:
>
>> This pic will help to see how it works: http://i.imgur.com/AFfIjrI.png
>
> That supposes that for each cut, the date range is in the cut. I.e. it
> could just as easily be twice as many cuts.
>
> Also, given restriction 3 above, there is no way for you to know what ID
> is at the middle of a sample. If you have 500k rows of data, you'd think
> that ID 250,000 would be in the middle, but since it was stipulated that
> ID could contain caps and not be continuous, you may have cut two thirds
> up in the series. In fact, ID 250,000 may be the very last ID according
> to rule #3. Or the first, for that matter.
Two things:
A binary seek doesn't need recursion, just two loops.
First loop needs a left bracket, a right bracket, and a test point (which
I like to call the "gripping bracket").
$lb,$rb,$gb
set $lb to position 0, $rb to end of range
LOOP1:
$gb = split the difference
if $gb is within the desired range, break and do the next loop
if the desired range is to the left (msg-id less than) $gb,
then make $gb = $rb
similarly, if the desired range is to the right of $gb, make $lb = $gb
LOOP2:
this is where you have found the "neighborhood" of your data. At this
point you'll have to use heuristics to figure out the range, possibly by
making guesses based on the average number of posts within a time period.
> In the end, traversing 500k lines of data is just as quick with todays
> CPU's.
It sounds like he's building leet spy tools that won't hammer his nntp
server. The XOVER (or the new "OVER") NNTP command can grab the metadata
for a group, and servers are optimized to hand this data over very
quickly. Thinking in those terms might yield more fruitful results. For
example (because I'm a perl guy), I'd be looking into:
https://tools.ietf.org/html/rfc3977#page-81
and
https://metacpan.org/pod/Net::NNTP
BTW: back when Webster servers were in short supply (c. 1996), I wrote a
tool to perform a dictionary which included a binary seek.
Here is what the output looks like when I run the debugging version:
/opt/sdict/lib]_(scott@xxx)_
$ ./show.pl jerk
{ * } span
0 187018 374036 374036
0 93509 187018 187018
93509 140263 187018 93509
140263 163640 187018 46755
163640 175329 187018 23378
175329 181173 187018 11689
175329 178251 181173 5844
178251 179712 181173 2922
178251 178981 179712 1461
178251 178616 178981 730
1 jerk \'j*rk\ vb
1: to give a sharp quick push, pull, or twist
2: to move in short abrupt motions
2 jerk n
1: a short quick pull or twist : TWITCH
2: a stupid, foolish, or eccentric person
-- jerk.i.ly adv
-- jerky adj
--
-v
Kernel:4.6.0-rc2-sd Desktop:Xfce 4.12.2 Distro:Linux Mint 17.3 Rosa
[toc] | [prev] | [next] | [standalone]
| From | owl <owl@rooftop.invalid> |
|---|---|
| Date | 2016-04-11 18:06 +0000 |
| Message-ID | <hgjei903a.jgi3@rooftop.invalid> |
| In reply to | #349365 |
DFS <nospam@dfs.com> wrote: > Take a list of data: > > ID Date > 1 2004-05-04 > 2 2004-05-04 > 3 2004-05-05 > 4 2004-05-06 > 5 2004-05-08 > ... > 500000 2014-05-03 > > I want to find the best approach to determine the unknown range of ID > numbers that correspond to a known range of dates. > > That is: > * you know the date range you're interested in (say all of Sep 2008) > * you want to know the corresponding ID range (say 124385 to 125008). > If this is usenet, you can use the xpat command. (And damn you if you have access to ten years worth of articles). from a `telnet <server> 119` prompt: xpat <header> <article range> <pattern> From my server, searching for old articles with a date pattern: https://vid.me/Y9rU It's not perfect, but it works ok.
[toc] | [prev] | [next] | [standalone]
| From | 7 <7@enemygadgets.com> |
|---|---|
| Date | 2016-04-11 22:46 +0000 |
| Message-ID | <dn2nq9Fm358U1@mid.individual.net> |
| In reply to | #349365 |
DFS wrote: > Take a list of data: > > ID Date > 1 2004-05-04 > 2 2004-05-04 > 3 2004-05-05 > 4 2004-05-06 > 5 2004-05-08 > ... > 500000 2014-05-03 > > I want to find the best approach to determine the unknown range of ID > numbers that correspond to a known range of dates. What a fscking pussy! Here pussy want to learn some real computing do you? ;) Sorry to hear microshat and appil did not educt you in matters of real computing. Various thoughts of not allowing trolls like you near computers had crossed my mind reading your long failed thesis on how micorhsafteess and appile retards solve problems with their illiterate ways. The solution is one paragraph or less with Linux.
[toc] | [prev] | [next] | [standalone]
| From | Omar <omarsayeed@linuxmail.org> |
|---|---|
| Date | 2016-04-11 18:52 -0400 |
| Message-ID | <neh9lu$hn1$1@dont-email.me> |
| In reply to | #349414 |
On 11 Apr 2016 22:46:33 GMT, 7 wrote: > DFS wrote: > >> Take a list of data: >> >> ID Date >> 1 2004-05-04 >> 2 2004-05-04 >> 3 2004-05-05 >> 4 2004-05-06 >> 5 2004-05-08 >> ... >> 500000 2014-05-03 >> >> I want to find the best approach to determine the unknown range of ID >> numbers that correspond to a known range of dates. > > What a fscking pussy! > > Here pussy want to learn some real computing do you? ;) > > Sorry to hear microshat and appil did not educt you in matters > of real computing. > > Various thoughts of not allowing trolls like you near computers > had crossed my mind reading your long failed thesis on > how micorhsafteess and appile retards solve problems with > their illiterate ways. > > The solution is one paragraph or less with Linux. So why don't you show us?
[toc] | [prev] | [next] | [standalone]
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-12 13:28 -0400 |
| Message-ID | <nejb2h$r1o$1@dont-email.me> |
| In reply to | #349415 |
On 4/11/2016 6:52 PM, Omar wrote: > On 11 Apr 2016 22:46:33 GMT, 7 wrote: > >> DFS wrote: >> >>> Take a list of data: >>> >>> ID Date >>> 1 2004-05-04 >>> 2 2004-05-04 >>> 3 2004-05-05 >>> 4 2004-05-06 >>> 5 2004-05-08 >>> ... >>> 500000 2014-05-03 >>> >>> I want to find the best approach to determine the unknown range of ID >>> numbers that correspond to a known range of dates. >> >> What a fscking pussy! >> >> Here pussy want to learn some real computing do you? ;) >> >> Sorry to hear microshat and appil did not educt you in matters >> of real computing. >> >> Various thoughts of not allowing trolls like you near computers >> had crossed my mind reading your long failed thesis on >> how micorhsafteess and appile retards solve problems with >> their illiterate ways. >> >> The solution is one paragraph or less with Linux. > > So why don't you show us? The same reason the lying fraud has never once posted a line of code from various fantastic programs he claims to have written: none of it exists - or it's of such low quality he's afraid for anyone else to see it. The crazy idiot said he "analyzed all the world's financial transactions in half-an-hour on a midrange PC". He claimed he processed 500,000,000 data points in 30 minutes. When you ask the wackjob what kind of processing it does, he runs away.
[toc] | [prev] | [next] | [standalone]
| From | Omar <omarsayeed@linuxmail.org> |
|---|---|
| Date | 2016-04-12 13:46 -0400 |
| Message-ID | <nejc36$64n$1@dont-email.me> |
| In reply to | #349540 |
On Tue, 12 Apr 2016 13:28:56 -0400, DFS wrote: > On 4/11/2016 6:52 PM, Omar wrote: >> On 11 Apr 2016 22:46:33 GMT, 7 wrote: >> >>> DFS wrote: >>> >>>> Take a list of data: >>>> >>>> ID Date >>>> 1 2004-05-04 >>>> 2 2004-05-04 >>>> 3 2004-05-05 >>>> 4 2004-05-06 >>>> 5 2004-05-08 >>>> ... >>>> 500000 2014-05-03 >>>> >>>> I want to find the best approach to determine the unknown range of ID >>>> numbers that correspond to a known range of dates. >>> >>> What a fscking pussy! >>> >>> Here pussy want to learn some real computing do you? ;) >>> >>> Sorry to hear microshat and appil did not educt you in matters >>> of real computing. >>> >>> Various thoughts of not allowing trolls like you near computers >>> had crossed my mind reading your long failed thesis on >>> how micorhsafteess and appile retards solve problems with >>> their illiterate ways. >>> >>> The solution is one paragraph or less with Linux. >> >> So why don't you show us? > > > The same reason the lying fraud has never once posted a line of code > from various fantastic programs he claims to have written: none of it > exists - or it's of such low quality he's afraid for anyone else to see it. > > The crazy idiot said he "analyzed all the world's financial transactions > in half-an-hour on a midrange PC". He claimed he processed 500,000,000 > data points in 30 minutes. When you ask the wackjob what kind of > processing it does, he runs away. As each day passes I become increasingly convinced that there is no such thing as an honest Linux advocate. They all seem to be pathological liars. And when they aren't lying they are supporting the lies of another member of the cult of Linux.
[toc] | [prev] | [next] | [standalone]
| From | DFS <nospam@dfs.com> |
|---|---|
| Date | 2016-04-11 22:13 -0400 |
| Message-ID | <nehldo$eju$1@dont-email.me> |
| In reply to | #349414 |
On 4/11/2016 6:46 PM, 7 wrote: > DFS wrote: > >> Take a list of data: >> >> ID Date >> 1 2004-05-04 >> 2 2004-05-04 >> 3 2004-05-05 >> 4 2004-05-06 >> 5 2004-05-08 >> ... >> 500000 2014-05-03 >> >> I want to find the best approach to determine the unknown range of ID >> numbers that correspond to a known range of dates. > > What a fscking pussy! > > Here pussy want to learn some real computing do you? ;) > > Sorry to hear microshat and appil did not educt you in matters > of real computing. > > Various thoughts of not allowing trolls like you near computers > had crossed my mind reading your long failed thesis on > how micorhsafteess and appile retards solve problems with > their illiterate ways. > > The solution is one paragraph or less with Linux. [insert moronic babbling paragraph here]
[toc] | [prev] | [next] | [standalone]
| From | Fabian Russell <fb@zen.info> |
|---|---|
| Date | 2016-04-11 23:22 +0000 |
| Message-ID | <nehbir0pis@news6.newsguy.com> |
| In reply to | #349365 |
On Mon, 11 Apr 2016 11:23:51 -0400, DFS wrote: > Take a list of data: > What the fuck does this have to do with Linux advocacy? The answer is that it has NOTHING to do with Linux advocacy. So why is it here? The answer is because you are a dumb-fuck asshole. But does is this search repetitive? I would assume that it is. Therefore do a preliminary sampling do obtain a series of intervals with upper/lower bounds. Then each subsequent search can be done using a binary tree on a vastly reduced range. GNU/Linux has superior tools for this purpose, as it does for everything else.
[toc] | [prev] | [next] | [standalone]
| From | vallor <vallor@cultnix.org> |
|---|---|
| Date | 2016-04-12 00:06 +0000 |
| Message-ID | <dn2sfpFhur3U1@mid.individual.net> |
| In reply to | #349422 |
On Mon, 11 Apr 2016 23:22:03 +0000, Fabian Russell wrote: > On Mon, 11 Apr 2016 11:23:51 -0400, DFS wrote: > >> Take a list of data: >> >> > What the fuck does this have to do with Linux advocacy? He probably wants it for his standard stalking-of-linux-advocate activities. -- -v Kernel:4.6.0-rc2-sd Desktop:Xfce 4.12.2 Distro:Linux Mint 17.3 Rosa
[toc] | [prev] | [next] | [standalone]
| From | Omar <omarsayeed@linuxmail.org> |
|---|---|
| Date | 2016-04-11 20:26 -0400 |
| Message-ID | <nehf59$uh$1@dont-email.me> |
| In reply to | #349429 |
On 12 Apr 2016 00:06:17 GMT, vallor wrote: > On Mon, 11 Apr 2016 23:22:03 +0000, Fabian Russell wrote: > >> On Mon, 11 Apr 2016 11:23:51 -0400, DFS wrote: >> >>> Take a list of data: >>> >>> >> What the fuck does this have to do with Linux advocacy? > > He probably wants it for his standard stalking-of-linux-advocate > activities. Is that you William Poaster?
[toc] | [prev] | [next] | [standalone]
| From | Chris Ahlstrom <OFeem1987@teleworm.us> |
|---|---|
| Date | 2016-04-12 05:30 -0400 |
| Message-ID | <neif8a$n0p$1@dont-email.me> |
| In reply to | #349429 |
vallor wrote this copyrighted missive and expects royalties: > On Mon, 11 Apr 2016 23:22:03 +0000, Fabian Russell wrote: > >> On Mon, 11 Apr 2016 11:23:51 -0400, DFS wrote: >> >>> Take a list of data: >>> >>> >> What the fuck does this have to do with Linux advocacy? > > He probably wants it for his standard stalking-of-linux-advocate > activities. Whahhh? : User-Agent: Pan/0.141 (Tarzan's Death; GIT fb7f2ee ...) Anyway, Doofus long ago claimed that COLA was his own private toilet. -- And I do lord over the lying Linux peons in cola. -- DFS
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | comp.os.linux.advocacy
csiph-web