Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > microsoft.public.excel.programming > #109214 > unrolled thread
| Started by | Robert Baer <robertbaer@localnet.com> |
|---|---|
| First post | 2016-08-26 10:57 -0800 |
| Last post | 2016-09-17 17:15 -0400 |
| Articles | 20 on this page of 92 — 6 participants |
Back to article view | Back to microsoft.public.excel.programming
Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-26 10:57 -0800
Re: Read (and parse) file on the web "Auric__" <not.my.real@email.address> - 2016-08-26 15:21 +0000
Re: Read (and parse) file on the web "Auric__" <not.my.real@email.address> - 2016-08-26 15:22 +0000
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-27 13:04 -0800
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-27 13:00 -0800
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-27 23:14 -0800
Re: Read (and parse) file on the web CORRECTION Robert Baer <robertbaer@localnet.com> - 2016-08-28 09:42 -0800
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-28 10:08 -0800
Re: Read (and parse) file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-08-29 01:51 +0000
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-28 21:32 -0800
Re: Read (and parse) file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-08-29 19:41 +0000
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-30 00:21 -0800
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-08-30 06:06 -0400
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-30 14:20 -0800
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-08-30 18:43 -0400
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-08-31 16:47 -0400
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-01 04:37 -0800
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-01 13:13 -0400
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-01 16:32 -0800
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-01 19:43 -0400
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-01 04:23 -0800
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-05 10:10 -0400
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-05 10:17 -0400
Re: Read (and parse) file on the web CORRECTION#3 Robert Baer <robertbaer@localnet.com> - 2016-09-05 20:50 -0800
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-05 20:32 -0800
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-06 10:49 -0400
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 12:05 -0800
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-06 15:15 -0400
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 22:32 -0800
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-07 11:44 -0400
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-07 09:33 -0800
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-08 20:18 -0800
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-09 06:07 -0400
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-09 22:05 -0400
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-14 17:16 -0800
Re: Read (and parse) file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-14 20:23 -0400
Re: Read (and parse) file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-09-07 02:37 +0000
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 22:43 -0800
Re: Read (and parse) file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-09-08 05:49 +0000
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-08 01:07 -0800
Re: Read (and parse) file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-09-09 06:40 +0000
Re: Read (and parse) file on the web WGET versions Robert Baer <robertbaer@localnet.com> - 2016-09-08 01:19 -0800
Re: Read (and parse) file on the web WGET versions Robert Baer <robertbaer@localnet.com> - 2016-09-08 20:09 -0800
Re: Read (and parse) file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 23:49 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-08-28 15:51 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-28 21:37 -0800
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-02 19:09 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-08-29 01:56 -0400
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-08-29 02:21 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 19:55 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-08-29 23:35 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 22:25 -0800
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 19:21 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-08-29 23:21 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 22:10 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-08-30 04:04 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-30 13:48 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-08-29 23:37 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 22:25 -0800
Re: Read (and parse) file on the web lovexes17816 <lovexes17816@gmail.com> - 2016-08-28 14:00 +0100
Re: Read (and parse) file on the web lovexes17816 <lovexes17816@gmail.com> - 2016-08-29 02:17 +0100
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-02 19:11 -0800
Re: Read (and parse) file on the web "Auric__" <not.my.real@email.address> - 2016-09-03 04:12 +0000
Re: Read (and parse) file on the web Tikivn23143 <Tikivn23143@gmail.com> - 2016-08-30 02:28 +0100
Re: Read (and parse) file on the web everonvietnam2016 <everonvietnam2016@gmail.com> - 2016-09-07 15:03 +0100
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-07 08:43 -0800
Re: Read (and parse) file on the web "Auric__" <not.my.real@email.address> - 2016-09-08 05:43 +0000
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-08 01:09 -0800
Re: Read (and parse) file on the web "Auric__" <not.my.real@email.address> - 2016-09-09 06:49 +0000
Re: Read (and parse) file on the web everonvietnam2016 <everonvietnam2016@gmail.com> - 2016-09-08 11:11 +0100
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-09-14 17:14 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-14 17:19 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-09-14 20:26 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 09:59 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-09-15 14:33 -0400
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-09-14 21:28 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 10:05 -0800
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-14 18:48 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-09-14 21:59 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 10:15 -0800
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 22:10 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-09-16 11:50 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-16 23:54 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-09-17 11:06 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-17 23:59 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-09-18 03:22 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-20 00:20 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-09-17 11:14 -0400
Re: Read (and parse) file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-18 00:18 -0800
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-09-18 03:30 -0400
Re: Read (and parse) file on the web GS <gs@v.invalid> - 2016-09-15 05:02 -0400
Re: Read (and parse) file on the web - revised GS <gs@v.invalid> - 2016-09-17 17:15 -0400
Page 2 of 5 — ← Prev page 1 [2] 3 4 5 Next page →
| From | Robert Baer <robertbaer@localnet.com> |
|---|---|
| Date | 2016-09-01 04:23 -0800 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <ddUxz.6599$091.3871@fx29.iad> |
| In reply to | #109253 |
GS wrote:
>> You are getting all of the right stuff from what i would call the
>> second file.
>> The first file is
>> "https://www.nsncenter.com/NSNSearch?q=5960%20regulator&PageNumber=" &
>> PageNum where PagNum (in ASCII) goes from "1" to "999".
>> Note the (implied?) space in the URL.
>
> I got Source, NSN Part#, Description from the 1st file. The NSN Item#
> links to the 2nd file.
>>
> <snip>
>> Would you be so kind as to share your working ADODB code?
>> Or did you hand-copy the source like i did?
>
> I use std VB file I/O not ADODB. Initial procedure was to copy/paste
> page source into Textpad and save as Tmp.txt. Then load the file into an
> array and parse from there.
>
> I thought I'd take a look at going with a userform and MS Web Browser
> control for more flexible programming opts, but haven't had the time. I
> assume this would definitely give you an advantage over trying to
> automate IE, but I need to research using it. I do have URL functions
> built into my fpSpread.ocx for doing this stuff, but that's an expensive
> 3rd party AX component. Otherwise, doing this from Excel isn't something
> I'm familiar with.
>
Check. I know QBASIC fairly well, so a lot of that knowledge crosses
over to VB.
Someone here was kind enough to give me a full working program that
can be used to copy a URL source to a temp file on the HD.
Once available all else is very simple and straight forward.
The rub is that function (or something it uses) does not allow a
space in the URL,AND also does not allow https.
So, i need two work-arounds, and the https part would seem to be the
worst.
I do not know how it works, what DLLs/libraries it calls; no useful
information is available.
It is:
Declare Function URLDownloadToFile Lib "urlmon" _
Alias "URLDownloadToFileA" (ByVal pCaller As Long, _
ByVal szURL As String, ByVal szFileName As String, _
ByVal dwReserved As Long, ByVal lpfnCB As Long) As Long
Only the well-known keywords can be found; 'urlmon', 'pCaller',
'szURL', and 'szFileName' are unknowns and not findable in the so-called
VB help.
And there are no examples; the few ranDUMB ones are incomplete and/or
do not work..
I do not see how you use std VB file I/O; AFAIK one cannot open a web
page as if it was a file.
[toc] | [prev] | [next] | [standalone]
| From | GS <gs@v.invalid> |
|---|---|
| Date | 2016-09-05 10:10 -0400 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <nqjuc4$7j8$1@dont-email.me> |
| In reply to | #109247 |
> I wish to read and parse every page of
> "https://www.nsncenter.com/NSNSearch?q=5960%20regulator&PageNumber="
> where the page number goes from 5 to 999.
> On each page, find "<a href="/NSN/5960 [it is longer, but that is
> the start].
> Given the full number (eg: <a href="/NSN/5960-00-831-8683"), open
> a new related page "https://www.nsncenter.com/NSN/5960-00-831-8683"
> and find the line ending "(MCRL)".
> Read abut 4 lines to <a href="/PartNumber/ which is <a
> href="/PartNumber/GV4S1400"> in this case.
> save/write that line plus the next three; close this secondary online
> URL and step to next "<a href="/NSN/5960 to process the same way.
> Continue to end of the page, close that URL and open the next
> page.
Robert,
Here's what I have after parsing 'parent' pages for a list of its
links:
N/A
5960-00-503-9529
5960-00-504-8401
5960-01-035-3901
5960-01-029-2766
5960-00-617-4105
5960-00-729-5602
5960-00-826-1280
5960-00-754-5316
5960-00-962-5391
5960-00-944-4671
This is pg1 where the 1st link doesn't contain "5960" and so will be
ignored.
Each link's text is appended to this URL to bring up its 'child' pg:
https://www.nsncenter.com/NSN/
Each child page is parsed for the following 4 lines:
<TD style="VERTICAL-ALIGN: middle" align=center><A
href="/PartNumber/GV3S2800">GV3S2800</A></TD>
<TD style="HEIGHT: 60px; WIDTH: 125px; VERTICAL-ALIGN: middle" noWrap
align=center> <A
href="/CAGE/63060">63060</A> </TD>
<TD style="VERTICAL-ALIGN: middle" align=center> <A
href="/CAGE/63060"><IMG class=img-thumbnail
src="https://placehold.it/90x45?text=No%0DImage%0DYet" width=90
height=45></A> </TD>
<TD style="VERTICAL-ALIGN: middle" text-align="center"><A title="CAGE
63060" href="/CAGE/63060">HEICO OHMITE LLC</A></TD>
I'm stripping html syntax to get this data:
Line1: PartNumber/GV3S2800
Line2: CAGE/63060
Line3: https://placehold.it/90x45?text=No%0DImage%0DYet
Line4: HEICO OHMITE LLC
The output file has these filenames in the 1st line:
NSN Item#,Description,Part#,MCRL,CAGE,Source
I left the 3rd line URL out since, outside its host webpage, it'll be
useless to you. I need to know from you if the 3rd line URL is needed!
Otherwise, the output file will have 1 line per item so it can be used
as the db file "NSN_5960_ElectronTube.dat". I invite your suggestion
for filename...
I could extend the collected data to include...
Reference Number/DRN_3570
Entity Code/DRN_9250
Category Code/DRN_2910
Variation Code/DRN_4780
..where the fieldnames would then be:
Item#,Part#,MCRL,CAGE,Source,Ref,Entity,Category,Variation
The 1st record will be:
5960-00-503-9529,GV3S2800,3302008,63060,HEICO OHMITE
LLC,DRN_3570,DRN_9250,DRN_2910,DRN_4780
Output file size for 1 parent pg is 1Kb; for 10 parent pgs is 10Kb. You
could have 1000 parent pgs of data stored in a 1Mb file.
Your feedback is appreciated...
--
Garry
Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
comp.lang.basic.visual.misc
microsoft.public.vb.general.discussion
---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus
[toc] | [prev] | [next] | [standalone]
| From | GS <gs@v.invalid> |
|---|---|
| Date | 2016-09-05 10:17 -0400 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <nqjuph$93t$1@dont-email.me> |
| In reply to | #109279 |
Typos... > > Robert, > Here's what I have after parsing 'parent' pages for a list of its > links: > > N/A > 5960-00-503-9529 > 5960-00-504-8401 > 5960-01-035-3901 > 5960-01-029-2766 > 5960-00-617-4105 > 5960-00-729-5602 > 5960-00-826-1280 > 5960-00-754-5316 > 5960-00-962-5391 > 5960-00-944-4671 > > This is pg1 where the 1st link doesn't contain "5960" and so will be > ignored. > > Each link's text is appended to this URL to bring up its 'child' pg: > > https://www.nsncenter.com/NSN/ > > Each child page is parsed for the following 4 lines: > > <TD style="VERTICAL-ALIGN: middle" align=center><A > href="/PartNumber/GV3S2800">GV3S2800</A></TD> > <TD style="HEIGHT: 60px; WIDTH: 125px; VERTICAL-ALIGN: middle" > noWrap align=center> <A > href="/CAGE/63060">63060</A> </TD> > <TD style="VERTICAL-ALIGN: middle" align=center> <A > href="/CAGE/63060"><IMG class=img-thumbnail > src="https://placehold.it/90x45?text=No%0DImage%0DYet" width=90 > height=45></A> </TD> > <TD style="VERTICAL-ALIGN: middle" text-align="center"><A > title="CAGE 63060" href="/CAGE/63060">HEICO OHMITE LLC</A></TD> > > I'm stripping html syntax to get this data: > > Line1: PartNumber/GV3S2800 > Line2: CAGE/63060 > Line3: https://placehold.it/90x45?text=No%0DImage%0DYet > Line4: HEICO OHMITE LLC > > The output file has these fieldnames in the 1st line: > > NSN Item#,Description,Part#,MCRL,CAGE,Source > > I left the 3rd line URL out since, outside its host webpage, it'll be > useless to you. I need to know from you if the 3rd line URL is > needed! > > Otherwise, the output file will have 1 line per item so it can be > used as the db file "NSN_5960_ElectronTube.dat". I invite your > suggestion for filename... > > I could extend the collected data to include... > > Reference Number/DRN_3570 > Entity Code/DRN_9250 > Category Code/DRN_2910 > Variation Code/DRN_4780 > > ..where the fieldnames would then be: > Item#,Part#,MCRL,CAGE,Source,REF,ENT,CAT,VAR > > The 1st record will be: > > 5960-00-503-9529,GV3S2800,3302008,63060,HEICO OHMITE > LLC,DRN_3570,DRN_9250,DRN_2910,DRN_4780 > > Output file size for 1 parent pg is 1Kb; for 10 parent pgs is 10Kb. > You could have 1000 parent pgs of data stored in a 1Mb file. > > Your feedback is appreciated... -- Garry Free usenet access at http://www.eternal-september.org Classic VB Users Regroup! comp.lang.basic.visual.misc microsoft.public.vb.general.discussion --- This email has been checked for viruses by Avast antivirus software. https://www.avast.com/antivirus
[toc] | [prev] | [next] | [standalone]
| From | Robert Baer <robertbaer@localnet.com> |
|---|---|
| Date | 2016-09-05 20:50 -0800 |
| Subject | Re: Read (and parse) file on the web CORRECTION#3 |
| Message-ID | <A2rzz.42429$e%3.31863@fx07.iad> |
| In reply to | #109280 |
GS wrote: > Typos... >> >> Robert, >> Here's what I have after parsing 'parent' pages for a list of its links: >> >> N/A >> 5960-00-503-9529 >> 5960-00-504-8401 >> 5960-01-035-3901 >> 5960-01-029-2766 >> 5960-00-617-4105 >> 5960-00-729-5602 >> 5960-00-826-1280 >> 5960-00-754-5316 >> 5960-00-962-5391 >> 5960-00-944-4671 >> >> This is pg1 where the 1st link doesn't contain "5960" and so will be >> ignored. >> >> Each link's text is appended to this URL to bring up its 'child' pg: >> >> https://www.nsncenter.com/NSN/ >> >> Each child page is parsed for the following 4 lines: >> >> <TD style="VERTICAL-ALIGN: middle" align=center><A >> href="/PartNumber/GV3S2800">GV3S2800</A></TD> >> <TD style="HEIGHT: 60px; WIDTH: 125px; VERTICAL-ALIGN: middle" noWrap >> align=center> <A href="/CAGE/63060">63060</A> </TD> >> <TD style="VERTICAL-ALIGN: middle" align=center> <A >> href="/CAGE/63060"><IMG class=img-thumbnail >> src="https://placehold.it/90x45?text=No%0DImage%0DYet" width=90 >> height=45></A> </TD> >> <TD style="VERTICAL-ALIGN: middle" text-align="center"><A title="CAGE >> 63060" href="/CAGE/63060">HEICO OHMITE LLC</A></TD> >> >> I'm stripping html syntax to get this data: >> >> Line1: PartNumber/GV3S2800 >> Line2: CAGE/63060 >> Line3: https://placehold.it/90x45?text=No%0DImage%0DYet >> Line4: HEICO OHMITE LLC >> >> > The output file has these fieldnames in the 1st line: >> >> NSN Item#,Description,Part#,MCRL,CAGE,Source >> >> I left the 3rd line URL out since, outside its host webpage, it'll be >> useless to you. I need to know from you if the 3rd line URL is needed! >> >> Otherwise, the output file will have 1 line per item so it can be used >> as the db file "NSN_5960_ElectronTube.dat". I invite your suggestion >> for filename... >> >> I could extend the collected data to include... >> >> Reference Number/DRN_3570 >> Entity Code/DRN_9250 >> Category Code/DRN_2910 >> Variation Code/DRN_4780 >> >> ..where the fieldnames would then be: >> > Item#,Part#,MCRL,CAGE,Source,REF,ENT,CAT,VAR >> >> The 1st record will be: >> >> 5960-00-503-9529,GV3S2800,3302008,63060,HEICO OHMITE >> LLC,DRN_3570,DRN_9250,DRN_2910,DRN_4780 >> >> Output file size for 1 parent pg is 1Kb; for 10 parent pgs is 10Kb. >> You could have 1000 parent pgs of data stored in a 1Mb file. >> >> Your feedback is appreciated... > Like i said, PERFECT! And you are correct, do not need line 3 nor the extended data. Please check my other answer for a corrected search term which needs a !corrected! human-readable version of the URL is: https://www.nsncenter.com/NSNSearch?q=5960 regulator and "ELECTRON TUBE"&PageNumber=1 The %20 is a virtual space, and the %22 is a virtual quote. ((guess that is the proper term))
[toc] | [prev] | [next] | [standalone]
| From | Robert Baer <robertbaer@localnet.com> |
|---|---|
| Date | 2016-09-05 20:32 -0800 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <JNqzz.334$OT5.113@fx19.iad> |
| In reply to | #109279 |
GS wrote:
>> I wish to read and parse every page of
>> "https://www.nsncenter.com/NSNSearch?q=5960%20regulator&PageNumber="
>> where the page number goes from 5 to 999.
>> On each page, find "<a href="/NSN/5960 [it is longer, but that is the
>> start].
>> Given the full number (eg: <a href="/NSN/5960-00-831-8683"), open a
>> new related page "https://www.nsncenter.com/NSN/5960-00-831-8683" and
>> find the line ending "(MCRL)".
>> Read abut 4 lines to <a href="/PartNumber/ which is <a
>> href="/PartNumber/GV4S1400"> in this case.
>> save/write that line plus the next three; close this secondary online
>> URL and step to next "<a href="/NSN/5960 to process the same way.
>> Continue to end of the page, close that URL and open the next page.
>
> Robert,
> Here's what I have after parsing 'parent' pages for a list of its links:
>
> N/A
> 5960-00-503-9529
> 5960-00-504-8401
> 5960-01-035-3901
> 5960-01-029-2766
> 5960-00-617-4105
> 5960-00-729-5602
> 5960-00-826-1280
> 5960-00-754-5316
> 5960-00-962-5391
> 5960-00-944-4671
>
> This is pg1 where the 1st link doesn't contain "5960" and so will be
> ignored.
>
> Each link's text is appended to this URL to bring up its 'child' pg:
>
> https://www.nsncenter.com/NSN/
>
> Each child page is parsed for the following 4 lines:
>
> <TD style="VERTICAL-ALIGN: middle" align=center><A
> href="/PartNumber/GV3S2800">GV3S2800</A></TD>
> <TD style="HEIGHT: 60px; WIDTH: 125px; VERTICAL-ALIGN: middle" noWrap
> align=center> <A href="/CAGE/63060">63060</A> </TD>
> <TD style="VERTICAL-ALIGN: middle" align=center> <A
> href="/CAGE/63060"><IMG class=img-thumbnail
> src="https://placehold.it/90x45?text=No%0DImage%0DYet" width=90
> height=45></A> </TD>
> <TD style="VERTICAL-ALIGN: middle" text-align="center"><A title="CAGE
> 63060" href="/CAGE/63060">HEICO OHMITE LLC</A></TD>
>
> I'm stripping html syntax to get this data:
>
> Line1: PartNumber/GV3S2800
> Line2: CAGE/63060
> Line3: https://placehold.it/90x45?text=No%0DImage%0DYet
> Line4: HEICO OHMITE LLC
>
>
> The output file has these filenames in the 1st line:
>
> NSN Item#,Description,Part#,MCRL,CAGE,Source
>
> I left the 3rd line URL out since, outside its host webpage, it'll be
> useless to you. I need to know from you if the 3rd line URL is needed!
>
> Otherwise, the output file will have 1 line per item so it can be used
> as the db file "NSN_5960_ElectronTube.dat". I invite your suggestion for
> filename...
>
> I could extend the collected data to include...
>
> Reference Number/DRN_3570
> Entity Code/DRN_9250
> Category Code/DRN_2910
> Variation Code/DRN_4780
>
> ..where the fieldnames would then be:
>
> Item#,Part#,MCRL,CAGE,Source,Ref,Entity,Category,Variation
>
> The 1st record will be:
>
> 5960-00-503-9529,GV3S2800,3302008,63060,HEICO OHMITE
> LLC,DRN_3570,DRN_9250,DRN_2910,DRN_4780
>
> Output file size for 1 parent pg is 1Kb; for 10 parent pgs is 10Kb. You
> could have 1000 parent pgs of data stored in a 1Mb file.
>
> Your feedback is appreciated...
>
WOW!
Absolutely PERFECT!
You are correct, #1) do not need that line 3, and #2) do not need the
extended info.
File name(s) for PageNumber=1 I would use 5960_001.TXT,..to
PageNumber=999 I would use 5960_999.TXT and that would preserve order.
*OR*
Reading & parsing from PageNumber=1 to PageNumber=999,one could
append to same file (name NSN_5960.TXT); might as well - makes it easier
to pour into a single Excel file.
Either way is fine.
I have found a way to get rid of items that are not strictly electron
tubes and/or not regulators; that way you do not have to parse out these
"unfit" items from first page description. Use:
"https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=1"
Naturally, PageNumber still goes from 1 to 999.
Note the implied "(", ")" and " "; human-readable "5960 regulator and
(ELECTRON TUBE)".
As far as i can tell, using that shows no undesirable parts.
Thanks!
PS: i found WGET to be non-useful (a) it truncates the filename (b) it
buggers it to partial gibberish.
[toc] | [prev] | [next] | [standalone]
| From | GS <gs@v.invalid> |
|---|---|
| Date | 2016-09-06 10:49 -0400 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <nqml2k$3t7$1@dont-email.me> |
| In reply to | #109281 |
> You are correct, #1) do not need that line 3, and #2) do not need
> the extended info.
Ok then, fieldnames will be: Item#,Part#,MCRL,CAGE,Source
>
> File name(s) for PageNumber=1 I would use 5960_001.TXT,..to
> PageNumber=999 I would use 5960_999.TXT and that would preserve
> order.
> *OR*
> Reading & parsing from PageNumber=1 to PageNumber=999,one could
> append to same file (name NSN_5960.TXT); might as well - makes it
> easier to pour into a single Excel file.
> Either way is fine.
Ok, then ouput filename will be: NSN_5960.txt
>
> I have found a way to get rid of items that are not strictly
> electron tubes and/or not regulators; that way you do not have to
> parse out these "unfit" items from first page description. Use:
> "https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=1"
> Naturally, PageNumber still goes from 1 to 999.
> Note the implied "(", ")" and " "; human-readable "5960 regulator
> and (ELECTRON TUBE)".
> As far as i can tell, using that shows no undesirable parts.
Works nice! Now I get 11 5960 items per parent page.
>
> Thanks!
> PS: i found WGET to be non-useful (a) it truncates the filename (b)
> it buggers it to partial gibberish
What is WGET?
--
Garry
Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
comp.lang.basic.visual.misc
microsoft.public.vb.general.discussion
---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus
[toc] | [prev] | [next] | [standalone]
| From | Robert Baer <robertbaer@localnet.com> |
|---|---|
| Date | 2016-09-06 12:05 -0800 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <WrEzz.32969$BN1.19549@fx02.iad> |
| In reply to | #109283 |
GS wrote:
>> You are correct, #1) do not need that line 3, and #2) do not need the
>> extended info.
>
> Ok then, fieldnames will be: Item#,Part#,MCRL,CAGE,Source
>>
>> File name(s) for PageNumber=1 I would use 5960_001.TXT,..to
>> PageNumber=999 I would use 5960_999.TXT and that would preserve order.
>> *OR*
>> Reading & parsing from PageNumber=1 to PageNumber=999,one could append
>> to same file (name NSN_5960.TXT); might as well - makes it easier to
>> pour into a single Excel file.
>> Either way is fine.
>
> Ok, then ouput filename will be: NSN_5960.txt
>>
>> I have found a way to get rid of items that are not strictly electron
>> tubes and/or not regulators; that way you do not have to parse out
>> these "unfit" items from first page description. Use:
>> "https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=1"
>>
>> Naturally, PageNumber still goes from 1 to 999.
>> Note the implied "(", ")" and " "; human-readable "5960 regulator and
>> (ELECTRON TUBE)".
>> As far as i can tell, using that shows no undesirable parts.
>
> Works nice! Now I get 11 5960 items per parent page.
>>
>> Thanks!
>> PS: i found WGET to be non-useful (a) it truncates the filename (b) it
>> buggers it to partial gibberish
>
> What is WGET?
>
WGET is a command line program that will copy contents of an URL to
the hard drive; it has various options, for SSL, i think for some
processing, for giving the output file a specific name, for recursion, etc.
Was still trying to find ways to copy the online file to the hard drive.
I still do not understand what magic you used.
Now, the nitty-gritty; in exchange for that nicely parsed file, what
do i owe you?
[toc] | [prev] | [next] | [standalone]
| From | GS <gs@v.invalid> |
|---|---|
| Date | 2016-09-06 15:15 -0400 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <nqn4k6$32u$1@dont-email.me> |
| In reply to | #109284 |
> I still do not understand what magic you used. I'm using the MS WebBrowser control and a textbox on a worksheet! > > Now, the nitty-gritty; in exchange for that nicely parsed file, > what do i owe you? A Timmies, straight up! -- Garry Free usenet access at http://www.eternal-september.org Classic VB Users Regroup! comp.lang.basic.visual.misc microsoft.public.vb.general.discussion --- This email has been checked for viruses by Avast antivirus software. https://www.avast.com/antivirus
[toc] | [prev] | [next] | [standalone]
| From | Robert Baer <robertbaer@localnet.com> |
|---|---|
| Date | 2016-09-06 22:32 -0800 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <vDNzz.60413$oU2.8638@fx09.iad> |
| In reply to | #109285 |
GS wrote: >> I still do not understand what magic you used. > > I'm using the MS WebBrowser control and a textbox on a worksheet! >> >> Now, the nitty-gritty; in exchange for that nicely parsed file, what >> do i owe you? > > A Timmies, straight up! > The search engine was not exactly forthcoming, to say the least; everything including the kitchen sink but NOT anything alcoholic. "Timmies drink" helped some; fifth "hit" down: "Timmy's Sweet and Sour mix Cocktails and Drink Recipes". Using "Timmies, straight up" was slightly better.."Average night at the Manotick Timmies... : ottawa" In all of this,a lot of "hits" mentioned something(always different) about Tim Hortons Franchise. Absolutely no clue regarding rum, scotch, vodka (or dare i say) milk.
[toc] | [prev] | [next] | [standalone]
| From | GS <gs@v.invalid> |
|---|---|
| Date | 2016-09-07 11:44 -0400 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <nqpclr$t1o$1@dont-email.me> |
| In reply to | #109287 |
> GS wrote: >>> I still do not understand what magic you used. >> >> I'm using the MS WebBrowser control and a textbox on a worksheet! >>> >>> Now, the nitty-gritty; in exchange for that nicely parsed file, >>> what >>> do i owe you? >> >> A Timmies, straight up! >> > The search engine was not exactly forthcoming, to say the least; > everything including the kitchen sink but NOT anything alcoholic. > "Timmies drink" helped some; fifth "hit" down: "Timmy's Sweet and > Sour mix Cocktails and Drink Recipes". > Using "Timmies, straight up" was slightly better.."Average night > at the Manotick Timmies... : ottawa" > > In all of this,a lot of "hits" mentioned something(always > different) about Tim Hortons Franchise. > > Absolutely no clue regarding rum, scotch, vodka (or dare i say) > milk. Ha-ha! Ok.., 'Timmies' is fan-speak for Tim Horton's coffee!<g> -- Garry Free usenet access at http://www.eternal-september.org Classic VB Users Regroup! comp.lang.basic.visual.misc microsoft.public.vb.general.discussion --- This email has been checked for viruses by Avast antivirus software. https://www.avast.com/antivirus
[toc] | [prev] | [next] | [standalone]
| From | Robert Baer <robertbaer@localnet.com> |
|---|---|
| Date | 2016-09-07 09:33 -0800 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <pjXzz.57656$SX5.14315@fx01.iad> |
| In reply to | #109292 |
GS wrote: >> GS wrote: >>>> I still do not understand what magic you used. >>> >>> I'm using the MS WebBrowser control and a textbox on a worksheet! >>>> >>>> Now, the nitty-gritty; in exchange for that nicely parsed file, what >>>> do i owe you? >>> >>> A Timmies, straight up! >>> >> The search engine was not exactly forthcoming, to say the least; >> everything including the kitchen sink but NOT anything alcoholic. >> "Timmies drink" helped some; fifth "hit" down: "Timmy's Sweet and Sour >> mix Cocktails and Drink Recipes". >> Using "Timmies, straight up" was slightly better.."Average night at >> the Manotick Timmies... : ottawa" >> >> In all of this,a lot of "hits" mentioned something(always different) >> about Tim Hortons Franchise. >> >> Absolutely no clue regarding rum, scotch, vodka (or dare i say) milk. > > Ha-ha! Ok.., 'Timmies' is fan-speak for Tim Horton's coffee!<g> > What i did WRT Wget, was uninstall it and checked that there were 'dregs' on the HD. Then i installed it from scratch, allowing all of the defaults. Finally, i modified the system path (shortened version): %SystemRoot%\system32;%SystemRoot%;%SystemRoot%\System32\Wbem;C:\Program Files\GnuWin32; No joy. Even at the root, the system insists wget does not exist (as an executable, etc).
[toc] | [prev] | [next] | [standalone]
| From | Robert Baer <robertbaer@localnet.com> |
|---|---|
| Date | 2016-09-08 20:18 -0800 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <gSpAz.36246$Q97.23130@fx06.iad> |
| In reply to | #109292 |
GS wrote: >> GS wrote: >>>> I still do not understand what magic you used. >>> >>> I'm using the MS WebBrowser control and a textbox on a worksheet! >>>> >>>> Now, the nitty-gritty; in exchange for that nicely parsed file, what >>>> do i owe you? >>> >>> A Timmies, straight up! >>> >> The search engine was not exactly forthcoming, to say the least; >> everything including the kitchen sink but NOT anything alcoholic. >> "Timmies drink" helped some; fifth "hit" down: "Timmy's Sweet and Sour >> mix Cocktails and Drink Recipes". >> Using "Timmies, straight up" was slightly better.."Average night at >> the Manotick Timmies... : ottawa" >> >> In all of this,a lot of "hits" mentioned something(always different) >> about Tim Hortons Franchise. >> >> Absolutely no clue regarding rum, scotch, vodka (or dare i say) milk. > > Ha-ha! Ok.., 'Timmies' is fan-speak for Tim Horton's coffee!<g> > Good show. You have an e-mail account,i presume? I could use PayPal to send some nickels for that parsed NSN file if you wish. Please let me know how soon i can get that file. Thanks.
[toc] | [prev] | [next] | [standalone]
| From | GS <gs@v.invalid> |
|---|---|
| Date | 2016-09-09 06:07 -0400 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <nqu1kb$9bu$1@dont-email.me> |
| In reply to | #109301 |
>> Ha-ha! Ok.., 'Timmies' is fan-speak for Tim Horton's coffee!<g> >> > Good show. > You have an e-mail account,i presume? > I could use PayPal to send some nickels for that parsed NSN file > if you wish. > Please let me know how soon i can get that file. > Thanks. No worries.., I'm just kidding! That's what I tell neighbors when they offer payment for helping them with computer issues! As I stated earlier, I'm commited right now and so only have time to work on this when I get a chance. Currently, it's ready to fully automate, but seems to have a snag writing past the 1st parent page's child pages. I'm likely going to have time to finish it this weekend, though, but I'll post a download link when done, ..regardless! -- Garry Free usenet access at http://www.eternal-september.org Classic VB Users Regroup! comp.lang.basic.visual.misc microsoft.public.vb.general.discussion --- This email has been checked for viruses by Avast antivirus software. https://www.avast.com/antivirus
[toc] | [prev] | [next] | [standalone]
| From | GS <gs@v.invalid> |
|---|---|
| Date | 2016-09-09 22:05 -0400 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <nqvpph$4pp$1@dont-email.me> |
| In reply to | #109304 |
> Currently, it's ready to fully automate, but seems to have a snag > writing past the 1st parent page's child pages. Turns out the problem was code not waiting for the browser not busy. I switched to using URLDownloadToFile() at this point because it's orders of magnitude faster. Using the browser/twxtbox on a sheet served well for getting the process code nailed down, but that was only a temp situation during dev. The links error out about mid pg7 thru pg10 as I time tested only pgs 1thru10: -this took 50.89 secs! I'll do some housekeeping of the code and post a download link to the file... -- Garry Free usenet access at http://www.eternal-september.org Classic VB Users Regroup! comp.lang.basic.visual.misc microsoft.public.vb.general.discussion --- This email has been checked for viruses by Avast antivirus software. https://www.avast.com/antivirus
[toc] | [prev] | [next] | [standalone]
| From | Robert Baer <robertbaer@localnet.com> |
|---|---|
| Date | 2016-09-14 17:16 -0800 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <pLlCz.24829$fL2.8044@fx20.iad> |
| In reply to | #109305 |
GS wrote: >> Currently, it's ready to fully automate, but seems to have a snag >> writing past the 1st parent page's child pages. > > Turns out the problem was code not waiting for the browser not busy. I > switched to using URLDownloadToFile() at this point because it's orders > of magnitude faster. Using the browser/twxtbox on a sheet served well > for getting the process code nailed down, but that was only a temp > situation during dev. "Problem" with Excel, is that there are MANY ways to get what is needed, and there is NO WAY of discovering _any_ of them; the "help" document is worse than useless in that manner. I have found that URLDownloadToFile() to be non-functional for https sources. > > The links error out about mid pg7 thru pg10 as I time tested only pgs > 1thru10: -this took 50.89 secs! > > I'll do some housekeeping of the code and post a download link to the > file... >
[toc] | [prev] | [next] | [standalone]
| From | GS <gs@v.invalid> |
|---|---|
| Date | 2016-09-14 20:23 -0400 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <nrcpm5$eu0$1@dont-email.me> |
| In reply to | #109308 |
> "Problem" with Excel, is that there are MANY ways to get what is > needed, and there is NO WAY of discovering _any_ of them; the "help" > document is worse than useless in that manner. > I have found that URLDownloadToFile() to be non-functional for > https > sources. I disagree because it's working in my project I posted the download for! -- Garry Free usenet access at http://www.eternal-september.org Classic VB Users Regroup! comp.lang.basic.visual.misc microsoft.public.vb.general.discussion --- This email has been checked for viruses by Avast antivirus software. https://www.avast.com/antivirus
[toc] | [prev] | [next] | [standalone]
| From | "Auric__" <not.my.real@email.address> |
|---|---|
| Date | 2016-09-07 02:37 +0000 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <XnsA67BC7911E470auricauricauricauric@213.239.209.88> |
| In reply to | #109281 |
Robert Baer wrote: > PS: i found WGET to be non-useful (a) it truncates the filename (b) it > buggers it to partial gibberish. Then you must be using a bad version, or perhaps have something wrong with your .wgetrc. I've been using wget for around 10 years, and never had anything like those issues unless I pass bad options. -- My life is richer, somehow, simply because I know that he exists.
[toc] | [prev] | [next] | [standalone]
| From | Robert Baer <robertbaer@localnet.com> |
|---|---|
| Date | 2016-09-06 22:43 -0800 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <8ONzz.56531$M27.26468@fx33.iad> |
| In reply to | #109286 |
Auric__ wrote: > Robert Baer wrote: > >> PS: i found WGET to be non-useful (a) it truncates the filename (b) it >> buggers it to partial gibberish. > > Then you must be using a bad version, or perhaps have something wrong with > your .wgetrc. I've been using wget for around 10 years, and never had > anything like those issues unless I pass bad options. > Know nothing about .wgetrc; am in Win2K cmd line, and the batch file used is: H: CD\Win2K_WORK\OIL4LESS\LLCDOCS\FED app\FBA stuff wget --no-check-certificate --output-document=5960_002.TXT --output-file=log002.TXT https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=2 PAUSE The SourceForge site offered a Zip which was supposed to be complete, but none of the created folders had an EXE (tried Win2K, WinXP, Win7). Found SofTonic offering only a plain jane wget.exe, which i am using, so that may be a buggered version. Suggestions?
[toc] | [prev] | [next] | [standalone]
| From | "Auric__" <not.my.real@email.address> |
|---|---|
| Date | 2016-09-08 05:49 +0000 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <XnsA67CE831F3B1Cauricauricauricauric@213.239.209.88> |
| In reply to | #109288 |
Robert Baer wrote:
> Auric__ wrote:
>> Robert Baer wrote:
>>
>>> PS: i found WGET to be non-useful (a) it truncates the filename (b) it
>>> buggers it to partial gibberish.
>>
>> Then you must be using a bad version, or perhaps have something wrong
>> with your .wgetrc. I've been using wget for around 10 years, and never
>> had anything like those issues unless I pass bad options.
>>
> Know nothing about .wgetrc;
Don't worry about it. It can be used to set default behaviors but every entry
can be replicated via switches.
> am in Win2K cmd line, and the batch file
> used is:
> H:
> CD\Win2K_WORK\OIL4LESS\LLCDOCS\FED app\FBA stuff
>
> wget --no-check-certificate --output-document=5960_002.TXT
> --output-file=log002.TXT
> https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRO
> N%20TUBE%22&PageNumber=2
That wget line performs as expected for me: 5960_002.TXT contains valid HTML
(although I haven't made any attempt to check the data; it looks like most of
the page is CSS) and log002.TXT is a typical wget log of a successful
transfer.
As for truncating the filenames, if I remove the --output-document switch,
the filename I get is
NSNSearch@q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=2
> PAUSE
>
> The SourceForge site offered a Zip which was supposed to be complete,
If you're talking about GNUwin32, that version is years out of date.
> but none of the created folders had an EXE (tried Win2K, WinXP, Win7).
> Found SofTonic offering only a plain jane wget.exe, which i am using,
> so that may be a buggered version.
Never even heard of them.
> Suggestions?
I'm using 1.16.3. No idea where I got it.
The batch file that I use for downloading looks like this:
call wget --no-check-certificate -x -c -e robots=off -i new.txt %*
-x Always create directories (e.g. http://a.b.c/1/2.txt -> .\a.b.c\1\2.txt).
-c Continue interrupted downloads.
-e Do this .wgetrc thing (in this case, ignore the robots.txt file).
-i Read list of filenames from the following file ("new.txt" because that's
the default name for a new file in my file manager).
I use the -i switch so I don't have to worry about escaping characters or %
vs %%. Whatever's in the text file is exactly what it looks for. (If you go
this route, it's one file per line.)
--
We have to stop letting George Lucas name our politicians.
[toc] | [prev] | [next] | [standalone]
| From | Robert Baer <robertbaer@localnet.com> |
|---|---|
| Date | 2016-09-08 01:07 -0800 |
| Subject | Re: Read (and parse) file on the web CORRECTION#2 |
| Message-ID | <8%8Az.103752$PM.41042@fx08.iad> |
| In reply to | #109295 |
Auric__ wrote:
> Robert Baer wrote:
>
>> Auric__ wrote:
>>> Robert Baer wrote:
>>>
>>>> PS: i found WGET to be non-useful (a) it truncates the filename (b) it
>>>> buggers it to partial gibberish.
>>>
>>> Then you must be using a bad version, or perhaps have something wrong
>>> with your .wgetrc. I've been using wget for around 10 years, and never
>>> had anything like those issues unless I pass bad options.
>>>
>> Know nothing about .wgetrc;
>
> Don't worry about it. It can be used to set default behaviors but every entry
> can be replicated via switches.
>
>> am in Win2K cmd line, and the batch file
>> used is:
>> H:
>> CD\Win2K_WORK\OIL4LESS\LLCDOCS\FED app\FBA stuff
>>
>> wget --no-check-certificate --output-document=5960_002.TXT
>> --output-file=log002.TXT
>> https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRO
>> N%20TUBE%22&PageNumber=2
>
> That wget line performs as expected for me: 5960_002.TXT contains valid HTML
> (although I haven't made any attempt to check the data; it looks like most of
> the page is CSS) and log002.TXT is a typical wget log of a successful
> transfer.
>
> As for truncating the filenames, if I remove the --output-document switch,
> the filename I get is
>
> NSNSearch@q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=2
>
>> PAUSE
>>
>> The SourceForge site offered a Zip which was supposed to be complete,
>
> If you're talking about GNUwin32, that version is years out of date.
>
>> but none of the created folders had an EXE (tried Win2K, WinXP, Win7).
>> Found SofTonic offering only a plain jane wget.exe, which i am using,
>> so that may be a buggered version.
>
> Never even heard of them.
>
>> Suggestions?
>
> I'm using 1.16.3. No idea where I got it.
>
> The batch file that I use for downloading looks like this:
>
> call wget --no-check-certificate -x -c -e robots=off -i new.txt %*
>
> -x Always create directories (e.g. http://a.b.c/1/2.txt -> .\a.b.c\1\2.txt).
> -c Continue interrupted downloads.
> -e Do this .wgetrc thing (in this case, ignore the robots.txt file).
> -i Read list of filenames from the following file ("new.txt" because that's
> the default name for a new file in my file manager).
>
> I use the -i switch so I don't have to worry about escaping characters or %
> vs %%. Whatever's in the text file is exactly what it looks for. (If you go
> this route, it's one file per line.)
>
You must have a different version of Wget; whatever i do on the
command line,including the "trick" of restrict-file-names=nocontrol, i
get a buggered path name plus the response &PageNumber not recognized.
Exactly same results in Win2K, WinXP or in Win7.
Yes, i used GNUwin32 as SourceForge "complete" of Wget had no EXE.
Is there some other (compiled, complete) source i should get?
[toc] | [prev] | [next] | [standalone]
Page 2 of 5 — ← Prev page 1 [2] 3 4 5 Next page →
Back to top | Article view | microsoft.public.excel.programming
csiph-web