Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > microsoft.public.excel.programming > #109214 > unrolled thread

Read (and parse) file on the web

Started byRobert Baer <robertbaer@localnet.com>
First post2016-08-26 10:57 -0800
Last post2016-09-17 17:15 -0400
Articles 20 on this page of 92 — 6 participants

Back to article view | Back to microsoft.public.excel.programming


Contents

  Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-26 10:57 -0800
    Re: Read (and parse)  file on the web "Auric__" <not.my.real@email.address> - 2016-08-26 15:21 +0000
      Re: Read (and parse)  file on the web "Auric__" <not.my.real@email.address> - 2016-08-26 15:22 +0000
        Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-27 13:04 -0800
      Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-27 13:00 -0800
      Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-27 23:14 -0800
        Re: Read (and parse)  file on the web CORRECTION Robert Baer <robertbaer@localnet.com> - 2016-08-28 09:42 -0800
          Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-28 10:08 -0800
            Re: Read (and parse)  file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-08-29 01:51 +0000
              Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-28 21:32 -0800
                Re: Read (and parse)  file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-08-29 19:41 +0000
                  Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-30 00:21 -0800
                    Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-08-30 06:06 -0400
                      Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-30 14:20 -0800
                        Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-08-30 18:43 -0400
                          Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-08-31 16:47 -0400
                            Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-01 04:37 -0800
                              Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-01 13:13 -0400
                                Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-01 16:32 -0800
                                  Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-01 19:43 -0400
                          Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-01 04:23 -0800
                    Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-05 10:10 -0400
                      Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-05 10:17 -0400
                        Re: Read (and parse)  file on the web CORRECTION#3 Robert Baer <robertbaer@localnet.com> - 2016-09-05 20:50 -0800
                      Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-05 20:32 -0800
                        Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-06 10:49 -0400
                          Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 12:05 -0800
                            Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-06 15:15 -0400
                              Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 22:32 -0800
                                Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-07 11:44 -0400
                                  Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-07 09:33 -0800
                                  Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-08 20:18 -0800
                                    Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-09 06:07 -0400
                                      Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-09 22:05 -0400
                                        Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-14 17:16 -0800
                                          Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-14 20:23 -0400
                        Re: Read (and parse)  file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-09-07 02:37 +0000
                          Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 22:43 -0800
                            Re: Read (and parse)  file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-09-08 05:49 +0000
                              Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-08 01:07 -0800
                                Re: Read (and parse)  file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-09-09 06:40 +0000
                              Re: Read (and parse)  file on the web WGET versions Robert Baer <robertbaer@localnet.com> - 2016-09-08 01:19 -0800
                                Re: Read (and parse)  file on the web WGET versions Robert Baer <robertbaer@localnet.com> - 2016-09-08 20:09 -0800
                          Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 23:49 -0800
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-28 15:51 -0400
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-28 21:37 -0800
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-02 19:09 -0800
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-29 01:56 -0400
          Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-29 02:21 -0400
            Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 19:55 -0800
              Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-29 23:35 -0400
                Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 22:25 -0800
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 19:21 -0800
            Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-29 23:21 -0400
              Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 22:10 -0800
                Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-30 04:04 -0400
                  Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-30 13:48 -0800
            Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-29 23:37 -0400
              Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 22:25 -0800
      Re: Read (and parse)  file on the web lovexes17816 <lovexes17816@gmail.com> - 2016-08-28 14:00 +0100
      Re: Read (and parse)  file on the web lovexes17816 <lovexes17816@gmail.com> - 2016-08-29 02:17 +0100
        Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-02 19:11 -0800
          Re: Read (and parse)  file on the web "Auric__" <not.my.real@email.address> - 2016-09-03 04:12 +0000
      Re: Read (and parse)  file on the web Tikivn23143 <Tikivn23143@gmail.com> - 2016-08-30 02:28 +0100
      Re: Read (and parse)  file on the web everonvietnam2016 <everonvietnam2016@gmail.com> - 2016-09-07 15:03 +0100
        Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-07 08:43 -0800
          Re: Read (and parse)  file on the web "Auric__" <not.my.real@email.address> - 2016-09-08 05:43 +0000
            Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-08 01:09 -0800
              Re: Read (and parse)  file on the web "Auric__" <not.my.real@email.address> - 2016-09-09 06:49 +0000
      Re: Read (and parse)  file on the web everonvietnam2016 <everonvietnam2016@gmail.com> - 2016-09-08 11:11 +0100
    Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-14 17:14 -0400
      Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-14 17:19 -0800
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-14 20:26 -0400
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 09:59 -0800
            Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-15 14:33 -0400
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-14 21:28 -0400
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 10:05 -0800
      Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-14 18:48 -0800
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-14 21:59 -0400
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 10:15 -0800
            Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 22:10 -0800
              Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-16 11:50 -0400
                Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-16 23:54 -0800
                  Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-17 11:06 -0400
                    Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-17 23:59 -0800
                      Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-18 03:22 -0400
                        Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-20 00:20 -0800
                  Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-17 11:14 -0400
                    Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-18 00:18 -0800
                      Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-18 03:30 -0400
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-15 05:02 -0400
          Re: Read (and parse)  file on the web - revised GS <gs@v.invalid> - 2016-09-17 17:15 -0400

Page 2 of 5 — ← Prev page 1 [2] 3 4 5  Next page →


#109258 — Re: Read (and parse) file on the web CORRECTION#2

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-01 04:23 -0800
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<ddUxz.6599$091.3871@fx29.iad>
In reply to#109253
GS wrote:
>> You are getting all of the right stuff from what i would call the
>> second file.
>> The first file is
>> "https://www.nsncenter.com/NSNSearch?q=5960%20regulator&PageNumber=" &
>> PageNum where PagNum (in ASCII) goes from "1" to "999".
>> Note the (implied?) space in the URL.
>
> I got Source, NSN Part#, Description from the 1st file. The NSN Item#
> links to the 2nd file.
>>
> <snip>
>> Would you be so kind as to share your working ADODB code?
>> Or did you hand-copy the source like i did?
>
> I use std VB file I/O not ADODB. Initial procedure was to copy/paste
> page source into Textpad and save as Tmp.txt. Then load the file into an
> array and parse from there.
>
> I thought I'd take a look at going with a userform and MS Web Browser
> control for more flexible programming opts, but haven't had the time. I
> assume this would definitely give you an advantage over trying to
> automate IE, but I need to research using it. I do have URL functions
> built into my fpSpread.ocx for doing this stuff, but that's an expensive
> 3rd party AX component. Otherwise, doing this from Excel isn't something
> I'm familiar with.
>
   Check. I know QBASIC fairly well, so a lot of that knowledge crosses 
over to VB.
   Someone here was kind enough to give me a full working program that 
can be used to copy a URL source to a temp file on the HD.
   Once available all else is very simple and straight forward.
   The rub is that function (or something it uses) does not allow a 
space in the URL,AND also does not allow https.
   So, i need two work-arounds, and the https part would seem to be the 
worst.
   I do not know how it works, what DLLs/libraries it calls; no useful 
information is available.
   It is:
   Declare Function URLDownloadToFile Lib "urlmon" _
       Alias "URLDownloadToFileA" (ByVal pCaller As Long, _
       ByVal szURL As String, ByVal szFileName As String, _
       ByVal dwReserved As Long, ByVal lpfnCB As Long) As Long

   Only the well-known keywords can be found; 'urlmon', 'pCaller', 
'szURL', and 'szFileName' are unknowns and not findable in the so-called 
VB help.
   And there are no examples; the few ranDUMB ones are incomplete and/or 
do not work..

   I do not see how you use std VB file I/O; AFAIK one cannot open a web 
page as if it was a file.



[toc] | [prev] | [next] | [standalone]


#109279 — Re: Read (and parse) file on the web CORRECTION#2

FromGS <gs@v.invalid>
Date2016-09-05 10:10 -0400
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<nqjuc4$7j8$1@dont-email.me>
In reply to#109247
>   I wish to read and parse every page of 
> "https://www.nsncenter.com/NSNSearch?q=5960%20regulator&PageNumber=" 
> where the page number goes from 5 to 999.
>    On each page, find "<a href="/NSN/5960 [it is longer, but that is 
> the start].
>    Given the full number (eg: <a href="/NSN/5960-00-831-8683"), open 
> a new related page "https://www.nsncenter.com/NSN/5960-00-831-8683" 
> and find the line ending "(MCRL)".
>    Read abut 4 lines to <a href="/PartNumber/ which is <a 
> href="/PartNumber/GV4S1400"> in this case.
> save/write that line plus the next three; close this secondary online 
> URL and step to next "<a href="/NSN/5960 to process the same way.
>    Continue to end of the page, close that URL and open the next 
> page.

Robert,
Here's what I have after parsing 'parent' pages for a list of its 
links:

    N/A
	5960-00-503-9529
	5960-00-504-8401
	5960-01-035-3901
	5960-01-029-2766
	5960-00-617-4105
	5960-00-729-5602
	5960-00-826-1280
	5960-00-754-5316
	5960-00-962-5391
	5960-00-944-4671

This is pg1 where the 1st link doesn't contain "5960" and so will be 
ignored.

Each link's text is appended to this URL to bring up its 'child' pg:

    https://www.nsncenter.com/NSN/

Each child page is parsed for the following 4 lines:

	<TD style="VERTICAL-ALIGN: middle" align=center><A 
href="/PartNumber/GV3S2800">GV3S2800</A></TD>
	<TD style="HEIGHT: 60px; WIDTH: 125px; VERTICAL-ALIGN: middle" noWrap 
align=center>&nbsp;&nbsp;<A 
href="/CAGE/63060">63060</A>&nbsp;&nbsp;</TD>
	<TD style="VERTICAL-ALIGN: middle" align=center>&nbsp;&nbsp;<A 
href="/CAGE/63060"><IMG class=img-thumbnail 
src="https://placehold.it/90x45?text=No%0DImage%0DYet" width=90 
height=45></A>&nbsp;&nbsp;</TD>
	<TD style="VERTICAL-ALIGN: middle" text-align="center"><A title="CAGE 
63060" href="/CAGE/63060">HEICO OHMITE LLC</A></TD>

I'm stripping html syntax to get this data:

  Line1:  PartNumber/GV3S2800
  Line2:  CAGE/63060
  Line3:  https://placehold.it/90x45?text=No%0DImage%0DYet
  Line4:  HEICO OHMITE LLC


The output file has these filenames in the 1st line:

    NSN Item#,Description,Part#,MCRL,CAGE,Source

I left the 3rd line URL out since, outside its host webpage, it'll be 
useless to you. I need to know from you if the 3rd line URL is needed!

Otherwise, the output file will have 1 line per item so it can be used 
as the db file "NSN_5960_ElectronTube.dat". I invite your suggestion 
for filename...

I could extend the collected data to include...

  Reference Number/DRN_3570
  Entity Code/DRN_9250
  Category Code/DRN_2910
  Variation Code/DRN_4780

..where the fieldnames would then be:

Item#,Part#,MCRL,CAGE,Source,Ref,Entity,Category,Variation

The 1st record will be:

5960-00-503-9529,GV3S2800,3302008,63060,HEICO OHMITE 
LLC,DRN_3570,DRN_9250,DRN_2910,DRN_4780

Output file size for 1 parent pg is 1Kb; for 10 parent pgs is 10Kb. You 
could have 1000 parent pgs of data stored in a 1Mb file.

Your feedback is appreciated...

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109280 — Re: Read (and parse) file on the web CORRECTION#2

FromGS <gs@v.invalid>
Date2016-09-05 10:17 -0400
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<nqjuph$93t$1@dont-email.me>
In reply to#109279
Typos...
>
> Robert,
> Here's what I have after parsing 'parent' pages for a list of its 
> links:
>
>     N/A
> 	5960-00-503-9529
> 	5960-00-504-8401
> 	5960-01-035-3901
> 	5960-01-029-2766
> 	5960-00-617-4105
> 	5960-00-729-5602
> 	5960-00-826-1280
> 	5960-00-754-5316
> 	5960-00-962-5391
> 	5960-00-944-4671
>
> This is pg1 where the 1st link doesn't contain "5960" and so will be 
> ignored.
>
> Each link's text is appended to this URL to bring up its 'child' pg:
>
>     https://www.nsncenter.com/NSN/
>
> Each child page is parsed for the following 4 lines:
>
> 	<TD style="VERTICAL-ALIGN: middle" align=center><A 
> href="/PartNumber/GV3S2800">GV3S2800</A></TD>
> 	<TD style="HEIGHT: 60px; WIDTH: 125px; VERTICAL-ALIGN: middle" 
> noWrap align=center>&nbsp;&nbsp;<A 
> href="/CAGE/63060">63060</A>&nbsp;&nbsp;</TD>
> 	<TD style="VERTICAL-ALIGN: middle" align=center>&nbsp;&nbsp;<A 
> href="/CAGE/63060"><IMG class=img-thumbnail 
> src="https://placehold.it/90x45?text=No%0DImage%0DYet" width=90 
> height=45></A>&nbsp;&nbsp;</TD>
> 	<TD style="VERTICAL-ALIGN: middle" text-align="center"><A 
> title="CAGE 63060" href="/CAGE/63060">HEICO OHMITE LLC</A></TD>
>
> I'm stripping html syntax to get this data:
>
>   Line1:  PartNumber/GV3S2800
>   Line2:  CAGE/63060
>   Line3:  https://placehold.it/90x45?text=No%0DImage%0DYet
>   Line4:  HEICO OHMITE LLC
>
>
  The output file has these fieldnames in the 1st line:
>
>     NSN Item#,Description,Part#,MCRL,CAGE,Source
>
> I left the 3rd line URL out since, outside its host webpage, it'll be 
> useless to you. I need to know from you if the 3rd line URL is 
> needed!
>
> Otherwise, the output file will have 1 line per item so it can be 
> used as the db file "NSN_5960_ElectronTube.dat". I invite your 
> suggestion for filename...
>
> I could extend the collected data to include...
>
>   Reference Number/DRN_3570
>   Entity Code/DRN_9250
>   Category Code/DRN_2910
>   Variation Code/DRN_4780
>
> ..where the fieldnames would then be:
>
  Item#,Part#,MCRL,CAGE,Source,REF,ENT,CAT,VAR
>
> The 1st record will be:
>
> 5960-00-503-9529,GV3S2800,3302008,63060,HEICO OHMITE 
> LLC,DRN_3570,DRN_9250,DRN_2910,DRN_4780
>
> Output file size for 1 parent pg is 1Kb; for 10 parent pgs is 10Kb. 
> You could have 1000 parent pgs of data stored in a 1Mb file.
>
> Your feedback is appreciated...

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109282 — Re: Read (and parse) file on the web CORRECTION#3

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-05 20:50 -0800
SubjectRe: Read (and parse) file on the web CORRECTION#3
Message-ID<A2rzz.42429$e%3.31863@fx07.iad>
In reply to#109280
GS wrote:
> Typos...
>>
>> Robert,
>> Here's what I have after parsing 'parent' pages for a list of its links:
>>
>> N/A
>> 5960-00-503-9529
>> 5960-00-504-8401
>> 5960-01-035-3901
>> 5960-01-029-2766
>> 5960-00-617-4105
>> 5960-00-729-5602
>> 5960-00-826-1280
>> 5960-00-754-5316
>> 5960-00-962-5391
>> 5960-00-944-4671
>>
>> This is pg1 where the 1st link doesn't contain "5960" and so will be
>> ignored.
>>
>> Each link's text is appended to this URL to bring up its 'child' pg:
>>
>> https://www.nsncenter.com/NSN/
>>
>> Each child page is parsed for the following 4 lines:
>>
>> <TD style="VERTICAL-ALIGN: middle" align=center><A
>> href="/PartNumber/GV3S2800">GV3S2800</A></TD>
>> <TD style="HEIGHT: 60px; WIDTH: 125px; VERTICAL-ALIGN: middle" noWrap
>> align=center>&nbsp;&nbsp;<A href="/CAGE/63060">63060</A>&nbsp;&nbsp;</TD>
>> <TD style="VERTICAL-ALIGN: middle" align=center>&nbsp;&nbsp;<A
>> href="/CAGE/63060"><IMG class=img-thumbnail
>> src="https://placehold.it/90x45?text=No%0DImage%0DYet" width=90
>> height=45></A>&nbsp;&nbsp;</TD>
>> <TD style="VERTICAL-ALIGN: middle" text-align="center"><A title="CAGE
>> 63060" href="/CAGE/63060">HEICO OHMITE LLC</A></TD>
>>
>> I'm stripping html syntax to get this data:
>>
>> Line1: PartNumber/GV3S2800
>> Line2: CAGE/63060
>> Line3: https://placehold.it/90x45?text=No%0DImage%0DYet
>> Line4: HEICO OHMITE LLC
>>
>>
> The output file has these fieldnames in the 1st line:
>>
>> NSN Item#,Description,Part#,MCRL,CAGE,Source
>>
>> I left the 3rd line URL out since, outside its host webpage, it'll be
>> useless to you. I need to know from you if the 3rd line URL is needed!
>>
>> Otherwise, the output file will have 1 line per item so it can be used
>> as the db file "NSN_5960_ElectronTube.dat". I invite your suggestion
>> for filename...
>>
>> I could extend the collected data to include...
>>
>> Reference Number/DRN_3570
>> Entity Code/DRN_9250
>> Category Code/DRN_2910
>> Variation Code/DRN_4780
>>
>> ..where the fieldnames would then be:
>>
> Item#,Part#,MCRL,CAGE,Source,REF,ENT,CAT,VAR
>>
>> The 1st record will be:
>>
>> 5960-00-503-9529,GV3S2800,3302008,63060,HEICO OHMITE
>> LLC,DRN_3570,DRN_9250,DRN_2910,DRN_4780
>>
>> Output file size for 1 parent pg is 1Kb; for 10 parent pgs is 10Kb.
>> You could have 1000 parent pgs of data stored in a 1Mb file.
>>
>> Your feedback is appreciated...
>
   Like i said, PERFECT!
   And you are correct, do not need line 3 nor the extended data.
   Please check my other answer for a corrected search term which needs 
a !corrected! human-readable version of the URL is:
https://www.nsncenter.com/NSNSearch?q=5960 regulator and "ELECTRON 
TUBE"&PageNumber=1
   The %20 is a virtual space, and the %22 is a virtual quote.
   ((guess that is the proper term))

[toc] | [prev] | [next] | [standalone]


#109281 — Re: Read (and parse) file on the web CORRECTION#2

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-05 20:32 -0800
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<JNqzz.334$OT5.113@fx19.iad>
In reply to#109279
GS wrote:
>> I wish to read and parse every page of
>> "https://www.nsncenter.com/NSNSearch?q=5960%20regulator&PageNumber="
>> where the page number goes from 5 to 999.
>> On each page, find "<a href="/NSN/5960 [it is longer, but that is the
>> start].
>> Given the full number (eg: <a href="/NSN/5960-00-831-8683"), open a
>> new related page "https://www.nsncenter.com/NSN/5960-00-831-8683" and
>> find the line ending "(MCRL)".
>> Read abut 4 lines to <a href="/PartNumber/ which is <a
>> href="/PartNumber/GV4S1400"> in this case.
>> save/write that line plus the next three; close this secondary online
>> URL and step to next "<a href="/NSN/5960 to process the same way.
>> Continue to end of the page, close that URL and open the next page.
>
> Robert,
> Here's what I have after parsing 'parent' pages for a list of its links:
>
> N/A
> 5960-00-503-9529
> 5960-00-504-8401
> 5960-01-035-3901
> 5960-01-029-2766
> 5960-00-617-4105
> 5960-00-729-5602
> 5960-00-826-1280
> 5960-00-754-5316
> 5960-00-962-5391
> 5960-00-944-4671
>
> This is pg1 where the 1st link doesn't contain "5960" and so will be
> ignored.
>
> Each link's text is appended to this URL to bring up its 'child' pg:
>
> https://www.nsncenter.com/NSN/
>
> Each child page is parsed for the following 4 lines:
>
> <TD style="VERTICAL-ALIGN: middle" align=center><A
> href="/PartNumber/GV3S2800">GV3S2800</A></TD>
> <TD style="HEIGHT: 60px; WIDTH: 125px; VERTICAL-ALIGN: middle" noWrap
> align=center>&nbsp;&nbsp;<A href="/CAGE/63060">63060</A>&nbsp;&nbsp;</TD>
> <TD style="VERTICAL-ALIGN: middle" align=center>&nbsp;&nbsp;<A
> href="/CAGE/63060"><IMG class=img-thumbnail
> src="https://placehold.it/90x45?text=No%0DImage%0DYet" width=90
> height=45></A>&nbsp;&nbsp;</TD>
> <TD style="VERTICAL-ALIGN: middle" text-align="center"><A title="CAGE
> 63060" href="/CAGE/63060">HEICO OHMITE LLC</A></TD>
>
> I'm stripping html syntax to get this data:
>
> Line1: PartNumber/GV3S2800
> Line2: CAGE/63060
> Line3: https://placehold.it/90x45?text=No%0DImage%0DYet
> Line4: HEICO OHMITE LLC
>
>
> The output file has these filenames in the 1st line:
>
> NSN Item#,Description,Part#,MCRL,CAGE,Source
>
> I left the 3rd line URL out since, outside its host webpage, it'll be
> useless to you. I need to know from you if the 3rd line URL is needed!
>
> Otherwise, the output file will have 1 line per item so it can be used
> as the db file "NSN_5960_ElectronTube.dat". I invite your suggestion for
> filename...
>
> I could extend the collected data to include...
>
> Reference Number/DRN_3570
> Entity Code/DRN_9250
> Category Code/DRN_2910
> Variation Code/DRN_4780
>
> ..where the fieldnames would then be:
>
> Item#,Part#,MCRL,CAGE,Source,Ref,Entity,Category,Variation
>
> The 1st record will be:
>
> 5960-00-503-9529,GV3S2800,3302008,63060,HEICO OHMITE
> LLC,DRN_3570,DRN_9250,DRN_2910,DRN_4780
>
> Output file size for 1 parent pg is 1Kb; for 10 parent pgs is 10Kb. You
> could have 1000 parent pgs of data stored in a 1Mb file.
>
> Your feedback is appreciated...
>
   WOW!
   Absolutely PERFECT!
   You are correct, #1) do not need that line 3, and #2) do not need the 
extended info.

   File name(s) for PageNumber=1 I would use 5960_001.TXT,..to 
PageNumber=999 I would use 5960_999.TXT and that would preserve order.
   *OR*
   Reading & parsing from PageNumber=1 to PageNumber=999,one could 
append to same file (name NSN_5960.TXT); might as well - makes it easier 
to pour into a single Excel file.
   Either way is fine.

   I have found a way to get rid of items that are not strictly electron 
tubes and/or not regulators; that way you do not have to parse out these 
"unfit" items from first page description. Use:
"https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=1"
   Naturally, PageNumber still goes from 1 to 999.
   Note the implied "(", ")" and " "; human-readable "5960 regulator and 
(ELECTRON TUBE)".
   As far as i can tell, using that shows no undesirable parts.

   Thanks!
PS: i found WGET to be non-useful (a) it truncates the filename (b) it 
buggers it to partial gibberish.

[toc] | [prev] | [next] | [standalone]


#109283 — Re: Read (and parse) file on the web CORRECTION#2

FromGS <gs@v.invalid>
Date2016-09-06 10:49 -0400
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<nqml2k$3t7$1@dont-email.me>
In reply to#109281
>    You are correct, #1) do not need that line 3, and #2) do not need 
> the extended info.

Ok then, fieldnames will be:  Item#,Part#,MCRL,CAGE,Source
>
>    File name(s) for PageNumber=1 I would use 5960_001.TXT,..to 
> PageNumber=999 I would use 5960_999.TXT and that would preserve 
> order.
>    *OR*
>    Reading & parsing from PageNumber=1 to PageNumber=999,one could 
> append to same file (name NSN_5960.TXT); might as well - makes it 
> easier to pour into a single Excel file.
>    Either way is fine.

Ok, then ouput filename will be:  NSN_5960.txt
>
>    I have found a way to get rid of items that are not strictly 
> electron tubes and/or not regulators; that way you do not have to 
> parse out these "unfit" items from first page description. Use:
> "https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=1"
>    Naturally, PageNumber still goes from 1 to 999.
>    Note the implied "(", ")" and " "; human-readable "5960 regulator 
> and (ELECTRON TUBE)".
>    As far as i can tell, using that shows no undesirable parts.

Works nice! Now I get 11 5960 items per parent page.
>
>    Thanks!
> PS: i found WGET to be non-useful (a) it truncates the filename (b) 
> it buggers it to partial gibberish

What is WGET?

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109284 — Re: Read (and parse) file on the web CORRECTION#2

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-06 12:05 -0800
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<WrEzz.32969$BN1.19549@fx02.iad>
In reply to#109283
GS wrote:
>> You are correct, #1) do not need that line 3, and #2) do not need the
>> extended info.
>
> Ok then, fieldnames will be: Item#,Part#,MCRL,CAGE,Source
>>
>> File name(s) for PageNumber=1 I would use 5960_001.TXT,..to
>> PageNumber=999 I would use 5960_999.TXT and that would preserve order.
>> *OR*
>> Reading & parsing from PageNumber=1 to PageNumber=999,one could append
>> to same file (name NSN_5960.TXT); might as well - makes it easier to
>> pour into a single Excel file.
>> Either way is fine.
>
> Ok, then ouput filename will be: NSN_5960.txt
>>
>> I have found a way to get rid of items that are not strictly electron
>> tubes and/or not regulators; that way you do not have to parse out
>> these "unfit" items from first page description. Use:
>> "https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=1"
>>
>> Naturally, PageNumber still goes from 1 to 999.
>> Note the implied "(", ")" and " "; human-readable "5960 regulator and
>> (ELECTRON TUBE)".
>> As far as i can tell, using that shows no undesirable parts.
>
> Works nice! Now I get 11 5960 items per parent page.
>>
>> Thanks!
>> PS: i found WGET to be non-useful (a) it truncates the filename (b) it
>> buggers it to partial gibberish
>
> What is WGET?
>
   WGET is a command line program that will copy contents of an URL to 
the hard drive; it has various options, for SSL, i think for some 
processing, for giving the output file a specific name, for recursion, etc.
   Was still trying to find ways to copy the online file to the hard drive.

   I still do not understand what magic you used.

   Now, the nitty-gritty; in exchange for that nicely parsed file, what 
do i owe you?

[toc] | [prev] | [next] | [standalone]


#109285 — Re: Read (and parse) file on the web CORRECTION#2

FromGS <gs@v.invalid>
Date2016-09-06 15:15 -0400
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<nqn4k6$32u$1@dont-email.me>
In reply to#109284
>  I still do not understand what magic you used.

I'm using the MS WebBrowser control and a textbox on a worksheet!
>
>    Now, the nitty-gritty; in exchange for that nicely parsed file, 
> what do i owe you?

A Timmies, straight up!

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109287 — Re: Read (and parse) file on the web CORRECTION#2

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-06 22:32 -0800
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<vDNzz.60413$oU2.8638@fx09.iad>
In reply to#109285
GS wrote:
>> I still do not understand what magic you used.
>
> I'm using the MS WebBrowser control and a textbox on a worksheet!
>>
>> Now, the nitty-gritty; in exchange for that nicely parsed file, what
>> do i owe you?
>
> A Timmies, straight up!
>
   The search engine was not exactly forthcoming, to say the least; 
everything including the kitchen sink but NOT anything alcoholic.
   "Timmies drink" helped some; fifth "hit" down: "Timmy's Sweet and 
Sour mix Cocktails and Drink Recipes".
   Using "Timmies, straight up" was slightly better.."Average night at 
the Manotick Timmies... : ottawa"

   In all of this,a lot of "hits" mentioned something(always different) 
about Tim Hortons Franchise.

   Absolutely no clue regarding rum, scotch, vodka (or dare i say) milk.


[toc] | [prev] | [next] | [standalone]


#109292 — Re: Read (and parse) file on the web CORRECTION#2

FromGS <gs@v.invalid>
Date2016-09-07 11:44 -0400
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<nqpclr$t1o$1@dont-email.me>
In reply to#109287
> GS wrote:
>>> I still do not understand what magic you used.
>>
>> I'm using the MS WebBrowser control and a textbox on a worksheet!
>>>
>>> Now, the nitty-gritty; in exchange for that nicely parsed file, 
>>> what
>>> do i owe you?
>>
>> A Timmies, straight up!
>>
>    The search engine was not exactly forthcoming, to say the least; 
> everything including the kitchen sink but NOT anything alcoholic.
>    "Timmies drink" helped some; fifth "hit" down: "Timmy's Sweet and 
> Sour mix Cocktails and Drink Recipes".
>    Using "Timmies, straight up" was slightly better.."Average night 
> at the Manotick Timmies... : ottawa"
>
>    In all of this,a lot of "hits" mentioned something(always 
> different) about Tim Hortons Franchise.
>
>    Absolutely no clue regarding rum, scotch, vodka (or dare i say) 
> milk.

Ha-ha! Ok.., 'Timmies' is fan-speak for Tim Horton's coffee!<g>

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109293 — Re: Read (and parse) file on the web CORRECTION#2

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-07 09:33 -0800
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<pjXzz.57656$SX5.14315@fx01.iad>
In reply to#109292
GS wrote:
>> GS wrote:
>>>> I still do not understand what magic you used.
>>>
>>> I'm using the MS WebBrowser control and a textbox on a worksheet!
>>>>
>>>> Now, the nitty-gritty; in exchange for that nicely parsed file, what
>>>> do i owe you?
>>>
>>> A Timmies, straight up!
>>>
>> The search engine was not exactly forthcoming, to say the least;
>> everything including the kitchen sink but NOT anything alcoholic.
>> "Timmies drink" helped some; fifth "hit" down: "Timmy's Sweet and Sour
>> mix Cocktails and Drink Recipes".
>> Using "Timmies, straight up" was slightly better.."Average night at
>> the Manotick Timmies... : ottawa"
>>
>> In all of this,a lot of "hits" mentioned something(always different)
>> about Tim Hortons Franchise.
>>
>> Absolutely no clue regarding rum, scotch, vodka (or dare i say) milk.
>
> Ha-ha! Ok.., 'Timmies' is fan-speak for Tim Horton's coffee!<g>
>
   What i did WRT Wget, was uninstall it and checked that there were 
'dregs' on the HD.
   Then i installed it from scratch, allowing all of the defaults.
   Finally, i modified the system path (shortened version):
%SystemRoot%\system32;%SystemRoot%;%SystemRoot%\System32\Wbem;C:\Program 
Files\GnuWin32;

   No joy.
   Even at the root, the system insists wget does not exist (as an 
executable, etc).

[toc] | [prev] | [next] | [standalone]


#109301 — Re: Read (and parse) file on the web CORRECTION#2

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-08 20:18 -0800
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<gSpAz.36246$Q97.23130@fx06.iad>
In reply to#109292
GS wrote:
>> GS wrote:
>>>> I still do not understand what magic you used.
>>>
>>> I'm using the MS WebBrowser control and a textbox on a worksheet!
>>>>
>>>> Now, the nitty-gritty; in exchange for that nicely parsed file, what
>>>> do i owe you?
>>>
>>> A Timmies, straight up!
>>>
>> The search engine was not exactly forthcoming, to say the least;
>> everything including the kitchen sink but NOT anything alcoholic.
>> "Timmies drink" helped some; fifth "hit" down: "Timmy's Sweet and Sour
>> mix Cocktails and Drink Recipes".
>> Using "Timmies, straight up" was slightly better.."Average night at
>> the Manotick Timmies... : ottawa"
>>
>> In all of this,a lot of "hits" mentioned something(always different)
>> about Tim Hortons Franchise.
>>
>> Absolutely no clue regarding rum, scotch, vodka (or dare i say) milk.
>
> Ha-ha! Ok.., 'Timmies' is fan-speak for Tim Horton's coffee!<g>
>
   Good show.
   You have an e-mail account,i presume?
   I could use PayPal to send some nickels for that parsed NSN file if 
you wish.
   Please let me know how soon i can get that file.
   Thanks.

[toc] | [prev] | [next] | [standalone]


#109304 — Re: Read (and parse) file on the web CORRECTION#2

FromGS <gs@v.invalid>
Date2016-09-09 06:07 -0400
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<nqu1kb$9bu$1@dont-email.me>
In reply to#109301
>> Ha-ha! Ok.., 'Timmies' is fan-speak for Tim Horton's coffee!<g>
>>
>    Good show.
>    You have an e-mail account,i presume?
>    I could use PayPal to send some nickels for that parsed NSN file 
> if you wish.
>    Please let me know how soon i can get that file.
>    Thanks.

No worries.., I'm just kidding! That's what I tell neighbors when they 
offer payment for helping them with computer issues!

As I stated earlier, I'm commited right now and so only have time to 
work on this when I get a chance. Currently, it's ready to fully 
automate, but seems to have a snag writing past the 1st parent page's 
child pages. I'm likely going to have time to finish it this weekend, 
though, but I'll post a download link when done, ..regardless!

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109305 — Re: Read (and parse) file on the web CORRECTION#2

FromGS <gs@v.invalid>
Date2016-09-09 22:05 -0400
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<nqvpph$4pp$1@dont-email.me>
In reply to#109304
> Currently, it's ready to fully automate, but seems to have a snag 
> writing past the 1st parent page's child pages.

Turns out the problem was code not waiting for the browser not busy. I 
switched to using URLDownloadToFile() at this point because it's orders 
of magnitude faster. Using the browser/twxtbox on a sheet served well 
for getting the process code nailed down, but that was only a temp 
situation during dev.

The links error out about mid pg7 thru pg10 as I time tested only pgs 
1thru10: -this took 50.89 secs!

I'll do some housekeeping of the code and post a download link to the 
file...

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109308 — Re: Read (and parse) file on the web CORRECTION#2

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-14 17:16 -0800
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<pLlCz.24829$fL2.8044@fx20.iad>
In reply to#109305
GS wrote:
>> Currently, it's ready to fully automate, but seems to have a snag
>> writing past the 1st parent page's child pages.
>
> Turns out the problem was code not waiting for the browser not busy. I
> switched to using URLDownloadToFile() at this point because it's orders
> of magnitude faster. Using the browser/twxtbox on a sheet served well
> for getting the process code nailed down, but that was only a temp
> situation during dev.
   "Problem" with Excel, is that there are MANY ways to get what is 
needed, and there is NO WAY of discovering _any_ of them; the "help" 
document is worse than useless in that manner.
   I have found that URLDownloadToFile() to be non-functional for https
  sources.

>
> The links error out about mid pg7 thru pg10 as I time tested only pgs
> 1thru10: -this took 50.89 secs!
>
> I'll do some housekeeping of the code and post a download link to the
> file...
>

[toc] | [prev] | [next] | [standalone]


#109310 — Re: Read (and parse) file on the web CORRECTION#2

FromGS <gs@v.invalid>
Date2016-09-14 20:23 -0400
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<nrcpm5$eu0$1@dont-email.me>
In reply to#109308
>    "Problem" with Excel, is that there are MANY ways to get what is 
> needed, and there is NO WAY of discovering _any_ of them; the "help" 
> document is worse than useless in that manner.
>    I have found that URLDownloadToFile() to be non-functional for 
> https
>   sources.

I disagree because it's working in my project I posted the download 
for!

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109286 — Re: Read (and parse) file on the web CORRECTION#2

From"Auric__" <not.my.real@email.address>
Date2016-09-07 02:37 +0000
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<XnsA67BC7911E470auricauricauricauric@213.239.209.88>
In reply to#109281
Robert Baer wrote:

> PS: i found WGET to be non-useful (a) it truncates the filename (b) it
> buggers it to partial gibberish.

Then you must be using a bad version, or perhaps have something wrong with 
your .wgetrc. I've been using wget for around 10 years, and never had 
anything like those issues unless I pass bad options.

-- 
My life is richer, somehow, simply because I know that he exists.

[toc] | [prev] | [next] | [standalone]


#109288 — Re: Read (and parse) file on the web CORRECTION#2

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-06 22:43 -0800
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<8ONzz.56531$M27.26468@fx33.iad>
In reply to#109286
Auric__ wrote:
> Robert Baer wrote:
>
>> PS: i found WGET to be non-useful (a) it truncates the filename (b) it
>> buggers it to partial gibberish.
>
> Then you must be using a bad version, or perhaps have something wrong with
> your .wgetrc. I've been using wget for around 10 years, and never had
> anything like those issues unless I pass bad options.
>
   Know nothing about .wgetrc; am in Win2K cmd line, and the batch file 
used is:
H:
CD\Win2K_WORK\OIL4LESS\LLCDOCS\FED app\FBA stuff

wget --no-check-certificate --output-document=5960_002.TXT 
--output-file=log002.TXT 
https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=2

PAUSE

   The SourceForge site offered a Zip which was supposed to be complete, 
but none of the created folders had an EXE (tried Win2K, WinXP, Win7).
   Found SofTonic offering only a plain jane wget.exe, which i am using, 
so that may be a buggered version.
   Suggestions?

[toc] | [prev] | [next] | [standalone]


#109295 — Re: Read (and parse) file on the web CORRECTION#2

From"Auric__" <not.my.real@email.address>
Date2016-09-08 05:49 +0000
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<XnsA67CE831F3B1Cauricauricauricauric@213.239.209.88>
In reply to#109288
Robert Baer wrote:

> Auric__ wrote:
>> Robert Baer wrote:
>>
>>> PS: i found WGET to be non-useful (a) it truncates the filename (b) it
>>> buggers it to partial gibberish.
>>
>> Then you must be using a bad version, or perhaps have something wrong
>> with your .wgetrc. I've been using wget for around 10 years, and never
>> had anything like those issues unless I pass bad options.
>>
> Know nothing about .wgetrc;

Don't worry about it. It can be used to set default behaviors but every entry 
can be replicated via switches.

> am in Win2K cmd line, and the batch file 
> used is:
> H:
> CD\Win2K_WORK\OIL4LESS\LLCDOCS\FED app\FBA stuff
> 
> wget --no-check-certificate --output-document=5960_002.TXT 
> --output-file=log002.TXT 
> https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRO
> N%20TUBE%22&PageNumber=2 

That wget line performs as expected for me: 5960_002.TXT contains valid HTML 
(although I haven't made any attempt to check the data; it looks like most of 
the page is CSS) and log002.TXT is a typical wget log of a successful 
transfer.

As for truncating the filenames, if I remove the --output-document switch, 
the filename I get is

  NSNSearch@q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=2

> PAUSE
> 
>    The SourceForge site offered a Zip which was supposed to be complete,

If you're talking about GNUwin32, that version is years out of date.

> but none of the created folders had an EXE (tried Win2K, WinXP, Win7).
>    Found SofTonic offering only a plain jane wget.exe, which i am using,
> so that may be a buggered version.

Never even heard of them.

>    Suggestions?

I'm using 1.16.3. No idea where I got it.

The batch file that I use for downloading looks like this:

  call wget --no-check-certificate -x -c -e robots=off -i new.txt %*

-x  Always create directories (e.g. http://a.b.c/1/2.txt -> .\a.b.c\1\2.txt).
-c  Continue interrupted downloads.
-e  Do this .wgetrc thing (in this case, ignore the robots.txt file).
-i  Read list of filenames from the following file ("new.txt" because that's
    the default name for a new file in my file manager).

I use the -i switch so I don't have to worry about escaping characters or % 
vs %%. Whatever's in the text file is exactly what it looks for. (If you go 
this route, it's one file per line.)

-- 
We have to stop letting George Lucas name our politicians.

[toc] | [prev] | [next] | [standalone]


#109296 — Re: Read (and parse) file on the web CORRECTION#2

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-08 01:07 -0800
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<8%8Az.103752$PM.41042@fx08.iad>
In reply to#109295
Auric__ wrote:
> Robert Baer wrote:
>
>> Auric__ wrote:
>>> Robert Baer wrote:
>>>
>>>> PS: i found WGET to be non-useful (a) it truncates the filename (b) it
>>>> buggers it to partial gibberish.
>>>
>>> Then you must be using a bad version, or perhaps have something wrong
>>> with your .wgetrc. I've been using wget for around 10 years, and never
>>> had anything like those issues unless I pass bad options.
>>>
>> Know nothing about .wgetrc;
>
> Don't worry about it. It can be used to set default behaviors but every entry
> can be replicated via switches.
>
>> am in Win2K cmd line, and the batch file
>> used is:
>> H:
>> CD\Win2K_WORK\OIL4LESS\LLCDOCS\FED app\FBA stuff
>>
>> wget --no-check-certificate --output-document=5960_002.TXT
>> --output-file=log002.TXT
>> https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECTRO
>> N%20TUBE%22&PageNumber=2
>
> That wget line performs as expected for me: 5960_002.TXT contains valid HTML
> (although I haven't made any attempt to check the data; it looks like most of
> the page is CSS) and log002.TXT is a typical wget log of a successful
> transfer.
>
> As for truncating the filenames, if I remove the --output-document switch,
> the filename I get is
>
>    NSNSearch@q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=2
>
>> PAUSE
>>
>>     The SourceForge site offered a Zip which was supposed to be complete,
>
> If you're talking about GNUwin32, that version is years out of date.
>
>> but none of the created folders had an EXE (tried Win2K, WinXP, Win7).
>>     Found SofTonic offering only a plain jane wget.exe, which i am using,
>> so that may be a buggered version.
>
> Never even heard of them.
>
>>     Suggestions?
>
> I'm using 1.16.3. No idea where I got it.
>
> The batch file that I use for downloading looks like this:
>
>    call wget --no-check-certificate -x -c -e robots=off -i new.txt %*
>
> -x  Always create directories (e.g. http://a.b.c/1/2.txt ->  .\a.b.c\1\2.txt).
> -c  Continue interrupted downloads.
> -e  Do this .wgetrc thing (in this case, ignore the robots.txt file).
> -i  Read list of filenames from the following file ("new.txt" because that's
>      the default name for a new file in my file manager).
>
> I use the -i switch so I don't have to worry about escaping characters or %
> vs %%. Whatever's in the text file is exactly what it looks for. (If you go
> this route, it's one file per line.)
>
   You must have a different version of Wget; whatever i do on the 
command line,including the "trick" of restrict-file-names=nocontrol, i 
get a buggered path name plus the response &PageNumber not recognized.
   Exactly same results in Win2K, WinXP or in Win7.

   Yes, i used GNUwin32 as SourceForge "complete" of Wget had no EXE.
   Is there some other (compiled, complete) source i should get?

[toc] | [prev] | [next] | [standalone]


Page 2 of 5 — ← Prev page 1 [2] 3 4 5  Next page →

Back to top | Article view | microsoft.public.excel.programming


csiph-web