Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > microsoft.public.excel.programming > #109214 > unrolled thread

Read (and parse) file on the web

Started byRobert Baer <robertbaer@localnet.com>
First post2016-08-26 10:57 -0800
Last post2016-09-17 17:15 -0400
Articles 20 on this page of 92 — 6 participants

Back to article view | Back to microsoft.public.excel.programming


Contents

  Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-26 10:57 -0800
    Re: Read (and parse)  file on the web "Auric__" <not.my.real@email.address> - 2016-08-26 15:21 +0000
      Re: Read (and parse)  file on the web "Auric__" <not.my.real@email.address> - 2016-08-26 15:22 +0000
        Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-27 13:04 -0800
      Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-27 13:00 -0800
      Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-27 23:14 -0800
        Re: Read (and parse)  file on the web CORRECTION Robert Baer <robertbaer@localnet.com> - 2016-08-28 09:42 -0800
          Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-28 10:08 -0800
            Re: Read (and parse)  file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-08-29 01:51 +0000
              Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-28 21:32 -0800
                Re: Read (and parse)  file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-08-29 19:41 +0000
                  Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-30 00:21 -0800
                    Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-08-30 06:06 -0400
                      Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-08-30 14:20 -0800
                        Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-08-30 18:43 -0400
                          Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-08-31 16:47 -0400
                            Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-01 04:37 -0800
                              Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-01 13:13 -0400
                                Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-01 16:32 -0800
                                  Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-01 19:43 -0400
                          Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-01 04:23 -0800
                    Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-05 10:10 -0400
                      Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-05 10:17 -0400
                        Re: Read (and parse)  file on the web CORRECTION#3 Robert Baer <robertbaer@localnet.com> - 2016-09-05 20:50 -0800
                      Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-05 20:32 -0800
                        Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-06 10:49 -0400
                          Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 12:05 -0800
                            Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-06 15:15 -0400
                              Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 22:32 -0800
                                Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-07 11:44 -0400
                                  Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-07 09:33 -0800
                                  Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-08 20:18 -0800
                                    Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-09 06:07 -0400
                                      Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-09 22:05 -0400
                                        Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-14 17:16 -0800
                                          Re: Read (and parse)  file on the web CORRECTION#2 GS <gs@v.invalid> - 2016-09-14 20:23 -0400
                        Re: Read (and parse)  file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-09-07 02:37 +0000
                          Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 22:43 -0800
                            Re: Read (and parse)  file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-09-08 05:49 +0000
                              Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-08 01:07 -0800
                                Re: Read (and parse)  file on the web CORRECTION#2 "Auric__" <not.my.real@email.address> - 2016-09-09 06:40 +0000
                              Re: Read (and parse)  file on the web WGET versions Robert Baer <robertbaer@localnet.com> - 2016-09-08 01:19 -0800
                                Re: Read (and parse)  file on the web WGET versions Robert Baer <robertbaer@localnet.com> - 2016-09-08 20:09 -0800
                          Re: Read (and parse)  file on the web CORRECTION#2 Robert Baer <robertbaer@localnet.com> - 2016-09-06 23:49 -0800
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-28 15:51 -0400
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-28 21:37 -0800
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-02 19:09 -0800
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-29 01:56 -0400
          Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-29 02:21 -0400
            Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 19:55 -0800
              Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-29 23:35 -0400
                Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 22:25 -0800
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 19:21 -0800
            Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-29 23:21 -0400
              Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 22:10 -0800
                Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-30 04:04 -0400
                  Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-30 13:48 -0800
            Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-08-29 23:37 -0400
              Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-08-29 22:25 -0800
      Re: Read (and parse)  file on the web lovexes17816 <lovexes17816@gmail.com> - 2016-08-28 14:00 +0100
      Re: Read (and parse)  file on the web lovexes17816 <lovexes17816@gmail.com> - 2016-08-29 02:17 +0100
        Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-02 19:11 -0800
          Re: Read (and parse)  file on the web "Auric__" <not.my.real@email.address> - 2016-09-03 04:12 +0000
      Re: Read (and parse)  file on the web Tikivn23143 <Tikivn23143@gmail.com> - 2016-08-30 02:28 +0100
      Re: Read (and parse)  file on the web everonvietnam2016 <everonvietnam2016@gmail.com> - 2016-09-07 15:03 +0100
        Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-07 08:43 -0800
          Re: Read (and parse)  file on the web "Auric__" <not.my.real@email.address> - 2016-09-08 05:43 +0000
            Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-08 01:09 -0800
              Re: Read (and parse)  file on the web "Auric__" <not.my.real@email.address> - 2016-09-09 06:49 +0000
      Re: Read (and parse)  file on the web everonvietnam2016 <everonvietnam2016@gmail.com> - 2016-09-08 11:11 +0100
    Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-14 17:14 -0400
      Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-14 17:19 -0800
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-14 20:26 -0400
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 09:59 -0800
            Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-15 14:33 -0400
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-14 21:28 -0400
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 10:05 -0800
      Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-14 18:48 -0800
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-14 21:59 -0400
          Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 10:15 -0800
            Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-15 22:10 -0800
              Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-16 11:50 -0400
                Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-16 23:54 -0800
                  Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-17 11:06 -0400
                    Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-17 23:59 -0800
                      Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-18 03:22 -0400
                        Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-20 00:20 -0800
                  Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-17 11:14 -0400
                    Re: Read (and parse)  file on the web Robert Baer <robertbaer@localnet.com> - 2016-09-18 00:18 -0800
                      Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-18 03:30 -0400
        Re: Read (and parse)  file on the web GS <gs@v.invalid> - 2016-09-15 05:02 -0400
          Re: Read (and parse)  file on the web - revised GS <gs@v.invalid> - 2016-09-17 17:15 -0400

Page 3 of 5 — ← Prev page 1 2 [3] 4 5  Next page →


#109302 — Re: Read (and parse) file on the web CORRECTION#2

From"Auric__" <not.my.real@email.address>
Date2016-09-09 06:40 +0000
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<XnsA67DF0C99B138auricauricauricauric@213.239.209.88>
In reply to#109296
Robert Baer wrote:

> Auric__ wrote:
>> Robert Baer wrote:
[snip]
>>> CD\Win2K_WORK\OIL4LESS\LLCDOCS\FED app\FBA stuff
>>>
>>> wget --no-check-certificate --output-document=5960_002.TXT
>>> --output-file=log002.TXT
>>> https://www.nsncenter.com/NSNSearch?q=5960%20regulator%20and%20%22ELECT
>>> RO N%20TUBE%22&PageNumber=2
>>
>> That wget line performs as expected for me: 5960_002.TXT contains valid
>> HTML (although I haven't made any attempt to check the data; it looks
>> like most of the page is CSS) and log002.TXT is a typical wget log of a
>> successful transfer.
>>
>> As for truncating the filenames, if I remove the --output-document
>> switch, the filename I get is
>>
>>   NSNSearch@q=5960%20regulator%20and%20%22ELECTRON%20TUBE%22&PageNumber=2 
[snip]
>> If you're talking about GNUwin32, that version is years out of date.
>>
>    You must have a different version of Wget; whatever i do on the 
> command line,including the "trick" of restrict-file-names=nocontrol, i 
> get a buggered path name plus the response &PageNumber not recognized.
>    Exactly same results in Win2K, WinXP or in Win7.

Hmm. Well... it could be that your copy of wget was compiled with old path 
length limits (260 characters). I suppose the best thing to do there is to 
try a different copy.

> Yes, i used GNUwin32 as SourceForge "complete" of Wget had no EXE.
> Is there some other (compiled, complete) source i should get?

Just google "wget windows" (without quotes) and start poking around. 
Download a few different versions and see if any of them work for you.

-- 
Stupid railroad plot.

[toc] | [prev] | [next] | [standalone]


#109298 — Re: Read (and parse) file on the web WGET versions

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-08 01:19 -0800
SubjectRe: Read (and parse) file on the web WGET versions
Message-ID<ha9Az.35953$Q97.20693@fx06.iad>
In reply to#109295
   Here are other versions of Wget i found; which do you recommend?

Windows binaries of GNU Wget - eternallybored.org
Windows binaries of GNU Wget A command-line utility for retrieving files 
using HTTP, HTTPS and FTP protocols. Warning: some antivirus tools 
recognise wget-1.18-win32 ...
[Search domain eternallybored.org] eternallybored.org/misc/wget/

WGET - Download
WGET, free and safe download. WGET latest version:.
[Search domain wget.en.softonic.com] wget.en.softonic.com
** This gives ONLY Wget.EXE

Wget - Free downloads and reviews - CNET Download.com
wget free download - SimpleWget, WinWGet, Nugget, and many more programs
[Search domain download.cnet.com] download.cnet.com/s/wget/

VisualWget Download - Softpedia
VisualWget was designed as a GUI front-end for Wget, a content retriever 
originally designed for GNU that has been on the market longer than we 
can remember.
[Search domain www.softpedia.com] 
softpedia.com/get/Internet/Download-Managers/VisualWg

Download WGET free - latest version
Download WGET now from Softonic: 100% safe and virus free. More than 327 
downloads this month. Download WGET latest version for free
[Search domain wget.en.softonic.com] wget.en.softonic.com/download
**
   I tried the eternallybored.org version, same results.

   Thanks.

[toc] | [prev] | [next] | [standalone]


#109300 — Re: Read (and parse) file on the web WGET versions

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-08 20:09 -0800
SubjectRe: Read (and parse) file on the web WGET versions
Message-ID<cKpAz.61129$eM3.52701@fx04.iad>
In reply to#109298
Robert Baer wrote:
> Here are other versions of Wget i found; which do you recommend?
>
> Windows binaries of GNU Wget - eternallybored.org
> Windows binaries of GNU Wget A command-line utility for retrieving files
> using HTTP, HTTPS and FTP protocols. Warning: some antivirus tools
> recognise wget-1.18-win32 ...
> [Search domain eternallybored.org] eternallybored.org/misc/wget/
* Well, this may be the most recent version, but even in Win7 SP1, it 
crashed--refuses to extract/store Wget.exe, making it useless. Error 
0x80004005 unspecified error.

>
> WGET - Download
> WGET, free and safe download. WGET latest version:.
> [Search domain wget.en.softonic.com] wget.en.softonic.com
> ** This gives ONLY Wget.EXE
>
> Wget - Free downloads and reviews - CNET Download.com
> wget free download - SimpleWget, WinWGet, Nugget, and many more programs
> [Search domain download.cnet.com] download.cnet.com/s/wget/
* i was not watching closely; think i got SourceForge version 1.11; 
truncates and clobbers URL.

>
> VisualWget Download - Softpedia
> VisualWget was designed as a GUI front-end for Wget, a content retriever
> originally designed for GNU that has been on the market longer than we
> can remember.
> [Search domain www.softpedia.com]
> softpedia.com/get/Internet/Download-Managers/VisualWg
>
> Download WGET free - latest version
> Download WGET now from Softonic: 100% safe and virus free. More than 327
> downloads this month. Download WGET latest version for free
> [Search domain wget.en.softonic.com] wget.en.softonic.com/download
> **
> I tried the eternallybored.org version, same results.
>
> Thanks.

   Now that i have my list saved and accessible via any OS, i can retry.

[toc] | [prev] | [next] | [standalone]


#109289 — Re: Read (and parse) file on the web CORRECTION#2

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-06 23:49 -0800
SubjectRe: Read (and parse) file on the web CORRECTION#2
Message-ID<_LOzz.58425$eM3.5723@fx04.iad>
In reply to#109286
Auric__ wrote:
> Robert Baer wrote:
>
>> PS: i found WGET to be non-useful (a) it truncates the filename (b) it
>> buggers it to partial gibberish.
>
> Then you must be using a bad version, or perhaps have something wrong with
> your .wgetrc. I've been using wget for around 10 years, and never had
> anything like those issues unless I pass bad options.
>
   I also tried versions 1.18 and 1.13 from 
https://eternallybored.org/misc/wget/.
   Exactly the same truncation and gibberish.
   At least, the 1.13 ZIP had wgetrc in the /etc folder; perhaps one 
step forward.
   No nobody said where to put the folder set, and certainly nothing 
about set path, which just maybe perhaps might be useful for operation.

[toc] | [prev] | [next] | [standalone]


#109225

FromGS <gs@v.invalid>
Date2016-08-28 15:51 -0400
Message-ID<npvfcq$52m$1@dont-email.me>
In reply to#109221
>    Well, i am in a pickle.
>    Firstly, i did a bit of experimenting, and i discovers a few 
> things.
> 1) The variable "tmp" is nice, but does not have to have the date and 
> time; that would fill the HD since i have thousands of files to 
> process.

So you did not 'catch' that the tmp file is deleted when this line...

  Kill tmp

..gets executed!

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109229

FromRobert Baer <robertbaer@localnet.com>
Date2016-08-28 21:37 -0800
Message-ID<RSPwz.1811$GX.452@fx13.iad>
In reply to#109225
GS wrote:
>> Well, i am in a pickle.
>> Firstly, i did a bit of experimenting, and i discovers a few things.
>> 1) The variable "tmp" is nice, but does not have to have the date and
>> time; that would fill the HD since i have thousands of files to process.
>
> So you did not 'catch' that the tmp file is deleted when this line...
>
> Kill tmp
>
> ..gets executed!
>
   I know about that; i have been stepping thru the execution (F8), and 
stop long before that.
   Then i go to CMD prompt and check files and DEL *.* when appropriate 
for next test.

[toc] | [prev] | [next] | [standalone]


#109276

FromRobert Baer <robertbaer@localnet.com>
Date2016-09-02 19:09 -0800
Message-ID<Rhqyz.45747$1b1.33436@fx43.iad>
In reply to#109225
GS wrote:
>> Well, i am in a pickle.
>> Firstly, i did a bit of experimenting, and i discovers a few things.
>> 1) The variable "tmp" is nice, but does not have to have the date and
>> time; that would fill the HD since i have thousands of files to process.
>
> So you did not 'catch' that the tmp file is deleted when this line...
>
> Kill tmp
>
> ..gets executed!
>
   Been thru this already..also makes no sense to have a name miles long.

[toc] | [prev] | [next] | [standalone]


#109230

FromGS <gs@v.invalid>
Date2016-08-29 01:56 -0400
Message-ID<nq0iqm$65d$1@dont-email.me>
In reply to#109221
>  but does not have to have the date and time; that would fill the HD 
> since i have thousands of files to process.

Filenames have nothing to do with storage space; -it's the file size! 
Given Auric_'s suggestion creates text files, the size of 999 txt files 
would hardly be more the 1MB total! If you append each page to the 1st 
file then all pages could be in 1 file...

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109231

FromGS <gs@v.invalid>
Date2016-08-29 02:21 -0400
Message-ID<nq0k9q$9rr$1@dont-email.me>
In reply to#109230
After looking at your link to p5, I see what you mean by the amount of 
storage space, but filename is not a factor. Parsing will certainly 
downsize each file considerably...

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109240

FromRobert Baer <robertbaer@localnet.com>
Date2016-08-29 19:55 -0800
Message-ID<1B6xz.14323$j44.12823@fx24.iad>
In reply to#109231
GS wrote:
> After looking at your link to p5, I see what you mean by the amount of
> storage space, but filename is not a factor. Parsing will certainly
> downsize each file considerably...
>
   One has to open the file to parse, so that is not logical.
   The function URLDownloadToFile gives zero options - it copies ALL of 
the source into a TEMP file; one hopes that the source is not equal to 
or larger than 4GB in size!

   For pages on the web, that is extremely unlikely; webpage size max 
limit prolly is 10MB; maybe 300K worst case on the average.

   So, once in TEMP, it can be opened for input (text), for random (may 
specify buffer size), or for binary (may specify buffer size).
   Here,one can optimize read speed VS string space used.

[toc] | [prev] | [next] | [standalone]


#109242

FromGS <gs@v.invalid>
Date2016-08-29 23:35 -0400
Message-ID<nq2utg$431$1@dont-email.me>
In reply to#109240
> GS wrote:
>> After looking at your link to p5, I see what you mean by the amount 
>> of
>> storage space, but filename is not a factor. Parsing will certainly
>> downsize each file considerably...
>>
>    One has to open the file to parse, so that is not logical.
>    The function URLDownloadToFile gives zero options - it copies ALL 
> of the source into a TEMP file; one hopes that the source is not 
> equal to or larger than 4GB in size!
>
>    For pages on the web, that is extremely unlikely; webpage size max 
> limit prolly is 10MB; maybe 300K worst case on the average.
>
>    So, once in TEMP, it can be opened for input (text), for random 
> (may specify buffer size), or for binary (may specify buffer size).
>    Here,one can optimize read speed VS string space used.

Your last statement contradicts your first statement. Sounds like you 
need to do some *extensive* research into standard VB file I/O 
procedures and general parsing techniques!

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109245

FromRobert Baer <robertbaer@localnet.com>
Date2016-08-29 22:25 -0800
Message-ID<_M8xz.11233$eM3.6466@fx04.iad>
In reply to#109242
GS wrote:
>> GS wrote:
>>> After looking at your link to p5, I see what you mean by the amount of
>>> storage space, but filename is not a factor. Parsing will certainly
>>> downsize each file considerably...
>>>
>> One has to open the file to parse, so that is not logical.
>> The function URLDownloadToFile gives zero options - it copies ALL of
>> the source into a TEMP file; one hopes that the source is not equal to
>> or larger than 4GB in size!
>>
>> For pages on the web, that is extremely unlikely; webpage size max
>> limit prolly is 10MB; maybe 300K worst case on the average.
>>
>> So, once in TEMP, it can be opened for input (text), for random (may
>> specify buffer size), or for binary (may specify buffer size).
>> Here,one can optimize read speed VS string space used.
>
> Your last statement contradicts your first statement. Sounds like you
> need to do some *extensive* research into standard VB file I/O
> procedures and general parsing techniques!
>
   Perhaps you have some things confused.
   In that program, "tmp" is a string used for the name of a (hopefully) 
to-be created file.

   THAT file can be large,as it MUST "hold" the contents of the URL 
being transferred.

   Like i said, the function URLDownloadToFile gives zero options - it 
copies ALL of the source into a TEMP file (named via "tmp").

   Then and only then one can do I/O. I am an expert on on file I/O and 
parsing; read backwards, read (and parse) TEXT files using the file in 
binary or random mode; "inserting" and/or "clipping" stuff into/out of 
the middle of a file, etc.  When pressed, i could even write a file 
backwards, but i have yet to see any reason to try that.


[toc] | [prev] | [next] | [standalone]


#109238

FromRobert Baer <robertbaer@localnet.com>
Date2016-08-29 19:21 -0800
Message-ID<C46xz.16588$oU2.7722@fx09.iad>
In reply to#109230
GS wrote:
>> but does not have to have the date and time; that would fill the HD
>> since i have thousands of files to process.
>
> Filenames have nothing to do with storage space; -it's the file size!
> Given Auric_'s suggestion creates text files, the size of 999 txt files
> would hardly be more the 1MB total! If you append each page to the 1st
> file then all pages could be in 1 file...
>
   True, BUT the files can be large:
"result = URLDownloadToFile(0, S$, tmp, 0, 0)" creates a file in TEMP 
the size of the source - which can be multi-megabtes; 999 of them can 
eat the HD space fast.
   Hopefully a URL file size does not exceed the space limit allowed in 
Excel 2003 string space (anyone know what that might be?).

   I have found that the stringvalue AKA TEMP filename can be fixed to 
anything reasonable, and does not have to include 
parts/substrings/subsets of the file one wants to download.

   I can be "a good thing" (quoting Martha Stewart) to delete the file 
when done.

   I have also found the following:
1) one does not have to use FreeFile for a file number (when all else is 
OK).
2) cannot use "contents" for string storage space.
3) one cannot mix use of "/" and "\" in a string for a given file name.
4) one cannot have a space in the file name, so that gives a serious 
problem for some web URLs (work-around anyone?)
5) method fails for "https:" (work-around anyone?)

[toc] | [prev] | [next] | [standalone]


#109241

FromGS <gs@v.invalid>
Date2016-08-29 23:21 -0400
Message-ID<nq2u36$224$1@dont-email.me>
In reply to#109238
>    I have also found the following:
> 1) one does not have to use FreeFile for a file number (when all else 
> is OK).

True, however not considered 'best practice'. Freefile() ensures a 
unique ID is assigned to your var.

> 2) cannot use "contents" for string storage space.
Why not? It's not a VB[A] keyword and so qualifies for use as a var.


> 3) one cannot mix use of "/" and "\" in a string for a given file 
> name.

Not sure why you'd use "/" in a path string! Forward slash is not a 
legal filename/path character. Backslash is the default Windows path 
delimiter. If choosing folders, the last backslah is not followed by a 
filename.

> 4) one cannot have a space in the file name, so that gives a serious 
> problem for some web URLs (work-around anyone?)

Web paths 'pad' spaces so the string is contiguous. I believe the pad 
string is "%20" OR "+".

> 5) method fails for "https:" (work-around anyone?)

What method?

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109244

FromRobert Baer <robertbaer@localnet.com>
Date2016-08-29 22:10 -0800
Message-ID<Oy8xz.31343$6d.20556@fx26.iad>
In reply to#109241
GS wrote:
>> I have also found the following:
>> 1) one does not have to use FreeFile for a file number (when all else
>> is OK).
>
> True, however not considered 'best practice'. Freefile() ensures a
> unique ID is assigned to your var.
>
>> 2) cannot use "contents" for string storage space.
> Why not? It's not a VB[A] keyword and so qualifies for use as a var.
* Get run-time error 458, "variable uses an Automation type not 
supported in Visual Basic".

>
>
>> 3) one cannot mix use of "/" and "\" in a string for a given file name.
>
> Not sure why you'd use "/" in a path string! Forward slash is not a
> legal filename/path character. Backslash is the default Windows path
> delimiter. If choosing folders, the last backslah is not followed by a
> filename.
   I note that in Windoze, that "\" is used,and on the web, "/" is used.

>
>> 4) one cannot have a space in the file name, so that gives a serious
>> problem for some web URLs (work-around anyone?)
>
> Web paths 'pad' spaces so the string is contiguous. I believe the pad
> string is "%20" OR "+".
   Yes; "%20" is used and seems to act like a space and seems to kill 
the URLDownloadToFile function usefulness.

>
>> 5) method fails for "https:" (work-around anyone?)
>
> What method?
* the function URLDownloadToFile. Is "method" the wrong term?


>

[toc] | [prev] | [next] | [standalone]


#109248

FromGS <gs@v.invalid>
Date2016-08-30 04:04 -0400
Message-ID<nq3els$as9$1@dont-email.me>
In reply to#109244
>  I note that in Windoze, that "\" is used,and on the web, "/" is 
> used.

For clarity.., a Windows file path is not the same as a URL.

  Windows file paths allow spaces; URLs do not.
  Windows path delimiter is "\"; Web path delimiter is "/".

The function URLDownloadToFile() downloads szURL to szFilename.

Once downloaded, szFilename needs to be opened, parsed, and result 
stored locally. Then the next file needs to be downloaded, parsed, and 
stored. And so on until all files have been downloaded and parsed.

Since the actual page contents comprise only a small portion of the 
files being downloaded, there size should be considerably smaller after 
parsing. If you extract the data (text) only (no images) and save this 
to a txt file you should be able to 'append' to a single file which 
would result in occupying far less disc space. (For example, pg5 is 
less than 1kb) I suspect, though, that you need the image to identify 
the item source (Raytheon, RCA, Lucent Tech, NAWC, MIL STD, etc) 
because this info is not stored in the image file metadata. Otherwise, 
the txt file after parsing pg5's text is the following 53 lines:

NSN 5960-00-509-3171
    5960-00-509-3171

    ELECTRON TUBE

NSN 5960-00-569-9531
    5960-00-569-9531

    ELECTRON TUBE

NSN 5960-00-553-3770
    5960-00-553-3770

    ELECTRON TUBE

NSN 5960-00-682-8624
    5960-00-682-8624

    ELECTRON TUBE

NSN 5960-00-808-6928
    5960-00-808-6928

    ELECTRON TUBE

NSN 5960-00-766-1953
    5960-00-766-1953

    ELECTRON TUBE

NSN 5960-00-850-6169
    5960-00-850-6169

    ELECTRON TUBE

NSN 5960-00-679-8153
    5960-00-679-8153

    ELECTRON TUBE

NSN 5960-00-134-6884
    5960-00-134-6884

    ELECTRON TUBE

NSN 5960-00-061-8610
    5960-00-061-8610

    ELECTRON TUBE

    5960-00-067-9636

    ELECTRON TUBE

The file size is 711 bytes, and lists 11 items. Note the last item has 
no image and so no filler text (NSN line). This inconsistency makes 
parsing the contents difficult since you don't know which items do not 
have images.

If you copy/paste pg5 into Excel you get both text and image. You could 
then do something to construct the info in a database fashion...

  Col Headers:
  Source :: PartNum :: Description

..and put the data in the respective columns. This seems very 
inefficient but is probably less daunting than what you've been doing 
manually thus far. Auto Complete should be helpful with this, and you 
could sort the list by Source. Note that clicking the image or part# on 
the worksheet takes you to the same page as does clicking it on the web 
page. In the case of pg5, the data will occupy 11 rows.

Seems like your approach is the long way; -I'd find a better data 
source myself! Perhaps subscribe to an electronics database utility 
(such as my CAD software would use) that I can update by downloading a 
single db file<g>

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109251

FromRobert Baer <robertbaer@localnet.com>
Date2016-08-30 13:48 -0800
Message-ID<Yimxz.32116$892.8406@fx34.iad>
In reply to#109248
GS wrote:
>> I note that in Windoze, that "\" is used,and on the web, "/" is used.
>
> For clarity.., a Windows file path is not the same as a URL.
>
> Windows file paths allow spaces; URLs do not.
* Incorrect! See -------------------------vvv
https://www.nsncenter.com/NSNSearch?q=5960%20regulator&PageNumber=5

> Windows path delimiter is "\"; Web path delimiter is "/".
>
> The function URLDownloadToFile() downloads szURL to szFilename.
* IF and when it works.

>
> Once downloaded, szFilename needs to be opened, parsed, and result
> stored locally. Then the next file needs to be downloaded, parsed, and
> stored. And so on until all files have been downloaded and parsed.
* This i knew from the git-go; nice to be clarified.

>
> Since the actual page contents comprise only a small portion of the
> files being downloaded, there size should be considerably smaller after
> parsing. If you extract the data (text) only (no images) and save this
> to a txt file you should be able to 'append' to a single file which
> would result in occupying far less disc space. (For example, pg5 is less
> than 1kb) I suspect, though, that you need the image to identify the
> item source (Raytheon, RCA, Lucent Tech, NAWC, MIL STD, etc) because
> this info is not stored in the image file metadata. Otherwise, the txt
> file after parsing pg5's text is the following 53 lines:
>
> NSN 5960-00-509-3171
> 5960-00-509-3171
>
> ELECTRON TUBE
>
> NSN 5960-00-569-9531
> 5960-00-569-9531
>
> ELECTRON TUBE
>
> NSN 5960-00-553-3770
> 5960-00-553-3770
>
> ELECTRON TUBE
>
> NSN 5960-00-682-8624
> 5960-00-682-8624
>
> ELECTRON TUBE
>
> NSN 5960-00-808-6928
> 5960-00-808-6928
>
> ELECTRON TUBE
>
> NSN 5960-00-766-1953
> 5960-00-766-1953
>
> ELECTRON TUBE
>
> NSN 5960-00-850-6169
> 5960-00-850-6169
>
> ELECTRON TUBE
>
> NSN 5960-00-679-8153
> 5960-00-679-8153
>
> ELECTRON TUBE
>
> NSN 5960-00-134-6884
> 5960-00-134-6884
>
> ELECTRON TUBE
>
> NSN 5960-00-061-8610
> 5960-00-061-8610
>
> ELECTRON TUBE
>
> 5960-00-067-9636
>
> ELECTRON TUBE
>
> The file size is 711 bytes, and lists 11 items. Note the last item has
> no image and so no filler text (NSN line). This inconsistency makes
> parsing the contents difficult since you don't know which items do not
> have images.
*  I think you may have pulled the info from what you saw on that page, 
and not from the source.
   In one of my responses, i gave QBASIC code for parsing, and as i 
remember, there were about 7760 lines of junk before one sees <a 
href="/NSN/5960; which gives the full NSN code.
   Use of that allows one to get the second URL, eg: 
https://www.nsncenter.com/NSN/5960-00-754-5782 NO image reliance at all.
   There are 11 entries per page,and no inconsistencies with my method 
of search in the page.

>
> If you copy/paste pg5 into Excel you get both text and image. You could
> then do something to construct the info in a database fashion...
* That would only make things more difficult. A copy to a local file is 
sufficient for a simple parsing as described here and elsewhere in this 
thread.

>
> Col Headers:
> Source :: PartNum :: Description
>
> ..and put the data in the respective columns. This seems very
> inefficient but is probably less daunting than what you've been doing
> manually thus far. Auto Complete should be helpful with this, and you
> could sort the list by Source. Note that clicking the image or part# on
> the worksheet takes you to the same page as does clicking it on the web
> page. In the case of pg5, the data will occupy 11 rows.
* Manual: Right click, select View Page Source, Save as to HD by 
changing Filetype from HTM to TXT and changing fiiename to add page 
number (013 for example).
   Like i said,parsing of that file is simple and easy; getting 35 pages 
copied that way did not take long, but there are 999 of them...

>
> Seems like your approach is the long way; -I'd find a better data source
> myself! Perhaps subscribe to an electronics database utility (such as my
> CAD software would use) that I can update by downloading a single db
> file<g>
>
* I have asked, and received zero response.

[toc] | [prev] | [next] | [standalone]


#109243

FromGS <gs@v.invalid>
Date2016-08-29 23:37 -0400
Message-ID<nq2v13$4a6$1@dont-email.me>
In reply to#109238
> 5) method fails for "https:" (work-around anyone?)

These URLs usually require some kind of 'login' be done, which needs to 
be included in the URL string.

-- 
Garry

Free usenet access at http://www.eternal-september.org
Classic VB Users Regroup!
  comp.lang.basic.visual.misc
  microsoft.public.vb.general.discussion

---
This email has been checked for viruses by Avast antivirus software.
https://www.avast.com/antivirus

[toc] | [prev] | [next] | [standalone]


#109246

FromRobert Baer <robertbaer@localnet.com>
Date2016-08-29 22:25 -0800
Message-ID<FN8xz.11234$eM3.7146@fx04.iad>
In reply to#109243
GS wrote:
>> 5) method fails for "https:" (work-around anyone?)
>
> These URLs usually require some kind of 'login' be done, which needs to
> be included in the URL string.
>
   NO login required; try it.

[toc] | [prev] | [next] | [standalone]


#109222

Fromlovexes17816 <lovexes17816@gmail.com>
Date2016-08-28 14:00 +0100
Message-ID<lovexes17816.1208a148@excelbanter.com>
In reply to#109215
Phim SEX không che , Phim SEX Nháº*t Bản , Phim SEX loạn luân ,
Phim SEX HD

>>>> 'Phim SEX HD' (http://sexchonloc.com/)  <<<<  Tuyển táº*p phim
sex chất lượng cao mới nhất, những bộ phim heo hay nhất
2016, -*xem phim sex*- HD full 1080p trên web không quảng cáo,
chúc các bạn xem ...




-- 
lovexes17816

[toc] | [prev] | [next] | [standalone]


Page 3 of 5 — ← Prev page 1 2 [3] 4 5  Next page →

Back to top | Article view | microsoft.public.excel.programming


csiph-web