Path: csiph.com!usenet.pasdenom.info!news.albasani.net!news2.arglkargh.de!noris.net!newsfeed.arcor.de!newsspool1.arcor-online.net!news.arcor.de.POSTED!not-for-mail Content-Type: text/plain; charset="UTF-8" Message-ID: <2030773.c2UJNlqvMH@PointedEars.de> From: Thomas 'PointedEars' Lahn Reply-To: Thomas 'PointedEars' Lahn Organization: PointedEars Software (PES) Date: Fri, 05 Oct 2012 01:25:55 +0200 User-Agent: KNode/4.4.11 Content-Transfer-Encoding: 8Bit X-Face: %i>XG-yXR'\"2P/C_aO%~;2o~?g0pPKmbOw^=NT`tprDEf++D.m7"}HW6.#=U:?2GGctkL,f89@H46O$ASoW&?s}.k+&. <2931754.sreIk3pbUD@PointedEars.de> Followup-To: comp.lang.javascript MIME-Version: 1.0 Lines: 36 NNTP-Posting-Date: 05 Oct 2012 01:25:55 CEST NNTP-Posting-Host: 398c025b.newsspool1.arcor-online.net X-Trace: DXC=[Ee`3FlXc2n_0Po7BmQ3]lic==]BZ:afn4Fo<]lROoRankgeX?EC@@`1CEWAg0_HNeDZm8W4\YJNlR=i2=[N6Y6jI8d2_XC2HW`LGMZoKNLEZo X-Complaints-To: usenet-abuse@arcor.de Xref: csiph.com comp.lang.javascript:16407 Thomas 'PointedEars' Lahn wrote: > Andrew Poulos wrote: >> If the server can return "arbitrary" but valid HTML such as >> […] >> how can you possibly parse it and use createElement/appendChild to >> recreate the HTML. Currently I'm using innerHTML. > > You can parse this efficiently using a regular expression containing token > patterns matching grammar atoms in alternation and > RegExp.prototype.exec(), then apply createElement/appendChild on the > result. CAVEAT: The RegExp must have the `global' flag set so that exec() looks for all matches and control can exit the loop in which it is used when exec() returns `null'. Otherwise the loop will be endless if there is a first match, and your script engine and browser (tab) will hang. Also, if there are two atoms with the same prefix, the first-match-wins rule inherent to RegExp's alternation (as opposed to the longest-match-wins rule usually found in parsers) becomes a problem. I have found two ways to work around that: If the relevant subpatterns match strings of fixed length, you can put the subpattern matching the longest string first. Otherwise you will have to loop through the patterns and match each from the same current offset in the string; the longest match wins, and if there are two or more matches of same length, the first match should win (so you should arrange your subpatterns accordingly or implement precedence otherwise). But this is probably not a concern for parsing HTML. PointedEars -- Use any version of Microsoft Frontpage to create your site. (This won't prevent people from viewing your source, but no one will want to steal it.) -- from (404-comp.)