Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.programming > #16325 > unrolled thread

Re: Scanning

Started byRichard Heathfield <rjh@cpax.org.uk>
First post2023-01-19 15:06 +0000
Last post2023-01-21 11:29 +1300
Articles 2 — 2 participants

Back to article view | Back to comp.programming

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: Scanning Richard Heathfield <rjh@cpax.org.uk> - 2023-01-19 15:06 +0000
    Re: Scanning Noel Duffy <uaigh@icloud.com> - 2023-01-21 11:29 +1300

#16325 — Re: Scanning

FromRichard Heathfield <rjh@cpax.org.uk>
Date2023-01-19 15:06 +0000
SubjectRe: Scanning
Message-ID<tqbm8u$1hm7b$1@dont-email.me>
On 19/01/2023 2:48 pm, Stefan Ram wrote:
> ram@zedat.fu-berlin.de (Stefan Ram) writes:
>> Let's take a very simple task: This scanner for text files
>> has nothing more to do than to return every character,
>> except to strip the spaces at the end of a line.
> 
>    Richard said that it matters what I need this for.
> 
>    I'd like to implement a tiny markup language

Okay, BIG job with lots of complicated, so strive to keep each 
part relatively simple if you ever hope to get it working. Do it 
in whatever way comes most natural to your programming style, 
because that's how /you/ can define 'simple'. You're using 
Python, so I guess you're not overly concerned by performance, so 
do it the way you personally find easiest. I'm guessing you'll go 
for line by line and lean on Python's memory management.

But write this down somewhere: if, further down the line, your 
parser turns out to be too slow and the profiler blames this bit, 
rewriting it to go byte by byte might well be one of the ways you 
could speed it up.

-- 
Richard Heathfield
Email: rjh at cpax dot org dot uk
"Usenet is a strange place" - dmr 29 July 1999
Sig line 4 vacant - apply within

[toc] | [next] | [standalone]


#16334

FromNoel Duffy <uaigh@icloud.com>
Date2023-01-21 11:29 +1300
Message-ID<tqf4jp$28cng$1@dont-email.me>
In reply to#16325
On 21/01/23 01:16, Stefan Ram wrote:
> ram@zedat.fu-berlin.de (Stefan Ram) writes:
[..]
> 
>    The output often ends with one space, because a '\n' is
>    added to the end of the input if it's missing, and this
>    then is being converted to a space. So, ironically, while
>    I set out to strip spaces at the end of lines, I now
>    sometimes add them to the end of lines!

While I don't have any great insight to offer, I did write a small 
markup engine a few years ago (in Object Pascal). What you say above 
brought back memories of struggles I had with my code too. The 
conclusion I came to at the time is that when it comes to things like 
spacing, there are several equally valid ways to do it, and you'll 
probably want different handling for different use-cases, so it's better 
to parameterize it so that users of your code can set which handling 
they prefer. I went with making it a parameter. It's a bit more work but 
the flexibility is usually worth it in the long run.

[toc] | [prev] | [standalone]


Back to top | Article view | comp.programming


csiph-web