Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.python > #198056

Re: Formatting

From ram@zedat.fu-berlin.de (Stefan Ram)
Newsgroups comp.lang.python
Subject Re: Formatting
Date 2026-09-23 12:42 +0000
Organization Stefan Ram
Message-ID <breaking-20260923133423@ram.dialup.fu-berlin.de> (permalink)
References <nhcb19FjpohU1@mid.individual.net> <nhg4lgF7nmhU1@mid.individual.net> <nnd$00ad2c5a$613e2476@d140bd1880d0a0a7> <formatting-20260923091605@ram.dialup.fu-berlin.de> <nnd$4ca00753$14137697@ed7836aeb5eb4d5a>

Show all headers | View raw


Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> wrote or quoted:
>On 9/23/2026 4:17 PM, Stefan Ram wrote:
>>I also have written a Python module that also does hyphenation, and
>>which I plan to publish one day, but so far the code has not yet
>>been cleaned up. This hyphenation uses TeX's hyphenation patterns.
>You mean you coded up TeX's line breaking algorithm?

  I did not write this above. I only wrote that I used TeX's
  hyphenation patterns. But these patterns can be used with
  any line-breaking algorithm.

  But in fact, I /did/ implement something very close to TeX's
  line breaking algorithm.

>                                                      Did you do it from
>the book /TeX: The Program/, or /Digital Typesetting/?

  This was in 2024, and I do non remember all details anymore.
  But, frankly, "TeX: The Program" in the end was too
  complicated for me to read, so I think I used more info from
  "Digital Typesetting" and used "TeX: The Program" only for some
  clarifications. I only implemented it for monospaced fonts, and
  omitted some features of TeX adding some features of my own.

>I have alas not coded up TeX's hyphenation patterns, and only thought to

  The hyphenation pattern files are out there for many languages,
  and I just needed code to read and interpret them. This then
  inserts possible hyphenation points into words before they are
  passed to the line breaker.

>Now to get perfect line breaks in fixed width text the best method is to
>construct the text carefully, and change the words with a thesaurus, and
>the word order, and possibly write a completely new sentence.  When I do
>this, I usually do this with hyphenation also, and do not care about the
>official rules, and just wing it.  The only paragraph I did this in this
>followup, is this one.

  Yeah, that's called "bricktext".

  Now, here are my two posts from 2024 with an early version of
  my line breaker.

Newsgroups: comp.text.tex
Subject: TeX's line breaking in the grub sesh
From: ram@zedat.fu-berlin.de (Stefan Ram)
Message-ID: <Lines-20240805135439@ram.dialup.fu-berlin.de>

  During my grub sesh today, I crushed out TeX's line breaking
  in Python in like a hot minute - 36 to be exact. Tryna use
  it for plain text, like, monospaced fonts and whatnot. Natch,
  I stripped it down to the bare bones, but it's still hella
  tight. Already got that "parshape" action goin' on (which
  Knuth-Plass can't hang with, if I'm not trippin'). Next up,
  I'm finna tackle those "discretionary items" - that's gonna be
  gnarly! For sure there's some janky bugs in there, but peep this:

  wrap.py 

source = 'Ich habe das gebackene Profi-Bettuch bereits gesehen. '

active0 =[ 0, 0, 0, 0 ] # previous, position, quality, line_number
active =[ active0 ]
parshape =[ 10, 20, 20, 20, 20, 20, 20, 20 ]

p = 1
while p < len( source ):
    ch = source[ p ]
    if ch == ' ':
        new_active = []
        best_quality = -10000
        best_act = active[ 0 ]
        for act in active:
            a = act[ 1 ]
            line_number = act[ 3 ] 
            target_length = parshape[ line_number ]
            this_length = p - a
            if this_length > target_length:
                pass
            else:
                quality = this_length - target_length
                new_active.append( act )
                if quality > best_quality:
                    best_quality = quality
                    best_act = act
            new_active.append( [ best_act, p, best_quality, best_act[ 3 ]+1 ] )
        active = new_active
    p += 1
    act = active[ -1 ]
buff = []
while act[ 0 ]:
    prev = act[ 0 ]
    buff.append( source[ prev[1]: act[1] ])
    act = prev
for line in reversed( buff ):
    print(line)

  output

Ich habe das
 gebackene Profi-Bettuch
 bereits gesehen.

Newsgroups: comp.text.tex
Subject: Re: TeX's line breaking in the grub sesh
References: <Lines-20240805135439@ram.dialup.fu-berlin.de>
From: ram@zedat.fu-berlin.de (Stefan Ram)
Message-ID: <wrap-20240828192402@ram.dialup.fu-berlin.de>

  Turns out there was still a glitch in the code! The latest
  build now shows a paragraph break with (fingers crossed)
  global optimization, taking parshape into account. Now that
  I've finally squashed the bug, discretionary items haven't
  been baked in yet. That's next on my to-do list though.

  Ironically, lines of the following Python 3.12 source code have NOT
  been wrapped to the 72 characters recommended for Usenet posts!

  main.py

from dataclasses import dataclass
from typing import Optional, List, Iterator
import bisect

source_text = list( ' Lorem ipsum dolor sit amet, consectetur adipiscing elit, '
'sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad '
'minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea '
'commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit '
'esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat '
'non proident, sunt in culpa qui officia deserunt mollit anim id est laborum. ' )
                    
parshape =[ 80, 80, 80, 40 ]

@dataclass
class ActiveEntry:
    previous:    Optional[ 'ActiveEntry' ]= None
    position:    int = 0 # position in the text 
    branch:      int = 0 # for future version
    sum_quality: int = 0 # the sum of all "merits" up to this point
    line_number: int = 0 # line started with this (first line = 0)

parshape_length = len( parshape )

def print_up_to( this ):
    '''prints the wrapped paragraph up to the point "this".
    It starts at the back point "this" and then goes forward
    in the text via the linked chain of points. Finally, for
    printing the text in the normal order, it then goes forward
    again.'''
    buff = []
    qual = []
    sum_quality = this.sum_quality
    while this.previous:
        previous = this.previous
        line = source_text[ previous.position: this.position ]
        # print( f'{line = }' )
        buff.append( line )
        qual.append( this.sum_quality )
        this = previous
    start_position = 0
    # we went backwards, but actually want to print in the normal direction
    first = 1
    for i,( line, qual ) in enumerate( zip( reversed( buff ), reversed( qual ))):
        text = ''.join( line[ start_position: ])
        target_length = parshape[ i ]if i < parshape_length else parshape[ -1 ]# dupe!
        output = text
        print( output[ first: ])
        first = 0
        start_position = 1 # skip an initial space or something
    print()
    print( 'Total merits:', sum_quality )
    print()

active0 = ActiveEntry()
    
active_list =[ active0 ]

current_position = 1
source_length = len( source_text )
while current_position < len( source_text ):
    ch = source_text[ current_position ]
    if ch == ' ': # possible breakpoint
        new_active_list = [] # next active list
        best_sum_quality = None # best quality summed across this and previous lines, not yet determined
        best_act = active_list[ 0 ] # preliminary choice
        for active in active_list:
            active_position = active.position
            line_number = active.line_number
            target_length = parshape[ line_number ]if line_number < parshape_length else parshape[ -1 ]# dupe!
            distance = current_position - active_position
            adjustment = target_length - distance
            if adjustment < 0:
                #   "When an active breakpoint a is encountered for which
                #      the line from a to b has an adjustment ratio less 
                #      than -1 (that is, when the line can't be shrunk to 
                #      fit the desired length), breakpoint a is removed 
                #      from the active list."
                pass # do not transfer into the new active list
            else:
                new_active_list.append( active )
                this_line_quality = -adjustment**2
                have_reached_final_space = current_position == source_length - 1 # final ' ' on end of last line 
                if have_reached_final_space: this_line_quality = 0 # arbitrary whitespace at end is accepted
                this_sum_quality = active.sum_quality + this_line_quality
                if \
                best_sum_quality is None or \
                this_sum_quality > best_sum_quality:
                    best_sum_quality = this_sum_quality
                    best_predecessor = active
        if best_sum_quality is not None:
            # make a new active point from current position, linking it to the best active point found
            new_active_list.append( ActiveEntry( previous=best_predecessor, position=current_position, branch=0, sum_quality=best_sum_quality, line_number=best_predecessor.line_number+1 ))
        active_list = new_active_list
    current_position += 1
active = active_list[ -1 ] # the final space

print_up_to( active )

  output

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor
incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis
nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat.
Duis aute irure dolor in reprehenderit
in voluptate velit esse cillum dolore
eu fugiat nulla pariatur. Excepteur
sint occaecat cupidatat non proident,
sunt in culpa qui officia deserunt
mollit anim id est laborum.

Total merits: -80

Back to comp.lang.python | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread


Thread

Best practise for managing Conda environments? Martin Schöön <martin.schoon@gmail.com> - 2026-09-21 09:17 +0000
  Re: Best practise for managing Conda environments? Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-09-21 17:32 +0800
    Re: Best practise for managing Conda environments? Martin Schöön <martin.schoon@gmail.com> - 2026-09-21 20:33 +0000
      Re: Best practise for managing Conda environments? Jon Ribbens <jon+usenet@unequivocal.eu> - 2026-09-21 22:05 +0000
        Re: Best practise for managing Conda environments? Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-09-22 14:37 +0800
          Re: Best practise for managing Conda environments? ram@zedat.fu-berlin.de (Stefan Ram) - 2026-09-22 07:33 +0000
          Re: Best practise for managing Conda environments? Jon Ribbens <jon+usenet@unequivocal.eu> - 2026-09-22 09:08 +0000
            Re: Best practise for managing Conda environments? Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-09-22 22:53 +0800
  Re: Best practise for managing Conda environments? Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-09-22 15:08 +0800
  Re: Best practise for managing Conda environments? Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-09-22 07:44 +0000
    Re: Best practise for managing Conda environments? "Loris Bennett" <loris.bennett@fu-berlin.de> - 2026-09-22 10:14 +0200
  Re: Best practise for managing Conda environments? Martin Schöön <martin.schoon@gmail.com> - 2026-09-22 19:53 +0000
    Re: Best practise for managing Conda environments? Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-09-23 04:57 +0800
      Formatting (was: Best practise for managing Conda environments?) ram@zedat.fu-berlin.de (Stefan Ram) - 2026-09-23 08:17 +0000
        Re: Formatting Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-09-23 20:08 +0800
          Re: Formatting ram@zedat.fu-berlin.de (Stefan Ram) - 2026-09-23 12:42 +0000
            Re: Formatting ram@zedat.fu-berlin.de (Stefan Ram) - 2026-09-23 14:43 +0000
              Re: Formatting ram@zedat.fu-berlin.de (Stefan Ram) - 2026-09-23 15:01 +0000
          Re: Formatting ram@zedat.fu-berlin.de (Stefan Ram) - 2026-09-23 13:38 +0000
        Re: Formatting (was: Best practise for managing Conda environments?) Martin Schöön <martin.schoon@gmail.com> - 2026-09-24 08:31 +0000
          Re: Formatting (was: Best practise for managing Conda environments?) Martin Schöön <martin.schoon@gmail.com> - 2026-09-25 08:53 +0000

csiph-web