Path: csiph.com!eternal-september.org!feeder.eternal-september.org!nntp.eternal-september.org!.POSTED!not-for-mail From: Kragen Javier Sitaker Newsgroups: alt.sys.pdp10,comp.text.pdf,comp.compression Subject: PDF file compactness (was Re: reprints of old AI memos) Followup-To: comp.text.pdf Date: Thu, 03 Sep 2026 12:06:18 -0300 Organization: Primarily biological and memetic Lines: 215 Message-ID: <87h5k61ij9.fsf_-_@debian> References: <874inbqdz7.fsf@mariorosell.es> <110nf0q$3vra6$3@dont-email.me> <87y0gg62cx.fsf@nightsong.com> <110rfqf$142ot$1@dont-email.me> <111ne5o$vsm3$3@dont-email.me> <111pbk6$36pi7$1@dont-email.me> <87qzlqgiyx.fsf@gmail.com> <111t96n$6sqq$1@dont-email.me> <111tev9$8hea$1@dont-email.me> <875x2zzres.fsf@nightsong.com> <86se5zr6ez.fsf@williamsburg.bawden.org> <87mru47iou.fsf_-_@debian> <7wecfgrtuh.fsf@junk.nocrew.org> <874igb7fmv.fsf_-_@debian> <7wwlt6rhw6.fsf@junk.nocrew.org> <874iga5jgr.fsf@debian> <7wld9l5xtc.fsf_-_@junk.nocrew.org> <87bjaf5tvl.fsf@debian> <117a5o2$35q9q$7@dont-email.me> MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 8bit Injection-Date: Thu, 03 Sep 2026 15:08:37 +0000 (UTC) Injection-Info: dont-email.me; logging-data="4000206"; mail-complaints-to="abuse@eternal-september.org"; posting-account="U2FsdGVkX1+3VNTUCoP0IyiLJGF+8nzT"; posting-host="9d8b1d041ffbfb92fa0be79acc76e4c5" User-Agent: Gnus/5.13 (Gnus v5.13) Emacs/28.2 (gnu/linux) Cancel-Lock: sha1:4XE3TF3Xc9jKl/3Eli6fJ51+dv0= sha1:XAmpcl1GQqgcLygEw3FXsfILdAM= sha256:nTYlv+z/4PhfyUP4TVnbBou2tLGAKQdnDUoVAxzf4QY= sha1:mtqlioC0RT8v0JMdPwBgaaVmw+4= sha256:aONzqDgNakB18jXIcII1W2QJ4IRwsXhxOmDalwO+/6o= Xref: csiph.com alt.sys.pdp10:9994 comp.text.pdf:2736 comp.compression:16290 Lawrence D’Oliveiro writes: > On Wed, 02 Sep 2026 16:35:26 -0300, Kragen Javier Sitaker wrote: >> Contrary to popular belief, PDF is a relatively compact file format. > > Only if you apply compression to it. This is ambiguous; you might be intending to say, “only if you use the PDF format’s compression features” or, “only if you compress the PDF file with an additional compressor after creating it.” I disagree with both of these interpretations. For text, PDF can be a relatively compact file format even if you do neither of these. There’s some file-format overhead of about a kilobyte — you need a header, a catalog, a page tree, an xrefs table, and trailer, even for a one-page document — and then you need to specify the font and the coordinates where your text starts. After that, I think the per-line overhead is about 4 bytes per line. I haven’t tested this content-stream example, but if it’s missing something, it’s not missing much: BT /F0 12 Tf 50 706 Td (For text, PDF is a relatively compact file format even if) ' (you do neither of these. There's file format overhead of) ' (about a kilobyte - you need a page tree and xrefs table even) ' (for a one-page document - and then you need to specify the) ' (font and the coordinates where your text starts. After) ' (that, the per-line overhead is a few bytes per line. I) ' (haven't tested this content-stream example, but if it's) ' (missing something, it's not missing much:) ' ET The `'` PDF content-stream operator is just `T* TJ`, if you’re familiar with those. There’s also a `"` shortcut operator that lets you set a different letter spacing and word spacing for each line: 1 1.4 (for a one-page document - and then you need to specify the) " That costs you about 11 bytes per line instead of 4. So, a simply formatted PDF file without using any kind of compression is about 5% bigger than a plain ASCII text file, plus about one kilobyte. Perhaps you don't consider 5% overhead to be “relatively compact”, but I do. Also, of course, PDF *does* support Deflate compression for content streams; and, since PDF 1.5, it also supports it for xrefs and object streams, so a PDF file can easily be half the size of a plain ASCII text file, down to a minimal size of a few hundred bytes. Amusingly, I’ve even seen PDF files that apply Paeth compression to the xrefs table. Now, in practice, PDF files are often not this compact. A much more typical example of a PDF content-stream (from the Derctuo PDF: ) looks like this (slightly reformatted): 1 0 0 1 0 0 cm BT /F1 12 Tf 14.4 TL ET 0 0 0 rg 0 0 0 rg BT 1 0 0 1 6 747.6 Tm .533333 0 0 rg /F2+0 24 Tf 28.8 TL (Derctuo) Tj T* ET 0 0 0 rg BT 1 0 0 1 6 714.84 Tm /F2+0 12 Tf 14.4 TL ( ) Tj T* ET 0 0 0 rg BT 1 0 0 1 6 700.44 Tm /F3+0 12 Tf 14.4 TL ( ) Tj /F4+0 12 Tf 14.4 TL (\200\201\201\200) Tj T* ET 0 0 0 rg BT 1 0 0 1 6 686.04 Tm /F3+0 12 Tf 14.4 TL ( ) Tj (Kragen ) Tj (Javier ) Tj (Sitaker) Tj T* ET 0 0 0 rg BT 1 0 0 1 6 671.64 Tm /F3+0 12 Tf 14.4 TL ( ) Tj (Buenos ) Tj (Aires) Tj T* ET 0 0 0 rg BT 1 0 0 1 6 657.24 Tm /F3+0 12 Tf 14.4 TL ( ) Tj (December, ) Tj (02020) Tj T* ET 0 0 0 rg BT 1 0 0 1 6 642.84 Tm /F3+0 12 Tf 14.4 TL ( ) Tj (Public ) Tj (domain ) Tj (work) Tj T* ET 0 0 0 rg BT 1 0 0 1 6 628.44 Tm /F3+0 12 Tf 14.4 TL ( ) Tj /F4+0 12 Tf 14.4 TL (\200\201\201\200) Tj T* ET 0 0 0 rg 0 0 0 rg BT 1 0 0 1 6 599.64 Tm /F2+0 12 Tf 14.4 TL ( ) Tj T* ET 0 0 0 rg BT 1 0 0 1 6 581.28 Tm /F2+0 12 Tf 14.4 TL ( ) Tj (Derctuo ) Tj (is ) Tj (a ) Tj (book ) Tj (of ) Tj (notes ) Tj (on ) Tj (various ) Tj (topics, ) Tj (mostly ) Tj (science ) Tj (and ) Tj T* ET It’s relatively straightforward to uncompress content-streams like this from PDF files in Python: zlib.decompress(base64.a85decode(a8.removesuffix(b'~>')) ).decode('utf-8') This is obviously inefficient in many different ways: - Reportlab decided to Ascii85Decode the compressed data for no real reason, even though I specified pageCompression=True. - There’s no need to Tj each word separately. You could Tj the whole line of text. I think this was my fault; I hacked together this PDF renderer in a week for a deadline. - It’s unnecessary to set the transformation matrix (cm) to the identity matrix. That’s the default. - Similarly, setting the RGB color to black for each line (and twice for the first line) is unnecessary. Black is the default. I think this is ReportLab’s fault. - In the one case where the color is set to a non-default color, it’s unnecessary to specify that color to six significant figures. - Displaying runs of spaces is generally unnecessary. - Changing fonts to display runs of spaces is extra unnecessary. - Changing fonts twice per line is unnecessary. Most of this text is in a single font. - Chanting to the same font again is also unnecessary. - Creating a new text object for every line (BT ET) is unnecessary and also counterproductive for copy-and-paste. Despite all this, the 47 lines of text on the page are 8202 bytes of uncompressed content-stream; FlateDecode encoded and Ascii85Decode encoded, they pack down to only 2471 bytes, plus 149 bytes of per-stream overhead (also mostly unnecessary), plus 309 bytes of the Page object containing the content stream (also mostly unnecessary), for a total of 3K per page, which is slightly more compact than plain ASCII text. (This doesn’t count the hyperlinks on the page, though.) It’s easy to see how small inefficiencies like these can pile up when people (like me) who don’t really understand what they’re doing get things to more or less work, and then stop. And that seems to be how most PDF files are built. Most of them are even worse than the Derctuo PDF. Derctuo is far from exemplary, but the PDF is 986 pages and 5.91 megabytes, roughly 5.9KiB per page; it divides up as follows: - bytes 569 to 2.23e6: intermixed page objects and link objects - bytes 2.23e6 to 2.50e6: embedded fonts, covering ASCII and a bunch of Unicode for things like math and Greek, in eight display styles - bytes 2.50e6 to 2.62e6: more document structure, including outline and page tree - bytes 2.62e6 to 5.74e6: page content streams - bytes 5.74e6 to 5.91e6: xrefs and trailer Due to the inept content-stream structure I demonstrated above, if the page content streams were uncompressed, they would be about 3.3× as large, going from about 3.12 megabytes to 10.3 megabytes. This would inflate the Derctuo PDF from 5.9 megabytes to 13.1 megabytes, which works out to about 13KiB per page. This is about three times bigger than plain ASCII text, but that’s only because of how badly I screwed the pooch in building the content streams. I was mostly using Edward Tufte’s “ET Book” TrueType version of Bembo, falling back to DejaVu Serif fonts for non-ASCII characters, and using Latin Modern Mono Light Condensed (a modified Computer Modern Typewriter) for typewriter text, falling back to FreeMono and DejaVu Sans Mono fonts. Embedding eight typefaces thus cost me 270K. If you want a PDF document to be much under 100K, you more or less have to restrict yourself to the 14 core PDF fonts instead of embedding your own, or hope that the fonts you want to use happen to be installed on the reader's system (prohibited in PDF/A and, I believe, PDF 2.0). Nearly half of the bytes in the Derctuo PDF are hyperlinks, which mostly look like this (I’ve elided the CRs ReportLab inserted before LFs): % 'Annot.NUMBER2006': class LinkAnnotation 2343 0 obj << /Border [ 0 0 .1 ] /C [ .6 .6 1 ] /Contents (notes/lithium-fuel.html) /Dest [ 2559 0 R /XYZ null null null ] /Rect [ 4.8 595.44 33.62578 609.84 ] /Subtype /Link /Type /Annot >> endobj 2559 0 obj is the /Page object for page 371, where the note on lithium fuel begins. I think there are a lot of opportunities for optimization here, including unnecessary whitespace, and I don’t think the PDF spec *requires* an /Annot to be a top-level object (I think you can embed it inside the /Page object), but honestly most of those would go away if you just used a PDF 1.5 deflated object stream. Kragen