Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #270421 > unrolled thread

Re: PDF Editor for Debian

Started byRichard <rrosner5@gmail.com>
First post2024-06-24 07:40 +0200
Last post2024-08-07 20:40 +0200
Articles 13 — 7 participants

Back to article view | Back to linux.debian.user

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: PDF Editor for Debian Richard <rrosner5@gmail.com> - 2024-06-24 07:40 +0200
    Re: PDF Editor for Debian jeremy ardley <jeremy.ardley@gmail.com> - 2024-06-24 10:40 +0200
      Publishing Formats (was: PDF Editor for Debian) Richard <rrosner5@gmail.com> - 2024-06-24 16:30 +0200
        Re: Publishing Formats jeremy ardley <jeremy.ardley@gmail.com> - 2024-06-24 23:30 +0200
          Re: Publishing Formats "Russell L. Harris" <russell@rlharris.org> - 2024-06-25 01:00 +0200
          Re: Publishing Formats Richard <rrosner5@gmail.com> - 2024-06-25 10:50 +0200
    Needed tool for vision-impaired - was [Re: PDF Editor for Debian] Richard Owlett <rowlett@access.net> - 2024-06-24 19:20 +0200
      Re: Needed tool for vision-impaired - was [Re: PDF Editor for Debian] Karen Lewellen <klewellen@shellworld.net> - 2024-06-24 19:30 +0200
        Re: Needed tool for vision-impaired - was [Re: PDF Editor for Debian] Nicolas George <george@nsup.org> - 2024-06-24 19:30 +0200
          Re: Needed tool for vision-impaired - was [Re: PDF Editor for Debian] Richard Owlett <rowlett@access.net> - 2024-08-08 14:20 +0200
        TARDY response -- [Re: Needed tool for vision-impaired - was [Re: PDF  Editor for Debian]] Richard Owlett <rowlett@access.net> - 2024-08-07 14:50 +0200
          Re: TARDY response -- [Re: Needed tool for vision-impaired Felix Miata <mrmazda@stanis.net> - 2024-08-07 19:50 +0200
            Re: TARDY response -- [Re: Needed tool for vision-impaired Richard Owlett <rowlett@access.net> - 2024-08-07 20:40 +0200

#270421 — Re: PDF Editor for Debian

FromRichard <rrosner5@gmail.com>
Date2024-06-24 07:40 +0200
SubjectRe: PDF Editor for Debian
Message-ID<ISKAN-59eV-1@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

Hello,
this very much depends on what you are expecting it to do. In general, PDFs
are only meant to be viewed - and printed - they where never meant for
anything else. Even filling out forms is just s bad hackjob through
JavaScript. That being said, there is software with PDF editing
capabilities on Linux, though it's much more basic than what you'll find on
Windows.

If you want to just make comments, Okular has some neat capabilities,
including signing PDFs. For handwritten notes on a PDF, Xournal++ is a
great tool. If you want to just want to reorder pages, rotate, delete or
add them, there are some tools like PDFSam. There's also the quite powerful
Ghostscript, though that's CLI only. At least I don't know of any GUI. For
more "editing" features, LibreOffice can import PDFs, but in my experience
it struggles quite a lot with layout. OnlyOffice also has that capability,
but I never used it. Also, Inkscape can do that. It can also import
multiple pages at once, but I recommend only importing single pages,
otherwise Inkscape quickly reaches its limits. It has two import modes, an
internal one and poppler. Use the internal one and see if that works for
you. It's easier to edit text boxes in there, but it's quite likely it
won't be able to use the right font, which will break the whole look. The
poppler import can preserve that, but that's because letters aren't
imported as letters but as paths. So you can't just edit text, you'd have
to delete letters and try to insert text in a way that looks decent.

Other than that, there are a few commercial tools, but they are not that
well known. So your best bet is just to try to never have to edit a PDF at
all. Always try to get a hand on the original file the PDF was delivered
from. Even if it's a docx - Microsofts infamous wannabe-open source format
that just nobody can handle properly, including their own software - it
will most likely be better handled by the software you use than a PDF made
editable.

Best
Richard

On Mon, Jun 24, 2024, 07:13 Arbol One <ArbolOne@hotmail.ca> wrote:

> Hello.
> Is there a PDF editor that would work with Debian 12?
>
> Thanks.
> --
> *ArbolOne.ca* Using Fire Fox and Thunderbird. ArbolOne is composed of
> students and volunteers dedicated to providing free services to charitable
> organizations. ArbolOne on Java Development is in progress [ í ]
>

[toc] | [next] | [standalone]


#270427

Fromjeremy ardley <jeremy.ardley@gmail.com>
Date2024-06-24 10:40 +0200
Message-ID<ISNoZ-5bn4-1@gated-at.bofh.it>
In reply to#270421
On 24/6/24 13:35, Richard wrote:
> So your best bet is just to try to never have to edit a PDF at all. 
> Always try to get a hand on the original file the PDF was delivered 
> from. Even if it's a docx 


In my view, pdf and docx shoud be regarded as publication formats for 
content managed in a professional content management system. HTML and 
odt and postscript also fall in to the category of publication formats.

Word documents suffer because back in the dim ages of the late 1980s 
Microsoft decided to merge content managing with content editing with 
content publishing and abysmally failed at all of them.

However, the easiest way to edit a pdf is convert it to word using say 
https://pdf2docx.com/ There are also plenty of ways in linux to do that 
but they all take time and effort to make work.

[toc] | [prev] | [next] | [standalone]


#270432 — Publishing Formats (was: PDF Editor for Debian)

FromRichard <rrosner5@gmail.com>
Date2024-06-24 16:30 +0200
SubjectPublishing Formats (was: PDF Editor for Debian)
Message-ID<ISSRH-5f1F-5@gated-at.bofh.it>
In reply to#270427
Since it's quite OT, starting a new thread for this.

I would most certainly never call formats like ooxml or odf “publishing formats”, they are content creation or editing formats. From a publishing format I expect to be able to show the content as intended — which actually neither of them can do 100 % can, the probability of messing up just isn't that big. Either you want a fixed format, e.g. for printing, what you get with the likes of PDF, PS, SVG or your various raster graphic formats. Or you want your content to adapt in a foreseeable way to the viewer, i.e. HTLM, usually with the help of CSS and worst case JS. Sure, ooxml and odf want to be the former, but due to technical caveats that's not necessarily possible. With ooxml, you have several incompatible versions you can't just easily tell apart, often making identical display impossible due to using but not embedding proprietary fonts by default — and being an abomination of a format spanning around 5500 pages plus another 1000 pages for their tranistional mode, that was 
only standardized by world-wide corruption. ODF usually does things way better, but support in software beyond LibreOffice is still often lacking — though that's not their fault since their format is much simpler, being documented in just around 1000 pages. But still, as it doesn't communicate fixed positions — and as far as I can tell doesn't imply those by telling the software explicitly how to render font, so the result will always look identical, and won't embed fonts — or the needed subset — by default, it's also kinda not fulfilling the needs.

And no, editing a PDF as docx isn't the easiest — not to mention best — way to edit a PDF, especially not with some ominous web tool. Maybe someone can write an AI for that, but even then it's most likely much easier to just go the OCR route to derive content and extract layout from the document. At least I don't know how strict PDF defines things, I only always hear that PDF is at least as much of an unholy mess as ooxml — which was supposed to be fixed by PDF 2.0, which still pretty much no software creates by default, even though most software seems to be supporting it — and writing tools like Ghostscript or Poppler is a royal pain. LaTeX can probably only circumvent this because they just have to create a PDF from a predefined set of functions — and be able to embed other PDFs into these PDFs. But the most reliable way to edit PDFs — as I have little to no experience with most commercial solutions — is Inkscapte. If the internal importer succeeds, you get creat text 
editing features, which obviously can't rival office suites, but at least you don't completely and almost guaranteed completely mess up the whole layout.

Richard


On 24.06.24 10:31, jeremy ardley wrote:
> In my view, pdf and docx shoud be regarded as publication formats for content managed in a professional content management system. HTML and odt and postscript also fall in to the category of publication formats.
>
> Word documents suffer because back in the dim ages of the late 1980s Microsoft decided to merge content managing with content editing with content publishing and abysmally failed at all of them.
>
> However, the easiest way to edit a pdf is convert it to word using say https://pdf2docx.com/ There are also plenty of ways in linux to do that but they all take time and effort to make work.
>

[toc] | [prev] | [next] | [standalone]


#270446 — Re: Publishing Formats

Fromjeremy ardley <jeremy.ardley@gmail.com>
Date2024-06-24 23:30 +0200
SubjectRe: Publishing Formats
Message-ID<ISZq9-5k0j-7@gated-at.bofh.it>
In reply to#270432
On 24/6/24 22:22, Richard wrote:
> Since it's quite OT, starting a new thread for this.
>
> I would most certainly never call formats like ooxml or odf 
> “publishing formats”, they are content creation or editing formats. 
> From a publishing format I expect to be able to show the content as 
> intended — which actually neither of them can do 100 % can, the 
> probability of messing up just isn't that big. Either you want a fixed 
> format, e.g. for printing, what you get with the likes of PDF, PS, SVG 
> or your various raster graphic formats. Or you want your content to 
> adapt in a foreseeable way to the viewer, i.e. HTLM, usually with the 
> help of CSS and worst case JS. Sure, ooxml and odf want to be the 
> former, but due to technical caveats that's not necessarily possible. 
> With ooxml, you have several incompatible versions you can't just 
> easily tell apart, often making identical display impossible due to 
> using but not embedding proprietary fonts by default — and being an 
> abomination of a format spanning around 5500 pages plus another 1000 
> pages for their tranistional mode, that was only standardized by 
> world-wide corruption. ODF usually does things way better, but support 
> in software beyond LibreOffice is still often lacking — though that's 
> not their fault since their format is much simpler, being documented 
> in just around 1000 pages. But still, as it doesn't communicate fixed 
> positions — and as far as I can tell doesn't imply those by telling 
> the software explicitly how to render font, so the result will always 
> look identical, and won't embed fonts — or the needed subset — by 
> default, it's also kinda not fulfilling the needs.
>
I triggered this by saying docx and pdf are publishing formats. In the 
world of professional content management that is exactly so. You have 
your content in a neutral format in a version controlled storage system, 
and you have choice to publish in pdf or docx or html or epub or 
whatever. What you don't do is use these output formats as your primary 
content.

Examples relevant to debian include package documentation such as man 
pages, markdown, doxygen, docbook, latex.

In fact I can't think of any project in debian that has pdf or docx as 
the primary source of documentation

Tools that can do this transform include pandoc, Visual Studio Code, 
ghostwriter, marktext and many many more.


> And no, editing a PDF as docx isn't the easiest — not to mention best 
> — way to edit a PDF, especially not with some ominous web tool. Maybe 
> someone can write an AI for that, but even then it's most likely much 
> easier to just go the OCR route to derive content and extract layout 
> from the document. At least I don't know how strict PDF defines 
> things, I only always hear that PDF is at least as much of an unholy 
> mess as ooxml — which was supposed to be fixed by PDF 2.0, which still 
> pretty much no software creates by default, even though most software 
> seems to be supporting it — and writing tools like Ghostscript or 
> Poppler is a royal pain. LaTeX can probably only circumvent this 
> because they just have to create a PDF from a predefined set of 
> functions — and be able to embed other PDFs into these PDFs. But the 
> most reliable way to edit PDFs — as I have little to no experience 
> with most commercial solutions — is Inkscapte. If the internal 
> importer succeeds, you get creat text editing features, which 
> obviously can't rival office suites, but at least you don't completely 
> and almost guaranteed completely mess up the whole layout.
>
> Richard


In my most recent experience, OCR of pdf documents is quite difficult if 
the layout is significant such as in bank statements. There are various 
tools to assist in extracting the content but it's quite marginal.

On the other hand, give a screenshot of a bank statement to the 'ai' 
GPT4 and ask it to extract all transactions in csv format and it is done 
perfectly.

On a sidenote, PDF is basically Postscript on steroids. Its entire 
purpose is to describe how content is to be placed on a printed page. On 
a side-side note Postscript is actually a programming language with 
specialty in text layout but quite capable of doing significant 
computation activities - so long as your output eventually gets rendered 
on a page.


>
>
> On 24.06.24 10:31, jeremy ardley wrote:
>> In my view, pdf and docx shoud be regarded as publication formats for 
>> content managed in a professional content management system. HTML and 
>> odt and postscript also fall in to the category of publication formats.
>>
>> Word documents suffer because back in the dim ages of the late 1980s 
>> Microsoft decided to merge content managing with content editing with 
>> content publishing and abysmally failed at all of them.
>>
>> However, the easiest way to edit a pdf is convert it to word using 
>> say https://pdf2docx.com/ There are also plenty of ways in linux to 
>> do that but they all take time and effort to make work.
>>

[toc] | [prev] | [next] | [standalone]


#270451 — Re: Publishing Formats

From"Russell L. Harris" <russell@rlharris.org>
Date2024-06-25 01:00 +0200
SubjectRe: Publishing Formats
Message-ID<IT0Pf-5kNg-1@gated-at.bofh.it>
In reply to#270446
Someone gave me an old SCEPTRE display with a screen 11.5 inch by 22
inch.  I never before saw the usefulness of a wide screen.

A reader such as Atril can take advantage of the wide screen, allowing
me to zoom in until the type size is comfortable, without the need to
scroll left and right to read each line.

RLH

[toc] | [prev] | [next] | [standalone]


#270460 — Re: Publishing Formats

FromRichard <rrosner5@gmail.com>
Date2024-06-25 10:50 +0200
SubjectRe: Publishing Formats
Message-ID<ITa2d-5qKe-3@gated-at.bofh.it>
In reply to#270446
On 24.06.24 23:28, jeremy ardley wrote:
> [...]You have your content in a neutral format [...]

ooxml is far from "neutral"...


> What you don't do is use these output formats as your primary content.

Obviously not. That's why they are publishing formats, as in you send that in to be published, not to be further edited. That's why formats like odf and ooxml are content creation and editing formats, not publishing formats.


> Examples relevant to debian include package documentation such as man pages, markdown, doxygen, docbook, latex.

That's just the same category like HTML, so that remark isn't adding anything to the discussion.


> In my most recent experience, OCR of pdf documents is quite difficult if the layout is significant such as in bank statements. There are various tools to assist in extracting the content but it's quite marginal.

Depends on the software you use. For all I know Abbyy has very capable OCR software (I think it's called FineReader) that is very much capable of handling various layouts and difficult to read - as in very old - fonts. That was already the case about a decade ago and I doubt the software has gotten any worse. But of course it's not available on Linux.


> On the other hand, give a screenshot of a bank statement to the 'ai' GPT4 and ask it to extract all transactions in csv format and it is done perfectly.

As I said, AI can help with that, Abbyy is using it too. Question only is, if locally run AI can do that too, as everything else is a guarantee for breaching data protection laws.


> On a sidenote, PDF is basically Postscript on steroids. Its entire purpose is to describe how content is to be placed on a printed page. On a side-side note Postscript is actually a programming language with specialty in text layout but quite capable of doing significant computation activities - so long as your output eventually gets rendered on a page.
Never said anything contrary to that. Just that it's a very difficult to handle format because many things weren't defined prior to PDF 2.0. You had a predefined feature set, but nobody told you how to implement it, so chances were high that things wouldn't work as intended with every reader. But it seems the community of programmers for PDF readers has found common ground long before PDF 2.0 was a thing, so at least the standardized things would work everywhere.

[toc] | [prev] | [next] | [standalone]


#270438 — Needed tool for vision-impaired - was [Re: PDF Editor for Debian]

FromRichard Owlett <rowlett@access.net>
Date2024-06-24 19:20 +0200
SubjectNeeded tool for vision-impaired - was [Re: PDF Editor for Debian]
Message-ID<ISVwd-5h49-1@gated-at.bofh.it>
In reply to#270421
On 06/24/2024 12:35 AM, Richard wrote:
> Hello,
> this very much depends on what you are expecting it to do. In general, PDFs
> are only meant to be viewed - and printed - they where never meant for
> anything else. ...

Second sentence should read:
> ... only meant to be viewed by those with *NORMAL* vision ...

I'm attempting to read a USDA document.[1]
The printed version of this document is marginally readable.

Tools such as "Atril Document Viewer" provide selected magnification.
For this particular document and monitor, 150% is comfortable. Requires 
re-positioning the viewpoint 500 to 600 times to read document.

For _this_ document, Atril can select all the text on a page in a manner 
that can be pasted in a "reasonable" manner to a Pluma document.

It will:
    a. ignore actual graphics.
    b. put title/headings/??? on a separate line.
    c. all text between full page-width title/headings/??? will be
       treated as a logical unit.
It will not:
    1. put a blank line between paragraphs.
    2. put a blank line above/below lines containing title/headings/???.
    3. identify superscripts in some manner.

All this suggests that it should be able to extract text from a PDF and 
create a HTML document likely using only <p>, <br>, <sup>, and <li> in 
its <body>.


[1] 
https://fns-prod.azureedge.us/sites/default/files/resource-files/TFP2021.pdf
     _Thrifty Food Plan, 2021_
     Food and Nutrition Service
     August 2021
     FNS-916

[toc] | [prev] | [next] | [standalone]


#270439 — Re: Needed tool for vision-impaired - was [Re: PDF Editor for Debian]

FromKaren Lewellen <klewellen@shellworld.net>
Date2024-06-24 19:30 +0200
SubjectRe: Needed tool for vision-impaired - was [Re: PDF Editor for Debian]
Message-ID<ISVFT-5hfa-5@gated-at.bofh.it>
In reply to#270438
Good afternoon.
I am providing another option that might help here.
robobraille,

www.robobraille.org
Provides services, free of charge, that will convert pdf files  to a 
number of different formats, including .html
They provide audio, mobi, and  convert epub files too..but I digress.
As a test, consider sending your file to
convert at robobraille.org
  correctly of course.
in the subjectline put html
leaving the body blank, and attach the file.
See if the .html file returned meets your needs.
Best,
Karen



On Mon, 24 Jun 2024, Richard Owlett wrote:

> On 06/24/2024 12:35 AM, Richard wrote:
>>  Hello,
>>  this very much depends on what you are expecting it to do. In general,
>>  PDFs
>>  are only meant to be viewed - and printed - they where never meant for
>>  anything else. ...
>
> Second sentence should read:
>>  ... only meant to be viewed by those with *NORMAL* vision ...
>
> I'm attempting to read a USDA document.[1]
> The printed version of this document is marginally readable.
>
> Tools such as "Atril Document Viewer" provide selected magnification.
> For this particular document and monitor, 150% is comfortable. Requires 
> re-positioning the viewpoint 500 to 600 times to read document.
>
> For _this_ document, Atril can select all the text on a page in a manner that 
> can be pasted in a "reasonable" manner to a Pluma document.
>
> It will:
>    a. ignore actual graphics.
>    b. put title/headings/??? on a separate line.
>    c. all text between full page-width title/headings/??? will be
>      treated as a logical unit.
> It will not:
>    1. put a blank line between paragraphs.
>    2. put a blank line above/below lines containing title/headings/???.
>    3. identify superscripts in some manner.
>
> All this suggests that it should be able to extract text from a PDF and 
> create a HTML document likely using only <p>, <br>, <sup>, and <li> in its 
> <body>.
>
>
> [1] 
> https://fns-prod.azureedge.us/sites/default/files/resource-files/TFP2021.pdf
>     _Thrifty Food Plan, 2021_
>     Food and Nutrition Service
>     August 2021
>     FNS-916
>
>

[toc] | [prev] | [next] | [standalone]


#270440 — Re: Needed tool for vision-impaired - was [Re: PDF Editor for Debian]

FromNicolas George <george@nsup.org>
Date2024-06-24 19:30 +0200
SubjectRe: Needed tool for vision-impaired - was [Re: PDF Editor for Debian]
Message-ID<ISVFT-5hfa-3@gated-at.bofh.it>
In reply to#270439
Karen Lewellen (12024-06-24):
> Good afternoon.
> I am providing another option that might help here.
> robobraille,
> 
> www.robobraille.org
> Provides services, free of charge, that will convert pdf files  to a number
> of different formats, including .html
> They provide audio, mobi, and  convert epub files too..but I digress.
> As a test, consider sending your file to
> convert at robobraille.org
>  correctly of course.
> in the subjectline put html
> leaving the body blank, and attach the file.
> See if the .html file returned meets your needs.

Interesting.

Do you know how they fare with math? I mean real, non-trivial formulas
produced by LaTeX like you would find in
https://arxiv.org/abs/1803.05929 ?

(I know, I could test. I will if you do not know the answer.)

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#272145 — Re: Needed tool for vision-impaired - was [Re: PDF Editor for Debian]

FromRichard Owlett <rowlett@access.net>
Date2024-08-08 14:20 +0200
SubjectRe: Needed tool for vision-impaired - was [Re: PDF Editor for Debian]
Message-ID<J9ahA-3yjh-15@gated-at.bofh.it>
In reply to#270440
On 06/24/2024 12:29 PM, Nicolas George wrote:
> Karen Lewellen (12024-06-24):
>> Good afternoon.
>> I am providing another option that might help here.
>> robobraille,
>>
>> www.robobraille.org
>> Provides services, free of charge, that will convert pdf files  to a number
>> of different formats, including .html
>> They provide audio, mobi, and  convert epub files too..but I digress.
>> As a test, consider sending your file to
>> convert at robobraille.org
>>   correctly of course.
>> in the subjectline put html
>> leaving the body blank, and attach the file.
>> See if the .html file returned meets your needs.
> 
> Interesting.
> 
> Do you know how they fare with math? I mean real, non-trivial formulas
> produced by LaTeX like you would find in
> https://arxiv.org/abs/1803.05929 ?
> 
> (I know, I could test. I will if you do not know the answer.)
> 
> Regards,
> 

While looking for something else I found
https://www.robobraille.org/resources/software-and-tools/#math

Relevant &/or useful?
HTH

[toc] | [prev] | [next] | [standalone]


#272111 — TARDY response -- [Re: Needed tool for vision-impaired - was [Re: PDF Editor for Debian]]

FromRichard Owlett <rowlett@access.net>
Date2024-08-07 14:50 +0200
SubjectTARDY response -- [Re: Needed tool for vision-impaired - was [Re: PDF Editor for Debian]]
Message-ID<J8Oh4-3kCi-9@gated-at.bofh.it>
In reply to#270439
On 06/24/2024 12:22 PM, Karen Lewellen wrote:
> Good afternoon.
> I am providing another option that might help here.
> robobraille,
> 
> www.robobraille.org
> Provides services, free of charge, that will convert pdf files  to a 
> number of different formats, including .html
> They provide audio, mobi, and  convert epub files too..but I digress.
> As a test, consider sending your file to
> convert at robobraille.org
>   correctly of course.
> in the subjectline put html
> leaving the body blank, and attach the file.
> See if the .html file returned meets your needs.
> Best,
> Karen
> 

I went to the site shortly after you posted.
*MY* browser (SeaMonkey 2.49.4  {32 bit Linux}) choked on it.
I didn't get a chance to visit local library to try another browser.
Forgot I had a copy of Firefox 68.10.0esr on my machine.
It ran fine.

I converted 
"https://fns-prod.azureedge.us/sites/default/files/resource-files/TFP2021.pdf" 
to both text and HTML.
The text version seems perfect.
The HTML version has problem of missing titles to several tables near 
end of file. There are 15 tables one after another. All table *contents* 
came thru OK. Only the last one had its associated title.

I'll give www.robobraille.org a heads-up about it.
As I've a peculiar local configuration of SeaMonkey, could another SM 
user run a quick check so that I can report if their site has a problem 
with SeaMonkey?

TIA



> 
> 
> On Mon, 24 Jun 2024, Richard Owlett wrote:
> 
>> On 06/24/2024 12:35 AM, Richard wrote:
>>>  Hello,
>>>  this very much depends on what you are expecting it to do. In general,
>>>  PDFs
>>>  are only meant to be viewed - and printed - they where never meant for
>>>  anything else. ...
>>
>> Second sentence should read:
>>>  ... only meant to be viewed by those with *NORMAL* vision ...
>>
>> I'm attempting to read a USDA document.[1]
>> The printed version of this document is marginally readable.
>>
>> Tools such as "Atril Document Viewer" provide selected magnification.
>> For this particular document and monitor, 150% is comfortable. 
>> Requires re-positioning the viewpoint 500 to 600 times to read document.
>>
>> For _this_ document, Atril can select all the text on a page in a 
>> manner that can be pasted in a "reasonable" manner to a Pluma document.
>>
>> It will:
>>    a. ignore actual graphics.
>>    b. put title/headings/??? on a separate line.
>>    c. all text between full page-width title/headings/??? will be
>>      treated as a logical unit.
>> It will not:
>>    1. put a blank line between paragraphs.
>>    2. put a blank line above/below lines containing title/headings/???.
>>    3. identify superscripts in some manner.
>>
>> All this suggests that it should be able to extract text from a PDF 
>> and create a HTML document likely using only <p>, <br>, <sup>, and 
>> <li> in its <body>.
>>
>>
>> [1] 
>> https://fns-prod.azureedge.us/sites/default/files/resource-files/TFP2021.pdf 
>>
>>     _Thrifty Food Plan, 2021_
>>     Food and Nutrition Service
>>     August 2021
>>     FNS-916
>>
>>
> 
> 

[toc] | [prev] | [next] | [standalone]


#272118 — Re: TARDY response -- [Re: Needed tool for vision-impaired

FromFelix Miata <mrmazda@stanis.net>
Date2024-08-07 19:50 +0200
SubjectRe: TARDY response -- [Re: Needed tool for vision-impaired
Message-ID<J8SXo-3nrG-13@gated-at.bofh.it>
In reply to#272111
Richard Owlett composed on 2024-08-07 07:45 (UTC-0500):

> I went to the site shortly after you posted.
> *MY* browser (SeaMonkey 2.49.4  {32 bit Linux}) choked on it.
> I didn't get a chance to visit local library to try another browser.
> Forgot I had a copy of Firefox 68.10.0esr on my machine.
> It ran fine.

> I converted 
> "https://fns-prod.azureedge.us/sites/default/files/resource-files/TFP2021.pdf" 
> to both text and HTML.
> The text version seems perfect.
> The HTML version has problem of missing titles to several tables near 
> end of file. There are 15 tables one after another. All table *contents* 
> came thru OK. Only the last one had its associated title.

> I'll give www.robobraille.org a heads-up about it.
> As I've a peculiar local configuration of SeaMonkey, could another SM 
> user run a quick check so that I can report if their site has a problem 
> with SeaMonkey?

The PDF loads fine here in current SeaMonkey 2.53.18.2 64bit. I haven't used
2.49.x in five or so years. I use the static build hosted on
http://archive.seamonkey-project.org/releases/ .
-- 
Evolution as taught in public schools is, like religion,
	based on faith, not based on science.

 Team OS/2 ** Reg. Linux User #211409 ** a11y rocks!

Felix Miata

[toc] | [prev] | [next] | [standalone]


#272119 — Re: TARDY response -- [Re: Needed tool for vision-impaired

FromRichard Owlett <rowlett@access.net>
Date2024-08-07 20:40 +0200
SubjectRe: TARDY response -- [Re: Needed tool for vision-impaired
Message-ID<J8TJL-3nWR-17@gated-at.bofh.it>
In reply to#272118
On 08/07/2024 12:44 PM, Felix Miata wrote:
> Richard Owlett composed on 2024-08-07 07:45 (UTC-0500):
> 
>> I went to the site shortly after you posted.
>> *MY* browser (SeaMonkey 2.49.4  {32 bit Linux}) choked on it.
>> I didn't get a chance to visit local library to try another browser.
>> Forgot I had a copy of Firefox 68.10.0esr on my machine.
>> It ran fine.
> 
>> I converted
>> "https://fns-prod.azureedge.us/sites/default/files/resource-files/TFP2021.pdf"
>> to both text and HTML.
>> The text version seems perfect.
>> The HTML version has problem of missing titles to several tables near
>> end of file. There are 15 tables one after another. All table *contents*
>> came thru OK. Only the last one had its associated title.
> 
>> I'll give www.robobraille.org a heads-up about it.
>> As I've a peculiar local configuration of SeaMonkey, could another SM
>> user run a quick check so that I can report if their site has a problem
>> with SeaMonkey?
> 
> The PDF loads fine here in current SeaMonkey 2.53.18.2 64bit. I haven't used
> 2.49.x in five or so years. I use the static build hosted on
> http://archive.seamonkey-project.org/releases/ .
> 

Thank you.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web