Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #244237 > unrolled thread

Suggestions for tesseract

Started byBob Bernstein <poobah@ruptured-duck.com>
First post2022-01-20 18:00 +0100
Last post2022-01-21 13:40 +0100
Articles 4 — 3 participants

Back to article view | Back to linux.debian.user


Contents

  Suggestions for tesseract Bob Bernstein <poobah@ruptured-duck.com> - 2022-01-20 18:00 +0100
    Re: Suggestions for tesseract Siard <shiems@mailbox.org> - 2022-01-20 18:30 +0100
      Re: Suggestions for tesseract Curt <curty@free.fr> - 2022-01-20 19:50 +0100
        Re: Suggestions for tesseract Siard <shiems@mailbox.org> - 2022-01-21 13:40 +0100

#244237 — Suggestions for tesseract

FromBob Bernstein <poobah@ruptured-duck.com>
Date2022-01-20 18:00 +0100
SubjectSuggestions for tesseract
Message-ID<DHIX0-3ly-7@gated-at.bofh.it>
Executing 'apt-cache search tesseract' brings up a multitude of 
packages.

My need is simple enough, I think: I like to scan (using an 
Epson scanner) pages of printed books -- almost one hundred per 
cent text -- and then use OCR to produce pages from which I can 
copy 'n paste snippets of text for note-taking purposes.

What do the assembled multitudes suggest for a tesseract package 
(that's the OCR I've been encouraged to use) on my bullseye 
system, which uname characterizes as:

Linux debian.localdomain 5.10.0-8-amd64 #1 SMP Debian 5.10.46-4 
(2021-08-03) x86_64 GNU/Linux

Thank you.

-- 
"The existence of God is not an experimental issue in the way it was."

                        John Wisdom - "Gods" (1944)

[toc] | [next] | [standalone]


#244240

FromSiard <shiems@mailbox.org>
Date2022-01-20 18:30 +0100
Message-ID<DHJq4-3KV-29@gated-at.bofh.it>
In reply to#244237
Bob Bernstein wrote:
> Executing 'apt-cache search tesseract' brings up a multitude of 
> packages.
> 
> My need is simple enough, I think: I like to scan (using an 
> Epson scanner) pages of printed books -- almost one hundred per 
> cent text -- and then use OCR to produce pages from which I can 
> copy 'n paste snippets of text for note-taking purposes.
> 
> What do the assembled multitudes suggest for a tesseract package 
> (that's the OCR I've been encouraged to use) on my bullseye 
> system, ...

Once you have a PDF containing the images (img2pdf may be used for
that), I think the cleverest way is to use ocrmypdf.
It adds an OCR text layer to the PDF file, so the PDF text becomes
selectable and can be copied.
It uses the Tesseract OCR engine.

$ ocrmypdf -f inputfile.pdf outputfile.pdf

[toc] | [prev] | [next] | [standalone]


#244242

FromCurt <curty@free.fr>
Date2022-01-20 19:50 +0100
Message-ID<DHKFr-4r5-3@gated-at.bofh.it>
In reply to#244240
On 2022-01-20, Siard <shiems@mailbox.org> wrote:
> Bob Bernstein wrote:
>> Executing 'apt-cache search tesseract' brings up a multitude of 
>> packages.
>> 
>> My need is simple enough, I think: I like to scan (using an 
>> Epson scanner) pages of printed books -- almost one hundred per 
>> cent text -- and then use OCR to produce pages from which I can 
>> copy 'n paste snippets of text for note-taking purposes.
>> 
>> What do the assembled multitudes suggest for a tesseract package 
>> (that's the OCR I've been encouraged to use) on my bullseye 
>> system, ...
>
> Once you have a PDF containing the images (img2pdf may be used for
> that), I think the cleverest way is to use ocrmypdf.
> It adds an OCR text layer to the PDF file, so the PDF text becomes
> selectable and can be copied.
> It uses the Tesseract OCR engine.
>
> $ ocrmypdf -f inputfile.pdf outputfile.pdf
>

ocrmypdf has quite a few dependencies on my machine.

The  multitude of packages corresponds more or less to the multiple
languages of the human multitude. I guess the OP's working in English
('tesseract-ocr-eng', pulled in with all the others here when installing
the above).

[toc] | [prev] | [next] | [standalone]


#244263

FromSiard <shiems@mailbox.org>
Date2022-01-21 13:40 +0100
Message-ID<DI1mX-6nv-17@gated-at.bofh.it>
In reply to#244242
On Thu, 20 Jan 2022, Curt wrote:
> On 2022-01-20, Siard <shiems@mailbox.org> wrote:
> > Bob Bernstein wrote:
> > > Executing 'apt-cache search tesseract' brings up a multitude of 
> > > packages.
> > >
> > > My need is simple enough, I think: I like to scan (using an 
> > > Epson scanner) pages of printed books -- almost one hundred per 
> > > cent text -- and then use OCR to produce pages from which I can 
> > > copy 'n paste snippets of text for note-taking purposes.
> > >
> > > What do the assembled multitudes suggest for a tesseract package 
> > > (that's the OCR I've been encouraged to use) on my bullseye 
> > > system, ...
> >
> > Once you have a PDF containing the images (img2pdf may be used for
> > that), I think the cleverest way is to use ocrmypdf.
> > It adds an OCR text layer to the PDF file, so the PDF text becomes
> > selectable and can be copied.
> > It uses the Tesseract OCR engine.
> >
> > $ ocrmypdf -f inputfile.pdf outputfile.pdf
>
> ocrmypdf has quite a few dependencies on my machine.
> 
> The  multitude of packages corresponds more or less to the multiple
> languages of the human multitude. I guess the OP's working in English
> ('tesseract-ocr-eng', pulled in with all the others here when installing
> the above).

With tesseract and one tesseract language package already installed,
installing ocrmypdf does not pull in more of them. At least, that's what
I see on my machine.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web