Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #244237 > unrolled thread
| Started by | Bob Bernstein <poobah@ruptured-duck.com> |
|---|---|
| First post | 2022-01-20 18:00 +0100 |
| Last post | 2022-01-21 13:40 +0100 |
| Articles | 4 — 3 participants |
Back to article view | Back to linux.debian.user
Suggestions for tesseract Bob Bernstein <poobah@ruptured-duck.com> - 2022-01-20 18:00 +0100
Re: Suggestions for tesseract Siard <shiems@mailbox.org> - 2022-01-20 18:30 +0100
Re: Suggestions for tesseract Curt <curty@free.fr> - 2022-01-20 19:50 +0100
Re: Suggestions for tesseract Siard <shiems@mailbox.org> - 2022-01-21 13:40 +0100
| From | Bob Bernstein <poobah@ruptured-duck.com> |
|---|---|
| Date | 2022-01-20 18:00 +0100 |
| Subject | Suggestions for tesseract |
| Message-ID | <DHIX0-3ly-7@gated-at.bofh.it> |
Executing 'apt-cache search tesseract' brings up a multitude of
packages.
My need is simple enough, I think: I like to scan (using an
Epson scanner) pages of printed books -- almost one hundred per
cent text -- and then use OCR to produce pages from which I can
copy 'n paste snippets of text for note-taking purposes.
What do the assembled multitudes suggest for a tesseract package
(that's the OCR I've been encouraged to use) on my bullseye
system, which uname characterizes as:
Linux debian.localdomain 5.10.0-8-amd64 #1 SMP Debian 5.10.46-4
(2021-08-03) x86_64 GNU/Linux
Thank you.
--
"The existence of God is not an experimental issue in the way it was."
John Wisdom - "Gods" (1944)
[toc] | [next] | [standalone]
| From | Siard <shiems@mailbox.org> |
|---|---|
| Date | 2022-01-20 18:30 +0100 |
| Message-ID | <DHJq4-3KV-29@gated-at.bofh.it> |
| In reply to | #244237 |
Bob Bernstein wrote: > Executing 'apt-cache search tesseract' brings up a multitude of > packages. > > My need is simple enough, I think: I like to scan (using an > Epson scanner) pages of printed books -- almost one hundred per > cent text -- and then use OCR to produce pages from which I can > copy 'n paste snippets of text for note-taking purposes. > > What do the assembled multitudes suggest for a tesseract package > (that's the OCR I've been encouraged to use) on my bullseye > system, ... Once you have a PDF containing the images (img2pdf may be used for that), I think the cleverest way is to use ocrmypdf. It adds an OCR text layer to the PDF file, so the PDF text becomes selectable and can be copied. It uses the Tesseract OCR engine. $ ocrmypdf -f inputfile.pdf outputfile.pdf
[toc] | [prev] | [next] | [standalone]
| From | Curt <curty@free.fr> |
|---|---|
| Date | 2022-01-20 19:50 +0100 |
| Message-ID | <DHKFr-4r5-3@gated-at.bofh.it> |
| In reply to | #244240 |
On 2022-01-20, Siard <shiems@mailbox.org> wrote:
> Bob Bernstein wrote:
>> Executing 'apt-cache search tesseract' brings up a multitude of
>> packages.
>>
>> My need is simple enough, I think: I like to scan (using an
>> Epson scanner) pages of printed books -- almost one hundred per
>> cent text -- and then use OCR to produce pages from which I can
>> copy 'n paste snippets of text for note-taking purposes.
>>
>> What do the assembled multitudes suggest for a tesseract package
>> (that's the OCR I've been encouraged to use) on my bullseye
>> system, ...
>
> Once you have a PDF containing the images (img2pdf may be used for
> that), I think the cleverest way is to use ocrmypdf.
> It adds an OCR text layer to the PDF file, so the PDF text becomes
> selectable and can be copied.
> It uses the Tesseract OCR engine.
>
> $ ocrmypdf -f inputfile.pdf outputfile.pdf
>
ocrmypdf has quite a few dependencies on my machine.
The multitude of packages corresponds more or less to the multiple
languages of the human multitude. I guess the OP's working in English
('tesseract-ocr-eng', pulled in with all the others here when installing
the above).
[toc] | [prev] | [next] | [standalone]
| From | Siard <shiems@mailbox.org> |
|---|---|
| Date | 2022-01-21 13:40 +0100 |
| Message-ID | <DI1mX-6nv-17@gated-at.bofh.it> |
| In reply to | #244242 |
On Thu, 20 Jan 2022, Curt wrote:
> On 2022-01-20, Siard <shiems@mailbox.org> wrote:
> > Bob Bernstein wrote:
> > > Executing 'apt-cache search tesseract' brings up a multitude of
> > > packages.
> > >
> > > My need is simple enough, I think: I like to scan (using an
> > > Epson scanner) pages of printed books -- almost one hundred per
> > > cent text -- and then use OCR to produce pages from which I can
> > > copy 'n paste snippets of text for note-taking purposes.
> > >
> > > What do the assembled multitudes suggest for a tesseract package
> > > (that's the OCR I've been encouraged to use) on my bullseye
> > > system, ...
> >
> > Once you have a PDF containing the images (img2pdf may be used for
> > that), I think the cleverest way is to use ocrmypdf.
> > It adds an OCR text layer to the PDF file, so the PDF text becomes
> > selectable and can be copied.
> > It uses the Tesseract OCR engine.
> >
> > $ ocrmypdf -f inputfile.pdf outputfile.pdf
>
> ocrmypdf has quite a few dependencies on my machine.
>
> The multitude of packages corresponds more or less to the multiple
> languages of the human multitude. I guess the OP's working in English
> ('tesseract-ocr-eng', pulled in with all the others here when installing
> the above).
With tesseract and one tesseract language package already installed,
installing ocrmypdf does not pull in more of them. At least, that's what
I see on my machine.
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.user
csiph-web