Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #244240

Re: Suggestions for tesseract

From Siard <shiems@mailbox.org>
Newsgroups linux.debian.user
Subject Re: Suggestions for tesseract
Date 2022-01-20 18:30 +0100
Message-ID <DHJq4-3KV-29@gated-at.bofh.it> (permalink)
References <DHIX0-3ly-7@gated-at.bofh.it>
Organization linux.* mail to news gateway

Show all headers | View raw


Bob Bernstein wrote:
> Executing 'apt-cache search tesseract' brings up a multitude of 
> packages.
> 
> My need is simple enough, I think: I like to scan (using an 
> Epson scanner) pages of printed books -- almost one hundred per 
> cent text -- and then use OCR to produce pages from which I can 
> copy 'n paste snippets of text for note-taking purposes.
> 
> What do the assembled multitudes suggest for a tesseract package 
> (that's the OCR I've been encouraged to use) on my bullseye 
> system, ...

Once you have a PDF containing the images (img2pdf may be used for
that), I think the cleverest way is to use ocrmypdf.
It adds an OCR text layer to the PDF file, so the PDF text becomes
selectable and can be copied.
It uses the Tesseract OCR engine.

$ ocrmypdf -f inputfile.pdf outputfile.pdf

Back to linux.debian.user | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread


Thread

Suggestions for tesseract Bob Bernstein <poobah@ruptured-duck.com> - 2022-01-20 18:00 +0100
  Re: Suggestions for tesseract Siard <shiems@mailbox.org> - 2022-01-20 18:30 +0100
    Re: Suggestions for tesseract Curt <curty@free.fr> - 2022-01-20 19:50 +0100
      Re: Suggestions for tesseract Siard <shiems@mailbox.org> - 2022-01-21 13:40 +0100

csiph-web