Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #244240
| From | Siard <shiems@mailbox.org> |
|---|---|
| Newsgroups | linux.debian.user |
| Subject | Re: Suggestions for tesseract |
| Date | 2022-01-20 18:30 +0100 |
| Message-ID | <DHJq4-3KV-29@gated-at.bofh.it> (permalink) |
| References | <DHIX0-3ly-7@gated-at.bofh.it> |
| Organization | linux.* mail to news gateway |
Bob Bernstein wrote: > Executing 'apt-cache search tesseract' brings up a multitude of > packages. > > My need is simple enough, I think: I like to scan (using an > Epson scanner) pages of printed books -- almost one hundred per > cent text -- and then use OCR to produce pages from which I can > copy 'n paste snippets of text for note-taking purposes. > > What do the assembled multitudes suggest for a tesseract package > (that's the OCR I've been encouraged to use) on my bullseye > system, ... Once you have a PDF containing the images (img2pdf may be used for that), I think the cleverest way is to use ocrmypdf. It adds an OCR text layer to the PDF file, so the PDF text becomes selectable and can be copied. It uses the Tesseract OCR engine. $ ocrmypdf -f inputfile.pdf outputfile.pdf
Back to linux.debian.user | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
Suggestions for tesseract Bob Bernstein <poobah@ruptured-duck.com> - 2022-01-20 18:00 +0100
Re: Suggestions for tesseract Siard <shiems@mailbox.org> - 2022-01-20 18:30 +0100
Re: Suggestions for tesseract Curt <curty@free.fr> - 2022-01-20 19:50 +0100
Re: Suggestions for tesseract Siard <shiems@mailbox.org> - 2022-01-21 13:40 +0100
csiph-web