Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #185974
| From | Siard <shiems146@kpnplanet.nl> |
|---|---|
| Newsgroups | linux.debian.user |
| Subject | Re: xsane & tesseract |
| Date | 2017-08-26 10:50 +0200 |
| Message-ID | <uiEX0-7nX-13@gated-at.bofh.it> (permalink) |
| References | <uiyox-3f2-3@gated-at.bofh.it> <uiz1f-3Ni-3@gated-at.bofh.it> <uiATn-4Zy-1@gated-at.bofh.it> |
| Organization | linux.* mail to news gateway |
Joe Pfeiffer wrote: > I scanned the document to ppm files, sent them to tesseract, put the > output of tesseract into a .txt file, and cleaned up from there. You could try gimagereader, a frontend for tesseract, making this process somewhat easier. Among others, it uses a spell checker, so errors are easily recognizable. If the resolution of the scanned image is (at least) 300 dpi, then my findings are that text recognition with tesseract is very good. There is also ocrmypdf, using tesseract, adding a text layer to a pdf consisting of scanned documents, making the pdf searchable. Also works very well.
Back to linux.debian.user | Previous | Next — Previous in thread | Find similar | Unroll thread
xsane & tesseract "Stephen Grant Brown" <steve.brown_nbn@iinet.net.au> - 2017-08-26 03:50 +0200
Re: xsane & tesseract Doug <dmcgarrett@optonline.net> - 2017-08-26 04:30 +0200
Re: xsane & tesseract Joe Pfeiffer <pfeiffer@cs.nmsu.edu> - 2017-08-26 06:30 +0200
Re: xsane & tesseract Siard <shiems146@kpnplanet.nl> - 2017-08-26 10:50 +0200
csiph-web