Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #185974

Re: xsane & tesseract

From Siard <shiems146@kpnplanet.nl>
Newsgroups linux.debian.user
Subject Re: xsane & tesseract
Date 2017-08-26 10:50 +0200
Message-ID <uiEX0-7nX-13@gated-at.bofh.it> (permalink)
References <uiyox-3f2-3@gated-at.bofh.it> <uiz1f-3Ni-3@gated-at.bofh.it> <uiATn-4Zy-1@gated-at.bofh.it>
Organization linux.* mail to news gateway

Show all headers | View raw


Joe Pfeiffer wrote:
> I scanned the document to ppm files, sent them to tesseract, put the
> output of tesseract into a .txt file, and cleaned up from there.

You could try gimagereader, a frontend for tesseract, making this
process somewhat easier. Among others, it uses a spell checker, so
errors are easily recognizable.

If the resolution of the scanned image is (at least) 300 dpi, then my
findings are that text recognition with tesseract is very good.

There is also ocrmypdf, using tesseract, adding a text layer to a pdf
consisting of scanned documents, making the pdf searchable. Also works
very well.

Back to linux.debian.user | Previous | Next — Previous in thread | Find similar | Unroll thread


Thread

xsane & tesseract "Stephen Grant Brown" <steve.brown_nbn@iinet.net.au> - 2017-08-26 03:50 +0200
  Re: xsane & tesseract Doug <dmcgarrett@optonline.net> - 2017-08-26 04:30 +0200
    Re: xsane & tesseract Joe Pfeiffer <pfeiffer@cs.nmsu.edu> - 2017-08-26 06:30 +0200
      Re: xsane & tesseract Siard <shiems146@kpnplanet.nl> - 2017-08-26 10:50 +0200

csiph-web