Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #185954 > unrolled thread

xsane & tesseract

Started by"Stephen Grant Brown" <steve.brown_nbn@iinet.net.au>
First post2017-08-26 03:50 +0200
Last post2017-08-26 10:50 +0200
Articles 4 — 4 participants

Back to article view | Back to linux.debian.user


Contents

  xsane & tesseract "Stephen Grant Brown" <steve.brown_nbn@iinet.net.au> - 2017-08-26 03:50 +0200
    Re: xsane & tesseract Doug <dmcgarrett@optonline.net> - 2017-08-26 04:30 +0200
      Re: xsane & tesseract Joe Pfeiffer <pfeiffer@cs.nmsu.edu> - 2017-08-26 06:30 +0200
        Re: xsane & tesseract Siard <shiems146@kpnplanet.nl> - 2017-08-26 10:50 +0200

#185954 — xsane & tesseract

From"Stephen Grant Brown" <steve.brown_nbn@iinet.net.au>
Date2017-08-26 03:50 +0200
Subjectxsane & tesseract
Message-ID<uiyox-3f2-3@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

Hi All,

How do I setup xsane to use the tesseract OCR engine?

I see gocr under preferences->setup->ocr.
Yours Sincerely
Stephen Grant Brown.

[toc] | [next] | [standalone]


#185956

FromDoug <dmcgarrett@optonline.net>
Date2017-08-26 04:30 +0200
Message-ID<uiz1f-3Ni-3@gated-at.bofh.it>
In reply to#185954

[Multipart message — attachments visible in raw view] — view raw

On 08/25/2017 08:31 PM, Stephen Grant Brown wrote:
> Hi All,
> How do I setup xsane to use the tesseract OCR engine?
> I see gocr under preferences->setup->ocr.
> Yours Sincerely
> Stephen Grant Brown.
Unless it has been vastly improved, you might as well copy the document 
by hand! Finding and fixing all the mistakes is not worth the trouble!
Abbyy for Windows does an excellent job.  One of only two programs I 
will boot Windows for. (The other one is a phono-to-CD program.)

--doug

[toc] | [prev] | [next] | [standalone]


#185963

FromJoe Pfeiffer <pfeiffer@cs.nmsu.edu>
Date2017-08-26 06:30 +0200
Message-ID<uiATn-4Zy-1@gated-at.bofh.it>
In reply to#185956
Doug <dmcgarrett@optonline.net> writes:

> On 08/25/2017 08:31 PM, Stephen Grant Brown wrote:
>
>  Hi All,
>  How do I setup xsane to use the tesseract OCR engine?
>  I see gocr under preferences->setup->ocr.
>  Yours Sincerely
>  Stephen Grant Brown.
>
> Unless it has been vastly improved, you might as well copy the document by hand! Finding and fixing all the mistakes is not worth the
> trouble!
> Abbyy for Windows does an excellent job. One of only two programs I will boot Windows for. (The other one is a phono-to-CD program.)

My experience OCRing a 16 page document with tesseract last spring was
quite good.  I didn't try to set xsane up to do it (as I thought it
would be a *long* time before I did it again), I scanned the document to
ppm files, sent them to tesseract, put the output of tesseract into a
.txt file, and cleaned up from there.  While it wasn't perfect, it was
far better than retyping the whole thing would have been.
-- 
"Erwin, have you seen the cat?" -- Mrs. Shrödinger

[toc] | [prev] | [next] | [standalone]


#185974

FromSiard <shiems146@kpnplanet.nl>
Date2017-08-26 10:50 +0200
Message-ID<uiEX0-7nX-13@gated-at.bofh.it>
In reply to#185963
Joe Pfeiffer wrote:
> I scanned the document to ppm files, sent them to tesseract, put the
> output of tesseract into a .txt file, and cleaned up from there.

You could try gimagereader, a frontend for tesseract, making this
process somewhat easier. Among others, it uses a spell checker, so
errors are easily recognizable.

If the resolution of the scanned image is (at least) 300 dpi, then my
findings are that text recognition with tesseract is very good.

There is also ocrmypdf, using tesseract, adding a text layer to a pdf
consisting of scanned documents, making the pdf searchable. Also works
very well.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web