Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #185954 > unrolled thread
| Started by | "Stephen Grant Brown" <steve.brown_nbn@iinet.net.au> |
|---|---|
| First post | 2017-08-26 03:50 +0200 |
| Last post | 2017-08-26 10:50 +0200 |
| Articles | 4 — 4 participants |
Back to article view | Back to linux.debian.user
xsane & tesseract "Stephen Grant Brown" <steve.brown_nbn@iinet.net.au> - 2017-08-26 03:50 +0200
Re: xsane & tesseract Doug <dmcgarrett@optonline.net> - 2017-08-26 04:30 +0200
Re: xsane & tesseract Joe Pfeiffer <pfeiffer@cs.nmsu.edu> - 2017-08-26 06:30 +0200
Re: xsane & tesseract Siard <shiems146@kpnplanet.nl> - 2017-08-26 10:50 +0200
| From | "Stephen Grant Brown" <steve.brown_nbn@iinet.net.au> |
|---|---|
| Date | 2017-08-26 03:50 +0200 |
| Subject | xsane & tesseract |
| Message-ID | <uiyox-3f2-3@gated-at.bofh.it> |
[Multipart message — attachments visible in raw view] — view raw
Hi All, How do I setup xsane to use the tesseract OCR engine? I see gocr under preferences->setup->ocr. Yours Sincerely Stephen Grant Brown.
[toc] | [next] | [standalone]
| From | Doug <dmcgarrett@optonline.net> |
|---|---|
| Date | 2017-08-26 04:30 +0200 |
| Message-ID | <uiz1f-3Ni-3@gated-at.bofh.it> |
| In reply to | #185954 |
[Multipart message — attachments visible in raw view] — view raw
On 08/25/2017 08:31 PM, Stephen Grant Brown wrote: > Hi All, > How do I setup xsane to use the tesseract OCR engine? > I see gocr under preferences->setup->ocr. > Yours Sincerely > Stephen Grant Brown. Unless it has been vastly improved, you might as well copy the document by hand! Finding and fixing all the mistakes is not worth the trouble! Abbyy for Windows does an excellent job. One of only two programs I will boot Windows for. (The other one is a phono-to-CD program.) --doug
[toc] | [prev] | [next] | [standalone]
| From | Joe Pfeiffer <pfeiffer@cs.nmsu.edu> |
|---|---|
| Date | 2017-08-26 06:30 +0200 |
| Message-ID | <uiATn-4Zy-1@gated-at.bofh.it> |
| In reply to | #185956 |
Doug <dmcgarrett@optonline.net> writes: > On 08/25/2017 08:31 PM, Stephen Grant Brown wrote: > > Hi All, > How do I setup xsane to use the tesseract OCR engine? > I see gocr under preferences->setup->ocr. > Yours Sincerely > Stephen Grant Brown. > > Unless it has been vastly improved, you might as well copy the document by hand! Finding and fixing all the mistakes is not worth the > trouble! > Abbyy for Windows does an excellent job. One of only two programs I will boot Windows for. (The other one is a phono-to-CD program.) My experience OCRing a 16 page document with tesseract last spring was quite good. I didn't try to set xsane up to do it (as I thought it would be a *long* time before I did it again), I scanned the document to ppm files, sent them to tesseract, put the output of tesseract into a .txt file, and cleaned up from there. While it wasn't perfect, it was far better than retyping the whole thing would have been. -- "Erwin, have you seen the cat?" -- Mrs. Shrödinger
[toc] | [prev] | [next] | [standalone]
| From | Siard <shiems146@kpnplanet.nl> |
|---|---|
| Date | 2017-08-26 10:50 +0200 |
| Message-ID | <uiEX0-7nX-13@gated-at.bofh.it> |
| In reply to | #185963 |
Joe Pfeiffer wrote: > I scanned the document to ppm files, sent them to tesseract, put the > output of tesseract into a .txt file, and cleaned up from there. You could try gimagereader, a frontend for tesseract, making this process somewhat easier. Among others, it uses a spell checker, so errors are easily recognizable. If the resolution of the scanned image is (at least) 300 dpi, then my findings are that text recognition with tesseract is very good. There is also ocrmypdf, using tesseract, adding a text layer to a pdf consisting of scanned documents, making the pdf searchable. Also works very well.
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.user
csiph-web