Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #203774 > unrolled thread
| Started by | kamaraju kusumanchi <raju.mailinglists@gmail.com> |
|---|---|
| First post | 2019-01-01 18:40 +0100 |
| Last post | 2019-01-04 11:00 +0100 |
| Articles | 11 on this page of 51 — 16 participants |
Back to article view | Back to linux.debian.user
[OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-01 18:40 +0100
Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-01 19:50 +0100
Re: [OT] scanned files are large in size Anders Andersson <pipatron@gmail.com> - 2019-01-01 20:00 +0100
Re: [OT] scanned files are large in size "Thomas Schmitt" <scdbackup@gmx.net> - 2019-01-01 20:30 +0100
Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-01 20:40 +0100
Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-02 04:50 +0100
Re: [OT] scanned files are large in size tomas@tuxteam.de - 2019-01-02 10:40 +0100
Re: [OT] scanned files are large in size Joe <joe@jretrading.com> - 2019-01-02 11:10 +0100
Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 11:30 +0100
Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-02 12:10 +0100
Re: [OT] scanned files are large in size Chris Ramsden <chris.ramsden@gmail.com> - 2019-01-02 12:10 +0100
Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 12:40 +0100
Re: [OT] scanned files are large in size mick crane <mick.crane@gmail.com> - 2019-01-02 20:20 +0100
Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 20:30 +0100
Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 03:50 +0100
Re: [OT] scanned files are large in size Siard <shiems146@kpnplanet.nl> - 2019-01-03 14:10 +0100
Re: [OT] scanned files are large in size Nicolas George <george@nsup.org> - 2019-01-03 14:20 +0100
Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-03 15:20 +0100
Re: [OT] scanned files are large in size Siard <shiems146@kpnplanet.nl> - 2019-01-03 16:00 +0100
Re: [OT] scanned files are large in size Nicolas George <george@nsup.org> - 2019-01-03 17:30 +0100
Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-03 17:30 +0100
Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 21:30 +0100
Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 15:50 +0100
Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 16:00 +0100
Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-02 16:20 +0100
Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 16:30 +0100
Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-02 17:20 +0100
Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 17:40 +0100
Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 20:10 +0100
Re: [OT] scanned files are large in size Nicolas George <george@nsup.org> - 2019-01-02 16:20 +0100
Re: [OT] scanned files are large in size tomas@tuxteam.de - 2019-01-02 16:20 +0100
Re: [OT] scanned files are large in size Joe <joe@jretrading.com> - 2019-01-02 18:00 +0100
Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 04:10 +0100
Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 03:30 +0100
Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-03 05:00 +0100
Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-04 18:30 +0100
Re: [OT] scanned files are large in size Gene Heskett <gheskett@shentel.net> - 2019-01-04 19:50 +0100
Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-04 20:50 +0100
Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-04 20:40 +0100
Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-04 21:00 +0100
Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-05 03:20 +0100
Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-01 21:10 +0100
Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-02 05:00 +0100
Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 20:20 +0100
Re: [OT] scanned files are large in size Jörg-Volker Peetz <jvpeetz@web.de> - 2019-01-02 12:40 +0100
Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-03 05:00 +0100
Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 06:10 +0100
Re: [OT] scanned files are large in size Jörg-Volker Peetz <jvpeetz@web.de> - 2019-01-03 09:40 +0100
Re: [OT] scanned files are large in size Jonathan Dowland <jmtd@debian.org> - 2019-01-03 14:50 +0100
Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 21:50 +0100
Re: [OT] scanned files are large in size Jonathan Dowland <jmtd@debian.org> - 2019-01-04 11:00 +0100
Page 3 of 3 — ← Prev page 1 2 [3]
| From | David Wright <deblis@lionunicorn.co.uk> |
|---|---|
| Date | 2019-01-05 03:20 +0100 |
| Message-ID | <xcJJ7-81R-5@gated-at.bofh.it> |
| In reply to | #203993 |
On Fri 04 Jan 2019 at 17:26:07 (+0000), Brian wrote: > On Wed 02 Jan 2019 at 22:56:22 -0500, kamaraju kusumanchi wrote: > > On Wed, Jan 2, 2019 at 9:23 PM David Wright <deblis@lionunicorn.co.uk> wrote: > > > On Wed 02 Jan 2019 at 14:44:14 (+0000), Brian wrote: > > > > > > > > I'm intrigued; I hadn't realised that conversion of the scanned image > > > > for some vendors' devices took place on the device itself. How do you > > > > know this happens? It is the frontend to SANE (xsane or scanimage, for > > > > example) which I've always associated with image aquisition conversion. > > > > > > It really is rather easy. You insert a USB stick into the scanner, > > > press scan, and later observe that a JPEG or PDF file has appeared > > > on the stick, as appropriate. > > > > Yes, that is precisely what I did. Stick a USB into the scanner and > > press the scan button. > > My HP Envy 4520 has no such button. There is an option for scanning to > the computer, but software is required on the computer to do that and > HPLIP does not provide it. > > Anyway, I managed to persuade the device to give me the PDF it would > have sent to a USB stick if the facility had existed (the device has > Apple's AirScan). If it matters, the PDF does not have any Creator or > Publisher information and doesn't contain any embedded or subset fonts. It sounds as if this is sufficient to make you confident that the device is doing the conversion and not the computer: anything that decouples the two from privately passing information to one another outside the delivered file. A USB stick, or email, is just the most obvious. > Scanned at a resolution of 600: > > brian@desktop:~$ pdfimages -list out.pdf > page num type width height color comp bpc enc interp object ID x-ppi y-ppi size ratio > -------------------------------------------------------------------------------------------- > 1 0 image 5100 6600 gray 1 8 jpeg no 1 0 600 600 2090K 6.4% > > ps2pdf reduces the 2090K by about 50% to 1051K. > > A different scanner device and source document, of course, and maybe > different methods of PDF production, so I wouldn't read too much into > this. Proving whether any compression applied is lossless is more difficult because pdfimages seems mute on what processes were carried out in extracting an image from the PDF. I have made the assumption that scanning compressed means that lossy compression is applied whereas scanning "uncompressed" means that lossless compression is applied. Cheers, David.
[toc] | [prev] | [next] | [standalone]
| From | Brian <ad44@cityscape.co.uk> |
|---|---|
| Date | 2019-01-01 21:10 +0100 |
| Message-ID | <xbywp-5xO-1@gated-at.bofh.it> |
| In reply to | #203774 |
On Tue 01 Jan 2019 at 12:34:38 -0500, kamaraju kusumanchi wrote: > A scanned document from Canon pixma mx870 printer is significantly > larger compared to the same document scanned on a different scanner. Which is...? > When I look at both the images side by side on a PC, there is no > visual difference between the two. I am trying to understand the > underlying cause and fix it if possible. You could mention which scanning software you used and what the setting for the output file format was. > As shown below, scanned_in_office.pdf is 332Kb, scanned_on_mx870.pdf is 1.7 Mb. > > % ls -al scanned_in_office.pdf scanned_on_mx870.pdf > -rw-r--r-- 1 rajulocal rajulocal 331796 Jan 1 11:54 scanned_in_office.pdf > -rw-r--r-- 1 rajulocal rajulocal 1775460 Jan 1 11:48 scanned_on_mx870.pdf > > Both are are scanned at 600 dpi. The only difference I see is in bpc, > enc fields. > > % pdfimages -list scanned_in_office.pdf > page num type width height color comp bpc enc interp object ID > x-ppi y-ppi size ratio > -------------------------------------------------------------------------------------------- > 1 0 image 5104 6600 gray 1 1 ccitt no 7 0 > 601 600 183K 4.5% > 2 1 image 5104 6600 gray 1 1 ccitt no 14 0 > 601 600 138K 3.4% > > % pdfimages -list scanned_on_mx870.pdf > page num type width height color comp bpc enc interp object ID > x-ppi y-ppi size ratio > -------------------------------------------------------------------------------------------- > 1 0 image 5100 6600 gray 1 8 jpeg no 8 0 > 600 600 1066K 3.2% > 2 1 image 5100 6600 gray 1 8 jpeg no 14 0 > 600 600 665K 2.0% > > Questions: > 1) Does the large file size have anything to do with the printer > itself? Is there anything I can do (ex:- update the driver/firmware or > something)? Not at all; the printer has nothing to do with it. Printing is printing. Scanning is scanning. > 2) Is the difference in image sizes due to the bpc (1 vs. 8) or > encoding (ccitt vs jped) fields? Could be. > 3) If yes, how to change them? One file is in (I think) tiff format. The other isn't. You didn't scan like and like from both devices. -- Brian.
[toc] | [prev] | [next] | [standalone]
| From | kamaraju kusumanchi <raju.mailinglists@gmail.com> |
|---|---|
| Date | 2019-01-02 05:00 +0100 |
| Message-ID | <xbFRf-1zt-1@gated-at.bofh.it> |
| In reply to | #203783 |
On Tue, Jan 1, 2019 at 3:04 PM Brian <ad44@cityscape.co.uk> wrote: > > On Tue 01 Jan 2019 at 12:34:38 -0500, kamaraju kusumanchi wrote: > > > A scanned document from Canon pixma mx870 printer is significantly > > larger compared to the same document scanned on a different scanner. > > Which is...? Do not have this information at the moment. Will provide it tomorrow. > > When I look at both the images side by side on a PC, there is no > > visual difference between the two. I am trying to understand the > > underlying cause and fix it if possible. > > You could mention which scanning software you used and what the > setting for the output file format was. > Both images are obtained from the scanners directly. I did not use any specific software per se. The only setting I had to choose was the dpi - which in both cases is set to 600. > > Questions: > > 1) Does the large file size have anything to do with the printer > > itself? Is there anything I can do (ex:- update the driver/firmware or > > something)? > > Not at all; the printer has nothing to do with it. Printing is printing. > Scanning is scanning. > Understood. This is an 'all in one' printer which has both printing and scanning capabilities. > > 2) Is the difference in image sizes due to the bpc (1 vs. 8) or > > encoding (ccitt vs jped) fields? > > Could be. > > > 3) If yes, how to change them? > > One file is in (I think) tiff format. The other isn't. You didn't scan > like and like from both devices. There are not that many options to choose from the scan settings. You just choose the dpi and that is about it. I understand that we can't change much on what the scanner produces. But there should be some software to further change the scanner's output files? -- Kamaraju S Kusumanchi | http://raju.shoutwiki.com/wiki/Blog
[toc] | [prev] | [next] | [standalone]
| From | Brian <ad44@cityscape.co.uk> |
|---|---|
| Date | 2019-01-02 20:20 +0100 |
| Message-ID | <xbUdA-2bP-1@gated-at.bofh.it> |
| In reply to | #203788 |
On Tue 01 Jan 2019 at 22:51:02 -0500, kamaraju kusumanchi wrote: > On Tue, Jan 1, 2019 at 3:04 PM Brian <ad44@cityscape.co.uk> wrote: > > > > On Tue 01 Jan 2019 at 12:34:38 -0500, kamaraju kusumanchi wrote: > > > > > A scanned document from Canon pixma mx870 printer is significantly > > > larger compared to the same document scanned on a different scanner. > > > > Which is...? > > Do not have this information at the moment. Will provide it tomorrow. > > > > When I look at both the images side by side on a PC, there is no > > > visual difference between the two. I am trying to understand the > > > underlying cause and fix it if possible. > > > > You could mention which scanning software you used and what the > > setting for the output file format was. > > Both images are obtained from the scanners directly. I did not use any > specific software per se. The only setting I had to choose was the dpi > - which in both cases is set to 600. Ah, I think I see now. You used a button on the device to initiate a scan. I was thinking in terms of something like xsane being used. > > > Questions: > > > 1) Does the large file size have anything to do with the printer > > > itself? Is there anything I can do (ex:- update the driver/firmware or > > > something)? > > > > Not at all; the printer has nothing to do with it. Printing is printing. > > Scanning is scanning. > > > > Understood. This is an 'all in one' printer which has both printing > and scanning capabilities. > > > > 2) Is the difference in image sizes due to the bpc (1 vs. 8) or > > > encoding (ccitt vs jped) fields? > > > > Could be. > > > > > 3) If yes, how to change them? > > > > One file is in (I think) tiff format. The other isn't. You didn't scan > > like and like from both devices. > > There are not that many options to choose from the scan settings. You > just choose the dpi and that is about it. > > I understand that we can't change much on what the scanner produces. > But there should be some software to further change the scanner's > output files? Jörg-Volker Peetz has indicated a technique; it can work. For a smaller file size you could also reduce the resolution from 600. -- Brian.
[toc] | [prev] | [next] | [standalone]
| From | Jörg-Volker Peetz <jvpeetz@web.de> |
|---|---|
| Date | 2019-01-02 12:40 +0100 |
| Message-ID | <xbN2p-659-5@gated-at.bofh.it> |
| In reply to | #203774 |
With the pdf-files from my Canon scanner, I did shrink them with the help of ghostscript: $ ps2pdf old.pdf new.pdf Documentation can be found in ghostscript-doc. Regards, Jörg.
[toc] | [prev] | [next] | [standalone]
| From | kamaraju kusumanchi <raju.mailinglists@gmail.com> |
|---|---|
| Date | 2019-01-03 05:00 +0100 |
| Message-ID | <xc2kN-74w-3@gated-at.bofh.it> |
| In reply to | #203796 |
On Wed, Jan 2, 2019 at 6:33 AM Jörg-Volker Peetz <jvpeetz@web.de> wrote: > > With the pdf-files from my Canon scanner, I did shrink them with the help of > ghostscript: > > $ ps2pdf old.pdf new.pdf > This does not help. The file sizes are more or less the same (if anything, they are slightly larger). Original files: % ls -al scanned_in_office.pdf scanned_on_mx870.pdf -rw-r--r-- 1 rajulocal rajulocal 331796 Jan 1 11:54 scanned_in_office.pdf -rw-r--r-- 1 rajulocal rajulocal 1775460 Jan 1 11:48 scanned_on_mx870.pdf Conversion: % ps2pdf scanned_in_office.pdf file1.pdf % ps2pdf scanned_on_mx870.pdf file2.pdf New file sizes: % ls -al file1.pdf file2.pdf -rw-r--r-- 1 rajulocal rajulocal 338539 Jan 2 22:48 file1.pdf -rw-r--r-- 1 rajulocal rajulocal 1775470 Jan 2 22:48 file2.pdf -- Kamaraju S Kusumanchi | http://raju.shoutwiki.com/wiki/Blog
[toc] | [prev] | [next] | [standalone]
| From | David Wright <deblis@lionunicorn.co.uk> |
|---|---|
| Date | 2019-01-03 06:10 +0100 |
| Message-ID | <xc3qx-7Wc-1@gated-at.bofh.it> |
| In reply to | #203847 |
On Wed 02 Jan 2019 at 22:50:10 (-0500), kamaraju kusumanchi wrote: > On Wed, Jan 2, 2019 at 6:33 AM Jörg-Volker Peetz <jvpeetz@web.de> wrote: > > > > With the pdf-files from my Canon scanner, I did shrink them with the help of > > ghostscript: > > > > $ ps2pdf old.pdf new.pdf > > > > This does not help. The file sizes are more or less the same (if > anything, they are slightly larger). That's my usual experience too. However, I looked around for an old PDF and found a magazine distributed by my old employer. I could halve the file size as above. Looking at the original, there's a lot more legible XML (like image metadata) though it's difficult to tell if that accounts for the difference. I think it was produced on a Mac; most of the images certainly were. $ pdfinfo /tmp/original.pdf Creator: Adobe InDesign CS3 (5.0.2) Producer: Adobe PDF Library 8.0 CreationDate: Thu May 1 05:47:22 2008 CDT ModDate: Thu May 1 06:21:43 2008 CDT Tagged: no UserProperties: no Suspects: no Form: AcroForm JavaScript: no Pages: 48 Encrypted: no Page size: 595.276 x 841.89 pts (A4) Page rot: 0 File size: 7297299 bytes Optimized: no PDF version: 1.6 $ pdfinfo /tmp/new.pdf Creator: Adobe InDesign CS3 (5.0.2) Producer: GPL Ghostscript 9.26 CreationDate: Wed Jan 2 22:21:22 2019 CST ModDate: Wed Jan 2 22:21:22 2019 CST Tagged: no UserProperties: no Suspects: no Form: none JavaScript: no Pages: 48 Encrypted: no Page size: 595.276 x 841.89 pts (A4) Page rot: 0 File size: 3769525 bytes Optimized: no PDF version: 1.4 $ Cheers, David.
[toc] | [prev] | [next] | [standalone]
| From | Jörg-Volker Peetz <jvpeetz@web.de> |
|---|---|
| Date | 2019-01-03 09:40 +0100 |
| Message-ID | <xc6HL-1ql-1@gated-at.bofh.it> |
| In reply to | #203847 |
Maybe you could then try some of the switches for ps2pdf, for example $ ps2pdf -dPDFSETTINGS=/printer old.pdf new.pdf "/printer" makes it 300dpi, "/ebook" 150 dpi, and "/screen" 72 dpi, the documentation can tell you more. Regards, Jörg.
[toc] | [prev] | [next] | [standalone]
| From | Jonathan Dowland <jmtd@debian.org> |
|---|---|
| Date | 2019-01-03 14:50 +0100 |
| Message-ID | <xcbxM-4ge-3@gated-at.bofh.it> |
| In reply to | #203774 |
I'm replying to the top-level of this thread because it's not a direct reply to any particular message, but the thread reminded me of something. I occasionally scan large piles of paperwork using an MFP belonging to a local University. It emails me the results and has several options for the format and quality. What I wanted was lossless files, so I selected TIFF instead of JPEG or PDF. But I later discovered that modern TIFF is a versatile container format, and the printer was sending me JPEG-in-TIFF. -- ⢀⣴⠾⠻⢶⣦⠀ ⣾⠁⢠⠒⠀⣿⡁ Jonathan Dowland ⢿⡄⠘⠷⠚⠋⠀ https://jmtd.net ⠈⠳⣄⠀⠀⠀⠀ Please do not CC me, I am subscribed to the list.
[toc] | [prev] | [next] | [standalone]
| From | David Wright <deblis@lionunicorn.co.uk> |
|---|---|
| Date | 2019-01-03 21:50 +0100 |
| Message-ID | <xci6d-89I-7@gated-at.bofh.it> |
| In reply to | #203864 |
On Thu 03 Jan 2019 at 13:43:40 (+0000), Jonathan Dowland wrote: > I'm replying to the top-level of this thread because it's not a direct > reply to any particular message, but the thread reminded me of > something. > > I occasionally scan large piles of paperwork using an MFP belonging to a > local University. It emails me the results and has several options for > the format and quality. > > What I wanted was lossless files, so I selected TIFF instead of JPEG or > PDF. But I later discovered that modern TIFF is a versatile container > format, and the printer was sending me JPEG-in-TIFF. I can understand the mistake. TIFF was a godsend discovery for me when I had raw image data that I wanted read by "conventional software" as it's very easy to package it up. And the tagging system means that you can choose what to read in the file without getting indigestion. But then you realise that the more modern tags let you package all sorts of strange things in the file and if you want to read those sections, you've got to do all the dirty work of unpacking it. (I neven did more complex than palette colours and runlength encoding.) But what a disappointment that you didn't get simple tags and raw data. BTW a lot of people make a similar mistake with WAV files, assuming that a .wav extension specifies certain properties of the file contents. Cheers, David.
[toc] | [prev] | [next] | [standalone]
| From | Jonathan Dowland <jmtd@debian.org> |
|---|---|
| Date | 2019-01-04 11:00 +0100 |
| Message-ID | <xcuqK-78Q-13@gated-at.bofh.it> |
| In reply to | #203911 |
On Thu, Jan 03, 2019 at 02:42:10PM -0600, David Wright wrote: >But what a disappointment that you didn't get simple tags and >raw data. Yes, indeed it seems these particular printers will *always* give you JPEG, what changes is the wrapper. -- ⢀⣴⠾⠻⢶⣦⠀ ⣾⠁⢠⠒⠀⣿⡁ Jonathan Dowland ⢿⡄⠘⠷⠚⠋⠀ https://jmtd.net ⠈⠳⣄⠀⠀⠀⠀ Please do not CC me, I am subscribed to the list.
[toc] | [prev] | [standalone]
Page 3 of 3 — ← Prev page 1 2 [3]
Back to top | Article view | linux.debian.user
csiph-web