Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #203774 > unrolled thread

[OT] scanned files are large in size

Started bykamaraju kusumanchi <raju.mailinglists@gmail.com>
First post2019-01-01 18:40 +0100
Last post2019-01-04 11:00 +0100
Articles 11 on this page of 51 — 16 participants

Back to article view | Back to linux.debian.user


Contents

  [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-01 18:40 +0100
    Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-01 19:50 +0100
      Re: [OT] scanned files are large in size Anders Andersson <pipatron@gmail.com> - 2019-01-01 20:00 +0100
        Re: [OT] scanned files are large in size "Thomas Schmitt" <scdbackup@gmx.net> - 2019-01-01 20:30 +0100
        Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-01 20:40 +0100
      Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-02 04:50 +0100
        Re: [OT] scanned files are large in size tomas@tuxteam.de - 2019-01-02 10:40 +0100
          Re: [OT] scanned files are large in size Joe <joe@jretrading.com> - 2019-01-02 11:10 +0100
            Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 11:30 +0100
              Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-02 12:10 +0100
              Re: [OT] scanned files are large in size Chris Ramsden <chris.ramsden@gmail.com> - 2019-01-02 12:10 +0100
                Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 12:40 +0100
              Re: [OT] scanned files are large in size mick crane <mick.crane@gmail.com> - 2019-01-02 20:20 +0100
                Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 20:30 +0100
                Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 03:50 +0100
                  Re: [OT] scanned files are large in size Siard <shiems146@kpnplanet.nl> - 2019-01-03 14:10 +0100
                    Re: [OT] scanned files are large in size Nicolas George <george@nsup.org> - 2019-01-03 14:20 +0100
                    Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-03 15:20 +0100
                      Re: [OT] scanned files are large in size Siard <shiems146@kpnplanet.nl> - 2019-01-03 16:00 +0100
                        Re: [OT] scanned files are large in size Nicolas George <george@nsup.org> - 2019-01-03 17:30 +0100
                        Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-03 17:30 +0100
                    Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 21:30 +0100
        Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 15:50 +0100
          Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 16:00 +0100
            Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-02 16:20 +0100
              Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 16:30 +0100
                Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-02 17:20 +0100
                  Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 17:40 +0100
              Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 20:10 +0100
            Re: [OT] scanned files are large in size Nicolas George <george@nsup.org> - 2019-01-02 16:20 +0100
              Re: [OT] scanned files are large in size tomas@tuxteam.de - 2019-01-02 16:20 +0100
            Re: [OT] scanned files are large in size Joe <joe@jretrading.com> - 2019-01-02 18:00 +0100
              Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 04:10 +0100
          Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 03:30 +0100
            Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-03 05:00 +0100
              Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-04 18:30 +0100
                Re: [OT] scanned files are large in size Gene Heskett <gheskett@shentel.net> - 2019-01-04 19:50 +0100
                  Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-04 20:50 +0100
                Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-04 20:40 +0100
                  Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-04 21:00 +0100
                Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-05 03:20 +0100
    Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-01 21:10 +0100
      Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-02 05:00 +0100
        Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 20:20 +0100
    Re: [OT] scanned files are large in size Jörg-Volker Peetz <jvpeetz@web.de> - 2019-01-02 12:40 +0100
      Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-03 05:00 +0100
        Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 06:10 +0100
        Re: [OT] scanned files are large in size Jörg-Volker Peetz <jvpeetz@web.de> - 2019-01-03 09:40 +0100
    Re: [OT] scanned files are large in size Jonathan Dowland <jmtd@debian.org> - 2019-01-03 14:50 +0100
      Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 21:50 +0100
        Re: [OT] scanned files are large in size Jonathan Dowland <jmtd@debian.org> - 2019-01-04 11:00 +0100

Page 3 of 3 — ← Prev page 1 2 [3]


#204058

FromDavid Wright <deblis@lionunicorn.co.uk>
Date2019-01-05 03:20 +0100
Message-ID<xcJJ7-81R-5@gated-at.bofh.it>
In reply to#203993
On Fri 04 Jan 2019 at 17:26:07 (+0000), Brian wrote:
> On Wed 02 Jan 2019 at 22:56:22 -0500, kamaraju kusumanchi wrote:
> > On Wed, Jan 2, 2019 at 9:23 PM David Wright <deblis@lionunicorn.co.uk> wrote:
> > > On Wed 02 Jan 2019 at 14:44:14 (+0000), Brian wrote:
> > > >
> > > > I'm intrigued; I hadn't realised that conversion of the scanned image
> > > > for some vendors' devices took place on the device itself. How do you
> > > > know this happens? It is the frontend to SANE (xsane or scanimage, for
> > > > example) which I've always associated with image aquisition conversion.
> > >
> > > It really is rather easy. You insert a USB stick into the scanner,
> > > press scan, and later observe that a JPEG or PDF file has appeared
> > > on the stick, as appropriate.
> > 
> > Yes, that is precisely what I did. Stick a USB into the scanner and
> > press the scan button.
> 
> My HP Envy 4520 has no such button. There is an option for scanning to
> the computer, but software is required on the computer to do that and
> HPLIP does not provide it.
> 
> Anyway, I managed to persuade the device to give me the PDF it would
> have sent to a USB stick if the facility had existed (the device has
> Apple's AirScan). If it matters, the PDF does not have any Creator or
> Publisher information and doesn't contain any embedded or subset fonts.

It sounds as if this is sufficient to make you confident that the
device is doing the conversion and not the computer: anything that
decouples the two from privately passing information to one another
outside the delivered file. A USB stick, or email, is just the most
obvious.

> Scanned at a resolution of 600:
> 
> brian@desktop:~$ pdfimages -list out.pdf
> page   num  type   width height color comp bpc  enc interp  object ID x-ppi y-ppi size ratio
> --------------------------------------------------------------------------------------------
>    1     0 image    5100  6600  gray    1   8  jpeg   no         1  0   600   600 2090K 6.4%
> 
> ps2pdf reduces the 2090K by about 50% to 1051K.
> 
> A different scanner device and source document, of course, and maybe
> different methods of PDF production, so I wouldn't read too much into
> this.

Proving whether any compression applied is lossless is more difficult
because pdfimages seems mute on what processes were carried out in
extracting an image from the PDF. I have made the assumption that
scanning compressed means that lossy compression is applied whereas
scanning "uncompressed" means that lossless compression is applied.

Cheers,
David.

[toc] | [prev] | [next] | [standalone]


#203783

FromBrian <ad44@cityscape.co.uk>
Date2019-01-01 21:10 +0100
Message-ID<xbywp-5xO-1@gated-at.bofh.it>
In reply to#203774
On Tue 01 Jan 2019 at 12:34:38 -0500, kamaraju kusumanchi wrote:

> A scanned document from Canon pixma mx870 printer is significantly
> larger compared to the same document scanned on a different scanner.

Which is...?

> When I look at both the images side by side on a PC, there is no
> visual difference between the two. I am trying to understand the
> underlying cause and fix it if possible.

You could mention which scanning software you used and what the
setting for the output file format was.

> As shown below, scanned_in_office.pdf is 332Kb, scanned_on_mx870.pdf is 1.7 Mb.
> 
> % ls -al scanned_in_office.pdf scanned_on_mx870.pdf
> -rw-r--r-- 1 rajulocal rajulocal  331796 Jan  1 11:54 scanned_in_office.pdf
> -rw-r--r-- 1 rajulocal rajulocal 1775460 Jan  1 11:48 scanned_on_mx870.pdf
> 
> Both are are scanned at 600 dpi. The only difference I see is in bpc,
> enc fields.
> 
> % pdfimages -list scanned_in_office.pdf
> page   num  type   width height color comp bpc  enc interp  object ID
> x-ppi y-ppi size ratio
> --------------------------------------------------------------------------------------------
>   1     0 image    5104  6600  gray    1   1  ccitt  no         7  0
> 601   600  183K 4.5%
>   2     1 image    5104  6600  gray    1   1  ccitt  no        14  0
> 601   600  138K 3.4%
> 
> % pdfimages -list scanned_on_mx870.pdf
> page   num  type   width height color comp bpc  enc interp  object ID
> x-ppi y-ppi size ratio
> --------------------------------------------------------------------------------------------
>   1     0 image    5100  6600  gray    1   8  jpeg   no         8  0
> 600   600 1066K 3.2%
>   2     1 image    5100  6600  gray    1   8  jpeg   no        14  0
> 600   600  665K 2.0%
> 
> Questions:
> 1) Does the large file size have anything to do with the printer
> itself? Is there anything I can do (ex:- update the driver/firmware or
> something)?

Not at all; the printer has nothing to do with it. Printing is printing.
Scanning is scanning.

> 2) Is the difference in image sizes due to the bpc (1 vs. 8) or
> encoding (ccitt vs jped) fields?

Could be.

> 3) If yes, how to change them?

One file is in (I think) tiff format. The other isn't. You didn't scan
like and like from both devices.

-- 
Brian.

[toc] | [prev] | [next] | [standalone]


#203788

Fromkamaraju kusumanchi <raju.mailinglists@gmail.com>
Date2019-01-02 05:00 +0100
Message-ID<xbFRf-1zt-1@gated-at.bofh.it>
In reply to#203783
On Tue, Jan 1, 2019 at 3:04 PM Brian <ad44@cityscape.co.uk> wrote:
>
> On Tue 01 Jan 2019 at 12:34:38 -0500, kamaraju kusumanchi wrote:
>
> > A scanned document from Canon pixma mx870 printer is significantly
> > larger compared to the same document scanned on a different scanner.
>
> Which is...?

Do not have this information at the moment. Will provide it tomorrow.

> > When I look at both the images side by side on a PC, there is no
> > visual difference between the two. I am trying to understand the
> > underlying cause and fix it if possible.
>
> You could mention which scanning software you used and what the
> setting for the output file format was.
>

Both images are obtained from the scanners directly. I did not use any
specific software per se. The only setting I had to choose was the dpi
- which in both cases is set to 600.

> > Questions:
> > 1) Does the large file size have anything to do with the printer
> > itself? Is there anything I can do (ex:- update the driver/firmware or
> > something)?
>
> Not at all; the printer has nothing to do with it. Printing is printing.
> Scanning is scanning.
>

Understood. This is an 'all in one' printer which has both printing
and scanning capabilities.

> > 2) Is the difference in image sizes due to the bpc (1 vs. 8) or
> > encoding (ccitt vs jped) fields?
>
> Could be.
>
> > 3) If yes, how to change them?
>
> One file is in (I think) tiff format. The other isn't. You didn't scan
> like and like from both devices.

There are not that many options to choose from the scan settings. You
just choose the dpi and that is about it.

I understand that we can't change much on what the scanner produces.
But there should be some software to further change the scanner's
output files?

-- 
Kamaraju S Kusumanchi | http://raju.shoutwiki.com/wiki/Blog

[toc] | [prev] | [next] | [standalone]


#203824

FromBrian <ad44@cityscape.co.uk>
Date2019-01-02 20:20 +0100
Message-ID<xbUdA-2bP-1@gated-at.bofh.it>
In reply to#203788
On Tue 01 Jan 2019 at 22:51:02 -0500, kamaraju kusumanchi wrote:

> On Tue, Jan 1, 2019 at 3:04 PM Brian <ad44@cityscape.co.uk> wrote:
> >
> > On Tue 01 Jan 2019 at 12:34:38 -0500, kamaraju kusumanchi wrote:
> >
> > > A scanned document from Canon pixma mx870 printer is significantly
> > > larger compared to the same document scanned on a different scanner.
> >
> > Which is...?
> 
> Do not have this information at the moment. Will provide it tomorrow.
> 
> > > When I look at both the images side by side on a PC, there is no
> > > visual difference between the two. I am trying to understand the
> > > underlying cause and fix it if possible.
> >
> > You could mention which scanning software you used and what the
> > setting for the output file format was.
> 
> Both images are obtained from the scanners directly. I did not use any
> specific software per se. The only setting I had to choose was the dpi
> - which in both cases is set to 600.

Ah, I think I see now. You used a button on the device to initiate a
scan. I was thinking in terms of something like xsane being used.

> > > Questions:
> > > 1) Does the large file size have anything to do with the printer
> > > itself? Is there anything I can do (ex:- update the driver/firmware or
> > > something)?
> >
> > Not at all; the printer has nothing to do with it. Printing is printing.
> > Scanning is scanning.
> >
> 
> Understood. This is an 'all in one' printer which has both printing
> and scanning capabilities.
> 
> > > 2) Is the difference in image sizes due to the bpc (1 vs. 8) or
> > > encoding (ccitt vs jped) fields?
> >
> > Could be.
> >
> > > 3) If yes, how to change them?
> >
> > One file is in (I think) tiff format. The other isn't. You didn't scan
> > like and like from both devices.
> 
> There are not that many options to choose from the scan settings. You
> just choose the dpi and that is about it.
> 
> I understand that we can't change much on what the scanner produces.
> But there should be some software to further change the scanner's
> output files?

Jörg-Volker Peetz has indicated a technique; it can work. For a smaller
file size you could also reduce the resolution from 600.

-- 
Brian.

[toc] | [prev] | [next] | [standalone]


#203796

FromJörg-Volker Peetz <jvpeetz@web.de>
Date2019-01-02 12:40 +0100
Message-ID<xbN2p-659-5@gated-at.bofh.it>
In reply to#203774
With the pdf-files from my Canon scanner, I did shrink them with the help of
ghostscript:

$ ps2pdf  old.pdf  new.pdf

Documentation can be found in ghostscript-doc.

Regards,
Jörg.

[toc] | [prev] | [next] | [standalone]


#203847

Fromkamaraju kusumanchi <raju.mailinglists@gmail.com>
Date2019-01-03 05:00 +0100
Message-ID<xc2kN-74w-3@gated-at.bofh.it>
In reply to#203796
On Wed, Jan 2, 2019 at 6:33 AM Jörg-Volker Peetz <jvpeetz@web.de> wrote:
>
> With the pdf-files from my Canon scanner, I did shrink them with the help of
> ghostscript:
>
> $ ps2pdf  old.pdf  new.pdf
>

This does not help. The file sizes are more or less the same (if
anything, they are slightly larger).

Original files:
% ls -al scanned_in_office.pdf scanned_on_mx870.pdf
-rw-r--r-- 1 rajulocal rajulocal  331796 Jan  1 11:54 scanned_in_office.pdf
-rw-r--r-- 1 rajulocal rajulocal 1775460 Jan  1 11:48 scanned_on_mx870.pdf

Conversion:
% ps2pdf scanned_in_office.pdf file1.pdf
% ps2pdf scanned_on_mx870.pdf file2.pdf

New file sizes:
% ls -al file1.pdf file2.pdf
-rw-r--r-- 1 rajulocal rajulocal  338539 Jan  2 22:48 file1.pdf
-rw-r--r-- 1 rajulocal rajulocal 1775470 Jan  2 22:48 file2.pdf

-- 
Kamaraju S Kusumanchi | http://raju.shoutwiki.com/wiki/Blog

[toc] | [prev] | [next] | [standalone]


#203850

FromDavid Wright <deblis@lionunicorn.co.uk>
Date2019-01-03 06:10 +0100
Message-ID<xc3qx-7Wc-1@gated-at.bofh.it>
In reply to#203847
On Wed 02 Jan 2019 at 22:50:10 (-0500), kamaraju kusumanchi wrote:
> On Wed, Jan 2, 2019 at 6:33 AM Jörg-Volker Peetz <jvpeetz@web.de> wrote:
> >
> > With the pdf-files from my Canon scanner, I did shrink them with the help of
> > ghostscript:
> >
> > $ ps2pdf  old.pdf  new.pdf
> >
> 
> This does not help. The file sizes are more or less the same (if
> anything, they are slightly larger).

That's my usual experience too. However, I looked around for an old
PDF and found a magazine distributed by my old employer. I could halve
the file size as above. Looking at the original, there's a lot more
legible XML (like image metadata) though it's difficult to tell if
that accounts for the difference. I think it was produced on a Mac;
most of the images certainly were.

$ pdfinfo /tmp/original.pdf
Creator:        Adobe InDesign CS3 (5.0.2)
Producer:       Adobe PDF Library 8.0
CreationDate:   Thu May  1 05:47:22 2008 CDT
ModDate:        Thu May  1 06:21:43 2008 CDT
Tagged:         no
UserProperties: no
Suspects:       no
Form:           AcroForm
JavaScript:     no
Pages:          48
Encrypted:      no
Page size:      595.276 x 841.89 pts (A4)
Page rot:       0
File size:      7297299 bytes
Optimized:      no
PDF version:    1.6
$ pdfinfo /tmp/new.pdf
Creator:        Adobe InDesign CS3 (5.0.2)
Producer:       GPL Ghostscript 9.26
CreationDate:   Wed Jan  2 22:21:22 2019 CST
ModDate:        Wed Jan  2 22:21:22 2019 CST
Tagged:         no
UserProperties: no
Suspects:       no
Form:           none
JavaScript:     no
Pages:          48
Encrypted:      no
Page size:      595.276 x 841.89 pts (A4)
Page rot:       0
File size:      3769525 bytes
Optimized:      no
PDF version:    1.4
$ 

Cheers,
David.

[toc] | [prev] | [next] | [standalone]


#203856

FromJörg-Volker Peetz <jvpeetz@web.de>
Date2019-01-03 09:40 +0100
Message-ID<xc6HL-1ql-1@gated-at.bofh.it>
In reply to#203847
Maybe you could then try some of the switches for ps2pdf, for example

$ ps2pdf -dPDFSETTINGS=/printer  old.pdf  new.pdf

"/printer" makes it 300dpi, "/ebook" 150 dpi, and "/screen" 72 dpi, the
documentation can tell you more.

Regards,
Jörg.

[toc] | [prev] | [next] | [standalone]


#203864

FromJonathan Dowland <jmtd@debian.org>
Date2019-01-03 14:50 +0100
Message-ID<xcbxM-4ge-3@gated-at.bofh.it>
In reply to#203774
I'm replying to the top-level of this thread because it's not a direct
reply to any particular message, but the thread reminded me of
something.

I occasionally scan large piles of paperwork using an MFP belonging to a
local University. It emails me the results and has several options for
the format and quality.

What I wanted was lossless files, so I selected TIFF instead of JPEG or
PDF. But I later discovered that modern TIFF is a versatile container
format, and the printer was sending me JPEG-in-TIFF.

-- 

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Jonathan Dowland
⢿⡄⠘⠷⠚⠋⠀ https://jmtd.net
⠈⠳⣄⠀⠀⠀⠀ Please do not CC me, I am subscribed to the list.

[toc] | [prev] | [next] | [standalone]


#203911

FromDavid Wright <deblis@lionunicorn.co.uk>
Date2019-01-03 21:50 +0100
Message-ID<xci6d-89I-7@gated-at.bofh.it>
In reply to#203864
On Thu 03 Jan 2019 at 13:43:40 (+0000), Jonathan Dowland wrote:
> I'm replying to the top-level of this thread because it's not a direct
> reply to any particular message, but the thread reminded me of
> something.
> 
> I occasionally scan large piles of paperwork using an MFP belonging to a
> local University. It emails me the results and has several options for
> the format and quality.
> 
> What I wanted was lossless files, so I selected TIFF instead of JPEG or
> PDF. But I later discovered that modern TIFF is a versatile container
> format, and the printer was sending me JPEG-in-TIFF.

I can understand the mistake. TIFF was a godsend discovery for me when
I had raw image data that I wanted read by "conventional software" as
it's very easy to package it up. And the tagging system means that you
can choose what to read in the file without getting indigestion.
But then you realise that the more modern tags let you package all
sorts of strange things in the file and if you want to read those
sections, you've got to do all the dirty work of unpacking it.
(I neven did more complex than palette colours and runlength encoding.)

But what a disappointment that you didn't get simple tags and
raw data.

BTW a lot of people make a similar mistake with WAV files, assuming
that a .wav extension specifies certain properties of the file contents.

Cheers,
David.

[toc] | [prev] | [next] | [standalone]


#203945

FromJonathan Dowland <jmtd@debian.org>
Date2019-01-04 11:00 +0100
Message-ID<xcuqK-78Q-13@gated-at.bofh.it>
In reply to#203911
On Thu, Jan 03, 2019 at 02:42:10PM -0600, David Wright wrote:
>But what a disappointment that you didn't get simple tags and
>raw data.

Yes, indeed it seems these particular printers will *always* give you
JPEG, what changes is the wrapper.

-- 

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Jonathan Dowland
⢿⡄⠘⠷⠚⠋⠀ https://jmtd.net
⠈⠳⣄⠀⠀⠀⠀ Please do not CC me, I am subscribed to the list.

[toc] | [prev] | [standalone]


Page 3 of 3 — ← Prev page 1 2 [3]

Back to top | Article view | linux.debian.user


csiph-web