Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #203774 > unrolled thread

[OT] scanned files are large in size

Started bykamaraju kusumanchi <raju.mailinglists@gmail.com>
First post2019-01-01 18:40 +0100
Last post2019-01-04 11:00 +0100
Articles 20 on this page of 51 — 16 participants

Back to article view | Back to linux.debian.user


Contents

  [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-01 18:40 +0100
    Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-01 19:50 +0100
      Re: [OT] scanned files are large in size Anders Andersson <pipatron@gmail.com> - 2019-01-01 20:00 +0100
        Re: [OT] scanned files are large in size "Thomas Schmitt" <scdbackup@gmx.net> - 2019-01-01 20:30 +0100
        Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-01 20:40 +0100
      Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-02 04:50 +0100
        Re: [OT] scanned files are large in size tomas@tuxteam.de - 2019-01-02 10:40 +0100
          Re: [OT] scanned files are large in size Joe <joe@jretrading.com> - 2019-01-02 11:10 +0100
            Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 11:30 +0100
              Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-02 12:10 +0100
              Re: [OT] scanned files are large in size Chris Ramsden <chris.ramsden@gmail.com> - 2019-01-02 12:10 +0100
                Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 12:40 +0100
              Re: [OT] scanned files are large in size mick crane <mick.crane@gmail.com> - 2019-01-02 20:20 +0100
                Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 20:30 +0100
                Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 03:50 +0100
                  Re: [OT] scanned files are large in size Siard <shiems146@kpnplanet.nl> - 2019-01-03 14:10 +0100
                    Re: [OT] scanned files are large in size Nicolas George <george@nsup.org> - 2019-01-03 14:20 +0100
                    Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-03 15:20 +0100
                      Re: [OT] scanned files are large in size Siard <shiems146@kpnplanet.nl> - 2019-01-03 16:00 +0100
                        Re: [OT] scanned files are large in size Nicolas George <george@nsup.org> - 2019-01-03 17:30 +0100
                        Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-03 17:30 +0100
                    Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 21:30 +0100
        Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 15:50 +0100
          Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 16:00 +0100
            Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-02 16:20 +0100
              Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 16:30 +0100
                Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-02 17:20 +0100
                  Re: [OT] scanned files are large in size <tomas@tuxteam.de> - 2019-01-02 17:40 +0100
              Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 20:10 +0100
            Re: [OT] scanned files are large in size Nicolas George <george@nsup.org> - 2019-01-02 16:20 +0100
              Re: [OT] scanned files are large in size tomas@tuxteam.de - 2019-01-02 16:20 +0100
            Re: [OT] scanned files are large in size Joe <joe@jretrading.com> - 2019-01-02 18:00 +0100
              Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 04:10 +0100
          Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 03:30 +0100
            Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-03 05:00 +0100
              Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-04 18:30 +0100
                Re: [OT] scanned files are large in size Gene Heskett <gheskett@shentel.net> - 2019-01-04 19:50 +0100
                  Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-04 20:50 +0100
                Re: [OT] scanned files are large in size deloptes <deloptes@gmail.com> - 2019-01-04 20:40 +0100
                  Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-04 21:00 +0100
                Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-05 03:20 +0100
    Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-01 21:10 +0100
      Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-02 05:00 +0100
        Re: [OT] scanned files are large in size Brian <ad44@cityscape.co.uk> - 2019-01-02 20:20 +0100
    Re: [OT] scanned files are large in size Jörg-Volker Peetz <jvpeetz@web.de> - 2019-01-02 12:40 +0100
      Re: [OT] scanned files are large in size kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2019-01-03 05:00 +0100
        Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 06:10 +0100
        Re: [OT] scanned files are large in size Jörg-Volker Peetz <jvpeetz@web.de> - 2019-01-03 09:40 +0100
    Re: [OT] scanned files are large in size Jonathan Dowland <jmtd@debian.org> - 2019-01-03 14:50 +0100
      Re: [OT] scanned files are large in size David Wright <deblis@lionunicorn.co.uk> - 2019-01-03 21:50 +0100
        Re: [OT] scanned files are large in size Jonathan Dowland <jmtd@debian.org> - 2019-01-04 11:00 +0100

Page 1 of 3  [1] 2 3  Next page →


#203774 — [OT] scanned files are large in size

Fromkamaraju kusumanchi <raju.mailinglists@gmail.com>
Date2019-01-01 18:40 +0100
Subject[OT] scanned files are large in size
Message-ID<xbwbf-3Xz-1@gated-at.bofh.it>
A scanned document from Canon pixma mx870 printer is significantly
larger compared to the same document scanned on a different scanner.
When I look at both the images side by side on a PC, there is no
visual difference between the two. I am trying to understand the
underlying cause and fix it if possible.

As shown below, scanned_in_office.pdf is 332Kb, scanned_on_mx870.pdf is 1.7 Mb.

% ls -al scanned_in_office.pdf scanned_on_mx870.pdf
-rw-r--r-- 1 rajulocal rajulocal  331796 Jan  1 11:54 scanned_in_office.pdf
-rw-r--r-- 1 rajulocal rajulocal 1775460 Jan  1 11:48 scanned_on_mx870.pdf

Both are are scanned at 600 dpi. The only difference I see is in bpc,
enc fields.

% pdfimages -list scanned_in_office.pdf
page   num  type   width height color comp bpc  enc interp  object ID
x-ppi y-ppi size ratio
--------------------------------------------------------------------------------------------
  1     0 image    5104  6600  gray    1   1  ccitt  no         7  0
601   600  183K 4.5%
  2     1 image    5104  6600  gray    1   1  ccitt  no        14  0
601   600  138K 3.4%

% pdfimages -list scanned_on_mx870.pdf
page   num  type   width height color comp bpc  enc interp  object ID
x-ppi y-ppi size ratio
--------------------------------------------------------------------------------------------
  1     0 image    5100  6600  gray    1   8  jpeg   no         8  0
600   600 1066K 3.2%
  2     1 image    5100  6600  gray    1   8  jpeg   no        14  0
600   600  665K 2.0%

Questions:
1) Does the large file size have anything to do with the printer
itself? Is there anything I can do (ex:- update the driver/firmware or
something)?
2) Is the difference in image sizes due to the bpc (1 vs. 8) or
encoding (ccitt vs jped) fields?
3) If yes, how to change them?

thanks
raju

-- 
Kamaraju S Kusumanchi | http://raju.shoutwiki.com/wiki/Blog

[toc] | [next] | [standalone]


#203778

From<tomas@tuxteam.de>
Date2019-01-01 19:50 +0100
Message-ID<xbxh1-4Bs-25@gated-at.bofh.it>
In reply to#203774

[Multipart message — attachments visible in raw view] — view raw

On Tue, Jan 01, 2019 at 12:34:38PM -0500, kamaraju kusumanchi wrote:
> A scanned document from Canon pixma mx870 printer is significantly
> larger compared to the same document scanned on a different scanner.
> When I look at both the images side by side on a PC, there is no
> visual difference between the two. I am trying to understand the
> underlying cause and fix it if possible.
> 
> As shown below, scanned_in_office.pdf is 332Kb, scanned_on_mx870.pdf is 1.7 Mb.
> 
> % ls -al scanned_in_office.pdf scanned_on_mx870.pdf
> -rw-r--r-- 1 rajulocal rajulocal  331796 Jan  1 11:54 scanned_in_office.pdf
> -rw-r--r-- 1 rajulocal rajulocal 1775460 Jan  1 11:48 scanned_on_mx870.pdf
> 
> Both are are scanned at 600 dpi. The only difference I see is in bpc,
> enc fields.

Yep. The one image is encoded as CCITT (aka Group 4, aka fax [1]), which is
passable for low res B&W images, but not that much for hi-res or color (or
gray scale). It compresses much worse than the other which is JPEG, which is
expressly made for hi-res and color (or grayscale) images.

OTOH, CCITT is lossless and JPEG lossy ;-)

> Questions:
> 1) Does the large file size have anything to do with the printer
> itself? Is there anything I can do (ex:- update the driver/firmware or
> something)?

That depends on what is encoding the images: does the scanner itself
"make" the PDF? Or some software, computer-side?

> 2) Is the difference in image sizes due to the bpc (1 vs. 8) or
> encoding (ccitt vs jped) fields?

CCITT vs JPEG, yes.

> 3) If yes, how to change them?

Hmmm. I don't know yet whether you have to talk to your scanner
or to your scan software...

Cheers

[1] https://en.wikipedia.org/wiki/Group_4_compression
-- tomás

[toc] | [prev] | [next] | [standalone]


#203779

FromAnders Andersson <pipatron@gmail.com>
Date2019-01-01 20:00 +0100
Message-ID<xbxqG-4F0-15@gated-at.bofh.it>
In reply to#203778

[Multipart message — attachments visible in raw view] — view raw

On Tue, Jan 1, 2019 at 7:40 PM <tomas@tuxteam.de> wrote:

> On Tue, Jan 01, 2019 at 12:34:38PM -0500, kamaraju kusumanchi wrote:
> > A scanned document from Canon pixma mx870 printer is significantly
> > larger compared to the same document scanned on a different scanner.
> > When I look at both the images side by side on a PC, there is no
> > visual difference between the two. I am trying to understand the
> > underlying cause and fix it if possible.
> >
> > As shown below, scanned_in_office.pdf is 332Kb, scanned_on_mx870.pdf is
> 1.7 Mb.
> >
> > % ls -al scanned_in_office.pdf scanned_on_mx870.pdf
> > -rw-r--r-- 1 rajulocal rajulocal  331796 Jan  1 11:54
> scanned_in_office.pdf
> > -rw-r--r-- 1 rajulocal rajulocal 1775460 Jan  1 11:48
> scanned_on_mx870.pdf
>
> Yep. The one image is encoded as CCITT (aka Group 4, aka fax [1]), which is
> passable for low res B&W images, but not that much for hi-res or color (or
> gray scale). It compresses much worse than the other which is JPEG, which
> is
> expressly made for hi-res and color (or grayscale) images.
>
> OTOH, CCITT is lossless and JPEG lossy ;-)
>

Not sure what you mean by "compresses much worse" here, but the CCITT
version is much smaller than the JPEG version. Maybe you meant that CCITT
looks worse after compression, which is weird when you also write that
CCITT is lossless!

[toc] | [prev] | [next] | [standalone]


#203780

From"Thomas Schmitt" <scdbackup@gmx.net>
Date2019-01-01 20:30 +0100
Message-ID<xbxTH-54c-7@gated-at.bofh.it>
In reply to#203779
Hi,

tomas wrote:
> > Yep. The one image is encoded as CCITT
> > It compresses much worse than the other which is JPEG,

Anders Andersson wrote:
> Not sure what you mean by "compresses much worse" here, but the CCITT
> version is much smaller than the JPEG version.

Because it uses 1 bit per channel/color whereas the other file uses
8 bpc. The fact that their size ratio is less than 1:8 demonstrates
the lossy compression power of the algorithm used in the 8 bpc image.

kamaraju kusumanchi, the OP, might see differences between both files if
using a viewer that can zoom-in far enough that scan pixels of 1/600 inch
become larger than the screen pixels.


Have a nice day :)

Thomas

[toc] | [prev] | [next] | [standalone]


#203781

From<tomas@tuxteam.de>
Date2019-01-01 20:40 +0100
Message-ID<xby3n-57E-1@gated-at.bofh.it>
In reply to#203779

[Multipart message — attachments visible in raw view] — view raw

On Tue, Jan 01, 2019 at 07:55:55PM +0100, Anders Andersson wrote:
> On Tue, Jan 1, 2019 at 7:40 PM <tomas@tuxteam.de> wrote:
> 
> > On Tue, Jan 01, 2019 at 12:34:38PM -0500, kamaraju kusumanchi wrote:

[...]

> > OTOH, CCITT is lossless and JPEG lossy ;-)
> >
> 
> Not sure what you mean by "compresses much worse" here, but the CCITT
> version is much smaller than the JPEG version. Maybe you meant that CCITT
> looks worse after compression, which is weird when you also write that
> CCITT is lossless!

You're right -- but Thomas has the solution for this riddle (the CCITT
throws away the grayscale information -- so, lossy too, after all :-)

Cheers
-- t

[toc] | [prev] | [next] | [standalone]


#203787

Fromkamaraju kusumanchi <raju.mailinglists@gmail.com>
Date2019-01-02 04:50 +0100
Message-ID<xbFHz-1wb-3@gated-at.bofh.it>
In reply to#203778
On Tue, Jan 1, 2019 at 1:40 PM <tomas@tuxteam.de> wrote:
>
> Yep. The one image is encoded as CCITT (aka Group 4, aka fax [1]), which is
> passable for low res B&W images, but not that much for hi-res or color (or
> gray scale). It compresses much worse than the other which is JPEG, which is
> expressly made for hi-res and color (or grayscale) images.
>
> OTOH, CCITT is lossless and JPEG lossy ;-)
>
ok, thanks.

> > Questions:
> > 1) Does the large file size have anything to do with the printer
> > itself? Is there anything I can do (ex:- update the driver/firmware or
> > something)?
>
> That depends on what is encoding the images: does the scanner itself
> "make" the PDF? Or some software, computer-side?
>

The scanner itself makes the pdf files.

> > 2) Is the difference in image sizes due to the bpc (1 vs. 8) or
> > encoding (ccitt vs jped) fields?
>
> CCITT vs JPEG, yes.
>

ok. What about bpc? Does that matter for file size?

> > 3) If yes, how to change them?
>
> Hmmm. I don't know yet whether you have to talk to your scanner
> or to your scan software...

I think there is not much I can do from the scanner's settings. But I
was wondering if I can use some software (like convert, gs etc.,) to
change the encoding, bpc etc., and reduce the file sizes without
sacrificing too much on the quality. To be honest, both files look the
same when viewed side by side.

-- 
Kamaraju S Kusumanchi | http://raju.shoutwiki.com/wiki/Blog

[toc] | [prev] | [next] | [standalone]


#203789

Fromtomas@tuxteam.de
Date2019-01-02 10:40 +0100
Message-ID<xbLah-4VR-5@gated-at.bofh.it>
In reply to#203787

[Multipart message — attachments visible in raw view] — view raw

On Tue, Jan 01, 2019 at 10:41:06PM -0500, kamaraju kusumanchi wrote:
> On Tue, Jan 1, 2019 at 1:40 PM <tomas@tuxteam.de> wrote:

[...]

> > OTOH, CCITT is lossless and JPEG lossy ;-)
> >
> ok, thanks.

But note that Anders observed that the larger files are actually
the JPEGs (somewhat to my surprise). A possible explanation would
be that the CCITT loses the grayscale information, as Thomas observed
(CCITT is bitonal), but then you should be able to see a difference
between both images.

In any case, if the scanners insist on producing PDFs, you'll have
to extract the images, convert them and, if necessary, repack them
again as PDFs. For an one-off job, the Gimp seems just about right;
if you want to automate it, I'd try with the ImageMagick suite, but
perhaps there are folks around who have more experience in those
things.

And next time, try to find a scanner which provides you with a raw
image. Wrapping images in PDFs is... not elegant.

Cheers
-- tomás

[toc] | [prev] | [next] | [standalone]


#203790

FromJoe <joe@jretrading.com>
Date2019-01-02 11:10 +0100
Message-ID<xbLDj-5lA-5@gated-at.bofh.it>
In reply to#203789
On Wed, 2 Jan 2019 09:59:48 +0100
tomas@tuxteam.de wrote:


> 
> And next time, try to find a scanner which provides you with a raw
> image. Wrapping images in PDFs is... not elegant.

They do this to cater for multiple pages, whereas in my experience,
most scanning is single-sheet. Even the Simple Scan program on Debian
defaults to pdf, something which cannot be configured.

-- 
Joe

[toc] | [prev] | [next] | [standalone]


#203791

From<tomas@tuxteam.de>
Date2019-01-02 11:30 +0100
Message-ID<xbLWG-5sg-7@gated-at.bofh.it>
In reply to#203790

[Multipart message — attachments visible in raw view] — view raw

On Wed, Jan 02, 2019 at 09:40:33AM +0000, Joe wrote:
> On Wed, 2 Jan 2019 09:59:48 +0100
> tomas@tuxteam.de wrote:
> 
> 
> > 
> > And next time, try to find a scanner which provides you with a raw
> > image. Wrapping images in PDFs is... not elegant.
> 
> They do this to cater for multiple pages, whereas in my experience,
> most scanning is single-sheet. Even the Simple Scan program on Debian
> defaults to pdf, something which cannot be configured.

I get this, and offering that option seems to make sense. But forcing
it (and forcing an image format like JPEG) doesn't make sense. So either
provide the knobs or let the host software do it.

My scanner just transfers the raw image. The scan program is responsible
for the transformation to the target format, which I can choose. This is
how /I/ want to be treated, as a paying customer.

Cheers
-- tomás

[toc] | [prev] | [next] | [standalone]


#203793

Fromdeloptes <deloptes@gmail.com>
Date2019-01-02 12:10 +0100
Message-ID<xbMzn-5Vl-1@gated-at.bofh.it>
In reply to#203791
tomas@tuxteam.de wrote:

> I get this, and offering that option seems to make sense. But forcing
> it (and forcing an image format like JPEG) doesn't make sense. So either
> provide the knobs or let the host software do it.
> 
> My scanner just transfers the raw image. The scan program is responsible
> for the transformation to the target format, which I can choose. This is
> how /I/ want to be treated, as a paying customer.

+1

I did some research before buying a scanner and I bought Epson Perfection
33-330. The iscan application let you choose the image format. I am not
quite sure what it uses if I say multiple pages in one PDF. I will try
this, but I assume it acts based on the configuration before scanning (i.e.
grayscale or black/white etc.)

regards

[toc] | [prev] | [next] | [standalone]


#203794

FromChris Ramsden <chris.ramsden@gmail.com>
Date2019-01-02 12:10 +0100
Message-ID<xbMzn-5Vl-3@gated-at.bofh.it>
In reply to#203791
On 2019-01-02 10:24, tomas@tuxteam.de wrote:
> My scanner just transfers the raw image. The scan program is responsible
> for the transformation to the target format, which I can choose. This is
> how /I/ want to be treated, as a paying customer.
>
> Cheers
> -- tomás

Would you mind sharing with us what make and model you use? I haven't
found it easy to identify which scanners offer this feature.

-- 
Chris

[toc] | [prev] | [next] | [standalone]


#203797

From<tomas@tuxteam.de>
Date2019-01-02 12:40 +0100
Message-ID<xbN2p-659-9@gated-at.bofh.it>
In reply to#203794

[Multipart message — attachments visible in raw view] — view raw

On Wed, Jan 02, 2019 at 11:04:07AM +0000, Chris Ramsden wrote:
> On 2019-01-02 10:24, tomas@tuxteam.de wrote:
> > My scanner just transfers the raw image. The scan program is responsible
> > for the transformation to the target format, which I can choose. This is
> > how /I/ want to be treated, as a paying customer.
> >
> > Cheers
> > -- tomás
> 
> Would you mind sharing with us what make and model you use? I haven't
> found it easy to identify which scanners offer this feature.

Sorry, I don't have access to it at the moment. I'll try to look it
up when I'm back home.

Anyway, it's now over 12 years old -- I doubt it's still on the
market. At that time I looked at the SANE databases [1] to help
me make a decision, together with my (then) computer dealer (ah,
I miss him, he knew what matters: having a computer dealer you
trust is worth gold).

Cheers
[1] http://sane-project.org/
-- tomás

[toc] | [prev] | [next] | [standalone]


#203825

Frommick crane <mick.crane@gmail.com>
Date2019-01-02 20:20 +0100
Message-ID<xbUdA-2bP-5@gated-at.bofh.it>
In reply to#203791
On 2019-01-02 10:24, tomas@tuxteam.de wrote:
> On Wed, Jan 02, 2019 at 09:40:33AM +0000, Joe wrote:
>> On Wed, 2 Jan 2019 09:59:48 +0100
>> tomas@tuxteam.de wrote:
>> 
>> 
>> >
>> > And next time, try to find a scanner which provides you with a raw
>> > image. Wrapping images in PDFs is... not elegant.
>> 
>> They do this to cater for multiple pages, whereas in my experience,
>> most scanning is single-sheet. Even the Simple Scan program on Debian
>> defaults to pdf, something which cannot be configured.
> 
> I get this, and offering that option seems to make sense. But forcing
> it (and forcing an image format like JPEG) doesn't make sense. So 
> either
> provide the knobs or let the host software do it.
> 
> My scanner just transfers the raw image. The scan program is 
> responsible
> for the transformation to the target format, which I can choose. This 
> is
> how /I/ want to be treated, as a paying customer.
> 

having a scanner do PDFs is weird, see Obama birth certificate, how do 
you know is a faithful copy ?
A piece of paper with marks on it is an image and should be treated as 
such.

mick



-- 
Key ID    4BFEBB31

[toc] | [prev] | [next] | [standalone]


#203826

FromBrian <ad44@cityscape.co.uk>
Date2019-01-02 20:30 +0100
Message-ID<xbUnf-2f0-3@gated-at.bofh.it>
In reply to#203825
On Wed 02 Jan 2019 at 19:17:00 +0000, mick crane wrote:

> On 2019-01-02 10:24, tomas@tuxteam.de wrote:
> > On Wed, Jan 02, 2019 at 09:40:33AM +0000, Joe wrote:
> > > On Wed, 2 Jan 2019 09:59:48 +0100
> > > tomas@tuxteam.de wrote:
> > > 
> > > 
> > > >
> > > > And next time, try to find a scanner which provides you with a raw
> > > > image. Wrapping images in PDFs is... not elegant.
> > > 
> > > They do this to cater for multiple pages, whereas in my experience,
> > > most scanning is single-sheet. Even the Simple Scan program on Debian
> > > defaults to pdf, something which cannot be configured.
> > 
> > I get this, and offering that option seems to make sense. But forcing
> > it (and forcing an image format like JPEG) doesn't make sense. So either
> > provide the knobs or let the host software do it.
> > 
> > My scanner just transfers the raw image. The scan program is responsible
> > for the transformation to the target format, which I can choose. This is
> > how /I/ want to be treated, as a paying customer.
> > 
> 
> having a scanner do PDFs is weird, see Obama birth certificate, how do you
> know is a faithful copy ?

Eh?

> A piece of paper with marks on it is an image and should be treated as such.

Egad. I wish I had thought of that.

-- 
Brian.

[toc] | [prev] | [next] | [standalone]


#203842

FromDavid Wright <deblis@lionunicorn.co.uk>
Date2019-01-03 03:50 +0100
Message-ID<xc1f3-6mc-1@gated-at.bofh.it>
In reply to#203825
On Wed 02 Jan 2019 at 19:17:00 (+0000), mick crane wrote:
> On 2019-01-02 10:24, tomas@tuxteam.de wrote:
> > On Wed, Jan 02, 2019 at 09:40:33AM +0000, Joe wrote:
> > > On Wed, 2 Jan 2019 09:59:48 +0100 tomas@tuxteam.de wrote:
> > > > And next time, try to find a scanner which provides you with a raw
> > > > image. Wrapping images in PDFs is... not elegant.
> > > 
> > > They do this to cater for multiple pages, whereas in my experience,
> > > most scanning is single-sheet. Even the Simple Scan program on Debian
> > > defaults to pdf, something which cannot be configured.
> > 
> > I get this, and offering that option seems to make sense. But forcing
> > it (and forcing an image format like JPEG) doesn't make sense. So
> > either
> > provide the knobs or let the host software do it.
> > 
> > My scanner just transfers the raw image. The scan program is
> > responsible
> > for the transformation to the target format, which I can choose.
> > This is
> > how /I/ want to be treated, as a paying customer.
> 
> having a scanner do PDFs is weird, see Obama birth certificate, how do
> you know is a faithful copy ?
> A piece of paper with marks on it is an image and should be treated as
> such.

If I could be bothered to look, I might be able to come across a raw
sound or video file on this Debian system. It's quite normal to wrap
such raw data in a container format of some sort.

So I can't understand your objection to wrapping a scanned image into
a PDF container, which makes a lot of data handling a lot easier than
would otherwise be the case. An obvious example was already mentioned:
put a document into the ADF, press the button, obtain one file
containing the entire document. Other examples would be postprocessing
with programs like pdftk and pdfjam. Would you really send a scanned
document to a company/institution as a multitude of image attachments
instead of a single PDF?

If you want the image back from a PDF, that's what the pdfimages
program is for. I would assume that tomás is pleased with the
packaging of the raw scan data into an image format of some
description, so what's the difference?

Cheers,
David.

[toc] | [prev] | [next] | [standalone]


#203862

FromSiard <shiems146@kpnplanet.nl>
Date2019-01-03 14:10 +0100
Message-ID<xcaV4-43B-5@gated-at.bofh.it>
In reply to#203842
David Wright wrote:
> So I can't understand your objection to wrapping a scanned image into
> a PDF container, which makes a lot of data handling a lot easier than
> would otherwise be the case.

After scanning, an image almost always needs editing. Crop, rotate to
correct a skew horizon, remove specks, adjust light and contrast.
Gimp can open a pdf, but not in its original resolution, so there is
loss of quality.  Pdfimages can extract the image first, but its
original format (tiff? jpg? pnm?) remains unclear then, so there is a
conversion, again causing loss of quality (AFAIU).

> Other examples would be postprocessing with programs like pdftk and
> pdfjam.

Those programs cannot edit images.

> An obvious example was already mentioned: put a document into the
> ADF, press the button, obtain one file containing the entire
> document. [...] Would you really send a scanned document to a
> company/institution as a multitude of image attachments instead of
> a single PDF?

That should be the final stage of the process, not the beginning!
You can use img2pdf to put the images in a pdf container, without
affecting the image quality.

[toc] | [prev] | [next] | [standalone]


#203863

FromNicolas George <george@nsup.org>
Date2019-01-03 14:20 +0100
Message-ID<xcb4J-46O-9@gated-at.bofh.it>
In reply to#203862

[Multipart message — attachments visible in raw view] — view raw

Siard (2019-01-03):
> After scanning, an image almost always needs editing. Crop, rotate to
> correct a skew horizon, remove specks, adjust light and contrast.

That depends on the purpose.

>		    Pdfimages can extract the image first, but its
> original format (tiff? jpg? pnm?) remains unclear then, so there is a
> conversion, again causing loss of quality (AFAIU).

You are mistaken. If the image is stored losslessly in the PDF, then
extracting it to any lossless format will not lose any quality. If the
image is stored with a lossy codec, pdfimages can be directed (RTFM) to
re-wrap the codec data in an adequate container, causing no extra loss
of quality either.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#203866

Fromdeloptes <deloptes@gmail.com>
Date2019-01-03 15:20 +0100
Message-ID<xcc0N-4EZ-5@gated-at.bofh.it>
In reply to#203862
Siard wrote:

> After scanning, an image almost always needs editing. Crop, rotate to
> correct a skew horizon, remove specks, adjust light and contrast.
> Gimp can open a pdf, but not in its original resolution, so there is
> loss of quality.  Pdfimages can extract the image first, but its
> original format (tiff? jpg? pnm?) remains unclear then, so there is a
> conversion, again causing loss of quality (AFAIU).

Don't know about you, but I usually press the button and get the image.
Sometimes I use PDF to scan multiple pages into one document, sometimes I
scan in PNG or JPG and create a PDF out of them.

As a user I do not want to spend time for correction. Put the image in the
scanner the way you want it to be scanned, select the options on software
side and just press a button.

[toc] | [prev] | [next] | [standalone]


#203869

FromSiard <shiems146@kpnplanet.nl>
Date2019-01-03 16:00 +0100
Message-ID<xccDv-4RQ-7@gated-at.bofh.it>
In reply to#203866
deloptes wrote:
> Siard wrote:
> > After scanning, an image almost always needs editing. Crop, rotate
> > to correct a skew horizon, remove specks, adjust light and contrast.
> 
> Don't know about you, but I usually press the button and get the
> image. Sometimes I use PDF to scan multiple pages into one document,
> sometimes I scan in PNG or JPG and create a PDF out of them.
> 
> As a user I do not want to spend time for correction. Put the image
> in the scanner the way you want it to be scanned, select the options
> on software side and just press a button.

Very different here. I scan from within Gimp:
File > Create > XSane > Device dialog...
Then the image scanned with XSane opens directly in Gimp.

[toc] | [prev] | [next] | [standalone]


#203875

FromNicolas George <george@nsup.org>
Date2019-01-03 17:30 +0100
Message-ID<xce2C-5Pg-11@gated-at.bofh.it>
In reply to#203869

[Multipart message — attachments visible in raw view] — view raw

deloptes (2019-01-03):
> If it is a document, why should I open it in Gimp?

The level of "my use case is the only use case" in this subthread is
frightening and staggering.

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | linux.debian.user


csiph-web