Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.compression > #1661

Re: new format .gz ?

From glen herrmannsfeldt <gah@ugcs.caltech.edu>
Newsgroups comp.compression
Subject Re: new format .gz ?
Date 2012-12-25 13:52 +0000
Organization Aioe.org NNTP Server
Message-ID <kbcb2s$ijp$1@speranza.aioe.org> (permalink)
References <ee881e51-fb1e-4699-83b9-4ab560c4f9aa@googlegroups.com> <f7576697-f420-4b2d-9d07-2dc9c03f02cd@googlegroups.com>

Show all headers | View raw


sterten@aol.com wrote:
> I don't know about BAM, I'm pretty new to this.
> These here were "vcf" files, normal text files , tables ,
> with ~4000 binary columns per ~100000 rows.
> So, basically a 4000*100000 binary array with some (partial) 
> rows and columns being similar especially if they are close.

Does look pretty wordy for something that is supposed to reduce
storage space.
 
> Compressing these could become important with the
> _lots_ of genetical data expected in the coming years

Yes. There is a lot of redundancy in many genetic data files that
compression algorithms aren't good at finding. 
 
> Their ftp site has already >400TB so you can understand
> why they need good compression.

DNA sequence data should compress to two bits per base with many
compression algorithms. If you add quality data, though, it will
be much worse. 

-- glen

Back to comp.compression | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

new format .gz ? sterten@aol.com - 2012-12-24 04:37 -0800
  Re: new format .gz ? sterten@aol.com - 2012-12-24 05:17 -0800
    Re: new format .gz ? sterten@aol.com - 2012-12-24 22:34 -0800
      Re: new format .gz ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2012-12-25 08:24 +0000
  Re: new format .gz ? sterten@aol.com - 2012-12-25 02:40 -0800
    Re: new format .gz ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2012-12-25 13:52 +0000
  Re: new format .gz ? sterten@aol.com - 2013-01-26 07:24 -0800

csiph-web