Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #270518 > unrolled thread

Re: Backup.

Started byEduardo M KALINOWSKI <eduardo@kalinowski.com.br>
First post2024-06-27 21:10 +0200
Last post2024-06-30 21:20 +0200
Articles 6 — 6 participants

Back to article view | Back to linux.debian.user

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: Backup. Eduardo M KALINOWSKI <eduardo@kalinowski.com.br> - 2024-06-27 21:10 +0200
    Re: Backup. Michael Kjörling <c9bc136c6063@ewoof.net> - 2024-06-27 22:00 +0200
    Re: Backup. Paul M Foster <paulf@quillandmouse.com> - 2024-06-27 22:50 +0200
    Re (3): Backup. peter@easthope.ca - 2024-06-30 17:00 +0200
      Re: Re (3): Backup. Andy Smith <andy@strugglers.net> - 2024-06-30 17:40 +0200
        Re: Re (3): Backup. David Christensen <dpchrist@holgerdanske.com> - 2024-06-30 21:20 +0200

#270518 — Re: Backup.

FromEduardo M KALINOWSKI <eduardo@kalinowski.com.br>
Date2024-06-27 21:10 +0200
SubjectRe: Backup.
Message-ID<IU2Fj-62Pv-3@gated-at.bofh.it>
On 27/06/2024 15:23, peter@easthope.ca wrote:
> Now I have a pair of 500 GB external USB drives.  Large compared to my
> working data of ~3 GB.  Please suggest improvements to my backup
> system by exploiting these drives.  I can imagine a complete copy of A
> onto an external drive for each backup; but with most files in A not
> changing during the backup interval, that is inefficient.

rnapshot


-- 
Mike:	"The Fourth Dimension is a shambles?"
Bernie:	"Nobody ever empties the ashtrays.  People are SO inconsiderate."
		-- Gary Trudeau, "Doonesbury"

Eduardo M KALINOWSKI
eduardo@kalinowski.com.br

[toc] | [next] | [standalone]


#270520

FromMichael Kjörling <c9bc136c6063@ewoof.net>
Date2024-06-27 22:00 +0200
Message-ID<IU3rH-636w-1@gated-at.bofh.it>
In reply to#270518
On 27 Jun 2024 16:06 -0300, from eduardo@kalinowski.com.br (Eduardo M KALINOWSKI):
>> Now I have a pair of 500 GB external USB drives.  Large compared to my
>> working data of ~3 GB.  Please suggest improvements to my backup
>> system by exploiting these drives.  I can imagine a complete copy of A
>> onto an external drive for each backup; but with most files in A not
>> changing during the backup interval, that is inefficient.
> 
> rnapshot

Yes, rsnapshot.

Which is essentially a front-end to rsync --link-dest; so, for
mostly-static data, very efficient.

-- 
Michael Kjörling                     🔗 https://michael.kjorling.se
“Remember when, on the Internet, nobody cared that you were a dog?”

[toc] | [prev] | [next] | [standalone]


#270524

FromPaul M Foster <paulf@quillandmouse.com>
Date2024-06-27 22:50 +0200
Message-ID<IU4e6-63Cw-5@gated-at.bofh.it>
In reply to#270518
On Thu, Jun 27, 2024 at 04:06:18PM -0300, Eduardo M KALINOWSKI wrote:

> On 27/06/2024 15:23, peter@easthope.ca wrote:
> > Now I have a pair of 500 GB external USB drives.  Large compared to my
> > working data of ~3 GB.  Please suggest improvements to my backup
> > system by exploiting these drives.  I can imagine a complete copy of A
> > onto an external drive for each backup; but with most files in A not
> > changing during the backup interval, that is inefficient.
> 
> rnapshot

Rsnapshot is written in Perl and is based on an article:
http://www.mikerubel.org/computers/rsync_snapshots/

I read that article a while back, and not knowing of the existence of
rsnapshot, I modified my existing bash/rsync backup script to use this
guy's methods. In essence, rather than backing up the same file again with
rsync, you just make a hard link to it. All "copies" of that file on the
backup link to that one copy. Saves a tremendous amount of room. I keep 7
backup directories on my backup drives (one for each day), and a cron job
fires of the backup each day. 

You can take a look at the script at:

https://gitlab.com/paulmfoster/bkp

Paul

-- 
Paul M. Foster
Personal Blog: http://noferblatz.com
Company Site: http://quillandmouse.com
Software Projects: https://gitlab.com/paulmfoster

[toc] | [prev] | [next] | [standalone]


#270658 — Re (3): Backup.

Frompeter@easthope.ca
Date2024-06-30 17:00 +0200
SubjectRe (3): Backup.
Message-ID<IV4c1-6HNA-3@gated-at.bofh.it>
In reply to#270518
    From: Eduardo M KALINOWSKI <eduardo@kalinowski.com.br>
    Date: Thu, 27 Jun 2024 16:06:18 -0300
> rnapshot

>From https://rsnapshot.org/
> rsnapshot is a filesystem snapshot utility ...

Rather than a snapshot of the extant file system, I want to keep a 
history of the files in the file system.

Thanks,                                ... P.

-- 
VoIP:   +1 604 670 0140
work: https://en.wikibooks.org/wiki/User:PeterEasthope

[toc] | [prev] | [next] | [standalone]


#270663 — Re: Re (3): Backup.

FromAndy Smith <andy@strugglers.net>
Date2024-06-30 17:40 +0200
SubjectRe: Re (3): Backup.
Message-ID<IV4OK-6Ige-29@gated-at.bofh.it>
In reply to#270658
Hi,

On Sun, Jun 30, 2024 at 07:36:58AM -0700, peter@easthope.ca wrote:
> >From https://rsnapshot.org/
> > rsnapshot is a filesystem snapshot utility ...
> 
> Rather than a snapshot of the extant file system, I want to keep a 
> history of the files in the file system.

You should read more than one line of a page. That is exactly what
it is intended for. Snapshots become history when you keep multiple
of them.

I have used rsnapshot a lot (decades worth of use) and it's good but
it is not perfect (nothing is). It is probably a much better backup
system than anything one can typically come up with by hand at short
order, but here are some of its downsides:

- No built in compression or encryption. You can implement these
  yourself using filesystem features.

- Since it uses hardlinks for deduplication, this brings with it
  some inherent limitations:

  - The filesystem you use must support hardlinks

  - All versions of a file will have the same metadata (mtime,
    permissions, ownership, etc) because hardlinks must have the
    same metadata. As a consequence, any change of metadata will
    result in two separate files being stored (not hardlinked
    together) in order to represent that change. Even if the files
    have identical content.

  - Changing one byte of a file results in the storage of two
    separate full copies of the two versions of the file. With
    hardlinks either the file is entirely the same or it needs to
    not be a hardlink. This makes rsnapshot and things like it
    particularly bad for backing up large append-only files like log
    files.

- rsnapshot only compares versions of a file at the same path and
  point in time. So for example /path/to/foo is only ever compared
  against /path/to/foo *from the previous backup run*. Other copies
  of foo anywhere else on the system being backed up, or from other
  systems being backed up, or from a backup run previous to the most
  recent, will not be considered so will not be hardlinked together.

  A typical system has a lot of duplicate files and once you start
  backing up multiple systems there tends to be an explosion of
  duplicate data. rsnapshot will not handle any of this specially
  and will just store it all.

  It is possible to improve this by for example running an external
  deduplication tool over the backups, or using deduplication
  facilities of a filesystem like zfs¹. This must be done carefully
  otherwise the workings of rsnapshot can be disrupted.

- rsnapshot must walk through the entire previous backup to compare
  all the content of the files to the content of the new files. This
  is quite expensive and will involve tons of random seeks which is
  a killer for rotational storage media. Once you get to several
  million inodes in a backup run, you may find a run of rsnapshot
  taking several hours.

On the other hand, rsnapshot's huge plus point is that everything is
stored in a tree of files and hardlinks so it can just be explored
and restored with normal filesystem tools. You don't need any part
of rsnapshot to access and restore your content. That is such a good
feature that many people feel able to overlook the negatives.

More featureful backup systems chunk backup content up and store it
by a has of its content, which tends to bring advantages like:

- Never needing to store the same chunk twice no matter where (or
  when) it came from

- Easy to compress and encrypt

- Locating which data is in which chunk gets done by a database,
  not by random access to a filesystem, so it's much faster. When
  you say "I want /path/to/foo from a week ago, but also show me
  every copy you have going back 3 years", that is a database query,
  not a walk of a filesystem with potentially several million inodes
  in it.

But, by doing that you lose the ability to just cp a file from your
backups.

Thanks,
Andy

¹ Though someone heavily in to an advanced filesystem like zfs may
  be more inclined to take advantage of zfs's proper snapshot
  capabilities (and zfs-send to move them off-site) than use
  rsnapshot on it.

-- 
https://bitfolk.com/ -- No-nonsense VPS hosting

[toc] | [prev] | [next] | [standalone]


#270672 — Re: Re (3): Backup.

FromDavid Christensen <dpchrist@holgerdanske.com>
Date2024-06-30 21:20 +0200
SubjectRe: Re (3): Backup.
Message-ID<IV8fD-6Krr-7@gated-at.bofh.it>
In reply to#270663
On 6/30/24 08:37, Andy Smith wrote:
<snip>


Thank you for that informative discussion of rsnapshot(1) and related.  :-)


My initial reaction to this thread was to recommend Preston [1].  I 
still think that is decent advice; both for noobs and for experienced 
people who missed it.


David


[1]  Preston, W., 2007, "Backup & Recovery", O'Reilly Media, Inc.
ISBN: 9780596102463, 
https://www.oreilly.com/library/view/backup-recovery/0596102461/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web