Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #185395 > unrolled thread

What tool can I use to make efficient incremental backups?

Started byMario Castelán Castro <marioxcc.MT@yandex.com>
First post2017-08-17 18:50 +0200
Last post2017-09-17 20:40 +0200
Articles 20 on this page of 22 — 9 participants

Back to article view | Back to linux.debian.user


Contents

  What tool can I use to make efficient incremental backups? Mario Castelán Castro <marioxcc.MT@yandex.com> - 2017-08-17 18:50 +0200
    Re: What tool can I use to make efficient incremental backups? Fungi4All <fungilife@protonmail.com> - 2017-08-17 19:20 +0200
      Re: What tool can I use to make efficient incremental backups? Mario Castelán Castro <marioxcc.MT@yandex.com> - 2017-08-17 19:40 +0200
    Re: What tool can I use to make efficient incremental backups? Nicolas George <george@nsup.org> - 2017-08-17 20:10 +0200
      Re: What tool can I use to make efficient incremental backups? Mario Castelán Castro <marioxcc.MT@yandex.com> - 2017-08-17 20:30 +0200
        Re: What tool can I use to make efficient incremental backups? Nicolas George <george@nsup.org> - 2017-08-17 20:50 +0200
          Re: What tool can I use to make efficient incremental backups? Mario Castelán Castro <marioxcc.MT@yandex.com> - 2017-08-17 22:30 +0200
            Re: What tool can I use to make efficient incremental backups? <tomas@tuxteam.de> - 2017-08-17 23:00 +0200
              Re: What tool can I use to make efficient incremental backups? Mario Castelán Castro <marioxcc.MT@yandex.com> - 2017-08-18 02:40 +0200
                Debian-user conventions [was: [...] incremental backups?] <tomas@tuxteam.de> - 2017-08-18 09:40 +0200
    Re: What tool can I use to make efficient incremental backups? Liam O'Toole <liam.p.otoole@gmail.com> - 2017-08-19 01:10 +0200
      Re: What tool can I use to make efficient incremental backups? Mario Castelán Castro <marioxcc.MT@yandex.com> - 2017-08-19 17:10 +0200
    Re: What tool can I use to make efficient incremental backups? Celejar <celejar@gmail.com> - 2017-08-20 05:10 +0200
      Re: What tool can I use to make efficient incremental backups? Gene Heskett <gheskett@shentel.net> - 2017-08-20 08:10 +0200
        Re: What tool can I use to make efficient incremental backups? Glenn English <ghe2001@gmail.com> - 2017-08-20 21:40 +0200
          Re: What tool can I use to make efficient incremental backups? Mario Castelán Castro <marioxcc.MT@yandex.com> - 2017-08-21 02:00 +0200
        Re: What tool can I use to make efficient incremental backups? Celejar <celejar@gmail.com> - 2017-08-22 05:50 +0200
          Re: What tool can I use to make efficient incremental backups? Gene Heskett <gheskett@shentel.net> - 2017-08-22 07:30 +0200
            Re: What tool can I use to make efficient incremental backups? Celejar <celejar@gmail.com> - 2017-08-22 15:20 +0200
      Re: What tool can I use to make efficient incremental backups? Mario Castelán Castro <marioxcc.MT@yandex.com> - 2017-08-21 03:10 +0200
        Re: What tool can I use to make efficient incremental backups? Celejar <celejar@gmail.com> - 2017-09-17 03:40 +0200
        Re: What tool can I use to make efficient incremental backups? Kushal Kumaran <kushal@locationd.net> - 2017-09-17 20:40 +0200

Page 1 of 2  [1] 2  Next page →


#185395 — What tool can I use to make efficient incremental backups?

FromMario Castelán Castro <marioxcc.MT@yandex.com>
Date2017-08-17 18:50 +0200
SubjectWhat tool can I use to make efficient incremental backups?
Message-ID<ufw9z-7up-7@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

Hello.

Currently I use rsync to make the backups of my personal data, including
some manually selected important files of system configuration. I keep
old backups to be more safe from the scenario where I have deleted
something important, I make a backup, and I only notice the deletion
afterwards.

Each backup snapshot is stored in its own directory. There is much
redundancy between subsequent backups. I use the option "--link-dest" to
make hard links and thus save space for files that are *identical* to an
already-existing file in the backup repository. but this is still
inefficient. Any change to a file, even to its metadata (permission,
modification time, etc.), will result in the file being saved at whole,
instead of a delta.

Can you suggest a more efficient alternative?

I know about bup <https://github.com/bup/bup> but I have not used it
because it warns that “This is a very early version. Therefore it will
most probably not work for you, but we don't know why. It is also
missing some probably-critical features.”.

I also know about obnam. Unfortunately, the main author it has been
announced that it will be unmaintained because it has become a piece of
engineering, with all the ugly consequences of that, and real
engineering is “not fun” for him.

Thanks.

[toc] | [next] | [standalone]


#185400

FromFungi4All <fungilife@protonmail.com>
Date2017-08-17 19:20 +0200
Message-ID<ufwCC-7Vp-17@gated-at.bofh.it>
In reply to#185395

[Multipart message — attachments visible in raw view] — view raw

> From: marioxcc.MT@yandex.com
> To: debian-user <debian-user@lists.debian.org>
>
> Hello.
>
> Currently I use rsync to make the backups of my personal data, including
> some manually selected important files of system configuration. I keep
> old backups to be more safe from the scenario where I have deleted
> something important, I make a backup, and I only notice the deletion
> afterwards.
>
> Each backup snapshot is stored in its own directory. There is much
> redundancy between subsequent backups. I use the option "--link-dest" to
> make hard links and thus save space for files that are *identical* to an
> already-existing file in the backup repository. but this is still
> inefficient. Any change to a file, even to its metadata (permission,
> modification time, etc.), will result in the file being saved at whole,
> instead of a delta.
>
> Can you suggest a more efficient alternative?
>
> I know about bup <https://github.com/bup/bup> but I have not used it
> because it warns that “This is a very early version. Therefore it will
> most probably not work for you, but we don"t know why. It is also
> missing some probably-critical features.”.
>
> I also know about obnam. Unfortunately, the main author it has been
> announced that it will be unmaintained because it has become a piece of
> engineering, with all the ugly consequences of that, and real
> engineering is “not fun” for him.
>
> Thanks.

Stay with rsync

[toc] | [prev] | [next] | [standalone]


#185405

FromMario Castelán Castro <marioxcc.MT@yandex.com>
Date2017-08-17 19:40 +0200
Message-ID<ufwVX-82Z-3@gated-at.bofh.it>
In reply to#185400

[Multipart message — attachments visible in raw view] — view raw

On 17/08/17 12:10, Fungi4All wrote:
> [[elided]]
> Stay with rsync

Why? Isn't there a more efficient alternative?

[toc] | [prev] | [next] | [standalone]


#185409

FromNicolas George <george@nsup.org>
Date2017-08-17 20:10 +0200
Message-ID<ufxp0-8vo-19@gated-at.bofh.it>
In reply to#185395
Le decadi 30 thermidor, an CCXXV, Mario Castelán Castro a écrit :
> Currently I use rsync to make the backups of my personal data, including
> some manually selected important files of system configuration. I keep
> old backups to be more safe from the scenario where I have deleted
> something important, I make a backup, and I only notice the deletion
> afterwards.
> 
> Each backup snapshot is stored in its own directory. There is much
> redundancy between subsequent backups. I use the option "--link-dest" to
> make hard links and thus save space for files that are *identical* to an
> already-existing file in the backup repository. but this is still
> inefficient. Any change to a file, even to its metadata (permission,
> modification time, etc.), will result in the file being saved at whole,
> instead of a delta.
> 
> Can you suggest a more efficient alternative?

We used a similar setup on a server, using the rsnapshot script. But we
have users with huge mbox files that were copied entirely each time. We
changed for a setup with normal rsync (no --link-dest) and btrfs
snapshots, it increased the efficiency (storage and disk bandwidth)
dramatically.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#185411

FromMario Castelán Castro <marioxcc.MT@yandex.com>
Date2017-08-17 20:30 +0200
Message-ID<ufxIl-bd-7@gated-at.bofh.it>
In reply to#185409

[Multipart message — attachments visible in raw view] — view raw

Thanks for your answer.

Let me know if I understood your approach correctly. You have a
directory in a btrfs filesystem that is the target of your backups. When
you make a backup, you take a brtfs snapshot of this directory and
*then* use rsync. Is this correct?

Regards.

On 17/08/17 12:50, Nicolas George wrote:
> [[elided]]
> 
> We used a similar setup on a server, using the rsnapshot script. But we
> have users with huge mbox files that were copied entirely each time. We
> changed for a setup with normal rsync (no --link-dest) and btrfs
> snapshots, it increased the efficiency (storage and disk bandwidth)
> dramatically.
> 
> Regards,

[toc] | [prev] | [next] | [standalone]


#185412

FromNicolas George <george@nsup.org>
Date2017-08-17 20:50 +0200
Message-ID<ufy1I-iq-15@gated-at.bofh.it>
In reply to#185411
Le decadi 30 thermidor, an CCXXV, Mario Castelán Castro a écrit :
> Let me know if I understood your approach correctly. You have a
> directory in a btrfs filesystem that is the target of your backups. When
> you make a backup, you take a brtfs snapshot of this directory and
> *then* use rsync. Is this correct?

No, it is the other way around: we rsync the data to a directory stored
on a btrfs filesystem, and then we make a snapshot of that directory.
With btrfs's CoW, only the parts of the files that have changed use
space.

Hum, I think it requires the --inplace option.

> On 17/08/17 12:50, Nicolas George wrote:

Please remember not to top-post.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#185414

FromMario Castelán Castro <marioxcc.MT@yandex.com>
Date2017-08-17 22:30 +0200
Message-ID<ufzAu-1r4-17@gated-at.bofh.it>
In reply to#185412

[Multipart message — attachments visible in raw view] — view raw

On 17/08/17 13:31, Nicolas George wrote:
> [[elided]]
> 
> No, it is the other way around: we rsync the data to a directory stored
> on a btrfs filesystem, and then we make a snapshot of that directory.
> With btrfs's CoW, only the parts of the files that have changed use
> space.

Thanks for the clarification.

> Please remember not to top-post.

Both bottom posting and top posting each have their own disadvantages.
Bottom posting requires scrolling past text that may be not needed. Top
posting puts the messages in reverse chronological order, which is not
something bad by itself.

When I explicitly want a quote to reply to a specific parts of a
message, I post after the parts, as in this message; I don't know if
that would still be considered bottom posting. When I am including the
previous message *only* for reference, I use top posting because the
previous message is also archived in the inbox of the other users, so
the quotation included for reference is of secondary importance, and
therefore IMO should go after the *important* (new) information, that
is, at the bottom.

Is there a rule, guideline or de-facto standard mandating either style
in debian-user?

[toc] | [prev] | [next] | [standalone]


#185417

From<tomas@tuxteam.de>
Date2017-08-17 23:00 +0200
Message-ID<ufA3w-1Eh-27@gated-at.bofh.it>
In reply to#185414
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Thu, Aug 17, 2017 at 03:24:35PM -0500, Mario Castelán Castro wrote:
> On 17/08/17 13:31, Nicolas George wrote:

[...]

> > Please remember not to top-post.
> 
> Both bottom posting and top posting each have their own disadvantages.

The general convention here is to only quote the relevant parts you
are replying to. For long threads this tends to work better.

Worst is, of course, mixing styles :-)

But in general, folks here tend to be tolerant. And yes, there's a
wiki entry encouraging "in-line" quoting [1].

> Bottom posting requires scrolling past text that may be not needed. Top
> posting puts the messages in reverse chronological order, which is not
> something bad by itself.

That's why you shoulnd't include the whole message, but snip the
relevant parts you are answering to. Believe me, for long threads,
this tends to work best.
 
> When I explicitly want a quote to reply to a specific parts of a
> message, I post after the parts, as in this message; I don't know if
> that would still be considered bottom posting.

Yes, exactly. If someone needs the unabridged original post, it's either
in her mailbox or in the archives.

>                                           When I am including the
> previous message *only* for reference, I use top posting because the
> previous message is also archived in the inbox of the other users, so
> the quotation included for reference is of secondary importance, and
> therefore IMO should go after the *important* (new) information, that
> is, at the bottom.

But that's why the original message isn't needed as a copy in the
first place.

> Is there a rule, guideline or de-facto standard mandating either style
> in debian-user?

See [1] (there are also other hints on that page). Also [2] is a good
reference.

[1] https://wiki.debian.org/FAQsFromDebianUser#What_is_top-posting_.28and_why_shouldn.27t_I_do_it.29.3F
[2] https://www.debian.org/MailingLists/

Cheers
- -- tomás
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlmWAegACgkQBcgs9XrR2kb2xgCffB4H+pR5piD29rDi2w3d6eqX
1egAnjrcCkv4wHQJzJxEk8JKMkoFVEgm
=3lRP
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#185423

FromMario Castelán Castro <marioxcc.MT@yandex.com>
Date2017-08-18 02:40 +0200
Message-ID<ufDuq-42y-7@gated-at.bofh.it>
In reply to#185417

[Multipart message — attachments visible in raw view] — view raw

On 17/08/17 15:51, tomas@tuxteam.de wrote:
> On Thu, Aug 17, 2017 at 03:24:35PM -0500, Mario Castelán Castro wrote:
> [...]
> 
> But in general, folks here tend to be tolerant. And yes, there's a
> wiki entry encouraging "in-line" quoting [1].

Ah, I see. I rarely check the Debian Wiki because it is almost
abandoned. I have never found something useful there. For an example of
its state of abandonment see this fragment from that page:

“You really should see _Where is the foo package?_ above, but Debian
ships with Iceweasel, a rebranded Firefox.”

But a non-rebranded Firefox package is available in both Debian 8 and
Debian 9 (at least in the later, this is the default browse installed).

>> Bottom posting requires scrolling past text that may be not needed. Top
>> posting puts the messages in reverse chronological order, which is not
>> something bad by itself.
> 
> That's why you shoulnd't include the whole message, but snip the
> relevant parts you are answering to. Believe me, for long threads,
> this tends to work best.

Yes, except when nobody deletes the nested quotations (I do when it is
appropriate). Eventually most of the text in the messages becomes quotes.

Also, it is a problem that one loses track of who do the nested (except
the topmost) quotations belong to. Do you have any recommendation about
that?

> Yes, exactly. If someone needs the unabridged original post, it's either
> in her mailbox or in the archives.

Right, but here is a note about your wording: The great majority of
people in debian-user are male (judging by the personal names), and
moreover “he” is established as the pronoun in English when the sex is
undetermined. The use of the female pronoun “she” is situations like
this is unjustified.

> See [1] (there are also other hints on that page). Also [2] is a good
> reference.

I had read [2] in the past, but I did not find anything about posting
styles.

-----

I already try to use the inline style when appropriate. I will avoid
quoting the previous message at all in the cases where formerly I would
have used top posting. This seems to be the only change necessary to
comply with your suggestions.

Regards.

[toc] | [prev] | [next] | [standalone]


#185431 — Debian-user conventions [was: [...] incremental backups?]

From<tomas@tuxteam.de>
Date2017-08-18 09:40 +0200
SubjectDebian-user conventions [was: [...] incremental backups?]
Message-ID<ufK2R-bj-9@gated-at.bofh.it>
In reply to#185423
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Thu, Aug 17, 2017 at 07:33:53PM -0500, Mario Castelán Castro wrote:
> On 17/08/17 15:51, tomas@tuxteam.de wrote:

[...]

> > [...] And yes, there's a wiki entry encouraging "in-line" quoting [1].
> 
> Ah, I see. I rarely check the Debian Wiki because it is almost
> abandoned. I have never found something useful there. For an example of
> its state of abandonment see this fragment from that page:
> 
> “You really should see _Where is the foo package?_ above, but Debian
> ships with Iceweasel, a rebranded Firefox.”
>
> But a non-rebranded Firefox package is available in both Debian 8 and
> Debian 9 (at least in the later, this is the default browse installed).

It's a wiki (hint, hint ;-)

> > That's why you shoulnd't include the whole message, but snip the
> > relevant parts you are answering to. Believe me, for long threads,
> > this tends to work best.
> 
> Yes, except when nobody deletes the nested quotations (I do when it is
> appropriate). Eventually most of the text in the messages becomes quotes.

Yes, some gardening is needed. Take into account that you are addressing
about 3000 readers here, so sacrificing ten seconds of your time may
save hours in total (yes, I'm dramatizing a bit, but you get the idea).

> Also, it is a problem that one loses track of who do the nested (except
> the topmost) quotations belong to. Do you have any recommendation about
> that?

I try to stay at three layers at most, two typically; a good mail user
agent will help you keeping the levels straight. Add in some judgement :)

[...]

> Right, but here is a note about your wording: The great majority of
> people in debian-user are male (judging by the personal names), and
> moreover “he” is established as the pronoun in English when the sex is
> undetermined. The use of the female pronoun “she” is situations like
> this is unjustified.

Aha, someone noticed. It might be unjustified (don't know, can't assess
that), but it is subversive :)

But getting into more details would be definitely off-topic, I fear:
take it as an idiosyncracy. Feel free to complain from time to time :-)

> I already try to use the inline style when appropriate. I will avoid
> quoting the previous message at all in the cases where formerly I would
> have used top posting. This seems to be the only change necessary to
> comply with your suggestions.

Don't take my word for that. I might be wrong, and all that.

Cheers
- -- tomás
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlmWmBUACgkQBcgs9XrR2kbMFQCcDIsnZMZavSJHyDVt3fmLc5Sm
wRwAn3D4P9f1ORgZ/q8o434A3iz7VrCQ
=hAgd
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#185468

FromLiam O'Toole <liam.p.otoole@gmail.com>
Date2017-08-19 01:10 +0200
Message-ID<ufYyS-1Tk-7@gated-at.bofh.it>
In reply to#185395
On 2017-08-17, Mario Castelán Castro <marioxcc.MT@yandex.com> wrote:
> Hello.
>
> Currently I use rsync to make the backups of my personal data, including
> some manually selected important files of system configuration. I keep
> old backups to be more safe from the scenario where I have deleted
> something important, I make a backup, and I only notice the deletion
> afterwards.
>
> Each backup snapshot is stored in its own directory. There is much
> redundancy between subsequent backups. I use the option "--link-dest" to
> make hard links and thus save space for files that are *identical* to an
> already-existing file in the backup repository. but this is still
> inefficient. Any change to a file, even to its metadata (permission,
> modification time, etc.), will result in the file being saved at whole,
> instead of a delta.
>
> Can you suggest a more efficient alternative?
>

(...)

I use duplicity for exactly this scenario. See the wiki page[1] to get
started.

1: https://wiki.debian.org/Duplicity

-- 

Liam

[toc] | [prev] | [next] | [standalone]


#185502

FromMario Castelán Castro <marioxcc.MT@yandex.com>
Date2017-08-19 17:10 +0200
Message-ID<ugdxT-2S1-7@gated-at.bofh.it>
In reply to#185468

[Multipart message — attachments visible in raw view] — view raw

On 2017-08-18 23:53 +0100 Liam O'Toole <liam.p.otoole@gmail.com> wrote:
>I use duplicity for exactly this scenario. See the wiki page[1] to get
>started.
>
>1: https://wiki.debian.org/Duplicity

Judging from a quick glance at that project's homepage in GNU Savannah,
this seem indeed to be the right tool for the job, but I have yet to try
it.

Thanks you very much.

[toc] | [prev] | [next] | [standalone]


#185542

FromCelejar <celejar@gmail.com>
Date2017-08-20 05:10 +0200
Message-ID<ugoMF-1qh-3@gated-at.bofh.it>
In reply to#185395
On Thu, 17 Aug 2017 11:47:34 -0500
Mario Castelán Castro <marioxcc.MT@yandex.com> wrote:

> Hello.
> 
> Currently I use rsync to make the backups of my personal data, including
> some manually selected important files of system configuration. I keep
> old backups to be more safe from the scenario where I have deleted
> something important, I make a backup, and I only notice the deletion
> afterwards.
> 
> Each backup snapshot is stored in its own directory. There is much
> redundancy between subsequent backups. I use the option "--link-dest" to
> make hard links and thus save space for files that are *identical* to an
> already-existing file in the backup repository. but this is still
> inefficient. Any change to a file, even to its metadata (permission,
> modification time, etc.), will result in the file being saved at whole,
> instead of a delta.
> 
> Can you suggest a more efficient alternative?

There's Borg, which apparently has good deduplication. I've just
started using it, but it's a very sophisticated and quite popular piece
of software, judging by chatter in various internet threads.

https://borgbackup.readthedocs.io/en/stable/

Celejar

[toc] | [prev] | [next] | [standalone]


#185546

FromGene Heskett <gheskett@shentel.net>
Date2017-08-20 08:10 +0200
Message-ID<ugrAR-3bi-13@gated-at.bofh.it>
In reply to#185542
On Saturday 19 August 2017 23:07:01 Celejar wrote:

> On Thu, 17 Aug 2017 11:47:34 -0500
>
> Mario Castelán Castro <marioxcc.MT@yandex.com> wrote:
> > Hello.
> >
> > Currently I use rsync to make the backups of my personal data,
> > including some manually selected important files of system
> > configuration. I keep old backups to be more safe from the scenario
> > where I have deleted something important, I make a backup, and I
> > only notice the deletion afterwards.
> >
> > Each backup snapshot is stored in its own directory. There is much
> > redundancy between subsequent backups. I use the option
> > "--link-dest" to make hard links and thus save space for files that
> > are *identical* to an already-existing file in the backup
> > repository. but this is still inefficient. Any change to a file,
> > even to its metadata (permission, modification time, etc.), will
> > result in the file being saved at whole, instead of a delta.
> >
> > Can you suggest a more efficient alternative?
>
> There's Borg, which apparently has good deduplication. I've just
> started using it, but it's a very sophisticated and quite popular
> piece of software, judging by chatter in various internet threads.
>
> https://borgbackup.readthedocs.io/en/stable/
>
> Celejar
Amanda has quite intelligent ways to do that. I run it nightly and have 
been since the late 90's. Storage in my case is in what are called 
v-tapes, which in fact are directories on a separate, terrabyte drive. 
However, unlike tapes which are time burning sequential reading devices, 
the terrabyte drive is true random access, so recovery operations are 
about 1000x faster than real tapes. Not to mention the terrabyte drive 
is about 1000 times more dependable than tape can ever be.

That drive had 25 re-allocated sectors when smartctl came out all those 
years ago, and still has that same 25 sectors marked bad and re-assigned 
right now.

5 Reallocated_Sector_Ct   0x0033   100   100   036    Pre-fail  
Always       -       25

240 Head_Flying_Hours       0x0000   100   253   000    Old_age   
Offline      -       66095 (35 137 0)

Pretty good for coming up on 67,000 head flying hours... :)

Seagate Barracuda's of course, pure "commodity drive" at less than a $70 
bill.

But when I do replace them, the first thing I do is goto the Seagate site 
and download the latest firmware for that drive to a cd, reboot to it 
and let it update the drives firmware if it finds old firmware.  What 
you buy from the box houses is often first run stuff and may be 20 
revisions out of date.  The drive will probably be faster, once going 
from 26 megabytes/second to a hair over 120 megabytes/second. Sata is 
slow, old Asus mainboard shows under 3GB/sec:
root@coyote:/etc/avahi# hdparm -tT /dev/sdc
/dev/sdc:
 Timing cached reads:   4882 MB in  2.00 seconds = 2441.26 MB/sec
 Timing buffered disk reads: 348 MB in  3.00 seconds = 115.84 MB/sec

Amanda and I have had our differences, but at the end of the day, it Just 
Works(TM).  And has been for around 18 years here.

Cheers, Gene Heskett
-- 
"There are four boxes to be used in defense of liberty:
 soap, ballot, jury, and ammo. Please use in that order."
-Ed Howdershelt (Author)
Genes Web page <http://geneslinuxbox.net:6309/gene>

[toc] | [prev] | [next] | [standalone]


#185599

FromGlenn English <ghe2001@gmail.com>
Date2017-08-20 21:40 +0200
Message-ID<ugEeJ-2HE-7@gated-at.bofh.it>
In reply to#185546
On Sun, Aug 20, 2017 at 6:05 AM, Gene Heskett <gheskett@shentel.net> wrote:

> Amanda has quite intelligent ways to do that. I run it nightly and have
> been since the late 90's. Storage in my case is in what are called
> v-tapes, which in fact are directories on a separate, terrabyte drive.
> However, unlike tapes which are time burning sequential reading devices,
> the terrabyte drive is true random access, so recovery operations are
> about 1000x faster than real tapes. Not to mention the terrabyte drive
> is about 1000 times more dependable than tape can ever be.

Amanda is a very well done collection of programs.

It very efficiently does incremental backups to several types of media
-- Gene goes to disk, I go to tape (takes forever, but there are
several little boxes containing backups that are nowhere near a
failure point).

It backs up in tar (or dump) files so you can restore data when the
program and all its data files have been lost (been there, and it
works). There's a program that will scan through a tape, tell you
what's on that tape, then fetch your selection(s) for you.

For me, the big drawback to Amanda was the initial configuration. It's
huge and complex (at least it was a couple decades ago). But after
it's all done, a cron job will run your backup(s) every night, while
you sleep, with no problems. If you ask it to, it'll even verify the
backup for you (an unverified backup isn't a backup, as they say).

Take a look. It's worth the trouble.

--
Glenn English

[toc] | [prev] | [next] | [standalone]


#185608

FromMario Castelán Castro <marioxcc.MT@yandex.com>
Date2017-08-21 02:00 +0200
Message-ID<ugIim-57p-5@gated-at.bofh.it>
In reply to#185599

[Multipart message — attachments visible in raw view] — view raw

On 2017-08-20 19:37 +0000 Glenn English <ghe2001@gmail.com> wrote:
>For me, the big drawback to Amanda was the initial configuration. It's
>huge and complex (at least it was a couple decades ago). But after
>it's all done, a cron job will run your backup(s) every night, while
>you sleep, with no problems. If you ask it to, it'll even verify the
>backup for you (an unverified backup isn't a backup, as they say).

I have taken a glance at AMANDA, and it seems indeed to be very complex.
It is great that it works for your use case, but it does not seem to be an
appropriate tool for my case. I do not need any highly sophisticated
tools. As I noted in the first message, I only want to backup a personal
computer to an USB drive.

Since I must manually connect the USB drive to make the backups, there is
no point in automatizing it with cron. Network backups are irrelevant
in my current case.

Regards and thanks.

-- 
Do not eat animals, respect them as you respect people.
https://duckduckgo.com/?q=how+to+(become+OR+eat)+vegan

[toc] | [prev] | [next] | [standalone]


#185678

FromCelejar <celejar@gmail.com>
Date2017-08-22 05:50 +0200
Message-ID<uh8mt-56r-7@gated-at.bofh.it>
In reply to#185546
On Sun, 20 Aug 2017 02:05:46 -0400
Gene Heskett <gheskett@shentel.net> wrote:

> On Saturday 19 August 2017 23:07:01 Celejar wrote:
> 
> > On Thu, 17 Aug 2017 11:47:34 -0500
> >
> > Mario Castelán Castro <marioxcc.MT@yandex.com> wrote:
> > > Hello.
> > >
> > > Currently I use rsync to make the backups of my personal data,
> > > including some manually selected important files of system
> > > configuration. I keep old backups to be more safe from the scenario
> > > where I have deleted something important, I make a backup, and I
> > > only notice the deletion afterwards.
> > >
> > > Each backup snapshot is stored in its own directory. There is much
> > > redundancy between subsequent backups. I use the option
> > > "--link-dest" to make hard links and thus save space for files that
> > > are *identical* to an already-existing file in the backup
> > > repository. but this is still inefficient. Any change to a file,
> > > even to its metadata (permission, modification time, etc.), will
> > > result in the file being saved at whole, instead of a delta.
> > >
> > > Can you suggest a more efficient alternative?
> >
> > There's Borg, which apparently has good deduplication. I've just
> > started using it, but it's a very sophisticated and quite popular
> > piece of software, judging by chatter in various internet threads.
> >
> > https://borgbackup.readthedocs.io/en/stable/
> >
> > Celejar

> Amanda has quite intelligent ways to do that. I run it nightly and have 

[Snipped lots of miscellaneous, but seemingly irrelevant, discussion about Amanda's virtues.]

Amanda does deduplication? Link?

Celejar

[toc] | [prev] | [next] | [standalone]


#185681

FromGene Heskett <gheskett@shentel.net>
Date2017-08-22 07:30 +0200
Message-ID<uh9Vg-6gP-13@gated-at.bofh.it>
In reply to#185678
On Monday 21 August 2017 23:43:09 Celejar wrote:

> On Sun, 20 Aug 2017 02:05:46 -0400
>
> Gene Heskett <gheskett@shentel.net> wrote:
> > On Saturday 19 August 2017 23:07:01 Celejar wrote:
> > > On Thu, 17 Aug 2017 11:47:34 -0500
> > >
> > > Mario Castelán Castro <marioxcc.MT@yandex.com> wrote:
> > > > Hello.
> > > >
> > > > Currently I use rsync to make the backups of my personal data,
> > > > including some manually selected important files of system
> > > > configuration. I keep old backups to be more safe from the
> > > > scenario where I have deleted something important, I make a
> > > > backup, and I only notice the deletion afterwards.
> > > >
> > > > Each backup snapshot is stored in its own directory. There is
> > > > much redundancy between subsequent backups. I use the option
> > > > "--link-dest" to make hard links and thus save space for files
> > > > that are *identical* to an already-existing file in the backup
> > > > repository. but this is still inefficient. Any change to a file,
> > > > even to its metadata (permission, modification time, etc.), will
> > > > result in the file being saved at whole, instead of a delta.
> > > >
> > > > Can you suggest a more efficient alternative?
> > >
> > > There's Borg, which apparently has good deduplication. I've just
> > > started using it, but it's a very sophisticated and quite popular
> > > piece of software, judging by chatter in various internet threads.
> > >
> > > https://borgbackup.readthedocs.io/en/stable/
> > >
> > > Celejar
> >
> > Amanda has quite intelligent ways to do that. I run it nightly and
> > have
>
> [Snipped lots of miscellaneous, but seemingly irrelevant, discussion
> about Amanda's virtues.]
>
> Amanda does deduplication? Link?
>
> Celejar

Amanda does not do this "deduplication" that I am aware of.

That is another aspect of data control that does not belong in the job 
discription of what a backup program should do, which is to be a 
repository on some other storage medium besides the day to day operating 
cache, of the data you will need to recover and restore normal 
operations should your main drive become unusable with no signs of ill 
health until its falls over.

The backup program should be a relatively simple, so dependable its 
boring, yet smart enough to adjust its internal schedule of backup 
levels so as to use as much of the storage media as it needs on a long 
time continuous use scenario. Amanda carries this to extremes but you 
may have to help it occasionally if a given entry in the disklist grows 
until a level 0 back no longer fits on the amount of media you allow it 
to use per run.  As I've added machines to the list as they've been 
added to my home network over the last 20 years, I find myself needing 
to either buy a bigger drive, or further breakup my home directory into 
smaller pieces to reduce the total size of that one entry.  But amanda 
will never throw you under the bus, it a full won't fit, it continues to 
do level 1's or even level 2's.

A level 1 is anything that has changed since the last level 0, a level 2 
is anything changed since the previous level 1, etc etc.

Amanda is an administrator program, useing, at the PFC level of the 
duty's, usually tar and gzip but can use other compressors, for the 
actual data moving.  It keeps records over the span time you set it up 
to use, so It knows where everything it has backed up is.  But because 
those records are stored on the daily use drive, they aren't of much 
utility if you need to do a bare metal recovery, so I wrote a wrapper 
that adds this database and a copy of the configuration that made the 
back to the end of every backup it makes, so I can. and have actually 
done a bare metal recovery to a new drive in around 8 hours. The only 
thing I lost was about 75 emails that had come in since the nightly 
backup a few hours before that failure.

Backups are so much a personal preferences thing its hard to define.

Some folks who are used to doing a full backup on friday night that may 
take 50 terrabytes worth of tapes and a tape library that costs $50,000 
that they simply cannot wrap their mind around a program that does a 
level 0 of a given disklist entry on any arbitrary nightly run. And 
keeps track of a system such as the NY State Health system, doing it on 
a tape a night.

They can't get that amanda keeps records, and if you need to recover the 
home directories of Joe and Jane Sixpack who work in sales, amanda will 
look up the last level 0, restore that, and restore over that from the 
various other level 1 or 2 backups made since until it arrives at and 
recovers anything of theirs in last nights backup. I am backing up 5 
machines here, using 20 to 32 GB worth of space a night on a separate 1 
TB drive thats currently about 78% full.

You can make up your own mind, but to me amanda has been a good thing.

Cheers, Gene Heskett
-- 
"There are four boxes to be used in defense of liberty:
 soap, ballot, jury, and ammo. Please use in that order."
-Ed Howdershelt (Author)
Genes Web page <http://geneslinuxbox.net:6309/gene>

[toc] | [prev] | [next] | [standalone]


#185705

FromCelejar <celejar@gmail.com>
Date2017-08-22 15:20 +0200
Message-ID<uhhg5-2TJ-19@gated-at.bofh.it>
In reply to#185681
On Tue, 22 Aug 2017 01:25:50 -0400
Gene Heskett <gheskett@shentel.net> wrote:

...

> Amanda does not do this "deduplication" that I am aware of.
> 
> That is another aspect of data control that does not belong in the job 
> discription of what a backup program should do, which is to be a 
> repository on some other storage medium besides the day to day operating 
> cache, of the data you will need to recover and restore normal 
> operations should your main drive become unusable with no signs of ill 
> health until its falls over.

That's not the only job of backup programs. I have backups of that
nature (rsnapshot), but I want some critical data to be stored offsite,
in the cloud, as well. This is a very bandwidth and storage limited
context, so deduplication is most welcome, even though there is, of
course, the tradeoff that you mention.

...

> Backups are so much a personal preferences thing its hard to
define.

...

> They can't get that amanda keeps records, and if you need to recover the 
> home directories of Joe and Jane Sixpack who work in sales, amanda will 
> look up the last level 0, restore that, and restore over that from the 
> various other level 1 or 2 backups made since until it arrives at and 
> recovers anything of theirs in last nights backup. I am backing up 5 
> machines here, using 20 to 32 GB worth of space a night on a separate 1 
> TB drive thats currently about 78% full.
> 
> You can make up your own mind, but to me amanda has been a good thing.

Sounds like a great program.

> Cheers, Gene Heskett

Celejar

[toc] | [prev] | [next] | [standalone]


#185612

FromMario Castelán Castro <marioxcc.MT@yandex.com>
Date2017-08-21 03:10 +0200
Message-ID<ugJo5-602-3@gated-at.bofh.it>
In reply to#185542

[Multipart message — attachments visible in raw view] — view raw

On 2017-08-19 23:07 -0400 Celejar <celejar@gmail.com> wrote:
>There's Borg, which apparently has good deduplication. I've just
>started using it, but it's a very sophisticated and quite popular piece
>of software, judging by chatter in various internet threads.

This seems like an excellent tool for my use case. It has an interface
very much like control version systems (which I am familiar with), makes
efficient use of space and is no more complex to use than required (I'm
referencing the saying “make things as simple as possible but not more
simple”).

I have been testing it with toy cases to have at least some experience
with it before using it for my real backups.

Using a Git checkout of the latest release I get this warning: “Using a
pure-python msgpack! This will result in lower performance.”. Yet I have
the Debian package “python3-msgpack“. Do you know what the problem is?

Thanks.

-- 
Do not eat animals, respect them as you respect people.
https://duckduckgo.com/?q=how+to+(become+OR+eat)+vegan

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.debian.user


csiph-web