Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #266119 > unrolled thread

rsync --delete vs rsync --delete-after

Started byDefault User <hunguponcontent@gmail.com>
First post2024-01-17 17:30 +0100
Last post2024-01-19 08:10 +0100
Articles 20 on this page of 73 — 25 participants

Back to article view | Back to linux.debian.user


Contents

  rsync --delete vs rsync --delete-after Default User <hunguponcontent@gmail.com> - 2024-01-17 17:30 +0100
    Re: rsync --delete vs rsync --delete-after David Christensen <dpchrist@holgerdanske.com> - 2024-01-17 18:30 +0100
      Re: rsync --delete vs rsync --delete-after Keith Bainbridgge <keithrbaugroups@gmail.com> - 2024-01-18 00:10 +0100
      Re: rsync --delete vs rsync --delete-after Default User <hunguponcontent@gmail.com> - 2024-01-18 02:30 +0100
        Re: rsync --delete vs rsync --delete-after David Christensen <dpchrist@holgerdanske.com> - 2024-01-18 08:40 +0100
        Re: rsync --delete vs rsync --delete-after Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-18 12:10 +0100
          Re: rsync --delete vs rsync --delete-after <tomas@tuxteam.de> - 2024-01-18 12:20 +0100
            Re: rsync --delete vs rsync --delete-after Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-18 12:50 +0100
        Re: rsync --delete vs rsync --delete-after Michel Verdier <mv524@free.fr> - 2024-01-18 14:00 +0100
    Re: rsync --delete vs rsync --delete-after Kushal Kumaran <kushal@locationd.net> - 2024-01-17 19:30 +0100
      Re: rsync --delete vs rsync --delete-after Default User <hunguponcontent@gmail.com> - 2024-01-17 21:00 +0100
        Re: rsync --delete vs rsync --delete-after Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-17 21:20 +0100
        Re: rsync --delete vs rsync --delete-after Andy Smith <andy@strugglers.net> - 2024-01-17 22:00 +0100
        Re: rsync --delete vs rsync --delete-after Michel Verdier <mv524@free.fr> - 2024-01-18 02:40 +0100
        Re: rsync --delete vs rsync --delete-after hw <hw@adminart.net> - 2024-01-18 12:20 +0100
          Re: rsync --delete vs rsync --delete-after Ralph Aichinger <ra@h5.or.at> - 2024-01-18 13:50 +0100
            Re: rsync --delete vs rsync --delete-after Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-18 14:10 +0100
              Re: rsync --delete vs rsync --delete-after Ralph Aichinger <ra@h5.or.at> - 2024-01-18 15:30 +0100
              Re: rsync --delete vs rsync --delete-after hw <hw@adminart.net> - 2024-01-26 16:40 +0100
                Re: rsync --delete vs rsync --delete-after Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-26 17:30 +0100
                  Re: rsync --delete vs rsync --delete-after hw <hw@adminart.net> - 2024-01-28 18:00 +0100
            Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) Andy Smith <andy@strugglers.net> - 2024-01-26 16:20 +0100
              Re: Home UPS recommendations (Was Re: rsync --delete vs rsync --delete-after) fxkl47BF@protonmail.com - 2024-01-26 16:30 +0100
              Re: Home UPS recommendations Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-26 16:40 +0100
                Re: Home UPS recommendations Tixy <tixy@yxit.co.uk> - 2024-01-26 17:20 +0100
                Re: Home UPS recommendations "James H. H. Lampert" <jamesl@touchtonecorp.com> - 2024-01-26 17:20 +0100
              Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) Roger Price <debian@rogerprice.org> - 2024-01-26 19:10 +0100
                Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) David Wright <deblis@lionunicorn.co.uk> - 2024-01-26 20:50 +0100
                  Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) Roger Price <debian@rogerprice.org> - 2024-01-27 11:00 +0100
              Re: Home UPS recommendations Felix Miata <mrmazda@earthlink.net> - 2024-01-26 20:10 +0100
                Re: Home UPS recommendations ghe2001 <ghe2001@protonmail.com> - 2024-01-26 20:30 +0100
                  Re: Home UPS recommendations debian-user@howorth.org.uk - 2024-01-26 21:50 +0100
              Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) hw <hw@adminart.net> - 2024-01-28 19:00 +0100
                Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) Andy Smith <andy@strugglers.net> - 2024-01-28 20:00 +0100
                  Re: Home UPS recommendations (Was Re: rsync --delete vs  rsync--delete-after) gene heskett <gheskett@shentel.net> - 2024-01-29 07:00 +0100
                  Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) Andy Smith <andy@strugglers.net> - 2024-02-08 16:30 +0100
                    Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) Charles Curley <charlescurley@charlescurley.com> - 2024-02-08 17:30 +0100
                      Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) Curt <curty@free.fr> - 2024-02-08 17:40 +0100
                        Re: Home UPS recommendations Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-02-08 18:00 +0100
                    Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) hw <hw@adminart.net> - 2024-02-09 12:10 +0100
                      Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) Dan Ritter <dsr@randomstring.org> - 2024-02-09 13:10 +0100
                        Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) hw <hw@adminart.net> - 2024-02-09 22:40 +0100
                          Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) Dan Ritter <dsr@randomstring.org> - 2024-02-09 23:30 +0100
                      Re: Home UPS recommendations (Was Re: rsync --delete vs rsync --delete-after) "Roy J. Tellason, Sr." <roy@rtellason.com> - 2024-02-09 17:40 +0100
                        Re: Home UPS recommendations (Was Re: rsync --delete vs rsync --delete-after) Stefan Monnier <monnier@iro.umontreal.ca> - 2024-02-09 18:20 +0100
                          Re: Home UPS recommendations Felix Miata <mrmazda@earthlink.net> - 2024-02-10 01:10 +0100
                        Re: Home UPS recommendations (Was Re: rsync --delete vs rsync  --delete-after) hw <hw@adminart.net> - 2024-02-09 22:50 +0100
                          Re: Home UPS recommendations (Was Re: rsync --delete vs rsync --delete-after) "Roy J. Tellason, Sr." <roy@rtellason.com> - 2024-02-11 19:10 +0100
                      Re: Home UPS recommendations Felix Miata <mrmazda@earthlink.net> - 2024-02-09 18:20 +0100
                        Re: Home UPS recommendations debian-user@howorth.org.uk - 2024-02-09 21:40 +0100
                        Re: Home UPS recommendations hw <hw@adminart.net> - 2024-02-09 22:50 +0100
                          Re: Home UPS recommendations Felix Miata <mrmazda@earthlink.net> - 2024-02-10 01:00 +0100
                            Re: Home UPS recommendations hw <hw@adminart.net> - 2024-02-10 03:20 +0100
                              Re: Home UPS recommendations Felix Miata <mrmazda@earthlink.net> - 2024-02-10 04:30 +0100
                                Re: Home UPS recommendations hw <hw@adminart.net> - 2024-02-10 11:10 +0100
                                  Re: Home UPS recommendations Felix Miata <mrmazda@earthlink.net> - 2024-02-10 15:00 +0100
                                    Re: Home UPS recommendations hw <hw@adminart.net> - 2024-02-10 16:50 +0100
                                      Re: Home UPS recommendations Joe <joe@jretrading.com> - 2024-02-10 19:50 +0100
                                        Re: Home UPS recommendations gene heskett <gheskett@shentel.net> - 2024-02-10 22:50 +0100
                                        Re: Home UPS recommendations hw <hw@adminart.net> - 2024-02-11 02:00 +0100
            Re: rsync --delete vs rsync --delete-after hw <hw@adminart.net> - 2024-01-26 16:20 +0100
              Re: Data and hardware protection measures; was: rsync --delete vs  rsync --delete-after Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-26 17:00 +0100
                Re: Data and hardware protection measures; was: rsync --delete vs  rsync --delete-after hw <hw@adminart.net> - 2024-01-28 19:30 +0100
                  Re: Data and hardware protection measures Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-28 20:30 +0100
                    Re: Data and hardware protection measures Felix Miata <mrmazda@earthlink.net> - 2024-01-28 21:50 +0100
              Re: rsync --delete vs rsync --delete-after Ralph Aichinger <ra@h5.or.at> - 2024-01-27 14:10 +0100
    Re: rsync --delete vs rsync --delete-after Stefan Monnier <monnier@iro.umontreal.ca> - 2024-01-18 05:20 +0100
      Re: rsync --delete vs rsync --delete-after Andy Smith <andy@strugglers.net> - 2024-01-18 15:40 +0100
        Re: rsync --delete vs rsync --delete-after Stefan Monnier <monnier@iro.umontreal.ca> - 2024-01-18 16:10 +0100
        Re: rsync --delete vs rsync --delete-after Michel Verdier <mv524@free.fr> - 2024-01-18 22:10 +0100
          Re: rsync --delete vs rsync --delete-after Andy Smith <andy@strugglers.net> - 2024-01-18 22:50 +0100
            Re: rsync --delete vs rsync --delete-after Default User <hunguponcontent@gmail.com> - 2024-01-19 03:40 +0100
            Re: rsync --delete vs rsync --delete-after Michel Verdier <mv524@free.fr> - 2024-01-19 08:10 +0100

Page 1 of 4  [1] 2 3 4  Next page →


#266119 — rsync --delete vs rsync --delete-after

FromDefault User <hunguponcontent@gmail.com>
Date2024-01-17 17:30 +0100
Subjectrsync --delete vs rsync --delete-after
Message-ID<HXgXD-3Jzs-5@gated-at.bofh.it>
Hello!

Opinions, please.

I use rsync to copy my primary backup drive to a secondary backup drive
, so that the secondary backup drive is theoretically always an exact
copy of the primary backup drive.  

Here is the rsync command I use:

time sudo rsync -aAXHxvv --delete-after --numeric-ids --
info=progress2,stats2,name2 --
exclude={"/dev/*","/proc/*","/sys/*","/tmp/*","/run/*","/mnt/*","/media
/*","/lost+found"} /media/default/MSD0001/ /media/default/MSD0002/

Question: 
I use rsync --delete-after because it might seem to be "safer", so in
case of a "glitch" of any kind, no file ever disappears from both the 
source drive and the destination drive.  

However, I have read that using rsync --delete instead of rsync --
delete-after is faster and uses less memory, and so is more efficient. 

Note: The current copy process time varies, but takes a long time -
last night 131 minutes.
:(

Disk space used is not currently an issue.

But, is rsync --delete AS SAFE as rsync --delete-after?

[toc] | [next] | [standalone]


#266128

FromDavid Christensen <dpchrist@holgerdanske.com>
Date2024-01-17 18:30 +0100
Message-ID<HXhTH-3K8d-5@gated-at.bofh.it>
In reply to#266119
On 1/17/24 08:19, Default User wrote:
> Hello!
> 
> Opinions, please.
> 
> I use rsync to copy my primary backup drive to a secondary backup drive
> , so that the secondary backup drive is theoretically always an exact
> copy of the primary backup drive.
> 
> Here is the rsync command I use:
> 
> time sudo rsync -aAXHxvv --delete-after --numeric-ids --
> info=progress2,stats2,name2 --
> exclude={"/dev/*","/proc/*","/sys/*","/tmp/*","/run/*","/mnt/*","/media
> /*","/lost+found"} /media/default/MSD0001/ /media/default/MSD0002/
> 
> Question:
> I use rsync --delete-after because it might seem to be "safer", so in
> case of a "glitch" of any kind, no file ever disappears from both the
> source drive and the destination drive.
> 
> However, I have read that using rsync --delete instead of rsync --
> delete-after is faster and uses less memory, and so is more efficient.
> 
> Note: The current copy process time varies, but takes a long time -
> last night 131 minutes.
> :(
> 
> Disk space used is not currently an issue.
> 
> But, is rsync --delete AS SAFE as rsync --delete-after?


In the past, I used the --backup and --backup-dir options to retain 
files on the destination.


Then I moved my primary backup to ZFS, implemented snapshots, and 
implemented replication to the secondard backup devices.


David

[toc] | [prev] | [next] | [standalone]


#266152

FromKeith Bainbridgge <keithrbaugroups@gmail.com>
Date2024-01-18 00:10 +0100
Message-ID<HXncJ-3Nr4-3@gated-at.bofh.it>
In reply to#266128

On 18/1/24 04:19, David Christensen wrote:
 > I use rsync to copy my primary backup drive to a secondary backup drive


Good morning

I wonder why both processes don't copy from the original data; so that 
you don't copy a potential glitch in the first backup?

on a separate matter
Glitch?  Power goes off inconveniently?
-- 
All the best

Keith Bainbridge

keithrbau@gmail.com
keith.bainbridge.3216@gmail.com
+61 (0)447 667 468

UTC + 10:00

[toc] | [prev] | [next] | [standalone]


#266157

FromDefault User <hunguponcontent@gmail.com>
Date2024-01-18 02:30 +0100
Message-ID<HXpod-3OBu-1@gated-at.bofh.it>
In reply to#266128
On Wed, 2024-01-17 at 09:19 -0800, David Christensen wrote:
> On 1/17/24 08:19, Default User wrote:
> > Hello!
> > 
> > Opinions, please.
> > 
> > I use rsync to copy my primary backup drive to a secondary backup
> > drive
> > , so that the secondary backup drive is theoretically always an
> > exact
> > copy of the primary backup drive.
> > 
> > Here is the rsync command I use:
> > 
> > time sudo rsync -aAXHxvv --delete-after --numeric-ids --
> > info=progress2,stats2,name2 --
> > exclude={"/dev/*","/proc/*","/sys/*","/tmp/*","/run/*","/mnt/*","/m
> > edia
> > /*","/lost+found"} /media/default/MSD0001/ /media/default/MSD0002/
> > 
> > Question:
> > I use rsync --delete-after because it might seem to be "safer", so
> > in
> > case of a "glitch" of any kind, no file ever disappears from both
> > the
> > source drive and the destination drive.
> > 
> > However, I have read that using rsync --delete instead of rsync --
> > delete-after is faster and uses less memory, and so is more
> > efficient.
> > 
> > Note: The current copy process time varies, but takes a long time -
> > last night 131 minutes.
> > :(
> > 
> > Disk space used is not currently an issue.
> > 
> > But, is rsync --delete AS SAFE as rsync --delete-after?
> 
> 
> In the past, I used the --backup and --backup-dir options to retain 
> files on the destination.
> 
> 
> Then I moved my primary backup to ZFS, implemented snapshots, and 
> implemented replication to the secondard backup devices.
> 
> 
> David
> 



Hi guys, thanks for the replies.

BTW, the two backup drives are external 4 Gb USB HDDs.  The secondary
backup drive is always kept away from the computer, in a locked steel
box, except when it is attached to the computer to have the primary
backup drive copied to it. 

The primary backup drive is almost always attached to the computer, so
that I can access files archived there, that are not on the computer.
Probably not good practice, but that's why I have the secondary backup
drive.

I guess in the back of my mind I was thinking of a scenario where a
file on the primary backup drive might be corrupted or deleted before
being copied to the secondary backup drive.  Then if it is not present
on the primary backup drive, rsync dutifully deletes it from the
secondary backup drive. If the file is no longer on the computer's
internal SSD, I am then SOL.

BTW(2), I do use rsnapshot with cron jobs to back up the internal SSD
to the primary backup drive daily (and weekly, monthly, yearly).  But I
am not sure if I could also use it to do copies of the primary backup
drive to the secondary backup drive (maybe using an additional
configuration file)? 

I have also thought of trying to use partclone to copy the data from
the primary backup drive to the secondary backup drive. Why not try,
since rsync takes an hour and a half, every day! 

As for ZFS . . .   I wish!  But I think the resource requirements would
be too high for my setup, so it probably would be impractical - or
impossible.  And  then there's the complexity.  And the learning curve.

Finally I really should have a third backup drive in the mix.  Yes, I
am familiar with the 1-2-3 backup theory.  But the third backup drive
could not be off-site, for various reasons.  And I do have other things
to do rather than spend all day, every day, managing backups. 

<Sigh.>

 

[toc] | [prev] | [next] | [standalone]


#266178

FromDavid Christensen <dpchrist@holgerdanske.com>
Date2024-01-18 08:40 +0100
Message-ID<HXvah-3S7G-1@gated-at.bofh.it>
In reply to#266157
On 1/17/24 17:23, Default User wrote:
> On Wed, 2024-01-17 at 09:19 -0800, David Christensen wrote:
>> On 1/17/24 08:19, Default User wrote:
>>> Opinions, please.
> ...
> Hi guys, thanks for the replies.


YW.  :-)


> BTW, the two backup drives are external 4 Gb USB HDDs.  The secondary
> backup drive is always kept away from the computer, in a locked steel
> box, except when it is attached to the computer to have the primary
> backup drive copied to it.


That means the live data and all the backups are exposed to the same 
computer if and when it is compromised.  It would be safer to do the 
copy on another computer.  If you only have one computer, then boot live 
media to do the copy.


> The primary backup drive is almost always attached to the computer, so
> that I can access files archived there, that are not on the computer.


Accidental file modification and deletion are probably the most common 
failure modes.


One of the features of ZFS is snapshots.  Snapshots can be accessed via 
the hidden ".zfs/snapshot/" directory in the root of each file system. 
Users can browse the snapshots and copy out whatever files and 
directories they need, subject to the file system permissions at the 
time the snapshot was taken.  So, recovery from the above failure is 
"self serve" -- no sysadmin required, no backup/ restore software required.


> Probably not good practice, but that's why I have the secondary backup
> drive.
> 
> I guess in the back of my mind I was thinking of a scenario where a
> file on the primary backup drive might be corrupted or deleted before
> being copied to the secondary backup drive.  Then if it is not present
> on the primary backup drive, rsync dutifully deletes it from the
> secondary backup drive. If the file is no longer on the computer's
> internal SSD, I am then SOL.


ZFS snapshots are read-only.  Individual files and directories within a 
snapshot cannot be created, updated, or deleted.


Another feature of ZFS is that both data and metadata are checksummed. 
So, if you can read a file, the metadata and data are good.  "bitrot" is 
extremely unlikely.


> I have also thought of trying to use partclone to copy the data from
> the primary backup drive to the secondary backup drive. Why not try,


Transferring 4 TB at 150 MB/s would require:

	4 TB / 150 MB/s = 26,667 s

Or, about 7.4 hours.


> since rsync takes an hour and a half, every day!


How much data is rsync(1) transferring?


Another feature of ZFS snapshots is that they can be replicated from one 
pool to another.  The first replication is full.  Subsequent 
replications are incremental.  Here are some recent transfers from my 
backup server to a removable HDD:

frequency  bytes           real time    bandwidth
=========  ==============  ===========  =========
weekly     11,125,755,324   25m08.848s  7.37 MB/s
monthly    75,218,102,216  140m53.710s  8.90 MB/s


> As for ZFS . . .   I wish!  But I think the resource requirements would
> be too high for my setup, so it probably would be impractical - or
> impossible. 


ZFS will use as much or as little resources as you provide.


(ZFS is known to consume most of available memory OOTB; this can be 
tuned.  The simple answer is to put it on a dedicated file server/ NAS.)


Please tell us about your setup -- hardware, stored data, and I/O workload.


> And  then there's the complexity.


I find ZFS has more conceptual integrity than cobbling together 
fdisk(8), mdadm(8), cryptsetup(8), lvm(8), mkfs(8), etc..


> And the learning curve.


The Lucas books got me up the learning curve:

https://mwl.io/nonfiction/os#fmzfs

https://mwl.io/nonfiction/os#fmaz


While both have "FreeBSD" in the title, most of the content applies 
equally to OpenZFS on Debian.


> Finally I really should have a third backup drive in the mix.  Yes, I
> am familiar with the 1-2-3 backup theory.  But the third backup drive
> could not be off-site, for various reasons.  And I do have other things
> to do rather than spend all day, every day, managing backups.


Disaster preparedness is an open-ended problem.  Only you can decide how 
much solution to give it.


David

[toc] | [prev] | [next] | [standalone]


#266182

FromMichael Kjörling <2695bd53d63c@ewoof.net>
Date2024-01-18 12:10 +0100
Message-ID<HXyrw-3UbG-1@gated-at.bofh.it>
In reply to#266157
On 17 Jan 2024 20:23 -0500, from hunguponcontent@gmail.com (Default User):
> BTW, the two backup drives are external 4 Gb USB HDDs.  The secondary
> backup drive is always kept away from the computer, in a locked steel
> box, except when it is attached to the computer to have the primary
> backup drive copied to it. 
> 
> The primary backup drive is almost always attached to the computer, so
> that I can access files archived there, that are not on the computer.
> Probably not good practice, but that's why I have the secondary backup
> drive.

Hold on. Let's pause right here.

**That "primary backup drive" is not a backup at all.**

If at any point it can legitimately contain the only copy of a file
that you want to keep, then conceptually it is not a backup.

If it is continuously accessible from the system that you're trying to
protect, then anything bad which happens to that system is liable to
also at least potentially affect your "primary backup drive" as well.

There is nothing wrong with using external storage media to expand the
storage capacity of your computer, but you really shouldn't treat or
consider it as being a backup, because a backup is a _second_ copy to
be used if the primary copy somehow becomes unusable, corrupt,
inaccessible or whatever else might be the case.

Solve the right problem.


> I guess in the back of my mind I was thinking of a scenario where a
> file on the primary backup drive might be corrupted or deleted before
> being copied to the secondary backup drive.  Then if it is not present
> on the primary backup drive, rsync dutifully deletes it from the
> secondary backup drive. If the file is no longer on the computer's
> internal SSD, I am then SOL.

That is why a backup scheme should include some form of retention;
ideally immutable. A single botched backup run should not risk the
integrity of an only backup.


> BTW(2), I do use rsnapshot with cron jobs to back up the internal SSD
> to the primary backup drive daily (and weekly, monthly, yearly).  But I
> am not sure if I could also use it to do copies of the primary backup
> drive to the secondary backup drive (maybe using an additional
> configuration file)? 

You can use rsnapshot with any number of different configuration
files. Just pass `-c` to it to name the configuration file other than
the default. (For a long time I had two, and a script selecting either
one based on the backup target drive in use for that particular
backup. Worked great.)


> As for ZFS . . .   I wish!  But I think the resource requirements would
> be too high for my setup, so it probably would be impractical - or
> impossible.  And  then there's the complexity.  And the learning curve.

ZFS works fine on low-spec'd systems. I use it on a VPS with 2 GB RAM
without any problems. You really want 64-bit, but that's about it.

And while ZFS _can_ be complex, and supports rather complex usage
scenarios (because it is, at its core, an enterprise solution), at the
basic level the biggest difference is that you write "zpool create"
instead of "mkfs.ext4".

-- 
Michael Kjörling                     🔗 https://michael.kjorling.se
“Remember when, on the Internet, nobody cared that you were a dog?”

[toc] | [prev] | [next] | [standalone]


#266184

From<tomas@tuxteam.de>
Date2024-01-18 12:20 +0100
Message-ID<HXyBb-3UeK-9@gated-at.bofh.it>
In reply to#266182

[Multipart message — attachments visible in raw view] — view raw

On Thu, Jan 18, 2024 at 11:05:01AM +0000, Michael Kjörling wrote:
> On 17 Jan 2024 20:23 -0500, from hunguponcontent@gmail.com (Default User):

[...]

> Hold on. Let's pause right here.
> 
> **That "primary backup drive" is not a backup at all.**

It is: against the situation you fat-finger something and react
before the next backup happens (this is a threat worth being taken
into account). For that case, a backup with more than one "back"
version would be more useful (see --link-dest in rsync for that).

It is not against the system running amok and thrashing all attached
disks (be it by accident -- improbable, or by malice -- more probable.
cf. crypto-malware). Disconnected back ups are better here.

It is not for against your shed going up in flames. Off site for this.

Cheers
-- 
t

[toc] | [prev] | [next] | [standalone]


#266188

FromMichael Kjörling <2695bd53d63c@ewoof.net>
Date2024-01-18 12:50 +0100
Message-ID<HXz4d-3Uox-11@gated-at.bofh.it>
In reply to#266184
On 18 Jan 2024 12:15 +0100, from tomas@tuxteam.de:
>> **That "primary backup drive" is not a backup at all.**
> 
> It is: against the situation you fat-finger something and react
> before the next backup happens (this is a threat worth being taken
> into account). For that case, a backup with more than one "back"
> version would be more useful (see --link-dest in rsync for that).
> 
> It is not against the system running amok and thrashing all attached
> disks (be it by accident -- improbable, or by malice -- more probable.
> cf. crypto-malware). Disconnected back ups are better here.
> 
> It is not for against your shed going up in flames. Off site for this.

OP has specified (17 Jan 19:52 UTC) that the threat model includes,
among many other things, "mechanical failure" and "lightning".

A single copy offers zero protection against that.

-- 
Michael Kjörling                     🔗 https://michael.kjorling.se
“Remember when, on the Internet, nobody cared that you were a dog?”

[toc] | [prev] | [next] | [standalone]


#266193

FromMichel Verdier <mv524@free.fr>
Date2024-01-18 14:00 +0100
Message-ID<HXA9X-3V4J-11@gated-at.bofh.it>
In reply to#266157
On 2024-01-17, Default User wrote:

> BTW(2), I do use rsnapshot with cron jobs to back up the internal SSD
> to the primary backup drive daily (and weekly, monthly, yearly).  But I
> am not sure if I could also use it to do copies of the primary backup
> drive to the secondary backup drive (maybe using an additional
> configuration file)? 

You can safely rsync entire rsnapshot backups if you use --hard-links
option to preserve links set by rsnapshot. You have to rsync "entire"
rsnapshot directory else it can't set correct hard links.

[toc] | [prev] | [next] | [standalone]


#266135

FromKushal Kumaran <kushal@locationd.net>
Date2024-01-17 19:30 +0100
Message-ID<HXiPL-3KGQ-5@gated-at.bofh.it>
In reply to#266119
On Wed, Jan 17 2024 at 11:19:39 AM, Default User <hunguponcontent@gmail.com> wrote:
> Hello!
>
> Opinions, please.
>
> I use rsync to copy my primary backup drive to a secondary backup drive
> , so that the secondary backup drive is theoretically always an exact
> copy of the primary backup drive.  
>
> Here is the rsync command I use:
>
> time sudo rsync -aAXHxvv --delete-after --numeric-ids --
> info=progress2,stats2,name2 --
> exclude={"/dev/*","/proc/*","/sys/*","/tmp/*","/run/*","/mnt/*","/media
> /*","/lost+found"} /media/default/MSD0001/ /media/default/MSD0002/
>
> Question: 
> I use rsync --delete-after because it might seem to be "safer", so in
> case of a "glitch" of any kind, no file ever disappears from both the 
> source drive and the destination drive.  
>

What do you mean by "glitch"?  Irrespective of whether you use --delete
or --delete-after, deleted files on the source are deleted on the
destination once your rsync is complete (which is what I'd assume you
want when you want an exact copy).  I'd presume if you're ok with that,
you are also fine with the deletion happening earlier in the rsync
process?

If you're concerned about accidental deletions, you should just not use
any of the `--delete*` options (and give up on the exact copy
requirement).  You can look at alternatives to bare rsync that keep
track of multiple backed-up images (rsnapshot is a very simple wrapper
over rsync that can do this, for example).

> However, I have read that using rsync --delete instead of rsync --
> delete-after is faster and uses less memory, and so is more efficient. 
>
> Note: The current copy process time varies, but takes a long time -
> last night 131 minutes.
> :(

You can try using --delete for a couple of runs and see if it actually
affects performance in your situation.

>
> Disk space used is not currently an issue.
>
> But, is rsync --delete AS SAFE as rsync --delete-after?

You'll need to define what safety means for you.

-- 
regards,
kushal

[toc] | [prev] | [next] | [standalone]


#266138

FromDefault User <hunguponcontent@gmail.com>
Date2024-01-17 21:00 +0100
Message-ID<HXkeR-3LpH-1@gated-at.bofh.it>
In reply to#266135
On Wed, 2024-01-17 at 10:29 -0800, Kushal Kumaran wrote:
> On Wed, Jan 17 2024 at 11:19:39 AM, Default User
> <hunguponcontent@gmail.com> wrote:
> > Hello!
> > 
> > Opinions, please.
> > 
> > I use rsync to copy my primary backup drive to a secondary backup
> > drive
> > , so that the secondary backup drive is theoretically always an
> > exact
> > copy of the primary backup drive.  
> > 
> > Here is the rsync command I use:
> > 
> > time sudo rsync -aAXHxvv --delete-after --numeric-ids --
> > info=progress2,stats2,name2 --
> > exclude={"/dev/*","/proc/*","/sys/*","/tmp/*","/run/*","/mnt/*","/m
> > edia
> > /*","/lost+found"} /media/default/MSD0001/ /media/default/MSD0002/
> > 
> > Question: 
> > I use rsync --delete-after because it might seem to be "safer", so
> > in
> > case of a "glitch" of any kind, no file ever disappears from both
> > the 
> > source drive and the destination drive.  
> > 
> 
> What do you mean by "glitch"?  Irrespective of whether you use --
> delete
> or --delete-after, deleted files on the source are deleted on the
> destination once your rsync is complete (which is what I'd assume you
> want when you want an exact copy).  I'd presume if you're ok with
> that,
> you are also fine with the deletion happening earlier in the rsync
> process?
> 
> If you're concerned about accidental deletions, you should just not
> use
> any of the `--delete*` options (and give up on the exact copy
> requirement).  You can look at alternatives to bare rsync that keep
> track of multiple backed-up images (rsnapshot is a very simple
> wrapper
> over rsync that can do this, for example).
> 
> > However, I have read that using rsync --delete instead of rsync --
> > delete-after is faster and uses less memory, and so is more
> > efficient. 
> > 
> > Note: The current copy process time varies, but takes a long time -
> > last night 131 minutes.
> > :(
> 
> You can try using --delete for a couple of runs and see if it
> actually
> affects performance in your situation.
> 
> > 
> > Disk space used is not currently an issue.
> > 
> > But, is rsync --delete AS SAFE as rsync --delete-after?
> 
> You'll need to define what safety means for you.
> 



Hi, Kushal! 
Thanks for replying. 

By "glitch", I mean anything that could interfere with the rsync copy
process.  Possible causes: 
- electrical outages, voltage spikes, voltage drops, "brownouts"
- mechanical failure
- earthquake
- lightning
- cat walking on keyboard
- out of memory errors
- out of disk space errors
- PEBKAC errors
- etc.

By safe, may I try to explain using a story? 

I once read that centuries ago, a king wanted his crown to be safe.  So
four guards watched his crown at all times.  The guards were not
allowed to take their eves off of the crown, even for a second, until
the the their replacements, the next group of guards, all said "I see
the crown". 

I am sorry if I can not come up with a better, technical explanation. 
I use this story because it has always been meaningful to me, and seems
to point to the essence of what I am getting at. 

I am writing as someone who has lost data more than once over time, for
various reasons.  The loss has ranged from slightly annoying, to soul-
rending catastrophe. It is NEVER appreciated. 

I do intend to try doing rsync --delete instead of rsync --delete
after, to see if it "seems to work".  

But I just wanted to ask first ask, as it would seem better to hear
someone say "How stupid are you? I can't believe you were going to do
that!", than to have them say "How stupid are you? I can't believe you
did that!" 

[toc] | [prev] | [next] | [standalone]


#266140

FromMichael Kjörling <2695bd53d63c@ewoof.net>
Date2024-01-17 21:20 +0100
Message-ID<HXkyd-3LLz-1@gated-at.bofh.it>
In reply to#266138
On 17 Jan 2024 14:52 -0500, from hunguponcontent@gmail.com (Default User):
> I am writing as someone who has lost data more than once over time, for
> various reasons.  The loss has ranged from slightly annoying, to soul-
> rending catastrophe. It is NEVER appreciated. 

I think this gets closer to the root of what you're trying to achieve:
it sounds to me as though no matter what happens, you want to have a
restorable backup which you can trust to (reasonably well) match the
state of the source at the time when that backup was taken. That is a
commendable goal.

Twiddling with options to rsync won't offer you that in light of
several of the scenarios you listed, though; not least a lightning
strike. A lightning strike will blow out the drive just as much no
matter what options you're using to invoke rsync.

What you need is really rather multiple copies.

I suggest to get a second drive to use for backup purposes along with
the one you currently have. Only ever connect one of them at the same
time to data or power anywhere. Ideally, always keep at least one of
them in a separate location, powered off and disconnected.

That way, if you mess up one, or if one of them fails (for any
reason), or whatever else might happen, you have the other. It might
not be quite as up to date, but it will be in a known-good state.

For less drastic failures, really, look at rsnapshot. It's a wrapper
around rsync which makes maintaining multiple revisions of a backup
much easier. (It essentially passes to rsync with --link-dest, and
manages the respective root directories.)

-- 
Michael Kjörling                     🔗 https://michael.kjorling.se
“Remember when, on the Internet, nobody cared that you were a dog?”

[toc] | [prev] | [next] | [standalone]


#266144

FromAndy Smith <andy@strugglers.net>
Date2024-01-17 22:00 +0100
Message-ID<HXlaV-3LYi-1@gated-at.bofh.it>
In reply to#266138
Hi,

On Wed, Jan 17, 2024 at 02:52:49PM -0500, Default User wrote:
> By "glitch", I mean anything that could interfere with the rsync copy
> process.  Possible causes: 
> - electrical outages, voltage spikes, voltage drops, "brownouts"
> - mechanical failure
> - earthquake
> - lightning
> - cat walking on keyboard
> - out of memory errors
> - out of disk space errors
> - PEBKAC errors
> - etc.

But, both --delete and --delete-after only delete things from the
destination that ALREADY are missing on the source, so in which of
the above situations would something that still exists somewhere be
accidentally deleted with either option?

In my view the only ones that apply would be human error / cat
standard behaviour but even then you'd have to notice it had
happened and abort the transfer before rsync has chance to do the
delete. Do you typically sit and watch these transfers? It doesn't
sound like any kind of backup to me if it doesn't have history,
i.e. the state of your data *before* the cat walked on 'r' 'm' ' '
'-' 'f' 'r' ' ' '.' '<return>'.

Thanks,
Andy

-- 
https://bitfolk.com/ -- No-nonsense VPS hosting

[toc] | [prev] | [next] | [standalone]


#266158

FromMichel Verdier <mv524@free.fr>
Date2024-01-18 02:40 +0100
Message-ID<HXpxT-3OEn-5@gated-at.bofh.it>
In reply to#266138
On 2024-01-17, Default User wrote:

> By "glitch", I mean anything that could interfere with the rsync copy
> process.  Possible causes: 

Whatever the cause you just have to get return code and restart rsync
until it complete succesfully. Then you are sure to have an exact copy.
To cope with errors in primary backup you must have multiple copies as
suggest Michael. Personnally I use rsnapshot (a tool using rsync itself)
for that.

[toc] | [prev] | [next] | [standalone]


#266183

Fromhw <hw@adminart.net>
Date2024-01-18 12:20 +0100
Message-ID<HXyBb-3UeK-5@gated-at.bofh.it>
In reply to#266138
On Wed, 2024-01-17 at 14:52 -0500, Default User wrote:
> On Wed, 2024-01-17 at 10:29 -0800, Kushal Kumaran wrote:
> > On Wed, Jan 17 2024 at 11:19:39 AM, Default User
> > <hunguponcontent@gmail.com> wrote:
> > > Hello!
> > > 
> > > Opinions, please.
> > > 
> > > I use rsync to copy my primary backup drive to a secondary backup
> > > drive
> > > , so that the secondary backup drive is theoretically always an
> > > exact
> > > copy of the primary backup drive.  
> > > 
> > > Here is the rsync command I use:
> > > 
> > > time sudo rsync -aAXHxvv --delete-after --numeric-ids --
> > > info=progress2,stats2,name2 --
> > > exclude={"/dev/*","/proc/*","/sys/*","/tmp/*","/run/*","/mnt/*","/m
> > > edia
> > > /*","/lost+found"} /media/default/MSD0001/ /media/default/MSD0002/
> > > 
> > > Question: 
> > > I use rsync --delete-after because it might seem to be "safer", so
> > > in
> > > case of a "glitch" of any kind, no file ever disappears from both
> > > the 
> > > source drive and the destination drive.  
> > > 
> > 
> > What do you mean by "glitch"?  Irrespective of whether you use --
> > delete
> > or --delete-after, deleted files on the source are deleted on the
> > destination once your rsync is complete (which is what I'd assume you
> > want when you want an exact copy).  I'd presume if you're ok with
> > that,
> > you are also fine with the deletion happening earlier in the rsync
> > process?
> > 
> > If you're concerned about accidental deletions, you should just not
> > use
> > any of the `--delete*` options (and give up on the exact copy
> > requirement).  You can look at alternatives to bare rsync that keep
> > track of multiple backed-up images (rsnapshot is a very simple
> > wrapper
> > over rsync that can do this, for example).
> > 
> > > However, I have read that using rsync --delete instead of rsync --
> > > delete-after is faster and uses less memory, and so is more
> > > efficient. 
> > > 
> > > Note: The current copy process time varies, but takes a long time -
> > > last night 131 minutes.
> > > :(
> > 
> > You can try using --delete for a couple of runs and see if it
> > actually
> > affects performance in your situation.
> > 
> > > 
> > > Disk space used is not currently an issue.
> > > 
> > > But, is rsync --delete AS SAFE as rsync --delete-after?
> > 
> > You'll need to define what safety means for you.
> > 
> 
> 
> 
> Hi, Kushal! 
> Thanks for replying. 
> 
> By "glitch", I mean anything that could interfere with the rsync copy
> process.  Possible causes: 
> - electrical outages, voltage spikes, voltage drops, "brownouts"
> - mechanical failure
> - earthquake
> - lightning
> - cat walking on keyboard
> - out of memory errors
> - out of disk space errors
> - PEBKAC errors
> - etc.
> 
> By safe, may I try to explain using a story? 
> 
> I once read that centuries ago, a king wanted his crown to be safe.  So
> four guards watched his crown at all times.  The guards were not
> allowed to take their eves off of the crown, even for a second, until
> the the their replacements, the next group of guards, all said "I see
> the crown". 
> 
> I am sorry if I can not come up with a better, technical explanation. 
> I use this story because it has always been meaningful to me, and seems
> to point to the essence of what I am getting at. 
> 
> I am writing as someone who has lost data more than once over time, for
> various reasons.  The loss has ranged from slightly annoying, to soul-
> rending catastrophe. It is NEVER appreciated. 
> 
> I do intend to try doing rsync --delete instead of rsync --delete
> after, to see if it "seems to work".  
> 
> But I just wanted to ask first ask, as it would seem better to hear
> someone say "How stupid are you? I can't believe you were going to do
> that!", than to have them say "How stupid are you? I can't believe you
> did that!" 
> 
> 

On Wed, 2024-01-17 at 14:52 -0500, Default User wrote:
> [...]
> By "glitch", I mean anything that could interfere with the rsync copy
> process.  Possible causes: 
> - electrical outages, voltage spikes, voltage drops, "brownouts"
> - mechanical failure
> - earthquake
> - lightning
> - cat walking on keyboard
> - out of memory errors
> - out of disk space errors
> - PEBKAC errors
> - etc.
> 
> By safe, may I try to explain using a story? 
> 
> I once read that centuries ago, a king wanted his crown to be safe.  So
> four guards watched his crown at all times.  The guards were not
> allowed to take their eves off of the crown, even for a second, until
> the the their replacements, the next group of guards, all said "I see
> the crown". 
> 
> I am sorry if I can not come up with a better, technical explanation. 
> I use this story because it has always been meaningful to me, and seems
> to point to the essence of what I am getting at. 
> 
> I am writing as someone who has lost data more than once over time, for
> various reasons.  The loss has ranged from slightly annoying, to soul-
> rending catastrophe. It is NEVER appreciated. 
> 
> I do intend to try doing rsync --delete instead of rsync --delete
> after, to see if it "seems to work".  
> 
> But I just wanted to ask first ask, as it would seem better to hear
> someone say "How stupid are you? I can't believe you were going to do
> that!", than to have them say "How stupid are you? I can't believe you
> did that!" 

Always use an UPS.

Always use ECC RAM, especially when having lots of data.

Always use redundancy to store data for a running system, like some
form of RAID.  It won't hurt to use RAID for backups as well, though I
don't think that's required when you use it for the data you're
backing up.

Make sure you're notified when a disk in a RAID fails so you can
replace it before more disks fail.  Otherwise you won't notice before
it's too late.

Use multiple generations of backups.  Store some of them offsite.

I use --delete-before when backup space is tight so it's less likely
to run out of space[1].  If you have plenty of space, you can use
--delete-after.  I don't use --delete for backups.

Keep in mind that the data of the running system tends to be more
relevant than the data in backups because the data of the running
system is likely more up to date than the data in backups.  Also,
RAID, while _not_ being a substitute for backups, doesn't only protect
your data (which may be replacable) against disk failures, it also
saves you from the hassle, time and effort involved with downtimes and
having to re-aquire the data.


[1]:

I'm assuming that --delete may first copy and then delete, and that it
probably will effectively delete after because it won't delete files
it encounters only later in the process before encountering them,
while it will copy files being encountered when they are found and not
some time later after it first deleted other files to make room for
the encountered ones, which is yet another way in which you can run
out of space.

If you either use --delete-before or --delete-after, at least you know
what you can expect.

And don't ignore extend attributes etc. in the backups.

[toc] | [prev] | [next] | [standalone]


#266191

FromRalph Aichinger <ra@h5.or.at>
Date2024-01-18 13:50 +0100
Message-ID<HXA0h-3V15-1@gated-at.bofh.it>
In reply to#266183
Hello fellow Debian users,

On Thu, 2024-01-18 at 12:18 +0100, hw wrote:

> Always use an UPS.


Here I have a somewhat contrarian view, I hope not to offend too much:

For countries with stable electricity supplies (like Austria where I
live) having a small UPS might actually lead to more problems instead
of less, unless you are putting a lot of effort into it. Very often
have I had problems with UPSes, e.g. batteries dying, the UPS going
into some self test mode and inadvertedly shutting down, etc.

I've had no external power outage in the last 5 or 10 years, but a UPS
often needs at least one battery replacement during that time.

Unless you have some sort of professional server rack and redundant 2
phase supply, in my opinion UPS make very little sense to the home or 
small office user. Also modern Linux systems with journalling
filesystems will survive the occasional hard shutdown. Yes, I have
pulled the plug out of running Linux boxes occasionally because I was
too lazy to shut it down correctly and never had one break beyond the
usual fsck on boot.

> Always use redundancy to store data for a running system, like some
> form of RAID.  It won't hurt to use RAID for backups as well, though
> I don't think that's required when you use it for the data you're
> backing up.

Here I also doubt if this is a wise suggestion for the typical home
or small office user. RAID leads to lots and lots of complexity, that
is often not needed in a home setup. I'd rather have a working backup
setup with many independent copies before I even start thinking about
RAID. Yes, disks can fail, but data loss often is due to user
error and malware. RAID helps very little with the latter two causes
of data loss. And all too often have I seen people mess up their
complicated RAID setups, because they pulled the wrong disk when
another one broke, or because they misinterpreted complicated error
messages, creating unnecessary data loss out of user error by
themselves.

As a home/SOHO user, I'd rather have a working backup every few hours
or every day than some RAID10 wonder that makes me lose more time on
reading RAID documentation, and ordering spare drives (you've got
one of those spares for each array, do you?) than is actually lost by
not being able to restore to the exact last minute before a hard disk
died.

/ralph -- no UPS at home, using RAID1 md mirroring though

[toc] | [prev] | [next] | [standalone]


#266194

FromMichael Kjörling <2695bd53d63c@ewoof.net>
Date2024-01-18 14:10 +0100
Message-ID<HXAjD-3Vnz-19@gated-at.bofh.it>
In reply to#266191
On 18 Jan 2024 13:26 +0100, from ra@h5.or.at (Ralph Aichinger):
> As a home/SOHO user, I'd rather have a working backup every few hours
> or every day than some RAID10 wonder

Definitely agree that a solid backup regimen (including regular
automated backups; at least one off-site copy _at least_ of critical,
hot data; and planning for the contingency that you need to restore
that backup onto a brand new system without access to anything on your
current system -- think "home burns down at night" or "burglar"
scenario) is the _first_ step, and one that a great deal of people
still fail at.

RAID is for uptime. If a week-long outage (to get replacement hardware
and restore the most recent backup) and a day's worth of data loss is
largely inconsequential, as quite frankly it likely is for most home
users save for the cost of replacement hardware, that's a very
different scenario from if that same outage costs $$€€¥¥ and could
destroy your livelihood; and consequently the choices made _should_
likely be different.

_Mirrored backups_ makes very little sense to me. If a storage device
used for storage of backups fails prematurely, just toss it and get a
new one and make a new backup. If the backup software goes haywire and
starts overwriting everything with random garbage, having the garbage
mirrored isn't going to help you. It's much better to have two
independent backup targets and switching between them, and figuring
the switching interval into your RPO. The only time when something
like mirrored backups will help you is when you have only one backup
set, the backup itself works fine, but a backup drive fails, _and_ the
source fails before you've been able to make a new backup. That's a
_very_ narrow scenario and easily solved by having two backup sets.

-- 
Michael Kjörling                     🔗 https://michael.kjorling.se
“Remember when, on the Internet, nobody cared that you were a dog?”

[toc] | [prev] | [next] | [standalone]


#266199

FromRalph Aichinger <ra@h5.or.at>
Date2024-01-18 15:30 +0100
Message-ID<HXBz4-3W1T-9@gated-at.bofh.it>
In reply to#266194
On Thu, 2024-01-18 at 13:09 +0000, Michael Kjörling wrote:
> 
> Definitely agree that a solid backup regimen (including regular
> automated backups; at least one off-site copy _at least_ of critical,
> hot data; and planning for the contingency that you need to restore
> that backup onto a brand new system without access to anything on
> your
> current system -- think "home burns down at night" or "burglar"
> scenario) is the _first_ step, and one that a great deal of people
> still fail at.

Absolutely. I use a Raspberry Pi with an external
USB drive for my off-site backups with Resitic. Seems to
work fine for now, draws very little power, and the 4TB of a small
2.5" disk is plenty for my personal backups, when deduplicated. Still
this setup probably is too complicated for many home users, where a 
cloud backup or similar makes more sense.

> RAID is for uptime. If a week-long outage (to get replacement
> hardware
> and restore the most recent backup) and a day's worth of data loss is
> largely inconsequential, as quite frankly it likely is for most home
> users save for the cost of replacement hardware,

For me the calculation is more or less "next workday to go to the local
shop for a replacement hard drive" and a few hours to restore backups.
Yes, if you depend on mail order, one week might be more realistic.
Then I probably would keep a spare drive around even as a home user.

>  that's a very
> different scenario from if that same outage costs $$€€¥¥ and could
> destroy your livelihood; and consequently the choices made _should_
> likely be different.

Of course. As soon as you have to pay several people's salaries 
needlessly while they sit around for access to their data, RAID
makes more sense quickly. Still, it makes sense to think about what
you can do yourself vs. what needs external work done, also because
somebody external to repair a RAID might not show up all that quickly
unless you've got some pre-negotiated contract.

> _Mirrored backups_ makes very little sense to me. If a storage device
> used for storage of backups fails prematurely, just toss it and get a
> new one and make a new backup.

Absolutely! Just make more backups, or more backups with different,
independent strategies. As much as possible I try to do two independent
systems (e.g. Restic doing time-based offsite backups, and a cron job
doing a simple tar.gz file into some local drive or storage).

/ralph

[toc] | [prev] | [next] | [standalone]


#266630

Fromhw <hw@adminart.net>
Date2024-01-26 16:40 +0100
Message-ID<I0wtb-5J3f-9@gated-at.bofh.it>
In reply to#266194
On Thu, 2024-01-18 at 13:09 +0000, Michael Kjörling wrote:
> On 18 Jan 2024 13:26 +0100, from ra@h5.or.at (Ralph Aichinger):
> > As a home/SOHO user, I'd rather have a working backup every few hours
> > or every day than some RAID10 wonder
> 
> Definitely agree that a solid backup regimen (including regular
> automated backups; at least one off-site copy _at least_ of critical,
> hot data; and planning for the contingency that you need to restore
> that backup onto a brand new system without access to anything on your
> current system -- think "home burns down at night" or "burglar"
> scenario) is the _first_ step, and one that a great deal of people
> still fail at.
> 
> RAID is for uptime.

It's also for saving you from the hassle involved with loosing data
when a disk fails.

> If a week-long outage (to get replacement hardware and restore the
> most recent backup) and a day's worth of data loss is largely
> inconsequential, as quite frankly it likely is for most home users
> save for the cost of replacement hardware, that's a very different
> scenario from if that same outage costs $$€€¥¥ and could destroy
> your livelihood; and consequently the choices made _should_ likely
> be different.

That's assuming your time isn't worth anything and ignores whatever
the loss of data may cost you.  If that isn't relevant to you, you
don't backups, either.

> _Mirrored backups_ makes very little sense to me. If a storage device
> used for storage of backups fails prematurely, just toss it and get a
> new one and make a new backup. If the backup software goes haywire and
> starts overwriting everything with random garbage, having the garbage
> mirrored isn't going to help you. It's much better to have two
> independent backup targets and switching between them, and figuring
> the switching interval into your RPO. The only time when something
> like mirrored backups will help you is when you have only one backup
> set, the backup itself works fine, but a backup drive fails, _and_ the
> source fails before you've been able to make a new backup. That's a
> _very_ narrow scenario and easily solved by having two backup sets.

Having multiple generations of backups already increases the needed
storage space by a bit more than half.  That makes it already arguable
if it's better to make (multiple generations of) backups on a single
RAID or on N single disks.  Any of the disks can fail at any time.  If
you go with N == 2, a RAID (with multiple generations of backups on
it) can be better because when a disk fails, the RAID will very likely
survive and the non-RAID may not.

So for daily use, I'd ideally make multiple generations on RAID and
copy the latest generation to somewhere off-site, using either single
disks, or another RAID.

Having no hassle and no downtime at the production system and instead
having hassle and downtime off-site may be much more desirable than
having it the other way round.

Trying to make things appear easier by pointing out that failed disks
can be replaced is not helpful.  Replacing a disk in a RAID isn't any
more difficult (and can be much easier) than replacing a disk that
isn't in a RAID.

[toc] | [prev] | [next] | [standalone]


#266639

FromMichael Kjörling <2695bd53d63c@ewoof.net>
Date2024-01-26 17:30 +0100
Message-ID<I0xfA-5Jzr-3@gated-at.bofh.it>
In reply to#266630
On 26 Jan 2024 16:39 +0100, from hw@adminart.net (hw):
>> RAID is for uptime.
> 
> It's also for saving you from the hassle involved with loosing data
> when a disk fails.

Which translates to more quickly fully recovering from the loss of a
storage device.

When used for redundancy and staying within the redundancy threshold
of your setup, _done right_, RAID reduces unplanned downtime to zero
in case of loss of a storage device. Depending on your setup, the
replacement can either be made online or planned for, and once
complete, (hopefully) no data whatsoever has been lost and everything
has happened at minimal time cost. I'm assuming that _the physical
act_ of replacing the storage device is similar regardless of how it
is configured in software: the same number of screws, clamps or
similar need removing and reinstalling; the same amount of time is
required to physically move the old and the new storage device; etc.


>> If a week-long outage (to get replacement hardware and restore the
>> most recent backup) and a day's worth of data loss is largely
>> inconsequential, as quite frankly it likely is for most home users
>> save for the cost of replacement hardware, that's a very different
>> scenario from if that same outage costs $$€€¥¥ and could destroy
>> your livelihood; and consequently the choices made _should_ likely
>> be different.
> 
> That's assuming your time isn't worth anything and ignores whatever
> the loss of data may cost you.  If that isn't relevant to you, you
> don't backups, either.

Sure, you can tune the RPO (recovery point objective) for your backup
solution as needed. If you need a RPO of five minutes after
catastrophic storage failure (meaning that you lose no more than five
minutes of data regardless of when failure happens), then you're going
to make different choices than if a 24-48 hour RPO is fine. But I'm
willing to say that _most_ home users can recover from a loss of a
day's worth of changes to their data without that being a major blow
to whatever they are doing; and if the user does something important,
nothing prevents running an extra backup to capture those changes.
(I've done that myself on occasion.)

Similarly, you are going to make different choices based on your RTO
(recovery time objective). RAID gives you essentially zero RTO as long
as you retain sufficient redundancy, but _not_ past that. Once your
storage drops below whatever level of redundancy you have, you're
looking at rebuilding from backups, at which point backup frequency
and backup restoration time are the minimums which will dictate your
RPO and RTO respectively. So even if you have RAID, you need to know
what your target RPO and RTO are, respectively, because again those
values are going to be a major factor in the design of your backup
regimen.


>> _Mirrored backups_ makes very little sense to me. [...]
> 
> Having multiple generations of backups already increases the needed
> storage space by a bit more than half.  That makes it already arguable
> if it's better to make (multiple generations of) backups on a single
> RAID or on N single disks.  Any of the disks can fail at any time.  If
> you go with N == 2, a RAID (with multiple generations of backups on
> it) can be better because when a disk fails, the RAID will very likely
> survive and the non-RAID may not.

I'm not sure how you figure that. To survive the loss of N > 0 storage
devices within a set, a storage solution needs to have a raw capacity
greater than the usable capacity. An illustrative case would be a
three-way mirror: it can survive the loss of two out of three storage
devices, but only has the usable storage capacity of a single
(typically the smallest) of the devices. It's not really any different
from that a commodity HDD or SSD probably won't survive the failure of
one of its platters or flash chips respectively, even ignoring cascade
effects of either the failure or the cause of failure.

How much extra storage space you need to keep multiple generations of
backups is going to depend a lot on the rate of churn of your dataset.
Again as an illustrative example only, suppose you have 1 TB of data
and modify 1 MB per day, and make backups daily; with two devices each
capable of storing 2 TB you can have two full copies plus 1M days'
worth of history on each. This is irrespective of how those storage
devices themselves are furnished.

> Trying to make things appear easier by pointing out that failed disks
> can be replaced is not helpful.

It's a _backup_. _By definition_, a backup is only critical once the
primary copy becomes inaccessible for some reason. Hence:

>> The only time when something
>> like mirrored backups will help you is when you have only one backup
>> set, the backup itself works fine, but a backup drive fails, _and_ the
>> source fails before you've been able to make a new backup.

For a primary copy, _of course_ the calculus is different.

-- 
Michael Kjörling                     🔗 https://michael.kjorling.se
“Remember when, on the Internet, nobody cared that you were a dog?”

[toc] | [prev] | [next] | [standalone]


Page 1 of 4  [1] 2 3 4  Next page →

Back to top | Article view | linux.debian.user


csiph-web