Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #244346 > unrolled thread

smartd

Started bypeter@easthope.ca
First post2022-01-22 19:30 +0100
Last post2022-01-22 22:50 +0100
Articles 17 — 8 participants

Back to article view | Back to linux.debian.user


Contents

  smartd peter@easthope.ca - 2022-01-22 19:30 +0100
    Re: smartd Dan Ritter <dsr@randomstring.org> - 2022-01-22 20:00 +0100
      Re: smartd peter@easthope.ca - 2022-01-23 16:50 +0100
      [SOLVED] Re: smartd peter@easthope.ca - 2022-02-04 18:20 +0100
    Re: smartd Andy Smith <andy@strugglers.net> - 2022-01-22 20:10 +0100
      Re: smartd peter@easthope.ca - 2022-01-23 07:00 +0100
        Re: smartd Andy Smith <andy@strugglers.net> - 2022-01-23 12:20 +0100
          Re: smartd Charles Curley <charlescurley@charlescurley.com> - 2022-01-23 16:20 +0100
            Re: smartd <tomas@tuxteam.de> - 2022-01-23 17:00 +0100
              Re: smartd rhkramer@gmail.com - 2022-01-24 18:20 +0100
      Re: smartd peter@easthope.ca - 2022-01-23 18:10 +0100
        Re: smartd Linux-Fan <Ma_Sys.ma@web.de> - 2022-01-23 19:20 +0100
          Re: smartd Andy Smith <andy@strugglers.net> - 2022-01-23 19:50 +0100
            Re: smartd Charles Curley <charlescurley@charlescurley.com> - 2022-01-23 22:20 +0100
          Re: smartd Andrei POPESCU <andreimpopescu@gmail.com> - 2022-01-25 10:50 +0100
      [SOLVED] Re: smartd peter@easthope.ca - 2022-02-04 19:10 +0100
    Re: smartd Charles Curley <charlescurley@charlescurley.com> - 2022-01-22 22:50 +0100

#244346 — smartd

Frompeter@easthope.ca
Date2022-01-22 19:30 +0100
Subjectsmartd
Message-ID<DItjb-6Bu-11@gated-at.bofh.it>
smartd reports to syslog.

Jan 22 08:49:17 joule smartd[563]: Device: /dev/sda [SAT], 155 Currently unreadable (pending) sectors
Jan 22 08:49:17 joule smartd[563]: Sending warning via /usr/share/smartmontools/smartd-runner to root ...
Jan 22 08:49:18 joule smartd[563]: Warning via /usr/share/smartmontools/smartd-runner to root: successful
Jan 22 08:49:18 joule smartd[563]: Device: /dev/sda [SAT], 132 Offline uncorrectable sectors

Two parts are available to mount /root; /root can be on /dev/sda1 or 
/dev/sda2. /home is used minimally.  If the errors are clustered, the 
bad area might be avoided easily in partitioning.

Feasible?  Can the locations of the errors be found?  Better to 
replace the drive?

Thx,                                ... P.

-- 
mobile: +1 778 951 5147
  VoIP: +1 604 670 0140
   48.7693 N 123.3053 W

[toc] | [next] | [standalone]


#244347

FromDan Ritter <dsr@randomstring.org>
Date2022-01-22 20:00 +0100
Message-ID<DItMd-6Lk-1@gated-at.bofh.it>
In reply to#244346
peter@easthope.ca wrote: 
> smartd reports to syslog.
> 
> Jan 22 08:49:17 joule smartd[563]: Device: /dev/sda [SAT], 155 Currently unreadable (pending) sectors
> Jan 22 08:49:17 joule smartd[563]: Sending warning via /usr/share/smartmontools/smartd-runner to root ...
> Jan 22 08:49:18 joule smartd[563]: Warning via /usr/share/smartmontools/smartd-runner to root: successful
> Jan 22 08:49:18 joule smartd[563]: Device: /dev/sda [SAT], 132 Offline uncorrectable sectors
> 
> Two parts are available to mount /root; /root can be on /dev/sda1 or 
> /dev/sda2. /home is used minimally.  If the errors are clustered, the 
> bad area might be avoided easily in partitioning.
> 
> Feasible?  Can the locations of the errors be found?  Better to 
> replace the drive?

Offline uncorrectable is bad.

You should do a backup ASAP.

Then, you have a choice: if the number doesn't increase over,
say, the  next week, it's just a bad patch. If it does increase,
the drive is bad and needs to be replaced.

-dsr-

[toc] | [prev] | [next] | [standalone]


#244423

Frompeter@easthope.ca
Date2022-01-23 16:50 +0100
Message-ID<DINhV-1HT-15@gated-at.bofh.it>
In reply to#244347
> You should do a backup ASAP.

Personal data is on a micro SD card.  After doing something worth 
saving it's backed to the host drive by me running this bash script .

Backup() { \
 if [ "$#" -gt 1 ]; then
   echo "Too many arguments.";
 else
   echo "0 or 1 arguments are OK.";
   if [ "$#" -eq 0 ]; then
     echo "0 arguments is OK.";
     destination=~/MY0.Bak;
     echo "destination is $destination.";
   else
     echo "1 argument is OK.";
     destination=~/MY1.Bak;
     echo "destination is $destination.";
   fi;
     echo "Executing rsync.";
     rsync \
      -auv /home/peter/MY/* $destination ;
     /bin/ls -ld ~/MY/MailMessages;
     printf "du -s $destination gives ";
     du -s $destination;
  fi;
}

"ls ... MailMessages" just reminds me to clean the mailbox.

The SD is used in multiple machines at two sites.  So my data is 
fairly well protected.  

If the SD fails, an inverse script restores data from a host drive, to 
a new SD.

If a meteorite goes through a machine, I get the holes in the case 
welded up and replace destroyed internals.  If the drive is replaced, 
I reinstall and configure the system.  A nuisance but not a catastrophe.

If an asteroid or meteorite shower or volcano destroys the SD card and 
all machines where it's backed, and I survive, data probably won't be 
a high priority but I can look for an old backup DVD.

Thx,                            ... P.



-- 
mobile: +1 778 951 5147
  VoIP: +1 604 670 0140
   48.7693 N 123.3053 W

[toc] | [prev] | [next] | [standalone]


#245004 — [SOLVED] Re: smartd

Frompeter@easthope.ca
Date2022-02-04 18:20 +0100
Subject[SOLVED] Re: smartd
Message-ID<DNapA-26h-3@gated-at.bofh.it>
In reply to#244347
    From: Dan Ritter <dsr@randomstring.org>
    Date: Sat, 22 Jan 2022 13:41:17 -0500
> Then, you have a choice: if the number doesn't increase over,
> say, the  next week, it's just a bad patch. 

That's the case.  The number isn't increasing.

> You should do a backup ASAP.

Backup function described here.
https://lists.debian.org/debian-user/2022/01/msg00863.html

Thanks for the information about the smartd.

                            ... P.
                            



-- 
mobile: +1 778 951 5147
  VoIP: +1 604 670 0140
   48.7693 N 123.3053 W

[toc] | [prev] | [next] | [standalone]


#244350

FromAndy Smith <andy@strugglers.net>
Date2022-01-22 20:10 +0100
Message-ID<DItVU-73X-11@gated-at.bofh.it>
In reply to#244346
Hello,

On Sat, Jan 22, 2022 at 09:18:27AM -0800, peter@easthope.ca wrote:
> smartd reports to syslog.
> 
> Jan 22 08:49:17 joule smartd[563]: Device: /dev/sda [SAT], 155 Currently unreadable (pending) sectors
> Jan 22 08:49:17 joule smartd[563]: Sending warning via /usr/share/smartmontools/smartd-runner to root ...
> Jan 22 08:49:18 joule smartd[563]: Warning via /usr/share/smartmontools/smartd-runner to root: successful
> Jan 22 08:49:18 joule smartd[563]: Device: /dev/sda [SAT], 132 Offline uncorrectable sectors
> 
> Two parts are available to mount /root; /root can be on /dev/sda1 or 
> /dev/sda2.

I don't understand what you mean by this statement. Either the disk
is already partitioned and / (you did mean "/", right, not "/root"?)
is on a known partition, or the disk isn't yet partitioned and / can
be on any partition you set it to be on.

> If the errors are clustered, the bad area might be avoided easily
> in partitioning.

You are better off finding the damaged sectors and causing the drive
to remap them by writing new content in there. Then you don't have
to keep track yourself of which sections of the disk are unusable.

> Feasible?  Can the locations of the errors be found?

Sure. Usually.

If the drive is currently not in use then it may be simpler to just
write over the entire drive with a simple

# dd if=/dev/zero of=/dev/sda

That should force a remap of any damaged sectors.

If you need to preserve what's currently on the drive then you can
use a SMART long self-test to try reading the whole drive. It should
report which LBA (sector) it got to when the test failed.

To start the test:

# smartctl -t long /dev/sda

To see the status of the test:

# smartctl -l selftest /dev/sda

You can instead do a "selective" test, to only test between certain
sector numbers.

Once you know the sector number you can verify that there's issues
by trying to read it with hdparm:

# hdparm --read-sector 9519790 /dev/sda

If that sector is truly damaged then this will show an error and
complaints in syslog.

You can force that sector to be written over with zeros, obviously
losing anything that was in it, again with hdparm:

# hdparm --yes-i-know-what-i-am-doing --write-sector 9519790 /dev/sda

This should force a remap and will complete successfully. If it
doesn't then the drive might be out of spare sectors, or is more
severely damaged, and it's done for.

If this drive is in use already then you possibly want to know which
files are affected by these bad sectors. I hope none, because you
use RAID. But if you need to know, I have done that before and can
dig out the scripts…

> Better to replace the drive?

Consumer HDDs usually have a few hundred spare sectors for
remapping. If I have a less important machine with a couple of bad
sectors I'll often be willing to force a remap like this. Seeing 155
bad sectors in a SMART report would worry me for any machine. But
it's your call.

Cheers,
Andy

-- 
https://bitfolk.com/ -- No-nonsense VPS hosting

[toc] | [prev] | [next] | [standalone]


#244390

Frompeter@easthope.ca
Date2022-01-23 07:00 +0100
Message-ID<DIE4V-4z6-1@gated-at.bofh.it>
In reply to#244350
    From: Andy Smith <andy@strugglers.net>
    Date: Sat, 22 Jan 2022 19:07:23 +0000
> > Two parts are available to mount /root; /root can be on /dev/sda1 or 
> > /dev/sda2.
>
> I don't understand what you mean by this statement. 

I should have referred to / rather than /root.

peter@joule:/home/peter$ lsblk --list | grep '\(N\|sda\)'
NAME MAJ:MIN RM   SIZE RO TYPE MOUNTPOINT
sda    8:0    0 149.1G  0 disk
sda1   8:1    0     7G  0 part
sda2   8:2    0     7G  0 part /
sda3   8:3    0     8G  0 part [SWAP]
sda4   8:4    0   127G  0 part /home

Currently / is in sda2 and sda1 is not used. If the faulty media is 
strictly in sda2, it can be avoided by shifting / to sda1.

> You are better off finding the damaged sectors and causing the drive
> to remap them by writing new content in there. Then you don't have
> to keep track yourself of which sections of the disk are unusable.

I don't understand how bad sectors are "remapped".  The process is 
internal to the drive?  Depends on Linux software? What about 
connecting the drive to another system and applying fsck to each part?
Then decide whether to scrap the drive.

> Consumer HDDs usually have a few hundred spare sectors for
> remapping.

What happens when all spare sectors are allocated?  Any indication  to 
prevent silent loss of data?

 Thanks,                    ... P.
 

-- 
mobile: +1 778 951 5147
  VoIP: +1 604 670 0140
   48.7693 N 123.3053 W

[toc] | [prev] | [next] | [standalone]


#244406

FromAndy Smith <andy@strugglers.net>
Date2022-01-23 12:20 +0100
Message-ID<DIJ4B-7HD-1@gated-at.bofh.it>
In reply to#244390
Hello,

On Sat, Jan 22, 2022 at 09:16:53PM -0800, peter@easthope.ca wrote:
>     From: Andy Smith <andy@strugglers.net>
>     Date: Sat, 22 Jan 2022 19:07:23 +0000
> > You are better off finding the damaged sectors and causing the drive
> > to remap them by writing new content in there. Then you don't have
> > to keep track yourself of which sections of the disk are unusable.
> 
> I don't understand how bad sectors are "remapped".  The process is 
> internal to the drive?

Yes. When a drive sector goes bad, the drive cannot read from it, so
you get an error in Linux when a read is attempted.

But if you are *writing* to it, if a modern drive can't do the write
it just writes the data to a spare sector and remaps that sector
location to the location of the formerly spare one.

The operating system is unaware that this has happened, though it is
recorded in SMART attributes (the reallocated sector count).

So overwriting bad sectors will make the problem go away until there
are no more spare sectors.

> Depends on Linux software?

No, anything that can write to the drive will work, which is why I
suggested dd over the whole drive if you aren't currently using it.

hdparm makes it easy to write a specific sector but it's also
possible with dd and its "skip" and "count" arguments. If you are
careful.

> What about connecting the drive to another system and applying
> fsck to each part?

What would be the goal? A SMART long self-test should tell you which
bits are unreadable.

> > Consumer HDDs usually have a few hundred spare sectors for
> > remapping.
> 
> What happens when all spare sectors are allocated?

The next time a sector goes bad it would not be fixable by writing
to it and there would be a part of the drive that is permanently
unusable. In the old days the "badblocks" tool would be used to find
these areas and avoid their use. These days we let drives remap bad
areas and replace either pro-actively or when they can't remap any
more.

Drives often encounter severe problems before they get as far as
using all their spare sectors. They send so many errors up to Linux
that Linux disconnects the whole device.

> Any indication to prevent silent loss of data?

When a sector goes bad, whatever data that was in there is now lost.
Since you cannot prevent drives from failing, appropriate
countermeasures include:

- Introducing redundancy with RAID or filesystems that have it built
  in, like btrfs or zfs

- Having good backups

Both are generally considered a good idea. With redundancy no data
would be lost and a tedious recovery process involving your backups
is turned into a more mundane process of replacing a failed drive.

You also need to monitor both of those to make sure they are
functioning properly.

Cheers,
Andy

-- 
https://bitfolk.com/ -- No-nonsense VPS hosting

[toc] | [prev] | [next] | [standalone]


#244421

FromCharles Curley <charlescurley@charlescurley.com>
Date2022-01-23 16:20 +0100
Message-ID<DIMOS-1xI-5@gated-at.bofh.it>
In reply to#244406
On Sun, 23 Jan 2022 11:09:47 +0000
Andy Smith <andy@strugglers.net> wrote:

> Yes. When a drive sector goes bad, the drive cannot read from it, so
> you get an error in Linux when a read is attempted.

As I understand things, that isn't entirely correct. From what I
understand:

If the drive can read a sector without error, it passes the data to the
OS and that's it.

If it gets an error, it uses cyclical redundancy check (CRC) data to
reconstruct the data. If that fails, it reports an error to the OS. If
the CRC reconstruction is successful, the drive re-writes the sector
and passes the reconstructed data back to the OS.

If the attempt to re-write the sector fails, the drive allocates a
spare sector, writes that, and notes the mapping in it sector
reallocation table.

There may be multiple efforts to re-write a sector, either in place or
reallocated.

And there's always the possibility that the sector reallocation table
will go bad.

-- 
Does anybody read signatures any more?

https://charlescurley.com
https://charlescurley.com/blog/

[toc] | [prev] | [next] | [standalone]


#244424

From<tomas@tuxteam.de>
Date2022-01-23 17:00 +0100
Message-ID<DINrz-1L2-1@gated-at.bofh.it>
In reply to#244421

[Multipart message — attachments visible in raw view] — view raw

On Sun, Jan 23, 2022 at 08:14:06AM -0700, Charles Curley wrote:
> On Sun, 23 Jan 2022 11:09:47 +0000
> Andy Smith <andy@strugglers.net> wrote:
> 
> > Yes. When a drive sector goes bad, the drive cannot read from it, so
> > you get an error in Linux when a read is attempted.
> 
> As I understand things, that isn't entirely correct. From what I
> understand:
> 
> If the drive can read a sector without error, it passes the data to the
> OS and that's it.
> 
> If it gets an error, it uses cyclical redundancy check (CRC) data to
> reconstruct the data. If that fails, it reports an error to the OS. If
> the CRC reconstruction is successful, the drive re-writes the sector
> and passes the reconstructed data back to the OS.

It is actually more complicated as this. As I understand this Wikipedia
entry [1], some errors while reading a block are to be expected: it
seems to be more profitable to push the density to the limit where error
correction picks up some rest. Only when the error rate surpasses some
threshold the block is remapped.

I guess SMART counts the latter events, but actually I have no idea :)

And the error correction codes are a bit more sophisticated than plain
CRC: Reed-Solomon or, more modern, low-density parity-check codes.

Cheers
-- 
tomás

[toc] | [prev] | [next] | [standalone]


#244525

Fromrhkramer@gmail.com
Date2022-01-24 18:20 +0100
Message-ID<DJbay-80P-11@gated-at.bofh.it>
In reply to#244424
On Sunday, January 23, 2022 10:57:53 AM tomas@tuxteam.de wrote:
> On Sun, Jan 23, 2022 at 08:14:06AM -0700, Charles Curley wrote:
> > On Sun, 23 Jan 2022 11:09:47 +0000
> > 
> > Andy Smith <andy@strugglers.net> wrote:
> > > Yes. When a drive sector goes bad, the drive cannot read from it, so
> > > you get an error in Linux when a read is attempted.
> > 
> > As I understand things, that isn't entirely correct. From what I
> > understand:
> > 
> > If the drive can read a sector without error, it passes the data to the
> > OS and that's it.
> > 
> > If it gets an error, it uses cyclical redundancy check (CRC) data to
> > reconstruct the data. If that fails, it reports an error to the OS. If
> > the CRC reconstruction is successful, the drive re-writes the sector
> > and passes the reconstructed data back to the OS.
> 
> It is actually more complicated as this. As I understand this Wikipedia
> entry [1], some errors while reading a block are to be expected: it
> seems to be more profitable to push the density to the limit where error
> correction picks up some rest. Only when the error rate surpasses some
> threshold the block is remapped.
> 
> I guess SMART counts the latter events, but actually I have no idea :)
> 
> And the error correction codes are a bit more sophisticated than plain
> CRC: Reed-Solomon or, more modern, low-density parity-check codes.

I would guess that the actual details vary depending  on the manufacturer and 
the revision level of the manufacturers firmware on the drive.

[toc] | [prev] | [next] | [standalone]


#244428

Frompeter@easthope.ca
Date2022-01-23 18:10 +0100
Message-ID<DIOxl-2C6-17@gated-at.bofh.it>
In reply to#244350
    From: Andy Smith <andy@strugglers.net>
    Date: Sat, 22 Jan 2022 19:07:23 +0000
> ... you use RAID.

I knew nothing of RAID.  Therefore read here.
https://en.wikipedia.org/wiki/RAID

Reliability is more valuable to me than speed.  RAID 0 won't help.  
For reliability I need a mirrored 2nd drive in the host; RAID 1 or 
higher.

Google of "site:wiki.debian.org raid" returned ten pages, each quite 
specialized and jargonified.  A few tips to establish mirroring can 
help.

> If this drive is in use already then you possibly want to know which
> files are affected by these bad sectors. I hope none, because you
> use RAID. But if you need to know, I have done that before and can
> dig out the scripts…

Seems more efficient to establish good reliability.  Then, if a drive 
fails, recycle and replace.

Thanks,                        ... P.



-- 
mobile: +1 778 951 5147
  VoIP: +1 604 670 0140
   48.7693 N 123.3053 W

[toc] | [prev] | [next] | [standalone]


#244430

FromLinux-Fan <Ma_Sys.ma@web.de>
Date2022-01-23 19:20 +0100
Message-ID<DIPD3-3dZ-5@gated-at.bofh.it>
In reply to#244428

[Multipart message — attachments visible in raw view] — view raw

peter@easthope.ca writes:

>     From: Andy Smith <andy@strugglers.net>
>     Date: Sat, 22 Jan 2022 19:07:23 +0000
> > ... you use RAID.
>
> I knew nothing of RAID.  Therefore read here.
> https://en.wikipedia.org/wiki/RAID
>
> Reliability is more valuable to me than speed.  RAID 0 won't help.
> For reliability I need a mirrored 2nd drive in the host; RAID 1 or
> higher.
>
> Google of "site:wiki.debian.org raid" returned ten pages, each quite
> specialized and jargonified.  A few tips to establish mirroring can
> help.

Here, it returns a few results, too. I think the most straight-forward is  
this one:

https://wiki.debian.org/SoftwareRAID

For most purposes, I recommend RAID1. If you have four HDDs of identical  
size, RAID10 might be tempting, too, but I'd still consider and possibly  
prefer just creating two independent RAID1 arrays.

If you want to configure it from the installer, these step-by-step  
instructions show all the relevant installer screens:

https://sleeplessbeastie.eu/2013/10/04/how-to-configure-software-raid1-during-installation-process/

Also, keep in mind that establishing the mirroring is not all you need to  
do. To really profit from the enhanced reliability, you need to play through  
the recovery scenario, too. I recommend doing this in a VM unless you have  
some dedicated machine with at least two HDDs to play with.

[...]

HTH and YMMV
Linux-Fan

öö

[toc] | [prev] | [next] | [standalone]


#244434

FromAndy Smith <andy@strugglers.net>
Date2022-01-23 19:50 +0100
Message-ID<DIQ66-3nO-13@gated-at.bofh.it>
In reply to#244430
Hello,

On Sun, Jan 23, 2022 at 07:09:48PM +0100, Linux-Fan wrote:
> To really profit from the enhanced reliability, you need to play
> through the recovery scenario, too. I recommend doing this in a VM
> unless you have some dedicated machine with at least two HDDs to
> play with.

If wanting to play around with mdraid you can do it with loop
devices created from image files on your regular filesystem.

$ cd /var/tmp
$ for i in a b; do fallocate -l 100M fake_disk_${i}.img; done
$ for i in a b; do sudo losetup -f fake_disk_{$i}.img; done
$ sudo mdadm --create --verbose /dev/md0 --level=1 --raid-devices=2 /dev/loop[01]
$ sudo mkfs.ext4 /dev/md0
$ sudo mount /dev/md0 /mnt

You can then practice removing, adding, failing etc. the loop devices.

When done playing around just unmount, stop array, losetup -d each loop
device then delete the files.

Cheers,
Andy

-- 
https://bitfolk.com/ -- No-nonsense VPS hosting

[toc] | [prev] | [next] | [standalone]


#244452

FromCharles Curley <charlescurley@charlescurley.com>
Date2022-01-23 22:20 +0100
Message-ID<DISrg-4U6-13@gated-at.bofh.it>
In reply to#244434
On Sun, 23 Jan 2022 18:41:36 +0000
Andy Smith <andy@strugglers.net> wrote:

> If wanting to play around with mdraid you can do it with loop
> devices created from image files on your regular filesystem.

Nice, thank you.

One would probably have to install mdadm:

# apt install mdadm


> for i in a b; do sudo losetup -f fake_disk_{$i}.img; done

Typo:

fake_disk_${i}.img

-- 
Does anybody read signatures any more?

https://charlescurley.com
https://charlescurley.com/blog/

[toc] | [prev] | [next] | [standalone]


#244574

FromAndrei POPESCU <andreimpopescu@gmail.com>
Date2022-01-25 10:50 +0100
Message-ID<DJqCC-n6-3@gated-at.bofh.it>
In reply to#244430

[Multipart message — attachments visible in raw view] — view raw

On Du, 23 ian 22, 19:09:48, Linux-Fan wrote:
> peter@easthope.ca writes:
> > 
> > I knew nothing of RAID.  Therefore read here.
> > https://en.wikipedia.org/wiki/RAID
> > 
> > Reliability is more valuable to me than speed.  RAID 0 won't help.
> > For reliability I need a mirrored 2nd drive in the host; RAID 1 or
> > higher.
> > 
> > Google of "site:wiki.debian.org raid" returned ten pages, each quite
> > specialized and jargonified.  A few tips to establish mirroring can
> > help.
> 
> Here, it returns a few results, too. I think the most straight-forward is
> this one:
> 
> https://wiki.debian.org/SoftwareRAID
> 
> For most purposes, I recommend RAID1. If you have four HDDs of identical
> size, RAID10 might be tempting, too, but I'd still consider and possibly
> prefer just creating two independent RAID1 arrays.
> 
> If you want to configure it from the installer, these step-by-step
> instructions show all the relevant installer screens:
> 
> https://sleeplessbeastie.eu/2013/10/04/how-to-configure-software-raid1-during-installation-process/
> 
> Also, keep in mind that establishing the mirroring is not all you need to
> do. To really profit from the enhanced reliability, you need to play through
> the recovery scenario, too. I recommend doing this in a VM unless you have
> some dedicated machine with at least two HDDs to play with.

Another thing to consider is that Linux Software RAID (also known as 
"md" or "mdadm" RAID) by itself doesn't have any integrity checking.

In case one of the drives returns bad data[1] it may end up overwriting 
the good data on the other drives[2][3].

It's possible to add an integrity checking layer, but in my opinion at 
that point the whole setup becomes so complex one might as well be using 
btrfs or ZFS instead.

Both have built in integrity checking and can recover the data provided 
there is at least one good copy available[4], in addition to the many 
other features they bring (logical volume management, snapshots, 
copy-on-write, etc.).


For the avoidance of doubt, neither is a replacement for backups[5].


[1] Cosmic rays flipped a bit, bad drive, bad cable, bad controller, 
etc.

[2] https://unixsheikh.com/articles/battle-testing-zfs-btrfs-and-mdadm-dm.html

[3] A RAID 1 can have more than just 2 drives, it's just uncommon in 
home setups because of cost reasons.

[4] It's possible to use both btrfs and ZFS without redundancy. They 
will be able to tell your data is corrupted, but won't be able to 
recover it, of course.

[5] http://taobackup.com/

Kind regards,
Andrei
-- 
http://wiki.debian.org/FAQsFromDebianUser

[toc] | [prev] | [next] | [standalone]


#245005 — [SOLVED] Re: smartd

Frompeter@easthope.ca
Date2022-02-04 19:10 +0100
Subject[SOLVED] Re: smartd
Message-ID<DNbbX-2C2-1@gated-at.bofh.it>
In reply to#244350
    From: Andy Smith <andy@strugglers.net>
    Date: Sat, 22 Jan 2022 19:07:23 +0000
> If the drive is currently not in use then it may be simpler to just
> write over the entire drive with a simple
> # dd if=/dev/zero of=/dev/sda

When convenient will get another drive and substitute in the machine.
Then the dodgy drive can be written over and I can decide whether to 
scrap it.

Meanwhile I have a nice laptop with Debian 11.1 installed.  If the 
desktop system fails I move the SD card to the laptop and carry on 
work as if nothing happened.  The desktop system can be resurrected 
when convenient.

> I hope none, because you use RAID. 

Being ignorant about RAID I had to read here.
https://en.wikipedia.org/wiki/RAID

No doubt invaluable for a server with a large quantity of dynamic 
data. For my work the SD card and spare machine seem adequate.  When 
the desktop system crashes I can lose a few hours of editing or an 
emessage. It's tolerable.

Thanks,                                ... P.


-- 
mobile: +1 778 951 5147
  VoIP: +1 604 670 0140
   48.7693 N 123.3053 W

[toc] | [prev] | [next] | [standalone]


#244359

FromCharles Curley <charlescurley@charlescurley.com>
Date2022-01-22 22:50 +0100
Message-ID<DIwqK-8oq-3@gated-at.bofh.it>
In reply to#244346
On Sat, 22 Jan 2022 09:18:27 -0800
peter@easthope.ca wrote:

> 

> Jan 22 08:49:17 joule smartd[563]: Device: /dev/sda [SAT], 155 Currently unreadable (pending) sectors
> Jan 22 08:49:17 joule smartd[563]: Sending warning via /usr/share/smartmontools/smartd-runner to root ...
> Jan 22 08:49:18 joule smartd[563]: Warning via /usr/share/smartmontools/smartd-runner to root: successful
> Jan 22 08:49:18 joule smartd[563]: Device: /dev/sda [SAT], 132
> Offline uncorrectable sectors

Unless you have a supply of replacement hard drives handy, I'd order a
new one now, then worry about the details of this one. My recent
experience with hard drives and lead times due to shipping times is not
encouraging.

-- 
Does anybody read signatures any more?

https://charlescurley.com
https://charlescurley.com/blog/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web