Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #265426 > unrolled thread

1 Currently unreadable (pending) sectors How worried should I be?

Started byCharles Curley <charlescurley@charlescurley.com>
First post2024-01-02 23:50 +0100
Last post2024-01-03 14:30 +0100
Articles 5 on this page of 25 — 9 participants

Back to article view | Back to linux.debian.user


Contents

  1 Currently unreadable (pending) sectors How worried should I be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-02 23:50 +0100
    Re: 1 Currently unreadable (pending) sectors How worried should I be? Dan Ritter <dsr@randomstring.org> - 2024-01-03 00:10 +0100
      Re: 1 Currently unreadable (pending) sectors How worried should I  be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-03 00:40 +0100
        Re: 1 Currently unreadable (pending) sectors How worried should I be? Dan Ritter <dsr@randomstring.org> - 2024-01-03 02:00 +0100
      Re: 1 Currently unreadable (pending) sectors How worried should I  be? Tixy <tixy@yxit.co.uk> - 2024-01-03 08:50 +0100
    Re: 1 Currently unreadable (pending) sectors How worried should I be? Dan Purgert <dan@djph.net> - 2024-01-03 00:10 +0100
      Re: 1 Currently unreadable (pending) sectors How worried should I  be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-03 00:50 +0100
        Re: 1 Currently unreadable (pending) sectors How worried should I be? Andy Smith <andy@strugglers.net> - 2024-01-03 01:40 +0100
          Re: 1 Currently unreadable (pending) sectors How worried should I  be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-03 04:20 +0100
            Re: 1 Currently unreadable (pending) sectors How worried should I be? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-03 12:10 +0100
              Re: 1 Currently unreadable (pending) sectors How worried should I  be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-03 21:30 +0100
                Re: 1 Currently unreadable (pending) sectors How worried should I  be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-04 00:30 +0100
                  Re: 1 Currently unreadable (pending) sectors How worried should I be? <tomas@tuxteam.de> - 2024-01-04 12:00 +0100
                    Re: 1 Currently unreadable (pending) sectors How worried should I  be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-04 15:10 +0100
                  Re: 1 Currently unreadable (pending) sectors How worried should I be? Andy Smith <andy@strugglers.net> - 2024-01-05 22:10 +0100
                    Re: 1 Currently unreadable (pending) sectors How worried should I  be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-06 00:30 +0100
                      Re: 1 Currently unreadable (pending) sectors How worried should I be? David Christensen <dpchrist@holgerdanske.com> - 2024-01-06 02:30 +0100
                        Secure erase [was: Re: 1 Currently unreadable (pending) sectors How  worried should I be?] Max Nikulin <manikulin@gmail.com> - 2024-01-06 03:50 +0100
                        Re: 1 Currently unreadable (pending) sectors How worried should I  be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-06 06:20 +0100
                          Re: 1 Currently unreadable (pending) sectors How worried should I be? David Christensen <dpchrist@holgerdanske.com> - 2024-01-06 09:40 +0100
                            Re: reinstallation and restore after catastrophic mistake or  failure; was: 1 Currently unreadable (pending) sectors How worried should I  be? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-06 13:40 +0100
                              Re: reinstallation and restore after catastrophic mistake or failure;  was: 1 Currently unreadable (pending) sectors How worried should I be? David Christensen <dpchrist@holgerdanske.com> - 2024-01-07 00:40 +0100
                Re: 1 Currently unreadable (pending) sectors How worried should I be? Max Nikulin <manikulin@gmail.com> - 2024-01-04 16:20 +0100
                Re: 1 Currently unreadable (pending) sectors How worried should I be? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-04 17:50 +0100
            Re: 1 Currently unreadable (pending) sectors How worried should I be? Andy Smith <andy@strugglers.net> - 2024-01-03 14:30 +0100

Page 2 of 2 — ← Prev page 1 [2]


#265581 — Re: reinstallation and restore after catastrophic mistake or failure; was: 1 Currently unreadable (pending) sectors How worried should I be?

FromMichael Kjörling <2695bd53d63c@ewoof.net>
Date2024-01-06 13:40 +0100
SubjectRe: reinstallation and restore after catastrophic mistake or failure; was: 1 Currently unreadable (pending) sectors How worried should I be?
Message-ID<HTe82-1deA-7@gated-at.bofh.it>
In reply to#265573
On 6 Jan 2024 00:37 -0800, from dpchrist@holgerdanske.com (David Christensen):
> I suggest taking an image (backup) with dd(1), Clonezilla, etc., when you're
> done.  This will allow you to restore the image later -- to roll-back a
> change you do not like, to recovery from a disaster, to clone the image to
> another device, to facilitate experiments, (such as doing a secure erase to
> see if it resolves the SSD pending sector issue), etc..
> 
> If you also keep your system configuration files in a version control
> system, restoring an image is faster than wipe/ fresh install/ configure/
> restore data.

I would go even farther. Backups should be designed such that
recovering from a catastrophic storage failure, such as getting hit by
ransomware, unintentionally doing a destructive badblocks write test
or the sudden failure of a storage device, is possible by at most
something very similar to:

* Boot some kind of live environment
* Set up file systems on the storage device to be restored onto
  (partitioning, setting up LUKS containers, formatting, whatever else
  might be called for)
* Within the live environment, install and configure the software
  needed to access the backup (if any) (this may include things like
  cryptographic keys, access passphrases and the likes)
* Perform the restoration from the most recent backup (this is the
  part that likely will take a significant amount of time)
* Update the restored copies of /etc/fstab, /etc/crypttab and any
  other files that directly reference the partitions or file systems
  by some kind of ID (UUID, /dev/disk/by-*/*, ...)
* Reinstall the boot loader
* Reboot
* Reinstall the boot loader again from within the restored environment
  to ensure that everything relating to it is in sync

Such recovery should _not_ need to involve significant reconfiguration
of anything. Any such requirements will massively increase your time
to recovery, as I think we're seeing an example of here. And yes,
pretty much all of this could be scripted, but I strongly suspect that
few people need to do a bare-metal restore of their most recent backup
often enough for _that_ to be worth the effort to create and maintain.

Which is not to say that keeping configuration files
version-controlled cannot provide benefits anyway; but given a proper,
frequent backup regime, the benefits even of that are reduced.

-- 
Michael Kjörling                     🔗 https://michael.kjorling.se
“Remember when, on the Internet, nobody cared that you were a dog?”

[toc] | [prev] | [next] | [standalone]


#265599 — Re: reinstallation and restore after catastrophic mistake or failure; was: 1 Currently unreadable (pending) sectors How worried should I be?

FromDavid Christensen <dpchrist@holgerdanske.com>
Date2024-01-07 00:40 +0100
SubjectRe: reinstallation and restore after catastrophic mistake or failure; was: 1 Currently unreadable (pending) sectors How worried should I be?
Message-ID<HToqJ-1jE4-3@gated-at.bofh.it>
In reply to#265581
On 1/6/24 04:36, Michael Kjörling wrote:
> On 6 Jan 2024 00:37 -0800, from dpchrist@holgerdanske.com (David Christensen):
>> I suggest taking an image (backup) with dd(1), Clonezilla, etc., when you're
>> done.  This will allow you to restore the image later -- to roll-back a
>> change you do not like, to recovery from a disaster, to clone the image to
>> another device, to facilitate experiments, (such as doing a secure erase to
>> see if it resolves the SSD pending sector issue), etc..
>>
>> If you also keep your system configuration files in a version control
>> system, restoring an image is faster than wipe/ fresh install/ configure/
>> restore data.
> 
> I would go even farther. Backups should be designed such that
> recovering from a catastrophic storage failure, such as getting hit by
> ransomware, unintentionally doing a destructive badblocks write test
> or the sudden failure of a storage device, is possible by at most
> something very similar to:
> 
> * Boot some kind of live environment


I wanted more tools than what the Debian installer rescue shell provides 
(e.g. BusyBox) and I am too lazy to learn yet another live system (e.g. 
Knoppix), so I installed Debian with Xfce onto two USB drives -- one 
with BIOS/MBR and the other with secure UEFI/GPT.  They are both 
complete installs, so they are familiar and I can add whatever I want.


> * Set up file systems on the storage device to be restored onto
>    (partitioning, setting up LUKS containers, formatting, whatever else
>    might be called for)
> * Within the live environment, install and configure the software
>    needed to access the backup (if any) (this may include things like
>    cryptographic keys, access passphrases and the likes)
> * Perform the restoration from the most recent backup (this is the
>    part that likely will take a significant amount of time)


I keep my Debian instances small, simple, and self-contained (1 GB ext4 
boot, 1 GB dm-crypt swap, and 12 GB LUKS ext4 root on one 16+ GB 2.5" 
SATA SSD).  dd(1) meets all of my imaging needs.  It's fast and requires 
minimal storage -- less than 10 minutes using an old-school USB 2.0 HDD; 
each 100 GB holds 6+ images.  (`apt-get autoremove`, `apt-get 
autoclean`, fstrim(8), and/or gzip(1) can reduce time and storage 
requirements.)


If my OS instances were larger, more complex, shared disk space, etc. -- 
e.g. multi-boot Windows, Debian, etc., with a shared data partition -- 
e.g. what the OP likely had -- I would think about a tool such as 
Clonezilla.  Then I would get a big USB 3.0+ HDD/RAID, boot one of my 
Debian USB instances, look at the partition table, and take dd(1) images 
in chunks -- block 0 to the last block before ESP, the ESP, then each 
partition or contiguous span of related partitions, and finally the 
secondary GPT header.


> * Update the restored copies of /etc/fstab, /etc/crypttab and any
>    other files that directly reference the partitions or file systems
>    by some kind of ID (UUID, /dev/disk/by-*/*, ...)
> * Reinstall the boot loader


When I take a dd(1) image of an MBR disk, I copy from block 0 through 
the end of the root partition.  So:

1.  UUID's are preserved.

2.  All boot loader stages are preserved.


When I take an dd(1) image of a GPT disk with lots of zeros (fresh wipe 
and install), I copy the whole thing.  Again, UUID's and boot loader 
stages are preserved.


Using live media for UUID and/or boot loader surgery is non-trivial, as 
discussed in more than a few posts to this list.  But, such may be 
required after restoring an image onto a different disk and/or hardware 
arrangement.


> * Reboot
> * Reinstall the boot loader again from within the restored environment
>    to ensure that everything relating to it is in sync


For the simple case of restoring an image onto the exact same hardware, 
a restored MBR image just works.  Same for GPT.  If a GPT disk was 
zeroed or secure erased, a secondary GPT header will need to be needed 
written.  I believe GRUB, Linux, or something on Debian did this 
automagically for me the last time I tried.


> Such recovery should _not_ need to involve significant reconfiguration
> of anything. Any such requirements will massively increase your time
> to recovery, as I think we're seeing an example of here. And yes,
> pretty much all of this could be scripted, but I strongly suspect that
> few people need to do a bare-metal restore of their most recent backup
> often enough for _that_ to be worth the effort to create and maintain.


AIUI the OP accidentally zeroed a Windows/ Debian multi-boot disk in a 
relatively new computer.  Rebuilding from scratch is going to involve 
more than twice the effort of rebuilding one OS from scratch, but 
hopefully there was no live data lost.


I have a half dozen computers in my SOHO network.  I trash my daily 
driver at least once a year and my workhorse more often than that.


I started with disaster preparedness/ recovery using 
lowest-common-denominator tools -- tar(1), gzip(1), rsync(1), dd(1), 
etc..  I am a coder, so I wrapped those with shell and Perl scripts. 
For better or worse, I have built my own backup, recovery, image, 
archive, etc., suite and have tailored my work flow to match.  The tool 
chain is Rube Goldberg, but the backup and archive products are 
identifiable as standard Unix tool outputs and accessible by hand.


> Which is not to say that keeping configuration files
> version-controlled cannot provide benefits anyway; but given a proper,
> frequent backup regime, the benefits even of that are reduced.


The goal is defense in depth -- version control, backup, restore, 
imaging, archive, zfs-auto-snapshot, replication, rotation, RAID, etc..


David

[toc] | [prev] | [next] | [standalone]


#265508

FromMax Nikulin <manikulin@gmail.com>
Date2024-01-04 16:20 +0100
Message-ID<HSxFL-LqJ-7@gated-at.bofh.it>
In reply to#265451
On 04/01/2024 03:25, Charles Curley wrote:
> I decided
> instead to boot to a USB stick and run badblocks. The read-only test
> took 12 minutes and reported no errors.
> 
> I now have a writing test (-w) running. It has reported no failures on
> its first pass.

Is badblock writing test useful for SSD taking into account wear 
leveling? Each write should be mapped to another physical address. All 
errors should be handled by firmware.

To test low-end USB pen drives and SD cards there is the f3 (Fight Flash 
Fraud or Fight Fake Flash) tool, however such test should not be 
necessary for a SATA SSD.

Have you checked that no firmware update is available for this drive?

I have experienced just a few failures of HDD. It may be irrelevant for 
SSD, but in the case of HDD I would replace the disk reporting 
Current_Pending_Sector as soon as possible. It seems, repeating reports 
from smartd are intentional.

On the other hand the "VALUE" has not decreased and is still 100, and 
the attribute is not marked as "pre-fail". Perhaps there is not reason 
to worry to much.

I am unsure concerning accounting if an error happens during reading a 
file then the file is deleted without overwriting and the address range 
is marked unused (trimmed).

[toc] | [prev] | [next] | [standalone]


#265512

FromMichael Kjörling <2695bd53d63c@ewoof.net>
Date2024-01-04 17:50 +0100
Message-ID<HSz4R-MvE-1@gated-at.bofh.it>
In reply to#265451
On 3 Jan 2024 13:25 -0700, from charlescurley@charlescurley.com (Charles Curley):
>> As a background process, try running something like
>> 
>> # ionice find / -xdev -type f -exec cat {} + >/dev/null
> 
> That would only reach files on the partition where it is run.

I covered that on the next few lines, which you choose not to quote.

> Since there is another operating system on this drive,

That kind of information might be helpful to include up front.

> and there are parts of
> the drive normally inaccessible to any operating system,

That's true, but it would have told you about any error in any
accessible data _and_ also told you which file or directory was
affected if that was the case. If the error is in an inaccessible
portion of the drive and the system is working normally aside from a
note in SMART data, then it would stand to reason that the error would
be in an unused location; thereby not likely to affect usage (because
a known bad spot would be reallocated elsewhere by the firmware on the
next write).

Also, badblocks too will only deal with user-accessible blocks; if the
drive already had remapped a bad location, for example as a part of a
write to an identified bad block, the original error would be
invisible to badblocks even if it still was an error.


On 3 Jan 2024 16:27 -0700, from charlescurley@charlescurley.com (Charles Curley):
> OOPS! -w is the destructive test. I now have a hard drive full of 0x00s.

After restoring your most recent backup, consider doing a fstrim to
TRIM unused blocks.

-- 
Michael Kjörling                     🔗 https://michael.kjorling.se
“Remember when, on the Internet, nobody cared that you were a dog?”

[toc] | [prev] | [next] | [standalone]


#265441

FromAndy Smith <andy@strugglers.net>
Date2024-01-03 14:30 +0100
Message-ID<HS9tL-vJK-3@gated-at.bofh.it>
In reply to#265433
Hi,

On Tue, Jan 02, 2024 at 08:17:55PM -0700, Charles Curley wrote:
> On Wed, 3 Jan 2024 00:29:42 +0000
> Andy Smith <andy@strugglers.net> wrote:
> > If a SMART long self-test came back clean then it already has been
> > re-mapped as a long self-test reads every user-accessible sector.
> 
> I'm not so sure about that. See the journalctl output at the bottom of
> this email.

I don't see anything but smartd repeatedly warning you about the 1
pending sector.

As I said, it's annoying when a drive doesn't decremement its
pending sector count after remapping. If you can read the whole
drive then it certainly has been remapped (or was a transient error
that isn't "pending" any more).

None of the logs you presented show any error coming from the drive,
but then they won't as you've only selected logs from smartd. smartd
will complain about that 1 pending sector count until the end of
time unless:

- Drive just decides to clear it, or;

- You reconfigure smartd

All smartd is doing here is reading the attributes of the drive and
reporting them to you. It will never show you the actual error that
caused those attributes to change.

> > Like to live dangerously, huh…
> 
> No. That's what fast networks, good and multiple backup programs, a
> good RAID array on another computer, and multiple off-site backups are
> for.

It's not my view as in my experience storage is one of the most
failure-prone parts of a computer, and an outage from non-redundant
storage typically annoys me more than making it redundant does.
That is in most cases really easy and cheap these days, so I do
regard going without it as living dangerously. Not always a good
cost-benefit trade off though, I grant you.

Thanks,
Andy

-- 
https://bitfolk.com/ -- No-nonsense VPS hosting

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.debian.user


csiph-web