Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #265426 > unrolled thread
| Started by | Charles Curley <charlescurley@charlescurley.com> |
|---|---|
| First post | 2024-01-02 23:50 +0100 |
| Last post | 2024-01-03 14:30 +0100 |
| Articles | 5 on this page of 25 — 9 participants |
Back to article view | Back to linux.debian.user
1 Currently unreadable (pending) sectors How worried should I be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-02 23:50 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Dan Ritter <dsr@randomstring.org> - 2024-01-03 00:10 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-03 00:40 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Dan Ritter <dsr@randomstring.org> - 2024-01-03 02:00 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Tixy <tixy@yxit.co.uk> - 2024-01-03 08:50 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Dan Purgert <dan@djph.net> - 2024-01-03 00:10 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-03 00:50 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Andy Smith <andy@strugglers.net> - 2024-01-03 01:40 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-03 04:20 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-03 12:10 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-03 21:30 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-04 00:30 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? <tomas@tuxteam.de> - 2024-01-04 12:00 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-04 15:10 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Andy Smith <andy@strugglers.net> - 2024-01-05 22:10 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-06 00:30 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? David Christensen <dpchrist@holgerdanske.com> - 2024-01-06 02:30 +0100
Secure erase [was: Re: 1 Currently unreadable (pending) sectors How worried should I be?] Max Nikulin <manikulin@gmail.com> - 2024-01-06 03:50 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Charles Curley <charlescurley@charlescurley.com> - 2024-01-06 06:20 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? David Christensen <dpchrist@holgerdanske.com> - 2024-01-06 09:40 +0100
Re: reinstallation and restore after catastrophic mistake or failure; was: 1 Currently unreadable (pending) sectors How worried should I be? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-06 13:40 +0100
Re: reinstallation and restore after catastrophic mistake or failure; was: 1 Currently unreadable (pending) sectors How worried should I be? David Christensen <dpchrist@holgerdanske.com> - 2024-01-07 00:40 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Max Nikulin <manikulin@gmail.com> - 2024-01-04 16:20 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Michael Kjörling <2695bd53d63c@ewoof.net> - 2024-01-04 17:50 +0100
Re: 1 Currently unreadable (pending) sectors How worried should I be? Andy Smith <andy@strugglers.net> - 2024-01-03 14:30 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Michael Kjörling <2695bd53d63c@ewoof.net> |
|---|---|
| Date | 2024-01-06 13:40 +0100 |
| Subject | Re: reinstallation and restore after catastrophic mistake or failure; was: 1 Currently unreadable (pending) sectors How worried should I be? |
| Message-ID | <HTe82-1deA-7@gated-at.bofh.it> |
| In reply to | #265573 |
On 6 Jan 2024 00:37 -0800, from dpchrist@holgerdanske.com (David Christensen): > I suggest taking an image (backup) with dd(1), Clonezilla, etc., when you're > done. This will allow you to restore the image later -- to roll-back a > change you do not like, to recovery from a disaster, to clone the image to > another device, to facilitate experiments, (such as doing a secure erase to > see if it resolves the SSD pending sector issue), etc.. > > If you also keep your system configuration files in a version control > system, restoring an image is faster than wipe/ fresh install/ configure/ > restore data. I would go even farther. Backups should be designed such that recovering from a catastrophic storage failure, such as getting hit by ransomware, unintentionally doing a destructive badblocks write test or the sudden failure of a storage device, is possible by at most something very similar to: * Boot some kind of live environment * Set up file systems on the storage device to be restored onto (partitioning, setting up LUKS containers, formatting, whatever else might be called for) * Within the live environment, install and configure the software needed to access the backup (if any) (this may include things like cryptographic keys, access passphrases and the likes) * Perform the restoration from the most recent backup (this is the part that likely will take a significant amount of time) * Update the restored copies of /etc/fstab, /etc/crypttab and any other files that directly reference the partitions or file systems by some kind of ID (UUID, /dev/disk/by-*/*, ...) * Reinstall the boot loader * Reboot * Reinstall the boot loader again from within the restored environment to ensure that everything relating to it is in sync Such recovery should _not_ need to involve significant reconfiguration of anything. Any such requirements will massively increase your time to recovery, as I think we're seeing an example of here. And yes, pretty much all of this could be scripted, but I strongly suspect that few people need to do a bare-metal restore of their most recent backup often enough for _that_ to be worth the effort to create and maintain. Which is not to say that keeping configuration files version-controlled cannot provide benefits anyway; but given a proper, frequent backup regime, the benefits even of that are reduced. -- Michael Kjörling 🔗 https://michael.kjorling.se “Remember when, on the Internet, nobody cared that you were a dog?”
[toc] | [prev] | [next] | [standalone]
| From | David Christensen <dpchrist@holgerdanske.com> |
|---|---|
| Date | 2024-01-07 00:40 +0100 |
| Subject | Re: reinstallation and restore after catastrophic mistake or failure; was: 1 Currently unreadable (pending) sectors How worried should I be? |
| Message-ID | <HToqJ-1jE4-3@gated-at.bofh.it> |
| In reply to | #265581 |
On 1/6/24 04:36, Michael Kjörling wrote: > On 6 Jan 2024 00:37 -0800, from dpchrist@holgerdanske.com (David Christensen): >> I suggest taking an image (backup) with dd(1), Clonezilla, etc., when you're >> done. This will allow you to restore the image later -- to roll-back a >> change you do not like, to recovery from a disaster, to clone the image to >> another device, to facilitate experiments, (such as doing a secure erase to >> see if it resolves the SSD pending sector issue), etc.. >> >> If you also keep your system configuration files in a version control >> system, restoring an image is faster than wipe/ fresh install/ configure/ >> restore data. > > I would go even farther. Backups should be designed such that > recovering from a catastrophic storage failure, such as getting hit by > ransomware, unintentionally doing a destructive badblocks write test > or the sudden failure of a storage device, is possible by at most > something very similar to: > > * Boot some kind of live environment I wanted more tools than what the Debian installer rescue shell provides (e.g. BusyBox) and I am too lazy to learn yet another live system (e.g. Knoppix), so I installed Debian with Xfce onto two USB drives -- one with BIOS/MBR and the other with secure UEFI/GPT. They are both complete installs, so they are familiar and I can add whatever I want. > * Set up file systems on the storage device to be restored onto > (partitioning, setting up LUKS containers, formatting, whatever else > might be called for) > * Within the live environment, install and configure the software > needed to access the backup (if any) (this may include things like > cryptographic keys, access passphrases and the likes) > * Perform the restoration from the most recent backup (this is the > part that likely will take a significant amount of time) I keep my Debian instances small, simple, and self-contained (1 GB ext4 boot, 1 GB dm-crypt swap, and 12 GB LUKS ext4 root on one 16+ GB 2.5" SATA SSD). dd(1) meets all of my imaging needs. It's fast and requires minimal storage -- less than 10 minutes using an old-school USB 2.0 HDD; each 100 GB holds 6+ images. (`apt-get autoremove`, `apt-get autoclean`, fstrim(8), and/or gzip(1) can reduce time and storage requirements.) If my OS instances were larger, more complex, shared disk space, etc. -- e.g. multi-boot Windows, Debian, etc., with a shared data partition -- e.g. what the OP likely had -- I would think about a tool such as Clonezilla. Then I would get a big USB 3.0+ HDD/RAID, boot one of my Debian USB instances, look at the partition table, and take dd(1) images in chunks -- block 0 to the last block before ESP, the ESP, then each partition or contiguous span of related partitions, and finally the secondary GPT header. > * Update the restored copies of /etc/fstab, /etc/crypttab and any > other files that directly reference the partitions or file systems > by some kind of ID (UUID, /dev/disk/by-*/*, ...) > * Reinstall the boot loader When I take a dd(1) image of an MBR disk, I copy from block 0 through the end of the root partition. So: 1. UUID's are preserved. 2. All boot loader stages are preserved. When I take an dd(1) image of a GPT disk with lots of zeros (fresh wipe and install), I copy the whole thing. Again, UUID's and boot loader stages are preserved. Using live media for UUID and/or boot loader surgery is non-trivial, as discussed in more than a few posts to this list. But, such may be required after restoring an image onto a different disk and/or hardware arrangement. > * Reboot > * Reinstall the boot loader again from within the restored environment > to ensure that everything relating to it is in sync For the simple case of restoring an image onto the exact same hardware, a restored MBR image just works. Same for GPT. If a GPT disk was zeroed or secure erased, a secondary GPT header will need to be needed written. I believe GRUB, Linux, or something on Debian did this automagically for me the last time I tried. > Such recovery should _not_ need to involve significant reconfiguration > of anything. Any such requirements will massively increase your time > to recovery, as I think we're seeing an example of here. And yes, > pretty much all of this could be scripted, but I strongly suspect that > few people need to do a bare-metal restore of their most recent backup > often enough for _that_ to be worth the effort to create and maintain. AIUI the OP accidentally zeroed a Windows/ Debian multi-boot disk in a relatively new computer. Rebuilding from scratch is going to involve more than twice the effort of rebuilding one OS from scratch, but hopefully there was no live data lost. I have a half dozen computers in my SOHO network. I trash my daily driver at least once a year and my workhorse more often than that. I started with disaster preparedness/ recovery using lowest-common-denominator tools -- tar(1), gzip(1), rsync(1), dd(1), etc.. I am a coder, so I wrapped those with shell and Perl scripts. For better or worse, I have built my own backup, recovery, image, archive, etc., suite and have tailored my work flow to match. The tool chain is Rube Goldberg, but the backup and archive products are identifiable as standard Unix tool outputs and accessible by hand. > Which is not to say that keeping configuration files > version-controlled cannot provide benefits anyway; but given a proper, > frequent backup regime, the benefits even of that are reduced. The goal is defense in depth -- version control, backup, restore, imaging, archive, zfs-auto-snapshot, replication, rotation, RAID, etc.. David
[toc] | [prev] | [next] | [standalone]
| From | Max Nikulin <manikulin@gmail.com> |
|---|---|
| Date | 2024-01-04 16:20 +0100 |
| Message-ID | <HSxFL-LqJ-7@gated-at.bofh.it> |
| In reply to | #265451 |
On 04/01/2024 03:25, Charles Curley wrote: > I decided > instead to boot to a USB stick and run badblocks. The read-only test > took 12 minutes and reported no errors. > > I now have a writing test (-w) running. It has reported no failures on > its first pass. Is badblock writing test useful for SSD taking into account wear leveling? Each write should be mapped to another physical address. All errors should be handled by firmware. To test low-end USB pen drives and SD cards there is the f3 (Fight Flash Fraud or Fight Fake Flash) tool, however such test should not be necessary for a SATA SSD. Have you checked that no firmware update is available for this drive? I have experienced just a few failures of HDD. It may be irrelevant for SSD, but in the case of HDD I would replace the disk reporting Current_Pending_Sector as soon as possible. It seems, repeating reports from smartd are intentional. On the other hand the "VALUE" has not decreased and is still 100, and the attribute is not marked as "pre-fail". Perhaps there is not reason to worry to much. I am unsure concerning accounting if an error happens during reading a file then the file is deleted without overwriting and the address range is marked unused (trimmed).
[toc] | [prev] | [next] | [standalone]
| From | Michael Kjörling <2695bd53d63c@ewoof.net> |
|---|---|
| Date | 2024-01-04 17:50 +0100 |
| Message-ID | <HSz4R-MvE-1@gated-at.bofh.it> |
| In reply to | #265451 |
On 3 Jan 2024 13:25 -0700, from charlescurley@charlescurley.com (Charles Curley):
>> As a background process, try running something like
>>
>> # ionice find / -xdev -type f -exec cat {} + >/dev/null
>
> That would only reach files on the partition where it is run.
I covered that on the next few lines, which you choose not to quote.
> Since there is another operating system on this drive,
That kind of information might be helpful to include up front.
> and there are parts of
> the drive normally inaccessible to any operating system,
That's true, but it would have told you about any error in any
accessible data _and_ also told you which file or directory was
affected if that was the case. If the error is in an inaccessible
portion of the drive and the system is working normally aside from a
note in SMART data, then it would stand to reason that the error would
be in an unused location; thereby not likely to affect usage (because
a known bad spot would be reallocated elsewhere by the firmware on the
next write).
Also, badblocks too will only deal with user-accessible blocks; if the
drive already had remapped a bad location, for example as a part of a
write to an identified bad block, the original error would be
invisible to badblocks even if it still was an error.
On 3 Jan 2024 16:27 -0700, from charlescurley@charlescurley.com (Charles Curley):
> OOPS! -w is the destructive test. I now have a hard drive full of 0x00s.
After restoring your most recent backup, consider doing a fstrim to
TRIM unused blocks.
--
Michael Kjörling 🔗 https://michael.kjorling.se
“Remember when, on the Internet, nobody cared that you were a dog?”
[toc] | [prev] | [next] | [standalone]
| From | Andy Smith <andy@strugglers.net> |
|---|---|
| Date | 2024-01-03 14:30 +0100 |
| Message-ID | <HS9tL-vJK-3@gated-at.bofh.it> |
| In reply to | #265433 |
Hi, On Tue, Jan 02, 2024 at 08:17:55PM -0700, Charles Curley wrote: > On Wed, 3 Jan 2024 00:29:42 +0000 > Andy Smith <andy@strugglers.net> wrote: > > If a SMART long self-test came back clean then it already has been > > re-mapped as a long self-test reads every user-accessible sector. > > I'm not so sure about that. See the journalctl output at the bottom of > this email. I don't see anything but smartd repeatedly warning you about the 1 pending sector. As I said, it's annoying when a drive doesn't decremement its pending sector count after remapping. If you can read the whole drive then it certainly has been remapped (or was a transient error that isn't "pending" any more). None of the logs you presented show any error coming from the drive, but then they won't as you've only selected logs from smartd. smartd will complain about that 1 pending sector count until the end of time unless: - Drive just decides to clear it, or; - You reconfigure smartd All smartd is doing here is reading the attributes of the drive and reporting them to you. It will never show you the actual error that caused those attributes to change. > > Like to live dangerously, huh… > > No. That's what fast networks, good and multiple backup programs, a > good RAID array on another computer, and multiple off-site backups are > for. It's not my view as in my experience storage is one of the most failure-prone parts of a computer, and an outage from non-redundant storage typically annoys me more than making it redundant does. That is in most cases really easy and cheap these days, so I do regard going without it as living dangerously. Not always a good cost-benefit trade off though, I grant you. Thanks, Andy -- https://bitfolk.com/ -- No-nonsense VPS hosting
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.debian.user
csiph-web