Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #178463 > unrolled thread
| Started by | Gregory Seidman <gsslist+debian@anthropohedron.net> |
|---|---|
| First post | 2017-03-05 22:20 +0100 |
| Last post | 2017-03-06 12:50 +0100 |
| Articles | 7 — 6 participants |
Back to article view | Back to linux.debian.user
Failing disk advice Gregory Seidman <gsslist+debian@anthropohedron.net> - 2017-03-05 22:20 +0100
Re: Failing disk advice The Wanderer <wanderer@fastmail.fm> - 2017-03-05 23:50 +0100
Re: Failing disk advice David Christensen <dpchrist@holgerdanske.com> - 2017-03-06 05:40 +0100
Re: Failing disk advice Mirko Parthey <mirko.parthey@web.de> - 2017-03-06 12:20 +0100
Re: Failing disk advice Gregory Seidman <gsslist+debian@anthropohedron.net> - 2017-03-07 02:30 +0100
Re: Failing disk advice Patrick Zaloum <pzaloum@gmail.com> - 2017-03-07 17:30 +0100
Re: Failing disk advice Andy Smith <andy@strugglers.net> - 2017-03-06 12:50 +0100
| From | Gregory Seidman <gsslist+debian@anthropohedron.net> |
|---|---|
| Date | 2017-03-05 22:20 +0100 |
| Subject | Failing disk advice |
| Message-ID | <thLJo-6ax-11@gated-at.bofh.it> |
I have a disk that is reporting SMART errors. It is an active disk in a (kernel, not hardware) RAID1 configuration. I also have a hot spare in the RAID1, and md hasn't decided it should fail the disk and switch to the hot spare. Should I proactively tell md to fail the disk (and let the hot spare take over), or should I just wait until md notices a problem? --Greg
[toc] | [next] | [standalone]
| From | The Wanderer <wanderer@fastmail.fm> |
|---|---|
| Date | 2017-03-05 23:50 +0100 |
| Message-ID | <thN8t-76h-5@gated-at.bofh.it> |
| In reply to | #178463 |
[Multipart message — attachments visible in raw view] — view raw
On 2017-03-05 at 16:02, Gregory Seidman wrote: > I have a disk that is reporting SMART errors. It is an active disk in > a (kernel, not hardware) RAID1 configuration. I also have a hot spare > in the RAID1, and md hasn't decided it should fail the disk and > switch to the hot spare. Should I proactively tell md to fail the > disk (and let the hot spare take over), or should I just wait until > md notices a problem? So, you're saying you have a two-disk RAID-1 array with a third disk as hot spare? Under those circumstances, I would be inclined to leave it alone until either md fails the one disk out or I start noticing visible symptoms, but I'm not an expert and I'm not sure what the best practice is. Certainly the paranoid, better-be-safe-than-sorry approach would be to fail it out, let the hot spare take over, then swap in a cold spare as the new hot spare. For my own main array (RAID-5 with no hot spares), I don't necessarily replace the disk as soon as I start noticing SMART errors - but I do start monitoring the situation more closely, and as soon as I start to see other indications (most prominently read- or write-related notices from dmesg), I arrange for a replacement, fail out the disk, and swap in the new one. (The only reason I have no hot spare is that there are no unused SATA ports in the system.) I initially expected that I would not fail out the disk manually at all, but the last time I saw drive errors md was not automatically failing a disk out of the array even when that disk was exhibiting read issues so severe that the entire UI was hanging for 15-to-60 seconds on any read attempt against the failed portions of the disk. -- The Wanderer The reasonable man adapts himself to the world; the unreasonable one persists in trying to adapt the world to himself. Therefore all progress depends on the unreasonable man. -- George Bernard Shaw
[toc] | [prev] | [next] | [standalone]
| From | David Christensen <dpchrist@holgerdanske.com> |
|---|---|
| Date | 2017-03-06 05:40 +0100 |
| Message-ID | <thSBc-2yj-5@gated-at.bofh.it> |
| In reply to | #178463 |
On 03/05/2017 01:02 PM, Gregory Seidman wrote: > I have a disk that is reporting SMART errors. It is an active disk in a > (kernel, not hardware) RAID1 configuration. I also have a hot spare in the > RAID1, and md hasn't decided it should fail the disk and switch to the hot > spare. Should I proactively tell md to fail the disk (and let the hot spare > take over), or should I just wait until md notices a problem? AFAIK desktop disks and "enterprise RAID" disks degrade differently. When a desktop disk is having trouble reading a sector, it will retry many times before giving up because it is likely the data does not exist anywhere else. But, an enterprise RAID disc will retry only a few times and then fail; because the data should exist elsewhere and hung reads are intolerable in enterprise environments. So, if you are using desktop disks in a RAID, you might need to manually intervene to compensate for the mismatch. I'm confused by "I also have a hot spare in the RAID1". Do you have a two-member RAID1 with a hot spare, or a three-member RAID1? I would prefer the latter: https://manpages.debian.org/jessie/mdadm/md.4.en.html If you're planning on buying a fourth disk and adding it after fixing the RAID, can you add it now as a fourth RAID1 member, let it resilver, remove the failing disk from the RAID (e.g. reconfigure as three-member RAID1), and then pull the failing disk? David
[toc] | [prev] | [next] | [standalone]
| From | Mirko Parthey <mirko.parthey@web.de> |
|---|---|
| Date | 2017-03-06 12:20 +0100 |
| Message-ID | <thYQi-7rD-19@gated-at.bofh.it> |
| In reply to | #178481 |
On Sun, Mar 05, 2017 at 08:38:27PM -0800, David Christensen wrote: > On 03/05/2017 01:02 PM, Gregory Seidman wrote: > >I have a disk that is reporting SMART errors. It is an active disk in a > >(kernel, not hardware) RAID1 configuration. I also have a hot spare in the > >RAID1, and md hasn't decided it should fail the disk and switch to the hot > >spare. Should I proactively tell md to fail the disk (and let the hot spare > >take over), or should I just wait until md notices a problem? > > I'm confused by "I also have a hot spare in the RAID1". Do you have a > two-member RAID1 with a hot spare, or a three-member RAID1? I would prefer > the latter: > > https://manpages.debian.org/jessie/mdadm/md.4.en.html Refining this advice a bit, I would convert the spare to a full RAID member now, without explicitly failing the disk that reports SMART errors first. Assuming you have a two-member RAID1 with a hot spare, the command should be similar to this (untested): mdadm -G /dev/mdX -n 3 This ensures you keep redundancy during further maintenance actions. Which SMART errors do you get, and who reports them? What is the output of the following command for the failing drive? smartctl -A /dev/sdY Regards, Mirko
[toc] | [prev] | [next] | [standalone]
| From | Gregory Seidman <gsslist+debian@anthropohedron.net> |
|---|---|
| Date | 2017-03-07 02:30 +0100 |
| Message-ID | <tic6S-8lE-11@gated-at.bofh.it> |
| In reply to | #178488 |
On Mon, Mar 06, 2017 at 12:17:03PM +0100, Mirko Parthey wrote: > On Sun, Mar 05, 2017 at 08:38:27PM -0800, David Christensen wrote: > > On 03/05/2017 01:02 PM, Gregory Seidman wrote: > > >I have a disk that is reporting SMART errors. It is an active disk in > > >a (kernel, not hardware) RAID1 configuration. I also have a hot spare > > >in the RAID1, and md hasn't decided it should fail the disk and switch > > >to the hot spare. Should I proactively tell md to fail the disk (and > > >let the hot spare take over), or should I just wait until md notices a > > >problem? > > > > I'm confused by "I also have a hot spare in the RAID1". Do you have a > > two-member RAID1 with a hot spare, or a three-member RAID1? I would > > prefer the latter: > > > > https://manpages.debian.org/jessie/mdadm/md.4.en.html > > Refining this advice a bit, I would convert the spare to a full RAID > member now, without explicitly failing the disk that reports SMART > errors first. > Assuming you have a two-member RAID1 with a hot spare, the command > should be similar to this (untested): > mdadm -G /dev/mdX -n 3 > This ensures you keep redundancy during further maintenance actions. I was unaware that this was possible. I've run it and mdadm -D reports that it is now in the "clean, degraded, rebuilding" state. Thank you! I wish I had room in my system to add the fourth (which I've ordered) without removing the failing disk, but I do not. > Which SMART errors do you get, and who reports them? I get emails sent to root: This message was generated by the smartd daemon running on: host name: XXXXXX DNS domain: YYYYYY The following warning/error was logged by the smartd daemon: Device: /dev/sdc [SAT], 8 Currently unreadable (pending) sectors Device info: ST31500341AS, S/N:9VS43CV9, WWN:5-000c50-0208aa9a3, FW:CC1H, 1.50 TB For details see host's SYSLOG. You can also use the smartctl utility for further investigation. The original message about this issue was sent at Wed Dec 14 00:51:36 2016 EST Another message will be sent in 24 hours if the problem persists. ...and... This message was generated by the smartd daemon running on: host name: XXXXXX DNS domain: YYYYYY The following warning/error was logged by the smartd daemon: Device: /dev/sdc [SAT], 8 Offline uncorrectable sectors Device info: ST31500341AS, S/N:9VS43CV9, WWN:5-000c50-0208aa9a3, FW:CC1H, 1.50 TB For details see host's SYSLOG. You can also use the smartctl utility for further investigation. The original message about this issue was sent at Wed Dec 14 00:51:37 2016 EST Another message will be sent in 24 hours if the problem persists. (Yes, I know, I've been letting it do this since mid-December, which is not great.) > What is the output of the following command for the failing drive? > smartctl -A /dev/sdY # smartctl -A /dev/sdc smartctl 6.4 2014-10-07 r4002 [i686-linux-3.16.0-4-686-pae] (local build) Copyright (C) 2002-14, Bruce Allen, Christian Franke, www.smartmontools.org === START OF READ SMART DATA SECTION === SMART Attributes Data Structure revision number: 10 Vendor Specific SMART Attributes with Thresholds: ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE 1 Raw_Read_Error_Rate 0x000f 119 099 006 Pre-fail Always - 205161943 3 Spin_Up_Time 0x0003 100 091 000 Pre-fail Always - 0 4 Start_Stop_Count 0x0032 099 099 020 Old_age Always - 1055 5 Reallocated_Sector_Ct 0x0033 099 099 036 Pre-fail Always - 41 7 Seek_Error_Rate 0x000f 092 060 030 Pre-fail Always - 1743842168 9 Power_On_Hours 0x0032 039 039 000 Old_age Always - 53898 10 Spin_Retry_Count 0x0013 100 100 097 Pre-fail Always - 0 12 Power_Cycle_Count 0x0032 100 100 020 Old_age Always - 85 184 End-to-End_Error 0x0032 100 100 099 Old_age Always - 0 187 Reported_Uncorrect 0x0032 097 097 000 Old_age Always - 3 188 Command_Timeout 0x0032 100 098 000 Old_age Always - 133146017827 189 High_Fly_Writes 0x003a 007 007 000 Old_age Always - 93 190 Airflow_Temperature_Cel 0x0022 060 040 045 Old_age Always In_the_past 40 (Min/Max 26/45 #502) 194 Temperature_Celsius 0x0022 040 060 000 Old_age Always - 40 (0 18 0 0 0) 195 Hardware_ECC_Recovered 0x001a 038 023 000 Old_age Always - 205161943 197 Current_Pending_Sector 0x0012 100 100 000 Old_age Always - 8 198 Offline_Uncorrectable 0x0010 100 100 000 Old_age Offline - 8 199 UDMA_CRC_Error_Count 0x003e 200 200 000 Old_age Always - 1 240 Head_Flying_Hours 0x0000 100 253 000 Old_age Offline - 53897 (15 186 0) 241 Total_LBAs_Written 0x0000 100 253 000 Old_age Offline - 917595486 242 Total_LBAs_Read 0x0000 100 253 000 Old_age Offline - 1262569510 > Regards, > Mirko Thanks for the help so far, --Greg
[toc] | [prev] | [next] | [standalone]
| From | Patrick Zaloum <pzaloum@gmail.com> |
|---|---|
| Date | 2017-03-07 17:30 +0100 |
| Message-ID | <tiq9P-1Dp-1@gated-at.bofh.it> |
| In reply to | #178512 |
[Multipart message — attachments visible in raw view] — view raw
In my experience, you will likely be able to pull a few more weeks / months of life out of the drive but it will die. Mirko's suggestion of migrating to a n=3 raid1 setup is also what I would recommend. You will notice in your smartctl output that Reallocated_Sector_Ct is 41. That means that there have already been 41 sectors remapped to the spare sectors of your drive. The 8 offline_uncorrectable / current_pending_sector are probably unpopulated sectors that haven't been rewritten to yet, but triggered an i/o error the last time there was data there. For me this was often on a swap partition since there are a lot of transient writes. The next time the system tries to write to those sectors it will either fail and mark it as permanently unusable, or succeed and clear the pending count. Good luck Patrick On Mon, Mar 6, 2017 at 8:27 PM, Gregory Seidman < gsslist+debian@anthropohedron.net> wrote: > On Mon, Mar 06, 2017 at 12:17:03PM +0100, Mirko Parthey wrote: > > On Sun, Mar 05, 2017 at 08:38:27PM -0800, David Christensen wrote: > > > On 03/05/2017 01:02 PM, Gregory Seidman wrote: > > > >I have a disk that is reporting SMART errors. It is an active disk in > > > >a (kernel, not hardware) RAID1 configuration. I also have a hot spare > > > >in the RAID1, and md hasn't decided it should fail the disk and switch > > > >to the hot spare. Should I proactively tell md to fail the disk (and > > > >let the hot spare take over), or should I just wait until md notices a > > > >problem? > > > > > > I'm confused by "I also have a hot spare in the RAID1". Do you have a > > > two-member RAID1 with a hot spare, or a three-member RAID1? I would > > > prefer the latter: > > > > > > https://manpages.debian.org/jessie/mdadm/md.4.en.html > > > > Refining this advice a bit, I would convert the spare to a full RAID > > member now, without explicitly failing the disk that reports SMART > > errors first. > > Assuming you have a two-member RAID1 with a hot spare, the command > > should be similar to this (untested): > > mdadm -G /dev/mdX -n 3 > > This ensures you keep redundancy during further maintenance actions. > > I was unaware that this was possible. I've run it and mdadm -D reports that > it is now in the "clean, degraded, rebuilding" state. Thank you! I wish I > had room in my system to add the fourth (which I've ordered) without > removing the failing disk, but I do not. > > > Which SMART errors do you get, and who reports them? > > I get emails sent to root: > > This message was generated by the smartd daemon running on: > > host name: XXXXXX > DNS domain: YYYYYY > > The following warning/error was logged by the smartd daemon: > > Device: /dev/sdc [SAT], 8 Currently unreadable (pending) sectors > > Device info: > ST31500341AS, S/N:9VS43CV9, WWN:5-000c50-0208aa9a3, FW:CC1H, 1.50 > TB > > For details see host's SYSLOG. > > You can also use the smartctl utility for further investigation. > The original message about this issue was sent at Wed Dec 14 > 00:51:36 2016 EST > Another message will be sent in 24 hours if the problem persists. > > ...and... > > This message was generated by the smartd daemon running on: > > host name: XXXXXX > DNS domain: YYYYYY > > The following warning/error was logged by the smartd daemon: > > Device: /dev/sdc [SAT], 8 Offline uncorrectable sectors > > Device info: > ST31500341AS, S/N:9VS43CV9, WWN:5-000c50-0208aa9a3, FW:CC1H, 1.50 > TB > > For details see host's SYSLOG. > > You can also use the smartctl utility for further investigation. > The original message about this issue was sent at Wed Dec 14 > 00:51:37 2016 EST > Another message will be sent in 24 hours if the problem persists. > > (Yes, I know, I've been letting it do this since mid-December, which is not > great.) > > > What is the output of the following command for the failing drive? > > smartctl -A /dev/sdY > > # smartctl -A /dev/sdc > smartctl 6.4 2014-10-07 r4002 [i686-linux-3.16.0-4-686-pae] (local > build) > Copyright (C) 2002-14, Bruce Allen, Christian Franke, > www.smartmontools.org > > === START OF READ SMART DATA SECTION === > SMART Attributes Data Structure revision number: 10 > Vendor Specific SMART Attributes with Thresholds: > ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE > UPDATED WHEN_FAILED RAW_VALUE > 1 Raw_Read_Error_Rate 0x000f 119 099 006 Pre-fail > Always - 205161943 > 3 Spin_Up_Time 0x0003 100 091 000 Pre-fail > Always - 0 > 4 Start_Stop_Count 0x0032 099 099 020 Old_age > Always - 1055 > 5 Reallocated_Sector_Ct 0x0033 099 099 036 Pre-fail > Always - 41 > 7 Seek_Error_Rate 0x000f 092 060 030 Pre-fail > Always - 1743842168 > 9 Power_On_Hours 0x0032 039 039 000 Old_age > Always - 53898 > 10 Spin_Retry_Count 0x0013 100 100 097 Pre-fail > Always - 0 > 12 Power_Cycle_Count 0x0032 100 100 020 Old_age > Always - 85 > 184 End-to-End_Error 0x0032 100 100 099 Old_age > Always - 0 > 187 Reported_Uncorrect 0x0032 097 097 000 Old_age > Always - 3 > 188 Command_Timeout 0x0032 100 098 000 Old_age > Always - 133146017827 > 189 High_Fly_Writes 0x003a 007 007 000 Old_age > Always - 93 > 190 Airflow_Temperature_Cel 0x0022 060 040 045 Old_age > Always In_the_past 40 (Min/Max 26/45 #502) > 194 Temperature_Celsius 0x0022 040 060 000 Old_age > Always - 40 (0 18 0 0 0) > 195 Hardware_ECC_Recovered 0x001a 038 023 000 Old_age > Always - 205161943 > 197 Current_Pending_Sector 0x0012 100 100 000 Old_age > Always - 8 > 198 Offline_Uncorrectable 0x0010 100 100 000 Old_age > Offline - 8 > 199 UDMA_CRC_Error_Count 0x003e 200 200 000 Old_age > Always - 1 > 240 Head_Flying_Hours 0x0000 100 253 000 Old_age > Offline - 53897 (15 186 0) > 241 Total_LBAs_Written 0x0000 100 253 000 Old_age > Offline - 917595486 > 242 Total_LBAs_Read 0x0000 100 253 000 Old_age > Offline - 1262569510 > > > Regards, > > Mirko > > Thanks for the help so far, > --Greg > >
[toc] | [prev] | [next] | [standalone]
| From | Andy Smith <andy@strugglers.net> |
|---|---|
| Date | 2017-03-06 12:50 +0100 |
| Message-ID | <thZjl-7C0-21@gated-at.bofh.it> |
| In reply to | #178481 |
Hello,
On Sun, Mar 05, 2017 at 08:38:27PM -0800, David Christensen wrote:
> On 03/05/2017 01:02 PM, Gregory Seidman wrote:
> >I have a disk that is reporting SMART errors.
What are the errors? Some are more serious, some less so.
> >It is an active disk in a (kernel, not hardware) RAID1
> >configuration. I also have a hot spare in the RAID1, and md
> >hasn't decided it should fail the disk and switch to the hot
> >spare. Should I proactively tell md to fail the disk (and let the
> >hot spare take over), or should I just wait until md notices a
> >problem?
>
> AFAIK desktop disks and "enterprise RAID" disks degrade differently.
> When a desktop disk is having trouble reading a sector, it will retry
> many times before giving up because it is likely the data does not
> exist anywhere else. But, an enterprise RAID disc will retry only a
> few times and then fail; because the data should exist elsewhere and
> hung reads are intolerable in enterprise environments.
What you're referring to here is SCT Error Recovery Control:
https://en.wikipedia.org/wiki/Error_recovery_control
At one point it was common for it to be a configurable timeout on
most drives, but defaulting to disabled on drive models designed for
desktop use. As you say, the rationale would be that a desktop drive
was probably not in a RAID, so holds the only copy of the data, and
must go to heroic lengths if necessary to read data.
As the drive vendors started being more aggressive about segmenting
their product ranges into "desktop" and "enterprise", they removed
the ability to change the timeout from drives in their desktop ranges.
This has had a very bad side effect for those using desktop drives
in their RAIDs. When SCTERC is not configurable, the timeout is
usually longer than Linux's own block layer timeout. The drive will
be unresponsive for so long that Linux will think the link has died
and reset it or the whole controller. That can cause multiple drives
to be kicked from the MD array though there is nothing wrong with
them, leading to the array becoming inoperable.
This is probably the number one cause of "my array broke and won't
assemble again" posts to linux-raid and so the first question asked
is usually, "what are your timeouts set to?"
It is imperative that anyone using MD RAID checks that their drive
timeouts are set sensibly.
You can check a drive's timeout like this:
# smartctl -l scterc /dev/sda
smartctl 6.4 2014-10-07 r4002 [x86_64-linux-3.16.0-4-amd64] (local build)
Copyright (C) 2002-14, Bruce Allen, Christian Franke, www.smartmontools.org
SCT Error Recovery Control:
Read: 70 (7.0 seconds)
Write: 70 (7.0 seconds)
If it comes back like this:
SCT Error Recovery Control:
Read: Disabled
Write: Disabled
then it means that SCTERC is supported but disabled, so just needs
setting, like so:
# smartctl -q errorsonly -l scterc,70,70 /dev/sda
but if it comes back like:
Warning: device does not support SCT Error Recovery Control command
then you have a problem as the drive does not support SCTERC and
will likely freeze up for several minutes trying to read a damaged
sector.
If you have drives that don't support SCTERC, and you can't replace
them for ones that do, then your next best course of action is to
increase Linux's own timeouts. 180 seconds seems to be enough:
# echo 180 > /sys/block/sda/device/timeout
The drive will still seem to freeze up for minutes when encountering
an unreadable sector, but Linux will give it longer and you'll avoid
a link/controller reset that could affect other drives.
If you needed to set SCTERC or Linux drive timeout then you must
re-apply those settings at every boot.
> So, if you are using desktop disks in a RAID, you might need to
> manually intervene to compensate for the mismatch.
Adjusting the timeouts is normally all that would be necessary.
If I had a drive that had SCTERC unsupported and it started showing
signs of impending failure, and I had no hot spare, then I'd
probably get a new drive and replace it ASAP just because of the
hassle involved when it does fail. Chances are that failure is going
to happen at an inconvenient time, whereas I could do the
replacement at a time convenient to me.
If, like OP, I had a hot spare in the array then really it is a
no-brainer to me: promote the hot spare then remove the suspect drive.
Since it's a spare there is no time where the array lacks
redundancy. If you wait for the drive to fail then there will be a
period of no redundancy while the spare is brought it.
This does depend on what kind of SMART failure it is though. Some of
them are a concern but do not imply total device failure in the near
future.
Cheers,
Andy
--
https://bitfolk.com/ -- No-nonsense VPS hosting
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.user
csiph-web