Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #205276 > unrolled thread

OT: Current_Pending_Sector on /dev/sd?

Started bybasti <mailinglist@unix-solution.de>
First post2019-02-13 14:30 +0100
Last post2019-02-14 14:50 +0100
Articles 10 — 6 participants

Back to article view | Back to linux.debian.user


Contents

  OT: Current_Pending_Sector on /dev/sd? basti <mailinglist@unix-solution.de> - 2019-02-13 14:30 +0100
    Re: OT: Current_Pending_Sector on /dev/sd? Dan Ritter <dsr@randomstring.org> - 2019-02-13 15:10 +0100
      Re: OT: Current_Pending_Sector on /dev/sd? Pascal Hambourg <pascal@plouf.fr.eu.org> - 2019-02-13 23:00 +0100
    Re: OT: Current_Pending_Sector on /dev/sd? "Alexander V. Makartsev" <avbetev@gmail.com> - 2019-02-13 15:10 +0100
      Re: OT: Current_Pending_Sector on /dev/sd? Dan Ritter <dsr@randomstring.org> - 2019-02-13 15:30 +0100
      Re: OT: Current_Pending_Sector on /dev/sd? Pascal Hambourg <pascal@plouf.fr.eu.org> - 2019-02-13 23:00 +0100
    Re: OT: Current_Pending_Sector on /dev/sd? Claudio Kuenzler <ck@claudiokuenzler.com> - 2019-02-13 20:20 +0100
      Re: OT: Current_Pending_Sector on /dev/sd? basti <mailinglist@unix-solution.de> - 2019-02-14 12:30 +0100
        Re: OT: Current_Pending_Sector on /dev/sd? Claudio Kuenzler <ck@claudiokuenzler.com> - 2019-02-14 13:00 +0100
          Re: OT: Current_Pending_Sector on /dev/sd? Greg Wooledge <wooledg@eeg.ccf.org> - 2019-02-14 14:50 +0100

#205276 — OT: Current_Pending_Sector on /dev/sd?

Frombasti <mailinglist@unix-solution.de>
Date2019-02-13 14:30 +0100
SubjectOT: Current_Pending_Sector on /dev/sd?
Message-ID<xr2LU-1wL-17@gated-at.bofh.it>
hello,
I have a raid6 with 4 disks. 2 of them show Current_Pending_Sector 1.
The disks has warranty till Apr. 2019 so I decide to replace them.

After I change the disk and install it on an other computer to overwrite
with zero it the Current_Pending_Sector is gone.

What should I do? Whats our experience?

Best Regards,

[toc] | [next] | [standalone]


#205285

FromDan Ritter <dsr@randomstring.org>
Date2019-02-13 15:10 +0100
Message-ID<xr3oB-1Zl-3@gated-at.bofh.it>
In reply to#205276
basti wrote: 
> hello,
> I have a raid6 with 4 disks. 2 of them show Current_Pending_Sector 1.
> The disks has warranty till Apr. 2019 so I decide to replace them.
> 
> After I change the disk and install it on an other computer to overwrite
> with zero it the Current_Pending_Sector is gone.
> 
> What should I do? Whats our experience?

Sometimes disk sectors go bad. When the disk detects that, it
reads whatever it can and writes it somewhere else.

You can expect the number to be reset to zero either on repair
or after a power-cycle.

A small number in CPS is fine. If it starts going up, or the
reallocated sector count starts increasing, you might have an 
impending disaster.

-dsr-

[toc] | [prev] | [next] | [standalone]


#205321

FromPascal Hambourg <pascal@plouf.fr.eu.org>
Date2019-02-13 23:00 +0100
Message-ID<xraJs-6xx-7@gated-at.bofh.it>
In reply to#205285
Le 13/02/2019 à 14:59, Dan Ritter a écrit :
> basti wrote:
>> hello,
>> I have a raid6 with 4 disks. 2 of them show Current_Pending_Sector 1.
(...)
> A small number in CPS is fine.

No it's not fine. It means that the the host requested to read 
unreadable data, so useful data has been lost.

"Offline uncorrectable" also means that some data are unreadable, but 
the host did not request to read them yet so no useful data may have 
been lost.

> If it starts going up, or the
> reallocated sector count starts increasing

Reallocated sectors are fine. It means that data could be moved to spare 
sectors. No data have been lost.

[toc] | [prev] | [next] | [standalone]


#205287

From"Alexander V. Makartsev" <avbetev@gmail.com>
Date2019-02-13 15:10 +0100
Message-ID<xr3oB-1Zl-5@gated-at.bofh.it>
In reply to#205276

[Multipart message — attachments visible in raw view] — view raw

On 13.02.2019 18:22, basti wrote:
> hello,
> I have a raid6 with 4 disks. 2 of them show Current_Pending_Sector 1.
> The disks has warranty till Apr. 2019 so I decide to replace them.
>
> After I change the disk and install it on an other computer to overwrite
> with zero it the Current_Pending_Sector is gone.
>
> What should I do? Whats our experience?
>
> Best Regards,
>
IMO, it's ok for disk devices to develop unreadable blocks (bad blocks),
as long as they don't progress.
Internal firmware of the disk takes care of bad blocks, marks them
internally and reallocates them. (makes sure that next read\write
request to that bad block will be redirected to a safe block)
When internal list of bad blocks and reallocations will be full,
firmware will mark HDD as failed in SMART.

As capacity of modern HDDs gets bigger and surface density of blocks
increases, so does margin for error of faulty blocks. It is rare for a
large disk to not have them, despite the fact that every disk was
scanned and faulty regions remapped beforehand during manufacturing
process at the factory.
HDDs that developed bad blocks should be monitored for progression and
replaced if bad blocks began to appear frequently. You can get one bad
block in 5 years or 30 in a few days and there is no guarantee that
brand new drive won't have bad blocks.

Most hardware RAID controllers perform automatic full surface scans to
ensure data consistency (also could be called "patrol scans") and mark
HDDs as "Expected to Fail Soon" to warn user.
You have to backup your data regularly and in case of large RAID5/6
level arrays you should have at least one Hot Spare drive available in
your enclosure at all times.
Precautions are almost the same for software RAID, but more difficult to
setup.

-- 
With kindest regards, Alexander.

⢀⣴⠾⠻⢶⣦⠀ 
⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org
⠈⠳⣄⠀⠀⠀⠀ 

[toc] | [prev] | [next] | [standalone]


#205293

FromDan Ritter <dsr@randomstring.org>
Date2019-02-13 15:30 +0100
Message-ID<xr3HY-26b-3@gated-at.bofh.it>
In reply to#205287
Alexander V. Makartsev wrote: 
> On 13.02.2019 18:22, basti wrote:
> 
> Most hardware RAID controllers perform automatic full surface scans to
> ensure data consistency (also could be called "patrol scans") and mark
> HDDs as "Expected to Fail Soon" to warn user.
> You have to backup your data regularly and in case of large RAID5/6
> level arrays you should have at least one Hot Spare drive available in
> your enclosure at all times.
> Precautions are almost the same for software RAID, but more difficult to
> setup.

Not really more difficult: Debian's default mdadm configuration
includes a weekly run of checkarray.

-dsr-

[toc] | [prev] | [next] | [standalone]


#205322

FromPascal Hambourg <pascal@plouf.fr.eu.org>
Date2019-02-13 23:00 +0100
Message-ID<xraJs-6xx-15@gated-at.bofh.it>
In reply to#205287
Le 13/02/2019 à 15:01, Alexander V. Makartsev a écrit :
>>
> IMO, it's ok for disk devices to develop unreadable blocks (bad blocks),

No it's not ok. Unreadable blocks means lost data.

> Internal firmware of the disk takes care of bad blocks, marks them
> internally and reallocates them. (makes sure that next read\write
> request to that bad block will be redirected to a safe block)

The next *successful* read/write.

At least this is how things should work. But in my experience, bad 
blocks are not reallocated so easily.

[toc] | [prev] | [next] | [standalone]


#205308

FromClaudio Kuenzler <ck@claudiokuenzler.com>
Date2019-02-13 20:20 +0100
Message-ID<xr8eB-5ai-11@gated-at.bofh.it>
In reply to#205276

[Multipart message — attachments visible in raw view] — view raw

On Wed, Feb 13, 2019 at 2:22 PM basti <mailinglist@unix-solution.de> wrote:

> hello,
> I have a raid6 with 4 disks. 2 of them show Current_Pending_Sector 1.
>

Hi Basti

are you using mdadm for the raid-6 or a hardware raid controller?


> The disks has warranty till Apr. 2019 so I decide to replace them.
>

If there's only 1 current pending sector it could be difficult to get a
full replacement. HDD's have spare sectors which are used in such events
and are (*should*) be capable to handle a few defect sectors. It doesn't
mean (yet) that the drive is defect.


>
> After I change the disk and install it on an other computer to overwrite
> with zero it the Current_Pending_Sector is gone.
>

Yes, I've seen this too a couple of months ago on a remote NAS server. I
probably had the same reaction as you: I couldn't believe it. Especially as
the Current_Pending_Sector went to 0 and Reallocated_Sectors and
Offline_Uncorrectable staid at the same value as before, too.


>
> What should I do? Whats our experience?
>

Continuously monitor your drive's SMART values and (if possible) store the
results in a database (RRD, Timeseries DB, you name it) to create graphs
from the values. You can use the check_smart.pl monitoring plugin as an
examle. This will show you if the number of defect sectors increase or if
they stay steady. If the bad sectors increase, it's just a matter of time
until the drive physically fails. You can see an example of such a graph
(rrd in this case) with increasing bad sectors over 5 weeks here:
https://www.claudiokuenzler.com/blog/469/multiple-several-ways-monitor-physical-hard-drive-disk

As helpful as SMART is, never rely 100% on it, as drives may also fail
without any bad values in SMART.

[toc] | [prev] | [next] | [standalone]


#205331

Frombasti <mailinglist@unix-solution.de>
Date2019-02-14 12:30 +0100
Message-ID<xrnnj-60w-3@gated-at.bofh.it>
In reply to#205308
Hello,
I use mdadm for raid and try your nagios check like:

./check_smart.pl -g /dev/sd[a-z] -i ata
OK: [/dev/sda] - Device is clean|

Is it ok that is only return one drive?

After resync the entire raid because of HDD change the CPS on the 2'nd
hard drive is also gone.

Reallocated_Sector_Ct and Offline_Uncorrectable is also 0 on all drives.

On 13.02.19 20:10, Claudio Kuenzler wrote:
> 
> On Wed, Feb 13, 2019 at 2:22 PM basti <mailinglist@unix-solution.de
> <mailto:mailinglist@unix-solution.de>> wrote:
> 
>     hello,
>     I have a raid6 with 4 disks. 2 of them show Current_Pending_Sector 1.
> 
> 

[toc] | [prev] | [next] | [standalone]


#205332

FromClaudio Kuenzler <ck@claudiokuenzler.com>
Date2019-02-14 13:00 +0100
Message-ID<xrnQm-6ag-3@gated-at.bofh.it>
In reply to#205331

[Multipart message — attachments visible in raw view] — view raw

> ./check_smart.pl -g /dev/sd[a-z] -i ata
> OK: [/dev/sda] - Device is clean|
>

> Is it ok that is only return one drive?
>

you have to use double-quotes because it's a regular expression within the
perl plugin:

./check_smart -g "/dev/sd[a-z]" -i ata
OK: [/dev/sda] - Device is clean --- [/dev/sdb] - Device is clean ---
[/dev/sdc] - Device is clean --- [/dev/sdd] - Device is clean|

But this is the variant for "lazy admins". ;-)
I suggest you use it on a single drive to obtain more information
(performance data).
Example:

./check_smart -d "/dev/sda" -i ata
OK: no SMART errors detected. |Raw_Read_Error_Rate=0 Spin_Up_Time=4000
Start_Stop_Count=30 Reallocated_Sector_Ct=0 Seek_Error_Rate=0
Power_On_Hours=18183 Spin_Retry_Count=0 Calibration_Retry_Count=0
Power_Cycle_Count=30 Power-Off_Retract_Count=20 Load_Cycle_Count=249397
Temperature_Celsius=32 Reallocated_Event_Count=0 Current_Pending_Sector=0
Offline_Uncorrectable=0 UDMA_CRC_Error_Count=0 Multi_Zone_Error_Rate=0

You can now parse these values and save it wherever you need to create a
history of the drive's SMART values.

[toc] | [prev] | [next] | [standalone]


#205334

FromGreg Wooledge <wooledg@eeg.ccf.org>
Date2019-02-14 14:50 +0100
Message-ID<xrpyN-7jz-9@gated-at.bofh.it>
In reply to#205332
> > ./check_smart.pl -g /dev/sd[a-z] -i ata
> > OK: [/dev/sda] - Device is clean|

> you have to use double-quotes because it's a regular expression within the
> perl plugin:
> 
> ./check_smart -g "/dev/sd[a-z]" -i ata

At the shell level, it's a glob which happens to match one or more
device nodes.  So you're quoting the glob to prevent the shell from
performing filename expansion.

Without the quotes, the glob is expanded, and you end up running a
command like:

./check_smart.pl -g /dev/sda /dev/sdb /dev/sdc -i ata

which doesn't do what you expected.

With the quotes, the glob is passed verbatim to the perl script, and
the perl script does whatever it has been programmed to do with it.

There's precedent for this behavior in a few other unix commands, like
find and tar, both of which can take (quoted) globs as arguments
and perform their own matching and/or expansion against them.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web