Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.os.linux.hardware > #2676

Re: A dying very old HDD confirmed?

From "Rod Speed" <rod.speed.aaa@gmail.com>
Newsgroups comp.os.linux.hardware, comp.sys.ibm.pc.hardware.storage, alt.comp.periphs.hdd
Subject Re: A dying very old HDD confirmed?
Date 2015-01-05 08:18 +1100
Message-ID <cgtp2qFiv9vU1@mid.individual.net> (permalink)
References <VLqdnVDGnrZrN8_OnZ2dnUVZ_sydnZ2d@earthlink.com> <B8SdnW5VSrXoJc_OnZ2dnUVZ_vednZ2d@earthlink.com> <brg69qF4vtjU1@mid.individual.net> <r82dnbnXB_n0MjXJnZ2dnUU7-T2dnZ2d@earthlink.com>

Cross-posted to 3 groups.

Show all headers | View raw



"Ant" <ant@zimage.comANT> wrote in message 
news:r82dnbnXB_n0MjXJnZ2dnUU7-T2dnZ2d@earthlink.com...
> On 4/19/2014 1:56 PM, Rod Speed wrote:
> ...
>>> From /var/log/syslog:
>>> ...
>>> Apr 19 09:45:01 MyLinuxBox /USR/SBIN/CRON[25663]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 09:55:01 MyLinuxBox /USR/SBIN/CRON[26267]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], 1
>>> Currently unreadable (pending) sectors
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], 1
>>> Offline uncorrectable sectors
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], SMART
>>> Prefailure Attribute: 1 Raw_Read_Error_Rate changed from 57 to 56
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], SMART
>>> Usage Attribute: 195 Hardware_ECC_Recovered changed from 57 to 56
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT],
>>> previous self-test completed with error (read test element)
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT],
>>> Self-Test Log error count increased from 0 to 1
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Sending warning via
>>> /usr/share/smartmontools/smartd-runner to root ...
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Warning via
>>> /usr/share/smartmontools/smartd-runner to root: successful
>>> Apr 19 10:05:01 MyLinuxBox /USR/SBIN/CRON[26877]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 10:15:01 MyLinuxBox /USR/SBIN/CRON[27449]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 10:17:01 MyLinuxBox /USR/SBIN/CRON[27560]: (root) CMD (   cd /
>>> && run-parts --report /etc/cron.hourly)
>>> Apr 19 10:25:01 MyLinuxBox /USR/SBIN/CRON[27986]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 10:25:40 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], 1
>>> Currently unreadable (pending) sectors
>>> Apr 19 10:25:40 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], 1
>>> Offline uncorrectable sectors
>>> Apr 19 10:25:40 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], SMART
>>> Usage Attribute: 194 Temperature_Celsius changed from 39 to 38
>>> Apr 19 10:35:01 MyLinuxBox /USR/SBIN/CRON[28518]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 10:45:01 MyLinuxBox /USR/SBIN/CRON[28950]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>>
>>> # smartctl -A /dev/sda
>>> smartctl 5.41 2011-06-09 r3365 [x86_64-linux-3.2.0-4-amd64] (local 
>>> build)
>>> Copyright (C) 2002-11 by Bruce Allen,
>>> http://smartmontools.sourceforge.net
>>>
>>> === START OF READ SMART DATA SECTION ===
>>> SMART Attributes Data Structure revision number: 10
>>> Vendor Specific SMART Attributes with Thresholds:
>>> ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE UPDATED
>>> WHEN_FAILED RAW_VALUE
>>>   1 Raw_Read_Error_Rate     0x000f   056   053   006    Pre-fail
>>>        -       208576601
>>>   3 Spin_Up_Time            0x0003   098   098   000    Pre-fail
>>>        -       0
>>>   4 Start_Stop_Count        0x0032   100   100   020    Old_age
>>> -       0
>>>   5 Reallocated_Sector_Ct   0x0033   100   100   036    Pre-fail
>>>        -       6
>>>   7 Seek_Error_Rate         0x000f   089   060   030    Pre-fail
>>>        -       928963346
>>>   9 Power_On_Hours          0x0032   019   019   000    Old_age
>>> -       71036
>>>  10 Spin_Retry_Count        0x0013   100   100   097    Pre-fail
>>>        -       0
>>>  12 Power_Cycle_Count       0x0032   100   100   020    Old_age
>>> -       362
>>> 194 Temperature_Celsius     0x0022   038   053   000    Old_age
>>> Always - 38
>>> 195 Hardware_ECC_Recovered  0x001a   056   053   000    Old_age
>>> Always - 208576601
>>> 197 Current_Pending_Sector  0x0012   100   100   000    Old_age
>>> Always - 1
>>> 198 Offline_Uncorrectable   0x0010   100   100   000    Old_age
>>> ne      -       1
>>> 199 UDMA_CRC_Error_Count    0x003e   200   200   000    Old_age
>>> Always - 0
>>> 200 Multi_Zone_Error_Rate   0x0000   100   253   000    Old_age
>>> ne      -       0
>>> 202 Data_Address_Mark_Errs  0x0032   100   253   000    Old_age
>>> Always - 0
>>>
>>> Those results look bad? :(
>>
>> Not too bad if they don't keep increasing.
>
> Hmm, I think they are now finally increasing a lot? From (py)dmesg's log:
> ...
> [2015-01-03 17:19:57] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0 
> action 0x0
> [2015-01-03 17:19:57] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:19:57] ata1.00: failed command: READ DMA
> [2015-01-03 17:19:57] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 
> 0 dma 4096 in
> [2015-01-03 17:19:57]          res 51/40:08:10:e0:3c/00:00:00:00:00/e0 
> Emask 0x9 (media error)
> [2015-01-03 17:19:57] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:19:57] ata1.00: error: { UNC }
> [2015-01-03 17:19:57] ata1.00: configured for UDMA/100
> [2015-01-03 17:19:57] ata1: EH complete
> [2015-01-03 17:20:01] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0 
> action 0x0
> [2015-01-03 17:20:01] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:20:01] ata1.00: failed command: READ DMA
> [2015-01-03 17:20:01] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 
> 0 dma 4096 in
> [2015-01-03 17:20:01]          res 51/40:08:10:e0:3c/00:00:00:00:00/e0 
> Emask 0x9 (media error)
> [2015-01-03 17:20:01] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:20:01] ata1.00: error: { UNC }
> [2015-01-03 17:20:01] ata1.00: configured for UDMA/100
> [2015-01-03 17:20:01] ata1: EH complete
> [2015-01-03 17:20:06] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0 
> action 0x0
> [2015-01-03 17:20:06] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:20:06] ata1.00: failed command: READ DMA
> [2015-01-03 17:20:06] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 
> 0 dma 4096 in
> [2015-01-03 17:20:06]          res 51/40:08:10:e0:3c/00:00:00:00:00/e0 
> Emask 0x9 (media error)
> [2015-01-03 17:20:06] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:20:06] ata1.00: error: { UNC }
> [2015-01-03 17:20:06] ata1.00: configured for UDMA/100
> [2015-01-03 17:20:06] ata1: EH complete
> [2015-01-03 17:20:11] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0 
> action 0x0
> [2015-01-03 17:20:11] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:20:11] ata1.00: failed command: READ DMA
> [2015-01-03 17:20:11] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 
> 0 dma 4096 in
> [2015-01-03 17:20:11]          res 51/40:08:10:e0:3c/00:00:00:00:00/e0 
> Emask 0x9 (media error)
> [2015-01-03 17:20:11] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:20:11] ata1.00: error: { UNC }
> [2015-01-03 17:20:11] ata1.00: configured for UDMA/100
> [2015-01-03 17:20:11] ata1: EH complete
> [2015-01-03 17:20:15] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0 
> action 0x0
> [2015-01-03 17:20:15] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:20:15] ata1.00: failed command: READ DMA
> [2015-01-03 17:20:15] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 
> 0 dma 4096 in
> [2015-01-03 17:20:15]          res 51/40:08:10:e0:3c/00:00:00:00:00/e0 
> Emask 0x9 (media error)
> [2015-01-03 17:20:15] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:20:15] ata1.00: error: { UNC }
> [2015-01-03 17:20:15] ata1.00: configured for UDMA/100
> [2015-01-03 17:20:15] ata1: EH complete
> [2015-01-03 17:20:20] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0 
> action 0x0
> [2015-01-03 17:20:20] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:20:20] ata1.00: failed command: READ DMA
> [2015-01-03 17:20:20] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 
> 0 dma 4096 in
> [2015-01-03 17:20:20]          res 51/40:08:10:e0:3c/00:00:00:00:00/e0 
> Emask 0x9 (media error)
> [2015-01-03 17:20:20] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:20:20] ata1.00: error: { UNC }
> [2015-01-03 17:20:20] ata1.00: configured for UDMA/100
> [2015-01-03 17:20:20] sd 0:0:0:0: [sda] Unhandled sense code
> [2015-01-03 17:20:20] sd 0:0:0:0: [sda]  Result: hostbyte=DID_OK 
> driverbyte=DRIVER_SENSE
> [2015-01-03 17:20:20] sd 0:0:0:0: [sda]  Sense Key : Medium Error 
> [current] [descriptor]
> [2015-01-03 17:20:20] Descriptor sense data with sense descriptors (in 
> hex):
> [2015-01-03 17:20:20]         72 03 11 04 00 00 00 0c 00 0a 80 00 00 00 00 
> 00
> [2015-01-03 17:20:20]         00 3c e0 10
> [2015-01-03 17:20:20] sd 0:0:0:0: [sda]  Add. Sense: Unrecovered read 
> error - auto reallocate failed
> [2015-01-03 17:20:20] sd 0:0:0:0: [sda] CDB: Read(10): 28 00 00 3c e0 10 
> 00 00 08 00
> [2015-01-03 17:20:20] end_request: I/O error, dev sda, sector 3989520
> [2015-01-03 17:20:20] EXT3-fs error (device sda5): ext3_get_inode_loc: 
> unable to read inode block - inode=212577, block=425986
> [2015-01-03 17:20:20] ata1: EH complete
> ...
>
> --
>
> From: root <root@MyBox>
> To: root@MyBox
> Subject: SMART error (ErrorCount) detected on host: MyBox
>
> This email was generated by the smartd daemon running on:
>
>    host name: MyBox
>   DNS domain: [Unknown]
>   NIS domain: (none)
>
> The following warning/error was logged by the smartd daemon:
>
> Device: /dev/sda [SAT], ATA error count increased from 0 to 6
>
> For details see host's SYSLOG.
>
> You can also use the smartctl utility for further investigation.
> Another email message will be sent in 24 hours if the problem persists.
>
> --
>
> From the /var/log/syslog:
> ...
> Jan  3 17:19:57 MyBox kernel: [292948.051212] ata1.00: exception Emask 0x0 
> SAct 0x0 SErr 0x0 action 0x0
> Jan  3 17:19:57 MyBox kernel: [292948.051217] ata1.00: BMDMA stat 0x25
> Jan  3 17:19:57 MyBox kernel: [292948.051220] ata1.00: failed command: 
> READ DMA
> Jan  3 17:19:57 MyBox kernel: [292948.051226] ata1.00: cmd 
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan  3 17:19:57 MyBox kernel: [292948.051228]          res 
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan  3 17:19:57 MyBox kernel: [292948.051231] ata1.00: status: { DRDY 
> ERR }
> Jan  3 17:19:57 MyBox kernel: [292948.051233] ata1.00: error: { UNC }
> Jan  3 17:19:57 MyBox kernel: [292948.136706] ata1.00: configured for 
> UDMA/100
> Jan  3 17:19:57 MyBox kernel: [292948.136715] ata1: EH complete
> Jan  3 17:20:02 MyBox kernel: [292952.671441] ata1.00: exception Emask 0x0 
> SAct 0x0 SErr 0x0 action 0x0
> Jan  3 17:20:02 MyBox kernel: [292952.671446] ata1.00: BMDMA stat 0x25
> Jan  3 17:20:02 MyBox kernel: [292952.671450] ata1.00: failed command: 
> READ DMA
> Jan  3 17:20:02 MyBox kernel: [292952.671456] ata1.00: cmd 
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan  3 17:20:02 MyBox kernel: [292952.671457]          res 
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan  3 17:20:02 MyBox kernel: [292952.671461] ata1.00: status: { DRDY 
> ERR }
> Jan  3 17:20:02 MyBox kernel: [292952.671463] ata1.00: error: { UNC }
> Jan  3 17:20:02 MyBox kernel: [292952.740715] ata1.00: configured for 
> UDMA/100
> Jan  3 17:20:02 MyBox kernel: [292952.740730] ata1: EH complete
> Jan  3 17:20:07 MyBox kernel: [292957.308333] ata1.00: exception Emask 0x0 
> SAct 0x0 SErr 0x0 action 0x0
> Jan  3 17:20:07 MyBox kernel: [292957.308338] ata1.00: BMDMA stat 0x25
> Jan  3 17:20:07 MyBox kernel: [292957.308341] ata1.00: failed command: 
> READ DMA
> Jan  3 17:20:07 MyBox kernel: [292957.308347] ata1.00: cmd 
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan  3 17:20:07 MyBox kernel: [292957.308349]          res 
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan  3 17:20:07 MyBox kernel: [292957.308352] ata1.00: status: { DRDY 
> ERR }
> Jan  3 17:20:07 MyBox kernel: [292957.308355] ata1.00: error: { UNC }
> Jan  3 17:20:07 MyBox kernel: [292957.420787] ata1.00: configured for 
> UDMA/100
> Jan  3 17:20:07 MyBox kernel: [292957.420801] ata1: EH complete
> Jan  3 17:20:11 MyBox kernel: [292962.001820] ata1.00: exception Emask 0x0 
> SAct 0x0 SErr 0x0 action 0x0
> Jan  3 17:20:11 MyBox kernel: [292962.001825] ata1.00: BMDMA stat 0x25
> Jan  3 17:20:11 MyBox kernel: [292962.001828] ata1.00: failed command: 
> READ DMA
> Jan  3 17:20:11 MyBox kernel: [292962.001835] ata1.00: cmd 
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan  3 17:20:11 MyBox kernel: [292962.001836]          res 
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan  3 17:20:11 MyBox kernel: [292962.001839] ata1.00: status: { DRDY 
> ERR }
> Jan  3 17:20:11 MyBox kernel: [292962.001842] ata1.00: error: { UNC }
> Jan  3 17:20:11 MyBox kernel: [292962.088713] ata1.00: configured for 
> UDMA/100
> Jan  3 17:20:11 MyBox kernel: [292962.088728] ata1: EH complete
> Jan  3 17:20:16 MyBox kernel: [292966.634963] ata1.00: exception Emask 0x0 
> SAct 0x0 SErr 0x0 action 0x0
> Jan  3 17:20:16 MyBox kernel: [292966.634968] ata1.00: BMDMA stat 0x25
> Jan  3 17:20:16 MyBox kernel: [292966.634971] ata1.00: failed command: 
> READ DMA
> Jan  3 17:20:16 MyBox kernel: [292966.634978] ata1.00: cmd 
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan  3 17:20:16 MyBox kernel: [292966.634979]          res 
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan  3 17:20:16 MyBox kernel: [292966.634982] ata1.00: status: { DRDY 
> ERR }
> Jan  3 17:20:16 MyBox kernel: [292966.634985] ata1.00: error: { UNC }
> Jan  3 17:20:16 MyBox kernel: [292966.720754] ata1.00: configured for 
> UDMA/100
> Jan  3 17:20:16 MyBox kernel: [292966.720769] ata1: EH complete
> Jan  3 17:20:21 MyBox kernel: [292971.391149] ata1.00: exception Emask 0x0 
> SAct 0x0 SErr 0x0 action 0x0
> Jan  3 17:20:21 MyBox kernel: [292971.391154] ata1.00: BMDMA stat 0x25
> Jan  3 17:20:21 MyBox kernel: [292971.391157] ata1.00: failed command: 
> READ DMA
> Jan  3 17:20:21 MyBox kernel: [292971.391164] ata1.00: cmd 
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan  3 17:20:21 MyBox kernel: [292971.391165]          res 
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan  3 17:20:21 MyBox kernel: [292971.391169] ata1.00: status: { DRDY 
> ERR }
> Jan  3 17:20:21 MyBox kernel: [292971.391171] ata1.00: error: { UNC }
> Jan  3 17:20:21 MyBox kernel: [292971.500707] ata1.00: configured for 
> UDMA/100
> Jan  3 17:20:21 MyBox kernel: [292971.500723] sd 0:0:0:0: [sda] Unhandled 
> sense code
> Jan  3 17:20:21 MyBox kernel: [292971.500726] sd 0:0:0:0: [sda]  Result: 
> hostbyte=DID_OK driverbyte=DRIVER_SENSE
> Jan  3 17:20:21 MyBox kernel: [292971.500729] sd 0:0:0:0: [sda]  Sense Key 
> : Medium Error [current] [descriptor]
> Jan  3 17:20:21 MyBox kernel: [292971.500734] Descriptor sense data with 
> sense descriptors (in hex):
> Jan  3 17:20:21 MyBox kernel: [292971.500736]         72 03 11 04 00 00 00 
> 0c 00 0a 80 00 00 00 00 00
> Jan  3 17:20:21 MyBox kernel: [292971.500745]         00 3c e0 10
> Jan  3 17:20:21 MyBox kernel: [292971.500749] sd 0:0:0:0: [sda]  Add. 
> Sense: Unrecovered read error - auto reallocate failed
> Jan  3 17:20:21 MyBox kernel: [292971.500754] sd 0:0:0:0: [sda] CDB: 
> Read(10): 28 00 00 3c e0 10 00 00 08 00
> Jan  3 17:20:21 MyBox kernel: [292971.500762] end_request: I/O error, dev 
> sda, sector 3989520
> Jan  3 17:20:21 MyBox kernel: [292971.500778] EXT3-fs error (device sda5): 
> ext3_get_inode_loc: unable to read inode block - inode=212577, 
> block=425986
> Jan  3 17:20:21 MyBox kernel: [292971.500782] ata1: EH complete
> Jan  3 17:25:01 MyBox /USR/SBIN/CRON[24006]: (root) CMD (command -v 
> debian-sa1 > /dev/null && debian-sa1 1 1)
> Jan  3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], 62448 
> Currently unreadable (pending) sectors (changed +55)
> Jan  3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], 62448 Offline 
> uncorrectable sectors (changed +55)
> Jan  3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART 
> Prefailure Attribute: 1 Raw_Read_Error_Rate changed from 74 to 71
> Jan  3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART Usage 
> Attribute: 194 Temperature_Celsius changed from 27 to 28
> Jan  3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART Usage 
> Attribute: 195 Hardware_ECC_Recovered changed from 74 to 71
> Jan  3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART Usage 
> Attribute: 202 Data_Address_Mark_Errs changed from 31 to 20
> Jan  3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], ATA error 
> count increased from 0 to 6
> Jan  3 17:27:39 MyBox smartd[3440]: Sending warning via 
> /usr/share/smartmontools/smartd-runner to root ...
> Jan  3 17:27:39 MyBox smartd[3440]: Warning via 
> /usr/share/smartmontools/smartd-runner to root: successful
> Jan  3 17:27:39 MyBox smartd[3440]: Device: /dev/sdb [SAT], SMART Usage 
> Attribute: 194 Temperature_Celsius changed from 25 to 26
> Jan  3 17:27:54 MyBox kernel: [293424.703922] device eth0 entered 
> promiscuous mode
> ...
>
> I was using it with my very old Windows XP Pro SP3 VM on it. It was fine 
> to me. No errors and slow downs. So, I ran smartctl's short and long tests 
> and they both failed quickly with 10%. I even got an e-mail:
>
> Date: Sat, 03 Jan 2015 18:57:42 -0800
> From: root <root@MyBox>
> To: root@MyBox
> Subject: SMART error (SelfTest) detected on host: MyBox
>
> This email was generated by the smartd daemon running on:
>
>    host name: MyBox
>   DNS domain: [Unknown]
>   NIS domain: (none)
>
> The following warning/error was logged by the smartd daemon:
>
> Device: /dev/sda [SAT], Self-Test Log error count increased from 3 to 5
>
> For details see host's SYSLOG.
>
> You can also use the smartctl utility for further investigation.
> The original email about this issue was sent at Sat Apr 19 09:55:41 2014 
> PDT
> Another email message will be sent in 24 hours if the problem persists.
>
> /var/log/syslog:
> ...
> Jan  3 18:57:42 MyBox smartd[3440]: Device: /dev/sda [SAT], 62485 Offline 
> uncorrectable sectors (changed +24)
> Jan  3 18:57:42 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART 
> Prefailure Attribute: 1 Raw_Read_Error_Rate changed from 66 to 63
> Jan  3 18:57:42 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART Usage 
> Attribute: 195 Hardware_ECC_Recovered changed from 66 to 63
> Jan  3 18:57:42 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART Usage 
> Attribute: 202 Data_Address_Mark_Errs changed from 0 to 249
> Jan  3 18:57:42 MyBox smartd[3440]: Device: /dev/sda [SAT], Self-Test Log 
> error count increased from 3 to 5
> Jan  3 18:57:42 MyBox smartd[3440]: Sending warning via 
> /usr/share/smartmontools/smartd-runner to root ...
> Jan  3 18:57:42 MyBox smartd[3440]: Warning via 
> /usr/share/smartmontools/smartd-runner to root: successful
> ...
>
> # smartctl -a /dev/sda
> smartctl 5.41 2011-06-09 r3365 [x86_64-linux-3.2.0-4-amd64] (local build)
> Copyright (C) 2002-11 by Bruce Allen, http://smartmontools.sourceforge.net
>
> === START OF INFORMATION SECTION ===
> Model Family:     Seagate Barracuda 7200.7 and 7200.7 Plus
> Device Model:     ST380011A
> Serial Number:    4JV5P7LN
> Firmware Version: 8.01
> User Capacity:    80,026,361,856 bytes [80.0 GB]
> Sector Size:      512 bytes logical/physical
> Device is:        In smartctl database [for details use: -P show]
> ATA Version is:   6
> ATA Standard is:  ATA/ATAPI-6 T13 1410D revision 2
> Local Time is:    Sat Jan  3 19:02:26 2015 PST
> SMART support is: Available - device has SMART capability.
> SMART support is: Enabled
>
> === START OF READ SMART DATA SECTION ===
> SMART overall-health self-assessment test result: PASSED
>
> General SMART Values:
> Offline data collection status:  (0x82) Offline data collection activity
>                                         was completed without error.
>                                         Auto Offline Data Collection: 
> Enabled.
> Self-test execution status:      ( 121) The previous self-test completed 
> having
>                                         the read element of the test 
> failed.
> Total time to complete Offline
> data collection:                (  430) seconds.
> Offline data collection
> capabilities:                    (0x5b) SMART execute Offline immediate.
>                                         Auto Offline data collection 
> on/off support.
>                                         Suspend Offline collection upon 
> new
>                                         command.
>                                         Offline surface scan supported.
>                                         Self-test supported.
>                                         No Conveyance Self-test supported.
>                                         Selective Self-test supported.
> SMART capabilities:            (0x0003) Saves SMART data before entering
>                                         power-saving mode.
>                                         Supports SMART auto save timer.
> Error logging capability:        (0x01) Error logging supported.
>                                         General Purpose Logging supported.
> Short self-test routine
> recommended polling time:        (   1) minutes.
> Extended self-test routine
> recommended polling time:        (  58) minutes.
>
> SMART Attributes Data Structure revision number: 10
> Vendor Specific SMART Attributes with Thresholds:
> ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE UPDATED 
> WHEN_FAILED RAW_VALUE
>   1 Raw_Read_Error_Rate     0x000f   063   046   006    Pre-fail 
>        -       165638319
>   3 Spin_Up_Time            0x0003   098   098   000    Pre-fail 
>        -       0
>   4 Start_Stop_Count        0x0032   100   100   020    Old_age 
>        -       0
>   5 Reallocated_Sector_Ct   0x0033   053   053   036    Pre-fail 
>        -       1884

That's very bad.

>   7 Seek_Error_Rate         0x000f   086   060   030    Pre-fail 
>        -       426361585
>   9 Power_On_Hours          0x0032   012   012   000    Old_age 
>        -       77139
>  10 Spin_Retry_Count        0x0013   100   100   097    Pre-fail 
>        -       0
>  12 Power_Cycle_Count       0x0032   100   100   020    Old_age 
>        -       370
> 194 Temperature_Celsius     0x0022   029   053   000    Old_age   Always - 
> 29
> 195 Hardware_ECC_Recovered  0x001a   063   046   000    Old_age   Always - 
> 165638319
> 197 Current_Pending_Sector  0x0012   001   001   000    Old_age   Always - 
> 62493

That's very bad.

> 198 Offline_Uncorrectable   0x0010   001   001   000    Old_age 
> ne      -       62493

Ditto.

> 199 UDMA_CRC_Error_Count    0x003e   200   200   000    Old_age   Always - 
> 0
> 200 Multi_Zone_Error_Rate   0x0000   100   253   000    Old_age 
> ne      -       0
> 202 Data_Address_Mark_Errs  0x0032   249   146   000    Old_age   Always - 
> 107
>
> SMART Error Log Version: 1
> ATA Error Count: 6 (device log contains only the most recent five errors)
>         CR = Command Register [HEX]
>         FR = Features Register [HEX]
>         SC = Sector Count Register [HEX]
>         SN = Sector Number Register [HEX]
>         CL = Cylinder Low Register [HEX]
>         CH = Cylinder High Register [HEX]
>         DH = Device/Head Register [HEX]
>         DC = Device Command Register [HEX]
>         ER = Error register [HEX]
>         ST = Status register [HEX]
> Powered_Up_Time is measured from power on, and printed as
> DDd+hh:mm:SS.sss where DD=days, hh=hours, mm=minutes,
> SS=sec, and sss=millisec. It "wraps" after 49.710 days.
>
> Error 6 occurred at disk power-on lifetime: 11601 hours (483 days + 9 
> hours)
>   When the command that caused the error occurred, the device was active 
> or idle.
>
>   After command completion occurred, registers were:
>   ER ST SC SN CL CH DH
>   -- -- -- -- -- -- --
>   40 51 08 10 e0 3c e0  Error: UNC 8 sectors at LBA = 0x003ce010 = 3989520
>
>   Commands leading to the command that caused the error were:
>   CR FR SC SN CL CH DH DC   Powered_Up_Time  Command/Feature_Name
>   -- -- -- -- -- -- -- --  ----------------  --------------------
>   c8 00 08 10 e0 3c e0 00      11:28:27.346  READ DMA
>   27 00 00 00 00 00 e0 00      11:28:27.334  READ NATIVE MAX ADDRESS EXT
>   ec 00 00 00 00 00 a0 02      11:28:36.815  IDENTIFY DEVICE
>   ef 03 45 00 00 00 a0 02      11:28:36.814  SET FEATURES [Set transfer 
> mode]
>   27 00 00 00 00 00 e0 00      11:28:36.807  READ NATIVE MAX ADDRESS EXT
>
> Error 5 occurred at disk power-on lifetime: 11601 hours (483 days + 9 
> hours)
>   When the command that caused the error occurred, the device was active 
> or idle.
>
>   After command completion occurred, registers were:
>   ER ST SC SN CL CH DH
>   -- -- -- -- -- -- --
>   40 51 08 10 e0 3c e0  Error: UNC 8 sectors at LBA = 0x003ce010 = 3989520
>
>   Commands leading to the command that caused the error were:
>   CR FR SC SN CL CH DH DC   Powered_Up_Time  Command/Feature_Name
>   -- -- -- -- -- -- -- --  ----------------  --------------------
>   c8 00 08 10 e0 3c e0 00      11:28:27.346  READ DMA
>   27 00 00 00 00 00 e0 00      11:28:27.334  READ NATIVE MAX ADDRESS EXT
>   ec 00 00 00 00 00 a0 02      11:28:27.241  IDENTIFY DEVICE
>   ef 03 45 00 00 00 a0 02      11:28:27.240  SET FEATURES [Set transfer 
> mode]
>   27 00 00 00 00 00 e0 00      11:28:27.232  READ NATIVE MAX ADDRESS EXT
>
> Error 4 occurred at disk power-on lifetime: 11601 hours (483 days + 9 
> hours)
>   When the command that caused the error occurred, the device was active 
> or idle.
>
>   After command completion occurred, registers were:
>   ER ST SC SN CL CH DH
>   -- -- -- -- -- -- --
>   40 51 08 10 e0 3c e0  Error: UNC 8 sectors at LBA = 0x003ce010 = 3989520
>
>   Commands leading to the command that caused the error were:
>   CR FR SC SN CL CH DH DC   Powered_Up_Time  Command/Feature_Name
>   -- -- -- -- -- -- -- --  ----------------  --------------------
>   c8 00 08 10 e0 3c e0 00      11:28:27.346  READ DMA
>   27 00 00 00 00 00 e0 00      11:28:27.334  READ NATIVE MAX ADDRESS EXT
>   ec 00 00 00 00 00 a0 02      11:28:27.241  IDENTIFY DEVICE
>   ef 03 45 00 00 00 a0 02      11:28:27.240  SET FEATURES [Set transfer 
> mode]
>   27 00 00 00 00 00 e0 00      11:28:27.232  READ NATIVE MAX ADDRESS EXT
>
> Error 3 occurred at disk power-on lifetime: 11601 hours (483 days + 9 
> hours)
>   When the command that caused the error occurred, the device was active 
> or idle.
>
>   After command completion occurred, registers were:
>   ER ST SC SN CL CH DH
>   -- -- -- -- -- -- --
>   40 51 08 10 e0 3c e0  Error: UNC 8 sectors at LBA = 0x003ce010 = 3989520
>
>   Commands leading to the command that caused the error were:
>   CR FR SC SN CL CH DH DC   Powered_Up_Time  Command/Feature_Name
>   -- -- -- -- -- -- -- --  ----------------  --------------------
>   c8 00 08 10 e0 3c e0 00      11:28:17.813  READ DMA
>   27 00 00 00 00 00 e0 00      11:28:17.808  READ NATIVE MAX ADDRESS EXT
>   ec 00 00 00 00 00 a0 02      11:28:13.111  IDENTIFY DEVICE
>   ef 03 45 00 00 00 a0 02      11:28:13.107  SET FEATURES [Set transfer 
> mode]
>   27 00 00 00 00 00 e0 00      11:28:13.092  READ NATIVE MAX ADDRESS EXT
>
> Error 2 occurred at disk power-on lifetime: 11601 hours (483 days + 9 
> hours)
>   When the command that caused the error occurred, the device was active 
> or idle.
>
>   After command completion occurred, registers were:
>   ER ST SC SN CL CH DH
>   -- -- -- -- -- -- --
>   40 51 08 10 e0 3c e0  Error: UNC 8 sectors at LBA = 0x003ce010 = 3989520
>
>   Commands leading to the command that caused the error were:
>   CR FR SC SN CL CH DH DC   Powered_Up_Time  Command/Feature_Name
>   -- -- -- -- -- -- -- --  ----------------  --------------------
>   c8 00 08 10 e0 3c e0 00      11:28:17.813  READ DMA
>   27 00 00 00 00 00 e0 00      11:28:17.808  READ NATIVE MAX ADDRESS EXT
>   ec 00 00 00 00 00 a0 02      11:28:13.111  IDENTIFY DEVICE
>   ef 03 45 00 00 00 a0 02      11:28:13.107  SET FEATURES [Set transfer 
> mode]
>   27 00 00 00 00 00 e0 00      11:28:13.092  READ NATIVE MAX ADDRESS EXT
>
> SMART Self-test log structure revision number 1
> Num  Test_Description    Status                  Remaining LifeTime(hours) 
> LBA_of_first_error
> # 1  Extended offline    Completed: read failure       90%     11603 
> 3989520
> # 2  Short offline       Completed: read failure       90%     11602 
> 3989520
> # 3  Extended offline    Completed: read failure       90%     10547 
> 616217
> # 4  Extended offline    Completed: read failure       90%      9989 
> 684994
> # 5  Extended offline    Completed: read failure       90%      9785 
> 3783997
> # 6  Short offline       Completed without error       00%      9785 -
> # 7  Extended offline    Completed without error       00%      5853 -
> # 8  Extended offline    Completed without error       00%      5815 -
> # 9  Extended offline    Completed without error       00%      5583 -
> #10  Short offline       Completed without error       00%      5583 -
> #11  Extended offline    Completed: read failure       60%      5502 
> 67691664
> #12  Extended offline    Completed: read failure       60%      5501 
> 67691664
> #13  Short offline       Completed without error       00%      5501 -
> #14  Short offline       Completed without error       00%      5501 -
> #15  Extended offline    Completed: read failure       60%      5501 
> 67691664
> #16  Short offline       Completed without error       00%      5501 -
> #17  Extended offline    Completed: read failure       60%      5499 
> 67691664
> #18  Extended offline    Completed without error       00%      3165 -
> #19  Extended offline    Completed without error       00%     48483 -
> #20  Extended offline    Completed without error       00%     43207 -
> #21  Extended offline    Interrupted (host reset)      90%     43201 -
> 4 of 9 failed self-tests are outdated by newer successful extended offline 
> self-test # 7
>
> SMART Selective self-test log data structure revision number 1
>  SPAN  MIN_LBA  MAX_LBA  CURRENT_TEST_STATUS
>     1        0        0  Not_testing
>     2        0        0  Not_testing
>     3        0        0  Not_testing
>     4        0        0  Not_testing
>     5        0        0  Not_testing
> Selective self-test flags (0x0):
>   After scanning selected spans, do NOT read-scan remainder of disk.
> If Selective self-test is pending on power-up, resume after 0 minute 
> delay.
>
>
> Is it finally dying to you now? :P

Yep, she's dead, Jim. You into necrophilia ? 

Back to comp.os.linux.hardware | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-19 09:57 -0700
  Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-19 10:55 -0700
    Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-19 11:56 -0700
    Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-20 06:56 +1000
      Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-19 17:34 -0700
        Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-20 13:02 +1000
          Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-19 23:16 -0700
            Re: A dying very old HDD confirmed? Robert Nichols <SEE_SIGNATURE@localhost.localdomain.invalid> - 2014-04-20 07:19 -0500
              Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-20 09:54 -0700
              Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-20 09:57 -0700
                Re: A dying very old HDD confirmed? Robert Nichols <SEE_SIGNATURE@localhost.localdomain.invalid> - 2014-04-21 07:56 -0500
                Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-21 08:08 -0700
                Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-22 01:31 +0200
                Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-21 19:02 -0700
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 05:24 +1000
                Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-04-22 19:03 -0500
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 14:42 +1000
                Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-22 23:27 -0700
                Re: A dying very old HDD confirmed? Robert Nichols <SEE_SIGNATURE@localhost.localdomain.invalid> - 2014-04-23 09:03 -0500
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-24 06:32 +1000
                Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-24 00:26 +0200
                Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-04-23 18:47 -0500
                Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-10-19 21:05 -0500
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-10-21 05:13 +1100
                Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-10-20 16:04 -0500
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-10-22 07:01 +1100
                Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-10-21 15:56 -0500
                Re: A dying very old HDD confirmed? Arno <me@privacy.net> - 2014-10-22 22:43 +0000
                Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-10-24 08:18 -0700
                Re: A dying very old HDD confirmed? Arno <me@privacy.net> - 2014-10-24 15:39 +0000
                Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-10-24 20:23 -0500
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-24 11:20 +1000
                Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-24 09:47 +0200
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-24 06:35 +1000
                Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-24 00:23 +0200
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-24 11:17 +1000
                Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-24 09:55 +0200
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-25 05:17 +1000
                Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-27 19:12 +0200
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-28 19:59 +1000
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 05:22 +1000
          Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-20 11:08 +0200
            Re: A dying very old HDD confirmed? Richard Kettlewell <rjk@greenend.org.uk> - 2014-04-20 10:13 +0100
              Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-20 12:28 +0200
              Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-21 05:52 +1000
            Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-21 05:50 +1000
              Re: A dying very old HDD confirmed? cjt <cheljuba@prodigy.net> - 2014-04-20 19:35 -0500
          Re: A dying very old HDD confirmed? Måns Rullgård <mans@mansr.com> - 2014-04-22 12:59 +0100
            Re: A dying very old HDD confirmed? Vic RR Garcia <VicGar007@at-gmail.dot.com> - 2014-04-22 13:23 -0400
              Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 05:30 +1000
                Re: A dying very old HDD confirmed? Måns Rullgård <mans@mansr.com> - 2014-04-22 20:57 +0100
                Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 14:37 +1000
            Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 05:28 +1000
      Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2015-01-03 19:11 -0800
        Re: A dying very old HDD confirmed? Robert Nichols <SEE_SIGNATURE@localhost.localdomain.invalid> - 2015-01-04 12:11 -0600
          Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2015-01-04 12:38 -0800
            Re: A dying very old HDD confirmed? mjb@signal11.invalid (Mike) - 2015-01-04 23:02 +0000
              Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2015-01-04 15:21 -0800
        Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2015-01-05 08:18 +1100
          Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2015-01-04 15:21 -0800
      Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2015-01-04 06:44 -0800
  Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-20 06:55 +1000

csiph-web