Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.os.linux.hardware > #2676
| From | "Rod Speed" <rod.speed.aaa@gmail.com> |
|---|---|
| Newsgroups | comp.os.linux.hardware, comp.sys.ibm.pc.hardware.storage, alt.comp.periphs.hdd |
| Subject | Re: A dying very old HDD confirmed? |
| Date | 2015-01-05 08:18 +1100 |
| Message-ID | <cgtp2qFiv9vU1@mid.individual.net> (permalink) |
| References | <VLqdnVDGnrZrN8_OnZ2dnUVZ_sydnZ2d@earthlink.com> <B8SdnW5VSrXoJc_OnZ2dnUVZ_vednZ2d@earthlink.com> <brg69qF4vtjU1@mid.individual.net> <r82dnbnXB_n0MjXJnZ2dnUU7-T2dnZ2d@earthlink.com> |
Cross-posted to 3 groups.
"Ant" <ant@zimage.comANT> wrote in message
news:r82dnbnXB_n0MjXJnZ2dnUU7-T2dnZ2d@earthlink.com...
> On 4/19/2014 1:56 PM, Rod Speed wrote:
> ...
>>> From /var/log/syslog:
>>> ...
>>> Apr 19 09:45:01 MyLinuxBox /USR/SBIN/CRON[25663]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 09:55:01 MyLinuxBox /USR/SBIN/CRON[26267]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], 1
>>> Currently unreadable (pending) sectors
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], 1
>>> Offline uncorrectable sectors
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], SMART
>>> Prefailure Attribute: 1 Raw_Read_Error_Rate changed from 57 to 56
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], SMART
>>> Usage Attribute: 195 Hardware_ECC_Recovered changed from 57 to 56
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT],
>>> previous self-test completed with error (read test element)
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT],
>>> Self-Test Log error count increased from 0 to 1
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Sending warning via
>>> /usr/share/smartmontools/smartd-runner to root ...
>>> Apr 19 09:55:41 MyLinuxBox smartd[3594]: Warning via
>>> /usr/share/smartmontools/smartd-runner to root: successful
>>> Apr 19 10:05:01 MyLinuxBox /USR/SBIN/CRON[26877]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 10:15:01 MyLinuxBox /USR/SBIN/CRON[27449]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 10:17:01 MyLinuxBox /USR/SBIN/CRON[27560]: (root) CMD ( cd /
>>> && run-parts --report /etc/cron.hourly)
>>> Apr 19 10:25:01 MyLinuxBox /USR/SBIN/CRON[27986]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 10:25:40 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], 1
>>> Currently unreadable (pending) sectors
>>> Apr 19 10:25:40 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], 1
>>> Offline uncorrectable sectors
>>> Apr 19 10:25:40 MyLinuxBox smartd[3594]: Device: /dev/sda [SAT], SMART
>>> Usage Attribute: 194 Temperature_Celsius changed from 39 to 38
>>> Apr 19 10:35:01 MyLinuxBox /USR/SBIN/CRON[28518]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>> Apr 19 10:45:01 MyLinuxBox /USR/SBIN/CRON[28950]: (root) CMD (command
>>> -v debian-sa1 > /dev/null && debian-sa1 1 1)
>>>
>>> # smartctl -A /dev/sda
>>> smartctl 5.41 2011-06-09 r3365 [x86_64-linux-3.2.0-4-amd64] (local
>>> build)
>>> Copyright (C) 2002-11 by Bruce Allen,
>>> http://smartmontools.sourceforge.net
>>>
>>> === START OF READ SMART DATA SECTION ===
>>> SMART Attributes Data Structure revision number: 10
>>> Vendor Specific SMART Attributes with Thresholds:
>>> ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED
>>> WHEN_FAILED RAW_VALUE
>>> 1 Raw_Read_Error_Rate 0x000f 056 053 006 Pre-fail
>>> - 208576601
>>> 3 Spin_Up_Time 0x0003 098 098 000 Pre-fail
>>> - 0
>>> 4 Start_Stop_Count 0x0032 100 100 020 Old_age
>>> - 0
>>> 5 Reallocated_Sector_Ct 0x0033 100 100 036 Pre-fail
>>> - 6
>>> 7 Seek_Error_Rate 0x000f 089 060 030 Pre-fail
>>> - 928963346
>>> 9 Power_On_Hours 0x0032 019 019 000 Old_age
>>> - 71036
>>> 10 Spin_Retry_Count 0x0013 100 100 097 Pre-fail
>>> - 0
>>> 12 Power_Cycle_Count 0x0032 100 100 020 Old_age
>>> - 362
>>> 194 Temperature_Celsius 0x0022 038 053 000 Old_age
>>> Always - 38
>>> 195 Hardware_ECC_Recovered 0x001a 056 053 000 Old_age
>>> Always - 208576601
>>> 197 Current_Pending_Sector 0x0012 100 100 000 Old_age
>>> Always - 1
>>> 198 Offline_Uncorrectable 0x0010 100 100 000 Old_age
>>> ne - 1
>>> 199 UDMA_CRC_Error_Count 0x003e 200 200 000 Old_age
>>> Always - 0
>>> 200 Multi_Zone_Error_Rate 0x0000 100 253 000 Old_age
>>> ne - 0
>>> 202 Data_Address_Mark_Errs 0x0032 100 253 000 Old_age
>>> Always - 0
>>>
>>> Those results look bad? :(
>>
>> Not too bad if they don't keep increasing.
>
> Hmm, I think they are now finally increasing a lot? From (py)dmesg's log:
> ...
> [2015-01-03 17:19:57] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0
> action 0x0
> [2015-01-03 17:19:57] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:19:57] ata1.00: failed command: READ DMA
> [2015-01-03 17:19:57] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag
> 0 dma 4096 in
> [2015-01-03 17:19:57] res 51/40:08:10:e0:3c/00:00:00:00:00/e0
> Emask 0x9 (media error)
> [2015-01-03 17:19:57] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:19:57] ata1.00: error: { UNC }
> [2015-01-03 17:19:57] ata1.00: configured for UDMA/100
> [2015-01-03 17:19:57] ata1: EH complete
> [2015-01-03 17:20:01] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0
> action 0x0
> [2015-01-03 17:20:01] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:20:01] ata1.00: failed command: READ DMA
> [2015-01-03 17:20:01] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag
> 0 dma 4096 in
> [2015-01-03 17:20:01] res 51/40:08:10:e0:3c/00:00:00:00:00/e0
> Emask 0x9 (media error)
> [2015-01-03 17:20:01] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:20:01] ata1.00: error: { UNC }
> [2015-01-03 17:20:01] ata1.00: configured for UDMA/100
> [2015-01-03 17:20:01] ata1: EH complete
> [2015-01-03 17:20:06] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0
> action 0x0
> [2015-01-03 17:20:06] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:20:06] ata1.00: failed command: READ DMA
> [2015-01-03 17:20:06] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag
> 0 dma 4096 in
> [2015-01-03 17:20:06] res 51/40:08:10:e0:3c/00:00:00:00:00/e0
> Emask 0x9 (media error)
> [2015-01-03 17:20:06] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:20:06] ata1.00: error: { UNC }
> [2015-01-03 17:20:06] ata1.00: configured for UDMA/100
> [2015-01-03 17:20:06] ata1: EH complete
> [2015-01-03 17:20:11] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0
> action 0x0
> [2015-01-03 17:20:11] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:20:11] ata1.00: failed command: READ DMA
> [2015-01-03 17:20:11] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag
> 0 dma 4096 in
> [2015-01-03 17:20:11] res 51/40:08:10:e0:3c/00:00:00:00:00/e0
> Emask 0x9 (media error)
> [2015-01-03 17:20:11] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:20:11] ata1.00: error: { UNC }
> [2015-01-03 17:20:11] ata1.00: configured for UDMA/100
> [2015-01-03 17:20:11] ata1: EH complete
> [2015-01-03 17:20:15] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0
> action 0x0
> [2015-01-03 17:20:15] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:20:15] ata1.00: failed command: READ DMA
> [2015-01-03 17:20:15] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag
> 0 dma 4096 in
> [2015-01-03 17:20:15] res 51/40:08:10:e0:3c/00:00:00:00:00/e0
> Emask 0x9 (media error)
> [2015-01-03 17:20:15] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:20:15] ata1.00: error: { UNC }
> [2015-01-03 17:20:15] ata1.00: configured for UDMA/100
> [2015-01-03 17:20:15] ata1: EH complete
> [2015-01-03 17:20:20] ata1.00: exception Emask 0x0 SAct 0x0 SErr 0x0
> action 0x0
> [2015-01-03 17:20:20] ata1.00: BMDMA stat 0x25
> [2015-01-03 17:20:20] ata1.00: failed command: READ DMA
> [2015-01-03 17:20:20] ata1.00: cmd c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag
> 0 dma 4096 in
> [2015-01-03 17:20:20] res 51/40:08:10:e0:3c/00:00:00:00:00/e0
> Emask 0x9 (media error)
> [2015-01-03 17:20:20] ata1.00: status: { DRDY ERR }
> [2015-01-03 17:20:20] ata1.00: error: { UNC }
> [2015-01-03 17:20:20] ata1.00: configured for UDMA/100
> [2015-01-03 17:20:20] sd 0:0:0:0: [sda] Unhandled sense code
> [2015-01-03 17:20:20] sd 0:0:0:0: [sda] Result: hostbyte=DID_OK
> driverbyte=DRIVER_SENSE
> [2015-01-03 17:20:20] sd 0:0:0:0: [sda] Sense Key : Medium Error
> [current] [descriptor]
> [2015-01-03 17:20:20] Descriptor sense data with sense descriptors (in
> hex):
> [2015-01-03 17:20:20] 72 03 11 04 00 00 00 0c 00 0a 80 00 00 00 00
> 00
> [2015-01-03 17:20:20] 00 3c e0 10
> [2015-01-03 17:20:20] sd 0:0:0:0: [sda] Add. Sense: Unrecovered read
> error - auto reallocate failed
> [2015-01-03 17:20:20] sd 0:0:0:0: [sda] CDB: Read(10): 28 00 00 3c e0 10
> 00 00 08 00
> [2015-01-03 17:20:20] end_request: I/O error, dev sda, sector 3989520
> [2015-01-03 17:20:20] EXT3-fs error (device sda5): ext3_get_inode_loc:
> unable to read inode block - inode=212577, block=425986
> [2015-01-03 17:20:20] ata1: EH complete
> ...
>
> --
>
> From: root <root@MyBox>
> To: root@MyBox
> Subject: SMART error (ErrorCount) detected on host: MyBox
>
> This email was generated by the smartd daemon running on:
>
> host name: MyBox
> DNS domain: [Unknown]
> NIS domain: (none)
>
> The following warning/error was logged by the smartd daemon:
>
> Device: /dev/sda [SAT], ATA error count increased from 0 to 6
>
> For details see host's SYSLOG.
>
> You can also use the smartctl utility for further investigation.
> Another email message will be sent in 24 hours if the problem persists.
>
> --
>
> From the /var/log/syslog:
> ...
> Jan 3 17:19:57 MyBox kernel: [292948.051212] ata1.00: exception Emask 0x0
> SAct 0x0 SErr 0x0 action 0x0
> Jan 3 17:19:57 MyBox kernel: [292948.051217] ata1.00: BMDMA stat 0x25
> Jan 3 17:19:57 MyBox kernel: [292948.051220] ata1.00: failed command:
> READ DMA
> Jan 3 17:19:57 MyBox kernel: [292948.051226] ata1.00: cmd
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan 3 17:19:57 MyBox kernel: [292948.051228] res
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan 3 17:19:57 MyBox kernel: [292948.051231] ata1.00: status: { DRDY
> ERR }
> Jan 3 17:19:57 MyBox kernel: [292948.051233] ata1.00: error: { UNC }
> Jan 3 17:19:57 MyBox kernel: [292948.136706] ata1.00: configured for
> UDMA/100
> Jan 3 17:19:57 MyBox kernel: [292948.136715] ata1: EH complete
> Jan 3 17:20:02 MyBox kernel: [292952.671441] ata1.00: exception Emask 0x0
> SAct 0x0 SErr 0x0 action 0x0
> Jan 3 17:20:02 MyBox kernel: [292952.671446] ata1.00: BMDMA stat 0x25
> Jan 3 17:20:02 MyBox kernel: [292952.671450] ata1.00: failed command:
> READ DMA
> Jan 3 17:20:02 MyBox kernel: [292952.671456] ata1.00: cmd
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan 3 17:20:02 MyBox kernel: [292952.671457] res
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan 3 17:20:02 MyBox kernel: [292952.671461] ata1.00: status: { DRDY
> ERR }
> Jan 3 17:20:02 MyBox kernel: [292952.671463] ata1.00: error: { UNC }
> Jan 3 17:20:02 MyBox kernel: [292952.740715] ata1.00: configured for
> UDMA/100
> Jan 3 17:20:02 MyBox kernel: [292952.740730] ata1: EH complete
> Jan 3 17:20:07 MyBox kernel: [292957.308333] ata1.00: exception Emask 0x0
> SAct 0x0 SErr 0x0 action 0x0
> Jan 3 17:20:07 MyBox kernel: [292957.308338] ata1.00: BMDMA stat 0x25
> Jan 3 17:20:07 MyBox kernel: [292957.308341] ata1.00: failed command:
> READ DMA
> Jan 3 17:20:07 MyBox kernel: [292957.308347] ata1.00: cmd
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan 3 17:20:07 MyBox kernel: [292957.308349] res
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan 3 17:20:07 MyBox kernel: [292957.308352] ata1.00: status: { DRDY
> ERR }
> Jan 3 17:20:07 MyBox kernel: [292957.308355] ata1.00: error: { UNC }
> Jan 3 17:20:07 MyBox kernel: [292957.420787] ata1.00: configured for
> UDMA/100
> Jan 3 17:20:07 MyBox kernel: [292957.420801] ata1: EH complete
> Jan 3 17:20:11 MyBox kernel: [292962.001820] ata1.00: exception Emask 0x0
> SAct 0x0 SErr 0x0 action 0x0
> Jan 3 17:20:11 MyBox kernel: [292962.001825] ata1.00: BMDMA stat 0x25
> Jan 3 17:20:11 MyBox kernel: [292962.001828] ata1.00: failed command:
> READ DMA
> Jan 3 17:20:11 MyBox kernel: [292962.001835] ata1.00: cmd
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan 3 17:20:11 MyBox kernel: [292962.001836] res
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan 3 17:20:11 MyBox kernel: [292962.001839] ata1.00: status: { DRDY
> ERR }
> Jan 3 17:20:11 MyBox kernel: [292962.001842] ata1.00: error: { UNC }
> Jan 3 17:20:11 MyBox kernel: [292962.088713] ata1.00: configured for
> UDMA/100
> Jan 3 17:20:11 MyBox kernel: [292962.088728] ata1: EH complete
> Jan 3 17:20:16 MyBox kernel: [292966.634963] ata1.00: exception Emask 0x0
> SAct 0x0 SErr 0x0 action 0x0
> Jan 3 17:20:16 MyBox kernel: [292966.634968] ata1.00: BMDMA stat 0x25
> Jan 3 17:20:16 MyBox kernel: [292966.634971] ata1.00: failed command:
> READ DMA
> Jan 3 17:20:16 MyBox kernel: [292966.634978] ata1.00: cmd
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan 3 17:20:16 MyBox kernel: [292966.634979] res
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan 3 17:20:16 MyBox kernel: [292966.634982] ata1.00: status: { DRDY
> ERR }
> Jan 3 17:20:16 MyBox kernel: [292966.634985] ata1.00: error: { UNC }
> Jan 3 17:20:16 MyBox kernel: [292966.720754] ata1.00: configured for
> UDMA/100
> Jan 3 17:20:16 MyBox kernel: [292966.720769] ata1: EH complete
> Jan 3 17:20:21 MyBox kernel: [292971.391149] ata1.00: exception Emask 0x0
> SAct 0x0 SErr 0x0 action 0x0
> Jan 3 17:20:21 MyBox kernel: [292971.391154] ata1.00: BMDMA stat 0x25
> Jan 3 17:20:21 MyBox kernel: [292971.391157] ata1.00: failed command:
> READ DMA
> Jan 3 17:20:21 MyBox kernel: [292971.391164] ata1.00: cmd
> c8/00:08:10:e0:3c/00:00:00:00:00/e0 tag 0 dma 4096 in
> Jan 3 17:20:21 MyBox kernel: [292971.391165] res
> 51/40:08:10:e0:3c/00:00:00:00:00/e0 Emask 0x9 (media error)
> Jan 3 17:20:21 MyBox kernel: [292971.391169] ata1.00: status: { DRDY
> ERR }
> Jan 3 17:20:21 MyBox kernel: [292971.391171] ata1.00: error: { UNC }
> Jan 3 17:20:21 MyBox kernel: [292971.500707] ata1.00: configured for
> UDMA/100
> Jan 3 17:20:21 MyBox kernel: [292971.500723] sd 0:0:0:0: [sda] Unhandled
> sense code
> Jan 3 17:20:21 MyBox kernel: [292971.500726] sd 0:0:0:0: [sda] Result:
> hostbyte=DID_OK driverbyte=DRIVER_SENSE
> Jan 3 17:20:21 MyBox kernel: [292971.500729] sd 0:0:0:0: [sda] Sense Key
> : Medium Error [current] [descriptor]
> Jan 3 17:20:21 MyBox kernel: [292971.500734] Descriptor sense data with
> sense descriptors (in hex):
> Jan 3 17:20:21 MyBox kernel: [292971.500736] 72 03 11 04 00 00 00
> 0c 00 0a 80 00 00 00 00 00
> Jan 3 17:20:21 MyBox kernel: [292971.500745] 00 3c e0 10
> Jan 3 17:20:21 MyBox kernel: [292971.500749] sd 0:0:0:0: [sda] Add.
> Sense: Unrecovered read error - auto reallocate failed
> Jan 3 17:20:21 MyBox kernel: [292971.500754] sd 0:0:0:0: [sda] CDB:
> Read(10): 28 00 00 3c e0 10 00 00 08 00
> Jan 3 17:20:21 MyBox kernel: [292971.500762] end_request: I/O error, dev
> sda, sector 3989520
> Jan 3 17:20:21 MyBox kernel: [292971.500778] EXT3-fs error (device sda5):
> ext3_get_inode_loc: unable to read inode block - inode=212577,
> block=425986
> Jan 3 17:20:21 MyBox kernel: [292971.500782] ata1: EH complete
> Jan 3 17:25:01 MyBox /USR/SBIN/CRON[24006]: (root) CMD (command -v
> debian-sa1 > /dev/null && debian-sa1 1 1)
> Jan 3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], 62448
> Currently unreadable (pending) sectors (changed +55)
> Jan 3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], 62448 Offline
> uncorrectable sectors (changed +55)
> Jan 3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART
> Prefailure Attribute: 1 Raw_Read_Error_Rate changed from 74 to 71
> Jan 3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART Usage
> Attribute: 194 Temperature_Celsius changed from 27 to 28
> Jan 3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART Usage
> Attribute: 195 Hardware_ECC_Recovered changed from 74 to 71
> Jan 3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART Usage
> Attribute: 202 Data_Address_Mark_Errs changed from 31 to 20
> Jan 3 17:27:39 MyBox smartd[3440]: Device: /dev/sda [SAT], ATA error
> count increased from 0 to 6
> Jan 3 17:27:39 MyBox smartd[3440]: Sending warning via
> /usr/share/smartmontools/smartd-runner to root ...
> Jan 3 17:27:39 MyBox smartd[3440]: Warning via
> /usr/share/smartmontools/smartd-runner to root: successful
> Jan 3 17:27:39 MyBox smartd[3440]: Device: /dev/sdb [SAT], SMART Usage
> Attribute: 194 Temperature_Celsius changed from 25 to 26
> Jan 3 17:27:54 MyBox kernel: [293424.703922] device eth0 entered
> promiscuous mode
> ...
>
> I was using it with my very old Windows XP Pro SP3 VM on it. It was fine
> to me. No errors and slow downs. So, I ran smartctl's short and long tests
> and they both failed quickly with 10%. I even got an e-mail:
>
> Date: Sat, 03 Jan 2015 18:57:42 -0800
> From: root <root@MyBox>
> To: root@MyBox
> Subject: SMART error (SelfTest) detected on host: MyBox
>
> This email was generated by the smartd daemon running on:
>
> host name: MyBox
> DNS domain: [Unknown]
> NIS domain: (none)
>
> The following warning/error was logged by the smartd daemon:
>
> Device: /dev/sda [SAT], Self-Test Log error count increased from 3 to 5
>
> For details see host's SYSLOG.
>
> You can also use the smartctl utility for further investigation.
> The original email about this issue was sent at Sat Apr 19 09:55:41 2014
> PDT
> Another email message will be sent in 24 hours if the problem persists.
>
> /var/log/syslog:
> ...
> Jan 3 18:57:42 MyBox smartd[3440]: Device: /dev/sda [SAT], 62485 Offline
> uncorrectable sectors (changed +24)
> Jan 3 18:57:42 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART
> Prefailure Attribute: 1 Raw_Read_Error_Rate changed from 66 to 63
> Jan 3 18:57:42 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART Usage
> Attribute: 195 Hardware_ECC_Recovered changed from 66 to 63
> Jan 3 18:57:42 MyBox smartd[3440]: Device: /dev/sda [SAT], SMART Usage
> Attribute: 202 Data_Address_Mark_Errs changed from 0 to 249
> Jan 3 18:57:42 MyBox smartd[3440]: Device: /dev/sda [SAT], Self-Test Log
> error count increased from 3 to 5
> Jan 3 18:57:42 MyBox smartd[3440]: Sending warning via
> /usr/share/smartmontools/smartd-runner to root ...
> Jan 3 18:57:42 MyBox smartd[3440]: Warning via
> /usr/share/smartmontools/smartd-runner to root: successful
> ...
>
> # smartctl -a /dev/sda
> smartctl 5.41 2011-06-09 r3365 [x86_64-linux-3.2.0-4-amd64] (local build)
> Copyright (C) 2002-11 by Bruce Allen, http://smartmontools.sourceforge.net
>
> === START OF INFORMATION SECTION ===
> Model Family: Seagate Barracuda 7200.7 and 7200.7 Plus
> Device Model: ST380011A
> Serial Number: 4JV5P7LN
> Firmware Version: 8.01
> User Capacity: 80,026,361,856 bytes [80.0 GB]
> Sector Size: 512 bytes logical/physical
> Device is: In smartctl database [for details use: -P show]
> ATA Version is: 6
> ATA Standard is: ATA/ATAPI-6 T13 1410D revision 2
> Local Time is: Sat Jan 3 19:02:26 2015 PST
> SMART support is: Available - device has SMART capability.
> SMART support is: Enabled
>
> === START OF READ SMART DATA SECTION ===
> SMART overall-health self-assessment test result: PASSED
>
> General SMART Values:
> Offline data collection status: (0x82) Offline data collection activity
> was completed without error.
> Auto Offline Data Collection:
> Enabled.
> Self-test execution status: ( 121) The previous self-test completed
> having
> the read element of the test
> failed.
> Total time to complete Offline
> data collection: ( 430) seconds.
> Offline data collection
> capabilities: (0x5b) SMART execute Offline immediate.
> Auto Offline data collection
> on/off support.
> Suspend Offline collection upon
> new
> command.
> Offline surface scan supported.
> Self-test supported.
> No Conveyance Self-test supported.
> Selective Self-test supported.
> SMART capabilities: (0x0003) Saves SMART data before entering
> power-saving mode.
> Supports SMART auto save timer.
> Error logging capability: (0x01) Error logging supported.
> General Purpose Logging supported.
> Short self-test routine
> recommended polling time: ( 1) minutes.
> Extended self-test routine
> recommended polling time: ( 58) minutes.
>
> SMART Attributes Data Structure revision number: 10
> Vendor Specific SMART Attributes with Thresholds:
> ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED
> WHEN_FAILED RAW_VALUE
> 1 Raw_Read_Error_Rate 0x000f 063 046 006 Pre-fail
> - 165638319
> 3 Spin_Up_Time 0x0003 098 098 000 Pre-fail
> - 0
> 4 Start_Stop_Count 0x0032 100 100 020 Old_age
> - 0
> 5 Reallocated_Sector_Ct 0x0033 053 053 036 Pre-fail
> - 1884
That's very bad.
> 7 Seek_Error_Rate 0x000f 086 060 030 Pre-fail
> - 426361585
> 9 Power_On_Hours 0x0032 012 012 000 Old_age
> - 77139
> 10 Spin_Retry_Count 0x0013 100 100 097 Pre-fail
> - 0
> 12 Power_Cycle_Count 0x0032 100 100 020 Old_age
> - 370
> 194 Temperature_Celsius 0x0022 029 053 000 Old_age Always -
> 29
> 195 Hardware_ECC_Recovered 0x001a 063 046 000 Old_age Always -
> 165638319
> 197 Current_Pending_Sector 0x0012 001 001 000 Old_age Always -
> 62493
That's very bad.
> 198 Offline_Uncorrectable 0x0010 001 001 000 Old_age
> ne - 62493
Ditto.
> 199 UDMA_CRC_Error_Count 0x003e 200 200 000 Old_age Always -
> 0
> 200 Multi_Zone_Error_Rate 0x0000 100 253 000 Old_age
> ne - 0
> 202 Data_Address_Mark_Errs 0x0032 249 146 000 Old_age Always -
> 107
>
> SMART Error Log Version: 1
> ATA Error Count: 6 (device log contains only the most recent five errors)
> CR = Command Register [HEX]
> FR = Features Register [HEX]
> SC = Sector Count Register [HEX]
> SN = Sector Number Register [HEX]
> CL = Cylinder Low Register [HEX]
> CH = Cylinder High Register [HEX]
> DH = Device/Head Register [HEX]
> DC = Device Command Register [HEX]
> ER = Error register [HEX]
> ST = Status register [HEX]
> Powered_Up_Time is measured from power on, and printed as
> DDd+hh:mm:SS.sss where DD=days, hh=hours, mm=minutes,
> SS=sec, and sss=millisec. It "wraps" after 49.710 days.
>
> Error 6 occurred at disk power-on lifetime: 11601 hours (483 days + 9
> hours)
> When the command that caused the error occurred, the device was active
> or idle.
>
> After command completion occurred, registers were:
> ER ST SC SN CL CH DH
> -- -- -- -- -- -- --
> 40 51 08 10 e0 3c e0 Error: UNC 8 sectors at LBA = 0x003ce010 = 3989520
>
> Commands leading to the command that caused the error were:
> CR FR SC SN CL CH DH DC Powered_Up_Time Command/Feature_Name
> -- -- -- -- -- -- -- -- ---------------- --------------------
> c8 00 08 10 e0 3c e0 00 11:28:27.346 READ DMA
> 27 00 00 00 00 00 e0 00 11:28:27.334 READ NATIVE MAX ADDRESS EXT
> ec 00 00 00 00 00 a0 02 11:28:36.815 IDENTIFY DEVICE
> ef 03 45 00 00 00 a0 02 11:28:36.814 SET FEATURES [Set transfer
> mode]
> 27 00 00 00 00 00 e0 00 11:28:36.807 READ NATIVE MAX ADDRESS EXT
>
> Error 5 occurred at disk power-on lifetime: 11601 hours (483 days + 9
> hours)
> When the command that caused the error occurred, the device was active
> or idle.
>
> After command completion occurred, registers were:
> ER ST SC SN CL CH DH
> -- -- -- -- -- -- --
> 40 51 08 10 e0 3c e0 Error: UNC 8 sectors at LBA = 0x003ce010 = 3989520
>
> Commands leading to the command that caused the error were:
> CR FR SC SN CL CH DH DC Powered_Up_Time Command/Feature_Name
> -- -- -- -- -- -- -- -- ---------------- --------------------
> c8 00 08 10 e0 3c e0 00 11:28:27.346 READ DMA
> 27 00 00 00 00 00 e0 00 11:28:27.334 READ NATIVE MAX ADDRESS EXT
> ec 00 00 00 00 00 a0 02 11:28:27.241 IDENTIFY DEVICE
> ef 03 45 00 00 00 a0 02 11:28:27.240 SET FEATURES [Set transfer
> mode]
> 27 00 00 00 00 00 e0 00 11:28:27.232 READ NATIVE MAX ADDRESS EXT
>
> Error 4 occurred at disk power-on lifetime: 11601 hours (483 days + 9
> hours)
> When the command that caused the error occurred, the device was active
> or idle.
>
> After command completion occurred, registers were:
> ER ST SC SN CL CH DH
> -- -- -- -- -- -- --
> 40 51 08 10 e0 3c e0 Error: UNC 8 sectors at LBA = 0x003ce010 = 3989520
>
> Commands leading to the command that caused the error were:
> CR FR SC SN CL CH DH DC Powered_Up_Time Command/Feature_Name
> -- -- -- -- -- -- -- -- ---------------- --------------------
> c8 00 08 10 e0 3c e0 00 11:28:27.346 READ DMA
> 27 00 00 00 00 00 e0 00 11:28:27.334 READ NATIVE MAX ADDRESS EXT
> ec 00 00 00 00 00 a0 02 11:28:27.241 IDENTIFY DEVICE
> ef 03 45 00 00 00 a0 02 11:28:27.240 SET FEATURES [Set transfer
> mode]
> 27 00 00 00 00 00 e0 00 11:28:27.232 READ NATIVE MAX ADDRESS EXT
>
> Error 3 occurred at disk power-on lifetime: 11601 hours (483 days + 9
> hours)
> When the command that caused the error occurred, the device was active
> or idle.
>
> After command completion occurred, registers were:
> ER ST SC SN CL CH DH
> -- -- -- -- -- -- --
> 40 51 08 10 e0 3c e0 Error: UNC 8 sectors at LBA = 0x003ce010 = 3989520
>
> Commands leading to the command that caused the error were:
> CR FR SC SN CL CH DH DC Powered_Up_Time Command/Feature_Name
> -- -- -- -- -- -- -- -- ---------------- --------------------
> c8 00 08 10 e0 3c e0 00 11:28:17.813 READ DMA
> 27 00 00 00 00 00 e0 00 11:28:17.808 READ NATIVE MAX ADDRESS EXT
> ec 00 00 00 00 00 a0 02 11:28:13.111 IDENTIFY DEVICE
> ef 03 45 00 00 00 a0 02 11:28:13.107 SET FEATURES [Set transfer
> mode]
> 27 00 00 00 00 00 e0 00 11:28:13.092 READ NATIVE MAX ADDRESS EXT
>
> Error 2 occurred at disk power-on lifetime: 11601 hours (483 days + 9
> hours)
> When the command that caused the error occurred, the device was active
> or idle.
>
> After command completion occurred, registers were:
> ER ST SC SN CL CH DH
> -- -- -- -- -- -- --
> 40 51 08 10 e0 3c e0 Error: UNC 8 sectors at LBA = 0x003ce010 = 3989520
>
> Commands leading to the command that caused the error were:
> CR FR SC SN CL CH DH DC Powered_Up_Time Command/Feature_Name
> -- -- -- -- -- -- -- -- ---------------- --------------------
> c8 00 08 10 e0 3c e0 00 11:28:17.813 READ DMA
> 27 00 00 00 00 00 e0 00 11:28:17.808 READ NATIVE MAX ADDRESS EXT
> ec 00 00 00 00 00 a0 02 11:28:13.111 IDENTIFY DEVICE
> ef 03 45 00 00 00 a0 02 11:28:13.107 SET FEATURES [Set transfer
> mode]
> 27 00 00 00 00 00 e0 00 11:28:13.092 READ NATIVE MAX ADDRESS EXT
>
> SMART Self-test log structure revision number 1
> Num Test_Description Status Remaining LifeTime(hours)
> LBA_of_first_error
> # 1 Extended offline Completed: read failure 90% 11603
> 3989520
> # 2 Short offline Completed: read failure 90% 11602
> 3989520
> # 3 Extended offline Completed: read failure 90% 10547
> 616217
> # 4 Extended offline Completed: read failure 90% 9989
> 684994
> # 5 Extended offline Completed: read failure 90% 9785
> 3783997
> # 6 Short offline Completed without error 00% 9785 -
> # 7 Extended offline Completed without error 00% 5853 -
> # 8 Extended offline Completed without error 00% 5815 -
> # 9 Extended offline Completed without error 00% 5583 -
> #10 Short offline Completed without error 00% 5583 -
> #11 Extended offline Completed: read failure 60% 5502
> 67691664
> #12 Extended offline Completed: read failure 60% 5501
> 67691664
> #13 Short offline Completed without error 00% 5501 -
> #14 Short offline Completed without error 00% 5501 -
> #15 Extended offline Completed: read failure 60% 5501
> 67691664
> #16 Short offline Completed without error 00% 5501 -
> #17 Extended offline Completed: read failure 60% 5499
> 67691664
> #18 Extended offline Completed without error 00% 3165 -
> #19 Extended offline Completed without error 00% 48483 -
> #20 Extended offline Completed without error 00% 43207 -
> #21 Extended offline Interrupted (host reset) 90% 43201 -
> 4 of 9 failed self-tests are outdated by newer successful extended offline
> self-test # 7
>
> SMART Selective self-test log data structure revision number 1
> SPAN MIN_LBA MAX_LBA CURRENT_TEST_STATUS
> 1 0 0 Not_testing
> 2 0 0 Not_testing
> 3 0 0 Not_testing
> 4 0 0 Not_testing
> 5 0 0 Not_testing
> Selective self-test flags (0x0):
> After scanning selected spans, do NOT read-scan remainder of disk.
> If Selective self-test is pending on power-up, resume after 0 minute
> delay.
>
>
> Is it finally dying to you now? :P
Yep, she's dead, Jim. You into necrophilia ?
Back to comp.os.linux.hardware | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-19 09:57 -0700
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-19 10:55 -0700
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-19 11:56 -0700
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-20 06:56 +1000
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-19 17:34 -0700
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-20 13:02 +1000
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-19 23:16 -0700
Re: A dying very old HDD confirmed? Robert Nichols <SEE_SIGNATURE@localhost.localdomain.invalid> - 2014-04-20 07:19 -0500
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-20 09:54 -0700
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-20 09:57 -0700
Re: A dying very old HDD confirmed? Robert Nichols <SEE_SIGNATURE@localhost.localdomain.invalid> - 2014-04-21 07:56 -0500
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-21 08:08 -0700
Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-22 01:31 +0200
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-21 19:02 -0700
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 05:24 +1000
Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-04-22 19:03 -0500
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 14:42 +1000
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-04-22 23:27 -0700
Re: A dying very old HDD confirmed? Robert Nichols <SEE_SIGNATURE@localhost.localdomain.invalid> - 2014-04-23 09:03 -0500
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-24 06:32 +1000
Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-24 00:26 +0200
Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-04-23 18:47 -0500
Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-10-19 21:05 -0500
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-10-21 05:13 +1100
Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-10-20 16:04 -0500
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-10-22 07:01 +1100
Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-10-21 15:56 -0500
Re: A dying very old HDD confirmed? Arno <me@privacy.net> - 2014-10-22 22:43 +0000
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2014-10-24 08:18 -0700
Re: A dying very old HDD confirmed? Arno <me@privacy.net> - 2014-10-24 15:39 +0000
Re: A dying very old HDD confirmed? ANTant@zimage.com (Ant) - 2014-10-24 20:23 -0500
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-24 11:20 +1000
Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-24 09:47 +0200
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-24 06:35 +1000
Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-24 00:23 +0200
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-24 11:17 +1000
Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-24 09:55 +0200
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-25 05:17 +1000
Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-27 19:12 +0200
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-28 19:59 +1000
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 05:22 +1000
Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-20 11:08 +0200
Re: A dying very old HDD confirmed? Richard Kettlewell <rjk@greenend.org.uk> - 2014-04-20 10:13 +0100
Re: A dying very old HDD confirmed? Pascal Hambourg <boite-a-spam@plouf.fr.eu.org> - 2014-04-20 12:28 +0200
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-21 05:52 +1000
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-21 05:50 +1000
Re: A dying very old HDD confirmed? cjt <cheljuba@prodigy.net> - 2014-04-20 19:35 -0500
Re: A dying very old HDD confirmed? Måns Rullgård <mans@mansr.com> - 2014-04-22 12:59 +0100
Re: A dying very old HDD confirmed? Vic RR Garcia <VicGar007@at-gmail.dot.com> - 2014-04-22 13:23 -0400
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 05:30 +1000
Re: A dying very old HDD confirmed? Måns Rullgård <mans@mansr.com> - 2014-04-22 20:57 +0100
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 14:37 +1000
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-23 05:28 +1000
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2015-01-03 19:11 -0800
Re: A dying very old HDD confirmed? Robert Nichols <SEE_SIGNATURE@localhost.localdomain.invalid> - 2015-01-04 12:11 -0600
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2015-01-04 12:38 -0800
Re: A dying very old HDD confirmed? mjb@signal11.invalid (Mike) - 2015-01-04 23:02 +0000
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2015-01-04 15:21 -0800
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2015-01-05 08:18 +1100
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2015-01-04 15:21 -0800
Re: A dying very old HDD confirmed? Ant <ant@zimage.comANT> - 2015-01-04 06:44 -0800
Re: A dying very old HDD confirmed? "Rod Speed" <rod.speed.aaa@gmail.com> - 2014-04-20 06:55 +1000
csiph-web