Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1664903
| From | John Garry <john.garry@huawei.com> |
|---|---|
| Newsgroups | linux.kernel |
| Subject | Re: [PATCH 20/22] scsi: hisi_sas: Add v3 code to support ECC and AXI bus fatal error |
| Date | 2017-06-13 16:10 +0200 |
| Message-ID | <tRUG6-w3-29@gated-at.bofh.it> (permalink) |
| References | <tI4dH-1LO-3@gated-at.bofh.it> <tI4np-1Pa-45@gated-at.bofh.it> <tI6fw-32g-5@gated-at.bofh.it> <tI6pb-35o-9@gated-at.bofh.it> |
| Organization | linux.* mail to news gateway |
On 17/05/2017 13:38, John Garry wrote: > On 17/05/2017 13:27, Arnd Bergmann wrote: >> On Wed, May 17, 2017 at 12:49 PM, John Garry <john.garry@huawei.com> >> wrote: >>> > From: Xiang Chen <chenxiang66@hisilicon.com> >>> > >>> > For ECC 1bit error, logic can recover it, so we only print a warning. >>> > For ECC multi-bit and AXI bus fatal error, we panic. >>> > >>> > Signed-off-by: John Garry <john.garry@huawei.com> >>> > Signed-off-by: Xiang Chen <chenxiang66@hisilicon.com> >> This one is tricky as there are conflicting requirements: >> >> - For debugging purposes, you want to continue running the system >> to figure out what exactly went wrong. Often enough, having the >> kernel panic means you don't get to see the panic message because >> console access is unavailable and you cannot log in any more >> >> - For data consistency purposes you want to stop the system as >> soon as there is any uncorrectable data error >> >> I see that most scsi drivers don't ever call panic or BUG(), though >> you already do so for v1 and v2 hw. >> >> Maybe the SCSI maintainers can provide some more guidance here. >> >> Arnd >> >> . >> > > Hi Arnd, > > Actually latest code for v2 has been updated to do a controller reset, > and not panic, for unrecoverable error: > https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/drivers/scsi/hisi_sas/hisi_sas_v2_hw.c?h=v4.12-rc1#n2926 > > > We never got around to implementing controller reset for v1 as this > platform (hip05) is not used much anymore. > > As for v3, we will change to to do same once controller reset is > implemented. I should have added this to the commit log. > > Thanks, > John It has come to light that hip08 RAS architecture requires handling certain errors with firmware first model. I am not sure on the flow of controller reset for fatal errors - I'm currently checking the details. But it is not worth adding this non-critical patch and reverting it later, so I'll omit this patch when sending the v6 patchset which includes the fix for sloppy spinlock usage. Thanks, John
Back to linux.kernel | Previous | Next | Find similar | Unroll thread
Re: [PATCH 20/22] scsi: hisi_sas: Add v3 code to support ECC and AXI bus fatal error John Garry <john.garry@huawei.com> - 2017-06-13 16:10 +0200
csiph-web