Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1425986 > unrolled thread
| Started by | Thorsten Leemhuis <regressions@leemhuis.info> |
|---|---|
| First post | 2016-06-19 17:00 +0200 |
| Last post | 2016-06-22 08:40 +0200 |
| Articles | 13 — 8 participants |
Back to article view | Back to linux.kernel
Reported regressions for 4.7 as of Sunday, 2016-06-19 Thorsten Leemhuis <regressions@leemhuis.info> - 2016-06-19 17:00 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 Christoph Hellwig <hch@infradead.org> - 2016-06-20 13:30 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 Josh Boyer <jwboyer@fedoraproject.org> - 2016-06-21 13:20 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-21 22:50 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 Josh Boyer <jwboyer@fedoraproject.org> - 2016-06-22 03:00 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 "Martin K. Petersen" <martin.petersen@oracle.com> - 2016-06-22 03:30 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 Quinn Tran <quinn.tran@qlogic.com> - 2016-06-22 06:10 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 Johannes Thumshirn <jthumshirn@suse.de> - 2016-06-22 14:10 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 Quinn Tran <quinn.tran@qlogic.com> - 2016-06-22 18:00 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 Johannes Thumshirn <jthumshirn@suse.de> - 2016-06-23 09:30 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 Quinn Tran <quinn.tran@qlogic.com> - 2016-06-23 18:20 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-23 18:40 +0200
Re: Reported regressions for 4.7 as of Sunday, 2016-06-19 Kalle Valo <kvalo@codeaurora.org> - 2016-06-22 08:40 +0200
| From | Thorsten Leemhuis <regressions@leemhuis.info> |
|---|---|
| Date | 2016-06-19 17:00 +0200 |
| Subject | Reported regressions for 4.7 as of Sunday, 2016-06-19 |
| Message-ID | <rLMmC-3I1-9@gated-at.bofh.it> |
Hi! Here is my second regression report for 4.7. It has 19 entries; 8 of them are new; 8 regressions were fixed since the last report (those are not included in this report) and I dropped 2 which turned out to not be regressions after all (at least that's what I think right now). FWIW, it's still a lot of work to generate this report (as expected). I'm still thinking about a plan how to make the whole tracking process easier and more attractive for everyone, but it will take a few weeks before I come up with a concrete plan. HTH, CU, Thorsten (¹) last weeks report was http://article.gmane.org/gmane.linux.kernel/2241805 P.S.: Please let me know if a regression is missing in the list; or if there is something on the list which shouldn't be there. ---- Description: ath10k no longer authenticates and freezes system Report: https://bugzilla.kernel.org/show_bug.cgi?id=119151 Latest status: http://thread.gmane.org/gmane.linux.kernel.wireless.general/152513/focus=152535 Date rep/stat: 2016-05-27 / 2016-06-02 Notes: forgotten? poked bug report on Friday Description: Bad flicker on skylake HQD due to code in the 4.7 merge window Report: http://thread.gmane.org/gmane.linux.kernel/2230377 Latest status: http://thread.gmane.org/gmane.linux.kernel/2230377/focus=92602 Date rep/stat: 2016-05-30 / 2016-06-18 Notes: investigation ongoing Description: we noticed reaim.jobs_per_min -49.1% regression Report: http://thread.gmane.org/gmane.linux.kernel/2231025/ Latest status: http://thread.gmane.org/gmane.linux.kernel/2231025/focus=2233571 Date rep/stat: 2016-05-31 / 2016-06-13 Notes: wip? http://article.gmane.org/gmane.linux.kernel/2241911 Description: NULL pointer dereference with BCM4350 wireless device Report: https://bugzilla.kernel.org/show_bug.cgi?id=119451 Latest status: 7.6. Date rep/stat: 2016-06-01 / 2016-06-07 Notes: poked bugzilla, likely fixed in mainline by https://git.kernel.org/torvalds/c/31143e2933 Description: 795ae7a0de: pixz.throughput -9.1% regression Report: http://thread.gmane.org/gmane.linux.kernel/2233056/ Latest status: http://thread.gmane.org/gmane.linux.kernel/2233056/focus=2238208 Date rep/stat: 2016-06-02 / 2016-06-08 Notes: @regression tracker: poke someone Description: RadeonSI get a huge performance dip with used with the nine state tracker Report: https://bugzilla.kernel.org/show_bug.cgi?id=119631 Latest status: https://bugzilla.kernel.org/show_bug.cgi?id=119631#c12 Date rep/stat: 2016-06-04 / 2016-06-15 Notes: investigation ongoing, waiting for reporter Description: 5c0a85fad9: unixbench.score -6.3% regression Report: http://thread.gmane.org/gmane.linux.kernel/2235794 Latest status: http://thread.gmane.org/gmane.linux.kernel.mm/153151/focus=153409 Date rep/stat: 2016-06-06 / 2016-06-17 Notes: wip, revert discussed Description: System hang possibly due to brcmfmac regression Report: https://bugzilla.kernel.org/show_bug.cgi?id=119761 Latest status: https://bugzilla.kernel.org/show_bug.cgi?id=119761#c1 Date rep/stat: 2016-06-07 / 2016-06-12 Notes: might be fixed, waiting for clarification from reporter Description: Regression in kbuild: fix if_change and friends to consider argument Report: http://thread.gmane.org/gmane.linux.kbuild.devel/14981/ Latest status: http://thread.gmane.org/gmane.linux.kbuild.devel/14981/focus=15000 Date rep/stat: 2016-06-07 / 2016-06-09 Notes: patch in linux-next Description: BUG: using smp_processor_id() in preemptible [00000000] code] when using a USB Mass Storage device Report: http://thread.gmane.org/gmane.linux.usb.general/143504 Latest status: http://thread.gmane.org/gmane.linux.usb.general/143504/focus=153154 https://lkml.org/lkml/2016/6/15/397 Date rep/stat: 2016-06-09 / 2016-06-15 Notes: investigation ongoing Description: Notebook Clevo N350DW i5-6500T freezes on shutdown (but reboots fine) Report: https://bugzilla.kernel.org/show_bug.cgi?id=119871 Latest status: https://bugzilla.kernel.org/show_bug.cgi?id=119871#c8 Date rep/stat: 2016-06-09 / 2016-06-16 Notes: reporter needs help to provide more details to debug the problem Description: BUG() in dmesg after loading nouveau module Report: https://bugzilla.kernel.org/show_bug.cgi?id=120591 Latest status: https://bugzilla.kernel.org/show_bug.cgi?id=120591#c3 Date rep/stat: 2016-06-18 / 2016-06-19 Notes: wip Description: BUG: unable to handle kernel NULL pointer dereference […] qla24xx_process_response_queue+0x49/0x4b0 [qla2xxx] Report: https://bugzilla.kernel.org/show_bug.cgi?id=120201 Latest status: n/a Date rep/stat: 2016-06-14 / n/a Notes: poked bugzilla, a bit unsure how to proceed Description: Performance drop 30-40% for SPECjbb2005 and SPECjvm2008 benchmarks Report: https://bugzilla.kernel.org/show_bug.cgi?id=120481 Latest status: n/a Date rep/stat: 2016-06-16 / n/a Notes: real reason unknown Description: performance drop on SFC interface around 30 % Report: https://bugzilla.kernel.org/show_bug.cgi?id=120461 Latest status: https://bugzilla.kernel.org/show_bug.cgi?id=120461#c9 Date rep/stat: 2016-06-17 / 2016-06-17 Notes: wip Description: System hang when plug/un-plug USB 3.1 key via thunderbolt port on Dell XPS 13 Report: https://bugzilla.kernel.org/show_bug.cgi?id=120241 Latest status: n/a Date rep/stat: 2016-06-14 / n/a Notes: waiting for reporter Description: performance regression on Jetson TK1 since 4.7-rc1: moving windows under X would become unsufferably slow, and graphical performance under X in general is seriously degraded Report: http://thread.gmane.org/gmane.linux.ports.tegra/26983/focus=2245415 Latest status: n/a Date rep/stat: 2016-06-16 / n/a Notes: wip, fix available Description: lk 4.7 regression: EDAC, amd64_edac: Drop pci_register_driver() use Report: http://thread.gmane.org/gmane.linux.kernel/2245115/ Latest status: http://thread.gmane.org/gmane.linux.kernel/2246008/focus=2246009 Date rep/stat: 2016-06-15 / 2016-06-16 Notes: wip, fix available Description: regression in 8250 uart driver Report: http://thread.gmane.org/gmane.linux.kernel/2243130/focus=2243653 Latest status: http://thread.gmane.org/gmane.linux.kernel/2243130/focus=2243653 Date rep/stat: 2016-06-14 / 2016-06-14 Notes: wip, fix available
[toc] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-06-20 13:30 +0200 |
| Message-ID | <rM5yW-7zy-7@gated-at.bofh.it> |
| In reply to | #1425986 |
Another important one is the rename regression in XFS and ext4 that I suspect is due the VFS changes in 4.7: http://oss.sgi.com/pipermail/xfs/2016-June/049138.html http://oss.sgi.com/pipermail/xfs/2016-June/049309.html possibly related: http://marc.info/?l=linux-kernel&m=146605889024559&w=2
[toc] | [prev] | [next] | [standalone]
| From | Josh Boyer <jwboyer@fedoraproject.org> |
|---|---|
| Date | 2016-06-21 13:20 +0200 |
| Message-ID | <rMrSN-579-11@gated-at.bofh.it> |
| In reply to | #1425986 |
On Sun, Jun 19, 2016 at 10:52 AM, Thorsten Leemhuis <regressions@leemhuis.info> wrote: > Description: BUG: unable to handle kernel NULL pointer dereference […] qla24xx_process_response_queue+0x49/0x4b0 [qla2xxx] > Report: https://bugzilla.kernel.org/show_bug.cgi?id=120201 > Latest status: n/a > Date rep/stat: 2016-06-14 / n/a > Notes: poked bugzilla, a bit unsure how to proceed We have two bug reports against 4.5.5 - 4.5.7 of this as well. So whatever commit caused this in 4.7 seems to have been pulled into the 4.5.y stable tree. I suspect it is in the 4.6.y stable tree as well, but we don't have that pushed out yet. https://bugzilla.redhat.com/show_bug.cgi?id=1348342 https://bugzilla.redhat.com/show_bug.cgi?id=1346753 josh
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2016-06-21 22:50 +0200 |
| Message-ID | <rMAMq-2gO-29@gated-at.bofh.it> |
| In reply to | #1427613 |
On Tue, Jun 21, 2016 at 4:11 AM, Josh Boyer <jwboyer@fedoraproject.org> wrote:
> On Sun, Jun 19, 2016 at 10:52 AM, Thorsten Leemhuis
> <regressions@leemhuis.info> wrote:
>> Description: BUG: unable to handle kernel NULL pointer dereference […] qla24xx_process_response_queue+0x49/0x4b0 [qla2xxx]
>> Report: https://bugzilla.kernel.org/show_bug.cgi?id=120201
>> Latest status: n/a
>> Date rep/stat: 2016-06-14 / n/a
>> Notes: poked bugzilla, a bit unsure how to proceed
>
> We have two bug reports against 4.5.5 - 4.5.7 of this as well. So
> whatever commit caused this in 4.7 seems to have been pulled into the
> 4.5.y stable tree. I suspect it is in the 4.6.y stable tree as well,
> but we don't have that pushed out yet.
>
> https://bugzilla.redhat.com/show_bug.cgi?id=1348342
> https://bugzilla.redhat.com/show_bug.cgi?id=1346753
That seems pretty unambiguous - 4.5.5 is fine, and 4.5.6 is bad. So
unless it's specific to whatever patches RH is carrying around, we
should be able to just look at the scsi-related stable tree patches in
that region. That seems simple enough.
But theres' really only two (trivial) patches in there:
- scsi: Add intermediate STARGET_REMOVE state to scsi_target_state
(f05795d3d771f30a7bdc3a138bf714b06d42aa95 upstream)
- Revert "scsi: fix soft lockup in scsi_remove_target() on module removal"
(305c2e71b3d733ec065cb716c76af7d554bd5571 upstream)
as far as I can tell. And neither of them looks very likely, but what
do I know. Adding Martin Petersen and Johannes Thumshirn to the
participants just in case they go "Ahh.."
Linus
[toc] | [prev] | [next] | [standalone]
| From | Josh Boyer <jwboyer@fedoraproject.org> |
|---|---|
| Date | 2016-06-22 03:00 +0200 |
| Message-ID | <rMEGm-4IG-3@gated-at.bofh.it> |
| In reply to | #1428153 |
On Tue, Jun 21, 2016 at 4:40 PM, Linus Torvalds <torvalds@linux-foundation.org> wrote: > On Tue, Jun 21, 2016 at 4:11 AM, Josh Boyer <jwboyer@fedoraproject.org> wrote: >> On Sun, Jun 19, 2016 at 10:52 AM, Thorsten Leemhuis >> <regressions@leemhuis.info> wrote: >>> Description: BUG: unable to handle kernel NULL pointer dereference […] qla24xx_process_response_queue+0x49/0x4b0 [qla2xxx] >>> Report: https://bugzilla.kernel.org/show_bug.cgi?id=120201 >>> Latest status: n/a >>> Date rep/stat: 2016-06-14 / n/a >>> Notes: poked bugzilla, a bit unsure how to proceed >> >> We have two bug reports against 4.5.5 - 4.5.7 of this as well. So >> whatever commit caused this in 4.7 seems to have been pulled into the >> 4.5.y stable tree. I suspect it is in the 4.6.y stable tree as well, >> but we don't have that pushed out yet. >> >> https://bugzilla.redhat.com/show_bug.cgi?id=1348342 >> https://bugzilla.redhat.com/show_bug.cgi?id=1346753 > > That seems pretty unambiguous - 4.5.5 is fine, and 4.5.6 is bad. So > unless it's specific to whatever patches RH is carrying around, we > should be able to just look at the scsi-related stable tree patches in > that region. That seems simple enough. I thought the same. We're only carrying one very very old scsi patch to revalidate a pointer. That shouldn't even been involved in this path and upstream 4.7-rcX is hitting the same issue anyway. Thus far we've only seen reports for qla2xxx devices as far as I'm aware. > But theres' really only two (trivial) patches in there: > > - scsi: Add intermediate STARGET_REMOVE state to scsi_target_state > (f05795d3d771f30a7bdc3a138bf714b06d42aa95 upstream) > > - Revert "scsi: fix soft lockup in scsi_remove_target() on module removal" > (305c2e71b3d733ec065cb716c76af7d554bd5571 upstream) > > as far as I can tell. And neither of them looks very likely, but what > do I know. Adding Martin Petersen and Johannes Thumshirn to the > participants just in case they go "Ahh.." Right, I had the same head scratching. josh
[toc] | [prev] | [next] | [standalone]
| From | "Martin K. Petersen" <martin.petersen@oracle.com> |
|---|---|
| Date | 2016-06-22 03:30 +0200 |
| Message-ID | <rMF9n-57L-23@gated-at.bofh.it> |
| In reply to | #1428153 |
>>>>> "Linus" == Linus Torvalds <torvalds@linux-foundation.org> writes:
>> https://bugzilla.redhat.com/show_bug.cgi?id=1348342
This first one appears to be a crash in a USB sound doodad and not
qla2xxx. Also, this appears to be where the 4.5.5 -> 4.5.6 notion comes
from. So we can probably ignore 4.5.5 as the last good revision.
Linus> as far as I can tell. And neither of them looks very likely, but
Linus> what do I know. Adding Martin Petersen and Johannes Thumshirn to
Linus> the participants just in case they go "Ahh.."
Doubt it's Johannes' tweak. The qla2xxx crash from the two other
bugzilla entries is in:
(gdb) list *qla24xx_process_response_queue+0x49
0x27e09 is in qla24xx_process_response_queue (drivers/scsi/qla2xxx/qla_isr.c:2560).
2555 if (rsp->msix->cpuid != smp_processor_id()) {
2556 /* if kernel does not notify qla of IRQ's CPU change,
2557 * then set it here.
2558 */
2559 rsp->msix->cpuid = smp_processor_id();
2560 ha->tgt.rspq_vector_cpuid = rsp->msix->cpuid;
2561 }
2562
2563 while (rsp->ring_ptr->signature != RESPONSE_PROCESSED) {
2564 pkt = (struct sts_entry_24xx *)rsp->ring_ptr;
That particular code went into 4.5 and comes from:
commit cdb898c52d1dfad4b4800b83a58b3fe5d352edde
Author: Quinn Tran <quinn.tran@qlogic.com>
Date: Thu Dec 17 14:57:05 2015 -0500
qla2xxx: Add irq affinity notification
Register to receive notification of when irq setting change
occured.
Signed-off-by: Quinn Tran <quinn.tran@qlogic.com>
Signed-off-by: Himanshu Madhani <himanshu.madhani@qlogic.com>
Signed-off-by: Nicholas Bellinger <nab@linux-iscsi.org>
Quinn?
--
Martin K. Petersen Oracle Linux Engineering
[toc] | [prev] | [next] | [standalone]
| From | Quinn Tran <quinn.tran@qlogic.com> |
|---|---|
| Date | 2016-06-22 06:10 +0200 |
| Message-ID | <rMHEd-6TC-3@gated-at.bofh.it> |
| In reply to | #1428320 |
Investigating.
Regards,
Quinn Tran
-----Original Message-----
From: "Martin K. Petersen" <martin.petersen@oracle.com>
Organization: Oracle Corporation
Date: Tuesday, June 21, 2016 at 6:25 PM
To: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Josh Boyer <jwboyer@fedoraproject.org>, "Martin K. Petersen" <martin.petersen@oracle.com>, Johannes Thumshirn <jthumshirn@suse.de>, Thorsten Leemhuis <regressions@leemhuis.info>, linux-kernel <linux-kernel@vger.kernel.org>, Quinn Tran <quinn.tran@qlogic.com>
Subject: Re: Reported regressions for 4.7 as of Sunday, 2016-06-19
>>>>>> "Linus" == Linus Torvalds <torvalds@linux-foundation.org> writes:
>
>>> https://bugzilla.redhat.com/show_bug.cgi?id=1348342
>
>This first one appears to be a crash in a USB sound doodad and not
>qla2xxx. Also, this appears to be where the 4.5.5 -> 4.5.6 notion comes
>from. So we can probably ignore 4.5.5 as the last good revision.
>
>Linus> as far as I can tell. And neither of them looks very likely, but
>Linus> what do I know. Adding Martin Petersen and Johannes Thumshirn to
>Linus> the participants just in case they go "Ahh.."
>
>Doubt it's Johannes' tweak. The qla2xxx crash from the two other
>bugzilla entries is in:
>
>(gdb) list *qla24xx_process_response_queue+0x49
>0x27e09 is in qla24xx_process_response_queue (drivers/scsi/qla2xxx/qla_isr.c:2560).
>2555 if (rsp->msix->cpuid != smp_processor_id()) {
>2556 /* if kernel does not notify qla of IRQ's CPU change,
>2557 * then set it here.
>2558 */
>2559 rsp->msix->cpuid = smp_processor_id();
>2560 ha->tgt.rspq_vector_cpuid = rsp->msix->cpuid;
>2561 }
>2562
>2563 while (rsp->ring_ptr->signature != RESPONSE_PROCESSED) {
>2564 pkt = (struct sts_entry_24xx *)rsp->ring_ptr;
>
>That particular code went into 4.5 and comes from:
>
>commit cdb898c52d1dfad4b4800b83a58b3fe5d352edde
>Author: Quinn Tran <quinn.tran@qlogic.com>
>Date: Thu Dec 17 14:57:05 2015 -0500
>
> qla2xxx: Add irq affinity notification
>
> Register to receive notification of when irq setting change
> occured.
>
> Signed-off-by: Quinn Tran <quinn.tran@qlogic.com>
> Signed-off-by: Himanshu Madhani <himanshu.madhani@qlogic.com>
> Signed-off-by: Nicholas Bellinger <nab@linux-iscsi.org>
>
>Quinn?
>
>--
>Martin K. Petersen Oracle Linux Engineering
[toc] | [prev] | [next] | [standalone]
| From | Johannes Thumshirn <jthumshirn@suse.de> |
|---|---|
| Date | 2016-06-22 14:10 +0200 |
| Message-ID | <rMP8K-3bX-27@gated-at.bofh.it> |
| In reply to | #1428320 |
On Tue, Jun 21, 2016 at 09:25:18PM -0400, Martin K. Petersen wrote:
> >>>>> "Linus" == Linus Torvalds <torvalds@linux-foundation.org> writes:
>
> >> https://bugzilla.redhat.com/show_bug.cgi?id=1348342
>
> This first one appears to be a crash in a USB sound doodad and not
> qla2xxx. Also, this appears to be where the 4.5.5 -> 4.5.6 notion comes
> from. So we can probably ignore 4.5.5 as the last good revision.
>
> Linus> as far as I can tell. And neither of them looks very likely, but
> Linus> what do I know. Adding Martin Petersen and Johannes Thumshirn to
> Linus> the participants just in case they go "Ahh.."
>
> Doubt it's Johannes' tweak. The qla2xxx crash from the two other
> bugzilla entries is in:
>
> (gdb) list *qla24xx_process_response_queue+0x49
> 0x27e09 is in qla24xx_process_response_queue (drivers/scsi/qla2xxx/qla_isr.c:2560).
> 2555 if (rsp->msix->cpuid != smp_processor_id()) {
> 2556 /* if kernel does not notify qla of IRQ's CPU change,
> 2557 * then set it here.
> 2558 */
> 2559 rsp->msix->cpuid = smp_processor_id();
> 2560 ha->tgt.rspq_vector_cpuid = rsp->msix->cpuid;
> 2561 }
> 2562
> 2563 while (rsp->ring_ptr->signature != RESPONSE_PROCESSED) {
> 2564 pkt = (struct sts_entry_24xx *)rsp->ring_ptr;
>
> That particular code went into 4.5 and comes from:
>
> commit cdb898c52d1dfad4b4800b83a58b3fe5d352edde
> Author: Quinn Tran <quinn.tran@qlogic.com>
> Date: Thu Dec 17 14:57:05 2015 -0500
>
> qla2xxx: Add irq affinity notification
>
> Register to receive notification of when irq setting change
> occured.
>
> Signed-off-by: Quinn Tran <quinn.tran@qlogic.com>
> Signed-off-by: Himanshu Madhani <himanshu.madhani@qlogic.com>
> Signed-off-by: Nicholas Bellinger <nab@linux-iscsi.org>
>
> Quinn?
>
> --
> Martin K. Petersen Oracle Linux Engineering
Having a quick look at it I _think_ this could be the problem.
We request the IRQ _before_ actually assigning the rsp->msix entry. Now If an
IRQ triggers, before the assignment we touch rsp->msix->cpuid, which is
probably the case. At least from what I conduct from Martin's mail.
diff --git a/drivers/scsi/qla2xxx/qla_isr.c b/drivers/scsi/qla2xxx/qla_isr.c
index 5649c20..20743a3 100644
--- a/drivers/scsi/qla2xxx/qla_isr.c
+++ b/drivers/scsi/qla2xxx/qla_isr.c
@@ -3086,6 +3086,8 @@ qla24xx_enable_msix(struct qla_hw_data *ha, struct rsp_que *rsp)
/* Enable MSI-X vectors for the base queue */
for (i = 0; i < 2; i++) {
qentry = &ha->msix_entries[i];
+ qentry->rsp = rsp;
+ rsp->msix = qentry;
if (IS_P3P_TYPE(ha))
ret = request_irq(qentry->vector,
qla82xx_msix_entries[i].handler,
@@ -3097,8 +3099,6 @@ qla24xx_enable_msix(struct qla_hw_data *ha, struct rsp_que *rsp)
if (ret)
goto msix_register_fail;
qentry->have_irq = 1;
- qentry->rsp = rsp;
- rsp->msix = qentry;
/* Register for CPU affinity notification. */
irq_set_affinity_notifier(qentry->vector, &qentry->irq_notify);
@@ -3119,12 +3119,12 @@ qla24xx_enable_msix(struct qla_hw_data *ha, struct rsp_que *rsp)
*/
if (QLA_TGT_MODE_ENABLED() && IS_ATIO_MSIX_CAPABLE(ha)) {
qentry = &ha->msix_entries[ATIO_VECTOR];
+ qentry->rsp = rsp;
+ rsp->msix = qentry;
ret = request_irq(qentry->vector,
qla83xx_msix_entries[ATIO_VECTOR].handler,
0, qla83xx_msix_entries[ATIO_VECTOR].name, rsp);
qentry->have_irq = 1;
- qentry->rsp = rsp;
- rsp->msix = qentry;
}
msix_register_fail:
I'm not sure if we need the qentry->have_irq assingment as well, I'm not
deep enough into the qla2xx driver yet, maybe Quinn can clarify.
Beware of the above change being untested.
Byte,
Johannes
--
Johannes Thumshirn Storage
jthumshirn@suse.de +49 911 74053 689
SUSE LINUX GmbH, Maxfeldstr. 5, 90409 Nürnberg
GF: Felix Imendörffer, Jane Smithard, Graham Norton
HRB 21284 (AG Nürnberg)
Key fingerprint = EC38 9CAB C2C4 F25D 8600 D0D0 0393 969D 2D76 0850
[toc] | [prev] | [next] | [standalone]
| From | Quinn Tran <quinn.tran@qlogic.com> |
|---|---|
| Date | 2016-06-22 18:00 +0200 |
| Message-ID | <rMSJk-5mw-23@gated-at.bofh.it> |
| In reply to | #1428729 |
Johannes, Martin,
Based on the screen shot/call trace, it looks like this adapter is not using MSIX. It defaulted back to MSI or INTx interrupt. The code made an assumption of MSIX is available. There is no point in go through that code segment.
Can you try this work around? It’s untested. Thanks.
diff --git a/drivers/scsi/qla2xxx/qla_isr.c b/drivers/scsi/qla2xxx/qla_isr.c
index 5649c20..e033ecb 100644
--- a/drivers/scsi/qla2xxx/qla_isr.c
+++ b/drivers/scsi/qla2xxx/qla_isr.c
@@ -2548,7 +2548,7 @@ void qla24xx_process_response_queue(struct scsi_qla_host *vha,
if (!vha->flags.online)
return;
- if (rsp->msix->cpuid != smp_processor_id()) {
+ if (rsp->msix && (rsp->msix->cpuid != smp_processor_id())) {
/* if kernel does not notify qla of IRQ's CPU change,
* then set it here.
*/
Regards,
Quinn Tran
-----Original Message-----
From: Johannes Thumshirn <jthumshirn@suse.de>
Date: Wednesday, June 22, 2016 at 4:51 AM
To: "Martin K. Petersen" <martin.petersen@oracle.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>, Josh Boyer <jwboyer@fedoraproject.org>, Thorsten Leemhuis <regressions@leemhuis.info>, linux-kernel <linux-kernel@vger.kernel.org>, Quinn Tran <quinn.tran@qlogic.com>
Subject: Re: Reported regressions for 4.7 as of Sunday, 2016-06-19
>On Tue, Jun 21, 2016 at 09:25:18PM -0400, Martin K. Petersen wrote:
>> >>>>> "Linus" == Linus Torvalds <torvalds@linux-foundation.org> writes:
>>
>> >> https://bugzilla.redhat.com/show_bug.cgi?id=1348342
>>
>> This first one appears to be a crash in a USB sound doodad and not
>> qla2xxx. Also, this appears to be where the 4.5.5 -> 4.5.6 notion comes
>> from. So we can probably ignore 4.5.5 as the last good revision.
>>
>> Linus> as far as I can tell. And neither of them looks very likely, but
>> Linus> what do I know. Adding Martin Petersen and Johannes Thumshirn to
>> Linus> the participants just in case they go "Ahh.."
>>
>> Doubt it's Johannes' tweak. The qla2xxx crash from the two other
>> bugzilla entries is in:
>>
>> (gdb) list *qla24xx_process_response_queue+0x49
>> 0x27e09 is in qla24xx_process_response_queue (drivers/scsi/qla2xxx/qla_isr.c:2560).
>> 2555 if (rsp->msix->cpuid != smp_processor_id()) {
>> 2556 /* if kernel does not notify qla of IRQ's CPU change,
>> 2557 * then set it here.
>> 2558 */
>> 2559 rsp->msix->cpuid = smp_processor_id();
>> 2560 ha->tgt.rspq_vector_cpuid = rsp->msix->cpuid;
>> 2561 }
>> 2562
>> 2563 while (rsp->ring_ptr->signature != RESPONSE_PROCESSED) {
>> 2564 pkt = (struct sts_entry_24xx *)rsp->ring_ptr;
>>
>> That particular code went into 4.5 and comes from:
>>
>> commit cdb898c52d1dfad4b4800b83a58b3fe5d352edde
>> Author: Quinn Tran <quinn.tran@qlogic.com>
>> Date: Thu Dec 17 14:57:05 2015 -0500
>>
>> qla2xxx: Add irq affinity notification
>>
>> Register to receive notification of when irq setting change
>> occured.
>>
>> Signed-off-by: Quinn Tran <quinn.tran@qlogic.com>
>> Signed-off-by: Himanshu Madhani <himanshu.madhani@qlogic.com>
>> Signed-off-by: Nicholas Bellinger <nab@linux-iscsi.org>
>>
>> Quinn?
>>
>> --
>> Martin K. Petersen Oracle Linux Engineering
>
>Having a quick look at it I _think_ this could be the problem.
>We request the IRQ _before_ actually assigning the rsp->msix entry. Now If an
>IRQ triggers, before the assignment we touch rsp->msix->cpuid, which is
>probably the case. At least from what I conduct from Martin's mail.
>
>diff --git a/drivers/scsi/qla2xxx/qla_isr.c b/drivers/scsi/qla2xxx/qla_isr.c
>index 5649c20..20743a3 100644
>--- a/drivers/scsi/qla2xxx/qla_isr.c
>+++ b/drivers/scsi/qla2xxx/qla_isr.c
>@@ -3086,6 +3086,8 @@ qla24xx_enable_msix(struct qla_hw_data *ha, struct rsp_que *rsp)
> /* Enable MSI-X vectors for the base queue */
> for (i = 0; i < 2; i++) {
> qentry = &ha->msix_entries[i];
>+ qentry->rsp = rsp;
>+ rsp->msix = qentry;
> if (IS_P3P_TYPE(ha))
> ret = request_irq(qentry->vector,
> qla82xx_msix_entries[i].handler,
>@@ -3097,8 +3099,6 @@ qla24xx_enable_msix(struct qla_hw_data *ha, struct rsp_que *rsp)
> if (ret)
> goto msix_register_fail;
> qentry->have_irq = 1;
>- qentry->rsp = rsp;
>- rsp->msix = qentry;
>
> /* Register for CPU affinity notification. */
> irq_set_affinity_notifier(qentry->vector, &qentry->irq_notify);
>@@ -3119,12 +3119,12 @@ qla24xx_enable_msix(struct qla_hw_data *ha, struct rsp_que *rsp)
> */
> if (QLA_TGT_MODE_ENABLED() && IS_ATIO_MSIX_CAPABLE(ha)) {
> qentry = &ha->msix_entries[ATIO_VECTOR];
>+ qentry->rsp = rsp;
>+ rsp->msix = qentry;
> ret = request_irq(qentry->vector,
> qla83xx_msix_entries[ATIO_VECTOR].handler,
> 0, qla83xx_msix_entries[ATIO_VECTOR].name, rsp);
> qentry->have_irq = 1;
>- qentry->rsp = rsp;
>- rsp->msix = qentry;
> }
>
> msix_register_fail:
>
>
>I'm not sure if we need the qentry->have_irq assingment as well, I'm not
>deep enough into the qla2xx driver yet, maybe Quinn can clarify.
>Beware of the above change being untested.
>
>Byte,
> Johannes
>--
>Johannes Thumshirn Storage
>jthumshirn@suse.de +49 911 74053 689
>SUSE LINUX GmbH, Maxfeldstr. 5, 90409 Nürnberg
>GF: Felix Imendörffer, Jane Smithard, Graham Norton
>HRB 21284 (AG Nürnberg)
>Key fingerprint = EC38 9CAB C2C4 F25D 8600 D0D0 0393 969D 2D76 0850
[toc] | [prev] | [next] | [standalone]
| From | Johannes Thumshirn <jthumshirn@suse.de> |
|---|---|
| Date | 2016-06-23 09:30 +0200 |
| Message-ID | <rN7fj-6FF-5@gated-at.bofh.it> |
| In reply to | #1428922 |
[+ Cc linux-scsi@vger.kernel.org ]
On Wed, Jun 22, 2016 at 03:57:35PM +0000, Quinn Tran wrote:
> Johannes, Martin,
>
> Based on the screen shot/call trace, it looks like this adapter is not using MSIX. It defaulted back to MSI or INTx interrupt. The code made an assumption of MSIX is available. There is no point in go through that code segment.
>
> Can you try this work around? It’s untested. Thanks.
>
>
> diff --git a/drivers/scsi/qla2xxx/qla_isr.c b/drivers/scsi/qla2xxx/qla_isr.c
> index 5649c20..e033ecb 100644
> --- a/drivers/scsi/qla2xxx/qla_isr.c
> +++ b/drivers/scsi/qla2xxx/qla_isr.c
> @@ -2548,7 +2548,7 @@ void qla24xx_process_response_queue(struct scsi_qla_host *vha,
> if (!vha->flags.online)
> return;
>
> - if (rsp->msix->cpuid != smp_processor_id()) {
> + if (rsp->msix && (rsp->msix->cpuid != smp_processor_id())) {
> /* if kernel does not notify qla of IRQ's CPU change,
> * then set it here.
> */
>
But this still does not fix the race which would be possible if the HBA is
using MSI-X but triggering IRQs early enough.
Have a look at this (I admit theoretical) path:
qla24xx_enable_msix(struct qla_hw_data *ha, struct rsp_que *rsp)
{
[...]
/* Enable MSI-X vectors for the base queue */
for (i = 0; i < 2; i++) {
qentry = &ha->msix_entries[i];
if (IS_P3P_TYPE(ha))
ret = request_irq(qentry->vector,
qla82xx_msix_entries[i].handler,
0, qla82xx_msix_entries[i].name, rsp);
else
ret = request_irq(qentry->vector,
msix_entries[i].handler,
0, msix_entries[i].name, rsp);
if (ret)
goto msix_register_fail;
<--- IRQ arrives here
qentry->have_irq = 1;
qentry->rsp = rsp;
rsp->msix = qentry;
[...]
void qla24xx_process_response_queue(struct scsi_qla_host *vha,
struct rsp_que *rsp)
{
[...]
if (rsp->msix->cpuid != smp_processor_id()) {
^
\--- rsp->msix == NULL
/* if kernel does not notify qla of IRQ's CPU change,
* then set it here.
*/
rsp->msix->cpuid = smp_processor_id();
ha->tgt.rspq_vector_cpuid = rsp->msix->cpuid;
--
Johannes Thumshirn Storage
jthumshirn@suse.de +49 911 74053 689
SUSE LINUX GmbH, Maxfeldstr. 5, 90409 Nürnberg
GF: Felix Imendörffer, Jane Smithard, Graham Norton
HRB 21284 (AG Nürnberg)
Key fingerprint = EC38 9CAB C2C4 F25D 8600 D0D0 0393 969D 2D76 0850
[toc] | [prev] | [next] | [standalone]
| From | Quinn Tran <quinn.tran@qlogic.com> |
|---|---|
| Date | 2016-06-23 18:20 +0200 |
| Message-ID | <rNfwe-45d-3@gated-at.bofh.it> |
| In reply to | #1429520 |
-----Original Message-----
From: Johannes Thumshirn <jthumshirn@suse.de>
Date: Thursday, June 23, 2016 at 12:22 AM
To: Quinn Tran <quinn.tran@qlogic.com>
Cc: "Martin K. Petersen" <martin.petersen@oracle.com>, Linus Torvalds <torvalds@linux-foundation.org>, Josh Boyer <jwboyer@fedoraproject.org>, Thorsten Leemhuis <regressions@leemhuis.info>, linux-kernel <linux-kernel@vger.kernel.org>, linux-scsi <linux-scsi@vger.kernel.org>
Subject: Re: Reported regressions for 4.7 as of Sunday, 2016-06-19
>[+ Cc linux-scsi@vger.kernel.org ]
>
>On Wed, Jun 22, 2016 at 03:57:35PM +0000, Quinn Tran wrote:
>> Johannes, Martin,
>>
>> Based on the screen shot/call trace, it looks like this adapter is not using MSIX. It defaulted back to MSI or INTx interrupt. The code made an assumption of MSIX is available. There is no point in go through that code segment.
>>
>> Can you try this work around? It’s untested. Thanks.
>>
>>
>> diff --git a/drivers/scsi/qla2xxx/qla_isr.c b/drivers/scsi/qla2xxx/qla_isr.c
>> index 5649c20..e033ecb 100644
>> --- a/drivers/scsi/qla2xxx/qla_isr.c
>> +++ b/drivers/scsi/qla2xxx/qla_isr.c
>> @@ -2548,7 +2548,7 @@ void qla24xx_process_response_queue(struct scsi_qla_host *vha,
>> if (!vha->flags.online)
>> return;
>>
>> - if (rsp->msix->cpuid != smp_processor_id()) {
>> + if (rsp->msix && (rsp->msix->cpuid != smp_processor_id())) {
>> /* if kernel does not notify qla of IRQ's CPU change,
>> * then set it here.
>> */
>>
>
>But this still does not fix the race which would be possible if the HBA is
>using MSI-X but triggering IRQs early enough.
>
>Have a look at this (I admit theoretical) path:
>qla24xx_enable_msix(struct qla_hw_data *ha, struct rsp_que *rsp)
>{
> [...]
> /* Enable MSI-X vectors for the base queue */
> for (i = 0; i < 2; i++) {
> qentry = &ha->msix_entries[i];
> if (IS_P3P_TYPE(ha))
> ret = request_irq(qentry->vector,
> qla82xx_msix_entries[i].handler,
> 0, qla82xx_msix_entries[i].name, rsp);
> else
> ret = request_irq(qentry->vector,
> msix_entries[i].handler,
> 0, msix_entries[i].name, rsp);
> if (ret)
> goto msix_register_fail;
> <--- IRQ arrives here
QT: setting up the interrupt vector does not mean the interrupt starts firing immediately. Interrupt starting firing when the driver is ready to accept the interrupt by enabling the interrupt (ha->isp_ops->enable_intrs(ha)) later on in time. In addition, that particular code path/qla24xx_process_response_queue is not executed until driver feeds commands to the hardware work queue.
IF there is a left over interrupt that happens to trigger the call immediately, there is another check that prevent the code from getting to the point of the “theoretical" race.
> qentry->have_irq = 1;
> qentry->rsp = rsp;
> rsp->msix = qentry;
>
> [...]
>
>
>void qla24xx_process_response_queue(struct scsi_qla_host *vha,
> struct rsp_que *rsp)
>{
--->8------
if (!vha->flags.online)
return;
---8<------
> if (rsp->msix->cpuid != smp_processor_id()) {
> ^
> \--- rsp->msix == NULL
>
> /* if kernel does not notify qla of IRQ's CPU change,
> * then set it here.
> */
> rsp->msix->cpuid = smp_processor_id();
> ha->tgt.rspq_vector_cpuid = rsp->msix->cpuid;
>
>--
>Johannes Thumshirn Storage
>jthumshirn@suse.de +49 911 74053 689
>SUSE LINUX GmbH, Maxfeldstr. 5, 90409 Nürnberg
>GF: Felix Imendörffer, Jane Smithard, Graham Norton
>HRB 21284 (AG Nürnberg)
>Key fingerprint = EC38 9CAB C2C4 F25D 8600 D0D0 0393 969D 2D76 0850
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2016-06-23 18:40 +0200 |
| Message-ID | <rNfPA-4cB-29@gated-at.bofh.it> |
| In reply to | #1429955 |
On Thu, Jun 23, 2016 at 9:13 AM, Quinn Tran <quinn.tran@qlogic.com> wrote:
>
>
> QT: setting up the interrupt vector does not mean the interrupt starts firing immediately.
Actually, it very much can mean that. If the interrupt can possibly be
shared, there is a very real possibility of it fiding immediately.
Now, with MSI(-X) I guess that isn't a worry, so I suspect your patch
that handles just the legacy INTx case anyway is sufficient, but in
general I would like people to always act as if interrupts can happen
immediately after request_irq().
We have had *tons* of situations where the firmware left a device
active, for example. Or where some random interrupt controller ended
up having stale interrupts pending, even.
So in general, it's just good practice to say "spurious interrupts can
and do happen" - the shared irq case is the most obvious case, but
there have been other sources of unexpected spurious interrupts
firing.
Linus
[toc] | [prev] | [next] | [standalone]
| From | Kalle Valo <kvalo@codeaurora.org> |
|---|---|
| Date | 2016-06-22 08:40 +0200 |
| Message-ID | <rMJZn-8iJ-3@gated-at.bofh.it> |
| In reply to | #1425986 |
Thorsten Leemhuis <regressions@leemhuis.info> writes: > Description: ath10k no longer authenticates and freezes system > Report: https://bugzilla.kernel.org/show_bug.cgi?id=119151 > Latest status: http://thread.gmane.org/gmane.linux.kernel.wireless.general/152513/focus=152535 > Date rep/stat: 2016-05-27 / 2016-06-02 > Notes: forgotten? poked bug report on Friday Here's the fix: ath10k: fix deadlock while processing rx_in_ord_ind https://git.kernel.org/cgit/linux/kernel/git/kvalo/wireless-drivers.git/commit/?id=e50525bef593c3dd0564df676c567d77f7c20322 -- Kalle Valo
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web