Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1716009 > unrolled thread
| Started by | Jeffrey Hugo <jhugo@codeaurora.org> |
|---|---|
| First post | 2017-08-20 21:40 +0200 |
| Last post | 2017-08-22 23:00 +0200 |
| Articles | 4 — 3 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [BUG] Deadlock due due to interactions of block, RCU, and cpu offline Jeffrey Hugo <jhugo@codeaurora.org> - 2017-08-20 21:40 +0200
Re: [BUG] Deadlock due due to interactions of block, RCU, and cpu offline "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-08-20 23:00 +0200
Re: [BUG] Deadlock due due to interactions of block, RCU, and cpu offline Paolo Bonzini <pbonzini@redhat.com> - 2017-08-22 18:20 +0200
Re: [BUG] Deadlock due due to interactions of block, RCU, and cpu offline Jeffrey Hugo <jhugo@codeaurora.org> - 2017-08-22 23:00 +0200
| From | Jeffrey Hugo <jhugo@codeaurora.org> |
|---|---|
| Date | 2017-08-20 21:40 +0200 |
| Subject | Re: [BUG] Deadlock due due to interactions of block, RCU, and cpu offline |
| Message-ID | <ugEeJ-2HE-1@gated-at.bofh.it> |
On 6/29/2017 6:18 PM, Paul E. McKenney wrote:
> On Thu, Jun 29, 2017 at 10:29:12AM -0600, Jeffrey Hugo wrote:
>> On 6/27/2017 6:11 PM, Paul E. McKenney wrote:
>>> On Tue, Jun 27, 2017 at 04:32:09PM -0600, Jeffrey Hugo wrote:
>>>> On 6/22/2017 9:34 PM, Paul E. McKenney wrote:
>>>>> On Wed, Jun 21, 2017 at 09:18:53AM -0700, Paul E. McKenney wrote:
>>>>>> No worries, and I am very much looking forward to seeing the results of
>>>>>> your testing.
>>>>>
>>>>> And please see below for an updated patch based on LKML review and
>>>>> more intensive testing.
>>>>>
>>>>
>>>> I spent some time on this today. It didn't go as I expected. I
>>>> validated the issue is reproducible as before on 4.11 and 4.12 rcs 1
>>>> through 4. However, the version of stress-ng that I was using ran
>>>> into constant errors starting with rc5, making it nearly impossible
>>>> to make progress toward reproduction. Upgrading stress-ng to tip
>>>> fixes the issue, however, I've still been unable to repro the issue.
>>>>
>>>> Its my unfounded suspicion that something went in between rc4 and
>>>> rc5 which changed the timing, and didn't actually fix the issue. I
>>>> will run the test overnight for 5 hours to try to repro.
>>>>
>>>> The patch you sent appears to be based on linux-next, and appears to
>>>> have a number of dependencies which prevent it from cleanly applying
>>>> on anything current that I'm able to repro on at this time. Do you
>>>> want to provide a rebased version of the patch which applies to say
>>>> 4.11? I could easily test that and report back.
>>>
>>> Here is a very lightly tested backport to v4.11.
>>>
>>
>> Works for me. Always reproduced the lockup within 2 minutes on stock
>> 4.11. With the change applied, I was able to test for 2 hours in
>> the same conditions, and 4 hours with the full system and not
>> encounter an issue.
>>
>> Feel free to add:
>> Tested-by: Jeffrey Hugo <jhugo@codeaurora.org>
>
> Applied, thank you!
>
>> I'm going to go back to 4.12-rc5 and see if I can get either repro
>> the issue, or identify what changed. Hopefully I can get to
>> linux-next and double check the original version of the change as
>> well.
>
> Looking forward to hearing what you find!
>
> Thanx, Paul
>
According to git bisect, the following is what "changed"
commit 9d0eb4624601ac978b9e89be4aeadbd51ab2c830
Merge: 5faab9e 9bc1f09
Author: Linus Torvalds <torvalds@linux-foundation.org>
Date: Sun Jun 11 11:07:25 2017 -0700
Merge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm
Pull KVM fixes from Paolo Bonzini:
"Bug fixes (ARM, s390, x86)"
* tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm:
KVM: async_pf: avoid async pf injection when in guest mode
KVM: cpuid: Fix read/write out-of-bounds vulnerability in cpuid
emulation
arm: KVM: Allow unaligned accesses at HYP
arm64: KVM: Allow unaligned accesses at EL2
arm64: KVM: Preserve RES1 bits in SCTLR_EL2
KVM: arm/arm64: Handle possible NULL stage2 pud when ageing pages
KVM: nVMX: Fix exception injection
kvm: async_pf: fix rcu_irq_enter() with irqs enabled
KVM: arm/arm64: vgic-v3: Fix nr_pre_bits bitfield extraction
KVM: s390: fix ais handling vs cpu model
KVM: arm/arm64: Fix isues with GICv2 on GICv3 migration
Nothing really stands out to me which would "fix" the issue.
--
Jeffrey Hugo
Qualcomm Datacenter Technologies as an affiliate of Qualcomm
Technologies, Inc.
Qualcomm Technologies, Inc. is a member of the
Code Aurora Forum, a Linux Foundation Collaborative Project.
[toc] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-08-20 23:00 +0200 |
| Message-ID | <ugFua-3mz-7@gated-at.bofh.it> |
| In reply to | #1716009 |
On Sun, Aug 20, 2017 at 01:31:01PM -0600, Jeffrey Hugo wrote: > On 6/29/2017 6:18 PM, Paul E. McKenney wrote: > >On Thu, Jun 29, 2017 at 10:29:12AM -0600, Jeffrey Hugo wrote: > >>On 6/27/2017 6:11 PM, Paul E. McKenney wrote: > >>>On Tue, Jun 27, 2017 at 04:32:09PM -0600, Jeffrey Hugo wrote: > >>>>On 6/22/2017 9:34 PM, Paul E. McKenney wrote: > >>>>>On Wed, Jun 21, 2017 at 09:18:53AM -0700, Paul E. McKenney wrote: > >>>>>>No worries, and I am very much looking forward to seeing the results of > >>>>>>your testing. > >>>>> > >>>>>And please see below for an updated patch based on LKML review and > >>>>>more intensive testing. > >>>>> > >>>> > >>>>I spent some time on this today. It didn't go as I expected. I > >>>>validated the issue is reproducible as before on 4.11 and 4.12 rcs 1 > >>>>through 4. However, the version of stress-ng that I was using ran > >>>>into constant errors starting with rc5, making it nearly impossible > >>>>to make progress toward reproduction. Upgrading stress-ng to tip > >>>>fixes the issue, however, I've still been unable to repro the issue. > >>>> > >>>>Its my unfounded suspicion that something went in between rc4 and > >>>>rc5 which changed the timing, and didn't actually fix the issue. I > >>>>will run the test overnight for 5 hours to try to repro. > >>>> > >>>>The patch you sent appears to be based on linux-next, and appears to > >>>>have a number of dependencies which prevent it from cleanly applying > >>>>on anything current that I'm able to repro on at this time. Do you > >>>>want to provide a rebased version of the patch which applies to say > >>>>4.11? I could easily test that and report back. > >>> > >>>Here is a very lightly tested backport to v4.11. > >>> > >> > >>Works for me. Always reproduced the lockup within 2 minutes on stock > >>4.11. With the change applied, I was able to test for 2 hours in > >>the same conditions, and 4 hours with the full system and not > >>encounter an issue. > >> > >>Feel free to add: > >>Tested-by: Jeffrey Hugo <jhugo@codeaurora.org> > > > >Applied, thank you! > > > >>I'm going to go back to 4.12-rc5 and see if I can get either repro > >>the issue, or identify what changed. Hopefully I can get to > >>linux-next and double check the original version of the change as > >>well. > > > >Looking forward to hearing what you find! > > > > Thanx, Paul > > > > According to git bisect, the following is what "changed" > > commit 9d0eb4624601ac978b9e89be4aeadbd51ab2c830 > Merge: 5faab9e 9bc1f09 > Author: Linus Torvalds <torvalds@linux-foundation.org> > Date: Sun Jun 11 11:07:25 2017 -0700 > > Merge tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm > > Pull KVM fixes from Paolo Bonzini: > "Bug fixes (ARM, s390, x86)" > > * tag 'for-linus' of git://git.kernel.org/pub/scm/virt/kvm/kvm: > KVM: async_pf: avoid async pf injection when in guest mode > KVM: cpuid: Fix read/write out-of-bounds vulnerability in > cpuid emulation > arm: KVM: Allow unaligned accesses at HYP > arm64: KVM: Allow unaligned accesses at EL2 > arm64: KVM: Preserve RES1 bits in SCTLR_EL2 > KVM: arm/arm64: Handle possible NULL stage2 pud when ageing pages > KVM: nVMX: Fix exception injection > kvm: async_pf: fix rcu_irq_enter() with irqs enabled > KVM: arm/arm64: vgic-v3: Fix nr_pre_bits bitfield extraction > KVM: s390: fix ais handling vs cpu model > KVM: arm/arm64: Fix isues with GICv2 on GICv3 migration > > Nothing really stands out to me which would "fix" the issue. My guess would be an undo of the change that provoked the problem in the first place. Did you try bisecting within the above group of commits? Either way, CCing Paolo for his thoughts? Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Paolo Bonzini <pbonzini@redhat.com> |
|---|---|
| Date | 2017-08-22 18:20 +0200 |
| Message-ID | <uhk4h-4NV-1@gated-at.bofh.it> |
| In reply to | #1716021 |
On 20/08/2017 22:56, Paul E. McKenney wrote: >> KVM: async_pf: avoid async pf injection when in guest mode >> KVM: cpuid: Fix read/write out-of-bounds vulnerability in cpuid emulation >> arm: KVM: Allow unaligned accesses at HYP >> arm64: KVM: Allow unaligned accesses at EL2 >> arm64: KVM: Preserve RES1 bits in SCTLR_EL2 >> KVM: arm/arm64: Handle possible NULL stage2 pud when ageing pages >> KVM: nVMX: Fix exception injection >> kvm: async_pf: fix rcu_irq_enter() with irqs enabled >> KVM: arm/arm64: vgic-v3: Fix nr_pre_bits bitfield extraction >> KVM: s390: fix ais handling vs cpu model >> KVM: arm/arm64: Fix isues with GICv2 on GICv3 migration >> >> Nothing really stands out to me which would "fix" the issue. > > My guess would be an undo of the change that provoked the problem > in the first place. Did you try bisecting within the above group > of commits? > > Either way, CCing Paolo for his thoughts? There is "kvm: async_pf: fix rcu_irq_enter() with irqs enabled", but it would have caused splats, not deadlocks. If you are using nested virtualization, "KVM: async_pf: avoid async pf injection when in guest mode" can be a wildcard, but only if you have memory pressure. My bet is still on the former changing the timing just a little bit. Paolo
[toc] | [prev] | [next] | [standalone]
| From | Jeffrey Hugo <jhugo@codeaurora.org> |
|---|---|
| Date | 2017-08-22 23:00 +0200 |
| Message-ID | <uhorg-7Cb-15@gated-at.bofh.it> |
| In reply to | #1717552 |
On 8/22/2017 10:12 AM, Paolo Bonzini wrote:
> On 20/08/2017 22:56, Paul E. McKenney wrote:
>>> KVM: async_pf: avoid async pf injection when in guest mode
>>> KVM: cpuid: Fix read/write out-of-bounds vulnerability in cpuid emulation
>>> arm: KVM: Allow unaligned accesses at HYP
>>> arm64: KVM: Allow unaligned accesses at EL2
>>> arm64: KVM: Preserve RES1 bits in SCTLR_EL2
>>> KVM: arm/arm64: Handle possible NULL stage2 pud when ageing pages
>>> KVM: nVMX: Fix exception injection
>>> kvm: async_pf: fix rcu_irq_enter() with irqs enabled
>>> KVM: arm/arm64: vgic-v3: Fix nr_pre_bits bitfield extraction
>>> KVM: s390: fix ais handling vs cpu model
>>> KVM: arm/arm64: Fix isues with GICv2 on GICv3 migration
>>>
>>> Nothing really stands out to me which would "fix" the issue.
>>
>> My guess would be an undo of the change that provoked the problem
>> in the first place. Did you try bisecting within the above group
>> of commits?
>>
>> Either way, CCing Paolo for his thoughts?
>
> There is "kvm: async_pf: fix rcu_irq_enter() with irqs enabled", but it
> would have caused splats, not deadlocks.
>
> If you are using nested virtualization, "KVM: async_pf: avoid async pf
> injection when in guest mode" can be a wildcard, but only if you have
> memory pressure.
>
> My bet is still on the former changing the timing just a little bit.
>
> Paolo
>
I'm sorry, I must have done the bisect incorrectly.
I attempted to bisect the KVM changes from the merge, but was seeing
that the issue didn't repro with any of them. I double checked the
merge commit, and found it did not introduce a "fix".
I redid the bisect, and it identified the following change this time. I
double checked that reverting the change reintroduces the deadlock, and
cherry-picking the change onto 4.12-rc4 (known to exhibit the issue)
causes the issue to disappear. I'm pretty sure (knock on wood) that the
bisect result is actually correct this time.
commit 6460495709aeb651896bc8e5c134b2e4ca7d34a8
Author: James Wang <jnwang@suse.com>
Date: Thu Jun 8 14:52:51 2017 +0800
Fix loop device flush before configure v3
While installing SLES-12 (based on v4.4), I found that the installer
will stall for 60+ seconds during LVM disk scan. The root cause was
determined to be the removal of a bound device check in loop_flush()
by commit b5dd2f6047ca ("block: loop: improve performance via blk-mq").
Restoring this check, examining ->lo_state as set by loop_set_fd()
eliminates the bad behavior.
Test method:
modprobe loop max_loop=64
dd if=/dev/zero of=disk bs=512 count=200K
for((i=0;i<4;i++))do losetup -f disk; done
mkfs.ext4 -F /dev/loop0
for((i=0;i<4;i++))do mkdir t$i; mount /dev/loop$i t$i;done
for f in `ls /dev/loop[0-9]*|sort`; do \
echo $f; dd if=$f of=/dev/null bs=512 count=1; \
done
Test output: stock patched
/dev/loop0 18.1217e-05 8.3842e-05
/dev/loop1 6.1114e-05 0.000147979
/dev/loop10 0.414701 0.000116564
/dev/loop11 0.7474 6.7942e-05
/dev/loop12 0.747986 8.9082e-05
/dev/loop13 0.746532 7.4799e-05
/dev/loop14 0.480041 9.3926e-05
/dev/loop15 1.26453 7.2522e-05
Note that from loop10 onward, the device is not mounted, yet the
stock kernel consumes several orders of magnitude more wall time
than it does for a mounted device.
(Thanks for Mike Galbraith <efault@gmx.de>, give a changelog review.)
Reviewed-by: Hannes Reinecke <hare@suse.com>
Reviewed-by: Ming Lei <ming.lei@redhat.com>
Signed-off-by: James Wang <jnwang@suse.com>
Fixes: b5dd2f6047ca ("block: loop: improve performance via blk-mq")
Signed-off-by: Jens Axboe <axboe@fb.com>
Considering the original analysis of the issue, it seems plausible that
this change could be fixing it.
--
Jeffrey Hugo
Qualcomm Datacenter Technologies as an affiliate of Qualcomm
Technologies, Inc.
Qualcomm Technologies, Inc. is a member of the
Code Aurora Forum, a Linux Foundation Collaborative Project.
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web