Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1716036 > unrolled thread

kvm splat in mmu_spte_clear_track_bits

Started byAdam Borowski <kilobyte@angband.pl>
First post2017-08-21 01:20 +0200
Last post2017-08-24 09:50 +0200
Articles 7 — 4 participants

Back to article view | Back to linux.kernel


Contents

  kvm splat in mmu_spte_clear_track_bits Adam Borowski <kilobyte@angband.pl> - 2017-08-21 01:20 +0200
    Re: kvm splat in mmu_spte_clear_track_bits Wanpeng Li <kernellwp@gmail.com> - 2017-08-21 03:30 +0200
      Re: kvm splat in mmu_spte_clear_track_bits Adam Borowski <kilobyte@angband.pl> - 2017-08-21 21:20 +0200
        Re: kvm splat in mmu_spte_clear_track_bits Radim Krčmář <rkrcmar@redhat.com> - 2017-08-21 22:00 +0200
          Re: kvm splat in mmu_spte_clear_track_bits Adam Borowski <kilobyte@angband.pl> - 2017-08-22 00:40 +0200
            Re: kvm splat in mmu_spte_clear_track_bits Paolo Bonzini <pbonzini@redhat.com> - 2017-08-23 14:30 +0200
              Re: kvm splat in mmu_spte_clear_track_bits Wanpeng Li <kernellwp@gmail.com> - 2017-08-24 09:50 +0200

#1716036 — kvm splat in mmu_spte_clear_track_bits

FromAdam Borowski <kilobyte@angband.pl>
Date2017-08-21 01:20 +0200
Subjectkvm splat in mmu_spte_clear_track_bits
Message-ID<ugHFE-4Td-9@gated-at.bofh.it>
Hi!
I'm afraid I keep getting a quite reliable, but random, splat when running
KVM:

------------[ cut here ]------------
WARNING: CPU: 5 PID: 5826 at arch/x86/kvm/mmu.c:717 mmu_spte_clear_track_bits+0x123/0x170
Modules linked in: tun nbd arc4 rtl8xxxu mac80211 cfg80211 rfkill nouveau video ttm
CPU: 5 PID: 5826 Comm: qemu-system-x86 Not tainted 4.13.0-rc5-vanilla-ubsan-00211-g7f680d7ec315 #1
Hardware name: System manufacturer System Product Name/M4A77T, BIOS 2401    05/18/2011
task: ffff880207ef0400 task.stack: ffffc900035e4000
RIP: 0010:mmu_spte_clear_track_bits+0x123/0x170
RSP: 0018:ffffc900035e7ab0 EFLAGS: 00010246
RAX: 0000000000000000 RBX: 000000010501cc67 RCX: 0000000000000001
RDX: dead0000000000ff RSI: ffff88020e501df8 RDI: 0000000004140700
RBP: ffffc900035e7ad8 R08: 0000000000000100 R09: 0000000000000003
R10: 0000000000000003 R11: 0000000000000005 R12: 000000000010501c
R13: ffffea0004140700 R14: ffff88020e1d0000 R15: 0000000000000000
FS:  00007f0213fbd700(0000) GS:ffff88022fd40000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 0000000000000000 CR3: 000000022187f000 CR4: 00000000000006e0
Call Trace:
 drop_spte+0x26/0x130
 mmu_page_zap_pte+0xc4/0x160
 kvm_mmu_prepare_zap_page+0x65/0x660
 kvm_mmu_invalidate_zap_all_pages+0xc5/0x1f0
 kvm_mmu_invalidate_zap_pages_in_memslot+0x9/0x10
 kvm_page_track_flush_slot+0x86/0xd0
 kvm_arch_flush_shadow_memslot+0x9/0x10
 __kvm_set_memory_region+0x8fb/0x14f0
 kvm_set_memory_region+0x2f/0x50
 kvm_vm_ioctl+0x559/0xcc0
 ? kvm_vcpu_ioctl+0x171/0x620
 ? __switch_to+0x30b/0x740
 do_vfs_ioctl+0xbb/0x8d0
 ? find_vma+0x23/0x100
 ? __fget_light+0x94/0x110
 SyS_ioctl+0x86/0xa0
 entry_SYSCALL_64_fastpath+0x17/0x98
RIP: 0033:0x7f021c80ddc7
RSP: 002b:00007f0213fbc518 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
RAX: ffffffffffffffda RBX: 0000000000000000 RCX: 00007f021c80ddc7
RDX: 00007f0213fbc5b0 RSI: 000000004020ae46 RDI: 000000000000000a
RBP: 0000000000000000 R08: 00007f020c1698a0 R09: 0000000000000000
R10: 00007f020c1698a0 R11: 0000000000000246 R12: 0000000000000006
R13: 00007f022201c000 R14: 0000000000000002 R15: 0000558c3899e550
Code: ae fc 01 48 85 c0 75 1c 4c 89 e7 e8 98 de fd ff 48 8b 05 81 ae fc 01 48 85 c0 74 ba 48 85 c3 0f 95 c3 eb b8 48 85 c3 74 e7 eb dd <0f> ff eb 97 4c 89 e7 66 0f 1f 44 00 00 e8 6b de fd ff eb 97 31 
---[ end trace 16c196134f0dd0a9 ]---

After this, there are hundreds of repeats and lots of secondary damage which
kills the host quickly.

Usually this happens within a few minutes, but sometimes it takes ~half an
hour to reproduce.  Because of this, it'd be unpleasant to bisect -- is this
problem already known?


Meow!
-- 
⢀⣴⠾⠻⢶⣦⠀ 
⣾⠁⢰⠒⠀⣿⡁ Vat kind uf sufficiently advanced technology iz dis!?
⢿⡄⠘⠷⠚⠋⠀                                        -- Genghis Ht'rok'din
⠈⠳⣄⠀⠀⠀⠀ 

[toc] | [next] | [standalone]


#1716054

FromWanpeng Li <kernellwp@gmail.com>
Date2017-08-21 03:30 +0200
Message-ID<ugJHs-66r-3@gated-at.bofh.it>
In reply to#1716036
2017-08-21 7:13 GMT+08:00 Adam Borowski <kilobyte@angband.pl>:
> Hi!
> I'm afraid I keep getting a quite reliable, but random, splat when running
> KVM:

I reported something similar before. https://lkml.org/lkml/2017/6/29/64

Regards,
Wanpeng Li

>
> ------------[ cut here ]------------
> WARNING: CPU: 5 PID: 5826 at arch/x86/kvm/mmu.c:717 mmu_spte_clear_track_bits+0x123/0x170
> Modules linked in: tun nbd arc4 rtl8xxxu mac80211 cfg80211 rfkill nouveau video ttm
> CPU: 5 PID: 5826 Comm: qemu-system-x86 Not tainted 4.13.0-rc5-vanilla-ubsan-00211-g7f680d7ec315 #1
> Hardware name: System manufacturer System Product Name/M4A77T, BIOS 2401    05/18/2011
> task: ffff880207ef0400 task.stack: ffffc900035e4000
> RIP: 0010:mmu_spte_clear_track_bits+0x123/0x170
> RSP: 0018:ffffc900035e7ab0 EFLAGS: 00010246
> RAX: 0000000000000000 RBX: 000000010501cc67 RCX: 0000000000000001
> RDX: dead0000000000ff RSI: ffff88020e501df8 RDI: 0000000004140700
> RBP: ffffc900035e7ad8 R08: 0000000000000100 R09: 0000000000000003
> R10: 0000000000000003 R11: 0000000000000005 R12: 000000000010501c
> R13: ffffea0004140700 R14: ffff88020e1d0000 R15: 0000000000000000
> FS:  00007f0213fbd700(0000) GS:ffff88022fd40000(0000) knlGS:0000000000000000
> CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> CR2: 0000000000000000 CR3: 000000022187f000 CR4: 00000000000006e0
> Call Trace:
>  drop_spte+0x26/0x130
>  mmu_page_zap_pte+0xc4/0x160
>  kvm_mmu_prepare_zap_page+0x65/0x660
>  kvm_mmu_invalidate_zap_all_pages+0xc5/0x1f0
>  kvm_mmu_invalidate_zap_pages_in_memslot+0x9/0x10
>  kvm_page_track_flush_slot+0x86/0xd0
>  kvm_arch_flush_shadow_memslot+0x9/0x10
>  __kvm_set_memory_region+0x8fb/0x14f0
>  kvm_set_memory_region+0x2f/0x50
>  kvm_vm_ioctl+0x559/0xcc0
>  ? kvm_vcpu_ioctl+0x171/0x620
>  ? __switch_to+0x30b/0x740
>  do_vfs_ioctl+0xbb/0x8d0
>  ? find_vma+0x23/0x100
>  ? __fget_light+0x94/0x110
>  SyS_ioctl+0x86/0xa0
>  entry_SYSCALL_64_fastpath+0x17/0x98
> RIP: 0033:0x7f021c80ddc7
> RSP: 002b:00007f0213fbc518 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
> RAX: ffffffffffffffda RBX: 0000000000000000 RCX: 00007f021c80ddc7
> RDX: 00007f0213fbc5b0 RSI: 000000004020ae46 RDI: 000000000000000a
> RBP: 0000000000000000 R08: 00007f020c1698a0 R09: 0000000000000000
> R10: 00007f020c1698a0 R11: 0000000000000246 R12: 0000000000000006
> R13: 00007f022201c000 R14: 0000000000000002 R15: 0000558c3899e550
> Code: ae fc 01 48 85 c0 75 1c 4c 89 e7 e8 98 de fd ff 48 8b 05 81 ae fc 01 48 85 c0 74 ba 48 85 c3 0f 95 c3 eb b8 48 85 c3 74 e7 eb dd <0f> ff eb 97 4c 89 e7 66 0f 1f 44 00 00 e8 6b de fd ff eb 97 31
> ---[ end trace 16c196134f0dd0a9 ]---
>
> After this, there are hundreds of repeats and lots of secondary damage which
> kills the host quickly.
>
> Usually this happens within a few minutes, but sometimes it takes ~half an
> hour to reproduce.  Because of this, it'd be unpleasant to bisect -- is this
> problem already known?
>
>
> Meow!
> --
> ⢀⣴⠾⠻⢶⣦⠀
> ⣾⠁⢰⠒⠀⣿⡁ Vat kind uf sufficiently advanced technology iz dis!?
> ⢿⡄⠘⠷⠚⠋⠀                                        -- Genghis Ht'rok'din
> ⠈⠳⣄⠀⠀⠀⠀

[toc] | [prev] | [next] | [standalone]


#1716792

FromAdam Borowski <kilobyte@angband.pl>
Date2017-08-21 21:20 +0200
Message-ID<uh0oW-8mG-21@gated-at.bofh.it>
In reply to#1716054
On Mon, Aug 21, 2017 at 09:26:57AM +0800, Wanpeng Li wrote:
> 2017-08-21 7:13 GMT+08:00 Adam Borowski <kilobyte@angband.pl>:
> > Hi!
> > I'm afraid I keep getting a quite reliable, but random, splat when running
> > KVM:
> 
> I reported something similar before. https://lkml.org/lkml/2017/6/29/64

Your problem seems to require OOM; I don't have any memory pressure at all:
running a single 2GB guest while there's nothing big on the host (bloatfox,
xfce, xorg, terminals + some minor junk); 8GB + (untouched) swap.  There's
no memory pressure inside the guest either -- none was Linux (I wanted to
test something on hurd, kfreebsd) and I doubt they even got to use all of
their frames.

Also, it doesn't reproduce for me on 4.12.

> > ------------[ cut here ]------------
> > WARNING: CPU: 5 PID: 5826 at arch/x86/kvm/mmu.c:717 mmu_spte_clear_track_bits+0x123/0x170
> > Modules linked in: tun nbd arc4 rtl8xxxu mac80211 cfg80211 rfkill nouveau video ttm
> > CPU: 5 PID: 5826 Comm: qemu-system-x86 Not tainted 4.13.0-rc5-vanilla-ubsan-00211-g7f680d7ec315 #1
> > Hardware name: System manufacturer System Product Name/M4A77T, BIOS 2401    05/18/2011
> > task: ffff880207ef0400 task.stack: ffffc900035e4000
> > RIP: 0010:mmu_spte_clear_track_bits+0x123/0x170
> > RSP: 0018:ffffc900035e7ab0 EFLAGS: 00010246
> > RAX: 0000000000000000 RBX: 000000010501cc67 RCX: 0000000000000001
> > RDX: dead0000000000ff RSI: ffff88020e501df8 RDI: 0000000004140700
> > RBP: ffffc900035e7ad8 R08: 0000000000000100 R09: 0000000000000003
> > R10: 0000000000000003 R11: 0000000000000005 R12: 000000000010501c
> > R13: ffffea0004140700 R14: ffff88020e1d0000 R15: 0000000000000000
> > FS:  00007f0213fbd700(0000) GS:ffff88022fd40000(0000) knlGS:0000000000000000
> > CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> > CR2: 0000000000000000 CR3: 000000022187f000 CR4: 00000000000006e0
> > Call Trace:
> >  drop_spte+0x26/0x130
> >  mmu_page_zap_pte+0xc4/0x160
> >  kvm_mmu_prepare_zap_page+0x65/0x660
> >  kvm_mmu_invalidate_zap_all_pages+0xc5/0x1f0
> >  kvm_mmu_invalidate_zap_pages_in_memslot+0x9/0x10
> >  kvm_page_track_flush_slot+0x86/0xd0
> >  kvm_arch_flush_shadow_memslot+0x9/0x10
> >  __kvm_set_memory_region+0x8fb/0x14f0
> >  kvm_set_memory_region+0x2f/0x50
> >  kvm_vm_ioctl+0x559/0xcc0
> >  ? kvm_vcpu_ioctl+0x171/0x620
> >  ? __switch_to+0x30b/0x740
> >  do_vfs_ioctl+0xbb/0x8d0
> >  ? find_vma+0x23/0x100
> >  ? __fget_light+0x94/0x110
> >  SyS_ioctl+0x86/0xa0
> >  entry_SYSCALL_64_fastpath+0x17/0x98
> > RIP: 0033:0x7f021c80ddc7
> > RSP: 002b:00007f0213fbc518 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
> > RAX: ffffffffffffffda RBX: 0000000000000000 RCX: 00007f021c80ddc7
> > RDX: 00007f0213fbc5b0 RSI: 000000004020ae46 RDI: 000000000000000a
> > RBP: 0000000000000000 R08: 00007f020c1698a0 R09: 0000000000000000
> > R10: 00007f020c1698a0 R11: 0000000000000246 R12: 0000000000000006
> > R13: 00007f022201c000 R14: 0000000000000002 R15: 0000558c3899e550
> > Code: ae fc 01 48 85 c0 75 1c 4c 89 e7 e8 98 de fd ff 48 8b 05 81 ae fc 01 48 85 c0 74 ba 48 85 c3 0f 95 c3 eb b8 48 85 c3 74 e7 eb dd <0f> ff eb 97 4c 89 e7 66 0f 1f 44 00 00 e8 6b de fd ff eb 97 31
> > ---[ end trace 16c196134f0dd0a9 ]---
> >
> > After this, there are hundreds of repeats and lots of secondary damage which
> > kills the host quickly.
> >
> > Usually this happens within a few minutes, but sometimes it takes ~half an
> > hour to reproduce.  Because of this, it'd be unpleasant to bisect -- is this
> > problem already known?

-- 
⢀⣴⠾⠻⢶⣦⠀ 
⣾⠁⢰⠒⠀⣿⡁ Vat kind uf sufficiently advanced technology iz dis!?
⢿⡄⠘⠷⠚⠋⠀                                        -- Genghis Ht'rok'din
⠈⠳⣄⠀⠀⠀⠀ 

[toc] | [prev] | [next] | [standalone]


#1716824

FromRadim Krčmář <rkrcmar@redhat.com>
Date2017-08-21 22:00 +0200
Message-ID<uh11E-bc-19@gated-at.bofh.it>
In reply to#1716792
2017-08-21 21:12+0200, Adam Borowski:
> On Mon, Aug 21, 2017 at 09:26:57AM +0800, Wanpeng Li wrote:
> > 2017-08-21 7:13 GMT+08:00 Adam Borowski <kilobyte@angband.pl>:
> > > Hi!
> > > I'm afraid I keep getting a quite reliable, but random, splat when running
> > > KVM:
> > 
> > I reported something similar before. https://lkml.org/lkml/2017/6/29/64
> 
> Your problem seems to require OOM; I don't have any memory pressure at all:
> running a single 2GB guest while there's nothing big on the host (bloatfox,
> xfce, xorg, terminals + some minor junk); 8GB + (untouched) swap.  There's
> no memory pressure inside the guest either -- none was Linux (I wanted to
> test something on hurd, kfreebsd) and I doubt they even got to use all of
> their frames.

I even tried hurd, but couldn't reproduce ... what is your qemu command
line and the output of host's `grep . /sys/module/kvm*/parameters/*`?

> Also, it doesn't reproduce for me on 4.12.

Great info ... the most suspicious between v4.12 and v4.13-rc5 is the
series with dcdca5fed5f6 ("x86: kvm: mmu: make spte mmio mask more
explicit"), does reverting it help?

`git revert ce00053b1cfca312c22e2a6465451f1862561eab~1..995f00a619584e65e53eff372d9b73b121a7bad5`

Thanks.

[toc] | [prev] | [next] | [standalone]


#1716924

FromAdam Borowski <kilobyte@angband.pl>
Date2017-08-22 00:40 +0200
Message-ID<uh3wt-1Pe-9@gated-at.bofh.it>
In reply to#1716824
On Mon, Aug 21, 2017 at 09:58:34PM +0200, Radim Krčmář wrote:
> 2017-08-21 21:12+0200, Adam Borowski:
> > On Mon, Aug 21, 2017 at 09:26:57AM +0800, Wanpeng Li wrote:
> > > 2017-08-21 7:13 GMT+08:00 Adam Borowski <kilobyte@angband.pl>:
> > > > I'm afraid I keep getting a quite reliable, but random, splat when running
> > > > KVM:
> > > 
> > > I reported something similar before. https://lkml.org/lkml/2017/6/29/64
> > 
> > Your problem seems to require OOM; I don't have any memory pressure at all:
> > running a single 2GB guest while there's nothing big on the host (bloatfox,
> > xfce, xorg, terminals + some minor junk); 8GB + (untouched) swap.  There's
> > no memory pressure inside the guest either -- none was Linux (I wanted to
> > test something on hurd, kfreebsd) and I doubt they even got to use all of
> > their frames.
> 
> I even tried hurd, but couldn't reproduce ...

Also happens with a win10 guest, and with multiple Linuxes.

> what is your qemu command
> line and the output of host's `grep . /sys/module/kvm*/parameters/*`?

qemu-system-x86_64 -enable-kvm -m 2048 -vga qxl -usbdevice tablet \
 -net bridge -net nic \
 -drive file="$DISK",cache=writeback,index=0,media=disk,discard=on

qemu-system-x86_64 -enable-kvm -m 2048 -vga qxl -usbdevice tablet \
 -net bridge -net nic \
 -drive file="$DISK",cache=unsafe,index=0,media=disk,discard=on,if=virtio,format=raw

/sys/module/kvm/parameters/halt_poll_ns:200000
/sys/module/kvm/parameters/halt_poll_ns_grow:2
/sys/module/kvm/parameters/halt_poll_ns_shrink:0
/sys/module/kvm/parameters/ignore_msrs:N
/sys/module/kvm/parameters/kvmclock_periodic_sync:Y
/sys/module/kvm/parameters/lapic_timer_advance_ns:0
/sys/module/kvm/parameters/min_timer_period_us:500
/sys/module/kvm/parameters/tsc_tolerance_ppm:250
/sys/module/kvm/parameters/vector_hashing:Y
/sys/module/kvm_amd/parameters/avic:0
/sys/module/kvm_amd/parameters/nested:1
/sys/module/kvm_amd/parameters/npt:1
/sys/module/kvm_amd/parameters/vls:0

> > Also, it doesn't reproduce for me on 4.12.
> 
> Great info ... the most suspicious between v4.12 and v4.13-rc5 is the
> series with dcdca5fed5f6 ("x86: kvm: mmu: make spte mmio mask more
> explicit"), does reverting it help?
> 
> `git revert ce00053b1cfca312c22e2a6465451f1862561eab~1..995f00a619584e65e53eff372d9b73b121a7bad5`

Alas, doesn't seem to help.

I've first installed a Debian stretch guest, the host survived both the
installation and subsequent fooling around.  But then I started a win10
guest which splatted as soon as the initial screen.


Meow!
-- 
⢀⣴⠾⠻⢶⣦⠀ 
⣾⠁⢰⠒⠀⣿⡁ Vat kind uf sufficiently advanced technology iz dis!?
⢿⡄⠘⠷⠚⠋⠀                                        -- Genghis Ht'rok'din
⠈⠳⣄⠀⠀⠀⠀ 

[toc] | [prev] | [next] | [standalone]


#1718308

FromPaolo Bonzini <pbonzini@redhat.com>
Date2017-08-23 14:30 +0200
Message-ID<uhCXg-u4-3@gated-at.bofh.it>
In reply to#1716924
On 22/08/2017 00:32, Adam Borowski wrote:
> On Mon, Aug 21, 2017 at 09:58:34PM +0200, Radim Krčmář wrote:
>> 2017-08-21 21:12+0200, Adam Borowski:
>>> On Mon, Aug 21, 2017 at 09:26:57AM +0800, Wanpeng Li wrote:
>>>> 2017-08-21 7:13 GMT+08:00 Adam Borowski <kilobyte@angband.pl>:
>>>>> I'm afraid I keep getting a quite reliable, but random, splat when running
>>>>> KVM:
>>>>
>>>> I reported something similar before. https://lkml.org/lkml/2017/6/29/64
>>>
>>> Your problem seems to require OOM; I don't have any memory pressure at all:
>>> running a single 2GB guest while there's nothing big on the host (bloatfox,
>>> xfce, xorg, terminals + some minor junk); 8GB + (untouched) swap.  There's
>>> no memory pressure inside the guest either -- none was Linux (I wanted to
>>> test something on hurd, kfreebsd) and I doubt they even got to use all of
>>> their frames.
>>
>> I even tried hurd, but couldn't reproduce ...
> 
> Also happens with a win10 guest, and with multiple Linuxes.
> 
>> what is your qemu command
>> line and the output of host's `grep . /sys/module/kvm*/parameters/*`?
> 
> qemu-system-x86_64 -enable-kvm -m 2048 -vga qxl -usbdevice tablet \
>  -net bridge -net nic \
>  -drive file="$DISK",cache=writeback,index=0,media=disk,discard=on
> 
> qemu-system-x86_64 -enable-kvm -m 2048 -vga qxl -usbdevice tablet \
>  -net bridge -net nic \
>  -drive file="$DISK",cache=unsafe,index=0,media=disk,discard=on,if=virtio,format=raw
> 
> /sys/module/kvm/parameters/halt_poll_ns:200000
> /sys/module/kvm/parameters/halt_poll_ns_grow:2
> /sys/module/kvm/parameters/halt_poll_ns_shrink:0
> /sys/module/kvm/parameters/ignore_msrs:N
> /sys/module/kvm/parameters/kvmclock_periodic_sync:Y
> /sys/module/kvm/parameters/lapic_timer_advance_ns:0
> /sys/module/kvm/parameters/min_timer_period_us:500
> /sys/module/kvm/parameters/tsc_tolerance_ppm:250
> /sys/module/kvm/parameters/vector_hashing:Y
> /sys/module/kvm_amd/parameters/avic:0
> /sys/module/kvm_amd/parameters/nested:1
> /sys/module/kvm_amd/parameters/npt:1
> /sys/module/kvm_amd/parameters/vls:0
> 
>>> Also, it doesn't reproduce for me on 4.12.
>>
>> Great info ... the most suspicious between v4.12 and v4.13-rc5 is the
>> series with dcdca5fed5f6 ("x86: kvm: mmu: make spte mmio mask more
>> explicit"), does reverting it help?
>>
>> `git revert ce00053b1cfca312c22e2a6465451f1862561eab~1..995f00a619584e65e53eff372d9b73b121a7bad5`
> 
> Alas, doesn't seem to help.
> 
> I've first installed a Debian stretch guest, the host survived both the
> installation and subsequent fooling around.  But then I started a win10
> guest which splatted as soon as the initial screen.

Can you check if disabling THP on the host also fixes it for you?  I
would also try commit 1372324b328cd5dabaef5e345e37ad48c63df2a9 to
identify whether it was caused by a KVM change in 4.13 or something else.

Paolo

[toc] | [prev] | [next] | [standalone]


#1718928

FromWanpeng Li <kernellwp@gmail.com>
Date2017-08-24 09:50 +0200
Message-ID<uhV3P-3x1-5@gated-at.bofh.it>
In reply to#1718308
2017-08-23 20:22 GMT+08:00 Paolo Bonzini <pbonzini@redhat.com>:
> On 22/08/2017 00:32, Adam Borowski wrote:
>> On Mon, Aug 21, 2017 at 09:58:34PM +0200, Radim Krčmář wrote:
>>> 2017-08-21 21:12+0200, Adam Borowski:
>>>> On Mon, Aug 21, 2017 at 09:26:57AM +0800, Wanpeng Li wrote:
>>>>> 2017-08-21 7:13 GMT+08:00 Adam Borowski <kilobyte@angband.pl>:
>>>>>> I'm afraid I keep getting a quite reliable, but random, splat when running
>>>>>> KVM:
>>>>>
>>>>> I reported something similar before. https://lkml.org/lkml/2017/6/29/64
>>>>
>>>> Your problem seems to require OOM; I don't have any memory pressure at all:
>>>> running a single 2GB guest while there's nothing big on the host (bloatfox,
>>>> xfce, xorg, terminals + some minor junk); 8GB + (untouched) swap.  There's
>>>> no memory pressure inside the guest either -- none was Linux (I wanted to
>>>> test something on hurd, kfreebsd) and I doubt they even got to use all of
>>>> their frames.
>>>
>>> I even tried hurd, but couldn't reproduce ...
>>
>> Also happens with a win10 guest, and with multiple Linuxes.
>>
>>> what is your qemu command
>>> line and the output of host's `grep . /sys/module/kvm*/parameters/*`?
>>
>> qemu-system-x86_64 -enable-kvm -m 2048 -vga qxl -usbdevice tablet \
>>  -net bridge -net nic \
>>  -drive file="$DISK",cache=writeback,index=0,media=disk,discard=on
>>
>> qemu-system-x86_64 -enable-kvm -m 2048 -vga qxl -usbdevice tablet \
>>  -net bridge -net nic \
>>  -drive file="$DISK",cache=unsafe,index=0,media=disk,discard=on,if=virtio,format=raw
>>
>> /sys/module/kvm/parameters/halt_poll_ns:200000
>> /sys/module/kvm/parameters/halt_poll_ns_grow:2
>> /sys/module/kvm/parameters/halt_poll_ns_shrink:0
>> /sys/module/kvm/parameters/ignore_msrs:N
>> /sys/module/kvm/parameters/kvmclock_periodic_sync:Y
>> /sys/module/kvm/parameters/lapic_timer_advance_ns:0
>> /sys/module/kvm/parameters/min_timer_period_us:500
>> /sys/module/kvm/parameters/tsc_tolerance_ppm:250
>> /sys/module/kvm/parameters/vector_hashing:Y
>> /sys/module/kvm_amd/parameters/avic:0
>> /sys/module/kvm_amd/parameters/nested:1
>> /sys/module/kvm_amd/parameters/npt:1
>> /sys/module/kvm_amd/parameters/vls:0
>>
>>>> Also, it doesn't reproduce for me on 4.12.
>>>
>>> Great info ... the most suspicious between v4.12 and v4.13-rc5 is the
>>> series with dcdca5fed5f6 ("x86: kvm: mmu: make spte mmio mask more
>>> explicit"), does reverting it help?
>>>
>>> `git revert ce00053b1cfca312c22e2a6465451f1862561eab~1..995f00a619584e65e53eff372d9b73b121a7bad5`
>>
>> Alas, doesn't seem to help.
>>
>> I've first installed a Debian stretch guest, the host survived both the
>> installation and subsequent fooling around.  But then I started a win10
>> guest which splatted as soon as the initial screen.
>
> Can you check if disabling THP on the host also fixes it for you?  I
> would also try commit 1372324b328cd5dabaef5e345e37ad48c63df2a9 to
> identify whether it was caused by a KVM change in 4.13 or something else.

For the OOM testcase, the splat will disappear if disabling THP.

Regards,
Wanpeng Li

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web