Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #57538 > unrolled thread

Bug#860236: xen pv domU crash with 3.16 kernel and xen 4.8

Started byVincent Legout <vincent.legout@gandi.net>
First post2017-04-13 11:30 +0200
Last post2017-04-20 09:10 +0200
Articles 7 — 3 participants

Back to article view | Back to linux.debian.kernel


Contents

  Bug#860236: xen pv domU crash with 3.16 kernel and xen 4.8 Vincent Legout <vincent.legout@gandi.net> - 2017-04-13 11:30 +0200
    Processed: Re: Bug#860236: xen pv domU crash with 3.16 kernel and  xen 4.8 owner@bugs.debian.org (Debian Bug Tracking System) - 2017-04-14 00:50 +0200
    Bug#860236: xen pv domU crash with 3.16 kernel and xen 4.8 Ben Hutchings <ben@decadent.org.uk> - 2017-04-14 00:50 +0200
      Bug#860236: xen pv domU crash with 3.16 kernel and xen 4.8 Vincent Legout <vincent.legout@gandi.net> - 2017-04-14 09:30 +0200
        Bug#860236: xen pv domU crash with 3.16 kernel and xen 4.8 Vincent Legout <vincent.legout@gandi.net> - 2017-04-14 11:40 +0200
          Bug#860236: xen pv domU crash with 3.16 kernel and xen 4.8 Ben Hutchings <ben@decadent.org.uk> - 2017-04-19 21:50 +0200
            Bug#860236: xen pv domU crash with 3.16 kernel and xen 4.8 Vincent Legout <vincent.legout@gandi.net> - 2017-04-20 09:10 +0200

#57538 — Bug#860236: xen pv domU crash with 3.16 kernel and xen 4.8

FromVincent Legout <vincent.legout@gandi.net>
Date2017-04-13 11:30 +0200
SubjectBug#860236: xen pv domU crash with 3.16 kernel and xen 4.8
Message-ID<tvJeH-Wr-29@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

Package: src:linux
Version: 3.16.39-1+deb8u2
Severity: normal

Hi,

A xen jessie domU crashes around 5 minutes after the boot with the
attached backtrace (at every boot). dom0 is also a Debian jessie running
Xen 4.8.

It only happens when the guest is in pv mode, it works fine with pvhvm.

It also crashes with older 3.16 kernels and 4.0.2-1, but not with
4.2.1-1 (last 2 kernels from snapshot.debian.org).

# uname -a
3.16.0-4-amd64 #1 SMP Debian 3.16.39-1+deb8u2 (2017-03-07) x86_64 GNU/Linux

Vincent

[toc] | [next] | [standalone]


#57544 — Processed: Re: Bug#860236: xen pv domU crash with 3.16 kernel and xen 4.8

Fromowner@bugs.debian.org (Debian Bug Tracking System)
Date2017-04-14 00:50 +0200
SubjectProcessed: Re: Bug#860236: xen pv domU crash with 3.16 kernel and xen 4.8
Message-ID<tvVIR-15E-3@gated-at.bofh.it>
In reply to#57538
Processing control commands:

> tag -1 moreinfo
Bug #860236 [src:linux] xen pv domU crash with 3.16 kernel and xen 4.8
Added tag(s) moreinfo.

-- 
860236: http://bugs.debian.org/cgi-bin/bugreport.cgi?bug=860236
Debian Bug Tracking System
Contact owner@bugs.debian.org with problems

[toc] | [prev] | [next] | [standalone]


#57545

FromBen Hutchings <ben@decadent.org.uk>
Date2017-04-14 00:50 +0200
Message-ID<tvVIR-15E-5@gated-at.bofh.it>
In reply to#57538

[Multipart message — attachments visible in raw view] — view raw

Control: tag -1 moreinfo

On Thu, 2017-04-13 at 11:18 +0200, Vincent Legout wrote:
> Package: src:linux
> Version: 3.16.39-1+deb8u2
> Severity: normal
> 
> Hi,
> 
> A xen jessie domU crashes around 5 minutes after the boot with the
> attached backtrace (at every boot). dom0 is also a Debian jessie running
> Xen 4.8.
> 
> It only happens when the guest is in pv mode, it works fine with pvhvm.
> 
> It also crashes with older 3.16 kernels and 4.0.2-1, but not with
> 4.2.1-1 (last 2 kernels from snapshot.debian.org).
> 
> # uname -a
> 3.16.0-4-amd64 #1 SMP Debian 3.16.39-1+deb8u2 (2017-03-07) x86_64 GNU/Linux

From the crash log:

> [  300.632389] CPU: 0 PID: 0 Comm: swapper/0 Tainted: G        W     3.16.0-4-amd64 #1 Debian 3.16.39-1+deb8u2

This indicates there was an earlier WARNING message; what was that?

Ben.

-- 
Ben Hutchings
Any sufficiently advanced bug is indistinguishable from a feature.

[toc] | [prev] | [next] | [standalone]


#57548

FromVincent Legout <vincent.legout@gandi.net>
Date2017-04-14 09:30 +0200
Message-ID<tw3Q5-6Fl-1@gated-at.bofh.it>
In reply to#57545

[Multipart message — attachments visible in raw view] — view raw

On Thu, Apr 13, 2017 at 11:41:37PM +0100, Ben Hutchings wrote :
> Control: tag -1 moreinfo
> 
> On Thu, 2017-04-13 at 11:18 +0200, Vincent Legout wrote:
> > Package: src:linux
> > Version: 3.16.39-1+deb8u2
> > Severity: normal
> > 
> > Hi,
> > 
> > A xen jessie domU crashes around 5 minutes after the boot with the
> > attached backtrace (at every boot). dom0 is also a Debian jessie running
> > Xen 4.8.
> > 
> > It only happens when the guest is in pv mode, it works fine with pvhvm.
> > 
> > It also crashes with older 3.16 kernels and 4.0.2-1, but not with
> > 4.2.1-1 (last 2 kernels from snapshot.debian.org).
> > 
> > # uname -a
> > 3.16.0-4-amd64 #1 SMP Debian 3.16.39-1+deb8u2 (2017-03-07) x86_64 GNU/Linux
> 
> From the crash log:
> 
> > [  300.632389] CPU: 0 PID: 0 Comm: swapper/0 Tainted: G        W     3.16.0-4-amd64 #1 Debian 3.16.39-1+deb8u2
> 
> This indicates there was an earlier WARNING message; what was that?

Thanks for the answer.

I got this WARNING after I increased verbosity in the command line:

[  300.636063] ------------[ cut here ]------------
[  300.636102] WARNING: CPU: 0 PID: 0 at /build/linux-GSgHvp/linux-3.16.39/arch/x86/kernel/cpu/mcheck/mce.c:1307 mce_timer_fn+0x132/0x140()
[  300.636116] Modules linked in: x86_pkg_temp_thermal thermal_sys intel_rapl coretemp crc32_pclmul evdev aesni_intel aes_x86_64 lrw gf128mul glue_helper ablk_helper pcspkr cryptd autofs4 ext4 crc16 mbcache jbd2 xen_netfront xen_blkfront crct10dif_pclmul crct10dif_common crc32c_intel
[  300.636167] CPU: 0 PID: 0 Comm: swapper/0 Not tainted 3.16.0-4-amd64 #1 Debian 3.16.39-1+deb8u2
[  300.636178]  0000000000000000 ffffffff81514c81 0000000000000000 0000000000000009
[  300.636188]  ffffffff81068867 ffff88003f80ca00 ffff88003f9eca00 0000000000000100
[  300.636199]  ffffffff81038a30 000000000000000f ffffffff81038b62 ffffffff81a66e00
[  300.636211] Call Trace:
[  300.636216]  <IRQ>  [<ffffffff81514c81>] ? dump_stack+0x5d/0x78
[  300.636242]  [<ffffffff81068867>] ? warn_slowpath_common+0x77/0x90
[  300.636250]  [<ffffffff81038a30>] ? mce_cpu_restart+0x40/0x40
[  300.636257]  [<ffffffff81038b62>] ? mce_timer_fn+0x132/0x140
[  300.636267]  [<ffffffff81073ea1>] ? call_timer_fn+0x31/0x140
[  300.636274]  [<ffffffff81038a30>] ? mce_cpu_restart+0x40/0x40
[  300.636284]  [<ffffffff81075559>] ? run_timer_softirq+0x1e9/0x2f0
[  300.636292]  [<ffffffff8106d911>] ? __do_softirq+0xf1/0x2d0
[  300.636299]  [<ffffffff8106dd25>] ? irq_exit+0x95/0xa0
[  300.636309]  [<ffffffff8135cca5>] ? xen_evtchn_do_upcall+0x35/0x50
[  300.636319]  [<ffffffff8151cade>] ? xen_do_hypervisor_callback+0x1e/0x30
[  300.636324]  <EOI>  [<ffffffff810013ac>] ? xen_hypercall_sched_op+0xc/0x20
[  300.636339]  [<ffffffff810013ac>] ? xen_hypercall_sched_op+0xc/0x20
[  300.636349]  [<ffffffff8100ad3c>] ? xen_safe_halt+0xc/0x20
[  300.636360]  [<ffffffff8101da69>] ? default_idle+0x19/0xd0
[  300.636370]  [<ffffffff810a9b74>] ? cpu_startup_entry+0x374/0x470
[  300.636384]  [<ffffffff81903076>] ? start_kernel+0x497/0x4a2
[  300.636392]  [<ffffffff81902a04>] ? set_init_arg+0x4e/0x4e
[  300.636400]  [<ffffffff81904f91>] ? xen_start_kernel+0x569/0x573
[  300.636413] ---[ end trace 7131ef713ca84161 ]---

Then, the same BUG as before. It always happens after 300 seconds.

Vincent

[toc] | [prev] | [next] | [standalone]


#57549

FromVincent Legout <vincent.legout@gandi.net>
Date2017-04-14 11:40 +0200
Message-ID<tw5RU-7Pb-35@gated-at.bofh.it>
In reply to#57548

[Multipart message — attachments visible in raw view] — view raw

On Fri, Apr 14, 2017 at 09:15:58AM +0200, Vincent Legout wrote :
> On Thu, Apr 13, 2017 at 11:41:37PM +0100, Ben Hutchings wrote :
> > Control: tag -1 moreinfo
> > 
> > On Thu, 2017-04-13 at 11:18 +0200, Vincent Legout wrote:
> > > Package: src:linux
> > > Version: 3.16.39-1+deb8u2
> > > Severity: normal
> > > 
> > > Hi,
> > > 
> > > A xen jessie domU crashes around 5 minutes after the boot with the
> > > attached backtrace (at every boot). dom0 is also a Debian jessie running
> > > Xen 4.8.
> > > 
> > > It only happens when the guest is in pv mode, it works fine with pvhvm.
> > > 
> > > It also crashes with older 3.16 kernels and 4.0.2-1, but not with
> > > 4.2.1-1 (last 2 kernels from snapshot.debian.org).
> > > 
> > > # uname -a
> > > 3.16.0-4-amd64 #1 SMP Debian 3.16.39-1+deb8u2 (2017-03-07) x86_64 GNU/Linux
> > 
> > From the crash log:
> > 
> > > [  300.632389] CPU: 0 PID: 0 Comm: swapper/0 Tainted: G        W     3.16.0-4-amd64 #1 Debian 3.16.39-1+deb8u2
> > 
> > This indicates there was an earlier WARNING message; what was that?
> 
> Thanks for the answer.
> 
> I got this WARNING after I increased verbosity in the command line:

The WARNING and BUG disappear if "maxvcpus" is disabled in the guest
configuration (which prevents adding or removing vcpus).

Could cpu hotplug be buggy in 3.16? And Xen triggers this bug after 5
minutes even without doing any 'xl vcpu-set'?

With "maxvcpus" set larger "vcpus", xl vcpu-set seems to work most of
the time (between 1 and 16 vcpus), but after several tries, I got the
attached trace.

Vincent

[toc] | [prev] | [next] | [standalone]


#57602

FromBen Hutchings <ben@decadent.org.uk>
Date2017-04-19 21:50 +0200
Message-ID<ty3LX-7rV-7@gated-at.bofh.it>
In reply to#57549

[Multipart message — attachments visible in raw view] — view raw

On Fri, 2017-04-14 at 11:18 +0200, Vincent Legout wrote:
[...]
> Could cpu hotplug be buggy in 3.16? And Xen triggers this bug after 5
> minutes even without doing any 'xl vcpu-set'?

The MCE polling timer for each CPU runs every 5 minutes, so this is
presumably the first time it runs.  Perhaps this domain is configured
such that CPUs are hot-removed shortly after boot?

In the first crash, it looks like the timer for CPU x!=0 is being
called on CPU 0.  In general this can happen if CPU x is hot-removed;
its timers are migrated to another CPU.  This should *not* be possible
with the MCE timer, as there is a hotplug callback that removes the
timer when a CPU is removed.  There is a check for the timer having
been migrated anyway, which triggers the WARNING.  The timer function
then tries to re-add the timer for the current CPU, but that's still
pending, which triggers the BUG.  Either the hotplug callback was not
called, or the timer was migrated before being removed resulting in a
race condition.

> With "maxvcpus" set larger "vcpus", xl vcpu-set seems to work most of
> the time (between 1 and 16 vcpus), but after several tries, I got the
> attached trace.

I'm not sure what's going on in this crash, but as it's a null
dereference in migrate_timer_list it seems somewhat related.

I didn't find any changes that would explain how this was fixed between
4.0 and 4.2.  I suggest you work around it by adding 'nomce' to the
kernel command line as I would expect Xen or dom0 to handle MCEs.

Ben.

-- 
Ben Hutchings
Man invented language to satisfy his deep need to complain. - Lily
Tomlin

[toc] | [prev] | [next] | [standalone]


#57606

FromVincent Legout <vincent.legout@gandi.net>
Date2017-04-20 09:10 +0200
Message-ID<tyeo1-5TJ-3@gated-at.bofh.it>
In reply to#57602

[Multipart message — attachments visible in raw view] — view raw

On Wed, Apr 19, 2017 at 08:39:05PM +0100, Ben Hutchings wrote :
> On Fri, 2017-04-14 at 11:18 +0200, Vincent Legout wrote:
> [...]
> > Could cpu hotplug be buggy in 3.16? And Xen triggers this bug after 5
> > minutes even without doing any 'xl vcpu-set'?
> 
> The MCE polling timer for each CPU runs every 5 minutes, so this is
> presumably the first time it runs.  Perhaps this domain is configured
> such that CPUs are hot-removed shortly after boot?

I didn't explicitly set anything like that, but I guess it could also be
a default configuration in Xen.

> In the first crash, it looks like the timer for CPU x!=0 is being
> called on CPU 0.  In general this can happen if CPU x is hot-removed;
> its timers are migrated to another CPU.  This should *not* be possible
> with the MCE timer, as there is a hotplug callback that removes the
> timer when a CPU is removed.  There is a check for the timer having
> been migrated anyway, which triggers the WARNING.  The timer function
> then tries to re-add the timer for the current CPU, but that's still
> pending, which triggers the BUG.  Either the hotplug callback was not
> called, or the timer was migrated before being removed resulting in a
> race condition.
> 
> > With "maxvcpus" set larger "vcpus", xl vcpu-set seems to work most of
> > the time (between 1 and 16 vcpus), but after several tries, I got the
> > attached trace.
> 
> I'm not sure what's going on in this crash, but as it's a null
> dereference in migrate_timer_list it seems somewhat related.
> 
> I didn't find any changes that would explain how this was fixed between
> 4.0 and 4.2.  I suggest you work around it by adding 'nomce' to the
> kernel command line as I would expect Xen or dom0 to handle MCEs.

Thanks a lot Ben, I can't reproduce the issue with 'nomce'.

Thanks,
Vincent

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web