Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1415972 > unrolled thread
| Started by | Paolo Bonzini <pbonzini@redhat.com> |
|---|---|
| First post | 2016-06-07 12:40 +0200 |
| Last post | 2016-06-07 20:30 +0200 |
| Articles | 4 — 3 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug Paolo Bonzini <pbonzini@redhat.com> - 2016-06-07 12:40 +0200
Re: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug Wanpeng Li <kernellwp@gmail.com> - 2016-06-07 13:20 +0200
Re: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug Wanpeng Li <kernellwp@gmail.com> - 2016-06-07 14:00 +0200
Re: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug John Stultz <john.stultz@linaro.org> - 2016-06-07 20:30 +0200
| From | Paolo Bonzini <pbonzini@redhat.com> |
|---|---|
| Date | 2016-06-07 12:40 +0200 |
| Subject | Re: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug |
| Message-ID | <rHmAp-1vx-39@gated-at.bofh.it> |
On 07/06/2016 09:59, Wanpeng Li wrote: > From: Wanpeng Li <wanpeng.li@hotmail.com> > > I observed that sometimes st is 100% instantaneous, then idle is 100% > even if there is a cpu hog on the guest cpu after the cpu hotplug comes > back(N.B. this can not always be readily reproduced). I add trace to > capture it as below: > > cpuhp/1-12 [001] d.h1 167.461657: account_process_tick: steal = 1291385514, prev_steal_time = 0 > cpuhp/1-12 [001] d.h1 167.461659: account_process_tick: steal_jiffies = 1291 > <idle>-0 [001] d.h1 167.462663: account_process_tick: steal = 18732255, prev_steal_time = 1291000000 > <idle>-0 [001] d.h1 167.462664: account_process_tick: steal_jiffies = 18446744072437 > > The steal clock warp and then steal_jiffies underflow. > > Rik also pointed out to me: > > | I have seen stuff like that with live migration too, in the past > > The root cause of steal clock warp during hotplug is kvm_steal_time reset > to 0 after cpu hotplug comes back which should be preexiting guest value. > This patch fix it by don't reset kvm_steal_time during guest cpu hotplug. Improved commit message: Sometimes, after CPU hotplug you can observe a spike in stolen time (100%) followed by the CPU being marked as 100% idle when it's actually busy with a CPU hog task. The trace looks like the following: cpuhp/1-12 [001] d.h1 167.461657: account_process_tick: steal = 1291385514, prev_steal_time = 0 cpuhp/1-12 [001] d.h1 167.461659: account_process_tick: steal_jiffies = 1291 <idle>-0 [001] d.h1 167.462663: account_process_tick: steal = 18732255, prev_steal_time = 1291000000 <idle>-0 [001] d.h1 167.462664: account_process_tick: steal_jiffies = 18446744072437 The sudden decrease of "steal" causes steal_jiffies to underflow. The root cause is kvm_steal_time being reset to 0 after hot-plugging back in a CPU. Instead, the preexisting value can be used, which is what the core scheduler code expects. John Stultz also reported a similar issue after guest S3. ------ Please also add Cc: John Stultz <john.stultz@linaro.org> Thanks, Paolo
[toc] | [next] | [standalone]
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2016-06-07 13:20 +0200 |
| Message-ID | <rHnd8-1YN-15@gated-at.bofh.it> |
| In reply to | #1415972 |
2016-06-07 18:39 GMT+08:00 Paolo Bonzini <pbonzini@redhat.com>: > > > On 07/06/2016 09:59, Wanpeng Li wrote: >> From: Wanpeng Li <wanpeng.li@hotmail.com> >> >> I observed that sometimes st is 100% instantaneous, then idle is 100% >> even if there is a cpu hog on the guest cpu after the cpu hotplug comes >> back(N.B. this can not always be readily reproduced). I add trace to >> capture it as below: >> >> cpuhp/1-12 [001] d.h1 167.461657: account_process_tick: steal = 1291385514, prev_steal_time = 0 >> cpuhp/1-12 [001] d.h1 167.461659: account_process_tick: steal_jiffies = 1291 >> <idle>-0 [001] d.h1 167.462663: account_process_tick: steal = 18732255, prev_steal_time = 1291000000 >> <idle>-0 [001] d.h1 167.462664: account_process_tick: steal_jiffies = 18446744072437 >> >> The steal clock warp and then steal_jiffies underflow. >> >> Rik also pointed out to me: >> >> | I have seen stuff like that with live migration too, in the past >> >> The root cause of steal clock warp during hotplug is kvm_steal_time reset >> to 0 after cpu hotplug comes back which should be preexiting guest value. >> This patch fix it by don't reset kvm_steal_time during guest cpu hotplug. > > Improved commit message: > > Sometimes, after CPU hotplug you can observe a spike in stolen time > (100%) followed by the CPU being marked as 100% idle when it's actually > busy with a CPU hog task. The trace looks like the following: > > cpuhp/1-12 [001] d.h1 167.461657: account_process_tick: steal = 1291385514, prev_steal_time = 0 > cpuhp/1-12 [001] d.h1 167.461659: account_process_tick: steal_jiffies = 1291 > <idle>-0 [001] d.h1 167.462663: account_process_tick: steal = 18732255, prev_steal_time = 1291000000 > <idle>-0 [001] d.h1 167.462664: account_process_tick: steal_jiffies = 18446744072437 > > The sudden decrease of "steal" causes steal_jiffies to underflow. > The root cause is kvm_steal_time being reset to 0 after hot-plugging > back in a CPU. Instead, the preexisting value can be used, which is > what the core scheduler code expects. > > John Stultz also reported a similar issue after guest S3. > ------ > > Please also add > > Cc: John Stultz <john.stultz@linaro.org> Thanks Paolo! Your help is always a great appreciated. :) Regards, Wanpeng Li
[toc] | [prev] | [next] | [standalone]
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2016-06-07 14:00 +0200 |
| Message-ID | <rHnPP-2e7-7@gated-at.bofh.it> |
| In reply to | #1415972 |
2016-06-07 18:39 GMT+08:00 Paolo Bonzini <pbonzini@redhat.com>: [...] > > John Stultz also reported a similar issue after guest S3. Since there is cpu hot-unplug during S3. Regards, Wanpeng Li
[toc] | [prev] | [next] | [standalone]
| From | John Stultz <john.stultz@linaro.org> |
|---|---|
| Date | 2016-06-07 20:30 +0200 |
| Message-ID | <rHtVf-686-1@gated-at.bofh.it> |
| In reply to | #1416073 |
On Tue, Jun 7, 2016 at 4:52 AM, Wanpeng Li <kernellwp@gmail.com> wrote: > 2016-06-07 18:39 GMT+08:00 Paolo Bonzini <pbonzini@redhat.com>: > [...] >> >> John Stultz also reported a similar issue after guest S3. > > Since there is cpu hot-unplug during S3. While I'm excited to finally see some progress on this, I unfortunately can't verify this fixes the issue I saw, as my qemu/kvm test environments have regressed far enough that suspend/resume no longer works. :P I'll see about trying to re-build qemu from source to see if the latest code works. thanks -john
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web