Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1415972 > unrolled thread

Re: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug

Started byPaolo Bonzini <pbonzini@redhat.com>
First post2016-06-07 12:40 +0200
Last post2016-06-07 20:30 +0200
Articles 4 — 3 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug Paolo Bonzini <pbonzini@redhat.com> - 2016-06-07 12:40 +0200
    Re: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug Wanpeng Li <kernellwp@gmail.com> - 2016-06-07 13:20 +0200
    Re: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug Wanpeng Li <kernellwp@gmail.com> - 2016-06-07 14:00 +0200
      Re: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug John Stultz <john.stultz@linaro.org> - 2016-06-07 20:30 +0200

#1415972 — Re: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug

FromPaolo Bonzini <pbonzini@redhat.com>
Date2016-06-07 12:40 +0200
SubjectRe: [PATCH v4 1/3] KVM: fix steal clock warp during guest cpu hotplug
Message-ID<rHmAp-1vx-39@gated-at.bofh.it>

On 07/06/2016 09:59, Wanpeng Li wrote:
> From: Wanpeng Li <wanpeng.li@hotmail.com>
> 
> I observed that sometimes st is 100% instantaneous, then idle is 100% 
> even if there is a cpu hog on the guest cpu after the cpu hotplug comes 
> back(N.B. this can not always be readily reproduced). I add trace to 
> capture it as below:
> 
> cpuhp/1-12    [001] d.h1   167.461657: account_process_tick: steal = 1291385514, prev_steal_time = 0         
> cpuhp/1-12    [001] d.h1   167.461659: account_process_tick: steal_jiffies = 1291          
> <idle>-0     [001] d.h1   167.462663: account_process_tick: steal = 18732255, prev_steal_time = 1291000000          
> <idle>-0     [001] d.h1   167.462664: account_process_tick: steal_jiffies = 18446744072437
> 
> The steal clock warp and then steal_jiffies underflow.
> 
> Rik also pointed out to me:
>  
> | I have seen stuff like that with live migration too, in the past 
> 
> The root cause of steal clock warp during hotplug is kvm_steal_time reset 
> to 0 after cpu hotplug comes back which should be preexiting guest value. 
> This patch fix it by don't reset kvm_steal_time during guest cpu hotplug.

Improved commit message:

Sometimes, after CPU hotplug you can observe a spike in stolen time 
(100%) followed by the CPU being marked as 100% idle when it's actually 
busy with a CPU hog task.  The trace looks like the following:

cpuhp/1-12    [001] d.h1   167.461657: account_process_tick: steal = 1291385514, prev_steal_time = 0         
cpuhp/1-12    [001] d.h1   167.461659: account_process_tick: steal_jiffies = 1291          
<idle>-0     [001] d.h1   167.462663: account_process_tick: steal = 18732255, prev_steal_time = 1291000000          
<idle>-0     [001] d.h1   167.462664: account_process_tick: steal_jiffies = 18446744072437

The sudden decrease of "steal" causes steal_jiffies to underflow.
The root cause is kvm_steal_time being reset to 0 after hot-plugging
back in a CPU.  Instead, the preexisting value can be used, which is
what the core scheduler code expects.

John Stultz also reported a similar issue after guest S3.
------

Please also add

Cc: John Stultz <john.stultz@linaro.org>

Thanks,

Paolo

[toc] | [next] | [standalone]


#1416030

FromWanpeng Li <kernellwp@gmail.com>
Date2016-06-07 13:20 +0200
Message-ID<rHnd8-1YN-15@gated-at.bofh.it>
In reply to#1415972
2016-06-07 18:39 GMT+08:00 Paolo Bonzini <pbonzini@redhat.com>:
>
>
> On 07/06/2016 09:59, Wanpeng Li wrote:
>> From: Wanpeng Li <wanpeng.li@hotmail.com>
>>
>> I observed that sometimes st is 100% instantaneous, then idle is 100%
>> even if there is a cpu hog on the guest cpu after the cpu hotplug comes
>> back(N.B. this can not always be readily reproduced). I add trace to
>> capture it as below:
>>
>> cpuhp/1-12    [001] d.h1   167.461657: account_process_tick: steal = 1291385514, prev_steal_time = 0
>> cpuhp/1-12    [001] d.h1   167.461659: account_process_tick: steal_jiffies = 1291
>> <idle>-0     [001] d.h1   167.462663: account_process_tick: steal = 18732255, prev_steal_time = 1291000000
>> <idle>-0     [001] d.h1   167.462664: account_process_tick: steal_jiffies = 18446744072437
>>
>> The steal clock warp and then steal_jiffies underflow.
>>
>> Rik also pointed out to me:
>>
>> | I have seen stuff like that with live migration too, in the past
>>
>> The root cause of steal clock warp during hotplug is kvm_steal_time reset
>> to 0 after cpu hotplug comes back which should be preexiting guest value.
>> This patch fix it by don't reset kvm_steal_time during guest cpu hotplug.
>
> Improved commit message:
>
> Sometimes, after CPU hotplug you can observe a spike in stolen time
> (100%) followed by the CPU being marked as 100% idle when it's actually
> busy with a CPU hog task.  The trace looks like the following:
>
> cpuhp/1-12    [001] d.h1   167.461657: account_process_tick: steal = 1291385514, prev_steal_time = 0
> cpuhp/1-12    [001] d.h1   167.461659: account_process_tick: steal_jiffies = 1291
> <idle>-0     [001] d.h1   167.462663: account_process_tick: steal = 18732255, prev_steal_time = 1291000000
> <idle>-0     [001] d.h1   167.462664: account_process_tick: steal_jiffies = 18446744072437
>
> The sudden decrease of "steal" causes steal_jiffies to underflow.
> The root cause is kvm_steal_time being reset to 0 after hot-plugging
> back in a CPU.  Instead, the preexisting value can be used, which is
> what the core scheduler code expects.
>
> John Stultz also reported a similar issue after guest S3.
> ------
>
> Please also add
>
> Cc: John Stultz <john.stultz@linaro.org>

Thanks Paolo! Your help is always a great appreciated. :)

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1416073

FromWanpeng Li <kernellwp@gmail.com>
Date2016-06-07 14:00 +0200
Message-ID<rHnPP-2e7-7@gated-at.bofh.it>
In reply to#1415972
2016-06-07 18:39 GMT+08:00 Paolo Bonzini <pbonzini@redhat.com>:
[...]
>
> John Stultz also reported a similar issue after guest S3.

Since there is cpu hot-unplug during S3.

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1416497

FromJohn Stultz <john.stultz@linaro.org>
Date2016-06-07 20:30 +0200
Message-ID<rHtVf-686-1@gated-at.bofh.it>
In reply to#1416073
On Tue, Jun 7, 2016 at 4:52 AM, Wanpeng Li <kernellwp@gmail.com> wrote:
> 2016-06-07 18:39 GMT+08:00 Paolo Bonzini <pbonzini@redhat.com>:
> [...]
>>
>> John Stultz also reported a similar issue after guest S3.
>
> Since there is cpu hot-unplug during S3.

While I'm excited to finally see some progress on this, I
unfortunately can't verify this fixes the issue I saw, as my qemu/kvm
test environments have regressed far enough that suspend/resume no
longer works. :P

I'll see about trying to re-build qemu from source to see if the
latest code works.

thanks
-john

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web