Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1681205
| Path | csiph.com!aioe.org!bofh.it!news.nic.it!robomod |
|---|---|
| From | Wanpeng Li <kernellwp@gmail.com> |
| Newsgroups | linux.kernel |
| Subject | Re: [PATCH 2/2] x86/idle: use dynamic halt poll |
| Date | Wed, 05 Jul 2017 00:40:01 +0200 |
| Message-ID | <tZEE9-666-1@gated-at.bofh.it> (permalink) |
| References | <tV8tb-3Zd-3@gated-at.bofh.it> <tV8tb-3Zd-7@gated-at.bofh.it> <tV8We-4b0-15@gated-at.bofh.it> <tVnVf-5pm-9@gated-at.bofh.it> <tWWQV-8f0-9@gated-at.bofh.it> <tWXtD-hF-3@gated-at.bofh.it> <tWXMZ-pJ-17@gated-at.bofh.it> <tWXN0-pJ-27@gated-at.bofh.it> <tWZ2q-1ch-13@gated-at.bofh.it> <tWZc5-1gg-5@gated-at.bofh.it> <tWZF9-1JQ-57@gated-at.bofh.it> <tZ5Q7-7za-25@gated-at.bofh.it> |
| X-Original-To | Yang Zhang <yang.zhang.wz@gmail.com> |
| Dkim-Signature | v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20161025; h=mime-version:in-reply-to:references:from:date:message-id:subject:to :cc:content-transfer-encoding; bh=NqXkJwwqDPFwuMlxJxwHPVDWGSS9zq3D+Ue1Z/ne/A4=; b=EwVmwJyiILajb3yadfxy5Ex6BPL3oNeyjFWWkZ1F9qfWgcZ2LLNKVEq5KX5Mt4T5V2 vlw5Z3p5phiDoo+1ZgUQaCwAtcB+SNJubG7lG9vV6mxg0/m99k/SAimvGCPM4xv1Kj8C dOjqeNH2HrIyCdsFuxhKygS1C+tjHJZ53iQzJkQTMvBz5mKorwT/xELxfEjhZ5UcO2sQ gweT5foN+EhT8oKcIWGhvA2XFQehuA7mMUxd375SD7mkUa3dKQN25rQMPrpIX+UNWQj9 5soh0lUeBmjj5jraiUmacfEebzYKHvLkRpYkeHB88NHv+WqjNuo38To7nbggyfY3Zj5e hk3Q== |
| X-Google-Dkim-Signature | v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:mime-version:in-reply-to:references:from:date :message-id:subject:to:cc:content-transfer-encoding; bh=NqXkJwwqDPFwuMlxJxwHPVDWGSS9zq3D+Ue1Z/ne/A4=; b=ByJGOcLX4Cl1VKhFjZVDA23KPvwXwq6Qdn3Pe/sUPzx7oy7/Rp3qhesgkhD0lYLvZ6 eu8Q7rUeNCA3pcMC6DUJgWtUsbEhmI2ts1Sl5DGwKhUadC2MTlkqGkp4pvnG4tG6Dj4u ozs3/6Cer9xpJ07x/6uAbxwPxWrdhlnzOJ4ErUOM8+5b29rKjLD3cVtR0w2ySr5losOr bI9PXPa+BP2jjjY22laVEghty3tMc0CO3+ty/FXzb88nUtg9rwaeI0ziVa8DC8Zd/3AR GsAyY0ON7yjs8lISamzZWwmhG9oe1J7cavIXxjZZrGuzKmSnNsFhvrP7kZDCj5qBDVYp JdMw== |
| X-Gm-Message-State | AKS2vOxjdf+WArbwJvycrjN4iRh64O4WGKkJuT+HxAvBpBhsig9jyKLo mUinUPUnbChOeBgzR/SGcpZ/5VrP9A== |
| X-Received | by 10.202.72.201 with SMTP id v192mr29963911oia.131.1499207329746; Tue, 04 Jul 2017 15:28:49 -0700 (PDT) |
| MIME-Version | 1.0 |
| Content-Type | text/plain; charset="UTF-8" |
| Content-Transfer-Encoding | quoted-printable |
| Sender | robomod@news.nic.it |
| List-ID | <linux-kernel.vger.kernel.org> |
| X-Mailing-List | linux-kernel@vger.kernel.org |
| Approved | robomod@news.nic.it |
| Lines | 119 |
| Organization | linux.* mail to news gateway |
| X-Original-Cc | Radim Krčmář <rkrcmar@redhat.com>, Paolo Bonzini <pbonzini@redhat.com>, Thomas Gleixner <tglx@linutronix.de>, Ingo Molnar <mingo@redhat.com>, "H. Peter Anvin" <hpa@zytor.com>, "the arch/x86 maintainers" <x86@kernel.org>, Jonathan Corbet <corbet@lwn.net>, tony.luck@intel.com, Borislav Petkov <bp@alien8.de>, Peter Zijlstra <peterz@infradead.org>, mchehab@kernel.org, Andrew Morton <akpm@linux-foundation.org>, krzk@kernel.org, jpoimboe@redhat.com, Andy Lutomirski <luto@kernel.org>, Christian Borntraeger <borntraeger@de.ibm.com>, Thomas Garnier <thgarnie@google.com>, Robert Gerst <rgerst@gmail.com>, Mathias Krause <minipli@googlemail.com>, douly.fnst@cn.fujitsu.com, Nicolai Stange <nicstange@gmail.com>, Frederic Weisbecker <fweisbec@gmail.com>, dvlasenk@redhat.com, Daniel Bristot de Oliveira <bristot@redhat.com>, yamada.masahiro@socionext.com, mika.westerberg@linux.intel.com, Chen Yu <yu.c.chen@intel.com>, aaron.lu@intel.com, Steven Rostedt <rostedt@goodmis.org>, Kyle Huey <me@kylehuey.com>, Len Brown <len.brown@intel.com>, Prarit Bhargava <prarit@redhat.com>, hidehiro.kawai.ez@hitachi.com, fengtiantian@huawei.com, pmladek@suse.com, jeyu@redhat.com, Larry.Finger@lwfinger.net, zijun_hu@htc.com, luisbg@osg.samsung.com, johannes.berg@intel.com, niklas.soderlund+renesas@ragnatech.se, zlpnobody@gmail.com, Alexey Dobriyan <adobriyan@gmail.com>, fgao@ikuai8.com, ebiederm@xmission.com, Subash Abhinov Kasiviswanathan <subashab@codeaurora.org>, Arnd Bergmann <arnd@arndb.de>, Matt Fleming <matt@codeblueprint.co.uk>, Mel Gorman <mgorman@techsingularity.net>, "linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>, linux-doc@vger.kernel.org, linux-edac@vger.kernel.org, kvm <kvm@vger.kernel.org> |
| X-Original-Date | Wed, 5 Jul 2017 06:28:48 +0800 |
| X-Original-Message-ID | <CANRm+Cx6H7pSNqiG7UMzvqD661AE1CC7_N6n6rJ6TdW5HY0tXQ@mail.gmail.com> |
| X-Original-References | <1498130534-26568-1-git-send-email-root@ip-172-31-39-62.us-west-2.compute.internal> <1498130534-26568-3-git-send-email-root@ip-172-31-39-62.us-west-2.compute.internal> <4444ffc8-9e7b-5bd2-20da-af422fe834cc@redhat.com> <2245bef7-b668-9265-f3f8-3b63d71b1033@gmail.com> <7d085956-2573-212f-44f4-86104beba9bb@gmail.com> <fd7acd49-9d37-4495-79cd-250e3b54cac2@redhat.com> <CANRm+Cyd4ycNdDN8=M66re5RhnrJF=3HPiPuJeMAUgLRb1qBcg@mail.gmail.com> <05ec7efc-fb9c-ae24-5770-66fc472545a4@redhat.com> <20170627134043.GA1487@potion> <2771f905-d1b0-b118-9ae9-db5fb87f877c@redhat.com> <20170627142251.GB1487@potion> <be2a5434-b990-75a9-9136-1a4519a6ca4d@gmail.com> |
| X-Original-Sender | linux-kernel-owner@vger.kernel.org |
| Xref | csiph.com linux.kernel:1681205 |
Show key headers only | View raw
2017-07-03 17:28 GMT+08:00 Yang Zhang <yang.zhang.wz@gmail.com>: > On 2017/6/27 22:22, Radim Krčmář wrote: >> >> 2017-06-27 15:56+0200, Paolo Bonzini: >>> >>> On 27/06/2017 15:40, Radim Krčmář wrote: >>>>> >>>>> ... which is not necessarily _wrong_. It's just a different heuristic. >>>> >>>> Right, it's just harder to use than host's single_task_running() -- the >>>> VCPU calling vcpu_is_preempted() is never preempted, so we have to look >>>> at other VCPUs that are not halted, but still preempted. >>>> >>>> If we see some ratio of preempted VCPUs (> 0?), then we stop polling and >>>> yield to the host. Working under the assumption that there is work for >>>> this PCPU if other VCPUs have stuff to do. The downside is that it >>>> misses information about host's topology, so it would be hard to make it >>>> work well. >>> >>> >>> I would just use vcpu_is_preempted on the current CPU. From guest POV >>> this option is really a "f*** everyone else" setting just like >>> idle=poll, only a little more polite. >> >> >> vcpu_is_preempted() on current cpu cannot return true, AFAIK. >> >>> If we've been preempted and we were polling, there are two cases. If an >>> interrupt was queued while the guest was preempted, the poll will be >>> treated as successful anyway. >> >> >> I think the poll should be treated as invalid if the window has expired >> while the VCPU was preempted -- the guest can't tell whether the >> interrupt arrived still within the poll window (unless we added paravirt >> for that), so it shouldn't be wasting time waiting for it. >> >>> If it hasn't, let others run---but really >>> that's not because the guest wants to be polite, it's to avoid that the >>> scheduler penalizes it excessively. >> >> >> This sounds like a VM entry just to do an immediate VM exit, so paravirt >> seems better here as well ... (the guest telling the host about its >> window -- which could also be used to rule it out as a target in the >> pause loop random kick.) >> >>> So until it's preempted, I think it's okay if the guest doesn't care >>> about others. You wouldn't use this option anyway in overcommitted >>> situations. >>> >>> (I'm still not very convinced about the idea). >> >> >> Me neither. (The same mechanism is applicable to bare-metal, but was >> never used there, so I would rather bring the guest behavior closer to >> bare-metal.) >> > > The background is that we(Alibaba Cloud) do get more and more complaints > from our customers in both KVM and Xen compare to bare-mental.After > investigations, the root cause is known to us: big cost in message passing > workload(David show it in KVM forum 2015) > > A typical message workload like below: > vcpu 0 vcpu 1 > 1. send ipi 2. doing hlt > 3. go into idle 4. receive ipi and wake up from hlt > 5. write APIC time twice 6. write APIC time twice to > to stop sched timer reprogram sched timer I didn't find these two scenarios will program APIC timer twice separately instead of once separately, could you point out the codes? Regards, Wanpeng Li > 7. doing hlt 8. handle task and send ipi to > vcpu 0 > 9. same to 4. 10. same to 3 > > One transaction will introduce about 12 vmexits(2 hlt and 10 msr write). The > cost of such vmexits will degrades performance severely. Linux kernel > already provide idle=poll to mitigate the trend. But it only eliminates the > IPI and hlt vmexit. It has nothing to do with start/stop sched timer. A > compromise would be to turn off NOHZ kernel, but it is not the default > config for new distributions. Same for halt-poll in KVM, it only solve the > cost from schedule in/out in host and can not help such workload much. > > The purpose of this patch we want to improve current idle=poll mechanism to > use dynamic polling and do poll before touch sched timer. It should not be a > virtualization specific feature but seems bare mental have low cost to > access the MSR. So i want to only enable it in VM. Though the idea below the > patch may not so perfect to fit all conditions, it looks no worse than now. > How about we keep current implementation and i integrate the patch to > para-virtualize part as Paolo suggested? We can continue discuss it and i > will continue to refine it if anyone has better suggestions? > > > > -- > Yang > Alibaba Cloud Computing
Back to linux.kernel | Previous | Next | Find similar | Unroll thread
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Wanpeng Li <kernellwp@gmail.com> - 2017-07-05 00:40 +0200
csiph-web