Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1313255 > unrolled thread
| Started by | Michal Hocko <mhocko@kernel.org> |
|---|---|
| First post | 2016-01-20 15:40 +0100 |
| Last post | 2016-01-20 16:30 +0100 |
| Articles | 20 on this page of 50 — 6 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-20 15:40 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Sasha Levin <sasha.levin@oracle.com> - 2016-01-20 16:00 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-20 16:20 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-20 16:30 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Sasha Levin <sasha.levin@oracle.com> - 2016-01-20 17:00 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-20 17:00 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-20 22:30 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-20 23:00 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-21 09:30 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-21 16:50 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-21 18:00 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-21 18:40 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Shiraz Hashim <shiraz.linux.kernel@gmail.com> - 2016-01-22 12:10 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-22 15:10 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-22 17:10 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-22 17:20 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-22 17:50 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-22 18:20 +0100
fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-23 17:30 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-24 01:40 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-24 03:50 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-24 04:50 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-24 06:40 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Michal Hocko <mhocko@kernel.org> - 2016-01-25 18:50 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-25 19:10 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Michal Hocko <mhocko@kernel.org> - 2016-01-25 21:20 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 17:30 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 19:40 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 19:50 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 20:30 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-27 04:20 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-27 05:20 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-27 17:30 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 19:40 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 03:20 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 03:30 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 17:30 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 18:40 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 19:20 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 17:30 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 18:10 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 19:30 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 20:10 +0100
Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 20:30 +0100
[PATCH] mm, vmstat: make quiet_vmstat lighter (was: Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable) again and shut down on idle) Michal Hocko <mhocko@kernel.org> - 2016-01-27 17:50 +0100
Re: [PATCH] mm, vmstat: make quiet_vmstat lighter (was: Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable) again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-27 18:10 +0100
Re: [PATCH] mm, vmstat: make quiet_vmstat lighter (was: Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable) again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-27 19:30 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Linus Torvalds <torvalds@linux-foundation.org> - 2016-01-24 18:00 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-20 16:20 +0100
Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-20 16:30 +0100
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-01-24 03:50 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qUiUy-6x1-7@gated-at.bofh.it> |
| In reply to | #1315795 |
On Sat, 2016-01-23 at 18:33 -0600, Christoph Lameter wrote:
> On Sat, 23 Jan 2016, Mike Galbraith wrote:
>
> > While you're fixing that commit up, can you perhaps find a better home
> > for quiet_vmstat()? It not only munches cycles when switching cross
> > -core mightily, for -rt it injects a sleeping lock into the idle task.
>
> Not sure what you are talking about. No sleeping locks are used in
> quiet_vmstat() nor does it switch across cores. It would be broken if it
> would do so.
By switching cross-core, I'm referring to scheduling of communicating
tasks.
The perf top snippet...
12.89% [kernel] [k] refresh_cpu_vm_stats.isra.12
4.75% [kernel] [k] __schedule
4.70% [kernel] [k] mutex_unlock
3.14% [kernel] [k] __switch_to
... was pipe-test, an unrealistic microbench, but the same will happen
at lower frequency in the real world.
Here's the sleeping lock for -rt:
[ 2.279582] CPU: 1 PID: 0 Comm: swapper/1 Not tainted 4.5.0-rt3 #7
[ 2.280444] Hardware name: MEDION MS-7848/MS-7848, BIOS M7848W08.20C 09/23/2013
[ 2.281316] ffff88040b00d640 ffff88040b01fe10 ffffffff812d20e2 0000000000000000
[ 2.282202] ffff88040b01fe30 ffffffff81081095 ffff88041ec4cee0 ffff88041ec501e0
[ 2.283073] ffff88040b01fe48 ffffffff815ff910 ffff88041ec4cee0 ffff88040b01fe88
[ 2.283941] Call Trace:
[ 2.284797] [<ffffffff812d20e2>] dump_stack+0x49/0x67
[ 2.285658] [<ffffffff81081095>] ___might_sleep+0xf5/0x180
[ 2.286521] [<ffffffff815ff910>] rt_spin_lock+0x20/0x50
[ 2.287382] [<ffffffff81075919>] try_to_grab_pending+0x69/0x240
[ 2.288239] [<ffffffff81075b16>] cancel_delayed_work+0x26/0xe0
[ 2.289094] [<ffffffff8115ec05>] quiet_vmstat+0x75/0xa0
[ 2.289949] [<ffffffff8109ab38>] cpu_idle_loop+0x38/0x3e0
[ 2.290800] [<ffffffff8109aef3>] cpu_startup_entry+0x13/0x20
[ 2.291647] [<ffffffff81036164>] start_secondary+0x114/0x140
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-01-24 04:50 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qUjQB-7iY-3@gated-at.bofh.it> |
| In reply to | #1315814 |
On Sun, 24 Jan 2016, Mike Galbraith wrote: > By switching cross-core, I'm referring to scheduling of communicating > tasks. ??? Its cancelling a work request. That is a "communicating task"? > Here's the sleeping lock for -rt: > > [ 2.279582] CPU: 1 PID: 0 Comm: swapper/1 Not tainted 4.5.0-rt3 #7 > [ 2.280444] Hardware name: MEDION MS-7848/MS-7848, BIOS M7848W08.20C 09/23/2013 > [ 2.281316] ffff88040b00d640 ffff88040b01fe10 ffffffff812d20e2 0000000000000000 > [ 2.282202] ffff88040b01fe30 ffffffff81081095 ffff88041ec4cee0 ffff88041ec501e0 > [ 2.283073] ffff88040b01fe48 ffffffff815ff910 ffff88041ec4cee0 ffff88040b01fe88 > [ 2.283941] Call Trace: > [ 2.284797] [<ffffffff812d20e2>] dump_stack+0x49/0x67 > [ 2.285658] [<ffffffff81081095>] ___might_sleep+0xf5/0x180 > [ 2.286521] [<ffffffff815ff910>] rt_spin_lock+0x20/0x50 > [ 2.287382] [<ffffffff81075919>] try_to_grab_pending+0x69/0x240 > [ 2.288239] [<ffffffff81075b16>] cancel_delayed_work+0x26/0xe0 OMG cancelling a work request causes a sleeping lock to be taken?
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-01-24 06:40 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qUlz3-bV-7@gated-at.bofh.it> |
| In reply to | #1315818 |
On Sat, 2016-01-23 at 21:46 -0600, Christoph Lameter wrote: > On Sun, 24 Jan 2016, Mike Galbraith wrote: > > > By switching cross-core, I'm referring to scheduling of communicating > > tasks. > > ??? Its cancelling a work request. That is a "communicating task"? No no no, pipe-test is two tasks playing ping-pong via a pipe. They are about as synchronous as it gets, so each task goes idle at high frequency. Idle is fastpath. > > Here's the sleeping lock for -rt: > > > > [ 2.279582] CPU: 1 PID: 0 Comm: swapper/1 Not tainted 4.5.0-rt3 #7 > > [ 2.280444] Hardware name: MEDION MS-7848/MS-7848, BIOS M7848W08.20C 09/23/2013 > > [ 2.281316] ffff88040b00d640 ffff88040b01fe10 ffffffff812d20e2 0000000000000000 > > [ 2.282202] ffff88040b01fe30 ffffffff81081095 ffff88041ec4cee0 ffff88041ec501e0 > > [ 2.283073] ffff88040b01fe48 ffffffff815ff910 ffff88041ec4cee0 ffff88040b01fe88 > > [ 2.283941] Call Trace: > > [ 2.284797] [] dump_stack+0x49/0x67 > > [ 2.285658] [] ___might_sleep+0xf5/0x180 > > [ 2.286521] [] rt_spin_lock+0x20/0x50 > > [ 2.287382] [] try_to_grab_pending+0x69/0x240 > > [ 2.288239] [] cancel_delayed_work+0x26/0xe0 > > OMG cancelling a work request causes a sleeping lock to be taken? Yup. In -rt, spinlocks that are not raw are transformed into rtmutex. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-01-25 18:50 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qUTr3-7vm-11@gated-at.bofh.it> |
| In reply to | #1315680 |
On Sat 23-01-16 17:21:55, Mike Galbraith wrote:
> Hi Christoph,
>
> While you're fixing that commit up, can you perhaps find a better home
> for quiet_vmstat()? It not only munches cycles when switching cross
> -core mightily, for -rt it injects a sleeping lock into the idle task.
>
> 12.89% [kernel] [k] refresh_cpu_vm_stats.isra.12
> 4.75% [kernel] [k] __schedule
> 4.70% [kernel] [k] mutex_unlock
> 3.14% [kernel] [k] __switch_to
Hmm, I wouldn't have expected that refresh_cpu_vm_stats could have
such a large footprint. I guess this would be just an expensive noop
because we have to check all the zones*counters and do an expensive
this_cpu_xchg. Is the whole deferred thing worth this overhead?
0eb77e988032 ("vmstat: make vmstat_updater deferrable again and
shut down on idle") doesn't talk about any numbers and neither does
39bf6270f524 ("VM statistics: Make timer deferrable").
But even when refresh_cpu_vm_stats is made more effective there would
still be the problem for RT and the work canceling, though. We would
need much more changes to make this RT ready (both timer and WQ code
would need some changes AFAIU).
Unless there is a clear and huge win from doing the vmstat update
deferrable then I think a revert is more appropriate IMHO.
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-01-25 19:10 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qUTKq-7UQ-17@gated-at.bofh.it> |
| In reply to | #1317126 |
On Mon, 25 Jan 2016, Michal Hocko wrote: > On Sat 23-01-16 17:21:55, Mike Galbraith wrote: > > Hi Christoph, > > > > While you're fixing that commit up, can you perhaps find a better home > > for quiet_vmstat()? It not only munches cycles when switching cross > > -core mightily, for -rt it injects a sleeping lock into the idle task. > > > > 12.89% [kernel] [k] refresh_cpu_vm_stats.isra.12 > > 4.75% [kernel] [k] __schedule > > 4.70% [kernel] [k] mutex_unlock > > 3.14% [kernel] [k] __switch_to > > Hmm, I wouldn't have expected that refresh_cpu_vm_stats could have > such a large footprint. I guess this would be just an expensive noop > because we have to check all the zones*counters and do an expensive > this_cpu_xchg. Is the whole deferred thing worth this overhead? Why would the deferring cause this overhead? Also there is no cross core activity from quiet_vmstat(). It simply disables the local vmstat updates. > Unless there is a clear and huge win from doing the vmstat update > deferrable then I think a revert is more appropriate IMHO. It reduces the OS events that the application experiences by folding it into the tick events. If its not deferrable then a timer event will be generated in addition to the tick. We do not want that. Workqueues are used in many places. If RT can sleep within workqueue management functions then spinlocks cannot be taken anymore and there may be issues with preemption. The regression that I know of (independent of "RT") is due as far as I know due to the switch of the parameters of some vmstat functions to 64 bit instead of 32 bit.
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-01-25 21:20 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qUVMf-PD-23@gated-at.bofh.it> |
| In reply to | #1317134 |
On Mon 25-01-16 12:02:06, Christoph Lameter wrote: > On Mon, 25 Jan 2016, Michal Hocko wrote: > > > On Sat 23-01-16 17:21:55, Mike Galbraith wrote: > > > Hi Christoph, > > > > > > While you're fixing that commit up, can you perhaps find a better home > > > for quiet_vmstat()? It not only munches cycles when switching cross > > > -core mightily, for -rt it injects a sleeping lock into the idle task. > > > > > > 12.89% [kernel] [k] refresh_cpu_vm_stats.isra.12 > > > 4.75% [kernel] [k] __schedule > > > 4.70% [kernel] [k] mutex_unlock > > > 3.14% [kernel] [k] __switch_to > > > > Hmm, I wouldn't have expected that refresh_cpu_vm_stats could have > > such a large footprint. I guess this would be just an expensive noop > > because we have to check all the zones*counters and do an expensive > > this_cpu_xchg. Is the whole deferred thing worth this overhead? > > Why would the deferring cause this overhead? I guess the profile speaks for itself, doesn't it? > Also there is no cross core activity from quiet_vmstat(). It simply > disables the local vmstat updates. It doesn't go cross core but it still does nr_zones * counters atomic ops. > > Unless there is a clear and huge win from doing the vmstat update > > deferrable then I think a revert is more appropriate IMHO. > > It reduces the OS events that the application experiences by folding it > into the tick events. If its not deferrable then a timer event will be > generated in addition to the tick. We do not want that. Yes this is what I have read in the changelog. But "how much" part is really missing. Is this even quantifiable? > Workqueues are used in many places. If RT can sleep within workqueue > management functions then spinlocks cannot be taken anymore and there may > be issues with preemption. RT can sleep in _any_ spinlock except for raw spin locks. Even though the !RT kernel is not sleeping doesn't really matter much because cancel_delayed_work is quite a heavy function which shouldn't be called from the idle context AFAIU. Sure most of the time it will boil down to del_timer but it can hit the slowpath as well if the timer got migrated to a different CPU and we have to race with the WQ pool management IIUC. Maybe this overhead can be reduced by outsourcing the functionality to vmstat_shepherd which can check idle CPUs, cancel the timer for them update the differentials and put them to cpu_stat_off? > The regression that I know of (independent of "RT") is due as far as I > know due to the switch of the parameters of some vmstat functions to 64 > bit instead of 32 bit. I am not sure I am following. -- Michal Hocko SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-01-26 17:30 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVeFd-6F7-19@gated-at.bofh.it> |
| In reply to | #1317274 |
On Mon, 25 Jan 2016, Michal Hocko wrote: > > Why would the deferring cause this overhead? > > I guess the profile speaks for itself, doesn't it? But the system is going idle? Why would this impact performance? > > Also there is no cross core activity from quiet_vmstat(). It simply > > disables the local vmstat updates. > > It doesn't go cross core but it still does nr_zones * counters atomic > ops. If there are updates then yes. > > It reduces the OS events that the application experiences by folding it > > into the tick events. If its not deferrable then a timer event will be > > generated in addition to the tick. We do not want that. > > Yes this is what I have read in the changelog. But "how much" part is > really missing. Is this even quantifiable? Oh yes. If you want to have low latency responses yes. And event like that causes a couple of microseconds delay which will cause a moneytary impact if you have to react to stock trading events. > Maybe this overhead can be reduced by outsourcing the functionality to > vmstat_shepherd which can check idle CPUs, cancel the timer for them > update the differentials and put them to cpu_stat_off? Remote updating of differentials is problematic due to the counters being per cpu values that are expected to only be updated from the cpu that "owns" them. > > The regression that I know of (independent of "RT") is due as far as I > > know due to the switch of the parameters of some vmstat functions to 64 > > bit instead of 32 bit. > > I am not sure I am following. An additional patch was merged for 4.5 that increases the arguments to the counter operations to 64 bit which is known for regressions.
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-01-26 19:40 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVgH0-88L-13@gated-at.bofh.it> |
| In reply to | #1318119 |
On Tue, 26 Jan 2016, Mike Galbraith wrote: > I disagree. You're burning electrons for no benefit at all to me on my > box. You want to do high speed trading, that's fine, but I expect my > box to be able to pop in and out of idle without having to pay a toll > to the high speed trading bandits of the world, thank you very much. > > This specialty thing does not belong in the generic fast path. The system going idle is a fastpath. Mind boogling.
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-01-26 19:50 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVgQG-8cd-9@gated-at.bofh.it> |
| In reply to | #1318230 |
On Tue, 2016-01-26 at 12:34 -0600, Christoph Lameter wrote: > On Tue, 26 Jan 2016, Mike Galbraith wrote: > > > I disagree. You're burning electrons for no benefit at all to me on my > > box. You want to do high speed trading, that's fine, but I expect my > > box to be able to pop in and out of idle without having to pay a toll > > to the high speed trading bandits of the world, thank you very much. > > > > This specialty thing does not belong in the generic fast path. > > The system going idle is a fastpath. Mind boogling. Hohum, noted. Now what about those cycles, and the sleeping lock you injected for -rt? -Mike
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-01-26 20:30 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVhtp-gQ-39@gated-at.bofh.it> |
| In reply to | #1318250 |
On Tue, 26 Jan 2016, Mike Galbraith wrote:
> > The system going idle is a fastpath. Mind boogling.
>
> Hohum, noted. Now what about those cycles, and the sleeping lock you
> injected for -rt?
Since we (the NOHZ people) care mostly about NOHZ then lets restrict
that to the NOHZ mode. Then it should not affect your load.
Subject: Move quiet_vmstat() to NOHZ code
quiet_vmstat() seems to cause regressions for some load because
the cpu going idle is a "fastpath". Mind boogling. Strange claim.
If the system goes idle then it has nothing to do after all
But anyways if we shift the quiet_vmstat() into the NOHZ logic
when it stops the tick then it will only affect those cores that
are setup for NOHZ mode. That is where we want this processing
after all to ensure that the OS keeps itself off those cores.
Signed-off-by: Christoph Lameter <cl@linux.com>
Index: linux/kernel/sched/idle.c
===================================================================
--- linux.orig/kernel/sched/idle.c
+++ linux/kernel/sched/idle.c
@@ -213,7 +213,6 @@ static void cpu_idle_loop(void)
*/
__current_set_polling();
- quiet_vmstat();
tick_nohz_idle_enter();
while (!need_resched()) {
Index: linux/kernel/time/tick-sched.c
===================================================================
--- linux.orig/kernel/time/tick-sched.c
+++ linux/kernel/time/tick-sched.c
@@ -811,6 +811,7 @@ static void __tick_nohz_idle_enter(struc
ts->idle_calls++;
expires = tick_nohz_stop_sched_tick(ts, now, cpu);
+ quiet_vmstat();
if (expires.tv64 > 0LL) {
ts->idle_sleeps++;
ts->idle_expires = expires;
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-01-27 04:20 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVoOe-5Ga-7@gated-at.bofh.it> |
| In reply to | #1318288 |
Good morning,
On Tue, 2016-01-26 at 13:20 -0600, Christoph Lameter wrote:
> On Tue, 26 Jan 2016, Mike Galbraith wrote:
>
> > > The system going idle is a fastpath. Mind boogling.
> >
> > Hohum, noted. Now what about those cycles, and the sleeping lock you
> > injected for -rt?
>
> Since we (the NOHZ people) care mostly about NOHZ then lets restrict
> that to the NOHZ mode. Then it should not affect your load.
Tons of folks do have NO_HZ enabled (including me). Isn't there a spot
somewhere in NO_HZ_FULL code where it can take up residence? (one with
a tad lower maximum call frequency would be good, a nohz_full cpu isn't
necessarily being used for pure compute exclusively)
> Subject: Move quiet_vmstat() to NOHZ code
>
> quiet_vmstat() seems to cause regressions for some load because
> the cpu going idle is a "fastpath". Mind boogling. Strange claim.
> If the system goes idle then it has nothing to do after all
>
> But anyways if we shift the quiet_vmstat() into the NOHZ logic
> when it stops the tick then it will only affect those cores that
> are setup for NOHZ mode. That is where we want this processing
> after all to ensure that the OS keeps itself off those cores.
>
> Signed-off-by: Christoph Lameter <cl@linux.com>
>
>
> Index: linux/kernel/sched/idle.c
> ===================================================================
> --- linux.orig/kernel/sched/idle.c
> +++ linux/kernel/sched/idle.c
> @@ -213,7 +213,6 @@ static void cpu_idle_loop(void)
> */
>
> __current_set_polling();
> - quiet_vmstat();
> tick_nohz_idle_enter();
>
> while (!need_resched()) {
> Index: linux/kernel/time/tick-sched.c
> ===================================================================
> --- linux.orig/kernel/time/tick-sched.c
> +++ linux/kernel/time/tick-sched.c
> @@ -811,6 +811,7 @@ static void __tick_nohz_idle_enter(struc
> ts->idle_calls++;
>
> expires = tick_nohz_stop_sched_tick(ts, now, cpu);
> + quiet_vmstat();
> if (expires.tv64 > 0LL) {
> ts->idle_sleeps++;
> ts->idle_expires = expires;
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-01-27 05:20 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVpKi-6jb-1@gated-at.bofh.it> |
| In reply to | #1318578 |
On Wed, 2016-01-27 at 04:12 +0100, Mike Galbraith wrote: > Good morning, > > On Tue, 2016-01-26 at 13:20 -0600, Christoph Lameter wrote: > > On Tue, 26 Jan 2016, Mike Galbraith wrote: > > > > > > The system going idle is a fastpath. Mind boogling. > > > > > > Hohum, noted. Now what about those cycles, and the sleeping lock you > > > injected for -rt? > > > > Since we (the NOHZ people) care mostly about NOHZ then lets restrict > > that to the NOHZ mode. Then it should not affect your load. > > Tons of folks do have NO_HZ enabled (including me). Isn't there a spot > somewhere in NO_HZ_FULL code where it can take up residence? (one with > a tad lower maximum call frequency would be good, a nohz_full cpu isn't > necessarily being used for pure compute exclusively) I forgot to mention that the spot you picked is called with irqs disabled. > > Subject: Move quiet_vmstat() to NOHZ code > > > > quiet_vmstat() seems to cause regressions for some load because > > the cpu going idle is a "fastpath". Mind boogling. Strange claim. > > If the system goes idle then it has nothing to do after all Hm, seems I also forgot to say "Hohum, noted..." again. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-01-27 17:30 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVB8L-6ft-25@gated-at.bofh.it> |
| In reply to | #1318578 |
On Wed, 27 Jan 2016, Mike Galbraith wrote:
> > Since we (the NOHZ people) care mostly about NOHZ then lets restrict
> > that to the NOHZ mode. Then it should not affect your load.
>
> Tons of folks do have NO_HZ enabled (including me). Isn't there a spot
> somewhere in NO_HZ_FULL code where it can take up residence? (one with
> a tad lower maximum call frequency would be good, a nohz_full cpu isn't
> necessarily being used for pure compute exclusively)
Frederic, any idea on where to put quiet_vmstat()?
> > Subject: Move quiet_vmstat() to NOHZ code
> >
> > quiet_vmstat() seems to cause regressions for some load because
> > the cpu going idle is a "fastpath". Mind boogling. Strange claim.
> > If the system goes idle then it has nothing to do after all
> >
> > But anyways if we shift the quiet_vmstat() into the NOHZ logic
> > when it stops the tick then it will only affect those cores that
> > are setup for NOHZ mode. That is where we want this processing
> > after all to ensure that the OS keeps itself off those cores.
> >
> > Signed-off-by: Christoph Lameter <cl@linux.com>
> >
> >
> > Index: linux/kernel/sched/idle.c
> > ===================================================================
> > --- linux.orig/kernel/sched/idle.c
> > +++ linux/kernel/sched/idle.c
> > @@ -213,7 +213,6 @@ static void cpu_idle_loop(void)
> > */
> >
> > __current_set_polling();
> > - quiet_vmstat();
> > tick_nohz_idle_enter();
> >
> > while (!need_resched()) {
> > Index: linux/kernel/time/tick-sched.c
> > ===================================================================
> > --- linux.orig/kernel/time/tick-sched.c
> > +++ linux/kernel/time/tick-sched.c
> > @@ -811,6 +811,7 @@ static void __tick_nohz_idle_enter(struc
> > ts->idle_calls++;
> >
> > expires = tick_nohz_stop_sched_tick(ts, now, cpu);
> > + quiet_vmstat();
> > if (expires.tv64 > 0LL) {
> > ts->idle_sleeps++;
> > ts->idle_expires = expires;
>
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-01-26 19:40 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVgH0-88L-15@gated-at.bofh.it> |
| In reply to | #1318119 |
On Tue, 2016-01-26 at 10:25 -0600, Christoph Lameter wrote: > On Mon, 25 Jan 2016, Michal Hocko wrote: > > > > Why would the deferring cause this overhead? > > > > I guess the profile speaks for itself, doesn't it? > > But the system is going idle? Why would this impact performance? We enter/exit idle a lot. Your reluctance to move it seem to suggest that 99.99% of CPUs on the planet chewing up cycles (measured) doing what for most is useless work on every micro-idle is a perfectly fine price to pay to ensure that .01% (or whatever tiny minority) get what they want. I disagree. You're burning electrons for no benefit at all to me on my box. You want to do high speed trading, that's fine, but I expect my box to be able to pop in and out of idle without having to pay a toll to the high speed trading bandits of the world, thank you very much. This specialty thing does not belong in the generic fast path. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-01-26 03:20 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qV1oB-4JB-1@gated-at.bofh.it> |
| In reply to | #1317134 |
On Mon, 2016-01-25 at 12:02 -0600, Christoph Lameter wrote: > On Mon, 25 Jan 2016, Michal Hocko wrote: > > > On Sat 23-01-16 17:21:55, Mike Galbraith wrote: > > > Hi Christoph, > > > > > > While you're fixing that commit up, can you perhaps find a better home > > > for quiet_vmstat()? It not only munches cycles when switching cross > > > -core mightily, for -rt it injects a sleeping lock into the idle task. > > > > > > 12.89% [kernel] [k] refresh_cpu_vm_stats.isra.12 > > > 4.75% [kernel] [k] __schedule > > > 4.70% [kernel] [k] mutex_unlock > > > 3.14% [kernel] [k] __switch_to > > > > Hmm, I wouldn't have expected that refresh_cpu_vm_stats could have > > such a large footprint. I guess this would be just an expensive noop > > because we have to check all the zones*counters and do an expensive > > this_cpu_xchg. Is the whole deferred thing worth this overhead? > > Why would the deferring cause this overhead? Because we schedule to idle cores aggressively, thus we may pop in and out of idle at high frequency. > Also there is no cross core activity from quiet_vmstat(). It simply > disables the local vmstat updates. Again, the cross core activity is not due to quiet_vmstat(), it is due to pipe-test threads running on two cores, and meeting quiet_vmstat() at high frequency. > > Unless there is a clear and huge win from doing the vmstat update > > deferrable then I think a revert is more appropriate IMHO. > > It reduces the OS events that the application experiences by folding it > into the tick events. If its not deferrable then a timer event will be > generated in addition to the tick. We do not want that. Perf and RT say we don't want quiet_vmstat() in the idle loop either. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-01-26 03:30 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qV1yj-4Nn-19@gated-at.bofh.it> |
| In reply to | #1317461 |
On Tue, 2016-01-26 at 03:14 +0100, Mike Galbraith wrote: > Perf and RT say we don't want quiet_vmstat() in the idle loop either. BTW, the perf numbers were not from an RT kernel, they were from my PREEMPT_VOLUNTARY desktop kernel. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-01-26 17:30 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVeFc-6F7-13@gated-at.bofh.it> |
| In reply to | #1317472 |
On Tue, 26 Jan 2016, Mike Galbraith wrote: > On Tue, 2016-01-26 at 03:14 +0100, Mike Galbraith wrote: > > > Perf and RT say we don't want quiet_vmstat() in the idle loop either. > > BTW, the perf numbers were not from an RT kernel, they were from my > PREEMPT_VOLUNTARY desktop kernel. Can we move quiet_vmstat() elsewhere after we have checked that really nothing else is going on soon?
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-01-26 18:40 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVfKX-7rW-23@gated-at.bofh.it> |
| In reply to | #1318116 |
On Tue, 2016-01-26 at 10:26 -0600, Christoph Lameter wrote: > On Tue, 26 Jan 2016, Mike Galbraith wrote: > > > On Tue, 2016-01-26 at 03:14 +0100, Mike Galbraith wrote: > > > > > Perf and RT say we don't want quiet_vmstat() in the idle loop > > > either. > > > > BTW, the perf numbers were not from an RT kernel, they were from my > > PREEMPT_VOLUNTARY desktop kernel. > > Can we move quiet_vmstat() elsewhere after we have checked that really > nothing else is going on soon? How would you check? Precognition doesn't work for mortals. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-01-26 19:20 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVgnD-7Zo-7@gated-at.bofh.it> |
| In reply to | #1318203 |
On Tue, 26 Jan 2016, Mike Galbraith wrote: > On Tue, 2016-01-26 at 10:26 -0600, Christoph Lameter wrote: > > On Tue, 26 Jan 2016, Mike Galbraith wrote: > > > > > On Tue, 2016-01-26 at 03:14 +0100, Mike Galbraith wrote: > > > > > > > Perf and RT say we don't want quiet_vmstat() in the idle loop > > > > either. > > > > > > BTW, the perf numbers were not from an RT kernel, they were from my > > > PREEMPT_VOLUNTARY desktop kernel. > > > > Can we move quiet_vmstat() elsewhere after we have checked that really > > nothing else is going on soon? > > How would you check? Precognition doesn't work for mortals. Dont we have some decision mechanism to go into higher levels of power savings when the system is idle for longer times?
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-01-26 17:30 +0100 |
| Subject | Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle) |
| Message-ID | <qVeFd-6F7-21@gated-at.bofh.it> |
| In reply to | #1317461 |
On Tue, 26 Jan 2016, Mike Galbraith wrote: > > Why would the deferring cause this overhead? > > Because we schedule to idle cores aggressively, thus we may pop in and > out of idle at high frequency. Whats the point of going idle if you have things to do soon?
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web