Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1313255 > unrolled thread

Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408!

Started byMichal Hocko <mhocko@kernel.org>
First post2016-01-20 15:40 +0100
Last post2016-01-20 16:30 +0100
Articles 20 on this page of 50 — 6 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-20 15:40 +0100
    Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Sasha Levin <sasha.levin@oracle.com> - 2016-01-20 16:00 +0100
      Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-20 16:20 +0100
        Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-20 16:30 +0100
          Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Sasha Levin <sasha.levin@oracle.com> - 2016-01-20 17:00 +0100
            Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-20 17:00 +0100
              Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-20 22:30 +0100
                Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-20 23:00 +0100
                  Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-21 09:30 +0100
                    Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-21 16:50 +0100
                      Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-21 18:00 +0100
                        Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-21 18:40 +0100
                          Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Shiraz Hashim <shiraz.linux.kernel@gmail.com> - 2016-01-22 12:10 +0100
                          Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-22 15:10 +0100
                            Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-22 17:10 +0100
                              Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-22 17:20 +0100
                                Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-22 17:50 +0100
                                  Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-22 18:20 +0100
                                  fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-23 17:30 +0100
                                    Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-24 01:40 +0100
                                      Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-24 03:50 +0100
                                        Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-24 04:50 +0100
                                          Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-24 06:40 +0100
                                    Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Michal Hocko <mhocko@kernel.org> - 2016-01-25 18:50 +0100
                                      Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-25 19:10 +0100
                                        Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Michal Hocko <mhocko@kernel.org> - 2016-01-25 21:20 +0100
                                          Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 17:30 +0100
                                            Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 19:40 +0100
                                              Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 19:50 +0100
                                                Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 20:30 +0100
                                                  Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-27 04:20 +0100
                                                    Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-27 05:20 +0100
                                                    Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-27 17:30 +0100
                                            Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 19:40 +0100
                                        Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 03:20 +0100
                                          Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 03:30 +0100
                                            Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 17:30 +0100
                                              Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 18:40 +0100
                                                Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 19:20 +0100
                                          Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 17:30 +0100
                                            Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 18:10 +0100
                                              Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 19:30 +0100
                                                Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-26 20:10 +0100
                                                  Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable  again and shut down on idle) Christoph Lameter <cl@linux.com> - 2016-01-26 20:30 +0100
                                    [PATCH] mm, vmstat: make quiet_vmstat lighter (was: Re: fast path  cycle muncher (vmstat: make vmstat_updater deferrable) again and shut down  on idle) Michal Hocko <mhocko@kernel.org> - 2016-01-27 17:50 +0100
                                      Re: [PATCH] mm, vmstat: make quiet_vmstat lighter (was: Re: fast  path cycle muncher (vmstat: make vmstat_updater deferrable) again and shut  down on idle) Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-01-27 18:10 +0100
                                      Re: [PATCH] mm, vmstat: make quiet_vmstat lighter (was: Re: fast  path cycle muncher (vmstat: make vmstat_updater deferrable) again and shut  down on idle) Christoph Lameter <cl@linux.com> - 2016-01-27 19:30 +0100
                                  Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Linus Torvalds <torvalds@linux-foundation.org> - 2016-01-24 18:00 +0100
    Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Christoph Lameter <cl@linux.com> - 2016-01-20 16:20 +0100
      Re: mm, vmstat: kernel BUG at mm/vmstat.c:1408! Michal Hocko <mhocko@kernel.org> - 2016-01-20 16:30 +0100

Page 2 of 3 — ← Prev page 1 [2] 3  Next page →


#1315814 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-01-24 03:50 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qUiUy-6x1-7@gated-at.bofh.it>
In reply to#1315795
On Sat, 2016-01-23 at 18:33 -0600, Christoph Lameter wrote:
> On Sat, 23 Jan 2016, Mike Galbraith wrote:
> 
> > While you're fixing that commit up, can you perhaps find a better home
> > for quiet_vmstat()?  It not only munches cycles when switching cross
> > -core mightily, for -rt it injects a sleeping lock into the idle task.
> 
> Not sure what you are talking about. No sleeping locks are used in
> quiet_vmstat() nor does it switch across cores. It would be broken if it
> would do so.

By switching cross-core, I'm referring to scheduling of communicating
tasks.

The perf top snippet...

    12.89%  [kernel]       [k] refresh_cpu_vm_stats.isra.12
     4.75%  [kernel]       [k] __schedule                  
     4.70%  [kernel]       [k] mutex_unlock                
     3.14%  [kernel]       [k] __switch_to

... was pipe-test, an unrealistic microbench, but the same will happen
at lower frequency in the real world.

Here's the sleeping lock for -rt:

[    2.279582] CPU: 1 PID: 0 Comm: swapper/1 Not tainted 4.5.0-rt3 #7
[    2.280444] Hardware name: MEDION MS-7848/MS-7848, BIOS M7848W08.20C 09/23/2013
[    2.281316]  ffff88040b00d640 ffff88040b01fe10 ffffffff812d20e2 0000000000000000
[    2.282202]  ffff88040b01fe30 ffffffff81081095 ffff88041ec4cee0 ffff88041ec501e0
[    2.283073]  ffff88040b01fe48 ffffffff815ff910 ffff88041ec4cee0 ffff88040b01fe88
[    2.283941] Call Trace:
[    2.284797]  [<ffffffff812d20e2>] dump_stack+0x49/0x67
[    2.285658]  [<ffffffff81081095>] ___might_sleep+0xf5/0x180
[    2.286521]  [<ffffffff815ff910>] rt_spin_lock+0x20/0x50
[    2.287382]  [<ffffffff81075919>] try_to_grab_pending+0x69/0x240
[    2.288239]  [<ffffffff81075b16>] cancel_delayed_work+0x26/0xe0
[    2.289094]  [<ffffffff8115ec05>] quiet_vmstat+0x75/0xa0
[    2.289949]  [<ffffffff8109ab38>] cpu_idle_loop+0x38/0x3e0
[    2.290800]  [<ffffffff8109aef3>] cpu_startup_entry+0x13/0x20
[    2.291647]  [<ffffffff81036164>] start_secondary+0x114/0x140

[toc] | [prev] | [next] | [standalone]


#1315818 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromChristoph Lameter <cl@linux.com>
Date2016-01-24 04:50 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qUjQB-7iY-3@gated-at.bofh.it>
In reply to#1315814
On Sun, 24 Jan 2016, Mike Galbraith wrote:

> By switching cross-core, I'm referring to scheduling of communicating
> tasks.

??? Its cancelling a work request. That is a "communicating task"?

> Here's the sleeping lock for -rt:
>
> [    2.279582] CPU: 1 PID: 0 Comm: swapper/1 Not tainted 4.5.0-rt3 #7
> [    2.280444] Hardware name: MEDION MS-7848/MS-7848, BIOS M7848W08.20C 09/23/2013
> [    2.281316]  ffff88040b00d640 ffff88040b01fe10 ffffffff812d20e2 0000000000000000
> [    2.282202]  ffff88040b01fe30 ffffffff81081095 ffff88041ec4cee0 ffff88041ec501e0
> [    2.283073]  ffff88040b01fe48 ffffffff815ff910 ffff88041ec4cee0 ffff88040b01fe88
> [    2.283941] Call Trace:
> [    2.284797]  [<ffffffff812d20e2>] dump_stack+0x49/0x67
> [    2.285658]  [<ffffffff81081095>] ___might_sleep+0xf5/0x180
> [    2.286521]  [<ffffffff815ff910>] rt_spin_lock+0x20/0x50
> [    2.287382]  [<ffffffff81075919>] try_to_grab_pending+0x69/0x240
> [    2.288239]  [<ffffffff81075b16>] cancel_delayed_work+0x26/0xe0

OMG cancelling a work request causes a sleeping lock to be taken?

[toc] | [prev] | [next] | [standalone]


#1315828 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-01-24 06:40 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qUlz3-bV-7@gated-at.bofh.it>
In reply to#1315818
On Sat, 2016-01-23 at 21:46 -0600, Christoph Lameter wrote:
> On Sun, 24 Jan 2016, Mike Galbraith wrote:
> 
> > By switching cross-core, I'm referring to scheduling of communicating
> > tasks.
> 
> ??? Its cancelling a work request. That is a "communicating task"?

No no no, pipe-test is two tasks playing ping-pong via a pipe.  They
are about as synchronous as it gets, so each task goes idle at high
frequency.  Idle is fastpath.

> > Here's the sleeping lock for -rt:
> > 
> > [    2.279582] CPU: 1 PID: 0 Comm: swapper/1 Not tainted 4.5.0-rt3 #7
> > [    2.280444] Hardware name: MEDION MS-7848/MS-7848, BIOS M7848W08.20C 09/23/2013
> > [    2.281316]  ffff88040b00d640 ffff88040b01fe10 ffffffff812d20e2 0000000000000000
> > [    2.282202]  ffff88040b01fe30 ffffffff81081095 ffff88041ec4cee0 ffff88041ec501e0
> > [    2.283073]  ffff88040b01fe48 ffffffff815ff910 ffff88041ec4cee0 ffff88040b01fe88
> > [    2.283941] Call Trace:
> > [    2.284797]  [] dump_stack+0x49/0x67
> > [    2.285658]  [] ___might_sleep+0xf5/0x180
> > [    2.286521]  [] rt_spin_lock+0x20/0x50
> > [    2.287382]  [] try_to_grab_pending+0x69/0x240
> > [    2.288239]  [] cancel_delayed_work+0x26/0xe0
> 
> OMG cancelling a work request causes a sleeping lock to be taken?

Yup.  In -rt, spinlocks that are not raw are transformed into rtmutex.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1317126 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromMichal Hocko <mhocko@kernel.org>
Date2016-01-25 18:50 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qUTr3-7vm-11@gated-at.bofh.it>
In reply to#1315680
On Sat 23-01-16 17:21:55, Mike Galbraith wrote:
> Hi Christoph,
> 
> While you're fixing that commit up, can you perhaps find a better home
> for quiet_vmstat()?  It not only munches cycles when switching cross
> -core mightily, for -rt it injects a sleeping lock into the idle task.
> 
>     12.89%  [kernel]       [k] refresh_cpu_vm_stats.isra.12
>      4.75%  [kernel]       [k] __schedule                  
>      4.70%  [kernel]       [k] mutex_unlock                
>      3.14%  [kernel]       [k] __switch_to                 

Hmm, I wouldn't have expected that refresh_cpu_vm_stats could have
such a large footprint. I guess this would be just an expensive noop
because we have to check all the zones*counters and do an expensive
this_cpu_xchg. Is the whole deferred thing worth this overhead?

0eb77e988032 ("vmstat: make vmstat_updater deferrable again and
shut down on idle") doesn't talk about any numbers and neither does
39bf6270f524 ("VM statistics: Make timer deferrable").

But even when refresh_cpu_vm_stats is made more effective there would
still be the problem for RT and the work canceling, though. We would
need much more changes to make this RT ready (both timer and WQ code
would need some changes AFAIU).

Unless there is a clear and huge win from doing the vmstat update
deferrable then I think a revert is more appropriate IMHO.
-- 
Michal Hocko
SUSE Labs 

[toc] | [prev] | [next] | [standalone]


#1317134 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromChristoph Lameter <cl@linux.com>
Date2016-01-25 19:10 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qUTKq-7UQ-17@gated-at.bofh.it>
In reply to#1317126
On Mon, 25 Jan 2016, Michal Hocko wrote:

> On Sat 23-01-16 17:21:55, Mike Galbraith wrote:
> > Hi Christoph,
> >
> > While you're fixing that commit up, can you perhaps find a better home
> > for quiet_vmstat()?  It not only munches cycles when switching cross
> > -core mightily, for -rt it injects a sleeping lock into the idle task.
> >
> >     12.89%  [kernel]       [k] refresh_cpu_vm_stats.isra.12
> >      4.75%  [kernel]       [k] __schedule
> >      4.70%  [kernel]       [k] mutex_unlock
> >      3.14%  [kernel]       [k] __switch_to
>
> Hmm, I wouldn't have expected that refresh_cpu_vm_stats could have
> such a large footprint. I guess this would be just an expensive noop
> because we have to check all the zones*counters and do an expensive
> this_cpu_xchg. Is the whole deferred thing worth this overhead?

Why would the deferring cause this overhead?

Also there is no cross core activity from quiet_vmstat(). It simply
disables the local vmstat updates.

> Unless there is a clear and huge win from doing the vmstat update
> deferrable then I think a revert is more appropriate IMHO.

It reduces the OS events that the application experiences by folding it
into the tick events. If its not deferrable then a timer event will be
generated in addition to the tick. We do not want that.

Workqueues are used in many places. If RT can sleep within workqueue
management functions then spinlocks cannot be taken anymore and there may
be issues with preemption.

The regression that I know of (independent of "RT") is due as far as I
know due to the switch of the parameters of some vmstat functions to 64
bit instead of 32 bit.

[toc] | [prev] | [next] | [standalone]


#1317274 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromMichal Hocko <mhocko@kernel.org>
Date2016-01-25 21:20 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qUVMf-PD-23@gated-at.bofh.it>
In reply to#1317134
On Mon 25-01-16 12:02:06, Christoph Lameter wrote:
> On Mon, 25 Jan 2016, Michal Hocko wrote:
> 
> > On Sat 23-01-16 17:21:55, Mike Galbraith wrote:
> > > Hi Christoph,
> > >
> > > While you're fixing that commit up, can you perhaps find a better home
> > > for quiet_vmstat()?  It not only munches cycles when switching cross
> > > -core mightily, for -rt it injects a sleeping lock into the idle task.
> > >
> > >     12.89%  [kernel]       [k] refresh_cpu_vm_stats.isra.12
> > >      4.75%  [kernel]       [k] __schedule
> > >      4.70%  [kernel]       [k] mutex_unlock
> > >      3.14%  [kernel]       [k] __switch_to
> >
> > Hmm, I wouldn't have expected that refresh_cpu_vm_stats could have
> > such a large footprint. I guess this would be just an expensive noop
> > because we have to check all the zones*counters and do an expensive
> > this_cpu_xchg. Is the whole deferred thing worth this overhead?
> 
> Why would the deferring cause this overhead?

I guess the profile speaks for itself, doesn't it?

> Also there is no cross core activity from quiet_vmstat(). It simply
> disables the local vmstat updates.

It doesn't go cross core but it still does nr_zones * counters atomic
ops.

> > Unless there is a clear and huge win from doing the vmstat update
> > deferrable then I think a revert is more appropriate IMHO.
> 
> It reduces the OS events that the application experiences by folding it
> into the tick events. If its not deferrable then a timer event will be
> generated in addition to the tick. We do not want that.

Yes this is what I have read in the changelog. But "how much" part is
really missing. Is this even quantifiable?

> Workqueues are used in many places. If RT can sleep within workqueue
> management functions then spinlocks cannot be taken anymore and there may
> be issues with preemption.

RT can sleep in _any_ spinlock except for raw spin locks. Even though
the !RT kernel is not sleeping doesn't really matter much because
cancel_delayed_work is quite a heavy function which shouldn't be called
from the idle context AFAIU. Sure most of the time it will boil down to
del_timer but it can hit the slowpath as well if the timer got migrated
to a different CPU and we have to race with the WQ pool management IIUC.

Maybe this overhead can be reduced by outsourcing the functionality to
vmstat_shepherd which can check idle CPUs, cancel the timer for them
update the differentials and put them to cpu_stat_off? 

> The regression that I know of (independent of "RT") is due as far as I
> know due to the switch of the parameters of some vmstat functions to 64
> bit instead of 32 bit.

I am not sure I am following.

-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1318119 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromChristoph Lameter <cl@linux.com>
Date2016-01-26 17:30 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVeFd-6F7-19@gated-at.bofh.it>
In reply to#1317274
On Mon, 25 Jan 2016, Michal Hocko wrote:

> > Why would the deferring cause this overhead?
>
> I guess the profile speaks for itself, doesn't it?

But the system is going idle? Why would this impact performance?

> > Also there is no cross core activity from quiet_vmstat(). It simply
> > disables the local vmstat updates.
>
> It doesn't go cross core but it still does nr_zones * counters atomic
> ops.

If there are updates then yes.

> > It reduces the OS events that the application experiences by folding it
> > into the tick events. If its not deferrable then a timer event will be
> > generated in addition to the tick. We do not want that.
>
> Yes this is what I have read in the changelog. But "how much" part is
> really missing. Is this even quantifiable?

Oh yes. If you want to have low latency responses yes. And event like that
causes a couple of microseconds delay which will cause a moneytary impact
if you have to react to stock trading events.

> Maybe this overhead can be reduced by outsourcing the functionality to
> vmstat_shepherd which can check idle CPUs, cancel the timer for them
> update the differentials and put them to cpu_stat_off?

Remote updating of differentials is problematic due to the counters being
per cpu values that are expected to only be updated from the cpu that
"owns" them.

> > The regression that I know of (independent of "RT") is due as far as I
> > know due to the switch of the parameters of some vmstat functions to 64
> > bit instead of 32 bit.
>
> I am not sure I am following.

An additional patch was merged for 4.5 that increases the arguments to the
counter operations to 64 bit which is known for regressions.

[toc] | [prev] | [next] | [standalone]


#1318230 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromChristoph Lameter <cl@linux.com>
Date2016-01-26 19:40 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVgH0-88L-13@gated-at.bofh.it>
In reply to#1318119
On Tue, 26 Jan 2016, Mike Galbraith wrote:

> I disagree.  You're burning electrons for no benefit at all to me on my
> box.  You want to do high speed trading, that's fine, but I expect my
> box to be able to pop in and out of idle without having to pay a toll
> to the high speed trading bandits of the world, thank you very much.
>
> This specialty thing does not belong in the generic fast path.

The system going idle is a fastpath. Mind boogling.

[toc] | [prev] | [next] | [standalone]


#1318250 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-01-26 19:50 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVgQG-8cd-9@gated-at.bofh.it>
In reply to#1318230
On Tue, 2016-01-26 at 12:34 -0600, Christoph Lameter wrote:
> On Tue, 26 Jan 2016, Mike Galbraith wrote:
> 
> > I disagree.  You're burning electrons for no benefit at all to me on my
> > box.  You want to do high speed trading, that's fine, but I expect my
> > box to be able to pop in and out of idle without having to pay a toll
> > to the high speed trading bandits of the world, thank you very much.
> > 
> > This specialty thing does not belong in the generic fast path.
> 
> The system going idle is a fastpath. Mind boogling.

Hohum, noted.  Now what about those cycles, and the sleeping lock you
injected for -rt?

	-Mike

[toc] | [prev] | [next] | [standalone]


#1318288 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromChristoph Lameter <cl@linux.com>
Date2016-01-26 20:30 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVhtp-gQ-39@gated-at.bofh.it>
In reply to#1318250
On Tue, 26 Jan 2016, Mike Galbraith wrote:

> > The system going idle is a fastpath. Mind boogling.
>
> Hohum, noted.  Now what about those cycles, and the sleeping lock you
> injected for -rt?

Since we (the NOHZ people) care mostly about NOHZ then lets restrict
that to the NOHZ mode. Then it should not affect your load.


Subject: Move quiet_vmstat() to NOHZ code

quiet_vmstat() seems to cause regressions for some load because
the cpu going idle is a "fastpath". Mind boogling. Strange claim.
If the system goes idle then it has nothing to do after all

But anyways if we shift the quiet_vmstat() into the NOHZ logic
when it stops the tick then it will only affect those cores that
are setup for NOHZ mode. That is where we want this processing
after all to ensure that the OS keeps itself off those cores.

Signed-off-by: Christoph Lameter <cl@linux.com>


Index: linux/kernel/sched/idle.c
===================================================================
--- linux.orig/kernel/sched/idle.c
+++ linux/kernel/sched/idle.c
@@ -213,7 +213,6 @@ static void cpu_idle_loop(void)
 		 */

 		__current_set_polling();
-		quiet_vmstat();
 		tick_nohz_idle_enter();

 		while (!need_resched()) {
Index: linux/kernel/time/tick-sched.c
===================================================================
--- linux.orig/kernel/time/tick-sched.c
+++ linux/kernel/time/tick-sched.c
@@ -811,6 +811,7 @@ static void __tick_nohz_idle_enter(struc
 		ts->idle_calls++;

 		expires = tick_nohz_stop_sched_tick(ts, now, cpu);
+		quiet_vmstat();
 		if (expires.tv64 > 0LL) {
 			ts->idle_sleeps++;
 			ts->idle_expires = expires;

[toc] | [prev] | [next] | [standalone]


#1318578 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-01-27 04:20 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVoOe-5Ga-7@gated-at.bofh.it>
In reply to#1318288
Good morning,

On Tue, 2016-01-26 at 13:20 -0600, Christoph Lameter wrote:
> On Tue, 26 Jan 2016, Mike Galbraith wrote:
> 
> > > The system going idle is a fastpath. Mind boogling.
> > 
> > Hohum, noted.  Now what about those cycles, and the sleeping lock you
> > injected for -rt?
> 
> Since we (the NOHZ people) care mostly about NOHZ then lets restrict
> that to the NOHZ mode. Then it should not affect your load.

Tons of folks do have NO_HZ enabled (including me).  Isn't there a spot
somewhere in NO_HZ_FULL code where it can take up residence?  (one with
a tad lower maximum call frequency would be good, a nohz_full cpu isn't
necessarily being used for pure compute exclusively) 

> Subject: Move quiet_vmstat() to NOHZ code
> 
> quiet_vmstat() seems to cause regressions for some load because
> the cpu going idle is a "fastpath". Mind boogling. Strange claim.
> If the system goes idle then it has nothing to do after all
> 
> But anyways if we shift the quiet_vmstat() into the NOHZ logic
> when it stops the tick then it will only affect those cores that
> are setup for NOHZ mode. That is where we want this processing
> after all to ensure that the OS keeps itself off those cores.
> 
> Signed-off-by: Christoph Lameter <cl@linux.com>
> 
> 
> Index: linux/kernel/sched/idle.c
> ===================================================================
> --- linux.orig/kernel/sched/idle.c
> +++ linux/kernel/sched/idle.c
> @@ -213,7 +213,6 @@ static void cpu_idle_loop(void)
>  		 */
> 
>  		__current_set_polling();
> -		quiet_vmstat();
>  		tick_nohz_idle_enter();
> 
>  		while (!need_resched()) {
> Index: linux/kernel/time/tick-sched.c
> ===================================================================
> --- linux.orig/kernel/time/tick-sched.c
> +++ linux/kernel/time/tick-sched.c
> @@ -811,6 +811,7 @@ static void __tick_nohz_idle_enter(struc
>  		ts->idle_calls++;
> 
>  		expires = tick_nohz_stop_sched_tick(ts, now, cpu);
> +		quiet_vmstat();
>  		if (expires.tv64 > 0LL) {
>  			ts->idle_sleeps++;
>  			ts->idle_expires = expires;

[toc] | [prev] | [next] | [standalone]


#1318599 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-01-27 05:20 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVpKi-6jb-1@gated-at.bofh.it>
In reply to#1318578
On Wed, 2016-01-27 at 04:12 +0100, Mike Galbraith wrote:
> Good morning,
> 
> On Tue, 2016-01-26 at 13:20 -0600, Christoph Lameter wrote:
> > On Tue, 26 Jan 2016, Mike Galbraith wrote:
> > 
> > > > The system going idle is a fastpath. Mind boogling.
> > > 
> > > Hohum, noted.  Now what about those cycles, and the sleeping lock you
> > > injected for -rt?
> > 
> > Since we (the NOHZ people) care mostly about NOHZ then lets restrict
> > that to the NOHZ mode. Then it should not affect your load.
> 
> Tons of folks do have NO_HZ enabled (including me).  Isn't there a spot
> somewhere in NO_HZ_FULL code where it can take up residence?  (one with
> a tad lower maximum call frequency would be good, a nohz_full cpu isn't
> necessarily being used for pure compute exclusively)

I forgot to mention that the spot you picked is called with irqs
disabled.

> > Subject: Move quiet_vmstat() to NOHZ code
> > 
> > quiet_vmstat() seems to cause regressions for some load because
> > the cpu going idle is a "fastpath". Mind boogling. Strange claim.
> > If the system goes idle then it has nothing to do after all

Hm, seems I also forgot to say "Hohum, noted..." again.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1319140 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromChristoph Lameter <cl@linux.com>
Date2016-01-27 17:30 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVB8L-6ft-25@gated-at.bofh.it>
In reply to#1318578
On Wed, 27 Jan 2016, Mike Galbraith wrote:

> > Since we (the NOHZ people) care mostly about NOHZ then lets restrict
> > that to the NOHZ mode. Then it should not affect your load.
>
> Tons of folks do have NO_HZ enabled (including me).  Isn't there a spot
> somewhere in NO_HZ_FULL code where it can take up residence?  (one with
> a tad lower maximum call frequency would be good, a nohz_full cpu isn't
> necessarily being used for pure compute exclusively)

Frederic, any idea on where to put quiet_vmstat()?


> > Subject: Move quiet_vmstat() to NOHZ code
> >
> > quiet_vmstat() seems to cause regressions for some load because
> > the cpu going idle is a "fastpath". Mind boogling. Strange claim.
> > If the system goes idle then it has nothing to do after all
> >
> > But anyways if we shift the quiet_vmstat() into the NOHZ logic
> > when it stops the tick then it will only affect those cores that
> > are setup for NOHZ mode. That is where we want this processing
> > after all to ensure that the OS keeps itself off those cores.
> >
> > Signed-off-by: Christoph Lameter <cl@linux.com>
> >
> >
> > Index: linux/kernel/sched/idle.c
> > ===================================================================
> > --- linux.orig/kernel/sched/idle.c
> > +++ linux/kernel/sched/idle.c
> > @@ -213,7 +213,6 @@ static void cpu_idle_loop(void)
> >  		 */
> >
> >  		__current_set_polling();
> > -		quiet_vmstat();
> >  		tick_nohz_idle_enter();
> >
> >  		while (!need_resched()) {
> > Index: linux/kernel/time/tick-sched.c
> > ===================================================================
> > --- linux.orig/kernel/time/tick-sched.c
> > +++ linux/kernel/time/tick-sched.c
> > @@ -811,6 +811,7 @@ static void __tick_nohz_idle_enter(struc
> >  		ts->idle_calls++;
> >
> >  		expires = tick_nohz_stop_sched_tick(ts, now, cpu);
> > +		quiet_vmstat();
> >  		if (expires.tv64 > 0LL) {
> >  			ts->idle_sleeps++;
> >  			ts->idle_expires = expires;
>

[toc] | [prev] | [next] | [standalone]


#1318241 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-01-26 19:40 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVgH0-88L-15@gated-at.bofh.it>
In reply to#1318119
On Tue, 2016-01-26 at 10:25 -0600, Christoph Lameter wrote:
> On Mon, 25 Jan 2016, Michal Hocko wrote:
> 
> > > Why would the deferring cause this overhead?
> > 
> > I guess the profile speaks for itself, doesn't it?
> 
> But the system is going idle? Why would this impact performance?

We enter/exit idle a lot.

Your reluctance to move it seem to suggest that 99.99% of CPUs on the
planet chewing up cycles (measured) doing what for most is useless work
on every micro-idle is a perfectly fine price to pay to ensure that
.01% (or whatever tiny minority) get what they want.

I disagree.  You're burning electrons for no benefit at all to me on my
box.  You want to do high speed trading, that's fine, but I expect my
box to be able to pop in and out of idle without having to pay a toll
to the high speed trading bandits of the world, thank you very much.

This specialty thing does not belong in the generic fast path.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1317461 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-01-26 03:20 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qV1oB-4JB-1@gated-at.bofh.it>
In reply to#1317134
On Mon, 2016-01-25 at 12:02 -0600, Christoph Lameter wrote:
> On Mon, 25 Jan 2016, Michal Hocko wrote:
> 
> > On Sat 23-01-16 17:21:55, Mike Galbraith wrote:
> > > Hi Christoph,
> > > 
> > > While you're fixing that commit up, can you perhaps find a better home
> > > for quiet_vmstat()?  It not only munches cycles when switching cross
> > > -core mightily, for -rt it injects a sleeping lock into the idle task.
> > > 
> > >     12.89%  [kernel]       [k] refresh_cpu_vm_stats.isra.12
> > >      4.75%  [kernel]       [k] __schedule
> > >      4.70%  [kernel]       [k] mutex_unlock
> > >      3.14%  [kernel]       [k] __switch_to
> > 
> > Hmm, I wouldn't have expected that refresh_cpu_vm_stats could have
> > such a large footprint. I guess this would be just an expensive noop
> > because we have to check all the zones*counters and do an expensive
> > this_cpu_xchg. Is the whole deferred thing worth this overhead?
> 
> Why would the deferring cause this overhead?

Because we schedule to idle cores aggressively, thus we may pop in and
out of idle at high frequency.

> Also there is no cross core activity from quiet_vmstat(). It simply
> disables the local vmstat updates.

Again, the cross core activity is not due to quiet_vmstat(), it is due
to pipe-test threads running on two cores, and meeting quiet_vmstat()
at high frequency.

> > Unless there is a clear and huge win from doing the vmstat update
> > deferrable then I think a revert is more appropriate IMHO.
> 
> It reduces the OS events that the application experiences by folding it
> into the tick events. If its not deferrable then a timer event will be
> generated in addition to the tick. We do not want that.

Perf and RT say we don't want quiet_vmstat() in the idle loop either.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1317472 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-01-26 03:30 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qV1yj-4Nn-19@gated-at.bofh.it>
In reply to#1317461
On Tue, 2016-01-26 at 03:14 +0100, Mike Galbraith wrote:

> Perf and RT say we don't want quiet_vmstat() in the idle loop either.

BTW, the perf numbers were not from an RT kernel, they were from my
PREEMPT_VOLUNTARY desktop kernel.

	-Mike 

[toc] | [prev] | [next] | [standalone]


#1318116 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromChristoph Lameter <cl@linux.com>
Date2016-01-26 17:30 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVeFc-6F7-13@gated-at.bofh.it>
In reply to#1317472
On Tue, 26 Jan 2016, Mike Galbraith wrote:

> On Tue, 2016-01-26 at 03:14 +0100, Mike Galbraith wrote:
>
> > Perf and RT say we don't want quiet_vmstat() in the idle loop either.
>
> BTW, the perf numbers were not from an RT kernel, they were from my
> PREEMPT_VOLUNTARY desktop kernel.

Can we move quiet_vmstat() elsewhere after we have checked that really
nothing else is going on soon?

[toc] | [prev] | [next] | [standalone]


#1318203 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-01-26 18:40 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVfKX-7rW-23@gated-at.bofh.it>
In reply to#1318116
On Tue, 2016-01-26 at 10:26 -0600, Christoph Lameter wrote:
> On Tue, 26 Jan 2016, Mike Galbraith wrote:
> 
> > On Tue, 2016-01-26 at 03:14 +0100, Mike Galbraith wrote:
> > 
> > > Perf and RT say we don't want quiet_vmstat() in the idle loop
> > > either.
> > 
> > BTW, the perf numbers were not from an RT kernel, they were from my
> > PREEMPT_VOLUNTARY desktop kernel.
> 
> Can we move quiet_vmstat() elsewhere after we have checked that really
> nothing else is going on soon?

How would you check?  Precognition doesn't work for mortals.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1318216 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromChristoph Lameter <cl@linux.com>
Date2016-01-26 19:20 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVgnD-7Zo-7@gated-at.bofh.it>
In reply to#1318203
On Tue, 26 Jan 2016, Mike Galbraith wrote:

> On Tue, 2016-01-26 at 10:26 -0600, Christoph Lameter wrote:
> > On Tue, 26 Jan 2016, Mike Galbraith wrote:
> >
> > > On Tue, 2016-01-26 at 03:14 +0100, Mike Galbraith wrote:
> > >
> > > > Perf and RT say we don't want quiet_vmstat() in the idle loop
> > > > either.
> > >
> > > BTW, the perf numbers were not from an RT kernel, they were from my
> > > PREEMPT_VOLUNTARY desktop kernel.
> >
> > Can we move quiet_vmstat() elsewhere after we have checked that really
> > nothing else is going on soon?
>
> How would you check?  Precognition doesn't work for mortals.

Dont we have some decision mechanism to go into higher levels of power
savings when the system is idle for longer times?

[toc] | [prev] | [next] | [standalone]


#1318120 — Re: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)

FromChristoph Lameter <cl@linux.com>
Date2016-01-26 17:30 +0100
SubjectRe: fast path cycle muncher (vmstat: make vmstat_updater deferrable again and shut down on idle)
Message-ID<qVeFd-6F7-21@gated-at.bofh.it>
In reply to#1317461
On Tue, 26 Jan 2016, Mike Galbraith wrote:

> > Why would the deferring cause this overhead?
>
> Because we schedule to idle cores aggressively, thus we may pop in and
> out of idle at high frequency.

Whats the point of going idle if you have things to do soon?

[toc] | [prev] | [next] | [standalone]


Page 2 of 3 — ← Prev page 1 [2] 3  Next page →

Back to top | Article view | linux.kernel


csiph-web