Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1610014 > unrolled thread

Re: [BUG nohz]: wrong user and system time accounting

Started byRik van Riel <riel@redhat.com>
First post2017-03-27 19:40 +0200
Last post2017-03-30 14:30 +0200
Articles 18 on this page of 38 — 5 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [BUG nohz]: wrong user and system time accounting Rik van Riel <riel@redhat.com> - 2017-03-27 19:40 +0200
    Re: [BUG nohz]: wrong user and system time accounting Wanpeng Li <kernellwp@gmail.com> - 2017-03-28 09:30 +0200
    Re: [BUG nohz]: wrong user and system time accounting Rik van Riel <riel@redhat.com> - 2017-03-28 23:10 +0200
      Re: [BUG nohz]: wrong user and system time accounting Luiz Capitulino <lcapitulino@redhat.com> - 2017-03-28 23:30 +0200
        Re: [BUG nohz]: wrong user and system time accounting Wanpeng Li <kernellwp@gmail.com> - 2017-03-29 12:00 +0200
          Re: [BUG nohz]: wrong user and system time accounting Frederic Weisbecker <fweisbec@gmail.com> - 2017-03-29 15:00 +0200
    Re: [BUG nohz]: wrong user and system time accounting Rik van Riel <riel@redhat.com> - 2017-03-28 23:30 +0200
      Re: [BUG nohz]: wrong user and system time accounting Luiz Capitulino <lcapitulino@redhat.com> - 2017-03-28 23:40 +0200
    Re: [BUG nohz]: wrong user and system time accounting Rik van Riel <riel@redhat.com> - 2017-03-29 22:20 +0200
      Re: [BUG nohz]: wrong user and system time accounting Frederic Weisbecker <fweisbec@gmail.com> - 2017-03-30 01:00 +0200
        Re: [BUG nohz]: wrong user and system time accounting Rik van Riel <riel@redhat.com> - 2017-03-30 15:00 +0200
      Re: [BUG nohz]: wrong user and system time accounting Wanpeng Li <kernellwp@gmail.com> - 2017-03-30 04:00 +0200
        Re: [BUG nohz]: wrong user and system time accounting Frederic Weisbecker <fweisbec@gmail.com> - 2017-03-30 14:50 +0200
          Re: [BUG nohz]: wrong user and system time accounting Mike Galbraith <efault@gmx.de> - 2017-03-30 15:20 +0200
      Re: [BUG nohz]: wrong user and system time accounting Mike Galbraith <efault@gmx.de> - 2017-03-30 06:30 +0200
        Re: [BUG nohz]: wrong user and system time accounting Wanpeng Li <kernellwp@gmail.com> - 2017-03-30 08:50 +0200
          Re: [BUG nohz]: wrong user and system time accounting Wanpeng Li <kernellwp@gmail.com> - 2017-03-30 14:00 +0200
            Re: [BUG nohz]: wrong user and system time accounting Mike Galbraith <efault@gmx.de> - 2017-03-30 14:40 +0200
          Re: [BUG nohz]: wrong user and system time accounting Frederic Weisbecker <fweisbec@gmail.com> - 2017-03-30 15:40 +0200
            Re: [BUG nohz]: wrong user and system time accounting Wanpeng Li <kernellwp@gmail.com> - 2017-03-30 16:10 +0200
              Re: [BUG nohz]: wrong user and system time accounting Frederic Weisbecker <fweisbec@gmail.com> - 2017-03-30 16:30 +0200
                Re: [BUG nohz]: wrong user and system time accounting Luiz Capitulino <lcapitulino@redhat.com> - 2017-03-30 23:30 +0200
                  Re: [BUG nohz]: wrong user and system time accounting Luiz Capitulino <lcapitulino@redhat.com> - 2017-03-31 22:10 +0200
                    Re: [BUG nohz]: wrong user and system time accounting Frederic Weisbecker <fweisbec@gmail.com> - 2017-04-01 01:30 +0200
                      Re: [BUG nohz]: wrong user and system time accounting Luiz Capitulino <lcapitulino@redhat.com> - 2017-04-01 05:20 +0200
                        Re: [BUG nohz]: wrong user and system time accounting Frederic Weisbecker <fweisbec@gmail.com> - 2017-04-03 17:30 +0200
                          Re: [BUG nohz]: wrong user and system time accounting Luiz Capitulino <lcapitulino@redhat.com> - 2017-04-03 21:10 +0200
                            Re: [BUG nohz]: wrong user and system time accounting Luiz Capitulino <lcapitulino@redhat.com> - 2017-04-04 20:10 +0200
                              Re: [BUG nohz]: wrong user and system time accounting Rik van Riel <riel@redhat.com> - 2017-04-05 16:30 +0200
        Re: [BUG nohz]: wrong user and system time accounting Frederic Weisbecker <fweisbec@gmail.com> - 2017-03-30 15:00 +0200
          Re: [BUG nohz]: wrong user and system time accounting Rik van Riel <riel@redhat.com> - 2017-03-30 15:10 +0200
            Re: [BUG nohz]: wrong user and system time accounting Mike Galbraith <efault@gmx.de> - 2017-03-30 15:40 +0200
              Re: [BUG nohz]: wrong user and system time accounting Frederic Weisbecker <fweisbec@gmail.com> - 2017-04-03 16:50 +0200
                Re: [BUG nohz]: wrong user and system time accounting Mike Galbraith <efault@gmx.de> - 2017-04-04 09:40 +0200
            Re: [BUG nohz]: wrong user and system time accounting Frederic Weisbecker <fweisbec@gmail.com> - 2017-03-30 15:50 +0200
    Re: [BUG nohz]: wrong user and system time accounting Wanpeng Li <kernellwp@gmail.com> - 2017-03-30 00:50 +0200
      Re: [BUG nohz]: wrong user and system time accounting Luiz Capitulino <lcapitulino@redhat.com> - 2017-03-30 04:20 +0200
        Re: [BUG nohz]: wrong user and system time accounting Wanpeng Li <kernellwp@gmail.com> - 2017-03-30 14:30 +0200

Page 2 of 2 — ← Prev page 1 [2]


#1613120

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-03-30 16:30 +0200
Message-ID<tqJfk-7FK-17@gated-at.bofh.it>
In reply to#1613106
On Thu, Mar 30, 2017 at 09:59:54PM +0800, Wanpeng Li wrote:
> 2017-03-30 21:38 GMT+08:00 Frederic Weisbecker <fweisbec@gmail.com>:
> > If it works, we may want to take that solution, likely less performance sensitive
> > than using sched_clock(). In fact sched_clock() is fast, especially as we require it to
> > be stable for nohz_full, but using it involves costly conversion back and forth to jiffies.
> 
> So both Rik and you agree with the skew tick solution, I will try it
> tomorrow. Btw, if we should just add random offset to the cpu in the
> nohz_full mode or add random offset to all cpus like the codes above?

Lets just keep it to all CPUs for simplicty.
Also please add a comment that explains why we need that skew_tick on nohz_full.

Thanks!

[toc] | [prev] | [next] | [standalone]


#1613488

FromLuiz Capitulino <lcapitulino@redhat.com>
Date2017-03-30 23:30 +0200
Message-ID<tqPNL-3S3-5@gated-at.bofh.it>
In reply to#1613120
On Thu, 30 Mar 2017 16:18:17 +0200
Frederic Weisbecker <fweisbec@gmail.com> wrote:

> On Thu, Mar 30, 2017 at 09:59:54PM +0800, Wanpeng Li wrote:
> > 2017-03-30 21:38 GMT+08:00 Frederic Weisbecker <fweisbec@gmail.com>:  
> > > If it works, we may want to take that solution, likely less performance sensitive
> > > than using sched_clock(). In fact sched_clock() is fast, especially as we require it to
> > > be stable for nohz_full, but using it involves costly conversion back and forth to jiffies.  
> > 
> > So both Rik and you agree with the skew tick solution, I will try it
> > tomorrow. Btw, if we should just add random offset to the cpu in the
> > nohz_full mode or add random offset to all cpus like the codes above?  
> 
> Lets just keep it to all CPUs for simplicty.
> Also please add a comment that explains why we need that skew_tick on nohz_full.

I've tried all the test-cases we discussed in this thread with skew_tick=1
and it worked as expected in bare-metal and KVM guests.

However, I found a test-case that works in bare-metal but show problems
in KVM guests. It could something that's KVM specific, or it could be
something that's harder to reproduce in bare-metal.

The reproducer is (not sure all the steps are necessary):

1. Isolate 8 cores in the host with isolcpus= and nohz_full= (and skew_tick=1)

2. Create a KVM guest with 8 vCPUs and pin each vCPU to an isolated
   host core

3. Boot the guest with isolcpus=2,3,4,5,6,7 nohz_full=2,3,4,5,6,7 skew_tick=1

4. Once the guest is booted, run:

# for i in $(seq 2 7); do taskset -c $i hog& ;done
# taskset -c 2,3,4,5,6,7 \
  cyclictest -m -n -q -p95 -D 1m -h60 -i 200 -t 6 -a 2,3,4,5,6,7

  (where hog is a program taking 100% of the CPU, and cyclictest
   is RT's cyclictest)

5. Run top -d1

In a few minutes into this test-case, I see one isolated CPU in the
guest reporting around 95% system time (where the expected is close
to 100% user time, which the others isolated CPUs correctly report).

[toc] | [prev] | [next] | [standalone]


#1614279

FromLuiz Capitulino <lcapitulino@redhat.com>
Date2017-03-31 22:10 +0200
Message-ID<trb1T-QL-3@gated-at.bofh.it>
In reply to#1613488
On Thu, 30 Mar 2017 17:25:46 -0400
Luiz Capitulino <lcapitulino@redhat.com> wrote:

> On Thu, 30 Mar 2017 16:18:17 +0200
> Frederic Weisbecker <fweisbec@gmail.com> wrote:
> 
> > On Thu, Mar 30, 2017 at 09:59:54PM +0800, Wanpeng Li wrote:  
> > > 2017-03-30 21:38 GMT+08:00 Frederic Weisbecker <fweisbec@gmail.com>:    
> > > > If it works, we may want to take that solution, likely less performance sensitive
> > > > than using sched_clock(). In fact sched_clock() is fast, especially as we require it to
> > > > be stable for nohz_full, but using it involves costly conversion back and forth to jiffies.    
> > > 
> > > So both Rik and you agree with the skew tick solution, I will try it
> > > tomorrow. Btw, if we should just add random offset to the cpu in the
> > > nohz_full mode or add random offset to all cpus like the codes above?    
> > 
> > Lets just keep it to all CPUs for simplicty.
> > Also please add a comment that explains why we need that skew_tick on nohz_full.  
> 
> I've tried all the test-cases we discussed in this thread with skew_tick=1
> and it worked as expected in bare-metal and KVM guests.
> 
> However, I found a test-case that works in bare-metal but show problems
> in KVM guests. It could something that's KVM specific, or it could be
> something that's harder to reproduce in bare-metal.

After discussing some findings on this issue with Rik, I realized that
we don't add the skew when restarting the tick in tick_nohz_restart().
Adding the offset there seems to solve this problem.

[toc] | [prev] | [next] | [standalone]


#1614325

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-04-01 01:30 +0200
Message-ID<tre9r-2KJ-3@gated-at.bofh.it>
In reply to#1614279
On Fri, Mar 31, 2017 at 04:09:10PM -0400, Luiz Capitulino wrote:
> On Thu, 30 Mar 2017 17:25:46 -0400
> Luiz Capitulino <lcapitulino@redhat.com> wrote:
> 
> > On Thu, 30 Mar 2017 16:18:17 +0200
> > Frederic Weisbecker <fweisbec@gmail.com> wrote:
> > 
> > > On Thu, Mar 30, 2017 at 09:59:54PM +0800, Wanpeng Li wrote:  
> > > > 2017-03-30 21:38 GMT+08:00 Frederic Weisbecker <fweisbec@gmail.com>:    
> > > > > If it works, we may want to take that solution, likely less performance sensitive
> > > > > than using sched_clock(). In fact sched_clock() is fast, especially as we require it to
> > > > > be stable for nohz_full, but using it involves costly conversion back and forth to jiffies.    
> > > > 
> > > > So both Rik and you agree with the skew tick solution, I will try it
> > > > tomorrow. Btw, if we should just add random offset to the cpu in the
> > > > nohz_full mode or add random offset to all cpus like the codes above?    
> > > 
> > > Lets just keep it to all CPUs for simplicty.
> > > Also please add a comment that explains why we need that skew_tick on nohz_full.  
> > 
> > I've tried all the test-cases we discussed in this thread with skew_tick=1
> > and it worked as expected in bare-metal and KVM guests.
> > 
> > However, I found a test-case that works in bare-metal but show problems
> > in KVM guests. It could something that's KVM specific, or it could be
> > something that's harder to reproduce in bare-metal.
> 
> After discussing some findings on this issue with Rik, I realized that
> we don't add the skew when restarting the tick in tick_nohz_restart().
> Adding the offset there seems to solve this problem.

Are you sure? tick_nohz_restart() doesn't seem to override the initial skew. It
always forwards the expiration time on top of the last tick.

[toc] | [prev] | [next] | [standalone]


#1614379

FromLuiz Capitulino <lcapitulino@redhat.com>
Date2017-04-01 05:20 +0200
Message-ID<trhK1-5pA-1@gated-at.bofh.it>
In reply to#1614325
On Sat, 1 Apr 2017 01:24:54 +0200
Frederic Weisbecker <fweisbec@gmail.com> wrote:

> On Fri, Mar 31, 2017 at 04:09:10PM -0400, Luiz Capitulino wrote:
> > On Thu, 30 Mar 2017 17:25:46 -0400
> > Luiz Capitulino <lcapitulino@redhat.com> wrote:
> >   
> > > On Thu, 30 Mar 2017 16:18:17 +0200
> > > Frederic Weisbecker <fweisbec@gmail.com> wrote:
> > >   
> > > > On Thu, Mar 30, 2017 at 09:59:54PM +0800, Wanpeng Li wrote:    
> > > > > 2017-03-30 21:38 GMT+08:00 Frederic Weisbecker <fweisbec@gmail.com>:      
> > > > > > If it works, we may want to take that solution, likely less performance sensitive
> > > > > > than using sched_clock(). In fact sched_clock() is fast, especially as we require it to
> > > > > > be stable for nohz_full, but using it involves costly conversion back and forth to jiffies.      
> > > > > 
> > > > > So both Rik and you agree with the skew tick solution, I will try it
> > > > > tomorrow. Btw, if we should just add random offset to the cpu in the
> > > > > nohz_full mode or add random offset to all cpus like the codes above?      
> > > > 
> > > > Lets just keep it to all CPUs for simplicty.
> > > > Also please add a comment that explains why we need that skew_tick on nohz_full.    
> > > 
> > > I've tried all the test-cases we discussed in this thread with skew_tick=1
> > > and it worked as expected in bare-metal and KVM guests.
> > > 
> > > However, I found a test-case that works in bare-metal but show problems
> > > in KVM guests. It could something that's KVM specific, or it could be
> > > something that's harder to reproduce in bare-metal.  
> > 
> > After discussing some findings on this issue with Rik, I realized that
> > we don't add the skew when restarting the tick in tick_nohz_restart().
> > Adding the offset there seems to solve this problem.  
> 
> Are you sure? tick_nohz_restart() doesn't seem to override the initial skew. It
> always forwards the expiration time on top of the last tick.

OK, I'll double check. Without my change the bug triggers almost
instantly with the described reproducer. With my change it didn't
trig for several minutes (but it does look wrong looking at it now).

[toc] | [prev] | [next] | [standalone]


#1615337

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-04-03 17:30 +0200
Message-ID<tsc5B-cE-45@gated-at.bofh.it>
In reply to#1614379
On Fri, Mar 31, 2017 at 11:11:19PM -0400, Luiz Capitulino wrote:
> On Sat, 1 Apr 2017 01:24:54 +0200
> Frederic Weisbecker <fweisbec@gmail.com> wrote:
> 
> > On Fri, Mar 31, 2017 at 04:09:10PM -0400, Luiz Capitulino wrote:
> > > On Thu, 30 Mar 2017 17:25:46 -0400
> > > Luiz Capitulino <lcapitulino@redhat.com> wrote:
> > >   
> > > > On Thu, 30 Mar 2017 16:18:17 +0200
> > > > Frederic Weisbecker <fweisbec@gmail.com> wrote:
> > > >   
> > > > > On Thu, Mar 30, 2017 at 09:59:54PM +0800, Wanpeng Li wrote:    
> > > > > > 2017-03-30 21:38 GMT+08:00 Frederic Weisbecker <fweisbec@gmail.com>:      
> > > > > > > If it works, we may want to take that solution, likely less performance sensitive
> > > > > > > than using sched_clock(). In fact sched_clock() is fast, especially as we require it to
> > > > > > > be stable for nohz_full, but using it involves costly conversion back and forth to jiffies.      
> > > > > > 
> > > > > > So both Rik and you agree with the skew tick solution, I will try it
> > > > > > tomorrow. Btw, if we should just add random offset to the cpu in the
> > > > > > nohz_full mode or add random offset to all cpus like the codes above?      
> > > > > 
> > > > > Lets just keep it to all CPUs for simplicty.
> > > > > Also please add a comment that explains why we need that skew_tick on nohz_full.    
> > > > 
> > > > I've tried all the test-cases we discussed in this thread with skew_tick=1
> > > > and it worked as expected in bare-metal and KVM guests.
> > > > 
> > > > However, I found a test-case that works in bare-metal but show problems
> > > > in KVM guests. It could something that's KVM specific, or it could be
> > > > something that's harder to reproduce in bare-metal.  
> > > 
> > > After discussing some findings on this issue with Rik, I realized that
> > > we don't add the skew when restarting the tick in tick_nohz_restart().
> > > Adding the offset there seems to solve this problem.  
> > 
> > Are you sure? tick_nohz_restart() doesn't seem to override the initial skew. It
> > always forwards the expiration time on top of the last tick.
> 
> OK, I'll double check. Without my change the bug triggers almost
> instantly with the described reproducer. With my change it didn't
> trig for several minutes (but it does look wrong looking at it now).

Do you observe aligned ticks with trace events (hrtimer_expire_entry)?

You might want to enforce the global clock to trace that:

    echo "global" > /sys/kernel/debug/tracing/trace_clock

[toc] | [prev] | [next] | [standalone]


#1615503

FromLuiz Capitulino <lcapitulino@redhat.com>
Date2017-04-03 21:10 +0200
Message-ID<tsfwv-2wG-49@gated-at.bofh.it>
In reply to#1615337
On Mon, 3 Apr 2017 17:23:17 +0200
Frederic Weisbecker <fweisbec@gmail.com> wrote:

> Do you observe aligned ticks with trace events (hrtimer_expire_entry)?
> 
> You might want to enforce the global clock to trace that:
> 
>     echo "global" > /sys/kernel/debug/tracing/trace_clock

I've used the same trace points & debugging code I've been using to debug
this issue, and this what I'm seeing:

    stress-25757 [002]  2742.717507: function:             enter_from_user_mode <-- apic_timer_interrupt
    stress-25757 [002]  2742.717508: function:             __context_tracking_exit <-- enter_from_user_mode
    stress-25757 [002]  2742.717508: bprint:               vtime_delta: diff=0 (now=4297409970 vtime_snap=4297409970)
    stress-25757 [002]  2742.717509: function:             smp_apic_timer_interrupt <-- apic_timer_interrupt
    stress-25757 [002]  2742.717509: function:             irq_enter <-- smp_apic_timer_interrupt
    stress-25757 [002]  2742.717510: hrtimer_expire_entry: hrtimer=0xffffc900039fbe58 function=hrtimer_wakeup now=2742674000776
    stress-25757 [002]  2742.717514: function:             irq_exit <-- smp_apic_timer_interrupt
cyclictest-25760 [002]  2742.717518: function:             vtime_account_system <-- vtime_common_task_switch
cyclictest-25760 [002]  2742.717518: bprint:               vtime_delta: diff=1000000 (now=4297409971 vtime_snap=4297409970)
cyclictest-25760 [002]  2742.717519: function:             __vtime_account_system <-- vtime_account_system
cyclictest-25760 [002]  2742.717519: bprint:               get_vtime_delta: vtime_snap=4297409970 now=4297409971
cyclictest-25760 [002]  2742.717520: function:             account_system_time <-- __vtime_account_system
cyclictest-25760 [002]  2742.717520: bprint:               account_system_time: cputime=961981
cyclictest-25760 [002]  2742.717521: function:             __context_tracking_enter <-- do_syscall_64
cyclictest-25760 [002]  2742.717522: function:             vtime_user_enter <-- __context_tracking_enter
cyclictest-25760 [002]  2742.717522: bprint:               vtime_delta: diff=0 (now=4297409971 vtime_snap=4297409971)

CPU2 shows 98% system time while the other CPUs (from CPU3 to CPU7)
show 98% user time (they're all running the same workload).

What's happening here is:

1. Timer interrupt
2. Transition from user-space to kernel-space, vtimer_delta()
   returns zero
3. Context switch from hog application to cyclictest
4. This time vtime_delta() returns != zero, which implies
   jiffies was updated between steps 2 and 3

This seems to be the pattern that accounts incorrectly,
and seem to suggest that the ticks are aligned because
this repeats over and over.

Please, let me know if you want me to run a different
trace-cmd command-line.

[toc] | [prev] | [next] | [standalone]


#1616279

FromLuiz Capitulino <lcapitulino@redhat.com>
Date2017-04-04 20:10 +0200
Message-ID<tsB3Y-8pK-7@gated-at.bofh.it>
In reply to#1615503
On Mon, 3 Apr 2017 15:06:13 -0400
Luiz Capitulino <lcapitulino@redhat.com> wrote:

> On Mon, 3 Apr 2017 17:23:17 +0200
> Frederic Weisbecker <fweisbec@gmail.com> wrote:
> 
> > Do you observe aligned ticks with trace events (hrtimer_expire_entry)?
> > 
> > You might want to enforce the global clock to trace that:
> > 
> >     echo "global" > /sys/kernel/debug/tracing/trace_clock  
> 
> I've used the same trace points & debugging code I've been using to debug
> this issue, and this what I'm seeing:
> 
>     stress-25757 [002]  2742.717507: function:             enter_from_user_mode <-- apic_timer_interrupt
>     stress-25757 [002]  2742.717508: function:             __context_tracking_exit <-- enter_from_user_mode
>     stress-25757 [002]  2742.717508: bprint:               vtime_delta: diff=0 (now=4297409970 vtime_snap=4297409970)
>     stress-25757 [002]  2742.717509: function:             smp_apic_timer_interrupt <-- apic_timer_interrupt
>     stress-25757 [002]  2742.717509: function:             irq_enter <-- smp_apic_timer_interrupt
>     stress-25757 [002]  2742.717510: hrtimer_expire_entry: hrtimer=0xffffc900039fbe58 function=hrtimer_wakeup now=2742674000776
>     stress-25757 [002]  2742.717514: function:             irq_exit <-- smp_apic_timer_interrupt
> cyclictest-25760 [002]  2742.717518: function:             vtime_account_system <-- vtime_common_task_switch
> cyclictest-25760 [002]  2742.717518: bprint:               vtime_delta: diff=1000000 (now=4297409971 vtime_snap=4297409970)
> cyclictest-25760 [002]  2742.717519: function:             __vtime_account_system <-- vtime_account_system
> cyclictest-25760 [002]  2742.717519: bprint:               get_vtime_delta: vtime_snap=4297409970 now=4297409971
> cyclictest-25760 [002]  2742.717520: function:             account_system_time <-- __vtime_account_system
> cyclictest-25760 [002]  2742.717520: bprint:               account_system_time: cputime=961981
> cyclictest-25760 [002]  2742.717521: function:             __context_tracking_enter <-- do_syscall_64
> cyclictest-25760 [002]  2742.717522: function:             vtime_user_enter <-- __context_tracking_enter
> cyclictest-25760 [002]  2742.717522: bprint:               vtime_delta: diff=0 (now=4297409971 vtime_snap=4297409971)
> 
> CPU2 shows 98% system time while the other CPUs (from CPU3 to CPU7)
> show 98% user time (they're all running the same workload).

On further debugging this, I realized that I had overlooked something:
the timer interrupt in this trace is not the tick, but cyclictest's timer
(remember that the test-case consists of pinning cyclictest and a task
hogging the CPU to the same CPU).

I'm running cyclictest with -i 200. If I increase this to -i 1000, then
I seem unable to reproduce the issue (caution: even with -i 200 it
doesn't always happen. But it does usually happen after I restart the
test-case a few times. However, I've never been able to reproduce
with -i 1000).

Now, if it's really cyclictest that's causing the timer interrupts to
get aligned, I guess this might not have a solution? (note: I haven't
been able to reproduce this on bare-metal).

[toc] | [prev] | [next] | [standalone]


#1616997

FromRik van Riel <riel@redhat.com>
Date2017-04-05 16:30 +0200
Message-ID<tsU6B-3Kb-15@gated-at.bofh.it>
In reply to#1616279

[Multipart message — attachments visible in raw view] — view raw

On Tue, 2017-04-04 at 13:36 -0400, Luiz Capitulino wrote:
> 
> On further debugging this, I realized that I had overlooked
> something:
> the timer interrupt in this trace is not the tick, but cyclictest's
> timer
> (remember that the test-case consists of pinning cyclictest and a
> task
> hogging the CPU to the same CPU).
> 
> I'm running cyclictest with -i 200. If I increase this to -i 1000,
> then
> I seem unable to reproduce the issue (caution: even with -i 200 it
> doesn't always happen. But it does usually happen after I restart the
> test-case a few times. However, I've never been able to reproduce
> with -i 1000).
> 
> Now, if it's really cyclictest that's causing the timer interrupts to
> get aligned, I guess this might not have a solution? (note: I haven't
> been able to reproduce this on bare-metal).

With any sample (tick) based timekeeping, it is possible
to construct workloads that avoid the sampling and result
in skewed statistics as a result.

However, given that local users can already DoS the system
in all kinds of ways, skewed statistics are probably not
that high up on the list of importance.

If there were a way to do accurate accounting (true vtime
accounting) without increasing the overhead of every
syscall and interrupt noticeably, that might be worth it,
but syscall overhead is likely to be a more important
factor than the accuracy of statistics.

I don't know if doing TSC reads and subtraction/addition
only, and delaying the conversion to cputime until a later
point would slow down system calls measurably, compared
with reading jiffies and comparing it against a cached
value of jiffies, nor do I know whether spending time
implementing and testing that would be worthwhile :)

-- 
All rights reversed

[toc] | [prev] | [next] | [standalone]


#1613040

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-03-30 15:00 +0200
Message-ID<tqHQd-6vZ-3@gated-at.bofh.it>
In reply to#1612558
On Thu, Mar 30, 2017 at 06:27:31AM +0200, Mike Galbraith wrote:
> On Wed, 2017-03-29 at 16:08 -0400, Rik van Riel wrote:
> 
> > A random offset, or better yet a somewhat randomized
> > tick length to make sure that simultaneous ticks are
> > fairly rare and the vtime sampling does not end up
> > "in phase" with the jiffies incrementing, could make
> > the accounting work right again.
> 
> That improves jitter, especially on big boxen.  I have an 8 socket box
> that thinks it's an extra large PC, there, collision avoidance matters
> hugely.  I couldn't reproduce bean counting woes, no idea if collision
> avoidance will help that.

Out of curiosity, where is the main contention between ticks? I indeed
know some locks that can be taken on special cases, such as posix cpu timers.

Also, why does it raise power consumption issues?

Thanks.

[toc] | [prev] | [next] | [standalone]


#1613045

FromRik van Riel <riel@redhat.com>
Date2017-03-30 15:10 +0200
Message-ID<tqHZU-6On-9@gated-at.bofh.it>
In reply to#1613040
On Thu, 2017-03-30 at 14:51 +0200, Frederic Weisbecker wrote:
> On Thu, Mar 30, 2017 at 06:27:31AM +0200, Mike Galbraith wrote:
> > On Wed, 2017-03-29 at 16:08 -0400, Rik van Riel wrote:
> > 
> > > A random offset, or better yet a somewhat randomized
> > > tick length to make sure that simultaneous ticks are
> > > fairly rare and the vtime sampling does not end up
> > > "in phase" with the jiffies incrementing, could make
> > > the accounting work right again.
> > 
> > That improves jitter, especially on big boxen.  I have an 8 socket
> > box
> > that thinks it's an extra large PC, there, collision avoidance
> > matters
> > hugely.  I couldn't reproduce bean counting woes, no idea if
> > collision
> > avoidance will help that.
> 
> Out of curiosity, where is the main contention between ticks? I
> indeed
> know some locks that can be taken on special cases, such as posix cpu
> timers.
> 
> Also, why does it raise power consumption issues?

On a system without either nohz_full or nohz idle
mode, skewed ticks result in CPU cores waking up
at different times, and keeping an idle system
consuming power for more time than it would if all
the ticks happened simultaneously.

This is not a factor at all on systems that switch
off the tick while idle, since the CPU will be busy
anyway while the tick is enabled.

[toc] | [prev] | [next] | [standalone]


#1613068

FromMike Galbraith <efault@gmx.de>
Date2017-03-30 15:40 +0200
Message-ID<tqIsW-70Q-27@gated-at.bofh.it>
In reply to#1613045
On Thu, 2017-03-30 at 09:02 -0400, Rik van Riel wrote:
> On Thu, 2017-03-30 at 14:51 +0200, Frederic Weisbecker wrote:

> > Also, why does it raise power consumption issues?
> 
> On a system without either nohz_full or nohz idle
> mode, skewed ticks result in CPU cores waking up
> at different times, and keeping an idle system
> consuming power for more time than it would if all
> the ticks happened simultaneously.

And if your server farm is mostly idle, that power savings may delay
your bankruptcy proceedings by a whole microsecond ;-)

Or more seriously, what skew does do on boxen of size X today is
something for perf to say.  At the time, removal was very bad for my 8
socket box, and allegedly caused huge SGI beasts in horrific pain.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1615274

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-04-03 16:50 +0200
Message-ID<tsbsS-89y-21@gated-at.bofh.it>
In reply to#1613068
On Thu, Mar 30, 2017 at 03:35:22PM +0200, Mike Galbraith wrote:
> On Thu, 2017-03-30 at 09:02 -0400, Rik van Riel wrote:
> > On Thu, 2017-03-30 at 14:51 +0200, Frederic Weisbecker wrote:
> 
> > > Also, why does it raise power consumption issues?
> > 
> > On a system without either nohz_full or nohz idle
> > mode, skewed ticks result in CPU cores waking up
> > at different times, and keeping an idle system
> > consuming power for more time than it would if all
> > the ticks happened simultaneously.
> 
> And if your server farm is mostly idle, that power savings may delay
> your bankruptcy proceedings by a whole microsecond ;-)
> 
> Or more seriously, what skew does do on boxen of size X today is
> something for perf to say.  At the time, removal was very bad for my 8
> socket box, and allegedly caused huge SGI beasts in horrific pain.

I see.
Nohz_full is already bad for powersavings anyway. CPU 0 always ticks :-)

[toc] | [prev] | [next] | [standalone]


#1615749

FromMike Galbraith <efault@gmx.de>
Date2017-04-04 09:40 +0200
Message-ID<tsrei-1N1-9@gated-at.bofh.it>
In reply to#1615274
On Mon, 2017-04-03 at 16:40 +0200, Frederic Weisbecker wrote:
> On Thu, Mar 30, 2017 at 03:35:22PM +0200, Mike Galbraith wrote:

> Nohz_full is already bad for powersavings anyway. CPU 0 always ticks :-)

OTOH, if a nohz_full set is doing what it was born to do, CPU0 tick
spikes won't be noticeable on your (pegged/glowing) watt meter :)

[toc] | [prev] | [next] | [standalone]


#1613078

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-03-30 15:50 +0200
Message-ID<tqICB-74s-5@gated-at.bofh.it>
In reply to#1613045
On Thu, Mar 30, 2017 at 09:02:31AM -0400, Rik van Riel wrote:
> On Thu, 2017-03-30 at 14:51 +0200, Frederic Weisbecker wrote:
> > On Thu, Mar 30, 2017 at 06:27:31AM +0200, Mike Galbraith wrote:
> > > On Wed, 2017-03-29 at 16:08 -0400, Rik van Riel wrote:
> > > 
> > > > A random offset, or better yet a somewhat randomized
> > > > tick length to make sure that simultaneous ticks are
> > > > fairly rare and the vtime sampling does not end up
> > > > "in phase" with the jiffies incrementing, could make
> > > > the accounting work right again.
> > > 
> > > That improves jitter, especially on big boxen.  I have an 8 socket
> > > box
> > > that thinks it's an extra large PC, there, collision avoidance
> > > matters
> > > hugely.  I couldn't reproduce bean counting woes, no idea if
> > > collision
> > > avoidance will help that.
> > 
> > Out of curiosity, where is the main contention between ticks? I
> > indeed
> > know some locks that can be taken on special cases, such as posix cpu
> > timers.
> > 
> > Also, why does it raise power consumption issues?
> 
> On a system without either nohz_full or nohz idle
> mode, skewed ticks result in CPU cores waking up
> at different times, and keeping an idle system
> consuming power for more time than it would if all
> the ticks happened simultaneously.

Ah fair point!

> 
> This is not a factor at all on systems that switch
> off the tick while idle, since the CPU will be busy
> anyway while the tick is enabled.

I see. Thanks!

[toc] | [prev] | [next] | [standalone]


#1612427

FromWanpeng Li <kernellwp@gmail.com>
Date2017-03-30 00:50 +0200
Message-ID<tquzD-5dc-5@gated-at.bofh.it>
In reply to#1610014
2017-03-30 6:17 GMT+08:00 Frederic Weisbecker <fweisbec@gmail.com>:
> On Wed, Mar 29, 2017 at 01:16:56PM -0400, Luiz Capitulino wrote:
>> On Tue, 28 Mar 2017 13:24:06 -0400
>> Luiz Capitulino <lcapitulino@redhat.com> wrote:
>>
>> >  1. In my tracing I'm seeing that sometimes (always?) the
>> >     time interval between two timer interrupts is less than 1ms
>>
>> I think that's the root cause.
>>
>> I'm getting traces like this:
>>
>>    hog-11980 [015]   341.494491: function:             enter_from_user_mode <-- apic_timer_interrupt
>> <idle>-0     [000]   341.494492: function:             smp_apic_timer_interrupt <-- apic_timer_interrupt
>>    hog-11980 [015]   341.494492: function:             __context_tracking_exit <-- enter_from_user_mode
>> <idle>-0     [000]   341.494492: function:             irq_enter <-- smp_apic_timer_interrupt
>>    hog-11980 [015]   341.494492: bprint:               vtime_delta: diff=0 (now=4295008339 vtime_snap=4295008339)
>>    hog-11980 [015]   341.494492: function:             smp_apic_timer_interrupt <-- apic_timer_interrupt
>>    hog-11980 [015]   341.494492: function:             irq_enter <-- smp_apic_timer_interrupt
>>    hog-11980 [015]   341.494493: function:             tick_sched_timer <-- __hrtimer_run_queues
>> <idle>-0     [000]   341.494493: function:             tick_sched_timer <-- __hrtimer_run_queues
>> <idle>-0     [000]   341.494493: function:             tick_do_update_jiffies64.part.14 <-- tick_sched_do_timer
>> <idle>-0     [000]   341.494494: function:             do_timer <-- tick_do_update_jiffies64.part.14
>>    hog-11980 [015]   341.494494: function:             irq_exit <-- smp_apic_timer_interrupt
>> <idle>-0     [000]   341.494494: bprint:               do_timer: updated jiffies_64=4295008340 ticks=1
>>    hog-11980 [015]   341.494494: function:             __context_tracking_enter <-- prepare_exit_to_usermode
>>    hog-11980 [015]   341.494494: function:             vtime_user_enter <-- __context_tracking_enter
>>    hog-11980 [015]   341.494495: bprint:               vtime_delta: diff=1000000 (now=4295008340 vtime_snap=4295008339)
>>    hog-11980 [015]   341.494495: function:             __vtime_account_system <-- vtime_user_enter
>>    hog-11980 [015]   341.494495: bprint:               get_vtime_delta: vtime_snap=4295008339 now=4295008340
>>    hog-11980 [015]   341.494495: function:             account_system_time <-- __vtime_account_system
>>    hog-11980 [015]   341.494495: bprint:               account_system_time: cputime=995488
>> <idle>-0     [000]   341.494497: function:             irq_exit <-- smp_apic_timer_interrupt
>>
>> In this trace, we see the following:
>>
>>  1. On CPU15, we transition from user-space to kernel-space because
>>     of a timer interrupt (it's the tick)
>>
>>  2. vtimer_delta() returns 0, because jiffies didn't change since the
>>     last accounting
>>
>>  3. While CPU15 is executing in kernel-space, jiffies is updated
>>     by CPU0
>>
>>  4. When going back to user-space, vtime_delta() returns non-zero
>>     and the whole time is accounted for system time (observe how
>>     the cputime parameter in account_system_time() is less than 1ms)
>
> Aah, so the issue can indeed happen if all CPUs fire their ticks at the same time:
>
>
>                  CPU 0                         CPU 1
>                  -----                         -----
>                                                exit_user() // no cputime update
> tick X           update_jiffies
>                                                enter_user() // cputime update
>
>
>                                                exit_user() //no cputime update
> tick X+1         update_jiffies
>                                                enter_user() // cputime update
>
>>
>> That's why my patch from yesterday fixed the issue, it increased the
>> tick period to more than 1ms. So vtime_delta() always evaluate to true
>> when transitioning from user-space to kernel-space (because we spend
>> more than 1ms in user-space between ticks). The patch below achieves
>> the same result by adding 10us to the tick period.
>>
>> diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
>> index 7fe53be..00e46df 100644
>> --- a/kernel/time/tick-sched.c
>> +++ b/kernel/time/tick-sched.c
>> @@ -1165,7 +1165,7 @@ static enum hrtimer_restart tick_sched_timer(struct hrtimer *timer)
>>         if (unlikely(ts->tick_stopped))
>>                 return HRTIMER_NORESTART;
>>
>> -       hrtimer_forward(timer, now, tick_period);
>> +       hrtimer_forward(timer, now, tick_period + 10000);
>
> I'm surprised it works though. If the 10us shift was only applied to CPU 0 and not the
> others then yes, but if it is applied to all CPUs, the ticks stay synchronized and the
> problem should stay...
>
> Ah wait! It can work because the nohz_full CPUs have their ticks sometimes scheduled
> by tick_nohz_stop_sched_tick() or tick_nohz_restart_sched_tick() which don't have the
> 10us shift. So a drift happens everytime the nohz_full CPUs have their tick stopped.
>
>> Now, why is the tick ticking at less than 1ms? I think it's the time
>> difference between "now" (that we pass to hrtimer_forward()) and the
>> time the timer hardware is actually programmed. That should account
>> for a few microseconds.
>
> Right, that's my feeling. And if it is the case, then it shouldn't matter.
>
> So! Now we need to find a proper fix :o)
>
> Hmm, how bad would it be to revert to sched_clock() instead of jiffies in vtime_delta()?
> We could use nanosecond granularity to check deltas but only perform an actual cputime update
> when that delta >= TICK_NSEC. That should keep the load ok.

Yeah, I mentioned something similar before.
https://lkml.org/lkml/2017/3/26/138 However, Rik's commit optimized
syscalls by not utilize sched_clock(), so if we should distinguish
between syscalls/exceptions and irqs?

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1612498

FromLuiz Capitulino <lcapitulino@redhat.com>
Date2017-03-30 04:20 +0200
Message-ID<tqxQR-7M1-1@gated-at.bofh.it>
In reply to#1612427
On Thu, 30 Mar 2017 06:46:30 +0800
Wanpeng Li <kernellwp@gmail.com> wrote:

> > So! Now we need to find a proper fix :o)
> >
> > Hmm, how bad would it be to revert to sched_clock() instead of jiffies in vtime_delta()?
> > We could use nanosecond granularity to check deltas but only perform an actual cputime update
> > when that delta >= TICK_NSEC. That should keep the load ok.  
> 
> Yeah, I mentioned something similar before.
> https://lkml.org/lkml/2017/3/26/138 However, Rik's commit optimized
> syscalls by not utilize sched_clock(), so if we should distinguish
> between syscalls/exceptions and irqs?

Why not use ktime_get()?

Here's the solution I was thinking about, it's mostly untested. I'm
rate limiting below TICK_NSEC because I want to avoid syncing with
the tick.

diff --git a/kernel/sched/cputime.c b/kernel/sched/cputime.c
index f3778e2b..a8b1e85 100644
--- a/kernel/sched/cputime.c
+++ b/kernel/sched/cputime.c
@@ -676,18 +676,20 @@ void thread_group_cputime_adjusted(struct task_struct *p, u64 *ut, u64 *st)
 #ifdef CONFIG_VIRT_CPU_ACCOUNTING_GEN
 static u64 vtime_delta(struct task_struct *tsk)
 {
-	unsigned long now = READ_ONCE(jiffies);
+	return ktime_sub(ktime_get(), tsk->vtime_snap);
+}
 
-	if (time_before(now, (unsigned long)tsk->vtime_snap))
-		return 0;
+/* A little bit less than the tick period */
+#define VTIME_RATE_LIMIT (TICK_NSEC - 200000)
 
-	return jiffies_to_nsecs(now - tsk->vtime_snap);
+static bool vtime_should_account(struct task_struct *tsk)
+{
+	return vtime_delta(tsk) > VTIME_RATE_LIMIT;
 }
 
 static u64 get_vtime_delta(struct task_struct *tsk)
 {
-	unsigned long now = READ_ONCE(jiffies);
-	u64 delta, other;
+	u64 delta, other, now = ktime_get();
 
 	/*
 	 * Unlike tick based timing, vtime based timing never has lost
@@ -696,7 +698,7 @@ static u64 get_vtime_delta(struct task_struct *tsk)
 	 * elapsed time. Limit account_other_time to prevent rounding
 	 * errors from causing elapsed vtime to go negative.
 	 */
-	delta = jiffies_to_nsecs(now - tsk->vtime_snap);
+	delta = ktime_sub(now, tsk->vtime_snap);
 	other = account_other_time(delta);
 	WARN_ON_ONCE(tsk->vtime_snap_whence == VTIME_INACTIVE);
 	tsk->vtime_snap = now;
@@ -711,7 +713,7 @@ static void __vtime_account_system(struct task_struct *tsk)
 
 void vtime_account_system(struct task_struct *tsk)
 {
-	if (!vtime_delta(tsk))
+	if (!vtime_should_account(tsk))
 		return;
 
 	write_seqcount_begin(&tsk->vtime_seqcount);
@@ -723,7 +725,7 @@ void vtime_account_user(struct task_struct *tsk)
 {
 	write_seqcount_begin(&tsk->vtime_seqcount);
 	tsk->vtime_snap_whence = VTIME_SYS;
-	if (vtime_delta(tsk))
+	if (vtime_should_account(tsk))
 		account_user_time(tsk, get_vtime_delta(tsk));
 	write_seqcount_end(&tsk->vtime_seqcount);
 }
@@ -731,7 +733,7 @@ void vtime_account_user(struct task_struct *tsk)
 void vtime_user_enter(struct task_struct *tsk)
 {
 	write_seqcount_begin(&tsk->vtime_seqcount);
-	if (vtime_delta(tsk))
+	if (vtime_should_account(tsk))
 		__vtime_account_system(tsk);
 	tsk->vtime_snap_whence = VTIME_USER;
 	write_seqcount_end(&tsk->vtime_seqcount);
@@ -747,7 +749,7 @@ void vtime_guest_enter(struct task_struct *tsk)
 	 * that can thus safely catch up with a tickless delta.
 	 */
 	write_seqcount_begin(&tsk->vtime_seqcount);
-	if (vtime_delta(tsk))
+	if (vtime_should_account(tsk))
 		__vtime_account_system(tsk);
 	current->flags |= PF_VCPU;
 	write_seqcount_end(&tsk->vtime_seqcount);
@@ -776,7 +778,7 @@ void arch_vtime_task_switch(struct task_struct *prev)
 
 	write_seqcount_begin(&current->vtime_seqcount);
 	current->vtime_snap_whence = VTIME_SYS;
-	current->vtime_snap = jiffies;
+	current->vtime_snap = ktime_get();
 	write_seqcount_end(&current->vtime_seqcount);
 }
 

[toc] | [prev] | [next] | [standalone]


#1613020

FromWanpeng Li <kernellwp@gmail.com>
Date2017-03-30 14:30 +0200
Message-ID<tqHnb-6iA-7@gated-at.bofh.it>
In reply to#1612498
2017-03-30 10:14 GMT+08:00 Luiz Capitulino <lcapitulino@redhat.com>:
> On Thu, 30 Mar 2017 06:46:30 +0800
> Wanpeng Li <kernellwp@gmail.com> wrote:
>
>> > So! Now we need to find a proper fix :o)
>> >
>> > Hmm, how bad would it be to revert to sched_clock() instead of jiffies in vtime_delta()?
>> > We could use nanosecond granularity to check deltas but only perform an actual cputime update
>> > when that delta >= TICK_NSEC. That should keep the load ok.
>>
>> Yeah, I mentioned something similar before.
>> https://lkml.org/lkml/2017/3/26/138 However, Rik's commit optimized
>> syscalls by not utilize sched_clock(), so if we should distinguish
>> between syscalls/exceptions and irqs?
>
> Why not use ktime_get()?

I believe ktime_get() is more heavy than local_clock() when sched
clock is stable. So we can cooperate to improve
https://lkml.org/lkml/2017/3/30/456.

Regards,
Wanpeng Li

>
> Here's the solution I was thinking about, it's mostly untested. I'm
> rate limiting below TICK_NSEC because I want to avoid syncing with
> the tick.
>
> diff --git a/kernel/sched/cputime.c b/kernel/sched/cputime.c
> index f3778e2b..a8b1e85 100644
> --- a/kernel/sched/cputime.c
> +++ b/kernel/sched/cputime.c
> @@ -676,18 +676,20 @@ void thread_group_cputime_adjusted(struct task_struct *p, u64 *ut, u64 *st)
>  #ifdef CONFIG_VIRT_CPU_ACCOUNTING_GEN
>  static u64 vtime_delta(struct task_struct *tsk)
>  {
> -       unsigned long now = READ_ONCE(jiffies);
> +       return ktime_sub(ktime_get(), tsk->vtime_snap);
> +}
>
> -       if (time_before(now, (unsigned long)tsk->vtime_snap))
> -               return 0;
> +/* A little bit less than the tick period */
> +#define VTIME_RATE_LIMIT (TICK_NSEC - 200000)
>
> -       return jiffies_to_nsecs(now - tsk->vtime_snap);
> +static bool vtime_should_account(struct task_struct *tsk)
> +{
> +       return vtime_delta(tsk) > VTIME_RATE_LIMIT;
>  }
>
>  static u64 get_vtime_delta(struct task_struct *tsk)
>  {
> -       unsigned long now = READ_ONCE(jiffies);
> -       u64 delta, other;
> +       u64 delta, other, now = ktime_get();
>
>         /*
>          * Unlike tick based timing, vtime based timing never has lost
> @@ -696,7 +698,7 @@ static u64 get_vtime_delta(struct task_struct *tsk)
>          * elapsed time. Limit account_other_time to prevent rounding
>          * errors from causing elapsed vtime to go negative.
>          */
> -       delta = jiffies_to_nsecs(now - tsk->vtime_snap);
> +       delta = ktime_sub(now, tsk->vtime_snap);
>         other = account_other_time(delta);
>         WARN_ON_ONCE(tsk->vtime_snap_whence == VTIME_INACTIVE);
>         tsk->vtime_snap = now;
> @@ -711,7 +713,7 @@ static void __vtime_account_system(struct task_struct *tsk)
>
>  void vtime_account_system(struct task_struct *tsk)
>  {
> -       if (!vtime_delta(tsk))
> +       if (!vtime_should_account(tsk))
>                 return;
>
>         write_seqcount_begin(&tsk->vtime_seqcount);
> @@ -723,7 +725,7 @@ void vtime_account_user(struct task_struct *tsk)
>  {
>         write_seqcount_begin(&tsk->vtime_seqcount);
>         tsk->vtime_snap_whence = VTIME_SYS;
> -       if (vtime_delta(tsk))
> +       if (vtime_should_account(tsk))
>                 account_user_time(tsk, get_vtime_delta(tsk));
>         write_seqcount_end(&tsk->vtime_seqcount);
>  }
> @@ -731,7 +733,7 @@ void vtime_account_user(struct task_struct *tsk)
>  void vtime_user_enter(struct task_struct *tsk)
>  {
>         write_seqcount_begin(&tsk->vtime_seqcount);
> -       if (vtime_delta(tsk))
> +       if (vtime_should_account(tsk))
>                 __vtime_account_system(tsk);
>         tsk->vtime_snap_whence = VTIME_USER;
>         write_seqcount_end(&tsk->vtime_seqcount);
> @@ -747,7 +749,7 @@ void vtime_guest_enter(struct task_struct *tsk)
>          * that can thus safely catch up with a tickless delta.
>          */
>         write_seqcount_begin(&tsk->vtime_seqcount);
> -       if (vtime_delta(tsk))
> +       if (vtime_should_account(tsk))
>                 __vtime_account_system(tsk);
>         current->flags |= PF_VCPU;
>         write_seqcount_end(&tsk->vtime_seqcount);
> @@ -776,7 +778,7 @@ void arch_vtime_task_switch(struct task_struct *prev)
>
>         write_seqcount_begin(&current->vtime_seqcount);
>         current->vtime_snap_whence = VTIME_SYS;
> -       current->vtime_snap = jiffies;
> +       current->vtime_snap = ktime_get();
>         write_seqcount_end(&current->vtime_seqcount);
>  }
>

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web