Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1393179 > unrolled thread

Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP

Started byPeter Zijlstra <peterz@infradead.org>
First post2016-05-03 10:40 +0200
Last post2016-05-04 14:00 +0200
Articles 20 on this page of 23 — 4 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP Peter Zijlstra <peterz@infradead.org> - 2016-05-03 10:40 +0200
    Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP Wanpeng Li <kernellwp@gmail.com> - 2016-05-03 11:20 +0200
      Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP Wanpeng Li <kernellwp@gmail.com> - 2016-05-03 11:30 +0200
        Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rafael@kernel.org> - 2016-05-03 15:40 +0200
          Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP Wanpeng Li <kernellwp@gmail.com> - 2016-05-04 03:00 +0200
            Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP Wanpeng Li <kernellwp@gmail.com> - 2016-05-04 13:50 +0200
            Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rafael@kernel.org> - 2016-05-04 13:50 +0200
    Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rafael@kernel.org> - 2016-05-03 14:20 +0200
      Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rafael@kernel.org> - 2016-05-03 15:00 +0200
        Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rafael@kernel.org> - 2016-05-03 15:00 +0200
          Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rafael@kernel.org> - 2016-05-03 15:30 +0200
            Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-05-03 16:00 +0200
              Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-05-03 17:10 +0200
                Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP Wanpeng Li <kernellwp@gmail.com> - 2016-05-05 07:10 +0200
                  Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-05-05 15:50 +0200
                    Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP Wanpeng Li <kernellwp@gmail.com> - 2016-05-06 09:10 +0200
                      Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-05-06 14:30 +0200
      Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP Wanpeng Li <kernellwp@gmail.com> - 2016-05-04 03:00 +0200
        Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rafael@kernel.org> - 2016-05-04 13:50 +0200
          Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP Wanpeng Li <kernellwp@gmail.com> - 2016-05-04 14:00 +0200
            Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP "Rafael J. Wysocki" <rafael@kernel.org> - 2016-05-04 14:00 +0200
              Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP Wanpeng Li <kernellwp@gmail.com> - 2016-05-04 14:10 +0200
    [PATCH] intel_pstate: Fix intel_pstate_get() "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-05-04 14:00 +0200

Page 1 of 2  [1] 2  Next page →


#1393179 — Re: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-03 10:40 +0200
SubjectRe: [lkp] [sched/fair] 41e0d37f7a: divide error: 0000 [#1] SMP
Message-ID<ruE25-4BC-5@gated-at.bofh.it>
On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
> FYI, we noticed the following commit:
> 
> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")


> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
> [   14.884474] random: systemd urandom read with 5 bits of entropy available
> [   14.903975] divide error: 0000 [#1] SMP 
> [   14.908375] Modules linked in:
> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
> [   15.018359] Stack:
> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
> [   15.045493] Call Trace:
> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77 
> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
> [   15.138875]  RSP <ffff88081ab23d70>
> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
> [   15.149323] Kernel panic - not syncing: Fatal exception
> 

That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
so.

Rafael?

[toc] | [next] | [standalone]


#1393222

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-03 11:20 +0200
Message-ID<ruEEO-5sx-17@gated-at.bofh.it>
In reply to#1393179
2016-05-03 16:32 GMT+08:00 Peter Zijlstra <peterz@infradead.org>:
> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>> FYI, we noticed the following commit:
>>
>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>
>
>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
>> [   14.903975] divide error: 0000 [#1] SMP
>> [   14.908375] Modules linked in:
>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
>> [   15.018359] Stack:
>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
>> [   15.045493] Call Trace:
>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>> [   15.138875]  RSP <ffff88081ab23d70>
>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
>> [   15.149323] Kernel panic - not syncing: Fatal exception
>>
>
> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
> so.

I think one sample should be called during intel_pstate driver
initialization, how about the below patch(untested)?

----snip----

diff --git a/drivers/cpufreq/intel_pstate.c b/drivers/cpufreq/intel_pstate.c
index 8b5a415..57b3843 100644
--- a/drivers/cpufreq/intel_pstate.c
+++ b/drivers/cpufreq/intel_pstate.c
@@ -1241,6 +1241,7 @@ static int intel_pstate_init_cpu(unsigned int cpunum)
        intel_pstate_get_cpu_pstates(cpu);

        intel_pstate_busy_pid_reset(cpu);
+       intel_pstate_sample(cpu);

        cpu->update_util.func = intel_pstate_update_util;

[toc] | [prev] | [next] | [standalone]


#1393228

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-03 11:30 +0200
Message-ID<ruEOu-5yj-11@gated-at.bofh.it>
In reply to#1393222
2016-05-03 17:19 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
> 2016-05-03 16:32 GMT+08:00 Peter Zijlstra <peterz@infradead.org>:
>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>>> FYI, we noticed the following commit:
>>>
>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>>
>>
>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
>>> [   14.903975] divide error: 0000 [#1] SMP
>>> [   14.908375] Modules linked in:
>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
>>> [   15.018359] Stack:
>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
>>> [   15.045493] Call Trace:
>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>> [   15.138875]  RSP <ffff88081ab23d70>
>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
>>> [   15.149323] Kernel panic - not syncing: Fatal exception
>>>
>>
>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
>> so.
>
> I think one sample should be called during intel_pstate driver
> initialization, how about the below patch(untested)?
>
> ----snip----
>
> diff --git a/drivers/cpufreq/intel_pstate.c b/drivers/cpufreq/intel_pstate.c
> index 8b5a415..57b3843 100644
> --- a/drivers/cpufreq/intel_pstate.c
> +++ b/drivers/cpufreq/intel_pstate.c
> @@ -1241,6 +1241,7 @@ static int intel_pstate_init_cpu(unsigned int cpunum)
>         intel_pstate_get_cpu_pstates(cpu);
>
>         intel_pstate_busy_pid_reset(cpu);
> +       intel_pstate_sample(cpu);

intel_pstate_sample(cpu, 0);

>
>         cpu->update_util.func = intel_pstate_update_util;

[toc] | [prev] | [next] | [standalone]


#1393420

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-05-03 15:40 +0200
Message-ID<ruIIr-Cs-21@gated-at.bofh.it>
In reply to#1393228
On Tue, May 3, 2016 at 11:25 AM, Wanpeng Li <kernellwp@gmail.com> wrote:
> 2016-05-03 17:19 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
>> 2016-05-03 16:32 GMT+08:00 Peter Zijlstra <peterz@infradead.org>:
>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>>>> FYI, we noticed the following commit:
>>>>
>>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>>>
>>>
>>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
>>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
>>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
>>>> [   14.903975] divide error: 0000 [#1] SMP
>>>> [   14.908375] Modules linked in:
>>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
>>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
>>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
>>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
>>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
>>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
>>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
>>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
>>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
>>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
>>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
>>>> [   15.018359] Stack:
>>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
>>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
>>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
>>>> [   15.045493] Call Trace:
>>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
>>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
>>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
>>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
>>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
>>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
>>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
>>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
>>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
>>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
>>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
>>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>>> [   15.138875]  RSP <ffff88081ab23d70>
>>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
>>>> [   15.149323] Kernel panic - not syncing: Fatal exception
>>>>
>>>
>>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
>>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
>>> so.
>>
>> I think one sample should be called during intel_pstate driver
>> initialization, how about the below patch(untested)?
>>
>> ----snip----
>>
>> diff --git a/drivers/cpufreq/intel_pstate.c b/drivers/cpufreq/intel_pstate.c
>> index 8b5a415..57b3843 100644
>> --- a/drivers/cpufreq/intel_pstate.c
>> +++ b/drivers/cpufreq/intel_pstate.c
>> @@ -1241,6 +1241,7 @@ static int intel_pstate_init_cpu(unsigned int cpunum)
>>         intel_pstate_get_cpu_pstates(cpu);
>>
>>         intel_pstate_busy_pid_reset(cpu);
>> +       intel_pstate_sample(cpu);
>
> intel_pstate_sample(cpu, 0);
>
>>
>>         cpu->update_util.func = intel_pstate_update_util;

That would avoid the divide by 0, but the value returned by
intel_pstate_get() would still be bogus.

[toc] | [prev] | [next] | [standalone]


#1393880

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-04 03:00 +0200
Message-ID<ruTku-1Li-11@gated-at.bofh.it>
In reply to#1393420
2016-05-03 21:33 GMT+08:00 Rafael J. Wysocki <rafael@kernel.org>:
> On Tue, May 3, 2016 at 11:25 AM, Wanpeng Li <kernellwp@gmail.com> wrote:
>> 2016-05-03 17:19 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
>>> 2016-05-03 16:32 GMT+08:00 Peter Zijlstra <peterz@infradead.org>:
>>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>>>>> FYI, we noticed the following commit:
>>>>>
>>>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>>>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>>>>
>>>>
>>>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
>>>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
>>>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
>>>>> [   14.903975] divide error: 0000 [#1] SMP
>>>>> [   14.908375] Modules linked in:
>>>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
>>>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
>>>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
>>>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
>>>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
>>>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
>>>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
>>>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
>>>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
>>>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
>>>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>>>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
>>>>> [   15.018359] Stack:
>>>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
>>>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
>>>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
>>>>> [   15.045493] Call Trace:
>>>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
>>>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
>>>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
>>>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
>>>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
>>>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
>>>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
>>>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
>>>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
>>>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
>>>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
>>>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>>>> [   15.138875]  RSP <ffff88081ab23d70>
>>>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
>>>>> [   15.149323] Kernel panic - not syncing: Fatal exception
>>>>>
>>>>
>>>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
>>>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
>>>> so.
>>>
>>> I think one sample should be called during intel_pstate driver
>>> initialization, how about the below patch(untested)?
>>>
>>> ----snip----
>>>
>>> diff --git a/drivers/cpufreq/intel_pstate.c b/drivers/cpufreq/intel_pstate.c
>>> index 8b5a415..57b3843 100644
>>> --- a/drivers/cpufreq/intel_pstate.c
>>> +++ b/drivers/cpufreq/intel_pstate.c
>>> @@ -1241,6 +1241,7 @@ static int intel_pstate_init_cpu(unsigned int cpunum)
>>>         intel_pstate_get_cpu_pstates(cpu);
>>>
>>>         intel_pstate_busy_pid_reset(cpu);
>>> +       intel_pstate_sample(cpu);
>>
>> intel_pstate_sample(cpu, 0);
>>
>>>
>>>         cpu->update_util.func = intel_pstate_update_util;
>
> That would avoid the divide by 0, but the value returned by
> intel_pstate_get() would still be bogus.

If your bogus means that some data is stale and could you explain more?

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1394174

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-04 13:50 +0200
Message-ID<rv3tw-33X-15@gated-at.bofh.it>
In reply to#1393880
2016-05-04 19:41 GMT+08:00 Rafael J. Wysocki <rafael@kernel.org>:
> On Wed, May 4, 2016 at 2:53 AM, Wanpeng Li <kernellwp@gmail.com> wrote:
>> 2016-05-03 21:33 GMT+08:00 Rafael J. Wysocki <rafael@kernel.org>:
>>> On Tue, May 3, 2016 at 11:25 AM, Wanpeng Li <kernellwp@gmail.com> wrote:
>>>> 2016-05-03 17:19 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
>>>>> 2016-05-03 16:32 GMT+08:00 Peter Zijlstra <peterz@infradead.org>:
>>>>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>
> [cut]
>
>>>>> ----snip----
>>>>>
>>>>> diff --git a/drivers/cpufreq/intel_pstate.c b/drivers/cpufreq/intel_pstate.c
>>>>> index 8b5a415..57b3843 100644
>>>>> --- a/drivers/cpufreq/intel_pstate.c
>>>>> +++ b/drivers/cpufreq/intel_pstate.c
>>>>> @@ -1241,6 +1241,7 @@ static int intel_pstate_init_cpu(unsigned int cpunum)
>>>>>         intel_pstate_get_cpu_pstates(cpu);
>>>>>
>>>>>         intel_pstate_busy_pid_reset(cpu);
>>>>> +       intel_pstate_sample(cpu);
>>>>
>>>> intel_pstate_sample(cpu, 0);
>>>>
>>>>>
>>>>>         cpu->update_util.func = intel_pstate_update_util;
>>>
>>> That would avoid the divide by 0, but the value returned by
>>> intel_pstate_get() would still be bogus.
>>
>> If your bogus means that some data is stale and could you explain more?
>
> get_avg_frequency() expects sample.aperf to be a delta between two
> different values of the APERF register obtained at two different
> instants of time, and analogously for sample.mperf, because that's
> when the formula used by it is guaranteed to be valid.  This means
> that it generally is not sufficient to read those registers just once
> to get a meaningful result, they need to be read at least twice for
> that (with some time between the reads to let the counters grow
> sufficiently).
>
> With your modification sample.aperf and sample.mperf would simply
> contain the values of APERF and MPERF, respectively, at the the
> intel_pstate_sample(cpu, 0) invocation time, so using them in the
> computation would not be guaranteed to lead to a meaningful result.

I see, thanks Rafael.

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1394180

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-05-04 13:50 +0200
Message-ID<rv3tw-33X-17@gated-at.bofh.it>
In reply to#1393880
On Wed, May 4, 2016 at 2:53 AM, Wanpeng Li <kernellwp@gmail.com> wrote:
> 2016-05-03 21:33 GMT+08:00 Rafael J. Wysocki <rafael@kernel.org>:
>> On Tue, May 3, 2016 at 11:25 AM, Wanpeng Li <kernellwp@gmail.com> wrote:
>>> 2016-05-03 17:19 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
>>>> 2016-05-03 16:32 GMT+08:00 Peter Zijlstra <peterz@infradead.org>:
>>>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:

[cut]

>>>> ----snip----
>>>>
>>>> diff --git a/drivers/cpufreq/intel_pstate.c b/drivers/cpufreq/intel_pstate.c
>>>> index 8b5a415..57b3843 100644
>>>> --- a/drivers/cpufreq/intel_pstate.c
>>>> +++ b/drivers/cpufreq/intel_pstate.c
>>>> @@ -1241,6 +1241,7 @@ static int intel_pstate_init_cpu(unsigned int cpunum)
>>>>         intel_pstate_get_cpu_pstates(cpu);
>>>>
>>>>         intel_pstate_busy_pid_reset(cpu);
>>>> +       intel_pstate_sample(cpu);
>>>
>>> intel_pstate_sample(cpu, 0);
>>>
>>>>
>>>>         cpu->update_util.func = intel_pstate_update_util;
>>
>> That would avoid the divide by 0, but the value returned by
>> intel_pstate_get() would still be bogus.
>
> If your bogus means that some data is stale and could you explain more?

get_avg_frequency() expects sample.aperf to be a delta between two
different values of the APERF register obtained at two different
instants of time, and analogously for sample.mperf, because that's
when the formula used by it is guaranteed to be valid.  This means
that it generally is not sufficient to read those registers just once
to get a meaningful result, they need to be read at least twice for
that (with some time between the reads to let the counters grow
sufficiently).

With your modification sample.aperf and sample.mperf would simply
contain the values of APERF and MPERF, respectively, at the the
intel_pstate_sample(cpu, 0) invocation time, so using them in the
computation would not be guaranteed to lead to a meaningful result.

[toc] | [prev] | [next] | [standalone]


#1393374

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-05-03 14:20 +0200
Message-ID<ruHt1-82C-17@gated-at.bofh.it>
In reply to#1393179
On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>> FYI, we noticed the following commit:
>>
>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>
>
>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
>> [   14.903975] divide error: 0000 [#1] SMP
>> [   14.908375] Modules linked in:
>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
>> [   15.018359] Stack:
>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
>> [   15.045493] Call Trace:
>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>> [   15.138875]  RSP <ffff88081ab23d70>
>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
>> [   15.149323] Kernel panic - not syncing: Fatal exception
>>
>
> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
> so.

Well, what's the tree based on?

The mainline does this:

bool sample_taken = intel_pstate_sample(cpu, time);

if (sample_taken && !hwp_active)
        intel_pstate_adjust_busy_pstate(cpu);

and (the mainline version of) intel_pstate_sample() returns false when
it is called for the first time after setting the update_util hook.

[toc] | [prev] | [next] | [standalone]


#1393402

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-05-03 15:00 +0200
Message-ID<ruI5J-8kl-19@gated-at.bofh.it>
In reply to#1393374
On Tue, May 3, 2016 at 2:15 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>>> FYI, we noticed the following commit:
>>>
>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>>
>>
>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
>>> [   14.903975] divide error: 0000 [#1] SMP
>>> [   14.908375] Modules linked in:
>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
>>> [   15.018359] Stack:
>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
>>> [   15.045493] Call Trace:
>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>> [   15.138875]  RSP <ffff88081ab23d70>
>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
>>> [   15.149323] Kernel panic - not syncing: Fatal exception
>>>
>>
>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
>> so.
>
> Well, what's the tree based on?
>
> The mainline does this:
>
> bool sample_taken = intel_pstate_sample(cpu, time);
>
> if (sample_taken && !hwp_active)
>         intel_pstate_adjust_busy_pstate(cpu);
>
> and (the mainline version of) intel_pstate_sample() returns false when
> it is called for the first time after setting the update_util hook.

If that helps, I can expose my pm-cpufreq-fixes branch to pull from.
It contains all cpufreq material that went into the Linus' tree to
date and is based on 4.5-rc3.

[toc] | [prev] | [next] | [standalone]


#1393403

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-05-03 15:00 +0200
Message-ID<ruI5J-8kl-21@gated-at.bofh.it>
In reply to#1393402
On Tue, May 3, 2016 at 2:54 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> On Tue, May 3, 2016 at 2:15 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
>> On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>>>> FYI, we noticed the following commit:
>>>>
>>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>>>
>>>
>>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
>>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
>>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
>>>> [   14.903975] divide error: 0000 [#1] SMP
>>>> [   14.908375] Modules linked in:
>>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
>>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
>>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
>>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
>>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
>>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
>>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
>>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
>>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
>>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
>>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
>>>> [   15.018359] Stack:
>>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
>>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
>>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
>>>> [   15.045493] Call Trace:
>>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
>>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
>>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
>>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
>>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
>>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
>>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
>>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
>>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
>>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
>>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
>>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>>> [   15.138875]  RSP <ffff88081ab23d70>
>>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
>>>> [   15.149323] Kernel panic - not syncing: Fatal exception
>>>>
>>>
>>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
>>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
>>> so.
>>
>> Well, what's the tree based on?
>>
>> The mainline does this:
>>
>> bool sample_taken = intel_pstate_sample(cpu, time);
>>
>> if (sample_taken && !hwp_active)
>>         intel_pstate_adjust_busy_pstate(cpu);
>>
>> and (the mainline version of) intel_pstate_sample() returns false when
>> it is called for the first time after setting the update_util hook.
>
> If that helps, I can expose my pm-cpufreq-fixes branch to pull from.
> It contains all cpufreq material that went into the Linus' tree to
> date and is based on 4.5-rc3.

In fact, it is exposed already:

git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm.git \
pm-cpufreq-fixes

and the top-most commit is 1becf03545a0859ceaaf9e8c2d9861882a71cb01
(cpufreq: intel_pstate: Fix processing for turbo activation ratio).

[toc] | [prev] | [next] | [standalone]


#1393418

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-05-03 15:30 +0200
Message-ID<ruIyJ-xb-7@gated-at.bofh.it>
In reply to#1393403
On Tue, May 3, 2016 at 2:58 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> On Tue, May 3, 2016 at 2:54 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
>> On Tue, May 3, 2016 at 2:15 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
>>> On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>>>>> FYI, we noticed the following commit:
>>>>>
>>>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>>>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>>>>
>>>>
>>>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
>>>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
>>>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
>>>>> [   14.903975] divide error: 0000 [#1] SMP
>>>>> [   14.908375] Modules linked in:
>>>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
>>>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
>>>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
>>>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
>>>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
>>>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
>>>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
>>>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
>>>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
>>>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
>>>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>>>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
>>>>> [   15.018359] Stack:
>>>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
>>>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
>>>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
>>>>> [   15.045493] Call Trace:
>>>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
>>>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
>>>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
>>>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
>>>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
>>>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
>>>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
>>>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
>>>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
>>>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
>>>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
>>>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>>>> [   15.138875]  RSP <ffff88081ab23d70>
>>>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
>>>>> [   15.149323] Kernel panic - not syncing: Fatal exception
>>>>>
>>>>
>>>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
>>>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
>>>> so.
>>>
>>> Well, what's the tree based on?
>>>
>>> The mainline does this:
>>>
>>> bool sample_taken = intel_pstate_sample(cpu, time);
>>>
>>> if (sample_taken && !hwp_active)
>>>         intel_pstate_adjust_busy_pstate(cpu);
>>>
>>> and (the mainline version of) intel_pstate_sample() returns false when
>>> it is called for the first time after setting the update_util hook.
>>
>> If that helps, I can expose my pm-cpufreq-fixes branch to pull from.
>> It contains all cpufreq material that went into the Linus' tree to
>> date and is based on 4.5-rc3.
>
> In fact, it is exposed already:
>
> git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm.git \
> pm-cpufreq-fixes
>
> and the top-most commit is 1becf03545a0859ceaaf9e8c2d9861882a71cb01
> (cpufreq: intel_pstate: Fix processing for turbo activation ratio).

Ah, that will fail as well.

The problem is that intel_pstate_get() can be called before we take
the first sample.

I need to think about how to fix that.

[toc] | [prev] | [next] | [standalone]


#1393436

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-05-03 16:00 +0200
Message-ID<ruJ1N-LG-13@gated-at.bofh.it>
In reply to#1393418
On Tuesday, May 03, 2016 03:22:24 PM Rafael J. Wysocki wrote:
> On Tue, May 3, 2016 at 2:58 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> > On Tue, May 3, 2016 at 2:54 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> >> On Tue, May 3, 2016 at 2:15 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> >>> On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
> >>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
> >>>>> FYI, we noticed the following commit:
> >>>>>
> >>>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
> >>>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
> >>>>
> >>>>
> >>>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
> >>>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
> >>>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
> >>>>> [   14.903975] divide error: 0000 [#1] SMP
> >>>>> [   14.908375] Modules linked in:
> >>>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
> >>>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
> >>>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
> >>>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
> >>>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
> >>>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
> >>>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
> >>>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
> >>>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
> >>>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
> >>>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
> >>>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> >>>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
> >>>>> [   15.018359] Stack:
> >>>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
> >>>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
> >>>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
> >>>>> [   15.045493] Call Trace:
> >>>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
> >>>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
> >>>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
> >>>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
> >>>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
> >>>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
> >>>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
> >>>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
> >>>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
> >>>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
> >>>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
> >>>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
> >>>>> [   15.138875]  RSP <ffff88081ab23d70>
> >>>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
> >>>>> [   15.149323] Kernel panic - not syncing: Fatal exception
> >>>>>
> >>>>
> >>>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
> >>>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
> >>>> so.
> >>>
> >>> Well, what's the tree based on?
> >>>
> >>> The mainline does this:
> >>>
> >>> bool sample_taken = intel_pstate_sample(cpu, time);
> >>>
> >>> if (sample_taken && !hwp_active)
> >>>         intel_pstate_adjust_busy_pstate(cpu);
> >>>
> >>> and (the mainline version of) intel_pstate_sample() returns false when
> >>> it is called for the first time after setting the update_util hook.
> >>
> >> If that helps, I can expose my pm-cpufreq-fixes branch to pull from.
> >> It contains all cpufreq material that went into the Linus' tree to
> >> date and is based on 4.5-rc3.
> >
> > In fact, it is exposed already:
> >
> > git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm.git \
> > pm-cpufreq-fixes
> >
> > and the top-most commit is 1becf03545a0859ceaaf9e8c2d9861882a71cb01
> > (cpufreq: intel_pstate: Fix processing for turbo activation ratio).
> 
> Ah, that will fail as well.
> 
> The problem is that intel_pstate_get() can be called before we take
> the first sample.
> 
> I need to think about how to fix that.

Maybe something like the below (untested, but builds).

It will make intel_pstate_get() return 0 until avg_frequency gets populated
which is actually OK.

---
 drivers/cpufreq/intel_pstate.c |    6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

Index: linux-pm/drivers/cpufreq/intel_pstate.c
===================================================================
--- linux-pm.orig/drivers/cpufreq/intel_pstate.c
+++ linux-pm/drivers/cpufreq/intel_pstate.c
@@ -114,6 +114,7 @@ struct cpudata {
 	u64	prev_mperf;
 	u64	prev_tsc;
 	u64	prev_cummulative_iowait;
+	int	avg_frequency;
 	struct sample sample;
 };
 
@@ -1037,6 +1038,7 @@ static inline void intel_pstate_adjust_b
 	intel_pstate_update_pstate(cpu, target_pstate);
 
 	sample = &cpu->sample;
+	cpu->avg_frequency = get_avg_frequency(cpu);
 	trace_pstate_sample(fp_toint(sample->core_pct_busy),
 		fp_toint(sample->busy_scaled),
 		from,
@@ -1044,7 +1046,7 @@ static inline void intel_pstate_adjust_b
 		sample->mperf,
 		sample->aperf,
 		sample->tsc,
-		get_avg_frequency(cpu));
+		cpu->avg_frequency);
 }
 
 static void intel_pstate_update_util(struct update_util_data *data, u64 time,
@@ -1130,7 +1132,7 @@ static unsigned int intel_pstate_get(uns
 	if (!cpu)
 		return 0;
 	sample = &cpu->sample;
-	return get_avg_frequency(cpu);
+	return cpu->avg_frequency;
 }
 
 static void intel_pstate_set_update_util_hook(unsigned int cpu_num)

[toc] | [prev] | [next] | [standalone]


#1393485

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-05-03 17:10 +0200
Message-ID<ruK7w-1Tc-35@gated-at.bofh.it>
In reply to#1393436
On Tuesday, May 03, 2016 03:53:12 PM Rafael J. Wysocki wrote:
> On Tuesday, May 03, 2016 03:22:24 PM Rafael J. Wysocki wrote:
> > On Tue, May 3, 2016 at 2:58 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> > > On Tue, May 3, 2016 at 2:54 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> > >> On Tue, May 3, 2016 at 2:15 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> > >>> On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
> > >>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
> > >>>>> FYI, we noticed the following commit:
> > >>>>>
> > >>>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
> > >>>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
> > >>>>
> > >>>>
> > >>>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
> > >>>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
> > >>>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
> > >>>>> [   14.903975] divide error: 0000 [#1] SMP
> > >>>>> [   14.908375] Modules linked in:
> > >>>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
> > >>>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
> > >>>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
> > >>>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
> > >>>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
> > >>>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
> > >>>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
> > >>>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
> > >>>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
> > >>>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
> > >>>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
> > >>>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> > >>>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
> > >>>>> [   15.018359] Stack:
> > >>>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
> > >>>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
> > >>>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
> > >>>>> [   15.045493] Call Trace:
> > >>>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
> > >>>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
> > >>>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
> > >>>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
> > >>>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
> > >>>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
> > >>>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
> > >>>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
> > >>>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
> > >>>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
> > >>>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
> > >>>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
> > >>>>> [   15.138875]  RSP <ffff88081ab23d70>
> > >>>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
> > >>>>> [   15.149323] Kernel panic - not syncing: Fatal exception
> > >>>>>
> > >>>>
> > >>>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
> > >>>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
> > >>>> so.
> > >>>
> > >>> Well, what's the tree based on?
> > >>>
> > >>> The mainline does this:
> > >>>
> > >>> bool sample_taken = intel_pstate_sample(cpu, time);
> > >>>
> > >>> if (sample_taken && !hwp_active)
> > >>>         intel_pstate_adjust_busy_pstate(cpu);
> > >>>
> > >>> and (the mainline version of) intel_pstate_sample() returns false when
> > >>> it is called for the first time after setting the update_util hook.
> > >>
> > >> If that helps, I can expose my pm-cpufreq-fixes branch to pull from.
> > >> It contains all cpufreq material that went into the Linus' tree to
> > >> date and is based on 4.5-rc3.
> > >
> > > In fact, it is exposed already:
> > >
> > > git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm.git \
> > > pm-cpufreq-fixes
> > >
> > > and the top-most commit is 1becf03545a0859ceaaf9e8c2d9861882a71cb01
> > > (cpufreq: intel_pstate: Fix processing for turbo activation ratio).
> > 
> > Ah, that will fail as well.
> > 
> > The problem is that intel_pstate_get() can be called before we take
> > the first sample.
> > 
> > I need to think about how to fix that.
> 
> Maybe something like the below (untested, but builds).
> 
> It will make intel_pstate_get() return 0 until avg_frequency gets populated
> which is actually OK.

The previous one would break the HWP case, so below is a new one (still
untested).

---
 drivers/cpufreq/intel_pstate.c |   19 +++++++++----------
 1 file changed, 9 insertions(+), 10 deletions(-)

Index: linux-pm/drivers/cpufreq/intel_pstate.c
===================================================================
--- linux-pm.orig/drivers/cpufreq/intel_pstate.c
+++ linux-pm/drivers/cpufreq/intel_pstate.c
@@ -114,6 +114,7 @@ struct cpudata {
 	u64	prev_mperf;
 	u64	prev_tsc;
 	u64	prev_cummulative_iowait;
+	int	avg_frequency;
 	struct sample sample;
 };
 
@@ -1044,7 +1045,7 @@ static inline void intel_pstate_adjust_b
 		sample->mperf,
 		sample->aperf,
 		sample->tsc,
-		get_avg_frequency(cpu));
+		cpu->avg_frequency);
 }
 
 static void intel_pstate_update_util(struct update_util_data *data, u64 time,
@@ -1056,8 +1057,11 @@ static void intel_pstate_update_util(str
 	if ((s64)delta_ns >= pid_params.sample_rate_ns) {
 		bool sample_taken = intel_pstate_sample(cpu, time);
 
-		if (sample_taken && !hwp_active)
-			intel_pstate_adjust_busy_pstate(cpu);
+		if (sample_taken) {
+			cpu->avg_frequency = get_avg_frequency(cpu);
+			if (!hwp_active)
+				intel_pstate_adjust_busy_pstate(cpu);
+		}
 	}
 }
 
@@ -1123,14 +1127,9 @@ static int intel_pstate_init_cpu(unsigne
 
 static unsigned int intel_pstate_get(unsigned int cpu_num)
 {
-	struct sample *sample;
-	struct cpudata *cpu;
+	struct cpudata *cpu = all_cpu_data[cpu_num];
 
-	cpu = all_cpu_data[cpu_num];
-	if (!cpu)
-		return 0;
-	sample = &cpu->sample;
-	return get_avg_frequency(cpu);
+	return cpu ? cpu->avg_frequency : 0;
 }
 
 static void intel_pstate_set_update_util_hook(unsigned int cpu_num)

[toc] | [prev] | [next] | [standalone]


#1394874

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-05 07:10 +0200
Message-ID<rvjHY-1MQ-15@gated-at.bofh.it>
In reply to#1393485
2016-05-03 23:10 GMT+08:00 Rafael J. Wysocki <rjw@rjwysocki.net>:
> On Tuesday, May 03, 2016 03:53:12 PM Rafael J. Wysocki wrote:
>> On Tuesday, May 03, 2016 03:22:24 PM Rafael J. Wysocki wrote:
>> > On Tue, May 3, 2016 at 2:58 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
>> > > On Tue, May 3, 2016 at 2:54 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
>> > >> On Tue, May 3, 2016 at 2:15 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
>> > >>> On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>> > >>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>> > >>>>> FYI, we noticed the following commit:
>> > >>>>>
>> > >>>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>> > >>>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>> > >>>>
>> > >>>>
>> > >>>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
>> > >>>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
>> > >>>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
>> > >>>>> [   14.903975] divide error: 0000 [#1] SMP
>> > >>>>> [   14.908375] Modules linked in:
>> > >>>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
>> > >>>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
>> > >>>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
>> > >>>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>> > >>>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
>> > >>>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
>> > >>>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
>> > >>>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
>> > >>>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
>> > >>>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
>> > >>>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
>> > >>>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>> > >>>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
>> > >>>>> [   15.018359] Stack:
>> > >>>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
>> > >>>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
>> > >>>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
>> > >>>>> [   15.045493] Call Trace:
>> > >>>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
>> > >>>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
>> > >>>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
>> > >>>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
>> > >>>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
>> > >>>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
>> > >>>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
>> > >>>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
>> > >>>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
>> > >>>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
>> > >>>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
>> > >>>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>> > >>>>> [   15.138875]  RSP <ffff88081ab23d70>
>> > >>>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
>> > >>>>> [   15.149323] Kernel panic - not syncing: Fatal exception
>> > >>>>>
>> > >>>>
>> > >>>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
>> > >>>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
>> > >>>> so.
>> > >>>
>> > >>> Well, what's the tree based on?
>> > >>>
>> > >>> The mainline does this:
>> > >>>
>> > >>> bool sample_taken = intel_pstate_sample(cpu, time);
>> > >>>
>> > >>> if (sample_taken && !hwp_active)
>> > >>>         intel_pstate_adjust_busy_pstate(cpu);
>> > >>>
>> > >>> and (the mainline version of) intel_pstate_sample() returns false when
>> > >>> it is called for the first time after setting the update_util hook.
>> > >>
>> > >> If that helps, I can expose my pm-cpufreq-fixes branch to pull from.
>> > >> It contains all cpufreq material that went into the Linus' tree to
>> > >> date and is based on 4.5-rc3.
>> > >
>> > > In fact, it is exposed already:
>> > >
>> > > git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm.git \
>> > > pm-cpufreq-fixes
>> > >
>> > > and the top-most commit is 1becf03545a0859ceaaf9e8c2d9861882a71cb01
>> > > (cpufreq: intel_pstate: Fix processing for turbo activation ratio).
>> >
>> > Ah, that will fail as well.
>> >
>> > The problem is that intel_pstate_get() can be called before we take
>> > the first sample.
>> >
>> > I need to think about how to fix that.
>>
>> Maybe something like the below (untested, but builds).
>>
>> It will make intel_pstate_get() return 0 until avg_frequency gets populated
>> which is actually OK.
>
> The previous one would break the HWP case, so below is a new one (still
> untested).

I can reproduce the bug and your patch fix it.

Tested-by: Wanpeng Li <wanpeng.li@hotmail.com>

>
> ---
>  drivers/cpufreq/intel_pstate.c |   19 +++++++++----------
>  1 file changed, 9 insertions(+), 10 deletions(-)
>
> Index: linux-pm/drivers/cpufreq/intel_pstate.c
> ===================================================================
> --- linux-pm.orig/drivers/cpufreq/intel_pstate.c
> +++ linux-pm/drivers/cpufreq/intel_pstate.c
> @@ -114,6 +114,7 @@ struct cpudata {
>         u64     prev_mperf;
>         u64     prev_tsc;
>         u64     prev_cummulative_iowait;
> +       int     avg_frequency;
>         struct sample sample;
>  };
>
> @@ -1044,7 +1045,7 @@ static inline void intel_pstate_adjust_b
>                 sample->mperf,
>                 sample->aperf,
>                 sample->tsc,
> -               get_avg_frequency(cpu));
> +               cpu->avg_frequency);
>  }
>
>  static void intel_pstate_update_util(struct update_util_data *data, u64 time,
> @@ -1056,8 +1057,11 @@ static void intel_pstate_update_util(str
>         if ((s64)delta_ns >= pid_params.sample_rate_ns) {
>                 bool sample_taken = intel_pstate_sample(cpu, time);
>
> -               if (sample_taken && !hwp_active)
> -                       intel_pstate_adjust_busy_pstate(cpu);
> +               if (sample_taken) {
> +                       cpu->avg_frequency = get_avg_frequency(cpu);
> +                       if (!hwp_active)
> +                               intel_pstate_adjust_busy_pstate(cpu);
> +               }
>         }
>  }
>
> @@ -1123,14 +1127,9 @@ static int intel_pstate_init_cpu(unsigne
>
>  static unsigned int intel_pstate_get(unsigned int cpu_num)
>  {
> -       struct sample *sample;
> -       struct cpudata *cpu;
> +       struct cpudata *cpu = all_cpu_data[cpu_num];
>
> -       cpu = all_cpu_data[cpu_num];
> -       if (!cpu)
> -               return 0;
> -       sample = &cpu->sample;
> -       return get_avg_frequency(cpu);
> +       return cpu ? cpu->avg_frequency : 0;
>  }
>
>  static void intel_pstate_set_update_util_hook(unsigned int cpu_num)
>



-- 
Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1395132

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-05-05 15:50 +0200
Message-ID<rvrPc-Wp-13@gated-at.bofh.it>
In reply to#1394874
On Thursday, May 05, 2016 01:05:41 PM Wanpeng Li wrote:
> 2016-05-03 23:10 GMT+08:00 Rafael J. Wysocki <rjw@rjwysocki.net>:
> > On Tuesday, May 03, 2016 03:53:12 PM Rafael J. Wysocki wrote:
> >> On Tuesday, May 03, 2016 03:22:24 PM Rafael J. Wysocki wrote:
> >> > On Tue, May 3, 2016 at 2:58 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> >> > > On Tue, May 3, 2016 at 2:54 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> >> > >> On Tue, May 3, 2016 at 2:15 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> >> > >>> On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
> >> > >>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
> >> > >>>>> FYI, we noticed the following commit:
> >> > >>>>>
> >> > >>>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
> >> > >>>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
> >> > >>>>
> >> > >>>>
> >> > >>>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
> >> > >>>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
> >> > >>>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
> >> > >>>>> [   14.903975] divide error: 0000 [#1] SMP
> >> > >>>>> [   14.908375] Modules linked in:
> >> > >>>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
> >> > >>>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
> >> > >>>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
> >> > >>>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
> >> > >>>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
> >> > >>>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
> >> > >>>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
> >> > >>>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
> >> > >>>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
> >> > >>>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
> >> > >>>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
> >> > >>>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> >> > >>>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
> >> > >>>>> [   15.018359] Stack:
> >> > >>>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
> >> > >>>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
> >> > >>>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
> >> > >>>>> [   15.045493] Call Trace:
> >> > >>>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
> >> > >>>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
> >> > >>>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
> >> > >>>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
> >> > >>>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
> >> > >>>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
> >> > >>>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
> >> > >>>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
> >> > >>>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
> >> > >>>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
> >> > >>>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
> >> > >>>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
> >> > >>>>> [   15.138875]  RSP <ffff88081ab23d70>
> >> > >>>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
> >> > >>>>> [   15.149323] Kernel panic - not syncing: Fatal exception
> >> > >>>>>
> >> > >>>>
> >> > >>>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
> >> > >>>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
> >> > >>>> so.
> >> > >>>
> >> > >>> Well, what's the tree based on?
> >> > >>>
> >> > >>> The mainline does this:
> >> > >>>
> >> > >>> bool sample_taken = intel_pstate_sample(cpu, time);
> >> > >>>
> >> > >>> if (sample_taken && !hwp_active)
> >> > >>>         intel_pstate_adjust_busy_pstate(cpu);
> >> > >>>
> >> > >>> and (the mainline version of) intel_pstate_sample() returns false when
> >> > >>> it is called for the first time after setting the update_util hook.
> >> > >>
> >> > >> If that helps, I can expose my pm-cpufreq-fixes branch to pull from.
> >> > >> It contains all cpufreq material that went into the Linus' tree to
> >> > >> date and is based on 4.5-rc3.
> >> > >
> >> > > In fact, it is exposed already:
> >> > >
> >> > > git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm.git \
> >> > > pm-cpufreq-fixes
> >> > >
> >> > > and the top-most commit is 1becf03545a0859ceaaf9e8c2d9861882a71cb01
> >> > > (cpufreq: intel_pstate: Fix processing for turbo activation ratio).
> >> >
> >> > Ah, that will fail as well.
> >> >
> >> > The problem is that intel_pstate_get() can be called before we take
> >> > the first sample.
> >> >
> >> > I need to think about how to fix that.
> >>
> >> Maybe something like the below (untested, but builds).
> >>
> >> It will make intel_pstate_get() return 0 until avg_frequency gets populated
> >> which is actually OK.
> >
> > The previous one would break the HWP case, so below is a new one (still
> > untested).
> 
> I can reproduce the bug and your patch fix it.
> 
> Tested-by: Wanpeng Li <wanpeng.li@hotmail.com>

Thanks!

Please also try this one:

https://patchwork.kernel.org/patch/9012861/

which is the final fix for this bug.

[toc] | [prev] | [next] | [standalone]


#1395651

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-06 09:10 +0200
Message-ID<rvI3E-FN-11@gated-at.bofh.it>
In reply to#1395132
2016-05-05 21:46 GMT+08:00 Rafael J. Wysocki <rjw@rjwysocki.net>:
> On Thursday, May 05, 2016 01:05:41 PM Wanpeng Li wrote:
>> 2016-05-03 23:10 GMT+08:00 Rafael J. Wysocki <rjw@rjwysocki.net>:
>> > On Tuesday, May 03, 2016 03:53:12 PM Rafael J. Wysocki wrote:
>> >> On Tuesday, May 03, 2016 03:22:24 PM Rafael J. Wysocki wrote:
>> >> > On Tue, May 3, 2016 at 2:58 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
>> >> > > On Tue, May 3, 2016 at 2:54 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
>> >> > >> On Tue, May 3, 2016 at 2:15 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
>> >> > >>> On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>> >> > >>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>> >> > >>>>> FYI, we noticed the following commit:
>> >> > >>>>>
>> >> > >>>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>> >> > >>>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>> >> > >>>>
>> >> > >>>>
>> >> > >>>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
>> >> > >>>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
>> >> > >>>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
>> >> > >>>>> [   14.903975] divide error: 0000 [#1] SMP
>> >> > >>>>> [   14.908375] Modules linked in:
>> >> > >>>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
>> >> > >>>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
>> >> > >>>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
>> >> > >>>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>> >> > >>>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
>> >> > >>>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
>> >> > >>>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
>> >> > >>>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
>> >> > >>>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
>> >> > >>>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
>> >> > >>>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
>> >> > >>>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>> >> > >>>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
>> >> > >>>>> [   15.018359] Stack:
>> >> > >>>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
>> >> > >>>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
>> >> > >>>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
>> >> > >>>>> [   15.045493] Call Trace:
>> >> > >>>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
>> >> > >>>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
>> >> > >>>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
>> >> > >>>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
>> >> > >>>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
>> >> > >>>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
>> >> > >>>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
>> >> > >>>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
>> >> > >>>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
>> >> > >>>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
>> >> > >>>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
>> >> > >>>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>> >> > >>>>> [   15.138875]  RSP <ffff88081ab23d70>
>> >> > >>>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
>> >> > >>>>> [   15.149323] Kernel panic - not syncing: Fatal exception
>> >> > >>>>>
>> >> > >>>>
>> >> > >>>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
>> >> > >>>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
>> >> > >>>> so.
>> >> > >>>
>> >> > >>> Well, what's the tree based on?
>> >> > >>>
>> >> > >>> The mainline does this:
>> >> > >>>
>> >> > >>> bool sample_taken = intel_pstate_sample(cpu, time);
>> >> > >>>
>> >> > >>> if (sample_taken && !hwp_active)
>> >> > >>>         intel_pstate_adjust_busy_pstate(cpu);
>> >> > >>>
>> >> > >>> and (the mainline version of) intel_pstate_sample() returns false when
>> >> > >>> it is called for the first time after setting the update_util hook.
>> >> > >>
>> >> > >> If that helps, I can expose my pm-cpufreq-fixes branch to pull from.
>> >> > >> It contains all cpufreq material that went into the Linus' tree to
>> >> > >> date and is based on 4.5-rc3.
>> >> > >
>> >> > > In fact, it is exposed already:
>> >> > >
>> >> > > git://git.kernel.org/pub/scm/linux/kernel/git/rafael/linux-pm.git \
>> >> > > pm-cpufreq-fixes
>> >> > >
>> >> > > and the top-most commit is 1becf03545a0859ceaaf9e8c2d9861882a71cb01
>> >> > > (cpufreq: intel_pstate: Fix processing for turbo activation ratio).
>> >> >
>> >> > Ah, that will fail as well.
>> >> >
>> >> > The problem is that intel_pstate_get() can be called before we take
>> >> > the first sample.
>> >> >
>> >> > I need to think about how to fix that.
>> >>
>> >> Maybe something like the below (untested, but builds).
>> >>
>> >> It will make intel_pstate_get() return 0 until avg_frequency gets populated
>> >> which is actually OK.
>> >
>> > The previous one would break the HWP case, so below is a new one (still
>> > untested).
>>
>> I can reproduce the bug and your patch fix it.
>>
>> Tested-by: Wanpeng Li <wanpeng.li@hotmail.com>
>
> Thanks!
>
> Please also try this one:
>
> https://patchwork.kernel.org/patch/9012861/
>
> which is the final fix for this bug.

The warning disappear.

Tested-by: Wanpeng Li <wanpeng.li@hotmail.com>

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1395824

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-05-06 14:30 +0200
Message-ID<rvN3j-51u-11@gated-at.bofh.it>
In reply to#1395651
On Friday, May 06, 2016 03:06:44 PM Wanpeng Li wrote:
> 2016-05-05 21:46 GMT+08:00 Rafael J. Wysocki <rjw@rjwysocki.net>:
> > On Thursday, May 05, 2016 01:05:41 PM Wanpeng Li wrote:
> >> 2016-05-03 23:10 GMT+08:00 Rafael J. Wysocki <rjw@rjwysocki.net>:
> >> > On Tuesday, May 03, 2016 03:53:12 PM Rafael J. Wysocki wrote:

[cut]

> >>
> >> I can reproduce the bug and your patch fix it.
> >>
> >> Tested-by: Wanpeng Li <wanpeng.li@hotmail.com>
> >
> > Thanks!
> >
> > Please also try this one:
> >
> > https://patchwork.kernel.org/patch/9012861/
> >
> > which is the final fix for this bug.
> 
> The warning disappear.
> 
> Tested-by: Wanpeng Li <wanpeng.li@hotmail.com>

Thank you!

[toc] | [prev] | [next] | [standalone]


#1393877

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-04 03:00 +0200
Message-ID<ruTku-1Li-5@gated-at.bofh.it>
In reply to#1393374
2016-05-03 20:15 GMT+08:00 Rafael J. Wysocki <rafael@kernel.org>:
> On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>>> FYI, we noticed the following commit:
>>>
>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>>
>>
>>> [   14.860950] Freeing unused kernel memory: 260K (ffff88103edbf000 - ffff88103ee00000)
>>> [   14.873013] systemd[1]: RTC configured in localtime, applying delta of 480 minutes to system time.
>>> [   14.884474] random: systemd urandom read with 5 bits of entropy available
>>> [   14.903975] divide error: 0000 [#1] SMP
>>> [   14.908375] Modules linked in:
>>> [   14.911793] CPU: 39 PID: 1 Comm: systemd Not tainted 4.6.0-rc4-00016-g41e0d37 #1
>>> [   14.920051] Hardware name: Intel Corporation S2600WP/S2600WP, BIOS SE5C600.86B.02.02.0002.122320131210 12/23/2013
>>> [   14.931509] task: ffff8810101d8000 ti: ffff88081ab20000 task.ti: ffff88081ab20000
>>> [   14.939862] RIP: 0010:[<ffffffff8176ad32>]  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>> [   14.949202] RSP: 0018:ffff88081ab23d70  EFLAGS: 00010006
>>> [   14.955129] RAX: 0000000000000000 RBX: 0000000000000024 RCX: ffff8808091e0300
>>> [   14.963094] RDX: 0000000000000000 RSI: 0000000000000100 RDI: 0000000000000024
>>> [   14.971057] RBP: ffff88081ab23d88 R08: 0000000000001000 R09: 00000000096a1000
>>> [   14.979022] R10: 0000000000ffff10 R11: 000000000000000f R12: 0000000000000202
>>> [   14.986984] R13: ffff88101390a040 R14: ffff88100e48e180 R15: ffff88101390a040
>>> [   14.994950] FS:  00007f66fe117880(0000) GS:ffff8810139c0000(0000) knlGS:0000000000000000
>>> [   15.003982] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>>> [   15.010393] CR2: 000055f78b760098 CR3: 000000103d759000 CR4: 00000000001406e0
>>> [   15.018359] Stack:
>>> [   15.020602]  ffffffff81764dad 0000000000000024 ffff88100e48e180 ffff88081ab23dc8
>>> [   15.028899]  ffffffff81040267 ffff88101390a0ac 0000000000000340 ffff88081ab23f20
>>> [   15.037197]  ffff88103cd7c400 ffff88100e48e180 ffff88101390a040 ffff88081ab23e30
>>> [   15.045493] Call Trace:
>>> [   15.048223]  [<ffffffff81764dad>] ? cpufreq_quick_get+0x3d/0x90
>>> [   15.054832]  [<ffffffff81040267>] show_cpuinfo+0x3c7/0x410
>>> [   15.060956]  [<ffffffff8121f5c4>] seq_read+0x2c4/0x3a0
>>> [   15.066685]  [<ffffffff81266ea8>] proc_reg_read+0x48/0x70
>>> [   15.072713]  [<ffffffff811f9d58>] __vfs_read+0x28/0xd0
>>> [   15.078451]  [<ffffffff813bab63>] ? security_file_permission+0xa3/0xc0
>>> [   15.085737]  [<ffffffff811faa97>] ? rw_verify_area+0x57/0xd0
>>> [   15.092054]  [<ffffffff811fab96>] vfs_read+0x86/0x130
>>> [   15.097691]  [<ffffffff811fbf96>] SyS_read+0x46/0xa0
>>> [   15.103234]  [<ffffffff818f71b2>] entry_SYSCALL_64_fastpath+0x1a/0xa4
>>> [   15.110421] Code: 05 dc 1b c3 00 89 ff 55 48 89 e5 48 8b 0c f8 48 85 c9 74 1f 48 63 51 1c 48 63 41 20 5d 48 0f af c2 31 d2 48 0f af 81 88 00 00 00 <48> f7 b1 90 00 00 00 c3 31 c0 5d c3 66 90 0f 1f 44 00 00 8b 77
>>> [   15.132161] RIP  [<ffffffff8176ad32>] intel_pstate_get+0x32/0x40
>>> [   15.138875]  RSP <ffff88081ab23d70>
>>> [   15.142770] ---[ end trace e5d5a8bedf5502e1 ]---
>>> [   15.149323] Kernel panic - not syncing: Fatal exception
>>>
>>
>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
>> so.
>
> Well, what's the tree based on?
>
> The mainline does this:
>
> bool sample_taken = intel_pstate_sample(cpu, time);
>
> if (sample_taken && !hwp_active)
>         intel_pstate_adjust_busy_pstate(cpu);
>
> and (the mainline version of) intel_pstate_sample() returns false when
> it is called for the first time after setting the update_util hook.

The callsites in scheduler will set time to rq_clock(rq) when trigger
sample, so when time 0 will be used even if it is set just before
setting the update_util hook?

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1394178

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-05-04 13:50 +0200
Message-ID<rv3tw-33X-31@gated-at.bofh.it>
In reply to#1393877
On Wed, May 4, 2016 at 2:58 AM, Wanpeng Li <kernellwp@gmail.com> wrote:
> 2016-05-03 20:15 GMT+08:00 Rafael J. Wysocki <rafael@kernel.org>:
>> On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>>>> FYI, we noticed the following commit:
>>>>
>>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>>>

[cut]

>>>
>>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
>>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
>>> so.
>>
>> Well, what's the tree based on?
>>
>> The mainline does this:
>>
>> bool sample_taken = intel_pstate_sample(cpu, time);
>>
>> if (sample_taken && !hwp_active)
>>         intel_pstate_adjust_busy_pstate(cpu);
>>
>> and (the mainline version of) intel_pstate_sample() returns false when
>> it is called for the first time after setting the update_util hook.
>
> The callsites in scheduler will set time to rq_clock(rq) when trigger
> sample, so when time 0 will be used even if it is set just before
> setting the update_util hook?

I'm not sure what you mean.

time=0 is special as it will cause intel_pstate_sample() to return
false on the next invocation.

[toc] | [prev] | [next] | [standalone]


#1394196

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-04 14:00 +0200
Message-ID<rv3Dd-38F-49@gated-at.bofh.it>
In reply to#1394178
2016-05-04 19:44 GMT+08:00 Rafael J. Wysocki <rafael@kernel.org>:
> On Wed, May 4, 2016 at 2:58 AM, Wanpeng Li <kernellwp@gmail.com> wrote:
>> 2016-05-03 20:15 GMT+08:00 Rafael J. Wysocki <rafael@kernel.org>:
>>> On Tue, May 3, 2016 at 10:32 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>>>> On Tue, May 03, 2016 at 09:10:51AM +0800, kernel test robot wrote:
>>>>> FYI, we noticed the following commit:
>>>>>
>>>>> https://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git sched/core
>>>>> commit 41e0d37f7ac81297c07ba311e4ad39465b8c8295 ("sched/fair: Do not call cpufreq hook unless util changed")
>>>>
>
> [cut]
>
>>>>
>>>> That's intel_pstate.c:get_avg_frequency(), which assumes mperf != 0. It
>>>> being 0 seems to suggest intel_pstate_sample() hasn't been called yet or
>>>> so.
>>>
>>> Well, what's the tree based on?
>>>
>>> The mainline does this:
>>>
>>> bool sample_taken = intel_pstate_sample(cpu, time);
>>>
>>> if (sample_taken && !hwp_active)
>>>         intel_pstate_adjust_busy_pstate(cpu);
>>>
>>> and (the mainline version of) intel_pstate_sample() returns false when
>>> it is called for the first time after setting the update_util hook.
>>
>> The callsites in scheduler will set time to rq_clock(rq) when trigger
>> sample, so when time 0 will be used even if it is set just before
>> setting the update_util hook?
>
> I'm not sure what you mean.
>
> time=0 is special as it will cause intel_pstate_sample() to return
> false on the next invocation.

Sample is driven by cpufreq_update_util() which uses rq_clock(rq) as
time parameter, so there is no opportunity to pass time 0 to
intel_pstate_sample().

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.kernel


csiph-web