Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1631226 > unrolled thread

x86-tip tsc/tick gripage

Started byMike Galbraith <efault@gmx.de>
First post2017-04-26 10:20 +0200
Last post2017-04-26 10:30 +0200
Articles 20 on this page of 26 — 5 participants

Back to article view | Back to linux.kernel


Contents

  x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 10:20 +0200
    Re: x86-tip tsc/tick gripage Ingo Molnar <mingo@kernel.org> - 2017-04-26 10:30 +0200
      Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 10:40 +0200
        Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 11:10 +0200
          Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 11:50 +0200
          Re: x86-tip tsc/tick gripage Peter Zijlstra <peterz@infradead.org> - 2017-04-26 12:30 +0200
            Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 13:50 +0200
              Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 14:40 +0200
                Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 20:20 +0200
                Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-27 07:30 +0200
                [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() Mike Galbraith <efault@gmx.de> - 2017-04-29 18:10 +0200
                  Re: [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-04-29 20:10 +0200
                    Re: [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() Mike Galbraith <efault@gmx.de> - 2017-04-29 20:30 +0200
                      Re: [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-04-29 23:50 +0200
                        Re: [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() Mike Galbraith <efault@gmx.de> - 2017-04-30 03:30 +0200
                          Re: [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-04-30 05:50 +0200
                            Re: [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() Mike Galbraith <efault@gmx.de> - 2017-04-30 06:30 +0200
                              Re: [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-04-30 06:40 +0200
                                Re: [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-05-01 00:50 +0200
                                  Re: [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() Thomas Gleixner <tglx@linutronix.de> - 2017-05-01 10:00 +0200
                                    Re: [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-05-01 21:40 +0200
                              Re: [patch] timer: Fix timers_update_migration(), and call it in  tmigr_init() Mike Galbraith <efault@gmx.de> - 2017-04-30 07:10 +0200
              [patch] Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-28 09:40 +0200
                Re: [patch] Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-28 10:50 +0200
                  Re: [patch] Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-28 16:20 +0200
    Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 10:30 +0200

Page 1 of 2  [1] 2  Next page →


#1631226 — x86-tip tsc/tick gripage

FromMike Galbraith <efault@gmx.de>
Date2017-04-26 10:20 +0200
Subjectx86-tip tsc/tick gripage
Message-ID<tAql4-84y-7@gated-at.bofh.it>
Greetings,

After picking up the pieces of my tip-rt tree, I'm seeing grumbling on
two boxen, and it ain't me, the below is virgin tip.  The second box
has crap BIOS (replacement ready to be installed), but works fine if
the sync code gets the things synchronized before giving up (bumping
loop count a wee bit makes that reliable).  This first box is my
primary RT test box, not a racehorse (old nag), but highly reliable.

tip v4.11-rc8-893-g8ec9e12aff06, trusty ole 8 socket (X7560) DL980 G7

[   11.856619] tsc: Refined TSC clocksource calibration: 2260.999 MHz
[   11.867703] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns
[   12.951432] clocksource: Switched to clocksource tsc
[  255.674776] clocksource: timekeeping watchdog on CPU54: Marking clocksource 'tsc' as unstable because the skew is too large:
[  255.895200] clocksource:                       'tsc' cs_now: 11b2a29f798 cs_last: c4faa1e90a mask: ffffffffffffffff
[  256.105724] tsc: Marking TSC unstable due to clocksource watchdog
[  316.976942] ------------[ cut here ]------------
[  316.980923] WARNING: CPU: 0 PID: 0 at kernel/time/tick-sched.c:874 tick_nohz_stop_sched_tick+0x339/0x380
[  316.980923] Modules linked in: autofs4(E) edd(E) af_packet(E) cpufreq_conservative(E) cpufreq_userspace(E) cpufreq_powersave(E) fuse(E) loop(E) md_mod(E) dm_mod(E) vhost_net(E) vhost(E) tap(E) tun(E) kvm_intel(E) iTCO_wdt(E) kvm(E) gpio_ich(E) iTCO_vendor_support(E) i7core_edac(E) ipmi_ssif(E) joydev(E) lpc_ich(E) edac_core(E) hpilo(E) bnx2(E) netxen_nic(E) sr_mod(E) hid_generic(E) hpwdt(E) irqbypass(E) mfd_core(E) shpchp(E) ipmi_si(E) pcc_cpufreq(E) ehci_pci(E) ipmi_msghandler(E) acpi_cpufreq(E) cdrom(E) sg(E) acpi_power_meter(E) pcspkr(E) button(E) ext4(E) mbcache(E) jbd2(E) crc16(E) usbhid(E) uhci_hcd(E) ehci_hcd(E) usbcore(E) thermal(E) sd_mod(E) scsi_dh_hp_sw(E) scsi_dh_emc(E) scsi_dh_rdac(E) scsi_dh_alua(E) ata_generic(E) ata_piix(E) libata(E) hpsa(E) scsi_transport_sas(E) cciss(E) scsi_mod(E)
[  316.980923] CPU: 0 PID: 0 Comm: swapper/0 Tainted: G            E   4.11.0-tip-default #57
[  316.980923] Hardware name: Hewlett-Packard ProLiant DL980 G7, BIOS P66 07/07/2010
[  316.980923] task: ffffffff81c104c0 task.stack: ffffffff81c00000
[  316.980923] RIP: 0010:tick_nohz_stop_sched_tick+0x339/0x380
[  316.980923] RSP: 0018:ffff88027fc03f40 EFLAGS: 00010093
[  316.980923] RAX: 0000000000000001 RBX: ffff88027fc15d20 RCX: 00000049cc0cd28a
[  316.980923] RDX: 00000068e05e6aa0 RSI: ffff88017e286010 RDI: ffff88027fc15e00
[  316.980923] RBP: 00000049cc0cbf00 R08: ffffffffffffffdd R09: 0000000040001000
[  316.980923] R10: 00000001000010b3 R11: 0000000000000000 R12: 00000049cc0cf0de
[  316.980923] R13: ffff88027fc0d640 R14: 00000049a9b7af00 R15: 00000049a9b7af00
[  316.980923] FS:  0000000000000000(0000) GS:ffff88027fc00000(0000) knlGS:0000000000000000
[  316.980923] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  316.980923] CR2: 00007f229a37e000 CR3: 0000000001c09000 CR4: 00000000000006f0
[  316.980923] Call Trace:
[  316.980923]  <IRQ>
[  316.980923]  ? __tick_nohz_idle_enter+0x8e/0x150
[  316.980923]  ? irq_exit+0x77/0xe0
[  316.980923]  ? do_IRQ+0x4c/0xd0
[  316.980923]  ? common_interrupt+0x8c/0x8c
[  316.980923]  </IRQ>
[  316.980923]  ? poll_idle+0x2d/0x58
[  316.980923]  ? cpuidle_enter_state+0x9d/0x260
[  316.980923]  ? do_idle+0x165/0x1d0
[  316.980923]  ? cpu_startup_entry+0x5d/0x70
[  316.980923]  ? start_kernel+0x481/0x48c
[  316.980923]  ? set_init_arg+0x50/0x50
[  316.980923]  ? early_idt_handler_array+0x120/0x120
[  316.980923]  ? x86_64_start_kernel+0x147/0x156
[  316.980923]  ? secondary_startup_64+0x9f/0x9f
[  316.980923] Code: 49 39 cf 7c 26 48 89 c8 e9 94 fe ff ff f3 90 e9 03 fd ff ff 83 7b 48 02 75 d9 48 89 df e8 d0 1a ff ff 49 8b 45 18 e9 76 fe ff ff <0f> ff 80 3d ee 59 cd 00 00 0f 85 8d fe ff ff 31 c0 4c 89 fa 48 
[  316.980923] ---[ end trace 996f71f82094d582 ]---
[  316.980923] basemono: 316956000000 ts->next_tick: 316380000000 dev->next_event: 316956005002

tip v4.11-rc8-887-gefb93ce741c6, 2 socket E7-8890 v3 (72 core) box

[   29.473651] tsc: Refined TSC clocksource calibration: 2493.986 MHz
[   29.473701] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x23f308a7388, max_idle_ns: 440795280393 ns
[   30.506471] clocksource: Switched to clocksource tsc
[  716.538171] clocksource: timekeeping watchdog on CPU67: Marking clocksource 'tsc' as unstable because the skew is too large:
[  716.584589] clocksource:                       'tsc' cs_now: d171cd5ce006e cs_last: d165a9ce422d1 mask: ffffffffffffffff
[  716.609048] tsc: Marking TSC unstable due to clocksource watchdog
[ 1067.093433] ------------[ cut here ]------------
[ 1067.097416] WARNING: CPU: 0 PID: 0 at kernel/time/tick-sched.c:874 tick_nohz_stop_sched_tick+0x339/0x380
[ 1067.097416] Modules linked in: af_packet(E) iscsi_ibft(E) intel_rapl(E) iscsi_boot_sysfs(E) sb_edac(E) edac_core(E) x86_pkg_temp_thermal(E) intel_powerclamp(E) coretemp(E) kvm_intel(E) kvm(E) ext4(E) crc16(E) irqbypass(E) jbd2(E) mbcache(E) joydev(E) crct10dif_pclmul(E) crc32_pclmul(E) ghash_clmulni_intel(E) pcbc(E) ipmi_ssif(E) ixgbe(E) iTCO_wdt(E) aesni_intel(E) i40e(E) iTCO_vendor_support(E) aes_x86_64(E) mdio(E) crypto_simd(E) lpc_ich(E) mptctl(E) ptp(E) glue_helper(E) ipmi_si(E) cryptd(E) pcspkr(E) pps_core(E) dca(E) mfd_core(E) i2c_i801(E) mptbase(E) ipmi_devintf(E) ipmi_msghandler(E) wmi(E) button(E) acpi_pad(E) shpchp(E) sunrpc(E) btrfs(E) xor(E) raid6_pq(E) sd_mod(E) hid_generic(E) usbhid(E) sr_mod(E) cdrom(E) i2c_algo_bit(E) drm_kms_helper(E) syscopyarea(E) sysfillrect(E) sysimgblt(E) ahci(E)
[ 1067.097416]  fb_sys_fops(E) ehci_pci(E) libahci(E) ttm(E) ehci_hcd(E) crc32c_intel(E) libata(E) drm(E) usbcore(E) mpt3sas(E) raid_class(E) scsi_transport_sas(E) sg(E) dm_multipath(E) dm_mod(E) scsi_dh_rdac(E) scsi_dh_emc(E) scsi_dh_alua(E) scsi_mod(E) autofs4(E)
[ 1067.097416] CPU: 0 PID: 0 Comm: swapper/0 Tainted: G            E   4.11.0-tip-default #226
[ 1067.097416] Hardware name: Intel Corporation BRICKLAND/BRICKLAND, BIOS BRHSXSD1.86B.0056.R01.1409242327 09/24/2014
[ 1067.097416] task: ffffffff81c104c0 task.stack: ffffffff81c00000
[ 1067.097416] RIP: 0010:tick_nohz_stop_sched_tick+0x339/0x380
[ 1067.097416] RSP: 0018:ffff88085f803f40 EFLAGS: 00010087
[ 1067.097416] RAX: 0000000000000001 RBX: ffff88085f815d20 RCX: 000000f86d81fe00
[ 1067.097416] RDX: 800000f86d70d2ff RSI: 0000000000000048 RDI: ffff88085f815e00
[ 1067.097416] RBP: 000000f86d70d300 R08: ffffffffffffff2c R09: 000000004002ed00
[ 1067.097416] R10: 000000010002edd8 R11: 0000000000000000 R12: 000000f86d82121a
[ 1067.097416] R13: ffff88085f80d640 R14: 000000de6fa7af00 R15: 000000de6fa7af00
[ 1067.097416] FS:  0000000000000000(0000) GS:ffff88085f800000(0000) knlGS:0000000000000000
[ 1067.097416] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 1067.097416] CR2: 00007fba363082a7 CR3: 000000017c3f9000 CR4: 00000000001406f0
[ 1067.097416] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
[ 1067.097416] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 0000000000000400
[ 1067.097416] Call Trace:
[ 1067.097416]  <IRQ>
[ 1067.097416]  ? __tick_nohz_idle_enter+0x8e/0x150
[ 1067.097416]  ? irq_exit+0x77/0xe0
[ 1067.097416]  ? do_IRQ+0x4c/0xd0
[ 1067.097416]  ? common_interrupt+0x8c/0x8c
[ 1067.097416]  </IRQ>
[ 1067.097416]  ? cpuidle_enter_state+0xe3/0x260
[ 1067.097416]  ? do_idle+0x165/0x1d0
[ 1067.097416]  ? cpu_startup_entry+0x5d/0x70
[ 1067.097416]  ? start_kernel+0x481/0x48c
[ 1067.097416]  ? set_init_arg+0x50/0x50
[ 1067.097416]  ? early_idt_handler_array+0x120/0x120
[ 1067.097416]  ? x86_64_start_kernel+0x147/0x156
[ 1067.097416]  ? secondary_startup_64+0x9f/0x9f
[ 1067.097416] Code: 49 39 cf 7c 26 48 89 c8 e9 94 fe ff ff f3 90 e9 03 fd ff ff 83 7b 48 02 75 d9 48 89 df e8 d0 1a ff ff 49 8b 45 18 e9 76 fe ff ff <0f> ff 80 3d 6e 5a cd 00 00 0f 85 8d fe ff ff 31 c0 4c 89 fa 48 
[ 1067.097416] ---[ end trace 65463b720a8bd23d ]---
[ 1067.097416] basemono: 1066988000000 ts->next_tick: 955356000000 dev->next_event: 1066989125120

[toc] | [next] | [standalone]


#1631235

FromIngo Molnar <mingo@kernel.org>
Date2017-04-26 10:30 +0200
Message-ID<tAquJ-87M-5@gated-at.bofh.it>
In reply to#1631226
* Mike Galbraith <efault@gmx.de> wrote:

> On Wed, 2017-04-26 at 10:02 +0200, Mike Galbraith wrote:
> 
> > tip v4.11-rc8-893-g8ec9e12aff06, trusty ole 8 socket (X7560) DL980 G7
> 
> Ew, DL980 then turned into unhappy RCU camper.
> 
> [  316.980923] basemono: 316956000000 ts->next_tick: 316380000000 dev->next_event: 316956005002
> [  689.893122] INFO: rcu_sched detected stalls on CPUs/tasks:

Probably RCU unrelated.

I have temporarily removed the current timers/urgent lineup from -tip:

 098991fccfc7: nohz: Print more debug info in tick_nohz_stop_sched_tick()
 22aa2ad45fd8: tick: Make sure tick timer is active when bypassing reprogramming
 d58bd60c773d: nohz: Fix again collision between tick and other hrtimers

... and have reintegrated tip:master, so it should be back to working I think.

Does this solve the warning and the RCU stalls on your boxes?

Thanks,

	Ingo

[toc] | [prev] | [next] | [standalone]


#1631246

FromMike Galbraith <efault@gmx.de>
Date2017-04-26 10:40 +0200
Message-ID<tAqEq-8cm-23@gated-at.bofh.it>
In reply to#1631235
On Wed, 2017-04-26 at 10:21 +0200, Ingo Molnar wrote:

> I have temporarily removed the current timers/urgent lineup from -tip:
> 
>  098991fccfc7: nohz: Print more debug info in tick_nohz_stop_sched_tick()
>  22aa2ad45fd8: tick: Make sure tick timer is active when bypassing reprogramming
>  d58bd60c773d: nohz: Fix again collision between tick and other hrtimers
> 
> ... and have reintegrated tip:master, so it should be back to working I think.
> 
> Does this solve the warning and the RCU stalls on your boxes?

(both boxen building)

[toc] | [prev] | [next] | [standalone]


#1631265

FromMike Galbraith <efault@gmx.de>
Date2017-04-26 11:10 +0200
Message-ID<tAr7s-bV-21@gated-at.bofh.it>
In reply to#1631246
On Wed, 2017-04-26 at 10:31 +0200, Mike Galbraith wrote:
> On Wed, 2017-04-26 at 10:21 +0200, Ingo Molnar wrote:
> 
> > I have temporarily removed the current timers/urgent lineup from -tip:
> > 
> >  098991fccfc7: nohz: Print more debug info in tick_nohz_stop_sched_tick()
> >  22aa2ad45fd8: tick: Make sure tick timer is active when bypassing reprogramming
> >  d58bd60c773d: nohz: Fix again collision between tick and other hrtimers
> > 
> > ... and have reintegrated tip:master, so it should be back to working I think.
> > 
> > Does this solve the warning and the RCU stalls on your boxes?
> 
> (both boxen building)

(oh come on, boot already... zzz)

Both still lose their TSC.

[   11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz
[   11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns
[   13.064172] clocksource: Switched to clocksource tsc
[  240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large:
[  240.462501] clocksource:                       'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff
[  240.675057] tsc: Marking TSC unstable due to clocksource watchdog

[  349.192460] clocksource: timekeeping watchdog on CPU63: Marking clocksource 'tsc' as unstable because the skew is too large:
[  349.217756] clocksource:                       'hpet' wd_now: 2bb0dad7 wd_last: 5d9ac241 mask: ffffffff
[  349.239021] clocksource:                       'tsc' cs_now: d2413ebfc1c28 cs_last: d2387b328a38c mask: ffffffffffffffff
[  349.263525] sched_clock: Marking unstable (349228856508, 34664619)<-(351146178141, -1882657014)
[  349.283245] tsc: Marking TSC unstable due to clocksource watchdog
[  349.298075] clocksource: Switched to clocksource hpet

[toc] | [prev] | [next] | [standalone]


#1631304

FromMike Galbraith <efault@gmx.de>
Date2017-04-26 11:50 +0200
Message-ID<tArKa-tv-41@gated-at.bofh.it>
In reply to#1631265
On Wed, 2017-04-26 at 10:57 +0200, Mike Galbraith wrote:
> On Wed, 2017-04-26 at 10:31 +0200, Mike Galbraith wrote:
> > On Wed, 2017-04-26 at 10:21 +0200, Ingo Molnar wrote:
> > 
> > > I have temporarily removed the current timers/urgent lineup from -tip:
> > > 
> > >  098991fccfc7: nohz: Print more debug info in tick_nohz_stop_sched_tick()
> > >  22aa2ad45fd8: tick: Make sure tick timer is active when bypassing reprogramming
> > >  d58bd60c773d: nohz: Fix again collision between tick and other hrtimers
> > > 
> > > ... and have reintegrated tip:master, so it should be back to working I think.
> > > 
> > > Does this solve the warning and the RCU stalls on your boxes?
> > 
> > (both boxen building)
> 
> (oh come on, boot already... zzz)
> 
> Both still lose their TSC.

Aw crap, and the DL980 RCU stalled while idling again.

> [   11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz
> [   11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns
> [   13.064172] clocksource: Switched to clocksource tsc
> [  240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large:
> [  240.462501] clocksource:                       'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff
> [  240.675057] tsc: Marking TSC unstable due to clocksource watchdog
> 
> [  349.192460] clocksource: timekeeping watchdog on CPU63: Marking clocksource 'tsc' as unstable because the skew is too large:
> [  349.217756] clocksource:                       'hpet' wd_now: 2bb0dad7 wd_last: 5d9ac241 mask: ffffffff
> [  349.239021] clocksource:                       'tsc' cs_now: d2413ebfc1c28 cs_last: d2387b328a38c mask: ffffffffffffffff
> [  349.263525] sched_clock: Marking unstable (349228856508, 34664619)<-(351146178141, -1882657014)
> [  349.283245] tsc: Marking TSC unstable due to clocksource watchdog
> [  349.298075] clocksource: Switched to clocksource hpet

[toc] | [prev] | [next] | [standalone]


#1631358

FromPeter Zijlstra <peterz@infradead.org>
Date2017-04-26 12:30 +0200
Message-ID<tAsmS-Y5-15@gated-at.bofh.it>
In reply to#1631265
On Wed, Apr 26, 2017 at 10:57:42AM +0200, Mike Galbraith wrote:

> Both still lose their TSC.
> 
> [   11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz
> [   11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns
> [   13.064172] clocksource: Switched to clocksource tsc
> [  240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large:
> [  240.462501] clocksource:                       'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff
> [  240.675057] tsc: Marking TSC unstable due to clocksource watchdog


And they didn't use to? We don't typically write to TSC or TSC_ADJUST
and thus would not cause such behaviour.

[toc] | [prev] | [next] | [standalone]


#1631412

FromMike Galbraith <efault@gmx.de>
Date2017-04-26 13:50 +0200
Message-ID<tAtCh-1In-1@gated-at.bofh.it>
In reply to#1631358
On Wed, 2017-04-26 at 12:26 +0200, Peter Zijlstra wrote:
> On Wed, Apr 26, 2017 at 10:57:42AM +0200, Mike Galbraith wrote:
> 
> > Both still lose their TSC.
> > 
> > [   11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz
> > [   11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns
> > [   13.064172] clocksource: Switched to clocksource tsc
> > [  240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large:
> > [  240.462501] clocksource:                       'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff
> > [  240.675057] tsc: Marking TSC unstable due to clocksource watchdog
> 
> 
> And they didn't use to? We don't typically write to TSC or TSC_ADJUST
> and thus would not cause such behaviour.

Nope.  The DL980 is my RT jitter test box, which I use all the time. 
 The other makes very spiffy jitter numbers when TSCs get synched up
(crap bios).. well used to before suse maintenance elves got around to
doing the BIOS replacement ;-/  (poor box, say hi to the goldfish)

	-Mike

[toc] | [prev] | [next] | [standalone]


#1631437

FromMike Galbraith <efault@gmx.de>
Date2017-04-26 14:40 +0200
Message-ID<tAuoF-2hj-13@gated-at.bofh.it>
In reply to#1631412
On Wed, 2017-04-26 at 13:39 +0200, Mike Galbraith wrote:
> On Wed, 2017-04-26 at 12:26 +0200, Peter Zijlstra wrote:
> > On Wed, Apr 26, 2017 at 10:57:42AM +0200, Mike Galbraith wrote:
> > 
> > > Both still lose their TSC.
> > > 
> > > [   11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz
> > > [   11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns
> > > [   13.064172] clocksource: Switched to clocksource tsc
> > > [  240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large:
> > > [  240.462501] clocksource:                       'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff
> > > [  240.675057] tsc: Marking TSC unstable due to clocksource watchdog
> > 
> > 
> > And they didn't use to? We don't typically write to TSC or TSC_ADJUST
> > and thus would not cause such behaviour.
> 
> Nope.

DL980 seems perfectly happy with master.today.. so off we go.

[toc] | [prev] | [next] | [standalone]


#1631638

FromMike Galbraith <efault@gmx.de>
Date2017-04-26 20:20 +0200
Message-ID<tAzHH-5SI-3@gated-at.bofh.it>
In reply to#1631437
On Wed, 2017-04-26 at 14:30 +0200, Mike Galbraith wrote:
> On Wed, 2017-04-26 at 13:39 +0200, Mike Galbraith wrote:
> > On Wed, 2017-04-26 at 12:26 +0200, Peter Zijlstra wrote:
> > > On Wed, Apr 26, 2017 at 10:57:42AM +0200, Mike Galbraith wrote:
> > > 
> > > > Both still lose their TSC.
> > > > 
> > > > [   11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz
> > > > [   11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns
> > > > [   13.064172] clocksource: Switched to clocksource tsc
> > > > [  240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large:
> > > > [  240.462501] clocksource:                       'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff
> > > > [  240.675057] tsc: Marking TSC unstable due to clocksource watchdog
> > > 
> > > 
> > > And they didn't use to? We don't typically write to TSC or TSC_ADJUST
> > > and thus would not cause such behaviour.
> > 
> > Nope.
> 
> DL980 seems perfectly happy with master.today.. so off we go.

ec2206b91d430da57c856869b5a37dc1e569e80d is the first bad commit
commit ec2206b91d430da57c856869b5a37dc1e569e80d
Author: Anna-Maria Gleixner <anna-maria@linutronix.de>
Date:   Mon Mar 20 10:34:20 2017 +0100

    timer: Implement the hierarchical pull model
    
    Placing timers at enqueue time on a target CPU based on dubious heuristics
    does not make any sense:
    
     1) Most timer wheel timers are canceled or rearmed before they expire.
    
     2) The heuristics to predict which CPU will be busy when the timer expires
        are wrong by definition.
    
    So we waste precious cycles to place timers at enqueue time.
    
    The proper solution to this problem is to always queue the timers on the
    local CPU and allow the non pinned timers to be pulled onto a busy CPU at
    expiry time.
    
    To achieve this the timer storage has been split into local pinned and
    global timers. Local pinned timers are always expired on the CPU on which
    they have been queued. Global timers can be expired on any CPU.
    
    As long as a CPU is busy it expires both local and global timers. When a
    CPU goes idle it arms for the first expiring local timer. If the first
    expiring pinned (local) timer is before the first expiring movable timer,
    then no action is required because the CPU will wake up before the first
    movable timer expires. If the first expiring movable timer is before the
    first expiring pinned (local) timer, then this timer is queued into a idle
    timerqueue and eventually expired by some other active CPU.
    
    To avoid global locking the timerqueues are implemented as a hierarchy. The
    lowest level of the hierarchy holds the CPUs. The CPUs are associated to
    groups of 8, which are seperated per node. If more than one CPU group
    exist, then a second level in the hierarchy collects the groups. Depending
    on the size of the system more than 2 levels are required. Each group has a
    "migrator" which checks the timerqueue during the tick for remote expirable
    timers.
    
    If the last CPU in a group goes idle it reports the first expiring event in
    the group up to the next group(s) in the hierarchy. If the last CPU goes
    idle it arms its timer for the first system wide expiring timer to ensure
    that no timer event is missed.
    
    Signed-off-by: Anna-Maria Gleixner <anna-maria@linutronix.de>
    Signed-off-by: Thomas Gleixner <tglx@linutronix.de>

:040000 040000 76086a473919931466dfeea76a17566e2bb15130 5432a1bbc5ace9796e70e91447c6f26eb9eb51be M	include
:040000 040000 102292e74099b7c3e0b2125a4deef728267b4d9d 24385b3c68140fbb3362f32753f296d127e22097 M	kernel

git bisect start
# good: [ea839b41744dffe5c77b8d9842c9bb7073460901] Merge tag 'arc-4.11-final' of git://git.kernel.org/pub/scm/linux/kernel/git/vgupta/arc
git bisect good ea839b41744dffe5c77b8d9842c9bb7073460901
# bad: [d02c59825bbfd2726430f410321f9d0bfc9d1ef4] Merge branch 'WIP.x86/fpu'
git bisect bad d02c59825bbfd2726430f410321f9d0bfc9d1ef4
# good: [1212249de8142df8c64c4165d1dee8ed4c3d766c] Merge branch 'perf/core'
git bisect good 1212249de8142df8c64c4165d1dee8ed4c3d766c
# good: [28a68c6ee0837004c87845a09717a99b4e96f7d4] Merge branch 'x86-mm-for-xen'
git bisect good 28a68c6ee0837004c87845a09717a99b4e96f7d4
# good: [7fd018291abbc076dc8073d48bba9614ca67b818] Merge branch 'x86/cleanups'
git bisect good 7fd018291abbc076dc8073d48bba9614ca67b818
# good: [6f029d98bc76cf9aa21e73e4beba9014e7e3adbb] Merge branch 'x86/platform'
git bisect good 6f029d98bc76cf9aa21e73e4beba9014e7e3adbb
# bad: [4691eff5d6475014b4c04481b03aff89f5ec87c1] Merge branch 'WIP.timers'
git bisect bad 4691eff5d6475014b4c04481b03aff89f5ec87c1
# good: [0605ab6fd5aa0c78133fc6611fa5e4f17e46c396] sched/wait: Disambiguate wq_entry->task_list and wq_head->task_list naming
git bisect good 0605ab6fd5aa0c78133fc6611fa5e4f17e46c396
# bad: [1e1c48e6befca01abb2ed224a76fbe0ea8c73e30] timer: Always queue timers on the local CPU
git bisect bad 1e1c48e6befca01abb2ed224a76fbe0ea8c73e30
# good: [c82a8b6b216e6910f36a94939de7e8793b25feaa] timer: Retrieve next expiry of pinned/non-pinned timers seperately
git bisect good c82a8b6b216e6910f36a94939de7e8793b25feaa
# good: [e2e1214438bb25ec4e37b8aedb3d3e502535b09b] tick/sched: Split out jiffies update helper function
git bisect good e2e1214438bb25ec4e37b8aedb3d3e502535b09b
# bad: [270c4e558cc26a0d89e89dda46c35ac39457a896] timer_migration: Add tracepoints
git bisect bad 270c4e558cc26a0d89e89dda46c35ac39457a896
# bad: [ec2206b91d430da57c856869b5a37dc1e569e80d] timer: Implement the hierarchical pull model
git bisect bad ec2206b91d430da57c856869b5a37dc1e569e80d

[toc] | [prev] | [next] | [standalone]


#1631864

FromMike Galbraith <efault@gmx.de>
Date2017-04-27 07:30 +0200
Message-ID<tAKa5-4nJ-1@gated-at.bofh.it>
In reply to#1631437
On Wed, 2017-04-26 at 14:30 +0200, Mike Galbraith wrote:
> On Wed, 2017-04-26 at 13:39 +0200, Mike Galbraith wrote:
> > On Wed, 2017-04-26 at 12:26 +0200, Peter Zijlstra wrote:
> > > On Wed, Apr 26, 2017 at 10:57:42AM +0200, Mike Galbraith wrote:
> > > 
> > > > Both still lose their TSC.
> > > > 
> > > > [   11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz
> > > > [   11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns
> > > > [   13.064172] clocksource: Switched to clocksource tsc
> > > > [  240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large:
> > > > [  240.462501] clocksource:                       'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff
> > > > [  240.675057] tsc: Marking TSC unstable due to clocksource watchdog
> > > 
> > > 
> > > And they didn't use to? We don't typically write to TSC or TSC_ADJUST
> > > and thus would not cause such behaviour.
> > 
> > Nope.
> 
> DL980 seems perfectly happy with master.today.. so off we go.

hm, this bit of huge trace looks less than wonderful.

          <idle>-0     [041] ..s2   317.304657: timer_expire_entry: timer=ffffffff820d6600 function=clocksource_watchdog now=4294971392
          <idle>-0     [041] d.s4   317.304660: timer_start: timer=ffffffff820d6600 function=clocksource_watchdog expires=4294916631 [timeout=-54761] cpu=42 idx=19 flags=
          <idle>-0     [041] ..s2   317.304660: timer_expire_exit: timer=ffffffff820d6600                                             ^^^^^^^^^^^^^^

1.1 megalines later, we finally meet function=clocksource_watchdog
again, and have a cow.
          <idle>-0     [043] d.s3   489.511620: timer_cancel: timer=ffffffff820d6600
          <idle>-0     [043] ..s2   489.511621: timer_expire_entry: timer=ffffffff820d6600 function=clocksource_watchdog now=4295014443
          <idle>-0     [043] ..s2   489.511628: clocksource_watchdog: timekeeping watchdog on CPU43: Marking clocksource 'tsc' as unstable because the skew is too large:
          <idle>-0     [043] ..s2   489.511630: clocksource_watchdog:                       'hpet' wd_now: a1cbfa1a wd_last: ed7acfe mask: ffffffff
          <idle>-0     [043] ..s2   489.511630: clocksource_watchdog:                       'tsc' cs_now: 1c24ee60f22 cs_last: 167a92e836f mask: ffffffffffffffff

[toc] | [prev] | [next] | [standalone]


#1633307 — [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()

FromMike Galbraith <efault@gmx.de>
Date2017-04-29 18:10 +0200
Subject[patch] timer: Fix timers_update_migration(), and call it in tmigr_init()
Message-ID<tBD6y-8qU-15@gated-at.bofh.it>
In reply to#1631437
Note: there is more.  With this applied, my desktop box will no longer
reproduce when booted to init 3 with nowatchdog on the command line. 
 My 8 socket DL980 OTOH still will, though it takes longer, and is
seemingly no longer interested in following up  with a permanent RCU
stall after the tsc clocksource is killed, as it does in virgin source.

---

timers_update_migration() is called by tick_nohz_activate() before
the late initcall tmigr_init() sets tmigr_enabled to true, resulting
in it updating neither timer_base.nohz_active nor .migration_enabled,
meaning we'll not kick an idling cpu in add_timer_on().

Remove redundant loop avoidance such that tick_nohz_activate() updates
timer_bases[].nohz_active as intended, and call it in tmigr_init() to
update timer_bases[].migration_enabled.

Signed-off-by: Mike Galbraith <efault@gmx.de>
Fixes: ec2206b91d43 timer: Implement the hierarchical pull model
---
 kernel/time/timer.c           |    4 ----
 kernel/time/timer_migration.c |    1 +
 2 files changed, 1 insertion(+), 4 deletions(-)

--- a/kernel/time/timer.c
+++ b/kernel/time/timer.c
@@ -224,10 +224,6 @@ void timers_update_migration(bool update
 	bool on = sysctl_timer_migration && tick_nohz_active && tmigr_enabled;
 	unsigned int cpu;
 
-	/* Avoid the loop, if nothing to update */
-	if (this_cpu_read(timer_bases[BASE_GLOBAL].migration_enabled) == on)
-		return;
-
 	for_each_possible_cpu(cpu) {
 		per_cpu(timer_bases[BASE_LOCAL].migration_enabled, cpu) = on;
 		per_cpu(timer_bases[BASE_GLOBAL].migration_enabled, cpu) = on;
--- a/kernel/time/timer_migration.c
+++ b/kernel/time/timer_migration.c
@@ -649,6 +649,7 @@ static int __init tmigr_init(void)
 		goto hp_err;
 
 	tmigr_enabled = true;
+	timers_update_migration(false);
 	pr_info("Timer migration: %d hierarchy levels\n", tmigr_hierarchy_levels);
 	return 0;
 

[toc] | [prev] | [next] | [standalone]


#1633317 — Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-04-29 20:10 +0200
SubjectRe: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()
Message-ID<tBEYG-15g-13@gated-at.bofh.it>
In reply to#1633307
On Sat, Apr 29, 2017 at 06:06:40PM +0200, Mike Galbraith wrote:
> Note: there is more.  With this applied, my desktop box will no longer
> reproduce when booted to init 3 with nowatchdog on the command line. 
>  My 8 socket DL980 OTOH still will, though it takes longer, and is
> seemingly no longer interested in following up  with a permanent RCU
> stall after the tsc clocksource is killed, as it does in virgin source.
> 
> ---
> 
> timers_update_migration() is called by tick_nohz_activate() before
> the late initcall tmigr_init() sets tmigr_enabled to true, resulting
> in it updating neither timer_base.nohz_active nor .migration_enabled,
> meaning we'll not kick an idling cpu in add_timer_on().
> 
> Remove redundant loop avoidance such that tick_nohz_activate() updates
> timer_bases[].nohz_active as intended, and call it in tmigr_init() to
> update timer_bases[].migration_enabled.
> 
> Signed-off-by: Mike Galbraith <efault@gmx.de>
> Fixes: ec2206b91d43 timer: Implement the hierarchical pull model

If someone will either repost a fresh series or point me at exactly
the set of patches to use, I will run it through rcutorture again.

							Thanx, Paul

> ---
>  kernel/time/timer.c           |    4 ----
>  kernel/time/timer_migration.c |    1 +
>  2 files changed, 1 insertion(+), 4 deletions(-)
> 
> --- a/kernel/time/timer.c
> +++ b/kernel/time/timer.c
> @@ -224,10 +224,6 @@ void timers_update_migration(bool update
>  	bool on = sysctl_timer_migration && tick_nohz_active && tmigr_enabled;
>  	unsigned int cpu;
> 
> -	/* Avoid the loop, if nothing to update */
> -	if (this_cpu_read(timer_bases[BASE_GLOBAL].migration_enabled) == on)
> -		return;
> -
>  	for_each_possible_cpu(cpu) {
>  		per_cpu(timer_bases[BASE_LOCAL].migration_enabled, cpu) = on;
>  		per_cpu(timer_bases[BASE_GLOBAL].migration_enabled, cpu) = on;
> --- a/kernel/time/timer_migration.c
> +++ b/kernel/time/timer_migration.c
> @@ -649,6 +649,7 @@ static int __init tmigr_init(void)
>  		goto hp_err;
> 
>  	tmigr_enabled = true;
> +	timers_update_migration(false);
>  	pr_info("Timer migration: %d hierarchy levels\n", tmigr_hierarchy_levels);
>  	return 0;
> 
> 

[toc] | [prev] | [next] | [standalone]


#1633320 — Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()

FromMike Galbraith <efault@gmx.de>
Date2017-04-29 20:30 +0200
SubjectRe: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()
Message-ID<tBFi2-1bq-3@gated-at.bofh.it>
In reply to#1633317
On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote:

> If someone will either repost a fresh series or point me at exactly
> the set of patches to use, I will run it through rcutorture again.

Patchlet is against x86-tip/master.today.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1633351 — Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-04-29 23:50 +0200
SubjectRe: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()
Message-ID<tBIpz-30B-9@gated-at.bofh.it>
In reply to#1633320
On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote:
> On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote:
> 
> > If someone will either repost a fresh series or point me at exactly
> > the set of patches to use, I will run it through rcutorture again.
> 
> Patchlet is against x86-tip/master.today.

So today's (as in Saturday April 29) x86-tip/master with the following
patch applied?

							Thanx, Paul

------------------------------------------------------------------------

timers_update_migration() is called by tick_nohz_activate() before
the late initcall tmigr_init() sets tmigr_enabled to true, resulting
in it updating neither timer_base.nohz_active nor .migration_enabled,
meaning we'll not kick an idling cpu in add_timer_on().

Remove redundant loop avoidance such that tick_nohz_activate() updates
timer_bases[].nohz_active as intended, and call it in tmigr_init() to
update timer_bases[].migration_enabled.

Signed-off-by: Mike Galbraith <efault@gmx.de>
Fixes: ec2206b91d43 timer: Implement the hierarchical pull model
---
 kernel/time/timer.c           |    4 ----
 kernel/time/timer_migration.c |    1 +
 2 files changed, 1 insertion(+), 4 deletions(-)

--- a/kernel/time/timer.c
+++ b/kernel/time/timer.c
@@ -224,10 +224,6 @@ void timers_update_migration(bool update
 	bool on = sysctl_timer_migration && tick_nohz_active && tmigr_enabled;
 	unsigned int cpu;

-	/* Avoid the loop, if nothing to update */
-	if (this_cpu_read(timer_bases[BASE_GLOBAL].migration_enabled) == on)
-		return;
-
 	for_each_possible_cpu(cpu) {
 		per_cpu(timer_bases[BASE_LOCAL].migration_enabled, cpu) = on;
 		per_cpu(timer_bases[BASE_GLOBAL].migration_enabled, cpu) = on;
--- a/kernel/time/timer_migration.c
+++ b/kernel/time/timer_migration.c
@@ -649,6 +649,7 @@ static int __init tmigr_init(void)
 		goto hp_err;

 	tmigr_enabled = true;
+	timers_update_migration(false);
 	pr_info("Timer migration: %d hierarchy levels\n", tmigr_hierarchy_levels);
 	return 0;

[toc] | [prev] | [next] | [standalone]


#1633368 — Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()

FromMike Galbraith <efault@gmx.de>
Date2017-04-30 03:30 +0200
SubjectRe: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()
Message-ID<tBLQu-5bK-5@gated-at.bofh.it>
In reply to#1633351
On Sat, 2017-04-29 at 14:45 -0700, Paul E. McKenney wrote:
> On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote:
> > On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote:
> > 
> > > If someone will either repost a fresh series or point me at exactly
> > > the set of patches to use, I will run it through rcutorture again.
> > 
> > Patchlet is against x86-tip/master.today.
> 
> So today's (as in Saturday April 29) x86-tip/master with the following
> patch applied?

Yeah.

[toc] | [prev] | [next] | [standalone]


#1633374 — Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-04-30 05:50 +0200
SubjectRe: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()
Message-ID<tBO1X-6vb-1@gated-at.bofh.it>
In reply to#1633368
On Sun, Apr 30, 2017 at 03:21:58AM +0200, Mike Galbraith wrote:
> On Sat, 2017-04-29 at 14:45 -0700, Paul E. McKenney wrote:
> > On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote:
> > > On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote:
> > > 
> > > > If someone will either repost a fresh series or point me at exactly
> > > > the set of patches to use, I will run it through rcutorture again.
> > > 
> > > Patchlet is against x86-tip/master.today.
> > 
> > So today's (as in Saturday April 29) x86-tip/master with the following
> > patch applied?
> 
> Yeah.

OK, will fire it up once the current set of overnight tests complete.

							Thanx, Paul

[toc] | [prev] | [next] | [standalone]


#1633375 — Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()

FromMike Galbraith <efault@gmx.de>
Date2017-04-30 06:30 +0200
SubjectRe: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()
Message-ID<tBOEF-72b-1@gated-at.bofh.it>
In reply to#1633374
On Sat, 2017-04-29 at 20:43 -0700, Paul E. McKenney wrote:
> On Sun, Apr 30, 2017 at 03:21:58AM +0200, Mike Galbraith wrote:
> > On Sat, 2017-04-29 at 14:45 -0700, Paul E. McKenney wrote:
> > > On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote:
> > > > On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote:
> > > > 
> > > > > If someone will either repost a fresh series or point me at exactly
> > > > > the set of patches to use, I will run it through rcutorture again.
> > > > 
> > > > Patchlet is against x86-tip/master.today.
> > > 
> > > So today's (as in Saturday April 29) x86-tip/master with the following
> > > patch applied?
> > 
> > Yeah.
> 
> OK, will fire it up once the current set of overnight tests complete.

I certainly don't want to discourage you from beating hell outta tip,
just want to make sure you know that I'm seeing zero RCU woes, only
late timer expiry (sharpening rocks/sticks to focus trace).

	-Mike

[toc] | [prev] | [next] | [standalone]


#1633376 — Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-04-30 06:40 +0200
SubjectRe: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()
Message-ID<tBOOl-765-3@gated-at.bofh.it>
In reply to#1633375
On Sun, Apr 30, 2017 at 06:20:15AM +0200, Mike Galbraith wrote:
> On Sat, 2017-04-29 at 20:43 -0700, Paul E. McKenney wrote:
> > On Sun, Apr 30, 2017 at 03:21:58AM +0200, Mike Galbraith wrote:
> > > On Sat, 2017-04-29 at 14:45 -0700, Paul E. McKenney wrote:
> > > > On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote:
> > > > > On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote:
> > > > > 
> > > > > > If someone will either repost a fresh series or point me at exactly
> > > > > > the set of patches to use, I will run it through rcutorture again.
> > > > > 
> > > > > Patchlet is against x86-tip/master.today.
> > > > 
> > > > So today's (as in Saturday April 29) x86-tip/master with the following
> > > > patch applied?
> > > 
> > > Yeah.
> > 
> > OK, will fire it up once the current set of overnight tests complete.
> 
> I certainly don't want to discourage you from beating hell outta tip,
> just want to make sure you know that I'm seeing zero RCU woes, only
> late timer expiry (sharpening rocks/sticks to focus trace).

I got timer_migration splats from an earlier rcutorture run.  Please see
message-ID <20170421192853.GD3956@linux.vnet.ibm.com> on LKML on April
21st in reply to Thomas's V2 00/10 cover letter.  So I am curious to
learn if your patches fix them.

							Thanx, Paul

[toc] | [prev] | [next] | [standalone]


#1633503 — Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-05-01 00:50 +0200
SubjectRe: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()
Message-ID<tC5Pc-F1-7@gated-at.bofh.it>
In reply to#1633376
On Sat, Apr 29, 2017 at 09:36:37PM -0700, Paul E. McKenney wrote:
> On Sun, Apr 30, 2017 at 06:20:15AM +0200, Mike Galbraith wrote:
> > On Sat, 2017-04-29 at 20:43 -0700, Paul E. McKenney wrote:
> > > On Sun, Apr 30, 2017 at 03:21:58AM +0200, Mike Galbraith wrote:
> > > > On Sat, 2017-04-29 at 14:45 -0700, Paul E. McKenney wrote:
> > > > > On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote:
> > > > > > On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote:
> > > > > > 
> > > > > > > If someone will either repost a fresh series or point me at exactly
> > > > > > > the set of patches to use, I will run it through rcutorture again.
> > > > > > 
> > > > > > Patchlet is against x86-tip/master.today.
> > > > > 
> > > > > So today's (as in Saturday April 29) x86-tip/master with the following
> > > > > patch applied?
> > > > 
> > > > Yeah.
> > > 
> > > OK, will fire it up once the current set of overnight tests complete.
> > 
> > I certainly don't want to discourage you from beating hell outta tip,
> > just want to make sure you know that I'm seeing zero RCU woes, only
> > late timer expiry (sharpening rocks/sticks to focus trace).
> 
> I got timer_migration splats from an earlier rcutorture run.  Please see
> message-ID <20170421192853.GD3956@linux.vnet.ibm.com> on LKML on April
> 21st in reply to Thomas's V2 00/10 cover letter.  So I am curious to
> learn if your patches fix them.

And sadly, the splats are still there.  Please see the following for
the relevant console output and .config files:

http://www2.rdrop.com/users/paulmck/submission/TREE04.2017.04.30a.config
http://www2.rdrop.com/users/paulmck/submission/TREE04.2017.04.30a.console.log
http://www2.rdrop.com/users/paulmck/submission/TREE04.3.2017.04.30a.console.log

http://www2.rdrop.com/users/paulmck/submission/TREE07.2017.04.30a.config
http://www2.rdrop.com/users/paulmck/submission/TREE07.2.2017.04.30a.bzImage
http://www2.rdrop.com/users/paulmck/submission/TREE07.2017.04.30a.console.log
http://www2.rdrop.com/users/paulmck/submission/TREE07.2.2017.04.30a.console.log

Please let me know if you have any trouble accessing these.

Here is the first splat from the first TREE04 run:

[    3.310642] WARNING: CPU: 1 PID: 0 at /home/paulmck/public_git/timer-tip/kernel/time/timer_migration.c:387 tmigr_set_cpu_active+0xc6/0xe0
[    3.313210] Modules linked in:
[    3.313861] CPU: 1 PID: 0 Comm: swapper/1 Not tainted 4.11.0-rc8+ #1
[    3.315196] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS Ubuntu-1.8.2-1ubuntu1 04/01/2014
[    3.317196] task: ffff8c1fde9f0000 task.stack: ffff8e4fc0124000
[    3.318433] RIP: 0010:tmigr_set_cpu_active+0xc6/0xe0
[    3.319464] RSP: 0000:ffff8e4fc0127e90 EFLAGS: 00010046
[    3.320598] RAX: 0000000000000004 RBX: 0000000000000001 RCX: 000000000000001f
[    3.322146] RDX: 0000000000000001 RSI: ffff8c1fdfc54cc8 RDI: ffff8c1fdeb26f80
[    3.323652] RBP: ffff8e4fc0127ea8 R08: 0000000000000000 R09: 0000000000000008
[    3.325237] R10: ffff8e4fc0127e80 R11: 0000000000000400 R12: ffff8c1fdeb26f80
[    3.326699] R13: ffff8c1fdfc54cc8 R14: 0000000000000000 R15: 0000000000000000
[    3.328149] FS:  0000000000000000(0000) GS:ffff8c1fdfc40000(0000) knlGS:0000000000000000
[    3.329845] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[    3.331078] CR2: ffff8e4fc02f0000 CR3: 0000000015e0a000 CR4: 00000000000006e0
[    3.332509] Call Trace:
[    3.333107]  tmigr_cpu_activate+0x36/0x40
[    3.333972]  tick_nohz_idle_exit+0xd1/0xf0
[    3.334845]  do_idle+0x113/0x170
[    3.335501]  cpu_startup_entry+0x18/0x20
[    3.336338]  start_secondary+0xe8/0xf0
[    3.337147]  secondary_startup_64+0x9f/0x9f
[    3.337998] Code: d0 48 8b 03 48 85 c0 75 eb eb a0 49 8b 7c 24 50 41 89 5c 24 08 48 85 ff 74 8c 49 8d 74 24 20 89 da e8 3f ff ff ff e9 7b ff ff ff <0f> ff 41 c6 04 24 00 5b 41 5c 41 5d 5d c3 66 90 66 2e 0f 1f 8

This is the first WARN_ON() in tmigr_set_cpu_active().  I got four splats
in 12 hours of running the rcutorture TREE04 test scenario, that is, three
runs of four hours each.

The TREE07 runs fared worse, with many more splats, starting with a
page fault.  The scripting claimed a hang, but that looks to have instead
been so many splats that the test failed to terminate itself in time.
I ran two TREE07 runs of four hours each.

							Thanx, Paul

[toc] | [prev] | [next] | [standalone]


#1633571 — Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()

FromThomas Gleixner <tglx@linutronix.de>
Date2017-05-01 10:00 +0200
SubjectRe: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init()
Message-ID<tCeps-6id-3@gated-at.bofh.it>
In reply to#1633503
On Sun, 30 Apr 2017, Paul E. McKenney wrote:
> 
> And sadly, the splats are still there.  Please see the following for
> the relevant console output and .config files:

Right. We are working on fixing the issues. I'll zap it from tip for now.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.kernel


csiph-web