Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1631226 > unrolled thread
| Started by | Mike Galbraith <efault@gmx.de> |
|---|---|
| First post | 2017-04-26 10:20 +0200 |
| Last post | 2017-04-26 10:30 +0200 |
| Articles | 20 on this page of 26 — 5 participants |
Back to article view | Back to linux.kernel
x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 10:20 +0200
Re: x86-tip tsc/tick gripage Ingo Molnar <mingo@kernel.org> - 2017-04-26 10:30 +0200
Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 10:40 +0200
Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 11:10 +0200
Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 11:50 +0200
Re: x86-tip tsc/tick gripage Peter Zijlstra <peterz@infradead.org> - 2017-04-26 12:30 +0200
Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 13:50 +0200
Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 14:40 +0200
Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 20:20 +0200
Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-27 07:30 +0200
[patch] timer: Fix timers_update_migration(), and call it in tmigr_init() Mike Galbraith <efault@gmx.de> - 2017-04-29 18:10 +0200
Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-04-29 20:10 +0200
Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() Mike Galbraith <efault@gmx.de> - 2017-04-29 20:30 +0200
Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-04-29 23:50 +0200
Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() Mike Galbraith <efault@gmx.de> - 2017-04-30 03:30 +0200
Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-04-30 05:50 +0200
Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() Mike Galbraith <efault@gmx.de> - 2017-04-30 06:30 +0200
Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-04-30 06:40 +0200
Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-05-01 00:50 +0200
Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() Thomas Gleixner <tglx@linutronix.de> - 2017-05-01 10:00 +0200
Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-05-01 21:40 +0200
Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() Mike Galbraith <efault@gmx.de> - 2017-04-30 07:10 +0200
[patch] Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-28 09:40 +0200
Re: [patch] Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-28 10:50 +0200
Re: [patch] Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-28 16:20 +0200
Re: x86-tip tsc/tick gripage Mike Galbraith <efault@gmx.de> - 2017-04-26 10:30 +0200
Page 1 of 2 [1] 2 Next page →
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-26 10:20 +0200 |
| Subject | x86-tip tsc/tick gripage |
| Message-ID | <tAql4-84y-7@gated-at.bofh.it> |
Greetings, After picking up the pieces of my tip-rt tree, I'm seeing grumbling on two boxen, and it ain't me, the below is virgin tip. The second box has crap BIOS (replacement ready to be installed), but works fine if the sync code gets the things synchronized before giving up (bumping loop count a wee bit makes that reliable). This first box is my primary RT test box, not a racehorse (old nag), but highly reliable. tip v4.11-rc8-893-g8ec9e12aff06, trusty ole 8 socket (X7560) DL980 G7 [ 11.856619] tsc: Refined TSC clocksource calibration: 2260.999 MHz [ 11.867703] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns [ 12.951432] clocksource: Switched to clocksource tsc [ 255.674776] clocksource: timekeeping watchdog on CPU54: Marking clocksource 'tsc' as unstable because the skew is too large: [ 255.895200] clocksource: 'tsc' cs_now: 11b2a29f798 cs_last: c4faa1e90a mask: ffffffffffffffff [ 256.105724] tsc: Marking TSC unstable due to clocksource watchdog [ 316.976942] ------------[ cut here ]------------ [ 316.980923] WARNING: CPU: 0 PID: 0 at kernel/time/tick-sched.c:874 tick_nohz_stop_sched_tick+0x339/0x380 [ 316.980923] Modules linked in: autofs4(E) edd(E) af_packet(E) cpufreq_conservative(E) cpufreq_userspace(E) cpufreq_powersave(E) fuse(E) loop(E) md_mod(E) dm_mod(E) vhost_net(E) vhost(E) tap(E) tun(E) kvm_intel(E) iTCO_wdt(E) kvm(E) gpio_ich(E) iTCO_vendor_support(E) i7core_edac(E) ipmi_ssif(E) joydev(E) lpc_ich(E) edac_core(E) hpilo(E) bnx2(E) netxen_nic(E) sr_mod(E) hid_generic(E) hpwdt(E) irqbypass(E) mfd_core(E) shpchp(E) ipmi_si(E) pcc_cpufreq(E) ehci_pci(E) ipmi_msghandler(E) acpi_cpufreq(E) cdrom(E) sg(E) acpi_power_meter(E) pcspkr(E) button(E) ext4(E) mbcache(E) jbd2(E) crc16(E) usbhid(E) uhci_hcd(E) ehci_hcd(E) usbcore(E) thermal(E) sd_mod(E) scsi_dh_hp_sw(E) scsi_dh_emc(E) scsi_dh_rdac(E) scsi_dh_alua(E) ata_generic(E) ata_piix(E) libata(E) hpsa(E) scsi_transport_sas(E) cciss(E) scsi_mod(E) [ 316.980923] CPU: 0 PID: 0 Comm: swapper/0 Tainted: G E 4.11.0-tip-default #57 [ 316.980923] Hardware name: Hewlett-Packard ProLiant DL980 G7, BIOS P66 07/07/2010 [ 316.980923] task: ffffffff81c104c0 task.stack: ffffffff81c00000 [ 316.980923] RIP: 0010:tick_nohz_stop_sched_tick+0x339/0x380 [ 316.980923] RSP: 0018:ffff88027fc03f40 EFLAGS: 00010093 [ 316.980923] RAX: 0000000000000001 RBX: ffff88027fc15d20 RCX: 00000049cc0cd28a [ 316.980923] RDX: 00000068e05e6aa0 RSI: ffff88017e286010 RDI: ffff88027fc15e00 [ 316.980923] RBP: 00000049cc0cbf00 R08: ffffffffffffffdd R09: 0000000040001000 [ 316.980923] R10: 00000001000010b3 R11: 0000000000000000 R12: 00000049cc0cf0de [ 316.980923] R13: ffff88027fc0d640 R14: 00000049a9b7af00 R15: 00000049a9b7af00 [ 316.980923] FS: 0000000000000000(0000) GS:ffff88027fc00000(0000) knlGS:0000000000000000 [ 316.980923] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [ 316.980923] CR2: 00007f229a37e000 CR3: 0000000001c09000 CR4: 00000000000006f0 [ 316.980923] Call Trace: [ 316.980923] <IRQ> [ 316.980923] ? __tick_nohz_idle_enter+0x8e/0x150 [ 316.980923] ? irq_exit+0x77/0xe0 [ 316.980923] ? do_IRQ+0x4c/0xd0 [ 316.980923] ? common_interrupt+0x8c/0x8c [ 316.980923] </IRQ> [ 316.980923] ? poll_idle+0x2d/0x58 [ 316.980923] ? cpuidle_enter_state+0x9d/0x260 [ 316.980923] ? do_idle+0x165/0x1d0 [ 316.980923] ? cpu_startup_entry+0x5d/0x70 [ 316.980923] ? start_kernel+0x481/0x48c [ 316.980923] ? set_init_arg+0x50/0x50 [ 316.980923] ? early_idt_handler_array+0x120/0x120 [ 316.980923] ? x86_64_start_kernel+0x147/0x156 [ 316.980923] ? secondary_startup_64+0x9f/0x9f [ 316.980923] Code: 49 39 cf 7c 26 48 89 c8 e9 94 fe ff ff f3 90 e9 03 fd ff ff 83 7b 48 02 75 d9 48 89 df e8 d0 1a ff ff 49 8b 45 18 e9 76 fe ff ff <0f> ff 80 3d ee 59 cd 00 00 0f 85 8d fe ff ff 31 c0 4c 89 fa 48 [ 316.980923] ---[ end trace 996f71f82094d582 ]--- [ 316.980923] basemono: 316956000000 ts->next_tick: 316380000000 dev->next_event: 316956005002 tip v4.11-rc8-887-gefb93ce741c6, 2 socket E7-8890 v3 (72 core) box [ 29.473651] tsc: Refined TSC clocksource calibration: 2493.986 MHz [ 29.473701] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x23f308a7388, max_idle_ns: 440795280393 ns [ 30.506471] clocksource: Switched to clocksource tsc [ 716.538171] clocksource: timekeeping watchdog on CPU67: Marking clocksource 'tsc' as unstable because the skew is too large: [ 716.584589] clocksource: 'tsc' cs_now: d171cd5ce006e cs_last: d165a9ce422d1 mask: ffffffffffffffff [ 716.609048] tsc: Marking TSC unstable due to clocksource watchdog [ 1067.093433] ------------[ cut here ]------------ [ 1067.097416] WARNING: CPU: 0 PID: 0 at kernel/time/tick-sched.c:874 tick_nohz_stop_sched_tick+0x339/0x380 [ 1067.097416] Modules linked in: af_packet(E) iscsi_ibft(E) intel_rapl(E) iscsi_boot_sysfs(E) sb_edac(E) edac_core(E) x86_pkg_temp_thermal(E) intel_powerclamp(E) coretemp(E) kvm_intel(E) kvm(E) ext4(E) crc16(E) irqbypass(E) jbd2(E) mbcache(E) joydev(E) crct10dif_pclmul(E) crc32_pclmul(E) ghash_clmulni_intel(E) pcbc(E) ipmi_ssif(E) ixgbe(E) iTCO_wdt(E) aesni_intel(E) i40e(E) iTCO_vendor_support(E) aes_x86_64(E) mdio(E) crypto_simd(E) lpc_ich(E) mptctl(E) ptp(E) glue_helper(E) ipmi_si(E) cryptd(E) pcspkr(E) pps_core(E) dca(E) mfd_core(E) i2c_i801(E) mptbase(E) ipmi_devintf(E) ipmi_msghandler(E) wmi(E) button(E) acpi_pad(E) shpchp(E) sunrpc(E) btrfs(E) xor(E) raid6_pq(E) sd_mod(E) hid_generic(E) usbhid(E) sr_mod(E) cdrom(E) i2c_algo_bit(E) drm_kms_helper(E) syscopyarea(E) sysfillrect(E) sysimgblt(E) ahci(E) [ 1067.097416] fb_sys_fops(E) ehci_pci(E) libahci(E) ttm(E) ehci_hcd(E) crc32c_intel(E) libata(E) drm(E) usbcore(E) mpt3sas(E) raid_class(E) scsi_transport_sas(E) sg(E) dm_multipath(E) dm_mod(E) scsi_dh_rdac(E) scsi_dh_emc(E) scsi_dh_alua(E) scsi_mod(E) autofs4(E) [ 1067.097416] CPU: 0 PID: 0 Comm: swapper/0 Tainted: G E 4.11.0-tip-default #226 [ 1067.097416] Hardware name: Intel Corporation BRICKLAND/BRICKLAND, BIOS BRHSXSD1.86B.0056.R01.1409242327 09/24/2014 [ 1067.097416] task: ffffffff81c104c0 task.stack: ffffffff81c00000 [ 1067.097416] RIP: 0010:tick_nohz_stop_sched_tick+0x339/0x380 [ 1067.097416] RSP: 0018:ffff88085f803f40 EFLAGS: 00010087 [ 1067.097416] RAX: 0000000000000001 RBX: ffff88085f815d20 RCX: 000000f86d81fe00 [ 1067.097416] RDX: 800000f86d70d2ff RSI: 0000000000000048 RDI: ffff88085f815e00 [ 1067.097416] RBP: 000000f86d70d300 R08: ffffffffffffff2c R09: 000000004002ed00 [ 1067.097416] R10: 000000010002edd8 R11: 0000000000000000 R12: 000000f86d82121a [ 1067.097416] R13: ffff88085f80d640 R14: 000000de6fa7af00 R15: 000000de6fa7af00 [ 1067.097416] FS: 0000000000000000(0000) GS:ffff88085f800000(0000) knlGS:0000000000000000 [ 1067.097416] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [ 1067.097416] CR2: 00007fba363082a7 CR3: 000000017c3f9000 CR4: 00000000001406f0 [ 1067.097416] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000 [ 1067.097416] DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 0000000000000400 [ 1067.097416] Call Trace: [ 1067.097416] <IRQ> [ 1067.097416] ? __tick_nohz_idle_enter+0x8e/0x150 [ 1067.097416] ? irq_exit+0x77/0xe0 [ 1067.097416] ? do_IRQ+0x4c/0xd0 [ 1067.097416] ? common_interrupt+0x8c/0x8c [ 1067.097416] </IRQ> [ 1067.097416] ? cpuidle_enter_state+0xe3/0x260 [ 1067.097416] ? do_idle+0x165/0x1d0 [ 1067.097416] ? cpu_startup_entry+0x5d/0x70 [ 1067.097416] ? start_kernel+0x481/0x48c [ 1067.097416] ? set_init_arg+0x50/0x50 [ 1067.097416] ? early_idt_handler_array+0x120/0x120 [ 1067.097416] ? x86_64_start_kernel+0x147/0x156 [ 1067.097416] ? secondary_startup_64+0x9f/0x9f [ 1067.097416] Code: 49 39 cf 7c 26 48 89 c8 e9 94 fe ff ff f3 90 e9 03 fd ff ff 83 7b 48 02 75 d9 48 89 df e8 d0 1a ff ff 49 8b 45 18 e9 76 fe ff ff <0f> ff 80 3d 6e 5a cd 00 00 0f 85 8d fe ff ff 31 c0 4c 89 fa 48 [ 1067.097416] ---[ end trace 65463b720a8bd23d ]--- [ 1067.097416] basemono: 1066988000000 ts->next_tick: 955356000000 dev->next_event: 1066989125120
[toc] | [next] | [standalone]
| From | Ingo Molnar <mingo@kernel.org> |
|---|---|
| Date | 2017-04-26 10:30 +0200 |
| Message-ID | <tAquJ-87M-5@gated-at.bofh.it> |
| In reply to | #1631226 |
* Mike Galbraith <efault@gmx.de> wrote: > On Wed, 2017-04-26 at 10:02 +0200, Mike Galbraith wrote: > > > tip v4.11-rc8-893-g8ec9e12aff06, trusty ole 8 socket (X7560) DL980 G7 > > Ew, DL980 then turned into unhappy RCU camper. > > [ 316.980923] basemono: 316956000000 ts->next_tick: 316380000000 dev->next_event: 316956005002 > [ 689.893122] INFO: rcu_sched detected stalls on CPUs/tasks: Probably RCU unrelated. I have temporarily removed the current timers/urgent lineup from -tip: 098991fccfc7: nohz: Print more debug info in tick_nohz_stop_sched_tick() 22aa2ad45fd8: tick: Make sure tick timer is active when bypassing reprogramming d58bd60c773d: nohz: Fix again collision between tick and other hrtimers ... and have reintegrated tip:master, so it should be back to working I think. Does this solve the warning and the RCU stalls on your boxes? Thanks, Ingo
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-26 10:40 +0200 |
| Message-ID | <tAqEq-8cm-23@gated-at.bofh.it> |
| In reply to | #1631235 |
On Wed, 2017-04-26 at 10:21 +0200, Ingo Molnar wrote: > I have temporarily removed the current timers/urgent lineup from -tip: > > 098991fccfc7: nohz: Print more debug info in tick_nohz_stop_sched_tick() > 22aa2ad45fd8: tick: Make sure tick timer is active when bypassing reprogramming > d58bd60c773d: nohz: Fix again collision between tick and other hrtimers > > ... and have reintegrated tip:master, so it should be back to working I think. > > Does this solve the warning and the RCU stalls on your boxes? (both boxen building)
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-26 11:10 +0200 |
| Message-ID | <tAr7s-bV-21@gated-at.bofh.it> |
| In reply to | #1631246 |
On Wed, 2017-04-26 at 10:31 +0200, Mike Galbraith wrote: > On Wed, 2017-04-26 at 10:21 +0200, Ingo Molnar wrote: > > > I have temporarily removed the current timers/urgent lineup from -tip: > > > > 098991fccfc7: nohz: Print more debug info in tick_nohz_stop_sched_tick() > > 22aa2ad45fd8: tick: Make sure tick timer is active when bypassing reprogramming > > d58bd60c773d: nohz: Fix again collision between tick and other hrtimers > > > > ... and have reintegrated tip:master, so it should be back to working I think. > > > > Does this solve the warning and the RCU stalls on your boxes? > > (both boxen building) (oh come on, boot already... zzz) Both still lose their TSC. [ 11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz [ 11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns [ 13.064172] clocksource: Switched to clocksource tsc [ 240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large: [ 240.462501] clocksource: 'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff [ 240.675057] tsc: Marking TSC unstable due to clocksource watchdog [ 349.192460] clocksource: timekeeping watchdog on CPU63: Marking clocksource 'tsc' as unstable because the skew is too large: [ 349.217756] clocksource: 'hpet' wd_now: 2bb0dad7 wd_last: 5d9ac241 mask: ffffffff [ 349.239021] clocksource: 'tsc' cs_now: d2413ebfc1c28 cs_last: d2387b328a38c mask: ffffffffffffffff [ 349.263525] sched_clock: Marking unstable (349228856508, 34664619)<-(351146178141, -1882657014) [ 349.283245] tsc: Marking TSC unstable due to clocksource watchdog [ 349.298075] clocksource: Switched to clocksource hpet
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-26 11:50 +0200 |
| Message-ID | <tArKa-tv-41@gated-at.bofh.it> |
| In reply to | #1631265 |
On Wed, 2017-04-26 at 10:57 +0200, Mike Galbraith wrote: > On Wed, 2017-04-26 at 10:31 +0200, Mike Galbraith wrote: > > On Wed, 2017-04-26 at 10:21 +0200, Ingo Molnar wrote: > > > > > I have temporarily removed the current timers/urgent lineup from -tip: > > > > > > 098991fccfc7: nohz: Print more debug info in tick_nohz_stop_sched_tick() > > > 22aa2ad45fd8: tick: Make sure tick timer is active when bypassing reprogramming > > > d58bd60c773d: nohz: Fix again collision between tick and other hrtimers > > > > > > ... and have reintegrated tip:master, so it should be back to working I think. > > > > > > Does this solve the warning and the RCU stalls on your boxes? > > > > (both boxen building) > > (oh come on, boot already... zzz) > > Both still lose their TSC. Aw crap, and the DL980 RCU stalled while idling again. > [ 11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz > [ 11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns > [ 13.064172] clocksource: Switched to clocksource tsc > [ 240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large: > [ 240.462501] clocksource: 'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff > [ 240.675057] tsc: Marking TSC unstable due to clocksource watchdog > > [ 349.192460] clocksource: timekeeping watchdog on CPU63: Marking clocksource 'tsc' as unstable because the skew is too large: > [ 349.217756] clocksource: 'hpet' wd_now: 2bb0dad7 wd_last: 5d9ac241 mask: ffffffff > [ 349.239021] clocksource: 'tsc' cs_now: d2413ebfc1c28 cs_last: d2387b328a38c mask: ffffffffffffffff > [ 349.263525] sched_clock: Marking unstable (349228856508, 34664619)<-(351146178141, -1882657014) > [ 349.283245] tsc: Marking TSC unstable due to clocksource watchdog > [ 349.298075] clocksource: Switched to clocksource hpet
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-04-26 12:30 +0200 |
| Message-ID | <tAsmS-Y5-15@gated-at.bofh.it> |
| In reply to | #1631265 |
On Wed, Apr 26, 2017 at 10:57:42AM +0200, Mike Galbraith wrote: > Both still lose their TSC. > > [ 11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz > [ 11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns > [ 13.064172] clocksource: Switched to clocksource tsc > [ 240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large: > [ 240.462501] clocksource: 'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff > [ 240.675057] tsc: Marking TSC unstable due to clocksource watchdog And they didn't use to? We don't typically write to TSC or TSC_ADJUST and thus would not cause such behaviour.
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-26 13:50 +0200 |
| Message-ID | <tAtCh-1In-1@gated-at.bofh.it> |
| In reply to | #1631358 |
On Wed, 2017-04-26 at 12:26 +0200, Peter Zijlstra wrote: > On Wed, Apr 26, 2017 at 10:57:42AM +0200, Mike Galbraith wrote: > > > Both still lose their TSC. > > > > [ 11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz > > [ 11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns > > [ 13.064172] clocksource: Switched to clocksource tsc > > [ 240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large: > > [ 240.462501] clocksource: 'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff > > [ 240.675057] tsc: Marking TSC unstable due to clocksource watchdog > > > And they didn't use to? We don't typically write to TSC or TSC_ADJUST > and thus would not cause such behaviour. Nope. The DL980 is my RT jitter test box, which I use all the time. The other makes very spiffy jitter numbers when TSCs get synched up (crap bios).. well used to before suse maintenance elves got around to doing the BIOS replacement ;-/ (poor box, say hi to the goldfish) -Mike
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-26 14:40 +0200 |
| Message-ID | <tAuoF-2hj-13@gated-at.bofh.it> |
| In reply to | #1631412 |
On Wed, 2017-04-26 at 13:39 +0200, Mike Galbraith wrote: > On Wed, 2017-04-26 at 12:26 +0200, Peter Zijlstra wrote: > > On Wed, Apr 26, 2017 at 10:57:42AM +0200, Mike Galbraith wrote: > > > > > Both still lose their TSC. > > > > > > [ 11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz > > > [ 11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns > > > [ 13.064172] clocksource: Switched to clocksource tsc > > > [ 240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large: > > > [ 240.462501] clocksource: 'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff > > > [ 240.675057] tsc: Marking TSC unstable due to clocksource watchdog > > > > > > And they didn't use to? We don't typically write to TSC or TSC_ADJUST > > and thus would not cause such behaviour. > > Nope. DL980 seems perfectly happy with master.today.. so off we go.
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-26 20:20 +0200 |
| Message-ID | <tAzHH-5SI-3@gated-at.bofh.it> |
| In reply to | #1631437 |
On Wed, 2017-04-26 at 14:30 +0200, Mike Galbraith wrote:
> On Wed, 2017-04-26 at 13:39 +0200, Mike Galbraith wrote:
> > On Wed, 2017-04-26 at 12:26 +0200, Peter Zijlstra wrote:
> > > On Wed, Apr 26, 2017 at 10:57:42AM +0200, Mike Galbraith wrote:
> > >
> > > > Both still lose their TSC.
> > > >
> > > > [ 11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz
> > > > [ 11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns
> > > > [ 13.064172] clocksource: Switched to clocksource tsc
> > > > [ 240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large:
> > > > [ 240.462501] clocksource: 'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff
> > > > [ 240.675057] tsc: Marking TSC unstable due to clocksource watchdog
> > >
> > >
> > > And they didn't use to? We don't typically write to TSC or TSC_ADJUST
> > > and thus would not cause such behaviour.
> >
> > Nope.
>
> DL980 seems perfectly happy with master.today.. so off we go.
ec2206b91d430da57c856869b5a37dc1e569e80d is the first bad commit
commit ec2206b91d430da57c856869b5a37dc1e569e80d
Author: Anna-Maria Gleixner <anna-maria@linutronix.de>
Date: Mon Mar 20 10:34:20 2017 +0100
timer: Implement the hierarchical pull model
Placing timers at enqueue time on a target CPU based on dubious heuristics
does not make any sense:
1) Most timer wheel timers are canceled or rearmed before they expire.
2) The heuristics to predict which CPU will be busy when the timer expires
are wrong by definition.
So we waste precious cycles to place timers at enqueue time.
The proper solution to this problem is to always queue the timers on the
local CPU and allow the non pinned timers to be pulled onto a busy CPU at
expiry time.
To achieve this the timer storage has been split into local pinned and
global timers. Local pinned timers are always expired on the CPU on which
they have been queued. Global timers can be expired on any CPU.
As long as a CPU is busy it expires both local and global timers. When a
CPU goes idle it arms for the first expiring local timer. If the first
expiring pinned (local) timer is before the first expiring movable timer,
then no action is required because the CPU will wake up before the first
movable timer expires. If the first expiring movable timer is before the
first expiring pinned (local) timer, then this timer is queued into a idle
timerqueue and eventually expired by some other active CPU.
To avoid global locking the timerqueues are implemented as a hierarchy. The
lowest level of the hierarchy holds the CPUs. The CPUs are associated to
groups of 8, which are seperated per node. If more than one CPU group
exist, then a second level in the hierarchy collects the groups. Depending
on the size of the system more than 2 levels are required. Each group has a
"migrator" which checks the timerqueue during the tick for remote expirable
timers.
If the last CPU in a group goes idle it reports the first expiring event in
the group up to the next group(s) in the hierarchy. If the last CPU goes
idle it arms its timer for the first system wide expiring timer to ensure
that no timer event is missed.
Signed-off-by: Anna-Maria Gleixner <anna-maria@linutronix.de>
Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
:040000 040000 76086a473919931466dfeea76a17566e2bb15130 5432a1bbc5ace9796e70e91447c6f26eb9eb51be M include
:040000 040000 102292e74099b7c3e0b2125a4deef728267b4d9d 24385b3c68140fbb3362f32753f296d127e22097 M kernel
git bisect start
# good: [ea839b41744dffe5c77b8d9842c9bb7073460901] Merge tag 'arc-4.11-final' of git://git.kernel.org/pub/scm/linux/kernel/git/vgupta/arc
git bisect good ea839b41744dffe5c77b8d9842c9bb7073460901
# bad: [d02c59825bbfd2726430f410321f9d0bfc9d1ef4] Merge branch 'WIP.x86/fpu'
git bisect bad d02c59825bbfd2726430f410321f9d0bfc9d1ef4
# good: [1212249de8142df8c64c4165d1dee8ed4c3d766c] Merge branch 'perf/core'
git bisect good 1212249de8142df8c64c4165d1dee8ed4c3d766c
# good: [28a68c6ee0837004c87845a09717a99b4e96f7d4] Merge branch 'x86-mm-for-xen'
git bisect good 28a68c6ee0837004c87845a09717a99b4e96f7d4
# good: [7fd018291abbc076dc8073d48bba9614ca67b818] Merge branch 'x86/cleanups'
git bisect good 7fd018291abbc076dc8073d48bba9614ca67b818
# good: [6f029d98bc76cf9aa21e73e4beba9014e7e3adbb] Merge branch 'x86/platform'
git bisect good 6f029d98bc76cf9aa21e73e4beba9014e7e3adbb
# bad: [4691eff5d6475014b4c04481b03aff89f5ec87c1] Merge branch 'WIP.timers'
git bisect bad 4691eff5d6475014b4c04481b03aff89f5ec87c1
# good: [0605ab6fd5aa0c78133fc6611fa5e4f17e46c396] sched/wait: Disambiguate wq_entry->task_list and wq_head->task_list naming
git bisect good 0605ab6fd5aa0c78133fc6611fa5e4f17e46c396
# bad: [1e1c48e6befca01abb2ed224a76fbe0ea8c73e30] timer: Always queue timers on the local CPU
git bisect bad 1e1c48e6befca01abb2ed224a76fbe0ea8c73e30
# good: [c82a8b6b216e6910f36a94939de7e8793b25feaa] timer: Retrieve next expiry of pinned/non-pinned timers seperately
git bisect good c82a8b6b216e6910f36a94939de7e8793b25feaa
# good: [e2e1214438bb25ec4e37b8aedb3d3e502535b09b] tick/sched: Split out jiffies update helper function
git bisect good e2e1214438bb25ec4e37b8aedb3d3e502535b09b
# bad: [270c4e558cc26a0d89e89dda46c35ac39457a896] timer_migration: Add tracepoints
git bisect bad 270c4e558cc26a0d89e89dda46c35ac39457a896
# bad: [ec2206b91d430da57c856869b5a37dc1e569e80d] timer: Implement the hierarchical pull model
git bisect bad ec2206b91d430da57c856869b5a37dc1e569e80d
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-27 07:30 +0200 |
| Message-ID | <tAKa5-4nJ-1@gated-at.bofh.it> |
| In reply to | #1631437 |
On Wed, 2017-04-26 at 14:30 +0200, Mike Galbraith wrote:
> On Wed, 2017-04-26 at 13:39 +0200, Mike Galbraith wrote:
> > On Wed, 2017-04-26 at 12:26 +0200, Peter Zijlstra wrote:
> > > On Wed, Apr 26, 2017 at 10:57:42AM +0200, Mike Galbraith wrote:
> > >
> > > > Both still lose their TSC.
> > > >
> > > > [ 11.982468] tsc: Refined TSC clocksource calibration: 2260.999 MHz
> > > > [ 11.994275] clocksource: tsc: mask: 0xffffffffffffffff max_cycles: 0x20974a4d8bb, max_idle_ns: 440795246623 ns
> > > > [ 13.064172] clocksource: Switched to clocksource tsc
> > > > [ 240.247851] clocksource: timekeeping watchdog on CPU23: Marking clocksource 'tsc' as unstable because the skew is too large:
> > > > [ 240.462501] clocksource: 'tsc' cs_now: 108fe5be09f cs_last: b90a6a0676 mask: ffffffffffffffff
> > > > [ 240.675057] tsc: Marking TSC unstable due to clocksource watchdog
> > >
> > >
> > > And they didn't use to? We don't typically write to TSC or TSC_ADJUST
> > > and thus would not cause such behaviour.
> >
> > Nope.
>
> DL980 seems perfectly happy with master.today.. so off we go.
hm, this bit of huge trace looks less than wonderful.
<idle>-0 [041] ..s2 317.304657: timer_expire_entry: timer=ffffffff820d6600 function=clocksource_watchdog now=4294971392
<idle>-0 [041] d.s4 317.304660: timer_start: timer=ffffffff820d6600 function=clocksource_watchdog expires=4294916631 [timeout=-54761] cpu=42 idx=19 flags=
<idle>-0 [041] ..s2 317.304660: timer_expire_exit: timer=ffffffff820d6600 ^^^^^^^^^^^^^^
1.1 megalines later, we finally meet function=clocksource_watchdog
again, and have a cow.
<idle>-0 [043] d.s3 489.511620: timer_cancel: timer=ffffffff820d6600
<idle>-0 [043] ..s2 489.511621: timer_expire_entry: timer=ffffffff820d6600 function=clocksource_watchdog now=4295014443
<idle>-0 [043] ..s2 489.511628: clocksource_watchdog: timekeeping watchdog on CPU43: Marking clocksource 'tsc' as unstable because the skew is too large:
<idle>-0 [043] ..s2 489.511630: clocksource_watchdog: 'hpet' wd_now: a1cbfa1a wd_last: ed7acfe mask: ffffffff
<idle>-0 [043] ..s2 489.511630: clocksource_watchdog: 'tsc' cs_now: 1c24ee60f22 cs_last: 167a92e836f mask: ffffffffffffffff
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-29 18:10 +0200 |
| Subject | [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() |
| Message-ID | <tBD6y-8qU-15@gated-at.bofh.it> |
| In reply to | #1631437 |
Note: there is more. With this applied, my desktop box will no longer
reproduce when booted to init 3 with nowatchdog on the command line.
My 8 socket DL980 OTOH still will, though it takes longer, and is
seemingly no longer interested in following up with a permanent RCU
stall after the tsc clocksource is killed, as it does in virgin source.
---
timers_update_migration() is called by tick_nohz_activate() before
the late initcall tmigr_init() sets tmigr_enabled to true, resulting
in it updating neither timer_base.nohz_active nor .migration_enabled,
meaning we'll not kick an idling cpu in add_timer_on().
Remove redundant loop avoidance such that tick_nohz_activate() updates
timer_bases[].nohz_active as intended, and call it in tmigr_init() to
update timer_bases[].migration_enabled.
Signed-off-by: Mike Galbraith <efault@gmx.de>
Fixes: ec2206b91d43 timer: Implement the hierarchical pull model
---
kernel/time/timer.c | 4 ----
kernel/time/timer_migration.c | 1 +
2 files changed, 1 insertion(+), 4 deletions(-)
--- a/kernel/time/timer.c
+++ b/kernel/time/timer.c
@@ -224,10 +224,6 @@ void timers_update_migration(bool update
bool on = sysctl_timer_migration && tick_nohz_active && tmigr_enabled;
unsigned int cpu;
- /* Avoid the loop, if nothing to update */
- if (this_cpu_read(timer_bases[BASE_GLOBAL].migration_enabled) == on)
- return;
-
for_each_possible_cpu(cpu) {
per_cpu(timer_bases[BASE_LOCAL].migration_enabled, cpu) = on;
per_cpu(timer_bases[BASE_GLOBAL].migration_enabled, cpu) = on;
--- a/kernel/time/timer_migration.c
+++ b/kernel/time/timer_migration.c
@@ -649,6 +649,7 @@ static int __init tmigr_init(void)
goto hp_err;
tmigr_enabled = true;
+ timers_update_migration(false);
pr_info("Timer migration: %d hierarchy levels\n", tmigr_hierarchy_levels);
return 0;
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-04-29 20:10 +0200 |
| Subject | Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() |
| Message-ID | <tBEYG-15g-13@gated-at.bofh.it> |
| In reply to | #1633307 |
On Sat, Apr 29, 2017 at 06:06:40PM +0200, Mike Galbraith wrote:
> Note: there is more. With this applied, my desktop box will no longer
> reproduce when booted to init 3 with nowatchdog on the command line.
> My 8 socket DL980 OTOH still will, though it takes longer, and is
> seemingly no longer interested in following up with a permanent RCU
> stall after the tsc clocksource is killed, as it does in virgin source.
>
> ---
>
> timers_update_migration() is called by tick_nohz_activate() before
> the late initcall tmigr_init() sets tmigr_enabled to true, resulting
> in it updating neither timer_base.nohz_active nor .migration_enabled,
> meaning we'll not kick an idling cpu in add_timer_on().
>
> Remove redundant loop avoidance such that tick_nohz_activate() updates
> timer_bases[].nohz_active as intended, and call it in tmigr_init() to
> update timer_bases[].migration_enabled.
>
> Signed-off-by: Mike Galbraith <efault@gmx.de>
> Fixes: ec2206b91d43 timer: Implement the hierarchical pull model
If someone will either repost a fresh series or point me at exactly
the set of patches to use, I will run it through rcutorture again.
Thanx, Paul
> ---
> kernel/time/timer.c | 4 ----
> kernel/time/timer_migration.c | 1 +
> 2 files changed, 1 insertion(+), 4 deletions(-)
>
> --- a/kernel/time/timer.c
> +++ b/kernel/time/timer.c
> @@ -224,10 +224,6 @@ void timers_update_migration(bool update
> bool on = sysctl_timer_migration && tick_nohz_active && tmigr_enabled;
> unsigned int cpu;
>
> - /* Avoid the loop, if nothing to update */
> - if (this_cpu_read(timer_bases[BASE_GLOBAL].migration_enabled) == on)
> - return;
> -
> for_each_possible_cpu(cpu) {
> per_cpu(timer_bases[BASE_LOCAL].migration_enabled, cpu) = on;
> per_cpu(timer_bases[BASE_GLOBAL].migration_enabled, cpu) = on;
> --- a/kernel/time/timer_migration.c
> +++ b/kernel/time/timer_migration.c
> @@ -649,6 +649,7 @@ static int __init tmigr_init(void)
> goto hp_err;
>
> tmigr_enabled = true;
> + timers_update_migration(false);
> pr_info("Timer migration: %d hierarchy levels\n", tmigr_hierarchy_levels);
> return 0;
>
>
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-29 20:30 +0200 |
| Subject | Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() |
| Message-ID | <tBFi2-1bq-3@gated-at.bofh.it> |
| In reply to | #1633317 |
On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote: > If someone will either repost a fresh series or point me at exactly > the set of patches to use, I will run it through rcutorture again. Patchlet is against x86-tip/master.today. -Mike
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-04-29 23:50 +0200 |
| Subject | Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() |
| Message-ID | <tBIpz-30B-9@gated-at.bofh.it> |
| In reply to | #1633320 |
On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote:
> On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote:
>
> > If someone will either repost a fresh series or point me at exactly
> > the set of patches to use, I will run it through rcutorture again.
>
> Patchlet is against x86-tip/master.today.
So today's (as in Saturday April 29) x86-tip/master with the following
patch applied?
Thanx, Paul
------------------------------------------------------------------------
timers_update_migration() is called by tick_nohz_activate() before
the late initcall tmigr_init() sets tmigr_enabled to true, resulting
in it updating neither timer_base.nohz_active nor .migration_enabled,
meaning we'll not kick an idling cpu in add_timer_on().
Remove redundant loop avoidance such that tick_nohz_activate() updates
timer_bases[].nohz_active as intended, and call it in tmigr_init() to
update timer_bases[].migration_enabled.
Signed-off-by: Mike Galbraith <efault@gmx.de>
Fixes: ec2206b91d43 timer: Implement the hierarchical pull model
---
kernel/time/timer.c | 4 ----
kernel/time/timer_migration.c | 1 +
2 files changed, 1 insertion(+), 4 deletions(-)
--- a/kernel/time/timer.c
+++ b/kernel/time/timer.c
@@ -224,10 +224,6 @@ void timers_update_migration(bool update
bool on = sysctl_timer_migration && tick_nohz_active && tmigr_enabled;
unsigned int cpu;
- /* Avoid the loop, if nothing to update */
- if (this_cpu_read(timer_bases[BASE_GLOBAL].migration_enabled) == on)
- return;
-
for_each_possible_cpu(cpu) {
per_cpu(timer_bases[BASE_LOCAL].migration_enabled, cpu) = on;
per_cpu(timer_bases[BASE_GLOBAL].migration_enabled, cpu) = on;
--- a/kernel/time/timer_migration.c
+++ b/kernel/time/timer_migration.c
@@ -649,6 +649,7 @@ static int __init tmigr_init(void)
goto hp_err;
tmigr_enabled = true;
+ timers_update_migration(false);
pr_info("Timer migration: %d hierarchy levels\n", tmigr_hierarchy_levels);
return 0;
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-30 03:30 +0200 |
| Subject | Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() |
| Message-ID | <tBLQu-5bK-5@gated-at.bofh.it> |
| In reply to | #1633351 |
On Sat, 2017-04-29 at 14:45 -0700, Paul E. McKenney wrote: > On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote: > > On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote: > > > > > If someone will either repost a fresh series or point me at exactly > > > the set of patches to use, I will run it through rcutorture again. > > > > Patchlet is against x86-tip/master.today. > > So today's (as in Saturday April 29) x86-tip/master with the following > patch applied? Yeah.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-04-30 05:50 +0200 |
| Subject | Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() |
| Message-ID | <tBO1X-6vb-1@gated-at.bofh.it> |
| In reply to | #1633368 |
On Sun, Apr 30, 2017 at 03:21:58AM +0200, Mike Galbraith wrote: > On Sat, 2017-04-29 at 14:45 -0700, Paul E. McKenney wrote: > > On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote: > > > On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote: > > > > > > > If someone will either repost a fresh series or point me at exactly > > > > the set of patches to use, I will run it through rcutorture again. > > > > > > Patchlet is against x86-tip/master.today. > > > > So today's (as in Saturday April 29) x86-tip/master with the following > > patch applied? > > Yeah. OK, will fire it up once the current set of overnight tests complete. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-04-30 06:30 +0200 |
| Subject | Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() |
| Message-ID | <tBOEF-72b-1@gated-at.bofh.it> |
| In reply to | #1633374 |
On Sat, 2017-04-29 at 20:43 -0700, Paul E. McKenney wrote: > On Sun, Apr 30, 2017 at 03:21:58AM +0200, Mike Galbraith wrote: > > On Sat, 2017-04-29 at 14:45 -0700, Paul E. McKenney wrote: > > > On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote: > > > > On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote: > > > > > > > > > If someone will either repost a fresh series or point me at exactly > > > > > the set of patches to use, I will run it through rcutorture again. > > > > > > > > Patchlet is against x86-tip/master.today. > > > > > > So today's (as in Saturday April 29) x86-tip/master with the following > > > patch applied? > > > > Yeah. > > OK, will fire it up once the current set of overnight tests complete. I certainly don't want to discourage you from beating hell outta tip, just want to make sure you know that I'm seeing zero RCU woes, only late timer expiry (sharpening rocks/sticks to focus trace). -Mike
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-04-30 06:40 +0200 |
| Subject | Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() |
| Message-ID | <tBOOl-765-3@gated-at.bofh.it> |
| In reply to | #1633375 |
On Sun, Apr 30, 2017 at 06:20:15AM +0200, Mike Galbraith wrote: > On Sat, 2017-04-29 at 20:43 -0700, Paul E. McKenney wrote: > > On Sun, Apr 30, 2017 at 03:21:58AM +0200, Mike Galbraith wrote: > > > On Sat, 2017-04-29 at 14:45 -0700, Paul E. McKenney wrote: > > > > On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote: > > > > > On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote: > > > > > > > > > > > If someone will either repost a fresh series or point me at exactly > > > > > > the set of patches to use, I will run it through rcutorture again. > > > > > > > > > > Patchlet is against x86-tip/master.today. > > > > > > > > So today's (as in Saturday April 29) x86-tip/master with the following > > > > patch applied? > > > > > > Yeah. > > > > OK, will fire it up once the current set of overnight tests complete. > > I certainly don't want to discourage you from beating hell outta tip, > just want to make sure you know that I'm seeing zero RCU woes, only > late timer expiry (sharpening rocks/sticks to focus trace). I got timer_migration splats from an earlier rcutorture run. Please see message-ID <20170421192853.GD3956@linux.vnet.ibm.com> on LKML on April 21st in reply to Thomas's V2 00/10 cover letter. So I am curious to learn if your patches fix them. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-05-01 00:50 +0200 |
| Subject | Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() |
| Message-ID | <tC5Pc-F1-7@gated-at.bofh.it> |
| In reply to | #1633376 |
On Sat, Apr 29, 2017 at 09:36:37PM -0700, Paul E. McKenney wrote: > On Sun, Apr 30, 2017 at 06:20:15AM +0200, Mike Galbraith wrote: > > On Sat, 2017-04-29 at 20:43 -0700, Paul E. McKenney wrote: > > > On Sun, Apr 30, 2017 at 03:21:58AM +0200, Mike Galbraith wrote: > > > > On Sat, 2017-04-29 at 14:45 -0700, Paul E. McKenney wrote: > > > > > On Sat, Apr 29, 2017 at 08:20:33PM +0200, Mike Galbraith wrote: > > > > > > On Sat, 2017-04-29 at 11:06 -0700, Paul E. McKenney wrote: > > > > > > > > > > > > > If someone will either repost a fresh series or point me at exactly > > > > > > > the set of patches to use, I will run it through rcutorture again. > > > > > > > > > > > > Patchlet is against x86-tip/master.today. > > > > > > > > > > So today's (as in Saturday April 29) x86-tip/master with the following > > > > > patch applied? > > > > > > > > Yeah. > > > > > > OK, will fire it up once the current set of overnight tests complete. > > > > I certainly don't want to discourage you from beating hell outta tip, > > just want to make sure you know that I'm seeing zero RCU woes, only > > late timer expiry (sharpening rocks/sticks to focus trace). > > I got timer_migration splats from an earlier rcutorture run. Please see > message-ID <20170421192853.GD3956@linux.vnet.ibm.com> on LKML on April > 21st in reply to Thomas's V2 00/10 cover letter. So I am curious to > learn if your patches fix them. And sadly, the splats are still there. Please see the following for the relevant console output and .config files: http://www2.rdrop.com/users/paulmck/submission/TREE04.2017.04.30a.config http://www2.rdrop.com/users/paulmck/submission/TREE04.2017.04.30a.console.log http://www2.rdrop.com/users/paulmck/submission/TREE04.3.2017.04.30a.console.log http://www2.rdrop.com/users/paulmck/submission/TREE07.2017.04.30a.config http://www2.rdrop.com/users/paulmck/submission/TREE07.2.2017.04.30a.bzImage http://www2.rdrop.com/users/paulmck/submission/TREE07.2017.04.30a.console.log http://www2.rdrop.com/users/paulmck/submission/TREE07.2.2017.04.30a.console.log Please let me know if you have any trouble accessing these. Here is the first splat from the first TREE04 run: [ 3.310642] WARNING: CPU: 1 PID: 0 at /home/paulmck/public_git/timer-tip/kernel/time/timer_migration.c:387 tmigr_set_cpu_active+0xc6/0xe0 [ 3.313210] Modules linked in: [ 3.313861] CPU: 1 PID: 0 Comm: swapper/1 Not tainted 4.11.0-rc8+ #1 [ 3.315196] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS Ubuntu-1.8.2-1ubuntu1 04/01/2014 [ 3.317196] task: ffff8c1fde9f0000 task.stack: ffff8e4fc0124000 [ 3.318433] RIP: 0010:tmigr_set_cpu_active+0xc6/0xe0 [ 3.319464] RSP: 0000:ffff8e4fc0127e90 EFLAGS: 00010046 [ 3.320598] RAX: 0000000000000004 RBX: 0000000000000001 RCX: 000000000000001f [ 3.322146] RDX: 0000000000000001 RSI: ffff8c1fdfc54cc8 RDI: ffff8c1fdeb26f80 [ 3.323652] RBP: ffff8e4fc0127ea8 R08: 0000000000000000 R09: 0000000000000008 [ 3.325237] R10: ffff8e4fc0127e80 R11: 0000000000000400 R12: ffff8c1fdeb26f80 [ 3.326699] R13: ffff8c1fdfc54cc8 R14: 0000000000000000 R15: 0000000000000000 [ 3.328149] FS: 0000000000000000(0000) GS:ffff8c1fdfc40000(0000) knlGS:0000000000000000 [ 3.329845] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [ 3.331078] CR2: ffff8e4fc02f0000 CR3: 0000000015e0a000 CR4: 00000000000006e0 [ 3.332509] Call Trace: [ 3.333107] tmigr_cpu_activate+0x36/0x40 [ 3.333972] tick_nohz_idle_exit+0xd1/0xf0 [ 3.334845] do_idle+0x113/0x170 [ 3.335501] cpu_startup_entry+0x18/0x20 [ 3.336338] start_secondary+0xe8/0xf0 [ 3.337147] secondary_startup_64+0x9f/0x9f [ 3.337998] Code: d0 48 8b 03 48 85 c0 75 eb eb a0 49 8b 7c 24 50 41 89 5c 24 08 48 85 ff 74 8c 49 8d 74 24 20 89 da e8 3f ff ff ff e9 7b ff ff ff <0f> ff 41 c6 04 24 00 5b 41 5c 41 5d 5d c3 66 90 66 2e 0f 1f 8 This is the first WARN_ON() in tmigr_set_cpu_active(). I got four splats in 12 hours of running the rcutorture TREE04 test scenario, that is, three runs of four hours each. The TREE07 runs fared worse, with many more splats, starting with a page fault. The scripting claimed a hang, but that looks to have instead been so many splats that the test failed to terminate itself in time. I ran two TREE07 runs of four hours each. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-05-01 10:00 +0200 |
| Subject | Re: [patch] timer: Fix timers_update_migration(), and call it in tmigr_init() |
| Message-ID | <tCeps-6id-3@gated-at.bofh.it> |
| In reply to | #1633503 |
On Sun, 30 Apr 2017, Paul E. McKenney wrote: > > And sadly, the splats are still there. Please see the following for > the relevant console output and .config files: Right. We are working on fixing the issues. I'll zap it from tip for now. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web