Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1176295 > unrolled thread
| Started by | Simon Horman <horms@verge.net.au> |
|---|---|
| First post | 2015-07-03 04:50 +0200 |
| Last post | 2015-07-03 19:10 +0200 |
| Articles | 18 — 6 participants |
Back to article view | Back to linux.kernel
Possible regression due to "tick: broadcast: Prevent livelock from event handler" Simon Horman <horms@verge.net.au> - 2015-07-03 04:50 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Thomas Gleixner <tglx@linutronix.de> - 2015-07-03 11:30 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Wolfram Sang <wsa@the-dreams.de> - 2015-07-03 13:00 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Russell King - ARM Linux <linux@arm.linux.org.uk> - 2015-07-03 13:10 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Thomas Gleixner <tglx@linutronix.de> - 2015-07-03 15:30 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Wolfram Sang <wsa@the-dreams.de> - 2015-07-03 15:40 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Thomas Gleixner <tglx@linutronix.de> - 2015-07-03 16:20 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Wolfram Sang <wsa@the-dreams.de> - 2015-07-03 16:40 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Sudeep Holla <sudeep.holla@arm.com> - 2015-07-03 16:50 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Sudeep Holla <sudeep.holla@arm.com> - 2015-07-03 17:00 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Thomas Gleixner <tglx@linutronix.de> - 2015-07-03 17:00 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Thomas Gleixner <tglx@linutronix.de> - 2015-07-03 17:00 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Wolfram Sang <wsa@the-dreams.de> - 2015-07-03 17:20 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Thomas Gleixner <tglx@linutronix.de> - 2015-07-03 17:30 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Wolfram Sang <wsa@the-dreams.de> - 2015-07-03 17:50 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Thomas Gleixner <tglx@linutronix.de> - 2015-07-03 18:00 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Wolfram Sang <wsa@the-dreams.de> - 2015-07-03 18:10 +0200
Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" Geert Uytterhoeven <geert@linux-m68k.org> - 2015-07-03 19:10 +0200
| From | Simon Horman <horms@verge.net.au> |
|---|---|
| Date | 2015-07-03 04:50 +0200 |
| Subject | Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pHZd7-2Kt-1@gated-at.bofh.it> |
Hi Thomas,
I have observed what appears to be a regression while testing next-20150702
which seems to be caused by 2951d5c031a3 ("tick: broadcast: Prevent
livelock from event handler").
The problem manifests on the emev2/kzm9d board as per the boot log below.
The problem manifests when booting using the shmobile_defconfig,
which uses multiplatform and enables all devices using DT.
The problem does not appear to always manifest but anecdotally it
seems to manifest more often of late (yes, I know that is vague).
This problem was reported to me by Geert Uytterhoeven.
Kevin Hillman has also reported problems reliably booting the emev2/kzm9d board.
Please note that in order to boot 2951d5c031a3 on the emev2/kzm9d board
using the shmobile_defconfig the following is required:
6b442bc81337 ("nohz: Fix !HIGH_RES_TIMERS hang").
Booting Linux on physical CPU 0x0
Linux version 4.1.0-next-20150702 (horms@ayumi.isobedori.kobe.vergenet.net) (gcc version 4.6.3 (GCC) ) #4556 SMP Fri Jul 3 11:31:38 JST 2015
CPU: ARMv7 Processor [411fc093] revision 3 (ARMv7), cr=10c5307d
CPU: PIPT / VIPT nonaliasing data cache, VIPT aliasing instruction cache
Machine model: EMEV2 KZM9D Board
debug: ignoring loglevel setting.
Memory policy: Data cache writealloc
On node 0 totalpages: 32768
free_area_init_node: node 0, pgdat c0817500, node_mem_map c7ef9000
Normal zone: 256 pages used for memmap
Normal zone: 0 pages reserved
Normal zone: 32768 pages, LIFO batch:7
PERCPU: Embedded 9 pages/cpu @c7ee0000 s13824 r0 d23040 u36864
pcpu-alloc: s13824 r0 d23040 u36864 alloc=9*4096
pcpu-alloc: [0] 0 [0] 1
Built 1 zonelists in Zone order, mobility grouping on. Total pages: 32512
Kernel command line: console=ttyS1,115200n81 ignore_loglevel root=/dev/nfs ip=dhcp
PID hash table entries: 512 (order: -1, 2048 bytes)
Dentry cache hash table entries: 16384 (order: 4, 65536 bytes)
Inode-cache hash table entries: 8192 (order: 3, 32768 bytes)
Memory: 121328K/131072K available (4888K kernel code, 283K rwdata, 1352K rodata, 1728K init, 204K bss, 9744K reserved, 0K cma-reserved, 0K highmem)
Virtual kernel memory layout:
vector : 0xffff0000 - 0xffff1000 ( 4 kB)
fixmap : 0xffc00000 - 0xfff00000 (3072 kB)
vmalloc : 0xc8800000 - 0xff000000 ( 872 MB)
lowmem : 0xc0000000 - 0xc8000000 ( 128 MB)
pkmap : 0xbfe00000 - 0xc0000000 ( 2 MB)
.text : 0xc0008000 - 0xc0621044 (6245 kB)
.init : 0xc0622000 - 0xc07d2000 (1728 kB)
.data : 0xc07d2000 - 0xc0818e40 ( 284 kB)
.bss : 0xc081b000 - 0xc084e24c ( 205 kB)
Hierarchical RCU implementation.
Additional per-CPU info printed with stalls.
Build-time adjustment of leaf fanout to 32.
RCU restricting CPUs from NR_CPUS=8 to nr_cpu_ids=2.
RCU: Adjusting geometry for rcu_fanout_leaf=32, nr_cpu_ids=2
NR_IRQS:16 nr_irqs:16 16
clocksource_of_init: no matching clocksources found
sched_clock: 32 bits at 100 Hz, resolution 10000000ns, wraps every 21474836475000000ns
Console: colour dummy device 80x30
Calibrating delay loop (skipped) preset value.. 355.33 BogoMIPS (lpj=1776666)
pid_max: default: 32768 minimum: 301
Mount-cache hash table entries: 1024 (order: 0, 4096 bytes)
Mountpoint-cache hash table entries: 1024 (order: 0, 4096 bytes)
CPU: Testing write buffer coherency: ok
CPU0: thread -1, cpu 0, socket 0, mpidr 80000000
Setting up static identity map for 0x40009000 - 0x40009058
CPU1: thread -1, cpu 1, socket 0, mpidr 80000001
Brought up 2 CPUs
SMP: Total of 2 processors activated (710.66 BogoMIPS).
CPU: All CPU(s) started in SVC mode.
devtmpfs: initialized
VFP support v0.3: implementor 41 architecture 3 part 30 variant 9 rev 1
clocksource: jiffies: mask: 0xffffffff max_cycles: 0xffffffff, max_idle_ns: 19112604462750000 ns
pinctrl core: initialized pinctrl subsystem
NET: Registered protocol family 16
DMA: preallocated 256 KiB pool for atomic coherent allocations
sh-pfc e0140200.pfc: emev2_pfc support registered
No ATAGs?
hw-breakpoint: found 5 (+1 reserved) breakpoint and 1 watchpoint registers.
hw-breakpoint: maximum watchpoint size is 4 bytes.
vgaarb: loaded
SCSI subsystem initialized
libata version 3.00 loaded.
usbcore: registered new interface driver usbfs
usbcore: registered new interface driver hub
usbcore: registered new device driver usb
media: Linux media interface: v0.10
Linux video capture interface: v2.00
em_sti e0180000.timer: used for clock events
em_sti e0180000.timer: used for oneshot clock events
em_sti e0180000.timer: used as clock source
clocksource: e0180000.timer: mask: 0xffffffffffff max_cycles: 0x1ef4687b1, max_idle_ns: 3697658158765000000 ns
Advanced Linux Sound Architecture Driver Initialized.
clocksource: e0180000.timer: mask: 0xffffffffffff max_cycles: 0x1ef4687b1, max_idle_ns: 112843571739654 ns
clocksource: Switched to clocksource e0180000.timer
NET: Registered protocol family 2
TCP established hash table entries: 1024 (order: 0, 4096 bytes)
TCP bind hash table entries: 1024 (order: 1, 8192 bytes)
TCP: Hash tables configured (established 1024 bind 1024)
UDP hash table entries: 256 (order: 1, 8192 bytes)
UDP-Lite hash table entries: 256 (order: 1, 8192 bytes)
NET: Registered protocol family 1
RPC: Registered named UNIX socket transport module.
RPC: Registered udp transport module.
RPC: Registered tcp transport module.
RPC: Registered tcp NFSv4.1 backchannel transport module.
PCI: CLS 0 bytes, default 64
Clockevents: could not switch to one-shot mode: dummy_timer is not functional.
Clockevents: could not switch to one-shot mode: dummy_timer is not functional.
hw perfevents: Failed to parse /pmu/interrupt-affinity[0]
hw perfevents: enabled with armv7_cortex_a9 PMU driver, 7 counters available
futex hash table entries: 512 (order: 3, 32768 bytes)
NFS: Registering the id_resolver key type
Key type id_resolver registered
Key type id_legacy registered
nfs4filelayout_init: NFSv4 File Layout Driver Registering...
nfs4flexfilelayout_init: NFSv4 Flexfile Layout Driver Registering...
jitterentropy: Initialization failed with host not compliant with requirements: 2
Block layer SCSI generic (bsg) driver version 0.4 loaded (major 250)
io scheduler noop registered
io scheduler deadline registered
io scheduler cfq registered (default)
Serial: 8250/16550 driver, 4 ports, IRQ sharing disabled
e1020000.serial: ttyS0 at MMIO 0xe1020000 (irq = 19, base_baud = 796444) is a 16550A
console [ttyS1] disabled
e1030000.serial: ttyS1 at MMIO 0xe1030000 (irq = 20, base_baud = 7168000) is a 16550A
console [ttyS1] enabled
e1040000.serial: ttyS2 at MMIO 0xe1040000 (irq = 21, base_baud = 14336000) is a 16550A
e1050000.serial: ttyS3 at MMIO 0xe1050000 (irq = 22, base_baud = 2389333) is a 16550A
SuperH (H)SCI(F) driver initialized
[drm] Initialized drm 1.1.0 20060810
libphy: smsc911x-mdio: probed
smsc911x 20000000.ethernet eth0: attached PHY driver [SMSC LAN8700] (mii_bus:phy_addr=20000000.etherne:01, irq=-1)
smsc911x 20000000.ethernet eth0: MAC Address: 00:01:9b:04:03:cf
ehci_hcd: USB 2.0 'Enhanced' Host Controller (EHCI) Driver
ehci-pci: EHCI PCI platform driver
ohci_hcd: USB 1.1 'Open' Host Controller (OHCI) Driver
ohci-pci: OHCI PCI platform driver
mousedev: PS/2 mouse device common for all mice
i2c /dev entries driver
usbcore: registered new interface driver usbhid
usbhid: USB HID core driver
NET: Registered protocol family 10
sit: IPv6 over IPv4 tunneling driver
NET: Registered protocol family 17
Key type dns_resolver registered
cpu cpu0: failed to get cpu0 clock: -2
cpufreq-dt: probe of cpufreq-dt failed with error -2
Registering SWP/SWPB emulation handler
input: gpio_keys as /devices/platform/gpio_keys/input/input0
hctosys: unable to open rtc device (rtc0)
The boot hangs here.
The next line should be:
smsc911x 20000000.ethernet eth0: SMSC911x/921x identified at 0xc8880000, IRQ: 33
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2015-07-03 11:30 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pI5sd-6JS-1@gated-at.bofh.it> |
| In reply to | #1176295 |
On Fri, 3 Jul 2015, Geert Uytterhoeven wrote:
> Hi Simon,
>
> On Fri, Jul 3, 2015 at 4:40 AM, Simon Horman <horms@verge.net.au> wrote:
> > I have observed what appears to be a regression while testing next-20150702
> > which seems to be caused by 2951d5c031a3 ("tick: broadcast: Prevent
> > livelock from event handler").
> >
> > The problem manifests on the emev2/kzm9d board as per the boot log below.
> >
> > The problem manifests when booting using the shmobile_defconfig,
> > which uses multiplatform and enables all devices using DT.
> >
> > The problem does not appear to always manifest but anecdotally it
> > seems to manifest more often of late (yes, I know that is vague).
>
> > hctosys: unable to open rtc device (rtc0)
> >
> > The boot hangs here.
> > The next line should be:
> >
> > smsc911x 20000000.ethernet eth0: SMSC911x/921x identified at 0xc8880000, IRQ: 33
>
> As you can reproduce it, can you please try enabling lockdep debugging?
Just looking at the em_sti driver. It calls clk_prepare/unprepare from
interrupt disabled regions ...
But that's not the problem at hand I think. The above commit is moving
the call to the event handler on the local cpu out of the broadcast
lock region to prevent a live lock. The only real change is the
timing.
Before:
bc_handler()
lock(bc_lock);
call_local_handler();
send_ipis();
reprogramm_bc_device();
unlock(bc_lock);
After:
bc_handler()
lock(bc_lock);
send_ipis();
reprogramm_bc_device();
unlock(bc_lock);
call_local_handler();
As this runs in hard interrupt context with interrupts disabled, I
really cannot figure out how that makes a difference.
Can you add some debugging to figure out whether the broadcast timer
interrupt still fires?
Thanks,
tglx
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Wolfram Sang <wsa@the-dreams.de> |
|---|---|
| Date | 2015-07-03 13:00 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pI6Rk-7tK-5@gated-at.bofh.it> |
| In reply to | #1176295 |
[Multipart message — attachments visible in raw view] — view raw
Hi Simon,
> I have observed what appears to be a regression while testing next-20150702
> which seems to be caused by 2951d5c031a3 ("tick: broadcast: Prevent
> livelock from event handler").
Does reverting help? I see a similar case with your branch
renesas-devel-20150629-v4.1 which does not include both commits.
For this branch, if I select CONFIG_NO_HZ_IDLE and *no* CONFIG_HIGH_RES_TIMERS,
I see the case you described:
> input: gpio_keys as /devices/platform/gpio_keys/input/input0
> hctosys: unable to open rtc device (rtc0)
>
> The boot hangs here.
> The next line should be:
>
> smsc911x 20000000.ethernet eth0: SMSC911x/921x identified at 0xc8880000, IRQ: 33
There is no OOPS, it just stops here. With CONFIG_HIGH_RES_TIMERS, the
system boots to the prompt.
Enabling all lockdep features doesn't show anything for the non-working
case. However, I get warning for the working case:
[ 2.560000] input: gpio_keys as /devices/platform/gpio_keys/input/input0
[ 2.610000] smsc911x 20000000.ethernet eth0: SMSC911x/921x identified at 0xc8880000, IRQ: 33
[ 5.510000] Sending DHCP requests ., OK
[ 5.580000] IP-Config: Got DHCP answer from 192.168.64.1, my address is 192.168.64.14
[ 5.580000] IP-Config: Complete:
[ 5.590000] device=eth0, hwaddr=00:01:9b:04:03:ce, ipaddr=192.168.64.14, mask=255.255.255.0, gw=192.168.64.1
[ 5.600000] host=192.168.64.14, domain=local, nis-domain=(none)
[ 5.600000] bootserver=192.168.64.1, rootserver=192.168.64.1, rootpath=
[ 5.610000] nameserver0=192.168.64.1
[ 5.630000] Freeing unused kernel memory: 632K (c0306000 - c03a4000)
[ 5.660000] random: init urandom read with 14 bits of entropy available
[ 5.820000] ------------[ cut here ]------------
[ 5.820000] WARNING: CPU: 0 PID: 282 at kernel/locking/lockdep.c:3557 check_flags+0x84/0x1f4()
[ 5.820000] DEBUG_LOCKS_WARN_ON(current->hardirqs_enabled)
[ 5.820000] CPU: 0 PID: 282 Comm: rcS Tainted: G W 4.1.0-00002-g5b076054611833 #179
[ 5.820000] Hardware name: Generic Emma Mobile EV2 (Flattened Device Tree)
[ 5.820000] Backtrace:
[ 5.820000] [<c0012c94>] (dump_backtrace) from [<c0012e3c>] (show_stack+0x18/0x1c)
[ 5.820000] r6:c02dcc67 r5:00000009 r4:00000000 r3:00400000
[ 5.820000] [<c0012e24>] (show_stack) from [<c02510c8>] (dump_stack+0x20/0x28)
[ 5.820000] [<c02510a8>] (dump_stack) from [<c0022c44>] (warn_slowpath_common+0x8c/0xb4)
[ 5.820000] [<c0022bb8>] (warn_slowpath_common) from [<c0022cd8>] (warn_slowpath_fmt+0x38/0x40)
[ 5.820000] r8:c780f470 r7:00000000 r6:00000000 r5:c03b0570 r4:c0b7ec04
[ 5.820000] [<c0022ca4>] (warn_slowpath_fmt) from [<c004cd38>] (check_flags+0x84/0x1f4)
[ 5.820000] r3:c02e13d8 r2:c02dceaa
[ 5.820000] [<c004ccb4>] (check_flags) from [<c0050e50>] (lock_acquire+0x4c/0xbc)
[ 5.820000] r5:00000000 r4:60000193
[ 5.820000] [<c0050e04>] (lock_acquire) from [<c0256000>] (_raw_spin_lock+0x34/0x44)
[ 5.820000] r9:000a8d5c r8:00000001 r7:c7806000 r6:c780f460 r5:c03b06a0 r4:c780f460
[ 5.820000] [<c0255fcc>] (_raw_spin_lock) from [<c005a8cc>] (handle_fasteoi_irq+0x20/0x11c)
[ 5.820000] r4:c780f400
[ 5.820000] [<c005a8ac>] (handle_fasteoi_irq) from [<c0057a4c>] (generic_handle_irq+0x28/0x38)
[ 5.820000] r6:00000000 r5:c03b038c r4:00000012 r3:c005a8ac
[ 5.820000] [<c0057a24>] (generic_handle_irq) from [<c0057ae4>] (__handle_domain_irq+0x88/0xa8)
[ 5.820000] r4:00000000 r3:00000026
[ 5.820000] [<c0057a5c>] (__handle_domain_irq) from [<c000a3cc>] (gic_handle_irq+0x40/0x58)
[ 5.820000] r8:10c5347d r7:10c5347d r6:c35b1fb0 r5:c03a6304 r4:c8802000 r3:c35b1fb0
[ 5.820000] [<c000a38c>] (gic_handle_irq) from [<c0013bc8>] (__irq_usr+0x48/0x60)
[ 5.820000] Exception stack(0xc35b1fb0 to 0xc35b1ff8)
[ 5.820000] 1fa0: 00000061 00000000 000ab736 00000066
[ 5.820000] 1fc0: 00000061 000aa1f0 000a8d54 000a8d54 000a8d88 000a8d5c 000a8cc8 000a8d68
[ 5.820000] 1fe0: 72727272 bef8a528 000398c0 00031334 20000010 ffffffff
[ 5.820000] r6:ffffffff r5:20000010 r4:00031334 r3:00000061
[ 5.820000] ---[ end trace cb88537fdc8fa202 ]---
[ 5.820000] possible reason: unannotated irqs-off.
[ 5.820000] irq event stamp: 769
[ 5.820000] hardirqs last enabled at (769): [<c000f82c>] ret_fast_syscall+0x2c/0x54
[ 5.820000] hardirqs last disabled at (768): [<c000f80c>] ret_fast_syscall+0xc/0x54
[ 5.820000] softirqs last enabled at (0): [<c0020ec4>] copy_process.part.65+0x2e8/0x11dc
[ 5.820000] softirqs last disabled at (0): [< (null)>] (null)
I haven't further researched, but wanted to share this already.
All the best,
Wolfram#
[toc] | [prev] | [next] | [standalone]
| From | Russell King - ARM Linux <linux@arm.linux.org.uk> |
|---|---|
| Date | 2015-07-03 13:10 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pI710-7Mn-11@gated-at.bofh.it> |
| In reply to | #1176475 |
On Fri, Jul 03, 2015 at 12:54:41PM +0200, Wolfram Sang wrote: > Enabling all lockdep features doesn't show anything for the non-working > case. However, I get warning for the working case: Please share your .config file for this case. Thanks. -- FTTC broadband for 0.8mile line: currently at 10.5Mbps down 400kbps up according to speedtest.net. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2015-07-03 15:30 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pI9cu-AJ-7@gated-at.bofh.it> |
| In reply to | #1176475 |
On Fri, 3 Jul 2015, Wolfram Sang wrote:
> Hi Simon,
>
> > I have observed what appears to be a regression while testing next-20150702
> > which seems to be caused by 2951d5c031a3 ("tick: broadcast: Prevent
> > livelock from event handler").
>
> Does reverting help? I see a similar case with your branch
> renesas-devel-20150629-v4.1 which does not include both commits.
This branch has actually none of the clockevent/tick related changes
which are in next/linus tree. So 2951d5c031a3 is hardly the culprit.
> For this branch, if I select CONFIG_NO_HZ_IDLE and *no* CONFIG_HIGH_RES_TIMERS,
So with high res timers it boots. Can you please provide the output of
/proc/timer_list for that case?
Thanks,
tglx
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Wolfram Sang <wsa@the-dreams.de> |
|---|---|
| Date | 2015-07-03 15:40 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pI9ma-E1-31@gated-at.bofh.it> |
| In reply to | #1176578 |
[Multipart message — attachments visible in raw view] — view raw
> So with high res timers it boots. Can you please provide the output of > /proc/timer_list for that case? Sure: Timer List Version: v0.7 HRTIMER_MAX_CLOCK_BASES: 4 now at 40752899169 nsecs cpu: 0 clock 0: .base: c03aa990 .index: 0 .resolution: 1 nsecs .get_time: ktime_get .offset: 0 nsecs active timers: #0: tick_cpu_sched, tick_sched_timer, S:01 # expires at 40760000000-40760000000 nsecs [in 7100831 to 7100831 nsecs] #1: sched_clock_timer, sched_clock_poll, S:01 # expires at 21474836475000000-21474836475000000 nsecs [in 21474795722100831 to 21474795722100831 nsecs] clock 1: .base: c03aa9c8 .index: 1 .resolution: 1 nsecs .get_time: ktime_get_real .offset: 0 nsecs active timers: clock 2: .base: c03aaa00 .index: 2 .resolution: 1 nsecs .get_time: ktime_get_boottime .offset: 0 nsecs active timers: clock 3: .base: c03aaa38 .index: 3 .resolution: 1 nsecs .get_time: ktime_get_clocktai .offset: 0 nsecs active timers: .expires_next : 40760000000 nsecs .hres_active : 1 .nr_events : 770 .nr_retries : 29 .nr_hangs : 0 .max_hang_time : 0 nsecs .nohz_mode : 2 .last_tick : 40400000000 nsecs .tick_stopped : 0 .idle_jiffies : 4294941337 .idle_calls : 378 .idle_sleeps : 354 .idle_entrytime : 40743316650 nsecs .idle_waketime : 40743316650 nsecs .idle_exittime : 40743377685 nsecs .idle_sleeptime : 37586425789 nsecs .iowait_sleeptime: 0 nsecs .last_jiffies : 4294941368 .next_jiffies : 4294941389 .idle_expires : 40930000000 nsecs jiffies: 4294941371 Tick Device: mode: 1 Per CPU device: 0 Clock Event Device: e0180000.timer max_delta_ns: 131071523464982 min_delta_ns: 61035 mult: 70369 shift: 31 mode: 3 next_event: 40760000000 nsecs set_next_event: em_sti_clock_event_next set_mode: em_sti_clock_event_mode event_handler: hrtimer_interrupt retries: 0
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2015-07-03 16:20 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pI9YS-17f-9@gated-at.bofh.it> |
| In reply to | #1176582 |
On Fri, 3 Jul 2015, Wolfram Sang wrote: > > > So with high res timers it boots. Can you please provide the output of > > /proc/timer_list for that case? > Tick Device: mode: 1 > Per CPU device: 0 > Clock Event Device: e0180000.timer > max_delta_ns: 131071523464982 > min_delta_ns: 61035 > mult: 70369 > shift: 31 > mode: 3 > next_event: 40760000000 nsecs > set_next_event: em_sti_clock_event_next > set_mode: em_sti_clock_event_mode > event_handler: hrtimer_interrupt > retries: 0 So this is a single core machine and uses the em_sti timer w/o the broadcast nonsense. In Simons case it looks like em_sti is used as broadcast device. Though the issues you see in the highres=n case might be the same as the ones Simon is observing. So in that nohz=y highres=n case, does adding idle=poll on the command line fix the issue? Thanks, tglx -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Wolfram Sang <wsa@the-dreams.de> |
|---|---|
| Date | 2015-07-03 16:40 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pIaif-1dJ-47@gated-at.bofh.it> |
| In reply to | #1176595 |
[Multipart message — attachments visible in raw view] — view raw
> So this is a single core machine and uses the em_sti timer w/o the > broadcast nonsense. In Simons case it looks like em_sti is used as > broadcast device. We use the same board. Just my kernel has SMP=n. > Though the issues you see in the highres=n case might be the same as > the ones Simon is observing. I hope so. One good thing about the issue I see is that it is 100% reproducable. > So in that nohz=y highres=n case, does adding idle=poll on the command > line fix the issue? Nope, still hangs. Thanks for assisting! Wolfram
[toc] | [prev] | [next] | [standalone]
| From | Sudeep Holla <sudeep.holla@arm.com> |
|---|---|
| Date | 2015-07-03 16:50 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pIarV-1ha-39@gated-at.bofh.it> |
| In reply to | #1176627 |
On 03/07/15 15:37, Wolfram Sang wrote: > >> So this is a single core machine and uses the em_sti timer w/o the >> broadcast nonsense. In Simons case it looks like em_sti is used as >> broadcast device. > > We use the same board. Just my kernel has SMP=n. > If it's UP build, then GENERIC_CLOCKEVENTS_BROADCAST is disabled. You may be hitting the issue I reported[1] and is still under discussion. Regards, Sudeep [1] https://lkml.org/lkml/2015/6/25/271 -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Sudeep Holla <sudeep.holla@arm.com> |
|---|---|
| Date | 2015-07-03 17:00 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pIaBA-1kt-5@gated-at.bofh.it> |
| In reply to | #1176640 |
On 03/07/15 15:54, Thomas Gleixner wrote: > On Fri, 3 Jul 2015, Sudeep Holla wrote: >> On 03/07/15 15:37, Wolfram Sang wrote: >>> >>>> So this is a single core machine and uses the em_sti timer w/o the >>>> broadcast nonsense. In Simons case it looks like em_sti is used as >>>> broadcast device. >>> >>> We use the same board. Just my kernel has SMP=n. >>> >> >> If it's UP build, then GENERIC_CLOCKEVENTS_BROADCAST is disabled. You >> may be hitting the issue I reported[1] and is still under discussion. > > Nope. The UP version uses a non affected timer. > Ah OK, sorry for the noise then. Regards, Sudeep -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2015-07-03 17:00 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pIaBA-1kt-7@gated-at.bofh.it> |
| In reply to | #1176640 |
On Fri, 3 Jul 2015, Sudeep Holla wrote: > On 03/07/15 15:37, Wolfram Sang wrote: > > > > > So this is a single core machine and uses the em_sti timer w/o the > > > broadcast nonsense. In Simons case it looks like em_sti is used as > > > broadcast device. > > > > We use the same board. Just my kernel has SMP=n. > > > > If it's UP build, then GENERIC_CLOCKEVENTS_BROADCAST is disabled. You > may be hitting the issue I reported[1] and is still under discussion. Nope. The UP version uses a non affected timer. Thanks, tglx -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2015-07-03 17:00 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pIaBA-1kt-19@gated-at.bofh.it> |
| In reply to | #1176627 |
On Fri, 3 Jul 2015, Wolfram Sang wrote: > > So this is a single core machine and uses the em_sti timer w/o the > > broadcast nonsense. In Simons case it looks like em_sti is used as > > broadcast device. > > We use the same board. Just my kernel has SMP=n. > > > Though the issues you see in the highres=n case might be the same as > > the ones Simon is observing. > > I hope so. One good thing about the issue I see is that it is 100% > reproducable. > > > So in that nohz=y highres=n case, does adding idle=poll on the command > > line fix the issue? > > Nope, still hangs. Ok. So it's unrelated to deep idle states. Any chance of poking with JTAG at the frozen box? If not, are there GPIOs which you could use to monitor certain state? Thanks, tglx -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Wolfram Sang <wsa@the-dreams.de> |
|---|---|
| Date | 2015-07-03 17:20 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pIaUV-1GA-11@gated-at.bofh.it> |
| In reply to | #1176645 |
[Multipart message — attachments visible in raw view] — view raw
> Ok. So it's unrelated to deep idle states. Any chance of poking with > JTAG at the frozen box? If not, are there GPIOs which you could use to > monitor certain state? No JTAGger here at the moment. And no manual/schematics for this board :( I'll ask around...
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2015-07-03 17:30 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pIb4C-1JT-23@gated-at.bofh.it> |
| In reply to | #1176651 |
On Fri, 3 Jul 2015, Wolfram Sang wrote: > > Ok. So it's unrelated to deep idle states. Any chance of poking with > > JTAG at the frozen box? If not, are there GPIOs which you could use to > > monitor certain state? > > No JTAGger here at the moment. And no manual/schematics for this board > :( Ok. One more check please. Does nohz=off on the command line fix/hide the issue as well? Thanks, tglx -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Wolfram Sang <wsa@the-dreams.de> |
|---|---|
| Date | 2015-07-03 17:50 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pIbnY-1T7-21@gated-at.bofh.it> |
| In reply to | #1176667 |
[Multipart message — attachments visible in raw view] — view raw
> Ok. One more check please. Does nohz=off on the command line fix/hide > the issue as well? Yes, it does. It boots to the prompt then.
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2015-07-03 18:00 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pIbxE-1Wv-11@gated-at.bofh.it> |
| In reply to | #1176684 |
On Fri, 3 Jul 2015, Wolfram Sang wrote: > > Ok. One more check please. Does nohz=off on the command line fix/hide > > the issue as well? > > Yes, it does. It boots to the prompt then. So something is fishy with this timer. In your UP setting we don't have any interaction with the broadcast stuff. It's just the single em_sti timer involved. Now looking at the code I notice, that this is one of the overly clever designed compare register based trainwrecks. Can you try the patch below, whether it makes a difference? Thanks, tglx --- diff --git a/drivers/clocksource/em_sti.c b/drivers/clocksource/em_sti.c index dc3c6ee04aaa..41d8035d294e 100644 --- a/drivers/clocksource/em_sti.c +++ b/drivers/clocksource/em_sti.c @@ -308,7 +308,7 @@ static void em_sti_register_clockevent(struct em_sti_priv *p) dev_info(&p->pdev->dev, "used for clock events\n"); /* Register with dummy 1 Hz value, gets updated in ->set_mode() */ - clockevents_config_and_register(ced, 1, 2, 0xffffffff); + clockevents_config_and_register(ced, 1, 100, 0xffffffff); } static int em_sti_probe(struct platform_device *pdev) -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Wolfram Sang <wsa@the-dreams.de> |
|---|---|
| Date | 2015-07-03 18:10 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pIbHj-2fv-13@gated-at.bofh.it> |
| In reply to | #1176687 |
[Multipart message — attachments visible in raw view] — view raw
> Can you try the patch below, whether it makes a difference? It doesn't. Board locks again. > - clockevents_config_and_register(ced, 1, 2, 0xffffffff); > + clockevents_config_and_register(ced, 1, 100, 0xffffffff);
[toc] | [prev] | [next] | [standalone]
| From | Geert Uytterhoeven <geert@linux-m68k.org> |
|---|---|
| Date | 2015-07-03 19:10 +0200 |
| Subject | Re: Possible regression due to "tick: broadcast: Prevent livelock from event handler" |
| Message-ID | <pIcDn-2PO-5@gated-at.bofh.it> |
| In reply to | #1176627 |
Hi Wolfram,
On Fri, Jul 3, 2015 at 4:37 PM, Wolfram Sang <wsa@the-dreams.de> wrote:
>> So this is a single core machine and uses the em_sti timer w/o the
>> broadcast nonsense. In Simons case it looks like em_sti is used as
>> broadcast device.
>
> We use the same board. Just my kernel has SMP=n.
Unlike our other multi-core A9 SoCs, emev2.dtsi doesn't have a node
for the arm,cortex-a9-twd-timer?
Gr{oetje,eeting}s,
Geert
--
Geert Uytterhoeven -- There's lots of Linux beyond ia32 -- geert@linux-m68k.org
In personal conversations with technical people, I call myself a hacker. But
when I'm talking to journalists I just say "programmer" or something like that.
-- Linus Torvalds
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web