Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1580727 > unrolled thread

Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot

Started byPavel Machek <pavel@ucw.cz>
First post2017-02-14 19:10 +0100
Last post2017-02-17 15:50 +0100
Articles 12 on this page of 32 — 6 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-14 19:10 +0100
    Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-14 20:30 +0100
      Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Alan Stern <stern@rowland.harvard.edu> - 2017-02-14 21:00 +0100
      Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Frederic Weisbecker <fweisbec@gmail.com> - 2017-02-23 17:30 +0100
        Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-23 19:50 +0100
          Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Frederic Weisbecker <fweisbec@gmail.com> - 2017-02-25 04:30 +0100
    Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-15 18:30 +0100
      Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-16 00:30 +0100
        Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Linus Torvalds <torvalds@linux-foundation.org> - 2017-02-16 00:40 +0100
          Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-16 12:20 +0100
            Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-16 18:30 +0100
              Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Frederic Weisbecker <fweisbec@gmail.com> - 2017-02-16 19:20 +0100
                Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Linus Torvalds <torvalds@linux-foundation.org> - 2017-02-16 19:30 +0100
                  Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Frederic Weisbecker <fweisbec@gmail.com> - 2017-02-16 19:40 +0100
                    Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Thomas Gleixner <tglx@linutronix.de> - 2017-02-16 20:40 +0100
                      Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-16 21:10 +0100
                        Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Linus Torvalds <torvalds@linux-foundation.org> - 2017-02-16 21:30 +0100
                          Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-16 21:50 +0100
                          Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-18 10:00 +0100
                        Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Greg Kroah-Hartman <gregkh@linuxfoundation.org> - 2017-02-17 02:20 +0100
                      Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Frederic Weisbecker <fweisbec@gmail.com> - 2017-02-17 15:10 +0100
                        Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Thomas Gleixner <tglx@linutronix.de> - 2017-02-17 17:40 +0100
                          Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-17 18:10 +0100
                            Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Frederic Weisbecker <fweisbec@gmail.com> - 2017-02-17 19:50 +0100
                              next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel  desktop, does not boot after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-18 10:40 +0100
                                Re: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel  desktop, does not boot after cold boots, boots after reboot Frederic Weisbecker <fweisbec@gmail.com> - 2017-02-18 16:00 +0100
                                  Re: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel  desktop, does not boot after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-18 19:10 +0100
                                    Re: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel  desktop, does not boot after cold boots, boots after reboot Frederic Weisbecker <fweisbec@gmail.com> - 2017-02-20 15:10 +0100
                                Re: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel  desktop, does not boot after cold boots, boots after reboot Frederic Weisbecker <fweisbec@gmail.com> - 2017-02-22 04:10 +0100
                                  Re: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel  desktop, does not boot after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-23 15:30 +0100
              Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Pavel Machek <pavel@ucw.cz> - 2017-02-16 20:10 +0100
                Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot  after cold boots, boots after reboot Frederic Weisbecker <fweisbec@gmail.com> - 2017-02-17 15:50 +0100

Page 2 of 2 — ← Prev page 1 [2]


#1583427

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-02-17 15:10 +0100
Message-ID<tbRou-5iA-19@gated-at.bofh.it>
In reply to#1582890
On Thu, Feb 16, 2017 at 08:34:45PM +0100, Thomas Gleixner wrote:
> On Thu, 16 Feb 2017, Frederic Weisbecker wrote:
> > On Thu, Feb 16, 2017 at 10:20:14AM -0800, Linus Torvalds wrote:
> > > On Thu, Feb 16, 2017 at 10:13 AM, Frederic Weisbecker
> > > <fweisbec@gmail.com> wrote:
> > > >
> > > > I haven't followed the discussion but this patch has a known issue which is fixed
> > > > with:
> > > >     7bdb59f1ad474bd7161adc8f923cdef10f2638d1
> > > >     "tick/nohz: Fix possible missing clock reprog after tick soft restart"
> > > >
> > > > I hope this fixes your issue.
> > > 
> > > No, Pavel saw the problem with rc8 too, which already has that fix.
> > > 
> > > So I think we'll just need to revert that original patch (and that
> > > means that we have to revert the commit you point to as well, since
> > > that ->next_tick field was added by the original commit).
> > 
> > Aw too bad, but indeed that late we don't have the choice.
> 
> Hint: Look for CPU hotplug interaction of these patches. I bet something
> becomes stale when the CPU goes down and does not get reset when it comes
> back online.

Indeed I should check that. But Pavel is seeing this on boot, where the
only hotplug operations that happen are CPU UP without preceding CPU DOWN
that may have retained stale values. I think the value of ts->next_tick should
be initially 0 for all CPUs. So perhaps that 0 value confuses stuff. But
looking at the code I don't see how. It maybe something more subtle.

[toc] | [prev] | [next] | [standalone]


#1583587

FromThomas Gleixner <tglx@linutronix.de>
Date2017-02-17 17:40 +0100
Message-ID<tbTJE-6GS-5@gated-at.bofh.it>
In reply to#1583427
On Fri, 17 Feb 2017, Frederic Weisbecker wrote:
> On Thu, Feb 16, 2017 at 08:34:45PM +0100, Thomas Gleixner wrote:
> > On Thu, 16 Feb 2017, Frederic Weisbecker wrote:
> > > On Thu, Feb 16, 2017 at 10:20:14AM -0800, Linus Torvalds wrote:
> > > > On Thu, Feb 16, 2017 at 10:13 AM, Frederic Weisbecker
> > > > <fweisbec@gmail.com> wrote:
> > > > >
> > > > > I haven't followed the discussion but this patch has a known issue which is fixed
> > > > > with:
> > > > >     7bdb59f1ad474bd7161adc8f923cdef10f2638d1
> > > > >     "tick/nohz: Fix possible missing clock reprog after tick soft restart"
> > > > >
> > > > > I hope this fixes your issue.
> > > > 
> > > > No, Pavel saw the problem with rc8 too, which already has that fix.
> > > > 
> > > > So I think we'll just need to revert that original patch (and that
> > > > means that we have to revert the commit you point to as well, since
> > > > that ->next_tick field was added by the original commit).
> > > 
> > > Aw too bad, but indeed that late we don't have the choice.
> > 
> > Hint: Look for CPU hotplug interaction of these patches. I bet something
> > becomes stale when the CPU goes down and does not get reset when it comes
> > back online.
> 
> Indeed I should check that. But Pavel is seeing this on boot, where the

I don't think so. He observed it on suspend resume and by doing hotplug
operations in a loop. But I might be wrong as usual.

> only hotplug operations that happen are CPU UP without preceding CPU DOWN
> that may have retained stale values. I think the value of ts->next_tick should
> be initially 0 for all CPUs. So perhaps that 0 value confuses stuff. But
> looking at the code I don't see how. It maybe something more subtle.
> 

[toc] | [prev] | [next] | [standalone]


#1583599

FromPavel Machek <pavel@ucw.cz>
Date2017-02-17 18:10 +0100
Message-ID<tbUcG-76j-11@gated-at.bofh.it>
In reply to#1583587

[Multipart message — attachments visible in raw view] — view raw

On Fri 2017-02-17 17:37:47, Thomas Gleixner wrote:
> On Fri, 17 Feb 2017, Frederic Weisbecker wrote:
> > On Thu, Feb 16, 2017 at 08:34:45PM +0100, Thomas Gleixner wrote:
> > > On Thu, 16 Feb 2017, Frederic Weisbecker wrote:
> > > > On Thu, Feb 16, 2017 at 10:20:14AM -0800, Linus Torvalds wrote:
> > > > > On Thu, Feb 16, 2017 at 10:13 AM, Frederic Weisbecker
> > > > > <fweisbec@gmail.com> wrote:
> > > > > >
> > > > > > I haven't followed the discussion but this patch has a known issue which is fixed
> > > > > > with:
> > > > > >     7bdb59f1ad474bd7161adc8f923cdef10f2638d1
> > > > > >     "tick/nohz: Fix possible missing clock reprog after tick soft restart"
> > > > > >
> > > > > > I hope this fixes your issue.
> > > > > 
> > > > > No, Pavel saw the problem with rc8 too, which already has that fix.
> > > > > 
> > > > > So I think we'll just need to revert that original patch (and that
> > > > > means that we have to revert the commit you point to as well, since
> > > > > that ->next_tick field was added by the original commit).
> > > > 
> > > > Aw too bad, but indeed that late we don't have the choice.
> > > 
> > > Hint: Look for CPU hotplug interaction of these patches. I bet something
> > > becomes stale when the CPU goes down and does not get reset when it comes
> > > back online.
> > 
> > Indeed I should check that. But Pavel is seeing this on boot, where the
> 
> I don't think so. He observed it on suspend resume and by doing hotplug
> operations in a loop. But I might be wrong as usual.

These are different bugs.

On x60, I see failures doing hotplug/unplug in a loop, or lot of
suspends. Someone seen it in v4.8-stable etc. Old bug. Rare to hit.

Desktop machine was failing to boot, and had some fun with
suspend/resume too. Boot hang was reproducible with right
procedure. (Hard poweroff, cold boot.). That one was introduced in
4.10-rc cycle.


									Pavel

-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1583651

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-02-17 19:50 +0100
Message-ID<tbVLs-7V6-15@gated-at.bofh.it>
In reply to#1583599
On Fri, Feb 17, 2017 at 06:05:08PM +0100, Pavel Machek wrote:
> On Fri 2017-02-17 17:37:47, Thomas Gleixner wrote:
> > On Fri, 17 Feb 2017, Frederic Weisbecker wrote:
> > > On Thu, Feb 16, 2017 at 08:34:45PM +0100, Thomas Gleixner wrote:
> > > > On Thu, 16 Feb 2017, Frederic Weisbecker wrote:
> > > > > On Thu, Feb 16, 2017 at 10:20:14AM -0800, Linus Torvalds wrote:
> > > > > > On Thu, Feb 16, 2017 at 10:13 AM, Frederic Weisbecker
> > > > > > <fweisbec@gmail.com> wrote:
> > > > > > >
> > > > > > > I haven't followed the discussion but this patch has a known issue which is fixed
> > > > > > > with:
> > > > > > >     7bdb59f1ad474bd7161adc8f923cdef10f2638d1
> > > > > > >     "tick/nohz: Fix possible missing clock reprog after tick soft restart"
> > > > > > >
> > > > > > > I hope this fixes your issue.
> > > > > > 
> > > > > > No, Pavel saw the problem with rc8 too, which already has that fix.
> > > > > > 
> > > > > > So I think we'll just need to revert that original patch (and that
> > > > > > means that we have to revert the commit you point to as well, since
> > > > > > that ->next_tick field was added by the original commit).
> > > > > 
> > > > > Aw too bad, but indeed that late we don't have the choice.
> > > > 
> > > > Hint: Look for CPU hotplug interaction of these patches. I bet something
> > > > becomes stale when the CPU goes down and does not get reset when it comes
> > > > back online.
> > > 
> > > Indeed I should check that. But Pavel is seeing this on boot, where the
> > 
> > I don't think so. He observed it on suspend resume and by doing hotplug
> > operations in a loop. But I might be wrong as usual.
> 
> These are different bugs.
> 
> On x60, I see failures doing hotplug/unplug in a loop, or lot of
> suspends. Someone seen it in v4.8-stable etc. Old bug. Rare to hit.
> 
> Desktop machine was failing to boot, and had some fun with
> suspend/resume too. Boot hang was reproducible with right
> procedure. (Hard poweroff, cold boot.). That one was introduced in
> 4.10-rc cycle.

Pavel, is there any chance you could apply this patch on top of latest linus tree
and send me your resulting dmesg log? This has the two reverted patches plus some
debugging code. The amount of printk shouldn't be too big, I tested it home without
issue.

If you can't manage to dump the dmesg, please try to take a picture of your screen
so that I can see the last messages starting with "NEXT_TICK_READ".

Thanks!

diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
index 2c115fd..504cb41 100644
--- a/kernel/time/tick-sched.c
+++ b/kernel/time/tick-sched.c
@@ -658,6 +658,8 @@ static void tick_nohz_restart(struct tick_sched *ts, ktime_t now)
 		tick_program_event(hrtimer_get_expires(&ts->sched_timer), 1);
 }
 
+static DEFINE_PER_CPU(u64, prev_next_tick);
+
 static ktime_t tick_nohz_stop_sched_tick(struct tick_sched *ts,
 					 ktime_t now, int cpu)
 {
@@ -725,6 +727,11 @@ static ktime_t tick_nohz_stop_sched_tick(struct tick_sched *ts,
 		 */
 		if (delta == 0) {
 			tick_nohz_restart(ts, now);
+			/*
+			 * Make sure next tick stop doesn't get fooled by past
+			 * clock deadline
+			 */
+			ts->next_tick = 0;
 			goto out;
 		}
 	}
@@ -767,8 +774,15 @@ static ktime_t tick_nohz_stop_sched_tick(struct tick_sched *ts,
 	tick = expires;
 
 	/* Skip reprogram of event if its not changed */
-	if (ts->tick_stopped && (expires == dev->next_event))
-		goto out;
+	if (ts->tick_stopped) {
+		if (system_state == SYSTEM_BOOTING) {
+			if (ts->next_tick != this_cpu_read(prev_next_tick))
+				printk("NEXT_TICK_READ: CPU: %d Expires: %llu ts->next_tick:%llu\n", smp_processor_id(), expires, ts->next_tick);
+			this_cpu_write(prev_next_tick, ts->next_tick);
+		}
+		if (expires == ts->next_tick)
+			goto out;
+	}
 
 	/*
 	 * nohz_stop_sched_tick can be called several times before
@@ -787,6 +801,8 @@ static ktime_t tick_nohz_stop_sched_tick(struct tick_sched *ts,
 		trace_tick_stop(1, TICK_DEP_MASK_NONE);
 	}
 
+	ts->next_tick = tick;
+
 	/*
 	 * If the expiration time == KTIME_MAX, then we simply stop
 	 * the tick timer.
@@ -802,7 +818,10 @@ static ktime_t tick_nohz_stop_sched_tick(struct tick_sched *ts,
 	else
 		tick_program_event(tick, 1);
 out:
-	/* Update the estimated sleep length */
+	/*
+	 * Update the estimated sleep length until the next timer
+	 * (not only the tick).
+	 */
 	ts->sleep_length = ktime_sub(dev->next_event, now);
 	return tick;
 }
diff --git a/kernel/time/tick-sched.h b/kernel/time/tick-sched.h
index bf38226..075444e 100644
--- a/kernel/time/tick-sched.h
+++ b/kernel/time/tick-sched.h
@@ -27,6 +27,7 @@ enum tick_nohz_mode {
  *			timer is modified for nohz sleeps. This is necessary
  *			to resume the tick timer operation in the timeline
  *			when the CPU returns from nohz sleep.
+ * @next_tick:		Next tick to be fired when in dynticks mode.
  * @tick_stopped:	Indicator that the idle tick has been stopped
  * @idle_jiffies:	jiffies at the entry to idle for idle time accounting
  * @idle_calls:		Total number of idle calls
@@ -44,6 +45,7 @@ struct tick_sched {
 	unsigned long			check_clocks;
 	enum tick_nohz_mode		nohz_mode;
 	ktime_t				last_tick;
+	ktime_t				next_tick;
 	int				inidle;
 	int				tick_stopped;
 	unsigned long			idle_jiffies;

[toc] | [prev] | [next] | [standalone]


#1583876 — next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot

FromPavel Machek <pavel@ucw.cz>
Date2017-02-18 10:40 +0100
Subjectnext_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot
Message-ID<tc9EK-5W-25@gated-at.bofh.it>
In reply to#1583651

[Multipart message — attachments visible in raw view] — view raw

Hi!

[I droped some CCs here, you may want to check the CC list].

> > These are different bugs.
> > 
> > On x60, I see failures doing hotplug/unplug in a loop, or lot of
> > suspends. Someone seen it in v4.8-stable etc. Old bug. Rare to hit.
> > 
> > Desktop machine was failing to boot, and had some fun with
> > suspend/resume too. Boot hang was reproducible with right
> > procedure. (Hard poweroff, cold boot.). That one was introduced in
> > 4.10-rc cycle.
> 
> Pavel, is there any chance you could apply this patch on top of latest linus tree
> and send me your resulting dmesg log? This has the two reverted patches plus some
> debugging code. The amount of printk shouldn't be too big, I tested it home without
> issue.
> 
> If you can't manage to dump the dmesg, please try to take a picture of your screen
> so that I can see the last messages starting with "NEXT_TICK_READ".
> 
> Thanks!

I guess I can. But I'll only have one 80x25 screen to look at...

.config is attached.
									Pavel

-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1583909 — Re: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-02-18 16:00 +0100
SubjectRe: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot
Message-ID<tceEq-3aL-3@gated-at.bofh.it>
In reply to#1583876
On Sat, Feb 18, 2017 at 10:39:17AM +0100, Pavel Machek wrote:
> Hi!
> 
> [I droped some CCs here, you may want to check the CC list].

Good!

> 
> > > These are different bugs.
> > > 
> > > On x60, I see failures doing hotplug/unplug in a loop, or lot of
> > > suspends. Someone seen it in v4.8-stable etc. Old bug. Rare to hit.
> > > 
> > > Desktop machine was failing to boot, and had some fun with
> > > suspend/resume too. Boot hang was reproducible with right
> > > procedure. (Hard poweroff, cold boot.). That one was introduced in
> > > 4.10-rc cycle.
> > 
> > Pavel, is there any chance you could apply this patch on top of latest linus tree
> > and send me your resulting dmesg log? This has the two reverted patches plus some
> > debugging code. The amount of printk shouldn't be too big, I tested it home without
> > issue.
> > 
> > If you can't manage to dump the dmesg, please try to take a picture of your screen
> > so that I can see the last messages starting with "NEXT_TICK_READ".
> > 
> > Thanks!
> 
> I guess I can. But I'll only have one 80x25 screen to look at...
> 
> .config is attached.

Ah this is x86-32, interesting! I'm going to try to boot that, we never know.

Thanks a lot!

[toc] | [prev] | [next] | [standalone]


#1583941 — Re: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot

FromPavel Machek <pavel@ucw.cz>
Date2017-02-18 19:10 +0100
SubjectRe: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot
Message-ID<tchCi-59d-5@gated-at.bofh.it>
In reply to#1583909

[Multipart message — attachments visible in raw view] — view raw

Hi!

> > I guess I can. But I'll only have one 80x25 screen to look at...
> > 
> > .config is attached.
> 
> Ah this is x86-32, interesting! I'm going to try to boot that, we never know.
> 
> Thanks a lot!

Happens on x86-64, too; I'm running that normally, but for testing,
32-bit kernel is easier.

thinkpad x60 works fine for me, so it is unlikely that .config is all
it takes... 
									Pavel
-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1584637 — Re: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-02-20 15:10 +0100
SubjectRe: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot
Message-ID<tcWP7-5GV-5@gated-at.bofh.it>
In reply to#1583941
On Sat, Feb 18, 2017 at 07:05:20PM +0100, Pavel Machek wrote:
> Hi!
> 
> > > I guess I can. But I'll only have one 80x25 screen to look at...
> > > 
> > > .config is attached.
> > 
> > Ah this is x86-32, interesting! I'm going to try to boot that, we never know.
> > 
> > Thanks a lot!
> 
> Happens on x86-64, too; I'm running that normally, but for testing,
> 32-bit kernel is easier.

Ah! And you've seen that on only one machine? What kind machine is it?

Ideally I would need a dump of all pending timer list timers (no sysrq key
for that though, but I can do a quick patch) and a stacktrace of all
tasks. But I guess you have no access to any serial port, right?

> 
> thinkpad x60 works fine for me, so it is unlikely that .config is all
> it takes...

Yeah I booted the .config and it reached the root filesystem mounting
without problem. So I think it's specific to some hardware.

[toc] | [prev] | [next] | [standalone]


#1585900 — Re: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-02-22 04:10 +0100
SubjectRe: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot
Message-ID<tdvtw-3w9-13@gated-at.bofh.it>
In reply to#1583876
On Sat, Feb 18, 2017 at 11:23:39AM +0100, Pavel Machek wrote:
> On Sat 2017-02-18 10:39:17, Pavel Machek wrote:
> > Hi!
> > 
> > [I droped some CCs here, you may want to check the CC list].
> > 
> > > > These are different bugs.
> > > > 
> > > > On x60, I see failures doing hotplug/unplug in a loop, or lot of
> > > > suspends. Someone seen it in v4.8-stable etc. Old bug. Rare to hit.
> > > > 
> > > > Desktop machine was failing to boot, and had some fun with
> > > > suspend/resume too. Boot hang was reproducible with right
> > > > procedure. (Hard poweroff, cold boot.). That one was introduced in
> > > > 4.10-rc cycle.
> > > 
> > > Pavel, is there any chance you could apply this patch on top of latest linus tree
> > > and send me your resulting dmesg log? This has the two reverted patches plus some
> > > debugging code. The amount of printk shouldn't be too big, I tested it home without
> > > issue.
> > > 
> > > If you can't manage to dump the dmesg, please try to take a picture of your screen
> > > so that I can see the last messages starting with "NEXT_TICK_READ".
> > > 
> > > Thanks!
> > 
> > I guess I can. But I'll only have one 80x25 screen to look at...
> 
> Ok, here it is.

Thanks, I haven't been able to deduce much though, except that the pending timer on CPU 0
looks quite far away.

Could you please add "initcall_debug" in your kernel parameters to identify if we are blocking in
a specific initcall? If so it should tell us which one.

Thanks!

[toc] | [prev] | [next] | [standalone]


#1586916 — Re: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot

FromPavel Machek <pavel@ucw.cz>
Date2017-02-23 15:30 +0100
SubjectRe: next_tick hang was Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does not boot after cold boots, boots after reboot
Message-ID<te2z7-2s9-13@gated-at.bofh.it>
In reply to#1585900

[Multipart message — attachments visible in raw view] — view raw

On Wed 2017-02-22 04:08:58, Frederic Weisbecker wrote:
> On Sat, Feb 18, 2017 at 11:23:39AM +0100, Pavel Machek wrote:
> > On Sat 2017-02-18 10:39:17, Pavel Machek wrote:
> > > Hi!
> > > 
> > > [I droped some CCs here, you may want to check the CC list].
> > > 
> > > > > These are different bugs.
> > > > > 
> > > > > On x60, I see failures doing hotplug/unplug in a loop, or lot of
> > > > > suspends. Someone seen it in v4.8-stable etc. Old bug. Rare to hit.
> > > > > 
> > > > > Desktop machine was failing to boot, and had some fun with
> > > > > suspend/resume too. Boot hang was reproducible with right
> > > > > procedure. (Hard poweroff, cold boot.). That one was introduced in
> > > > > 4.10-rc cycle.
> > > > 
> > > > Pavel, is there any chance you could apply this patch on top of latest linus tree
> > > > and send me your resulting dmesg log? This has the two reverted patches plus some
> > > > debugging code. The amount of printk shouldn't be too big, I tested it home without
> > > > issue.
> > > > 
> > > > If you can't manage to dump the dmesg, please try to take a picture of your screen
> > > > so that I can see the last messages starting with "NEXT_TICK_READ".
> > > > 
> > > > Thanks!
> > > 
> > > I guess I can. But I'll only have one 80x25 screen to look at...
> > 
> > Ok, here it is.
> 
> Thanks, I haven't been able to deduce much though, except that the pending timer on CPU 0
> looks quite far away.
> 
> Could you please add "initcall_debug" in your kernel parameters to identify if we are blocking in
> a specific initcall? If so it should tell us which one.

Please see Re: v4.10-rc8 (-rc6) boot regression on Intel desktop, does
not boot after cold boots, boots after reboot thread. Hang was traced
down to the USB handoff code.
									Pavel

-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1582871

FromPavel Machek <pavel@ucw.cz>
Date2017-02-16 20:10 +0100
Message-ID<tbzBg-2gu-13@gated-at.bofh.it>
In reply to#1582720

[Multipart message — attachments visible in raw view] — view raw

On Thu 2017-02-16 18:25:35, Pavel Machek wrote:
> Hi!
> 
> > > > 4.10-rc4 broken
> > > > 4.10-rc3 ok
> > > 
> > > Hmm. If those actually end up being reliable, then there's not a whole
> > > lot in between them wrt PCI or USB.
> > > 
> > > What looked like the most likely candidate seems to be xhci-specific, though.
> > > 
> > > But maybe it's something that isn't directly in drivers/{pci,usb}/ and
> > > just interacts badly.
> > 
> > Ok. I _hope_ my tests are ok. Bisect log so far is:
> 
> And the winner is:
> 
> pavel@half:/data/l/linux$ git bisect bad
> 24b91e360ef521a2808771633d76ebc68bd5604b is the first bad commit
> commit 24b91e360ef521a2808771633d76ebc68bd5604b
> Author: Frederic Weisbecker <fweisbec@gmail.com>
> Date:   Wed Jan 4 15:12:04 2017 +0100
> 
>     nohz: Fix collision between tick and other hrtimers
>     

I had to revert 7bdb59f1ad474bd7161adc8f923cdef10f2638d1, too,
otherwise -rc8 does not compile.

With 24b91e360ef521a28087716 and 7bdb59f1ad474 reverted, it seems to
boot ok. (I did few tries.)

Best regards,
								Pavel

-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1583496

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2017-02-17 15:50 +0100
Message-ID<tbS1c-5xl-43@gated-at.bofh.it>
In reply to#1582871
On Thu, Feb 16, 2017 at 08:06:04PM +0100, Pavel Machek wrote:
> On Thu 2017-02-16 18:25:35, Pavel Machek wrote:
> > Hi!
> > 
> > > > > 4.10-rc4 broken
> > > > > 4.10-rc3 ok
> > > > 
> > > > Hmm. If those actually end up being reliable, then there's not a whole
> > > > lot in between them wrt PCI or USB.
> > > > 
> > > > What looked like the most likely candidate seems to be xhci-specific, though.
> > > > 
> > > > But maybe it's something that isn't directly in drivers/{pci,usb}/ and
> > > > just interacts badly.
> > > 
> > > Ok. I _hope_ my tests are ok. Bisect log so far is:
> > 
> > And the winner is:
> > 
> > pavel@half:/data/l/linux$ git bisect bad
> > 24b91e360ef521a2808771633d76ebc68bd5604b is the first bad commit
> > commit 24b91e360ef521a2808771633d76ebc68bd5604b
> > Author: Frederic Weisbecker <fweisbec@gmail.com>
> > Date:   Wed Jan 4 15:12:04 2017 +0100
> > 
> >     nohz: Fix collision between tick and other hrtimers
> >     
> 
> I had to revert 7bdb59f1ad474bd7161adc8f923cdef10f2638d1, too,
> otherwise -rc8 does not compile.
> 
> With 24b91e360ef521a28087716 and 7bdb59f1ad474 reverted, it seems to
> boot ok. (I did few tries.)

Do you still have the config that triggered this? I don't have much expectations
about reproducing, this has almost never worked for me, but at least I could narrow
down the context.

Thanks.

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web