Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1726476 > unrolled thread

Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

Started byMarkus Trippelsdorf <markus@trippelsdorf.de>
First post2017-09-05 09:40 +0200
Last post2017-09-08 17:00 +0200
Articles 20 on this page of 53 — 8 participants

Back to article view | Back to linux.kernel


Contents

  Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-05 09:40 +0200
    Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Peter Zijlstra <peterz@infradead.org> - 2017-09-05 11:00 +0200
      Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-05 12:00 +0200
        Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Thomas Gleixner <tglx@linutronix.de> - 2017-09-06 15:00 +0200
          Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-06 15:20 +0200
            Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-07 08:30 +0200
              Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Thomas Gleixner <tglx@linutronix.de> - 2017-09-08 08:30 +0200
                Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-08 10:10 +0200
                  Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-08 11:20 +0200
                    Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 11:50 +0200
                      Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Ingo Molnar <mingo@kernel.org> - 2017-09-08 12:40 +0200
                        Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 12:40 +0200
                          Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 13:40 +0200
                            Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-08 18:20 +0200
                              Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 19:20 +0200
                                Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-08 23:50 +0200
                                  Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 00:00 +0200
                                    Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 01:10 +0200
                                      Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 01:30 +0200
                                        Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 02:10 +0200
                                          Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 03:10 +0200
                                            Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@amacapital.net> - 2017-09-09 03:40 +0200
                                              Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 20:00 +0200
                                                Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 20:10 +0200
                                    Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 08:40 +0200
                                      Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 12:20 +0200
                                        Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 13:10 +0200
                                          Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 15:10 +0200
                                            Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 15:40 +0200
                                              Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 15:50 +0200
                                                Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 16:10 +0200
                                                  Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 16:30 +0200
                                                    Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 16:40 +0200
                                                      Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 16:50 +0200
                                                        Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 18:40 +0200
                                                          Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 19:10 +0200
                                                            Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 19:30 +0200
                                                              Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 19:40 +0200
                                                                Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 20:20 +0200
                                                                  Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 20:30 +0200
                                                                    Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 20:50 +0200
                                                                      Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:20 +0200
                                                                      Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:20 +0200
                                                                  Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 20:30 +0200
                                                                    Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 20:40 +0200
                                                                      Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 20:50 +0200
                                                                        Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:20 +0200
                                                                          Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 21:30 +0200
                                                                            Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:40 +0200
                                                                              Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-10 06:50 +0200
                                                                          Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 21:30 +0200
                                  Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 10:20 +0200
                    Re: Current mainline git (24e700e291d52bd2) hangs when building e.g.  perf Borislav Petkov <bp@alien8.de> - 2017-09-08 17:00 +0200

Page 1 of 3  [1] 2 3  Next page →


#1726476 — Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromMarkus Trippelsdorf <markus@trippelsdorf.de>
Date2017-09-05 09:40 +0200
SubjectCurrent mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<umgCJ-8ei-3@gated-at.bofh.it>
Current mainline git (24e700e291d52bd2) hangs when building software
concurrently (for example perf).
The issue is not 100% reproducible (sometimes building perf succeeds),
so bisecting will not work.
Magic SysRq key doesn't work and there is nothing in the logs.
Enabling CONFIG_PROVE_LOCKING makes the issue go away.

Any ideas on how to debug this further?

-- 
Markus

[toc] | [next] | [standalone]


#1726561 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromPeter Zijlstra <peterz@infradead.org>
Date2017-09-05 11:00 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<umhSa-t5-9@gated-at.bofh.it>
In reply to#1726476
On Tue, Sep 05, 2017 at 09:27:38AM +0200, Markus Trippelsdorf wrote:
> Current mainline git (24e700e291d52bd2) hangs when building software
> concurrently (for example perf).
> The issue is not 100% reproducible (sometimes building perf succeeds),
> so bisecting will not work.

Sadly I cannot reproduce, I had:

  while :; do make clean; make; done

running on tools/perf for a while, and now have:

  while :; do make O=defconfig-build clean; make O=defconfig-build -j80; done

running, all smooth sailing, although there's the hope that the moment I
hit send on this email the box comes unstuck.

> Magic SysRq key doesn't work and there is nothing in the logs.
> Enabling CONFIG_PROVE_LOCKING makes the issue go away.

SysRq not working is suspicious.. and I take it the NMI watchdog also
isn't firing?

> Any ideas on how to debug this further?

So you have a (real) serial line on that box?

Could you try something like:

  debug ignore_loglevel sysrq_always_enabled earlyprintk=serial,ttyS0,115200 force_early_printk

with the below patch applied? That always gives me the most reliable
output.

---
 kernel/printk/printk.c | 119 +++++++++++++++++++++++++++++++++++--------------
 1 file changed, 86 insertions(+), 33 deletions(-)

diff --git a/kernel/printk/printk.c b/kernel/printk/printk.c
index fc47863f629c..b17099fbc7ce 100644
--- a/kernel/printk/printk.c
+++ b/kernel/printk/printk.c
@@ -365,6 +365,75 @@ __packed __aligned(4)
 #endif
 ;
 
+#ifdef CONFIG_EARLY_PRINTK
+struct console *early_console;
+
+static bool __read_mostly force_early_printk;
+
+static int __init force_early_printk_setup(char *str)
+{
+	force_early_printk = true;
+	return 0;
+}
+early_param("force_early_printk", force_early_printk_setup);
+
+static int early_printk_cpu = -1;
+
+static int early_vprintk(const char *fmt, va_list args)
+{
+	int n, cpu, old;
+	char buf[512];
+
+	cpu = get_cpu();
+	/*
+	 * Test-and-Set inter-cpu spinlock with recursion.
+	 */
+	for (;;) {
+		/*
+		 * c-cas to avoid the exclusive bouncing on spin.
+		 * Depends on the memory barrier implied by cmpxchg
+		 * for ACQUIRE semantics.
+		 */
+		old = READ_ONCE(early_printk_cpu);
+		if (old == -1) {
+			old = cmpxchg(&early_printk_cpu, -1, cpu);
+			if (old == -1)
+				break;
+		}
+		/*
+		 * Allow recursion for interrupts and the like.
+		 */
+		if (old == cpu)
+			break;
+
+		cpu_relax();
+	}
+
+	n = vscnprintf(buf, sizeof(buf), fmt, args);
+	early_console->write(early_console, buf, n);
+
+	/*
+	 * Unlock -- in case @old == @cpu, this is a no-op.
+	 */
+	smp_store_release(&early_printk_cpu, old);
+	put_cpu();
+
+	return n;
+}
+
+asmlinkage __visible void early_printk(const char *fmt, ...)
+{
+	va_list ap;
+
+	if (!early_console)
+		return;
+
+	va_start(ap, fmt);
+	early_vprintk(fmt, ap);
+	va_end(ap);
+}
+#endif
+
 /*
  * The logbuf_lock protects kmsg buffer, indices, counters.  This can be taken
  * within the scheduler's rq lock. It must be released before calling
@@ -1704,6 +1773,16 @@ asmlinkage int vprintk_emit(int facility, int level,
 	int printed_len = 0;
 	bool in_sched = false;
 
+#ifdef CONFIG_KGDB_KDB
+	if (unlikely(kdb_trap_printk && kdb_printf_cpu < 0))
+		return vkdb_printf(KDB_MSGSRC_PRINTK, fmt, args);
+#endif
+
+#ifdef CONFIG_EARLY_PRINTK
+	if (force_early_printk && early_console)
+		return early_vprintk(fmt, args);
+#endif
+
 	if (level == LOGLEVEL_SCHED) {
 		level = LOGLEVEL_DEFAULT;
 		in_sched = true;
@@ -1796,18 +1875,7 @@ EXPORT_SYMBOL(printk_emit);
 
 int vprintk_default(const char *fmt, va_list args)
 {
-	int r;
-
-#ifdef CONFIG_KGDB_KDB
-	/* Allow to pass printk() to kdb but avoid a recursion. */
-	if (unlikely(kdb_trap_printk && kdb_printf_cpu < 0)) {
-		r = vkdb_printf(KDB_MSGSRC_PRINTK, fmt, args);
-		return r;
-	}
-#endif
-	r = vprintk_emit(0, LOGLEVEL_DEFAULT, NULL, 0, fmt, args);
-
-	return r;
+	return vprintk_emit(0, LOGLEVEL_DEFAULT, NULL, 0, fmt, args);
 }
 EXPORT_SYMBOL_GPL(vprintk_default);
 
@@ -1838,7 +1906,12 @@ asmlinkage __visible int printk(const char *fmt, ...)
 	int r;
 
 	va_start(args, fmt);
-	r = vprintk_func(fmt, args);
+#ifdef CONFIG_EARLY_PRINTK
+	if (force_early_printk && early_console)
+		r = vprintk_default(fmt, args);
+	else
+#endif
+		r = vprintk_func(fmt, args);
 	va_end(args);
 
 	return r;
@@ -1875,26 +1948,6 @@ static bool suppress_message_printing(int level) { return false; }
 
 #endif /* CONFIG_PRINTK */
 
-#ifdef CONFIG_EARLY_PRINTK
-struct console *early_console;
-
-asmlinkage __visible void early_printk(const char *fmt, ...)
-{
-	va_list ap;
-	char buf[512];
-	int n;
-
-	if (!early_console)
-		return;
-
-	va_start(ap, fmt);
-	n = vscnprintf(buf, sizeof(buf), fmt, ap);
-	va_end(ap);
-
-	early_console->write(early_console, buf, n);
-}
-#endif
-
 static int __add_preferred_console(char *name, int idx, char *options,
 				   char *brl_options)
 {

[toc] | [prev] | [next] | [standalone]


#1726593 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromMarkus Trippelsdorf <markus@trippelsdorf.de>
Date2017-09-05 12:00 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<umiOe-13p-9@gated-at.bofh.it>
In reply to#1726561
On 2017.09.05 at 10:53 +0200, Peter Zijlstra wrote:
> On Tue, Sep 05, 2017 at 09:27:38AM +0200, Markus Trippelsdorf wrote:
> > Current mainline git (24e700e291d52bd2) hangs when building software
> > concurrently (for example perf).
> > The issue is not 100% reproducible (sometimes building perf succeeds),
> > so bisecting will not work.
> 
> Sadly I cannot reproduce, I had:
> 
>   while :; do make clean; make; done
> 
> running on tools/perf for a while, and now have:
> 
>   while :; do make O=defconfig-build clean; make O=defconfig-build -j80; done
> 
> running, all smooth sailing, although there's the hope that the moment I
> hit send on this email the box comes unstuck.
> 
> > Magic SysRq key doesn't work and there is nothing in the logs.
> > Enabling CONFIG_PROVE_LOCKING makes the issue go away.
> 
> SysRq not working is suspicious.. and I take it the NMI watchdog also
> isn't firing?

Yes.

> > Any ideas on how to debug this further?
> 
> So you have a (real) serial line on that box?

Sadly, no. But hopefully somebody else (with a proper kernel debugging
setup) will reproduce the issue soon.

-- 
Markus

[toc] | [prev] | [next] | [standalone]


#1727427 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromThomas Gleixner <tglx@linutronix.de>
Date2017-09-06 15:00 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<umI5Y-1TZ-33@gated-at.bofh.it>
In reply to#1726593
On Tue, 5 Sep 2017, Markus Trippelsdorf wrote:
> On 2017.09.05 at 10:53 +0200, Peter Zijlstra wrote:
> > > Any ideas on how to debug this further?
> > 
> > So you have a (real) serial line on that box?
> 
> Sadly, no. But hopefully somebody else (with a proper kernel debugging
> setup) will reproduce the issue soon.

Does the machine respond to ping or is it entirely dead?

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


#1727445 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromMarkus Trippelsdorf <markus@trippelsdorf.de>
Date2017-09-06 15:20 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<umIpj-2iw-9@gated-at.bofh.it>
In reply to#1727427
On 2017.09.06 at 14:52 +0200, Thomas Gleixner wrote:
> On Tue, 5 Sep 2017, Markus Trippelsdorf wrote:
> > On 2017.09.05 at 10:53 +0200, Peter Zijlstra wrote:
> > > > Any ideas on how to debug this further?
> > > 
> > > So you have a (real) serial line on that box?
> > 
> > Sadly, no. But hopefully somebody else (with a proper kernel debugging
> > setup) will reproduce the issue soon.
> 
> Does the machine respond to ping or is it entirely dead?

It is entirely dead and doesn't respond to ping.

-- 
Markus

[toc] | [prev] | [next] | [standalone]


#1727943 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromMarkus Trippelsdorf <markus@trippelsdorf.de>
Date2017-09-07 08:30 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<umYu6-4DP-7@gated-at.bofh.it>
In reply to#1727445
On 2017.09.06 at 15:15 +0200, Markus Trippelsdorf wrote:
> On 2017.09.06 at 14:52 +0200, Thomas Gleixner wrote:
> > On Tue, 5 Sep 2017, Markus Trippelsdorf wrote:
> > > On 2017.09.05 at 10:53 +0200, Peter Zijlstra wrote:
> > > > > Any ideas on how to debug this further?
> > > > 
> > > > So you have a (real) serial line on that box?
> > > 
> > > Sadly, no. But hopefully somebody else (with a proper kernel debugging
> > > setup) will reproduce the issue soon.
> > 
> > Does the machine respond to ping or is it entirely dead?
> 
> It is entirely dead and doesn't respond to ping.

The bug even kills the host (running 4.13) when running 24e700e2 in qemu
(kvm) and compiling stuff in parallel in the guest.
I see an RCU CPU stall in dmesg (on the host), but unfortunately cannot
save it, because nothing gets written to disk after the stall.
Connecting to qemu via gdb also doesn't work.

-- 
Markus

[toc] | [prev] | [next] | [standalone]


#1728628 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromThomas Gleixner <tglx@linutronix.de>
Date2017-09-08 08:30 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<unkXF-2YQ-9@gated-at.bofh.it>
In reply to#1727943
On Fri, 8 Sep 2017, Markus Trippelsdorf wrote:

CC+ Borislav. He might have access to such a beast

> On 2017.09.07 at 08:28 +0200, Markus Trippelsdorf wrote:
> > On 2017.09.06 at 15:15 +0200, Markus Trippelsdorf wrote:
> > > On 2017.09.06 at 14:52 +0200, Thomas Gleixner wrote:
> > > > On Tue, 5 Sep 2017, Markus Trippelsdorf wrote:
> > > > > On 2017.09.05 at 10:53 +0200, Peter Zijlstra wrote:
> > > > > > > Any ideas on how to debug this further?
> > > > > > 
> > > > > > So you have a (real) serial line on that box?
> > > > > 
> > > > > Sadly, no. But hopefully somebody else (with a proper kernel debugging
> > > > > setup) will reproduce the issue soon.
> > > > 
> > > > Does the machine respond to ping or is it entirely dead?
> > > 
> > > It is entirely dead and doesn't respond to ping.
> > 
> > The bug even kills the host (running 4.13) when running 24e700e2 in qemu
> > (kvm) and compiling stuff in parallel in the guest.
> > I see an RCU CPU stall in dmesg (on the host), but unfortunately cannot
> > save it, because nothing gets written to disk after the stall.
> > Connecting to qemu via gdb also doesn't work.
> 
> My guess would be a bug in a low level function (asm) that only hits AMD
> machines. I'm running an old Phenom II X4 processor. My config is
> attached.
> 
> -- 
> Markus
> 

[toc] | [prev] | [next] | [standalone]


#1728671 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromBorislav Petkov <bp@alien8.de>
Date2017-09-08 10:10 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<unmwq-48b-15@gated-at.bofh.it>
In reply to#1728628
On Fri, Sep 08, 2017 at 08:26:44AM +0200, Thomas Gleixner wrote:
> On Fri, 8 Sep 2017, Markus Trippelsdorf wrote:
> 
> CC+ Borislav. He might have access to such a beast

Can I have /proc/cpuinfo and dmesg pls, in order to see whether I have
something similar?

Private mail's fine too.

Thx.

-- 
Regards/Gruss,
    Boris.

Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1728713 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromBorislav Petkov <bp@alien8.de>
Date2017-09-08 11:20 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<unnC9-4LB-7@gated-at.bofh.it>
In reply to#1728671
On Fri, Sep 08, 2017 at 10:05:36AM +0200, Borislav Petkov wrote:
> On Fri, Sep 08, 2017 at 08:26:44AM +0200, Thomas Gleixner wrote:
> > On Fri, 8 Sep 2017, Markus Trippelsdorf wrote:
> > 
> > CC+ Borislav. He might have access to such a beast
> 
> Can I have /proc/cpuinfo and dmesg pls, in order to see whether I have
> something similar?
> 
> Private mail's fine too.

So I don't have exactly your model - mine is model 2, stepping 3 but I see
something strange too, in dmesg:

+blkid[981]: segfault at 8 ip 00007f8267a4a3fd sp 00007ffc77045de0 error 4 in ld-linux-x86-64.so.2[7f8267a3a000+22000]
+blkid[984]: segfault at 8 ip 00007fe6035583fd sp 00007ffe9456e5b0 error 4 in ld-linux-x86-64.so.2[7fe603548000+22000]
+blkid[987]: segfault at 8 ip 00007fc695f123fd sp 00007fff1838f5f0 error 4 in ld-linux-x86-64.so.2[7fc695f02000+22000]
+blkid[990]: segfault at 8 ip 00007fd2a8d7c3fd sp 00007ffe4f663870 error 4 in ld-linux-x86-64.so.2[7fd2a8d6c000+22000]
+blkid[993]: segfault at 8 ip 00007fdb865fb3fd sp 00007ffc2ff85050 error 4 in ld-linux-x86-64.so.2[7fdb865eb000+22000]
+blkid[996]: segfault at 8 ip 00007fc28b90a3fd sp 00007ffc7233e0e0 error 4 in ld-linux-x86-64.so.2[7fc28b8fa000+22000]
+blkid[999]: segfault at 8 ip 00007fb95934c3fd sp 00007ffed513d660 error 4 in ld-linux-x86-64.so.2[7fb95933c000+22000]
+blkid[1002]: segfault at 8 ip 00007fb5facc83fd sp 00007ffece190fe0 error 4 in ld-linux-x86-64.so.2[7fb5facb8000+22000]
+blkid[1005]: segfault at 8 ip 00007f113cc693fd sp 00007ffcf6bbd3f0 error 4 in ld-linux-x86-64.so.2[7f113cc59000+22000]
+blkid[1008]: segfault at 8 ip 00007f1ad0a593fd sp 00007ffeee162df0 error 4 in ld-linux-x86-64.so.2[7f1ad0a49000+22000]
+blkid[1011]: segfault at 8 ip 00007fdd003183fd sp 00007ffde1f69e60 error 4 in ld-linux-x86-64.so.2[7fdd00308000+22000]
+blkid[1014]: segfault at 8 ip 00007ffb3240c3fd sp 00007ffc75a88180 error 4 in ld-linux-x86-64.so.2[7ffb323fc000+22000]
+blkid[1017]: segfault at 8 ip 00007f88b6d683fd sp 00007ffef5dbe830 error 4 in ld-linux-x86-64.so.2[7f88b6d58000+22000]
+blkid[1020]: segfault at 8 ip 00007fec7760c3fd sp 00007ffc9dd05890 error 4 in ld-linux-x86-64.so.2[7fec775fc000+22000]
+blkid[1026]: segfault at 8 ip 00007f5a31ecc3fd sp 00007fffaf3604b0 error 4 in ld-linux-x86-64.so.2[7f5a31ebc000+22000]
+logsave[1027]: segfault at 8 ip 00007f237d2033fd sp 00007fff53933e60 error 4 in ld-linux-x86-64.so.2[7f237d1f3000+22000]

and then

git pull
...

Fast-forward
error: merge died of signal 7

Lemme try 4.13.

-- 
Regards/Gruss,
    Boris.

Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1728738 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromMarkus Trippelsdorf <markus@trippelsdorf.de>
Date2017-09-08 11:50 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<uno5d-4XX-29@gated-at.bofh.it>
In reply to#1728713
On 2017.09.08 at 11:16 +0200, Borislav Petkov wrote:
> On Fri, Sep 08, 2017 at 10:05:36AM +0200, Borislav Petkov wrote:
> > On Fri, Sep 08, 2017 at 08:26:44AM +0200, Thomas Gleixner wrote:
> > > On Fri, 8 Sep 2017, Markus Trippelsdorf wrote:
> > > 
> > > CC+ Borislav. He might have access to such a beast
> > 
> > Can I have /proc/cpuinfo and dmesg pls, in order to see whether I have
> > something similar?
> > 
> > Private mail's fine too.
> 
> So I don't have exactly your model - mine is model 2, stepping 3 but I see
> something strange too, in dmesg:

I'm pretty sure the bug is in the merged 'x86-mm-for-linus' branch:
Either Andy's "PCID optimized TLB flushing" (would be my guess) or
'encrypted memory' support by Tom Lendacky.

(Bisecting is hard, because sometimes I can compile stuff for over 15
minutes without hitting the bug. At other times the machine locks up
hard when starting X11 already.)

-- 
Markus

[toc] | [prev] | [next] | [standalone]


#1728753 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromIngo Molnar <mingo@kernel.org>
Date2017-09-08 12:40 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<unoRz-5F0-5@gated-at.bofh.it>
In reply to#1728738
* Markus Trippelsdorf <markus@trippelsdorf.de> wrote:

> On 2017.09.08 at 11:16 +0200, Borislav Petkov wrote:
> > On Fri, Sep 08, 2017 at 10:05:36AM +0200, Borislav Petkov wrote:
> > > On Fri, Sep 08, 2017 at 08:26:44AM +0200, Thomas Gleixner wrote:
> > > > On Fri, 8 Sep 2017, Markus Trippelsdorf wrote:
> > > > 
> > > > CC+ Borislav. He might have access to such a beast
> > > 
> > > Can I have /proc/cpuinfo and dmesg pls, in order to see whether I have
> > > something similar?
> > > 
> > > Private mail's fine too.
> > 
> > So I don't have exactly your model - mine is model 2, stepping 3 but I see
> > something strange too, in dmesg:
> 
> I'm pretty sure the bug is in the merged 'x86-mm-for-linus' branch:
> Either Andy's "PCID optimized TLB flushing" (would be my guess) or
> 'encrypted memory' support by Tom Lendacky.
> 
> (Bisecting is hard, because sometimes I can compile stuff for over 15
> minutes without hitting the bug. At other times the machine locks up
> hard when starting X11 already.)

Do you have the 72c0098d92ce fix?

Thanks,

	Ingo

[toc] | [prev] | [next] | [standalone]


#1728755 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromMarkus Trippelsdorf <markus@trippelsdorf.de>
Date2017-09-08 12:40 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<unoRA-5F0-13@gated-at.bofh.it>
In reply to#1728753
On 2017.09.08 at 12:35 +0200, Ingo Molnar wrote:
> 
> * Markus Trippelsdorf <markus@trippelsdorf.de> wrote:
> 
> > On 2017.09.08 at 11:16 +0200, Borislav Petkov wrote:
> > > On Fri, Sep 08, 2017 at 10:05:36AM +0200, Borislav Petkov wrote:
> > > > On Fri, Sep 08, 2017 at 08:26:44AM +0200, Thomas Gleixner wrote:
> > > > > On Fri, 8 Sep 2017, Markus Trippelsdorf wrote:
> > > > > 
> > > > > CC+ Borislav. He might have access to such a beast
> > > > 
> > > > Can I have /proc/cpuinfo and dmesg pls, in order to see whether I have
> > > > something similar?
> > > > 
> > > > Private mail's fine too.
> > > 
> > > So I don't have exactly your model - mine is model 2, stepping 3 but I see
> > > something strange too, in dmesg:
> > 
> > I'm pretty sure the bug is in the merged 'x86-mm-for-linus' branch:
> > Either Andy's "PCID optimized TLB flushing" (would be my guess) or
> > 'encrypted memory' support by Tom Lendacky.
> > 
> > (Bisecting is hard, because sometimes I can compile stuff for over 15
> > minutes without hitting the bug. At other times the machine locks up
> > hard when starting X11 already.)
> 
> Do you have the 72c0098d92ce fix?

Yes. The bug still happens on the current git tree (which has the fix
already):
 % git describe
v4.13-9217-g5969d1bb3082

-- 
Markus

[toc] | [prev] | [next] | [standalone]


#1728771 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromMarkus Trippelsdorf <markus@trippelsdorf.de>
Date2017-09-08 13:40 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<unpND-6r9-13@gated-at.bofh.it>
In reply to#1728755
On 2017.09.08 at 12:39 +0200, Markus Trippelsdorf wrote:
> On 2017.09.08 at 12:35 +0200, Ingo Molnar wrote:
> > 
> > * Markus Trippelsdorf <markus@trippelsdorf.de> wrote:
> > 
> > > On 2017.09.08 at 11:16 +0200, Borislav Petkov wrote:
> > > > On Fri, Sep 08, 2017 at 10:05:36AM +0200, Borislav Petkov wrote:
> > > > > On Fri, Sep 08, 2017 at 08:26:44AM +0200, Thomas Gleixner wrote:
> > > > > > On Fri, 8 Sep 2017, Markus Trippelsdorf wrote:
> > > > > > 
> > > > > > CC+ Borislav. He might have access to such a beast
> > > > > 
> > > > > Can I have /proc/cpuinfo and dmesg pls, in order to see whether I have
> > > > > something similar?
> > > > > 
> > > > > Private mail's fine too.
> > > > 
> > > > So I don't have exactly your model - mine is model 2, stepping 3 but I see
> > > > something strange too, in dmesg:
> > > 
> > > I'm pretty sure the bug is in the merged 'x86-mm-for-linus' branch:
> > > Either Andy's "PCID optimized TLB flushing" (would be my guess) or
> > > 'encrypted memory' support by Tom Lendacky.
> > > 
> > > (Bisecting is hard, because sometimes I can compile stuff for over 15
> > > minutes without hitting the bug. At other times the machine locks up
> > > hard when starting X11 already.)
> > 
> > Do you have the 72c0098d92ce fix?
> 
> Yes. The bug still happens on the current git tree (which has the fix
> already):

The bug is definitely caused by Andy Lutomirski's PCID optimized TLB
flushing" patches. Tom is off the hook.

-- 
Markus

[toc] | [prev] | [next] | [standalone]


#1729067

FromAndy Lutomirski <luto@kernel.org>
Date2017-09-08 18:20 +0200
Message-ID<unuaD-VM-41@gated-at.bofh.it>
In reply to#1728771
On Fri, Sep 8, 2017 at 4:30 AM, Markus Trippelsdorf
<markus@trippelsdorf.de> wrote:
> On 2017.09.08 at 12:39 +0200, Markus Trippelsdorf wrote:
>> On 2017.09.08 at 12:35 +0200, Ingo Molnar wrote:
>> >
>> > * Markus Trippelsdorf <markus@trippelsdorf.de> wrote:
>> >
>> > > On 2017.09.08 at 11:16 +0200, Borislav Petkov wrote:
>> > > > On Fri, Sep 08, 2017 at 10:05:36AM +0200, Borislav Petkov wrote:
>> > > > > On Fri, Sep 08, 2017 at 08:26:44AM +0200, Thomas Gleixner wrote:
>> > > > > > On Fri, 8 Sep 2017, Markus Trippelsdorf wrote:
>> > > > > >
>> > > > > > CC+ Borislav. He might have access to such a beast
>> > > > >
>> > > > > Can I have /proc/cpuinfo and dmesg pls, in order to see whether I have
>> > > > > something similar?
>> > > > >
>> > > > > Private mail's fine too.
>> > > >
>> > > > So I don't have exactly your model - mine is model 2, stepping 3 but I see
>> > > > something strange too, in dmesg:
>> > >
>> > > I'm pretty sure the bug is in the merged 'x86-mm-for-linus' branch:
>> > > Either Andy's "PCID optimized TLB flushing" (would be my guess) or
>> > > 'encrypted memory' support by Tom Lendacky.
>> > >
>> > > (Bisecting is hard, because sometimes I can compile stuff for over 15
>> > > minutes without hitting the bug. At other times the machine locks up
>> > > hard when starting X11 already.)
>> >
>> > Do you have the 72c0098d92ce fix?
>>
>> Yes. The bug still happens on the current git tree (which has the fix
>> already):
>
> The bug is definitely caused by Andy Lutomirski's PCID optimized TLB
> flushing" patches. Tom is off the hook.

I'm pretty sure it can't be PCID per se, since these CPUs are way too
old and are very unlikely to have PCID.

It could plausibly be the lazy TLB flushing changes.

[toc] | [prev] | [next] | [standalone]


#1729104 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromMarkus Trippelsdorf <markus@trippelsdorf.de>
Date2017-09-08 19:20 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<unv6G-1yj-21@gated-at.bofh.it>
In reply to#1729067
On 2017.09.08 at 09:12 -0700, Andy Lutomirski wrote:
> On Fri, Sep 8, 2017 at 4:30 AM, Markus Trippelsdorf
> <markus@trippelsdorf.de> wrote:
> > On 2017.09.08 at 12:39 +0200, Markus Trippelsdorf wrote:
> >> On 2017.09.08 at 12:35 +0200, Ingo Molnar wrote:
> >> >
> >> > * Markus Trippelsdorf <markus@trippelsdorf.de> wrote:
> >> >
> >> > > On 2017.09.08 at 11:16 +0200, Borislav Petkov wrote:
> >> > > > On Fri, Sep 08, 2017 at 10:05:36AM +0200, Borislav Petkov wrote:
> >> > > > > On Fri, Sep 08, 2017 at 08:26:44AM +0200, Thomas Gleixner wrote:
> >> > > > > > On Fri, 8 Sep 2017, Markus Trippelsdorf wrote:
> >> > > > > >
> >> > > > > > CC+ Borislav. He might have access to such a beast
> >> > > > >
> >> > > > > Can I have /proc/cpuinfo and dmesg pls, in order to see whether I have
> >> > > > > something similar?
> >> > > > >
> >> > > > > Private mail's fine too.
> >> > > >
> >> > > > So I don't have exactly your model - mine is model 2, stepping 3 but I see
> >> > > > something strange too, in dmesg:
> >> > >
> >> > > I'm pretty sure the bug is in the merged 'x86-mm-for-linus' branch:
> >> > > Either Andy's "PCID optimized TLB flushing" (would be my guess) or
> >> > > 'encrypted memory' support by Tom Lendacky.
> >> > >
> >> > > (Bisecting is hard, because sometimes I can compile stuff for over 15
> >> > > minutes without hitting the bug. At other times the machine locks up
> >> > > hard when starting X11 already.)
> >> >
> >> > Do you have the 72c0098d92ce fix?
> >>
> >> Yes. The bug still happens on the current git tree (which has the fix
> >> already):
> >
> > The bug is definitely caused by Andy Lutomirski's PCID optimized TLB
> > flushing" patches. Tom is off the hook.
> 
> I'm pretty sure it can't be PCID per se, since these CPUs are way too
> old and are very unlikely to have PCID.

Yes, the CPU doesn't support PCID (,but it does support PGE).

> It could plausibly be the lazy TLB flushing changes.

Yes, I've narrowed it down to:

commit 94b1b03b519b81c494900cb112aa00ed205cc2d9
Author: Andy Lutomirski <luto@kernel.org>
Date:   Thu Jun 29 08:53:17 2017 -0700

    x86/mm: Rework lazy TLB mode and TLB freshness tracking


Theoretically you guys should be able to reproduce the issue by using
the "nopcid" boot option. 

-- 
Markus

[toc] | [prev] | [next] | [standalone]


#1729310

FromAndy Lutomirski <luto@kernel.org>
Date2017-09-08 23:50 +0200
Message-ID<unzjZ-4iI-35@gated-at.bofh.it>
In reply to#1729104
On Fri, Sep 8, 2017 at 10:16 AM, Markus Trippelsdorf
<markus@trippelsdorf.de> wrote:
> On 2017.09.08 at 09:12 -0700, Andy Lutomirski wrote:
>> On Fri, Sep 8, 2017 at 4:30 AM, Markus Trippelsdorf
>> <markus@trippelsdorf.de> wrote:
>> > On 2017.09.08 at 12:39 +0200, Markus Trippelsdorf wrote:
>> >> On 2017.09.08 at 12:35 +0200, Ingo Molnar wrote:
>> >> >
>> >> > * Markus Trippelsdorf <markus@trippelsdorf.de> wrote:
>> >> >
>> >> > > On 2017.09.08 at 11:16 +0200, Borislav Petkov wrote:
>> >> > > > On Fri, Sep 08, 2017 at 10:05:36AM +0200, Borislav Petkov wrote:
>> >> > > > > On Fri, Sep 08, 2017 at 08:26:44AM +0200, Thomas Gleixner wrote:
>> >> > > > > > On Fri, 8 Sep 2017, Markus Trippelsdorf wrote:
>> >> > > > > >
>> >> > > > > > CC+ Borislav. He might have access to such a beast
>> >> > > > >
>> >> > > > > Can I have /proc/cpuinfo and dmesg pls, in order to see whether I have
>> >> > > > > something similar?
>> >> > > > >
>> >> > > > > Private mail's fine too.
>> >> > > >
>> >> > > > So I don't have exactly your model - mine is model 2, stepping 3 but I see
>> >> > > > something strange too, in dmesg:
>> >> > >
>> >> > > I'm pretty sure the bug is in the merged 'x86-mm-for-linus' branch:
>> >> > > Either Andy's "PCID optimized TLB flushing" (would be my guess) or
>> >> > > 'encrypted memory' support by Tom Lendacky.
>> >> > >
>> >> > > (Bisecting is hard, because sometimes I can compile stuff for over 15
>> >> > > minutes without hitting the bug. At other times the machine locks up
>> >> > > hard when starting X11 already.)
>> >> >
>> >> > Do you have the 72c0098d92ce fix?
>> >>
>> >> Yes. The bug still happens on the current git tree (which has the fix
>> >> already):
>> >
>> > The bug is definitely caused by Andy Lutomirski's PCID optimized TLB
>> > flushing" patches. Tom is off the hook.
>>
>> I'm pretty sure it can't be PCID per se, since these CPUs are way too
>> old and are very unlikely to have PCID.
>
> Yes, the CPU doesn't support PCID (,but it does support PGE).
>
>> It could plausibly be the lazy TLB flushing changes.
>
> Yes, I've narrowed it down to:
>
> commit 94b1b03b519b81c494900cb112aa00ed205cc2d9
> Author: Andy Lutomirski <luto@kernel.org>
> Date:   Thu Jun 29 08:53:17 2017 -0700
>
>     x86/mm: Rework lazy TLB mode and TLB freshness tracking
>
>
> Theoretically you guys should be able to reproduce the issue by using
> the "nopcid" boot option.
>

Any chance you could test with CONFIG_DEBUG_VM=y?  There are lots of
potentially useful assertions in that code.

Can you also post your /proc/cpuinfo?  And can you re-confirm that a
problematic guest kernel is causing problems in the *host*?

[toc] | [prev] | [next] | [standalone]


#1729323 — Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf

FromBorislav Petkov <bp@alien8.de>
Date2017-09-09 00:00 +0200
SubjectRe: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf
Message-ID<unztF-4m3-29@gated-at.bofh.it>
In reply to#1729310
On Fri, Sep 08, 2017 at 02:47:00PM -0700, Andy Lutomirski wrote:
> Any chance you could test with CONFIG_DEBUG_VM=y?  There are lots of
> potentially useful assertions in that code.
> 
> Can you also post your /proc/cpuinfo?  And can you re-confirm that a
> problematic guest kernel is causing problems in the *host*?

Also, have you seen any MCEs during early boot, after the freezes?

You probably wouldn't have because we don't log them on F10h due to
broken BIOSen. So add "mce=bootlog" to your grub and warm-reset your box
after one of those freezes and send me dmesg. It should have an MCE in
there, if it happens what I think it happens.

Thx.

-- 
Regards/Gruss,
    Boris.

Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1729341

FromAndy Lutomirski <luto@kernel.org>
Date2017-09-09 01:10 +0200
Message-ID<unAzo-5dI-9@gated-at.bofh.it>
In reply to#1729323
[Linus, I added you to get your opinion on whether the last bit here
is a problem.]

On Fri, Sep 8, 2017 at 2:56 PM, Borislav Petkov <bp@alien8.de> wrote:
> On Fri, Sep 08, 2017 at 02:47:00PM -0700, Andy Lutomirski wrote:
>> Any chance you could test with CONFIG_DEBUG_VM=y?  There are lots of
>> potentially useful assertions in that code.
>>
>> Can you also post your /proc/cpuinfo?  And can you re-confirm that a
>> problematic guest kernel is causing problems in the *host*?
>
> Also, have you seen any MCEs during early boot, after the freezes?
>
> You probably wouldn't have because we don't log them on F10h due to
> broken BIOSen. So add "mce=bootlog" to your grub and warm-reset your box
> after one of those freezes and send me dmesg. It should have an MCE in
> there, if it happens what I think it happens.
>

Here's my theory as to what's happening.

Before my patch, flush_tlb_mm_range() guaranteed that the range would
be flushed on all CPUs prior to returning.  With the patch, it only
promises that it will be flushed on all CPUs prior to anyone trying to
access it on the CPU in question.  This has two consequences:

1. A kernel thread that accidentally reads or writes a user address
could hit a stale TLB entry.  This seems harmless in the sense that
this can only happen if we already have a bug.

2. The CPU itself could see the TLB entry and do nefarious
architecturally invisible things with it.

I bet that #2 dramatically increases the chance that we hit erratum 383.

I can imagine a case where we have a problem even in the absence of an
erratum.  Specifically, suppose we have some page mapped.  CPU A
writes to it using combining (it's mapped WC or an explicit streaming
write is done).  CPU B removes the TLB entry and does
flush_tlb_mm_range().  CPU B would expect that all writes to the page
are done, but CPU A's write is still sitting in the streaming buffers.

I *think* this is impossible because CPU A's mm_cpumask manipulations
are atomic and should therefore force out the streaming write buffers,
but maybe there's some other scenario where this matters.

--Andy

[toc] | [prev] | [next] | [standalone]


#1729348

FromLinus Torvalds <torvalds@linux-foundation.org>
Date2017-09-09 01:30 +0200
Message-ID<unASK-5pn-5@gated-at.bofh.it>
In reply to#1729341
On Fri, Sep 8, 2017 at 4:07 PM, Andy Lutomirski <luto@kernel.org> wrote:
>
> I *think* this is impossible because CPU A's mm_cpumask manipulations
> are atomic and should therefore force out the streaming write buffers,
> but maybe there's some other scenario where this matters.

I don't think atomic memops do that.

They enforce globally visible ordering, but since they happen in the
cache and is not actually visible to outside, that doesn't actually
affect any streaming write buffers.

Then, if somebody else requests a cacheline that we have exclusive
ownership to, the write buffers just need to flush before we give up
that cacheline.

So a locked memory op is *not* serializing, it only enforces memory
ordering. Big difference.

Only fully serializing instructions will serialize with the write
buffers, and they are expensive as hell (partly exactly _due_ to these
kinds of issues).

So this change to delay invalidation does sound fairly scary..

              Linus

[toc] | [prev] | [next] | [standalone]


#1729358

FromAndy Lutomirski <luto@kernel.org>
Date2017-09-09 02:10 +0200
Message-ID<unBvr-5WK-1@gated-at.bofh.it>
In reply to#1729348
On Fri, Sep 8, 2017 at 4:23 PM, Linus Torvalds
<torvalds@linux-foundation.org> wrote:
> On Fri, Sep 8, 2017 at 4:07 PM, Andy Lutomirski <luto@kernel.org> wrote:
>>
>> I *think* this is impossible because CPU A's mm_cpumask manipulations
>> are atomic and should therefore force out the streaming write buffers,
>> but maybe there's some other scenario where this matters.
>
> I don't think atomic memops do that.
>
> They enforce globally visible ordering, but since they happen in the
> cache and is not actually visible to outside, that doesn't actually
> affect any streaming write buffers.
>
> Then, if somebody else requests a cacheline that we have exclusive
> ownership to, the write buffers just need to flush before we give up
> that cacheline.
>
> So a locked memory op is *not* serializing, it only enforces memory
> ordering. Big difference.
>
> Only fully serializing instructions will serialize with the write
> buffers, and they are expensive as hell (partly exactly _due_ to these
> kinds of issues).

I'm not convinced.  The SDM says (Vol 3, 11.3, under WC):

If the WC buffer is partially filled, the writes may be delayed until
the next occurrence of a serializing event; such as, an SFENCE or
MFENCE instruction, CPUID execution, a read or write to uncached
memory, an interrupt occurrence, or a LOCK instruction execution.

Thanks, Intel, for definiing "serializing event" differently here than
anywhere else in the whole manual.

Anyhow, I can think of two cases where this is relevant.

1. The kernel wants to reclaim a page of normal memory, so it unmaps
it and flushes.  Another CPU has an entry for that page in its WC
buffer.  I don't think we care whether the flush causes the WC write
to really hit RAM because it's unobservable -- we just need to make
sure it is ordered, as seen by software, before the flush operation
completes.  From the quote above, I think we're okay here.

2. The kernel is unmapping some IO memory (e.g. a GPU command buffer).
It wants a guarantee that, when flush_tlb_mm_range returns, all CPUs
are really done writing to it.  Here I'm less convinced.  The SDM
quote certainly suggests to me that we have a promise that the WC
write has *started* before flush_tlb_mm_range returns, but I'm not
sure I believe that it's guaranteed to have retired.  That being said,
I'm not sure that this is observable either -- anything the kernel
does that depends on the writes being done presumably has to involve
further IO to the same device, and I suspect that WC write; lock write
to memory; observe that write on another CPU; IO on other CPU really
does guarantee that everything hits the bus in order.

FWIW, I'm not sure that we ever had a guarantee that IO writes were
all fully done before flush_tlb_mm_range would return.  Can't they
still be hanging out in the PCI bridge or whatever until someone
*reads* that device?

>
> So this change to delay invalidation does sound fairly scary..

With PCID, we're fundamentally delaying invalidation.  I think the
worry is more that we're not guaranteeing that every CPU that could
have accessed the page being flushed has executed a real serializing
instruction.

If we want to force invalidation/serialization, I can see two
reasonable solutions:

1. Revert this behavior change: continue sending IPIs to lazy CPUs.
The problem is that this will totally wipe out the performance gain,
and that gain seemed to be substantial in some microbenchmarks at
least.

2. Get rid of lazy mode.  With PCID at least, switching to init_mm isn't so bad.

I'd prefer to leave it as is except on the buggy AMD CPUs, though,
since the current code is nice and fast.

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | linux.kernel


csiph-web