Path: csiph.com!weretis.net!feeder7.news.weretis.net!usenet.goja.nl.eu.org!aioe.org!bofh.it!news.nic.it!robomod From: Jann Horn Newsgroups: linux.debian.bugs.dist,linux.debian.kernel Subject: Bug#963493: Repeatable hard lockup running strace testsuite on 4.19.98+ onwards Date: Fri, 26 Jun 2020 20:40:03 +0200 Message-ID: References: X-Original-To: Steve McIntyre X-Mailbox-Line: From debian-bugs-dist-request@lists.debian.org Fri Jun 26 18:30:15 2020 Old-Return-Path: X-Spam-Flag: NO X-Spam-Score: 0.951 Reply-To: Jann Horn , 963493@bugs.debian.org Resent-To: debian-bugs-dist@lists.debian.org Resent-Cc: Debian Kernel Team X-Debian-Pr-Message: followup 963493 X-Debian-Pr-Package: src:linux X-Debian-Pr-Keywords: upstream X-Debian-Pr-Source: linux X-Spam-Bayes: score:0.0000 Tokens: new, 21; hammy, 149; neutral, 190; spammy, 1. spammytokens:0.993-1--john's hammytokens:0.000-+--UD:kernel.org, 0.000-+--sk:config_, 0.000-+--sk:CONFIG_, 0.000-+--wwwkernelorg, 0.000-+--UD:www.kernel.org Dkim-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20161025; h=mime-version:references:in-reply-to:from:date:message-id:subject:to :cc; bh=aW6KHpbtTRJJscuFSXZMdCFqV8eNv1XYkHhDBiBGbfE=; b=tjES2mfxkQ1y0nbq1Vs43yjDZVPu8njkTXVRAnhkeWzbkoYPRId9PQ5Nes/wNXpPiT otd8lynMQ+Lq1GrqRJbRTJ7Q2cqcRf/+KvfzGZ6ZTzKN6RPy57/PPBSvsmyZL8rJRxXu GYPe2kIPNF7ofdGf93Y9nSqBpCP1+Ksr3bVrK1pG/cG5+W+0BCJtbXy0Lo99DkX/UPqG 999J/GVTtf11HyjaMC+sDlwB4kivQFaOqetpMtPwj+l+Y3rKIKnE4WDruaX6yW7If2Jf hzdQFXLjIFBHayHuMDllrqsQdYlFUMpU0MlK7x1q7FKzswVUv0fqrGA9WN6F3d9E3v1V EYOg== X-Google-Dkim-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:mime-version:references:in-reply-to:from:date :message-id:subject:to:cc; bh=aW6KHpbtTRJJscuFSXZMdCFqV8eNv1XYkHhDBiBGbfE=; b=MYXnCEtPrSAqk9NXhiAf4yvFkdfrlOe55cK1DVLCfyh0JpRmbGG/zn+HYoOxZwFqEK y4EzGdSoYKjz5zCc71D1sB4S+0KQZG+COmtes25YcOpqKVBvJwYYx9mID+rK71fm/ZmB MFgSc9MucIF8aIGVrGLYkZgT/kKm+ZhQVrQXGgSAIsL156FzNF/BGZDnslTzPX9rs4sb zV5f49WEgT4WHfuSvnUHdCiWzrTWUjhDSmp9KYm8yQyJqesPzYk/SWtfevTJRmleBwfY VflV2Cq4gOTXR8Xkdd9qwml7k1T4hEFGcc7EvlJlzMV2OsWn9bnDAplJEY8QgmNxAPd3 AWmw== X-Gm-Message-State: AOAM530qOt7XZDijbd3cbEsUSmFh9G1ZhnQjTrqgiV3s4VKaRKzwfPvq U+bpZ+SpSk/Yal86AQ9D4zteDly7//wfaNJVf1P9bQ== X-Google-SMTP-Source: ABdhPJz+LIlK17Mz0UHfymuwSXh5M0ZiiXytdF0cyCwXgOe19o+BAJxPFbx/FwgooFmtoX3Sd+rLCOuZRIc1Tq+b+k4= X-Received: by 2002:a19:8b8a:: with SMTP id n132mr2528578lfd.45.1593196068639; Fri, 26 Jun 2020 11:27:48 -0700 (PDT) MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" X-Debian-Message: from BTS X-Mailing-List: archive/latest/1610484 List-ID: List-URL: Approved: robomod@news.nic.it Lines: 68 Organization: linux.* mail to news gateway Sender: robomod@news.nic.it X-Original-Cc: Greg KH , John Johansen , Sasha Levin , stable , 963493@bugs.debian.org X-Original-Date: Fri, 26 Jun 2020 20:27:22 +0200 X-Original-Message-ID: X-Original-References: <20200626113558.GA32542@unset.einval.com> <20200626134132.GB4024297@kroah.com> <20200626165000.GB2950@unset.einval.com> <20200626175201.GA9110@unset.einval.com> <159282711531.24396.10297981990474437360.reportbug@tack.local> <20200626175201.GA9110@unset.einval.com> Xref: csiph.com linux.debian.bugs.dist:1015410 linux.debian.kernel:67408 On Fri, Jun 26, 2020 at 7:52 PM Steve McIntyre wrote: > On Fri, Jun 26, 2020 at 05:50:00PM +0100, Steve McIntyre wrote: > >On Fri, Jun 26, 2020 at 04:25:59PM +0200, Jann Horn wrote: > >>On Fri, Jun 26, 2020 at 3:41 PM Greg KH wrote: > >>> On Fri, Jun 26, 2020 at 12:35:58PM +0100, Steve McIntyre wrote: > > > >... > > > >>> > Considering I'm running strace build tests to provoke this bug, > >>> > finding the failure in a commit talking about ptrace changes does look > >>> > very suspicious...! > >>> > > >>> > Annoyingly, I can't reproduce this on my disparate other machines > >>> > here, suggesting it's maybe(?) timing related. > >> > >>Does "hard lockup" mean that the HARDLOCKUP_DETECTOR infrastructure > >>prints a warning to dmesg? If so, can you share that warning? > > > >I mean the machine locks hard - X stops updating, the mouse/keyboard > >stop responding. No pings, etc. When I reboot, there's nothing in the > >logs. > > > >>If you don't have any way to see console output, and you don't have a > >>working serial console setup or such, you may want to try re-running > >>those tests while the kernel is booted with netconsole enabled to log > >>to a different machine over UDP (see > >>https://www.kernel.org/doc/Documentation/networking/netconsole.txt). > > > >ACK, will try that now for you. > > > >>You may want to try setting the sysctl kernel.sysrq=1 , then when the > >>system has locked up, press ALT+PRINT+L (to generate stack traces for > >>all active CPUs from NMI context), and maybe also ALT+PRINT+T and > >>ALT+PRINT+W (to collect more information about active tasks). > > > >Nod. > > > >>(If you share stack traces from these things with us, it would be > >>helpful if you could run them through scripts/decode_stacktrace.pl > >>from the kernel tree first, to add line number information.) > > > >ACK. > > Output passed through scripts/decode_stacktrace.sh attached. > > Just about to try John's suggestion next. Okay, so this is some sort of deadlock... Looking at the NMI backtraces, all the CPUs are blocked on spinlocks: CPU 3 is blocked on current->sighand->siglock, in tty_open_proc_set_tty() CPU 1 is blocked on... I'm not sure which lock, somewhere in do_wait() CPU 2 is blocked on something, somewhere in ptrace_stop() CPU 0 is stuck on a lock in do_exit() So I think it's probably something like a classic deadlock, or a sleeping-in-atomic issue, or a lock-balancing issue (or memory corruption, that can cause all kinds of weird errors)? If it really is a classic deadlock, CONFIG_PROVE_LOCKING=y should be able to pinpoint the issue. If it is a sleeping-in-atomic issue, CONFIG_DEBUG_ATOMIC_SLEEP=y should help. If it is memory corruption, CONFIG_KASAN=y should discover it... but that might majorly mess up the timing, so if this really is a race, that might not work. Maybe flip all of those on, and if it doesn't reproduce anymore, turn off CONFIG_KASAN and try again?