Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1259973

Re: timer code oops when calling mod_delayed_work

Path csiph.com!eternal-september.org!feeder.eternal-september.org!aioe.org!bofh.it!news.nic.it!robomod
From Jeff Layton <jlayton@poochiereds.net>
Newsgroups linux.kernel
Subject Re: timer code oops when calling mod_delayed_work
Date Sat, 31 Oct 2015 12:40:01 +0100
Message-ID <qpCFP-2fo-11@gated-at.bofh.it> (permalink)
References <qoWwW-1lC-23@gated-at.bofh.it> <qoZEt-3ix-1@gated-at.bofh.it> <qptMd-5iJ-1@gated-at.bofh.it>
X-Original-To Tejun Heo <tj@kernel.org>
Dkim-Signature v=1; a=rsa-sha256; c=relaxed/relaxed; d=poochiereds_net.20150623.gappssmtp.com; s=20150623; h=date:from:to:cc:subject:message-id:in-reply-to:references :mime-version:content-type:content-transfer-encoding; bh=uc02K6Ph59V4KLeBDqJ3bcZVtum3AcFHk28tDBl9ABs=; b=wmJwWAA4mc9HT0IA4MFXx7ZGk0ihBLck2amWIaXMylAAV7+Ah02qryVclSWCKWEJ6G MvZCHmxld1vKES9n7nAOI+X1wEBPAMp+FTXT45fgNFHBzfrmyEAlC6OfupCYmIlHPAB+ xcpKRGaLtmKwIdHk21EC8gDw3hvAf7EAjNupG71cTp/ZY1V1GfLCF5xxT1OeNUO+13Xy q4SlM4RZajLahQwTIfai+/frBJFzbuQBaHn1phRNy09U9ETvVvkUkN1AJUfbPdNIKNi3 GXUKOGwWjVRCptctmKnlvEJ+vTsbRabhVuYRKdBML5CLJ+e9WABkk5u8YTkiGu9Bz1cz lhaw==
X-Google-Dkim-Signature v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20130820; h=x-gm-message-state:date:from:to:cc:subject:message-id:in-reply-to :references:mime-version:content-type:content-transfer-encoding; bh=uc02K6Ph59V4KLeBDqJ3bcZVtum3AcFHk28tDBl9ABs=; b=Z7gFkIxpg51i9Vl/1fljMoqnuhyjls4Mawj2XoobiWb2yB+1RIZTbwuebEzkh/iGKg IWi+AUHDgVw6KnG2V2wN5Q1mWNs4bZ+IcTk9bujOzP8U/OdJbhBHxo6eRaVOURSiJIuP boRmPWWE8Xessa4LzGIkKSv1v+nzv9s+9mrAOAWbSO7cMOxH851a2MlCYgJWgnIvN+oh F9ZyAmG7e5Lsd3At/euMiW/VP4gvYGENl8wIpPeDfTChOgLiCTTtJeQCZlBdKP84ZGzC 82qSELwBjFYl3LecJr4Wy7pITYbfOHP9vi6V0S/LO8Fva9XpANC4XLRVYRfAEN4GTFtn nhIg==
X-Gm-Message-State ALoCoQnmf7UVLrZhGbNoLsv0KYRjsjpOx7Znwj4aTf9WHlZXRjDHM1eRN5+dhHZEln5LyETT1Rl4
X-Received by 10.140.223.17 with SMTP id t17mr17370364qhb.77.1446291244118; Sat, 31 Oct 2015 04:34:04 -0700 (PDT)
X-Mailer Claws Mail 3.12.0 (GTK+ 2.24.28; x86_64-redhat-linux-gnu)
MIME-Version 1.0
Content-Type text/plain; charset=US-ASCII
Content-Transfer-Encoding 7bit
Sender robomod@news.nic.it
List-ID <linux-kernel.vger.kernel.org>
X-Mailing-List linux-kernel@vger.kernel.org
Approved robomod@news.nic.it
Lines 116
Organization linux.* mail to news gateway
X-Original-Cc linux-kernel@vger.kernel.org, bfields@fieldses.org, Michael Skralivetsky <michael.skralivetsky@primarydata.com>, Chris Worley <chris.worley@primarydata.com>, Trond Myklebust <trond.myklebust@primarydata.com>, Lai Jiangshan <laijs@cn.fujitsu.com>
X-Original-Date Sat, 31 Oct 2015 07:34:00 -0400
X-Original-Message-ID <20151031073400.2cf05d77@tlielax.poochiereds.net>
X-Original-References <20151029103113.2f893924@tlielax.poochiereds.net> <20151029135836.02ad9000@synchrony.poochiereds.net> <20151031020012.GH3582@mtj.duckdns.org>
X-Original-Sender linux-kernel-owner@vger.kernel.org
Xref csiph.com linux.kernel:1259973

Show key headers only | View raw


On Sat, 31 Oct 2015 11:00:12 +0900
Tejun Heo <tj@kernel.org> wrote:

> (cc'ing Lai)
> 
> Hello, Jeff.
> 
> On Thu, Oct 29, 2015 at 01:58:36PM -0400, Jeff Layton wrote:
> > crash> p cache_cleaner
> > cache_cleaner = $12 = {
> >   work = {
> >     data = {
> >       counter = 0xfffffffe1
> 
> If I'm not mistaken, PENDING, flush color 14, OFFQ and POOL_NONE.
> 
> >     }, 
> >     entry = {
> >       next = 0xffffffffa03623c8 <cache_cleaner+8>, 
> >       prev = 0xffffffffa03623c8 <cache_cleaner+8>
> 
> Empty entry.
> 
> >     }, 
> >     func = 0xffffffffa03333c0 <cache_cleaner_func>
> >   }, 
> >   timer = {
> >     entry = {
> >       next = 0x0, 
> >       pprev = 0xffff88085fd0eaf8
> >     }, 
> >     expires = 0x100021e99, 
> >     function = 0xffffffff810b66a0 <delayed_work_timer_fn>, 
> >     data = 0xffffffffa03623c0, 
> >     flags = 0x200014, 
> >     slack = 0xffffffff, 
> >     start_pid = 0x0, 
> >     start_site = 0x0, 
> >     start_comm = "\000\000\000\000\000\000\000\000\000\000\000\000\000\000\000"
> >   }, 
> >   wq = 0xffff88085f48fc00, 
> >   cpu = 0x14
> > }
> > 
> > So the PENDING bit is set (lowest bit in data.counter), and timer->entry.pprev
> > pprev pointer is not NULL (so timer_pending is true). I also see that
> > there are several nfsd threads running the shrinker at the same time.
> > 
> > There is one potential problem that I see, but I'd appreciate someone
> > sanity checking me on this. Here is mod_delayed_work_on:
> ...
> > ...and here is the beginning of try_to_grab_pending:
> > 
> > ------------------[snip]------------------------
> >         /* try to steal the timer if it exists */
> >         if (is_dwork) {
> >                 struct delayed_work *dwork = to_delayed_work(work);
> > 
> >                 /*
> >                  * dwork->timer is irqsafe.  If del_timer() fails, it's
> >                  * guaranteed that the timer is not queued anywhere and not
> >                  * running on the local CPU.
> >                  */
> >                 if (likely(del_timer(&dwork->timer)))
> >                         return 1;
> >         }
> > 
> >         /* try to claim PENDING the normal way */
> >         if (!test_and_set_bit(WORK_STRUCT_PENDING_BIT, work_data_bits(work)))
> >                 return 0;
> > ------------------[snip]------------------------
> > 
> > 
> > ...so if del_timer returns true, we'll return 1 from
> > try_to_grab_pending without actually setting the
> > WORK_STRUCT_PENDING_BIT, and then will end up calling
> > __queue_delayed_work.
> > 
> > That seems wrong to me -- shouldn't we be ensuring that that bit is set
> > when returning 1 from try_to_grab_pending to guard against concurrent
> > queue_delayed_work_on calls?
> 
> But if try_to_grab_pending() succeeded at stealing dwork->timer, it's
> known that the PENDING bit must already be set.  IOW, the bit is
> stolen together with the timer.
> 
> Heh, this one is tricky.  Yeah, try_to_grab_pending() missing PENDING
> would explain the failure but I can't see how it'd leak at the moment.
> 

Thanks Tejun. Yeah, I realized that after sending the response above.

If you successfully delete the timer the timer then the PENDING bit
should already be set. Might be worth throwing in something like this,
just before the return 1:

    WARN_ON(!test_bit(WORK_STRUCT_PENDING_BIT, work_data_bits(work)))

...but I doubt it would fire. I think it's likely that the bug is
elsewhere.

The other thing is that we've had this code in place for a couple of
years now, and this is the first time I've seen an oops like this. I
suspect that this may be a recent regression, but I don't know that for
sure.

I have asked Chris and Michael to see if they can bisect it down, but
it may be a bit before they can get that done. Any insight you might
have in the meantime would helpful.
-- 
Jeff Layton <jlayton@poochiereds.net>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

Back to linux.kernel | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread


Thread

timer code oops when calling mod_delayed_work Jeff Layton <jlayton@poochiereds.net> - 2015-10-29 15:40 +0100
  Re: timer code oops when calling mod_delayed_work Jeff Layton <jlayton@poochiereds.net> - 2015-10-29 19:00 +0100
    Re: timer code oops when calling mod_delayed_work Tejun Heo <tj@kernel.org> - 2015-10-31 03:10 +0100
      Re: timer code oops when calling mod_delayed_work Jeff Layton <jlayton@poochiereds.net> - 2015-10-31 12:40 +0100
        Re: timer code oops when calling mod_delayed_work Tejun Heo <tj@kernel.org> - 2015-10-31 22:40 +0100
          Re: timer code oops when calling mod_delayed_work Jeff Layton <jlayton@poochiereds.net> - 2015-10-31 23:00 +0100
            Re: timer code oops when calling mod_delayed_work Chris Worley <chris.worley@primarydata.com> - 2015-11-02 20:50 +0100
              Re: timer code oops when calling mod_delayed_work Jeff Layton <jlayton@poochiereds.net> - 2015-11-02 21:00 +0100
                Re: timer code oops when calling mod_delayed_work Jeff Layton <jlayton@poochiereds.net> - 2015-11-03 02:40 +0100
                Re: timer code oops when calling mod_delayed_work Jeff Layton <jlayton@poochiereds.net> - 2015-11-03 19:00 +0100
                Re: timer code oops when calling mod_delayed_work Tejun Heo <tj@kernel.org> - 2015-11-04 00:00 +0100
                Re: timer code oops when calling mod_delayed_work Tejun Heo <tj@kernel.org> - 2015-11-04 01:10 +0100
                Re: timer code oops when calling mod_delayed_work Jeff Layton <jlayton@poochiereds.net> - 2015-11-04 12:50 +0100
                [PATCH] timer: add_timer_on() should perform proper migration Tejun Heo <tj@kernel.org> - 2015-11-04 18:20 +0100
                [tip:timers/urgent] timers:   Use proper base migration in add_timer_on() tip-bot for Tejun Heo <tipbot@zytor.com> - 2015-11-04 20:30 +0100
                Re: [PATCH] timer: add_timer_on() should perform proper migration Thomas Gleixner <tglx@linutronix.de> - 2015-11-04 20:40 +0100
                Re: [PATCH] timer: add_timer_on() should perform proper migration Tejun Heo <tj@kernel.org> - 2015-11-04 20:50 +0100

csiph-web