Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1180638
| Path | csiph.com!aioe.org!bofh.it!news.nic.it!robomod |
|---|---|
| From | Ian Kent <raven@themaw.net> |
| Newsgroups | linux.kernel |
| Subject | Re: [RFC] freeing unliked file indefinitely delayed |
| Date | Thu, 09 Jul 2015 13:30:02 +0200 |
| Message-ID | <pKibE-7YN-13@gated-at.bofh.it> (permalink) |
| References | <pJMEN-4OM-3@gated-at.bofh.it> <pKibD-7YN-1@gated-at.bofh.it> |
| X-Original-To | Al Viro <viro@ZenIV.linux.org.uk> |
| Dkim-Signature | v=1; a=rsa-sha1; c=relaxed/relaxed; d=themaw.net; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to:x-sasl-enc :x-sasl-enc; s=mesmtp; bh=BiisL/2WZJNgiTjNl8SKEDNFHbo=; b=ybcpeN QcuZTrnRwEYo2Nq9zzuPcDR6JdjU3LAhaM+OGau08QvyGfr/gE7Jj/8NKxlFE73X b0W46pUZL/ArdH8aajYTrjkTXwuWNqfPoVS1Gtg3/lWdlBO7zH6Y0zw4DKHIJdM1 fGP6Kk4mWwFYZ7+vPbYo7thZBv0C3WQXBDvG8= |
| Dkim-Signature | v=1; a=rsa-sha1; c=relaxed/relaxed; d= messagingengine.com; h=cc:content-transfer-encoding:content-type :date:from:in-reply-to:message-id:mime-version:references :subject:to:x-sasl-enc:x-sasl-enc; s=smtpout; bh=BiisL/2WZJNgiTj Nl8SKEDNFHbo=; b=VxPDsxMWqFIG3TKQsn3JCe1JUWRwkLbQlUqjxxKOtDYCuCH /OjAKxvI5S6hRM4a356y4gcygyeWBIoGTc6QUokslc7a2YW0RNTGZPT7JVUOhUOI V1UpXCagcF33YdGWiEx+B64lRf5NXkCCY52UhOUjKADfTKlLN3iTogMT+Ui4= |
| X-Sasl-Enc | qNwRVSGB+5KmWKp1lo+RBNR2Er3LLog/VrDkEYcWU6oN 1436441213 |
| Content-Type | text/plain; charset="UTF-8" |
| X-Mailer | Evolution 3.10.4 (3.10.4-4.fc20) |
| MIME-Version | 1.0 |
| Content-Transfer-Encoding | 7bit |
| Sender | robomod@news.nic.it |
| List-ID | <linux-kernel.vger.kernel.org> |
| X-Mailing-List | linux-kernel@vger.kernel.org |
| Approved | robomod@news.nic.it |
| Lines | 112 |
| Organization | linux.* mail to news gateway |
| X-Original-Cc | Linus Torvalds <torvalds@linux-foundation.org>, "J. Bruce Fields" <bfields@fieldses.org>, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org |
| X-Original-Date | Thu, 09 Jul 2015 19:26:44 +0800 |
| X-Original-Message-ID | <1436441204.2709.10.camel@pluto.fritz.box> |
| X-Original-References | <20150708014237.GC17109@ZenIV.linux.org.uk> <1436440655.2709.8.camel@pluto.fritz.box> |
| X-Original-Sender | linux-kernel-owner@vger.kernel.org |
| Xref | aioe.org linux.kernel:1180638 |
Show key headers only | View raw
On Thu, 2015-07-09 at 19:17 +0800, Ian Kent wrote:
> On Wed, 2015-07-08 at 02:42 +0100, Al Viro wrote:
> > Normally opening a file, unlinking it and then closing will have
> > the inode freed upon close() (provided that it's not otherwise busy and
> > has no remaining links, of course). However, there's one case where that
> > does *not* happen. Namely, if you open it by fhandle with cold dcache,
> > then unlink() and close().
> >
> > In normal case you get d_delete() in unlink(2) notice that dentry
> > is busy and unhash it; on the final dput() it will be forcibly evicted from
> > dcache, triggering iput() and inode removal. In this case, though, we end
> > up with *two* dentries - disconnected (created by open-by-fhandle) and
> > regular one (used by unlink()). The latter will have its reference to inode
> > dropped just fine, but the former will not - it's considered hashed (it
> > is on the ->s_anon list), so it will stay around until the memory pressure
> > will finally do it in. As the result, we have the final iput() delayed
> > indefinitely. It's trivial to reproduce -
> >
> > #define _GNU_SOURCE
> > #include <sys/types.h>
> > #include <sys/stat.h>
> > #include <fcntl.h>
> >
> > void flush_dcache(void)
> > {
> > system("mount -o remount,rw /");
> > }
> >
> > static char buf[20 * 1024 * 1024];
> >
> > main()
> > {
> > int fd;
> > union {
> > struct file_handle f;
> > char buf[MAX_HANDLE_SZ];
> > } x;
> > int m;
> >
> > x.f.handle_bytes = sizeof(x);
> > chdir("/root");
> > mkdir("foo", 0700);
> > fd = open("foo/bar", O_CREAT | O_RDWR, 0600);
> > close(fd);
> > name_to_handle_at(AT_FDCWD, "foo/bar", &x.f, &m, 0);
> > flush_dcache();
> > fd = open_by_handle_at(AT_FDCWD, &x.f, O_RDWR);
> > unlink("foo/bar");
> > write(fd, buf, sizeof(buf));
> > system("df ."); /* 20Mb eaten */
> > close(fd);
> > system("df ."); /* should've freed those 20Mb */
> > flush_dcache();
> > system("df ."); /* should be the same as #2 */
> > }
> >
> > will spit out something like
> > Filesystem 1K-blocks Used Available Use% Mounted on
> > /dev/root 322023 303843 1131 100% /
> > Filesystem 1K-blocks Used Available Use% Mounted on
> > /dev/root 322023 303843 1131 100% /
> > Filesystem 1K-blocks Used Available Use% Mounted on
> > /dev/root 322023 283282 21692 93% /
> > - inode gets freed only when dentry is finally evicted (here we trigger
> > than by remount; normally it would've happened in response to memory
> > pressure hell knows when).
> >
> > IMO it's a bug. Between the close() and final flush_dcache() the file has
> > no surviving links, is *not* busy, won't show up in fuser/lsof/whatnot
> > output, and yet it's still not freed. I'm not saying that it's realistically
> > exploitable (albeit with nfsd it might be), but it's a very unpleasant
> > self-LART.
> >
> > FWIW, my prefered fix would be simply to have the final dput() treat
> > disconnected dentries as "kill on sight"; checking for i_nlink of the
> > inode, as Bruce suggested several years ago, will *not* work, simply
> > because having another link to that file and unlinking it after close
> > will reproduce the situation - disconnected dentry sticks around in
> > dcache past its final dput() and past the last unlink() of our file.
> > Theoretically it might cause an overhead for nfsd (no_subtree_check v3
> > export might see more d_alloc()/d_free(); icache retention will still
> > prevent constant rereading the inode from disk). _IF_ that proves to
> > be noticable, we might need to come up with something more elaborate
> > (e.g. have unlink() and rename() kick disconnected aliases out if the link
> > count has reached 0), but it's more complex and needs careful ananlysis
> > of correctness - we need to prove that there's no way to miss the link
> > count reaching 0. I would prefer to treat all disconnected as unhashed
> > for dcache retention purposes - it's simpler and less brittle. Comments?
> > I mean something like this:
>
> Al, help me out here, I'm struggling to understand where these dentrys
> come from (apart from your reproducer).
>
> For example, on the heavily patched 2.6.32 kernel where this was first
> seen the problem dentry is annoymous, refcount 0, and unhashed.
>
> But the dentrys that will most likely face summary execution will be
> hashed, such as was the case on that 2.6.32 kernel at dput().
>
> Doesn't that mean that something dropped the dentry after the dput(),
> that will now also free the dentry, that took the refcount to 0?
Oh wait, think I get it now ... perhaps it's prune_one_dentry() doing
it ...
Ian
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
Back to linux.kernel | Previous | Next — Previous in thread | Find similar | Unroll thread
Re: [RFC] freeing unliked file indefinitely delayed Ian Kent <raven@themaw.net> - 2015-07-09 13:30 +0200 Re: [RFC] freeing unliked file indefinitely delayed Ian Kent <raven@themaw.net> - 2015-07-09 13:30 +0200
csiph-web