Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1271424 > unrolled thread

[PATCH v3 0/7] User namespace mount updates

Started bySeth Forshee <seth.forshee@canonical.com>
First post2015-11-17 17:50 +0100
Last post2015-11-18 19:50 +0100
Articles 20 on this page of 50 — 15 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-17 17:50 +0100
    [PATCH v3 1/7] block_dev: Support checking inode permissions in lookup_bdev() Seth Forshee <seth.forshee@canonical.com> - 2015-11-17 17:50 +0100
    [PATCH v3 2/7] block_dev: Check permissions towards block device inode when mounting Seth Forshee <seth.forshee@canonical.com> - 2015-11-17 17:50 +0100
    [PATCH v3 3/7] mtd: Check permissions towards mtd block device inode when mounting Seth Forshee <seth.forshee@canonical.com> - 2015-11-17 17:50 +0100
    Re: [PATCH v3 0/7] User namespace mount updates Al Viro <viro@ZenIV.linux.org.uk> - 2015-11-17 18:10 +0100
      Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-17 18:30 +0100
        Re: [PATCH v3 0/7] User namespace mount updates "Serge E. Hallyn" <serge@hallyn.com> - 2015-11-17 18:50 +0100
        Re: [PATCH v3 0/7] User namespace mount updates Al Viro <viro@ZenIV.linux.org.uk> - 2015-11-17 19:00 +0100
          Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-17 19:40 +0100
            Re: [PATCH v3 0/7] User namespace mount updates Richard Weinberger <richard.weinberger@gmail.com> - 2015-11-17 20:20 +0100
              Re: [PATCH v3 0/7] User namespace mount updates Octavian Purdila <octavian.purdila@intel.com> - 2015-11-17 20:30 +0100
                Re: [PATCH v3 0/7] User namespace mount updates Richard Weinberger <richard@nod.at> - 2015-11-17 21:20 +0100
                  Re: [PATCH v3 0/7] User namespace mount updates Octavian Purdila <octavian.purdila@intel.com> - 2015-11-17 23:10 +0100
                    Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-19 16:30 +0100
                      Re: [PATCH v3 0/7] User namespace mount updates Octavian Purdila <octavian.purdila@intel.com> - 2015-11-19 17:20 +0100
                        Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-19 17:40 +0100
                        Re: [PATCH v3 0/7] User namespace mount updates "Serge E. Hallyn" <serge.hallyn@ubuntu.com> - 2015-11-20 18:40 +0100
              Re: [PATCH v3 0/7] User namespace mount updates Richard Weinberger <richard@nod.at> - 2015-11-17 20:30 +0100
              Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-17 20:30 +0100
            Re: [PATCH v3 0/7] User namespace mount updates Theodore Ts'o <tytso@mit.edu> - 2015-11-18 20:20 +0100
              Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-18 20:30 +0100
              Re: [PATCH v3 0/7] User namespace mount updates Serge Hallyn <serge.hallyn@ubuntu.com> - 2015-11-18 20:40 +0100
          Re: [PATCH v3 0/7] User namespace mount updates Austin S Hemmelgarn <ahferroin7@gmail.com> - 2015-11-17 20:10 +0100
            Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-17 20:20 +0100
              Re: [PATCH v3 0/7] User namespace mount updates Austin S Hemmelgarn <ahferroin7@gmail.com> - 2015-11-17 22:00 +0100
                Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-17 22:40 +0100
                  Re: [PATCH v3 0/7] User namespace mount updates Austin S Hemmelgarn <ahferroin7@gmail.com> - 2015-11-18 13:30 +0100
                    Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-18 15:30 +0100
                      Re: [PATCH v3 0/7] User namespace mount updates Al Viro <viro@ZenIV.linux.org.uk> - 2015-11-18 16:00 +0100
                        Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-18 16:10 +0100
                          Re: [PATCH v3 0/7] User namespace mount updates Al Viro <viro@ZenIV.linux.org.uk> - 2015-11-18 16:20 +0100
                            Re: [PATCH v3 0/7] User namespace mount updates Richard Weinberger <richard.weinberger@gmail.com> - 2015-11-18 16:30 +0100
                              Re: [PATCH v3 0/7] User namespace mount updates James Morris <jmorris@namei.org> - 2015-11-19 08:50 +0100
                                Re: [PATCH v3 0/7] User namespace mount updates Richard Weinberger <richard@nod.at> - 2015-11-19 09:00 +0100
                                  Re: [PATCH v3 0/7] User namespace mount updates "Serge E. Hallyn" <serge.hallyn@ubuntu.com> - 2015-11-19 15:30 +0100
                                    Re: [PATCH v3 0/7] User namespace mount updates Richard Weinberger <richard@nod.at> - 2015-11-19 16:10 +0100
                                  Re: [PATCH v3 0/7] User namespace mount updates Colin Walters <walters@verbum.org> - 2015-11-19 15:40 +0100
                                    Re: [PATCH v3 0/7] User namespace mount updates Richard Weinberger <richard@nod.at> - 2015-11-19 15:50 +0100
                                      Re: [PATCH v3 0/7] User namespace mount updates "Richard W.M. Jones" <rjones@redhat.com> - 2015-11-19 16:20 +0100
                            Re: [PATCH v3 0/7] User namespace mount updates "Serge E. Hallyn" <serge.hallyn@ubuntu.com> - 2015-11-19 16:00 +0100
                        Re: [PATCH v3 0/7] User namespace mount updates Nikolay Borisov <kernel@kyup.com> - 2015-11-18 16:40 +0100
                        Re: [PATCH v3 0/7] User namespace mount updates Austin S Hemmelgarn <ahferroin7@gmail.com> - 2015-11-18 16:40 +0100
            Re: [PATCH v3 0/7] User namespace mount updates Al Viro <viro@ZenIV.linux.org.uk> - 2015-11-17 20:40 +0100
              Re: [PATCH v3 0/7] User namespace mount updates Austin S Hemmelgarn <ahferroin7@gmail.com> - 2015-11-17 21:40 +0100
                Re: [PATCH v3 0/7] User namespace mount updates Al Viro <viro@ZenIV.linux.org.uk> - 2015-11-17 22:10 +0100
                  Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-17 23:10 +0100
                    Re: [PATCH v3 0/7] User namespace mount updates Austin S Hemmelgarn <ahferroin7@gmail.com> - 2015-11-18 13:50 +0100
                      Re: [PATCH v3 0/7] User namespace mount updates Seth Forshee <seth.forshee@canonical.com> - 2015-11-18 15:40 +0100
                        Re: [PATCH v3 0/7] User namespace mount updates Austin S Hemmelgarn <ahferroin7@gmail.com> - 2015-11-18 16:40 +0100
              Re: [PATCH v3 0/7] User namespace mount updates bfields@fieldses.org (J. Bruce Fields) - 2015-11-18 19:50 +0100

Page 2 of 3 — ← Prev page 1 [2] 3  Next page →


#1272497

FromSeth Forshee <seth.forshee@canonical.com>
Date2015-11-18 20:30 +0100
Message-ID<qwgAy-7es-15@gated-at.bofh.it>
In reply to#1272492
On Wed, Nov 18, 2015 at 02:10:45PM -0500, Theodore Ts'o wrote:
> On Tue, Nov 17, 2015 at 12:34:44PM -0600, Seth Forshee wrote:
> > On Tue, Nov 17, 2015 at 05:55:06PM +0000, Al Viro wrote:
> > > On Tue, Nov 17, 2015 at 11:25:51AM -0600, Seth Forshee wrote:
> > > 
> > > > Shortly after that I plan to follow with support for ext4. I've been
> > > > fuzzing ext4 for a while now and it has held up well, and I'm currently
> > > > working on hand-crafted attacks. Ted has commented privately (to others,
> > > > not to me personally) that he will fix bugs for such attacks, though I
> > > > haven't seen any public comments to that effect.
> > > 
> > > _Static_ attacks, or change-image-under-mounted-fs attacks?
> > 
> > Right now only static attacks, change-image-under-mounted-fs attacks
> > will be next.
> 
> I will fix bugs about static attacks.  That is, it's interesting to me
> that a buggy file system (no matter how it is created), not cause the
> kernel to crash --- and privilege escalation attacks tend to be
> strongly related to those bugs where we're not doing strong enough
> checking.
> 
> Protecting against a malicious user which changes the image under the
> file system is a whole other kettle of fish.  I am not at all user you
> can do this without completely sacrificing performance or making the
> code impossible to maintain.  So my comments do *not* extend to
> protecting against a malicious user who is changing the block device
> underneath the kernel.
> 
> If you want to submit patches to make the kernel more robust against
> these attacks, I'm certainly willing to look at the patches.  But I'm
> certainly not guaranteeing that they will go in, and I'm certainly not
> promising to fix all vulnerabilities that you might find that are
> caused by a malicious block device.  Sorry, that's too much buying a
> pig in a poke....

Thanks Ted. My plan right now is to explore the possibility of blocking
writes to the backing store from userspace while it's mounted.

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1272506

FromSerge Hallyn <serge.hallyn@ubuntu.com>
Date2015-11-18 20:40 +0100
Message-ID<qwgKe-7hK-19@gated-at.bofh.it>
In reply to#1272492
Quoting Theodore Ts'o (tytso@mit.edu):
> On Tue, Nov 17, 2015 at 12:34:44PM -0600, Seth Forshee wrote:
> > On Tue, Nov 17, 2015 at 05:55:06PM +0000, Al Viro wrote:
> > > On Tue, Nov 17, 2015 at 11:25:51AM -0600, Seth Forshee wrote:
> > > 
> > > > Shortly after that I plan to follow with support for ext4. I've been
> > > > fuzzing ext4 for a while now and it has held up well, and I'm currently
> > > > working on hand-crafted attacks. Ted has commented privately (to others,
> > > > not to me personally) that he will fix bugs for such attacks, though I
> > > > haven't seen any public comments to that effect.
> > > 
> > > _Static_ attacks, or change-image-under-mounted-fs attacks?
> > 
> > Right now only static attacks, change-image-under-mounted-fs attacks
> > will be next.
> 
> I will fix bugs about static attacks.  That is, it's interesting to me
> that a buggy file system (no matter how it is created), not cause the
> kernel to crash --- and privilege escalation attacks tend to be
> strongly related to those bugs where we're not doing strong enough
> checking.
> 
> Protecting against a malicious user which changes the image under the
> file system is a whole other kettle of fish.  I am not at all user you
> can do this without completely sacrificing performance or making the
> code impossible to maintain.  So my comments do *not* extend to
> protecting against a malicious user who is changing the block device
> underneath the kernel.

Yup, thanks, Ted.  I think the only sane thing to do is work on making the
mounted files immutable.  Guarding against under-mounted-writes seems
crazy.  Well, actually it seems like a fascinating problem, and maybe
solvable without fs changes, but not in scope here.

> If you want to submit patches to make the kernel more robust against
> these attacks, I'm certainly willing to look at the patches.  But I'm
> certainly not guaranteeing that they will go in, and I'm certainly not
> promising to fix all vulnerabilities that you might find that are
> caused by a malicious block device.  Sorry, that's too much buying a
> pig in a poke....
> 
> 						- Ted
> 				
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1271577

FromAustin S Hemmelgarn <ahferroin7@gmail.com>
Date2015-11-17 20:10 +0100
Message-ID<qvTND-EM-9@gated-at.bofh.it>
In reply to#1271501

[Multipart message — attachments visible in raw view] — view raw

On 2015-11-17 12:55, Al Viro wrote:
> On Tue, Nov 17, 2015 at 11:25:51AM -0600, Seth Forshee wrote:
>
>> Shortly after that I plan to follow with support for ext4. I've been
>> fuzzing ext4 for a while now and it has held up well, and I'm currently
>> working on hand-crafted attacks. Ted has commented privately (to others,
>> not to me personally) that he will fix bugs for such attacks, though I
>> haven't seen any public comments to that effect.
>
> _Static_ attacks, or change-image-under-mounted-fs attacks?
To properly protect against attacks on mounted filesystems, we'd need 
some new concept of a userspace immutable file (that is, one where 
nobody can write to it except the kernel, and only the kernel can change 
it between regular access and this new state), and then have the kernel 
set an image (or block device) to this state when a filesystem is 
mounted from it (this introduces all kinds of other issues too however, 
for example stuff that allows an online fsck on the device will stop 
working, as will many un-deletion tools).

The only other option would be to force the FS to cache all metadata in 
memory, and validate between the cache and what's on disk on every 
access, which is not realistic for any real world system.

It's unfeasible from a practical standpoint to expect filesystems to 
assume that stuff they write might change under them due to malicious 
intent of a third party.  Some filesystems may be more resilient to this 
kind of attack (ZFS and BTRFS in some configurations come to mind), but 
a determined attacker can still circumvent those protections (on at 
least BTRFS, it's not all that hard to cause a sub-tree of the 
filesystem to disappear with at most two 64k blocks being written to the 
block device directly, and there is no way that this can be prevented 
short of what I suggest above).

[toc] | [prev] | [next] | [standalone]


#1271586

FromSeth Forshee <seth.forshee@canonical.com>
Date2015-11-17 20:20 +0100
Message-ID<qvTXk-J9-19@gated-at.bofh.it>
In reply to#1271577
On Tue, Nov 17, 2015 at 02:02:09PM -0500, Austin S Hemmelgarn wrote:
> On 2015-11-17 12:55, Al Viro wrote:
> >On Tue, Nov 17, 2015 at 11:25:51AM -0600, Seth Forshee wrote:
> >
> >>Shortly after that I plan to follow with support for ext4. I've been
> >>fuzzing ext4 for a while now and it has held up well, and I'm currently
> >>working on hand-crafted attacks. Ted has commented privately (to others,
> >>not to me personally) that he will fix bugs for such attacks, though I
> >>haven't seen any public comments to that effect.
> >
> >_Static_ attacks, or change-image-under-mounted-fs attacks?
> To properly protect against attacks on mounted filesystems, we'd
> need some new concept of a userspace immutable file (that is, one
> where nobody can write to it except the kernel, and only the kernel
> can change it between regular access and this new state), and then
> have the kernel set an image (or block device) to this state when a
> filesystem is mounted from it (this introduces all kinds of other
> issues too however, for example stuff that allows an online fsck on
> the device will stop working, as will many un-deletion tools).

Yeah, Serge and I were just tossing that idea around on irc. If we can
make that work then it's probably the best solution.

From a naive perspective it seems like all we really have to do is make
the block device inode immutable to userspace when it is mounted. And
the parent block device if it's a partition, which might be a bit
troublesome. We'd have to ensure that writes couldn't happen via any fds
already open when the device was mounted too.

We'd need some cooperation from the loop driver too I suppose, to make
sure the file backing the loop device is also immutable.

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1271645

FromAustin S Hemmelgarn <ahferroin7@gmail.com>
Date2015-11-17 22:00 +0100
Message-ID<qvVw5-1Ai-1@gated-at.bofh.it>
In reply to#1271586

[Multipart message — attachments visible in raw view] — view raw

On 2015-11-17 14:16, Seth Forshee wrote:
> On Tue, Nov 17, 2015 at 02:02:09PM -0500, Austin S Hemmelgarn wrote:
>> On 2015-11-17 12:55, Al Viro wrote:
>>> On Tue, Nov 17, 2015 at 11:25:51AM -0600, Seth Forshee wrote:
>>>
>>>> Shortly after that I plan to follow with support for ext4. I've been
>>>> fuzzing ext4 for a while now and it has held up well, and I'm currently
>>>> working on hand-crafted attacks. Ted has commented privately (to others,
>>>> not to me personally) that he will fix bugs for such attacks, though I
>>>> haven't seen any public comments to that effect.
>>>
>>> _Static_ attacks, or change-image-under-mounted-fs attacks?
>> To properly protect against attacks on mounted filesystems, we'd
>> need some new concept of a userspace immutable file (that is, one
>> where nobody can write to it except the kernel, and only the kernel
>> can change it between regular access and this new state), and then
>> have the kernel set an image (or block device) to this state when a
>> filesystem is mounted from it (this introduces all kinds of other
>> issues too however, for example stuff that allows an online fsck on
>> the device will stop working, as will many un-deletion tools).
>
> Yeah, Serge and I were just tossing that idea around on irc. If we can
> make that work then it's probably the best solution.
>
>  From a naive perspective it seems like all we really have to do is make
> the block device inode immutable to userspace when it is mounted. And
> the parent block device if it's a partition, which might be a bit
> troublesome. We'd have to ensure that writes couldn't happen via any fds
> already open when the device was mounted too.
>
> We'd need some cooperation from the loop driver too I suppose, to make
> sure the file backing the loop device is also immutable.
>
 From a completeness perspective, you'd also need to hook into DM, MD, 
and bcache to handle their backing devices.  There's not much we could 
do about iSCSI/ATAoE/NBD devices, and I think being able to restrict 
stuff calling mount from a userns to only be able to use FUSE would 
still be useful (FWIW, GRUB2 has a tool to use FUSE for testing it's own 
filesystem drivers, which I use regularly when I just need a read-only 
mount).

[toc] | [prev] | [next] | [standalone]


#1271680

FromSeth Forshee <seth.forshee@canonical.com>
Date2015-11-17 22:40 +0100
Message-ID<qvW8N-24z-1@gated-at.bofh.it>
In reply to#1271645
On Tue, Nov 17, 2015 at 03:54:50PM -0500, Austin S Hemmelgarn wrote:
> On 2015-11-17 14:16, Seth Forshee wrote:
> >On Tue, Nov 17, 2015 at 02:02:09PM -0500, Austin S Hemmelgarn wrote:
> >>On 2015-11-17 12:55, Al Viro wrote:
> >>>On Tue, Nov 17, 2015 at 11:25:51AM -0600, Seth Forshee wrote:
> >>>
> >>>>Shortly after that I plan to follow with support for ext4. I've been
> >>>>fuzzing ext4 for a while now and it has held up well, and I'm currently
> >>>>working on hand-crafted attacks. Ted has commented privately (to others,
> >>>>not to me personally) that he will fix bugs for such attacks, though I
> >>>>haven't seen any public comments to that effect.
> >>>
> >>>_Static_ attacks, or change-image-under-mounted-fs attacks?
> >>To properly protect against attacks on mounted filesystems, we'd
> >>need some new concept of a userspace immutable file (that is, one
> >>where nobody can write to it except the kernel, and only the kernel
> >>can change it between regular access and this new state), and then
> >>have the kernel set an image (or block device) to this state when a
> >>filesystem is mounted from it (this introduces all kinds of other
> >>issues too however, for example stuff that allows an online fsck on
> >>the device will stop working, as will many un-deletion tools).
> >
> >Yeah, Serge and I were just tossing that idea around on irc. If we can
> >make that work then it's probably the best solution.
> >
> > From a naive perspective it seems like all we really have to do is make
> >the block device inode immutable to userspace when it is mounted. And
> >the parent block device if it's a partition, which might be a bit
> >troublesome. We'd have to ensure that writes couldn't happen via any fds
> >already open when the device was mounted too.
> >
> >We'd need some cooperation from the loop driver too I suppose, to make
> >sure the file backing the loop device is also immutable.
> >
> From a completeness perspective, you'd also need to hook into DM,
> MD, and bcache to handle their backing devices.  There's not much we
> could do about iSCSI/ATAoE/NBD devices, and I think being able to

But really no one would be able to mount any of those without
intervention from a privileged user anyway. The same is true today of
loop devices, but I have some patches to change that.

> restrict stuff calling mount from a userns to only be able to use
> FUSE would still be useful (FWIW, GRUB2 has a tool to use FUSE for
> testing it's own filesystem drivers, which I use regularly when I
> just need a read-only mount).

Agreed, fuse alone is very useful, though there is a performance
penalty.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1272119

FromAustin S Hemmelgarn <ahferroin7@gmail.com>
Date2015-11-18 13:30 +0100
Message-ID<qwa26-2NS-5@gated-at.bofh.it>
In reply to#1271680

[Multipart message — attachments visible in raw view] — view raw

On 2015-11-17 16:32, Seth Forshee wrote:
> On Tue, Nov 17, 2015 at 03:54:50PM -0500, Austin S Hemmelgarn wrote:
>> On 2015-11-17 14:16, Seth Forshee wrote:
>>> On Tue, Nov 17, 2015 at 02:02:09PM -0500, Austin S Hemmelgarn wrote:
>>>> On 2015-11-17 12:55, Al Viro wrote:
>>>>> On Tue, Nov 17, 2015 at 11:25:51AM -0600, Seth Forshee wrote:
>>>>>
>>>>>> Shortly after that I plan to follow with support for ext4. I've been
>>>>>> fuzzing ext4 for a while now and it has held up well, and I'm currently
>>>>>> working on hand-crafted attacks. Ted has commented privately (to others,
>>>>>> not to me personally) that he will fix bugs for such attacks, though I
>>>>>> haven't seen any public comments to that effect.
>>>>>
>>>>> _Static_ attacks, or change-image-under-mounted-fs attacks?
>>>> To properly protect against attacks on mounted filesystems, we'd
>>>> need some new concept of a userspace immutable file (that is, one
>>>> where nobody can write to it except the kernel, and only the kernel
>>>> can change it between regular access and this new state), and then
>>>> have the kernel set an image (or block device) to this state when a
>>>> filesystem is mounted from it (this introduces all kinds of other
>>>> issues too however, for example stuff that allows an online fsck on
>>>> the device will stop working, as will many un-deletion tools).
>>>
>>> Yeah, Serge and I were just tossing that idea around on irc. If we can
>>> make that work then it's probably the best solution.
>>>
>>>  From a naive perspective it seems like all we really have to do is make
>>> the block device inode immutable to userspace when it is mounted. And
>>> the parent block device if it's a partition, which might be a bit
>>> troublesome. We'd have to ensure that writes couldn't happen via any fds
>>> already open when the device was mounted too.
>>>
>>> We'd need some cooperation from the loop driver too I suppose, to make
>>> sure the file backing the loop device is also immutable.
>>>
>>  From a completeness perspective, you'd also need to hook into DM,
>> MD, and bcache to handle their backing devices.  There's not much we
>> could do about iSCSI/ATAoE/NBD devices, and I think being able to
>
> But really no one would be able to mount any of those without
> intervention from a privileged user anyway. The same is true today of
> loop devices, but I have some patches to change that.
Um, no, depending on how your system is configured, it's fully possible 
to mount those as a regular user with no administrative interaction 
required.  All that needs done is some udev rules (or something 
equivalent) to add ACL's to the device nodes allowing regular users to 
access them (and there are systems I've seen that are naive enough to do 
this).  And I can almost assure you that there will be someone out there 
who decides to directly expose such block devices to a user namespace.

[toc] | [prev] | [next] | [standalone]


#1272220

FromSeth Forshee <seth.forshee@canonical.com>
Date2015-11-18 15:30 +0100
Message-ID<qwbUf-445-25@gated-at.bofh.it>
In reply to#1272119
On Wed, Nov 18, 2015 at 07:23:48AM -0500, Austin S Hemmelgarn wrote:
> On 2015-11-17 16:32, Seth Forshee wrote:
> >On Tue, Nov 17, 2015 at 03:54:50PM -0500, Austin S Hemmelgarn wrote:
> >>On 2015-11-17 14:16, Seth Forshee wrote:
> >>>On Tue, Nov 17, 2015 at 02:02:09PM -0500, Austin S Hemmelgarn wrote:
> >>>>On 2015-11-17 12:55, Al Viro wrote:
> >>>>>On Tue, Nov 17, 2015 at 11:25:51AM -0600, Seth Forshee wrote:
> >>>>>
> >>>>>>Shortly after that I plan to follow with support for ext4. I've been
> >>>>>>fuzzing ext4 for a while now and it has held up well, and I'm currently
> >>>>>>working on hand-crafted attacks. Ted has commented privately (to others,
> >>>>>>not to me personally) that he will fix bugs for such attacks, though I
> >>>>>>haven't seen any public comments to that effect.
> >>>>>
> >>>>>_Static_ attacks, or change-image-under-mounted-fs attacks?
> >>>>To properly protect against attacks on mounted filesystems, we'd
> >>>>need some new concept of a userspace immutable file (that is, one
> >>>>where nobody can write to it except the kernel, and only the kernel
> >>>>can change it between regular access and this new state), and then
> >>>>have the kernel set an image (or block device) to this state when a
> >>>>filesystem is mounted from it (this introduces all kinds of other
> >>>>issues too however, for example stuff that allows an online fsck on
> >>>>the device will stop working, as will many un-deletion tools).
> >>>
> >>>Yeah, Serge and I were just tossing that idea around on irc. If we can
> >>>make that work then it's probably the best solution.
> >>>
> >>> From a naive perspective it seems like all we really have to do is make
> >>>the block device inode immutable to userspace when it is mounted. And
> >>>the parent block device if it's a partition, which might be a bit
> >>>troublesome. We'd have to ensure that writes couldn't happen via any fds
> >>>already open when the device was mounted too.
> >>>
> >>>We'd need some cooperation from the loop driver too I suppose, to make
> >>>sure the file backing the loop device is also immutable.
> >>>
> >> From a completeness perspective, you'd also need to hook into DM,
> >>MD, and bcache to handle their backing devices.  There's not much we
> >>could do about iSCSI/ATAoE/NBD devices, and I think being able to
> >
> >But really no one would be able to mount any of those without
> >intervention from a privileged user anyway. The same is true today of
> >loop devices, but I have some patches to change that.
> Um, no, depending on how your system is configured, it's fully
> possible to mount those as a regular user with no administrative
> interaction required.  All that needs done is some udev rules (or
> something equivalent) to add ACL's to the device nodes allowing
> regular users to access them (and there are systems I've seen that
> are naive enough to do this).  And I can almost assure you that
> there will be someone out there who decides to directly expose such
> block devices to a user namespace.

But it still requires the admin set it up that way, no? And aren't
privileges required to set up those devices in the first place?

I'm not saying that it wouldn't be a good idea to lock down the backing
stores for those types of devices too, just that it isn't something that
a regular user could exploit without an admin doing something to
facilitate it.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1272262

FromAl Viro <viro@ZenIV.linux.org.uk>
Date2015-11-18 16:00 +0100
Message-ID<qwcng-4gQ-21@gated-at.bofh.it>
In reply to#1272220
On Wed, Nov 18, 2015 at 08:22:38AM -0600, Seth Forshee wrote:

> But it still requires the admin set it up that way, no? And aren't
> privileges required to set up those devices in the first place?
> 
> I'm not saying that it wouldn't be a good idea to lock down the backing
> stores for those types of devices too, just that it isn't something that
> a regular user could exploit without an admin doing something to
> facilitate it.

Sigh...  If it boils down to "all admins within all containers must be
trusted not to try and break out" (along with "roothole in any container
escalates to kernel-mode code execution on host"), then what the fuck
is the *point* of bothering with containers, userns, etc. in the first
place?  If your model is basically "you want isolation, just use kvm",
fine, but where's the place for userns in all that?

And if you are talking about the _host_ admin, then WTF not have him just
mount what's needed as part of setup and to hell with mounting those
inside the container?

Look at that from the hosting company POV - they are offering a bunch of
virtual machines on one physical system.  And you want the admins on those
virtual machines independent from the host admin.  Fine, but then you
really need to keep them unable to screw each other or gain kernel-mode
execution on the host.

Again, what's the point of all that?  I assumed the model where containers
do, you know, contain what's in them, regardless of trust.  You guys seem
to assume something different and I really wonder what it _is_...
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1272277

FromSeth Forshee <seth.forshee@canonical.com>
Date2015-11-18 16:10 +0100
Message-ID<qwcwW-4zt-7@gated-at.bofh.it>
In reply to#1272262
On Wed, Nov 18, 2015 at 02:58:18PM +0000, Al Viro wrote:
> On Wed, Nov 18, 2015 at 08:22:38AM -0600, Seth Forshee wrote:
> 
> > But it still requires the admin set it up that way, no? And aren't
> > privileges required to set up those devices in the first place?
> > 
> > I'm not saying that it wouldn't be a good idea to lock down the backing
> > stores for those types of devices too, just that it isn't something that
> > a regular user could exploit without an admin doing something to
> > facilitate it.
> 
> Sigh...  If it boils down to "all admins within all containers must be
> trusted not to try and break out" (along with "roothole in any container
> escalates to kernel-mode code execution on host"), then what the fuck
> is the *point* of bothering with containers, userns, etc. in the first
> place?  If your model is basically "you want isolation, just use kvm",
> fine, but where's the place for userns in all that?
> 
> And if you are talking about the _host_ admin, then WTF not have him just
> mount what's needed as part of setup and to hell with mounting those
> inside the container?

Yes, the host admin. I'm not talking about trusting the admin inside the
container at all.

From my perspective the idea is essentially to allow mounting with fuse
or with ext4 using "mount -o loop ..." within a container.

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1272281

FromAl Viro <viro@ZenIV.linux.org.uk>
Date2015-11-18 16:20 +0100
Message-ID<qwcGC-4D5-13@gated-at.bofh.it>
In reply to#1272277
On Wed, Nov 18, 2015 at 09:05:12AM -0600, Seth Forshee wrote:

> Yes, the host admin. I'm not talking about trusting the admin inside the
> container at all.

Then why not have the same host admin just plain mount it when setting the
container up and be done with that?  From the host namespace, before spawning
the docker instance or whatever framework you are using.  IDGI...
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1272291

FromRichard Weinberger <richard.weinberger@gmail.com>
Date2015-11-18 16:30 +0100
Message-ID<qwcQi-4HC-9@gated-at.bofh.it>
In reply to#1272281
On Wed, Nov 18, 2015 at 4:13 PM, Al Viro <viro@zeniv.linux.org.uk> wrote:
> On Wed, Nov 18, 2015 at 09:05:12AM -0600, Seth Forshee wrote:
>
>> Yes, the host admin. I'm not talking about trusting the admin inside the
>> container at all.
>
> Then why not have the same host admin just plain mount it when setting the
> container up and be done with that?  From the host namespace, before spawning
> the docker instance or whatever framework you are using.  IDGI...

Because hosting companies sell containers as "full virtual machines"
and customers expect to be able mount stuff like disk images they upload.

-- 
Thanks,
//richard
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1272902

FromJames Morris <jmorris@namei.org>
Date2015-11-19 08:50 +0100
Message-ID<qws8G-6mi-9@gated-at.bofh.it>
In reply to#1272291
On Wed, 18 Nov 2015, Richard Weinberger wrote:

> On Wed, Nov 18, 2015 at 4:13 PM, Al Viro <viro@zeniv.linux.org.uk> wrote:
> > On Wed, Nov 18, 2015 at 09:05:12AM -0600, Seth Forshee wrote:
> >
> >> Yes, the host admin. I'm not talking about trusting the admin inside the
> >> container at all.
> >
> > Then why not have the same host admin just plain mount it when setting the
> > container up and be done with that?  From the host namespace, before spawning
> > the docker instance or whatever framework you are using.  IDGI...
> 
> Because hosting companies sell containers as "full virtual machines"
> and customers expect to be able mount stuff like disk images they upload.

I don't think this is a valid reason for merging functionality into the 
kernel.


-- 
James Morris
<jmorris@namei.org>

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1272907

FromRichard Weinberger <richard@nod.at>
Date2015-11-19 09:00 +0100
Message-ID<qwsin-6pI-9@gated-at.bofh.it>
In reply to#1272902
Am 19.11.2015 um 08:47 schrieb James Morris:
> On Wed, 18 Nov 2015, Richard Weinberger wrote:
> 
>> On Wed, Nov 18, 2015 at 4:13 PM, Al Viro <viro@zeniv.linux.org.uk> wrote:
>>> On Wed, Nov 18, 2015 at 09:05:12AM -0600, Seth Forshee wrote:
>>>
>>>> Yes, the host admin. I'm not talking about trusting the admin inside the
>>>> container at all.
>>>
>>> Then why not have the same host admin just plain mount it when setting the
>>> container up and be done with that?  From the host namespace, before spawning
>>> the docker instance or whatever framework you are using.  IDGI...
>>
>> Because hosting companies sell containers as "full virtual machines"
>> and customers expect to be able mount stuff like disk images they upload.
> 
> I don't think this is a valid reason for merging functionality into the 
> kernel.

Erm, I don't want this in the kernel. That's why I've proposed the lklfuse approach.

Thanks,
//richard
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1273172

From"Serge E. Hallyn" <serge.hallyn@ubuntu.com>
Date2015-11-19 15:30 +0100
Message-ID<qwynN-22I-29@gated-at.bofh.it>
In reply to#1272907
On Thu, Nov 19, 2015 at 08:53:26AM +0100, Richard Weinberger wrote:
> Am 19.11.2015 um 08:47 schrieb James Morris:
> > On Wed, 18 Nov 2015, Richard Weinberger wrote:
> > 
> >> On Wed, Nov 18, 2015 at 4:13 PM, Al Viro <viro@zeniv.linux.org.uk> wrote:
> >>> On Wed, Nov 18, 2015 at 09:05:12AM -0600, Seth Forshee wrote:
> >>>
> >>>> Yes, the host admin. I'm not talking about trusting the admin inside the
> >>>> container at all.
> >>>
> >>> Then why not have the same host admin just plain mount it when setting the
> >>> container up and be done with that?  From the host namespace, before spawning
> >>> the docker instance or whatever framework you are using.  IDGI...
> >>
> >> Because hosting companies sell containers as "full virtual machines"
> >> and customers expect to be able mount stuff like disk images they upload.
> > 
> > I don't think this is a valid reason for merging functionality into the 
> > kernel.
> 
> Erm, I don't want this in the kernel. That's why I've proposed the lklfuse approach.

That would require the fuse-in-containers functionality, right?
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1273208

FromRichard Weinberger <richard@nod.at>
Date2015-11-19 16:10 +0100
Message-ID<qwz0t-2xX-13@gated-at.bofh.it>
In reply to#1273172
Am 19.11.2015 um 15:21 schrieb Serge E. Hallyn:
> On Thu, Nov 19, 2015 at 08:53:26AM +0100, Richard Weinberger wrote:
>> Am 19.11.2015 um 08:47 schrieb James Morris:
>>> On Wed, 18 Nov 2015, Richard Weinberger wrote:
>>>
>>>> On Wed, Nov 18, 2015 at 4:13 PM, Al Viro <viro@zeniv.linux.org.uk> wrote:
>>>>> On Wed, Nov 18, 2015 at 09:05:12AM -0600, Seth Forshee wrote:
>>>>>
>>>>>> Yes, the host admin. I'm not talking about trusting the admin inside the
>>>>>> container at all.
>>>>>
>>>>> Then why not have the same host admin just plain mount it when setting the
>>>>> container up and be done with that?  From the host namespace, before spawning
>>>>> the docker instance or whatever framework you are using.  IDGI...
>>>>
>>>> Because hosting companies sell containers as "full virtual machines"
>>>> and customers expect to be able mount stuff like disk images they upload.
>>>
>>> I don't think this is a valid reason for merging functionality into the 
>>> kernel.
>>
>> Erm, I don't want this in the kernel. That's why I've proposed the lklfuse approach.
> 
> That would require the fuse-in-containers functionality, right?

Correct.
Still a less large attack surface. :-)

Thanks,
//richard
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1273174

FromColin Walters <walters@verbum.org>
Date2015-11-19 15:40 +0100
Message-ID<qwyxs-25Z-7@gated-at.bofh.it>
In reply to#1272907
On Thu, Nov 19, 2015, at 02:53 AM, Richard Weinberger wrote:

> Erm, I don't want this in the kernel. That's why I've proposed the lklfuse approach.

I already said this before but just to repeat, since I'm confused:

How would "lklfuse" be different from http://libguestfs.org/
which we at Red Hat (and a number of other organizations)
use quite widely now for build systems, debugging etc.

In the end it's just running the kernel in KVM with a custom protocol,
with support for non-filesystem things like "install a bootloader",
and it already supports FUSE.

I'm pretty firmly with Al here - the attack surface increase here
is too great, and we'd likely turn this off if it even did make it
into the kernel.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1273183

FromRichard Weinberger <richard@nod.at>
Date2015-11-19 15:50 +0100
Message-ID<qwyH8-29G-23@gated-at.bofh.it>
In reply to#1273174
Am 19.11.2015 um 15:37 schrieb Colin Walters:
> On Thu, Nov 19, 2015, at 02:53 AM, Richard Weinberger wrote:
> 
>> Erm, I don't want this in the kernel. That's why I've proposed the lklfuse approach.
> 
> I already said this before but just to repeat, since I'm confused:
> 
> How would "lklfuse" be different from http://libguestfs.org/
> which we at Red Hat (and a number of other organizations)
> use quite widely now for build systems, debugging etc.

Currently libguestfs has a rather huge overhead because it
boots a full virtual machine and hence a lot of communication
is needed.
With LKL you can use Linux as Library and link it to fuse.
AFAIK Richard added already a LKL backend to libguestfs. :-)

> In the end it's just running the kernel in KVM with a custom protocol,
> with support for non-filesystem things like "install a bootloader",
> and it already supports FUSE.
> 
> I'm pretty firmly with Al here - the attack surface increase here
> is too great, and we'd likely turn this off if it even did make it
> into the kernel.

Agreed. This is why I'm promoting the fuse solution.

Thanks,
//richard
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1273211

From"Richard W.M. Jones" <rjones@redhat.com>
Date2015-11-19 16:20 +0100
Message-ID<qwzaa-2Bz-11@gated-at.bofh.it>
In reply to#1273183
On Thu, Nov 19, 2015 at 03:49:00PM +0100, Richard Weinberger wrote:
> Am 19.11.2015 um 15:37 schrieb Colin Walters:
> > On Thu, Nov 19, 2015, at 02:53 AM, Richard Weinberger wrote:
> > 
> >> Erm, I don't want this in the kernel. That's why I've proposed the lklfuse approach.
> > 
> > I already said this before but just to repeat, since I'm confused:
> > 
> > How would "lklfuse" be different from http://libguestfs.org/
> > which we at Red Hat (and a number of other organizations)
> > use quite widely now for build systems, debugging etc.
> 
> Currently libguestfs has a rather huge overhead because it
> boots a full virtual machine and hence a lot of communication
> is needed.
> With LKL you can use Linux as Library and link it to fuse.
> AFAIK Richard added already a LKL backend to libguestfs. :-)

Right.  For the longer story, see:

https://rwmj.wordpress.com/2015/11/07/linux-kernel-library-backend-for-libguestfs/#content

Rich.

-- 
Richard Jones, Virtualization Group, Red Hat http://people.redhat.com/~rjones
Read my programming and virtualization blog: http://rwmj.wordpress.com
virt-top is 'top' for virtual machines.  Tiny program with many
powerful monitoring features, net stats, disk stats, logging, etc.
http://people.redhat.com/~rjones/virt-top
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1273204

From"Serge E. Hallyn" <serge.hallyn@ubuntu.com>
Date2015-11-19 16:00 +0100
Message-ID<qwyQP-2dv-43@gated-at.bofh.it>
In reply to#1272281
On Wed, Nov 18, 2015 at 03:13:35PM +0000, Al Viro wrote:
> On Wed, Nov 18, 2015 at 09:05:12AM -0600, Seth Forshee wrote:
> 
> > Yes, the host admin. I'm not talking about trusting the admin inside the
> > container at all.
> 
> Then why not have the same host admin just plain mount it when setting the
> container up and be done with that?  From the host namespace, before spawning
> the docker instance or whatever framework you are using.  IDGI...

fwiw one example use case is building a livecd or vm image from inside
container, on a system where each piece of functionality is
compartmentalized in a separate container.

For specific programs we can engineer them to call out to a helper
in the init user_ns, but that doesn't work for existing tools.

James Bottomley a year or two ago had mentioned the idea of using
seccomp to have the mount syscall in a container trap into a
init_user_ns helper, but we'd have to ptrace the whole container
for its whole lifetime to do that, or use SECCOMP_RET_TRAP with
a LD_PRELOAD that does goes through a proxy to do the remote
request (which becomes infeasible bc LD_PRELOAD won't stick).

If we did enable anything but FUSE I would hope each fs would have
a separate sysctl to enable non-init-userns mounts.  Then the admin
could enable them just as they do the automount of whatever garbage
is on a usb stick found lying on the trade floor.

-serge
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


Page 2 of 3 — ← Prev page 1 [2] 3  Next page →

Back to top | Article view | linux.kernel


csiph-web