Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1437548 > unrolled thread

Re: Introspecting userns relationships to other namespaces?

Started by"Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
First post2016-07-06 10:50 +0200
Last post2016-07-09 05:30 +0200
Articles 20 on this page of 33 — 7 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: Introspecting userns relationships to other namespaces? "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-06 10:50 +0200
    Re: Introspecting userns relationships to other namespaces? "Serge E. Hallyn" <serge@hallyn.com> - 2016-07-06 16:20 +0200
      Re: Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-06 18:00 +0200
        Re: [CRIU] Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-08 10:00 +0200
          Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-08 16:40 +0200
            Re: [CRIU] Introspecting userns relationships to other namespaces? "W. Trevor King" <wking@tremily.us> - 2016-07-08 23:00 +0200
            Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@Hansenpartnership.com> - 2016-07-09 00:30 +0200
            Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@Hansenpartnership.com> - 2016-07-09 00:30 +0200
              Re: [CRIU] Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-09 02:10 +0200
                Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-09 02:20 +0200
                  Re: [CRIU] Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-09 05:20 +0200
                    Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@Hansenpartnership.com> - 2016-07-09 12:40 +0200
                    Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@Hansenpartnership.com> - 2016-07-09 12:40 +0200
                    Re: [CRIU] Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-09 20:30 +0200
                      Re: [CRIU] Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-09 20:50 +0200
      Re: Introspecting userns relationships to other namespaces? "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-07 10:20 +0200
        Re: Introspecting userns relationships to other namespaces? "Serge E. Hallyn" <serge@hallyn.com> - 2016-07-07 15:40 +0200
          Re: Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-07 17:10 +0200
            Re: Introspecting userns relationships to other namespaces? "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-07 20:30 +0200
              Re: Introspecting userns relationships to other namespaces? "Serge E. Hallyn" <serge@hallyn.com> - 2016-07-07 20:30 +0200
              Re: Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-07 21:20 +0200
                Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-08 05:30 +0200
                  Re: Introspecting userns relationships to other namespaces? "W. Trevor King" <wking@tremily.us> - 2016-07-08 07:40 +0200
                    Re: Introspecting userns relationships to other namespaces? "W. Trevor King" <wking@tremily.us> - 2016-07-08 08:20 +0200
                    Re: Introspecting userns relationships to other namespaces? "W. Trevor King" <wking@tremily.us> - 2016-07-08 09:20 +0200
                  Re: [CRIU] Introspecting userns relationships to other namespaces? Andrei Vagin <avagin@gmail.com> - 2016-07-08 07:50 +0200
                    Re: [CRIU] Introspecting userns relationships to other namespaces? Andrei Vagin <avagin@gmail.com> - 2016-07-08 07:50 +0200
                    Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-08 08:20 +0200
                  Re: [CRIU] Introspecting userns relationships to other namespaces? "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-08 13:20 +0200
                Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-08 05:30 +0200
                Re: Introspecting userns relationships to other namespaces? "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-08 13:20 +0200
            Re: Introspecting userns relationships to other namespaces? "W. Trevor King" <wking@tremily.us> - 2016-07-09 05:20 +0200
              Re: Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-09 05:30 +0200

Page 1 of 2  [1] 2  Next page →


#1437548 — Re: Introspecting userns relationships to other namespaces?

From"Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
Date2016-07-06 10:50 +0200
SubjectRe: Introspecting userns relationships to other namespaces?
Message-ID<rRQGS-6SL-5@gated-at.bofh.it>
[Rats! Doing now what I should have down to start with. Looping some
lists and CRIU and other possibly relevant people into this
conversation]

Hi Eric,

On 5 July 2016 at 23:47, Eric W. Biederman <ebiederm@xmission.com> wrote:
> "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> writes:
>
>> Hi Eric,
>>
>> I have a question. Is there any way currently to discover which
>> user namespace a particular nonuser namespace is governed by?
>> Maybe I am missing something, but there does not seem to be a
>> way to do this. Also, can one discover which userns is the
>> parent of a given userns? Again, I can't see a way to do this.
>>
>> The point here is introspecting so that a process might determine
>> what its capabilities are when operating on some resource governed
>> by a (nonuser) namespace.
>
> To the best of my knowledge that there is not an interface to get that
> information.  It would be good to have such an interface for no other
> reason than the CRIU folks are going to need it at some point.  I am a
> bit surprised they have not complained yet.
>
> That said in a normal use scenario I don't think that information is
> needed.
>
> Do you have a particular use case besides checkpoint/restart where this
> is useful?  That might help in coming up with a good userspace interface
> for this information.

So, I spend a moderate amount of time working with people to introduce
them to the namespaces infrastructure, and one topic that comes up now
and this introspection/visualization tools. For example,
nowadays--thanks to the (bizarrely misnamed) NStgid and NSpid fields
in /proc/PID--it's possible to (and someone I was working with did)
write tools that introspect the PID namespace hierarchy to show all of
process's and their PIDs in the various namespace instance. It's a
natural enough thing to want to do, when confronted with the
complexity of the namespaces.

Someone else then asked me a question that led me to wonder about
generally introspecting on the parental relationships between user
namespaces and the association of other namespaces types with user
namespaces. One use would be visualization, in order to understand the
running system. Another would be to answer the question I already
mentioned: what capability does process X have to perform operations
on a resource governed by namespace Y?

Cheers,

Michael




-- 
Michael Kerrisk
Linux man-pages maintainer; http://www.kernel.org/doc/man-pages/
Linux/UNIX System Programming Training: http://man7.org/training/

[toc] | [next] | [standalone]


#1437759

From"Serge E. Hallyn" <serge@hallyn.com>
Date2016-07-06 16:20 +0200
Message-ID<rRVQd-1PF-17@gated-at.bofh.it>
In reply to#1437548
On Wed, Jul 06, 2016 at 10:41:48AM +0200, Michael Kerrisk (man-pages) wrote:
> [Rats! Doing now what I should have down to start with. Looping some
> lists and CRIU and other possibly relevant people into this
> conversation]
> 
> Hi Eric,
> 
> On 5 July 2016 at 23:47, Eric W. Biederman <ebiederm@xmission.com> wrote:
> > "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> writes:
> >
> >> Hi Eric,
> >>
> >> I have a question. Is there any way currently to discover which
> >> user namespace a particular nonuser namespace is governed by?
> >> Maybe I am missing something, but there does not seem to be a
> >> way to do this. Also, can one discover which userns is the
> >> parent of a given userns? Again, I can't see a way to do this.
> >>
> >> The point here is introspecting so that a process might determine
> >> what its capabilities are when operating on some resource governed
> >> by a (nonuser) namespace.
> >
> > To the best of my knowledge that there is not an interface to get that
> > information.  It would be good to have such an interface for no other
> > reason than the CRIU folks are going to need it at some point.  I am a
> > bit surprised they have not complained yet.

I don't think they need it.  They do in fact have what they need.  Assume
you have tasks T1, T2, T1_1 and T2_1;  T1 and T2 are in init_user_ns;  T1
spawned T1_1 in a new userns;  T2 spawned T2_1 which setns()d to T1_1's ns.
There's some {handwave} uid mapping, does not matter.

At restart, it doesn't matter which task originally created the new userns.
criu knows T1_1 and T2_1 are in the same userns;  it creates the userns, sets
up the mapping, and T1_1 and T2_1 setns() to it.

> > That said in a normal use scenario I don't think that information is
> > needed.
> >
> > Do you have a particular use case besides checkpoint/restart where this
> > is useful?  That might help in coming up with a good userspace interface
> > for this information.
> 
> So, I spend a moderate amount of time working with people to introduce
> them to the namespaces infrastructure, and one topic that comes up now
> and this introspection/visualization tools. For example,
> nowadays--thanks to the (bizarrely misnamed) NStgid and NSpid fields
> in /proc/PID--it's possible to (and someone I was working with did)
> write tools that introspect the PID namespace hierarchy to show all of
> process's and their PIDs in the various namespace instance. It's a
> natural enough thing to want to do, when confronted with the
> complexity of the namespaces.
> 
> Someone else then asked me a question that led me to wonder about
> generally introspecting on the parental relationships between user
> namespaces and the association of other namespaces types with user
> namespaces. One use would be visualization, in order to understand the
> running system. Another would be to answer the question I already
> mentioned: what capability does process X have to perform operations
> on a resource governed by namespace Y?

I agree they'll probably want it, but if we want for a real need and
use case we can do a better job of providing what's needed.

-serge

[toc] | [prev] | [next] | [standalone]


#1437808

Fromebiederm@xmission.com (Eric W. Biederman)
Date2016-07-06 18:00 +0200
Message-ID<rRXp0-2E2-3@gated-at.bofh.it>
In reply to#1437759
"Serge E. Hallyn" <serge@hallyn.com> writes:

> On Wed, Jul 06, 2016 at 10:41:48AM +0200, Michael Kerrisk (man-pages) wrote:
>> [Rats! Doing now what I should have down to start with. Looping some
>> lists and CRIU and other possibly relevant people into this
>> conversation]
>> 
>> Hi Eric,
>> 
>> On 5 July 2016 at 23:47, Eric W. Biederman <ebiederm@xmission.com> wrote:
>> > "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> writes:
>> >
>> >> Hi Eric,
>> >>
>> >> I have a question. Is there any way currently to discover which
>> >> user namespace a particular nonuser namespace is governed by?
>> >> Maybe I am missing something, but there does not seem to be a
>> >> way to do this. Also, can one discover which userns is the
>> >> parent of a given userns? Again, I can't see a way to do this.
>> >>
>> >> The point here is introspecting so that a process might determine
>> >> what its capabilities are when operating on some resource governed
>> >> by a (nonuser) namespace.
>> >
>> > To the best of my knowledge that there is not an interface to get that
>> > information.  It would be good to have such an interface for no other
>> > reason than the CRIU folks are going to need it at some point.  I am a
>> > bit surprised they have not complained yet.
>
> I don't think they need it.  They do in fact have what they need.  Assume
> you have tasks T1, T2, T1_1 and T2_1;  T1 and T2 are in init_user_ns;  T1
> spawned T1_1 in a new userns;  T2 spawned T2_1 which setns()d to T1_1's ns.
> There's some {handwave} uid mapping, does not matter.
>
> At restart, it doesn't matter which task originally created the new userns.
> criu knows T1_1 and T2_1 are in the same userns;  it creates the userns, sets
> up the mapping, and T1_1 and T2_1 setns() to it.

Given that the simple cases are so easy it probably doesn't matter in
that sense.

However we now have the case where user namespaces own pid namespaces,
and uts namespaces, and network namespaces, and ipc namespaces, and
filesystems.  Throw in some mount propagation and use of setns and
things could get confusing.   It is something that will need to be
figured out if CRIU is going to properly checkpoint containers
containing containers containing containers containing containers.

Did I mention I like recursion?

>> > That said in a normal use scenario I don't think that information is
>> > needed.
>> >
>> > Do you have a particular use case besides checkpoint/restart where this
>> > is useful?  That might help in coming up with a good userspace interface
>> > for this information.
>> 
>> So, I spend a moderate amount of time working with people to introduce
>> them to the namespaces infrastructure, and one topic that comes up now
>> and this introspection/visualization tools. For example,
>> nowadays--thanks to the (bizarrely misnamed) NStgid and NSpid fields
>> in /proc/PID--it's possible to (and someone I was working with did)
>> write tools that introspect the PID namespace hierarchy to show all of
>> process's and their PIDs in the various namespace instance. It's a
>> natural enough thing to want to do, when confronted with the
>> complexity of the namespaces.
>> 
>> Someone else then asked me a question that led me to wonder about
>> generally introspecting on the parental relationships between user
>> namespaces and the association of other namespaces types with user
>> namespaces. One use would be visualization, in order to understand the
>> running system. Another would be to answer the question I already
>> mentioned: what capability does process X have to perform operations
>> on a resource governed by namespace Y?
>
> I agree they'll probably want it, but if we want for a real need and
> use case we can do a better job of providing what's needed.

That two which is why I mentioned CRIU.  But yeah it will probably take
a little while to get there.

Eric

[toc] | [prev] | [next] | [standalone]


#1439160 — Re: [CRIU] Introspecting userns relationships to other namespaces?

Fromebiederm@xmission.com (Eric W. Biederman)
Date2016-07-08 10:00 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSyRz-1NP-3@gated-at.bofh.it>
In reply to#1437808
Andrew Vagin <avagin@virtuozzo.com> writes:

> On Wed, Jul 06, 2016 at 10:46:33AM -0500, Eric W. Biederman wrote:
>> "Serge E. Hallyn" <serge@hallyn.com> writes:
>> 
>> > On Wed, Jul 06, 2016 at 10:41:48AM +0200, Michael Kerrisk (man-pages) wrote:
>> >> [Rats! Doing now what I should have down to start with. Looping some
>> >> lists and CRIU and other possibly relevant people into this
>> >> conversation]
>> >> 
>> >> Hi Eric,
>> >> 
>> >> On 5 July 2016 at 23:47, Eric W. Biederman <ebiederm@xmission.com> wrote:
>> >> > "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> writes:
>> >> >
>> >> >> Hi Eric,
>> >> >>
>> >> >> I have a question. Is there any way currently to discover which
>> >> >> user namespace a particular nonuser namespace is governed by?
>> >> >> Maybe I am missing something, but there does not seem to be a
>> >> >> way to do this. Also, can one discover which userns is the
>> >> >> parent of a given userns? Again, I can't see a way to do this.
>> >> >>
>> >> >> The point here is introspecting so that a process might determine
>> >> >> what its capabilities are when operating on some resource governed
>> >> >> by a (nonuser) namespace.
>> >> >
>> >> > To the best of my knowledge that there is not an interface to get that
>> >> > information.  It would be good to have such an interface for no other
>> >> > reason than the CRIU folks are going to need it at some point.  I am a
>> >> > bit surprised they have not complained yet.
>> >
>> > I don't think they need it.  They do in fact have what they need.  Assume
>> > you have tasks T1, T2, T1_1 and T2_1;  T1 and T2 are in init_user_ns;  T1
>> > spawned T1_1 in a new userns;  T2 spawned T2_1 which setns()d to T1_1's ns.
>> > There's some {handwave} uid mapping, does not matter.
>> >
>> > At restart, it doesn't matter which task originally created the new userns.
>> > criu knows T1_1 and T2_1 are in the same userns;  it creates the userns, sets
>> > up the mapping, and T1_1 and T2_1 setns() to it.
>> 
>> Given that the simple cases are so easy it probably doesn't matter in
>> that sense.
>> 
>> However we now have the case where user namespaces own pid namespaces,
>> and uts namespaces, and network namespaces, and ipc namespaces, and
>> filesystems.  Throw in some mount propagation and use of setns and
>> things could get confusing.   It is something that will need to be
>> figured out if CRIU is going to properly checkpoint containers
>> containing containers containing containers containing containers.
>
> It isn't a joke:). We have a few requests to support CR of containers with
> Docker containers inside. And we are going to start this task in a near
> future, so we would like to have interface to get dependencies between
> namespaces too.
>
> BTW: CRIU already supports nested mount namespaces, because systemd
> creates them for services.

The tricky part about this and what messes up James proposed plan is
that the interface needs to be something that returns a namespace file
descriptor.  So we can't print something out in a simple text file.
Well I suppose we could print an device number and inode number pair.
But then someone would still have to scour processes looking for a user
namespace so that is likely less than ideal.

Starting with 4.8 we are also going to need to be able to retrieve the
user namespace owner of filesystems.  That will be an interesting mix.

Eric

[toc] | [prev] | [next] | [standalone]


#1439525 — Re: [CRIU] Introspecting userns relationships to other namespaces?

FromJames Bottomley <James.Bottomley@HansenPartnership.com>
Date2016-07-08 16:40 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSF6G-60Q-47@gated-at.bofh.it>
In reply to#1439160
On Fri, 2016-07-08 at 02:44 -0500, Eric W. Biederman wrote:
> Andrew Vagin <avagin@virtuozzo.com> writes:
> 
> > On Wed, Jul 06, 2016 at 10:46:33AM -0500, Eric W. Biederman wrote:
> > > "Serge E. Hallyn" <serge@hallyn.com> writes:
> > > 
> > > > On Wed, Jul 06, 2016 at 10:41:48AM +0200, Michael Kerrisk (man
> > > > -pages) wrote:
> > > > > [Rats! Doing now what I should have down to start with.
> > > > > Looping some
> > > > > lists and CRIU and other possibly relevant people into this
> > > > > conversation]
> > > > > 
> > > > > Hi Eric,
> > > > > 
> > > > > On 5 July 2016 at 23:47, Eric W. Biederman <
> > > > > ebiederm@xmission.com> wrote:
> > > > > > "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
> > > > > > writes:
> > > > > > 
> > > > > > > Hi Eric,
> > > > > > > 
> > > > > > > I have a question. Is there any way currently to discover 
> > > > > > > which user namespace a particular nonuser namespace is 
> > > > > > > governed by? Maybe I am missing something, but there does 
> > > > > > > not seem to be a way to do this. Also, can one discover 
> > > > > > > which userns is the parent of a given userns? Again, I 
> > > > > > > can't see a way to do this.
> > > > > > > 
> > > > > > > The point here is introspecting so that a process might 
> > > > > > > determine what its capabilities are when operating on 
> > > > > > > some resource governed by a (nonuser) namespace.
> > > > > > 
> > > > > > To the best of my knowledge that there is not an interface 
> > > > > > to get that information.  It would be good to have such an 
> > > > > > interface for no other reason than the CRIU folks are going 
> > > > > > to need it at some point.  I am a bit surprised they have
> > > > > > not complained yet.
> > > > 
> > > > I don't think they need it.  They do in fact have what they 
> > > > need.  Assume you have tasks T1, T2, T1_1 and T2_1;  T1 and T2 
> > > > are in init_user_ns;  T1 spawned T1_1 in a new userns;  T2 
> > > > spawned T2_1 which setns()d to T1_1's ns. There's some
> > > > {handwave} uid mapping, does not matter.
> > > > 
> > > > At restart, it doesn't matter which task originally created the 
> > > > new userns. criu knows T1_1 and T2_1 are in the same userns; 
> > > >  it creates the userns, sets up the mapping, and T1_1 and T2_1
> > > > setns() to it.
> > > 
> > > Given that the simple cases are so easy it probably doesn't 
> > > matter in that sense.
> > > 
> > > However we now have the case where user namespaces own pid 
> > > namespaces, and uts namespaces, and network namespaces, and ipc 
> > > namespaces, and filesystems.  Throw in some mount propagation and 
> > > use of setns and things could get confusing.   It is something 
> > > that will need to be figured out if CRIU is going to properly 
> > > checkpoint containers containing containers containing containers 
> > > containing containers.
> > 
> > It isn't a joke:). We have a few requests to support CR of 
> > containers with Docker containers inside. And we are going to start 
> > this task in a near future, so we would like to have interface to 
> > get dependencies between namespaces too.
> > 
> > BTW: CRIU already supports nested mount namespaces, because systemd
> > creates them for services.
> 
> The tricky part about this and what messes up James proposed plan is
> that the interface needs to be something that returns a namespace 
> file descriptor.  So we can't print something out in a simple text
> file.

I actually described two problems: the first was how we get the
information in the first place.  Currently the owning or parent user_ns
is tucked inside an opaque structure.  I think we need to move that to
ns_common where it would be the owning userns for all non-user
namespaces and the parent for the userns.

Once we actually have the information, we can also add a set of proc
links, say either

/proc/<pid>/ns/X-userns

Which might be a bit messy since it doubles the number of files, or
perhaps in a simple directory.

> Well I suppose we could print an device number and inode number pair.
> But then someone would still have to scour processes looking for a 
> user namespace so that is likely less than ideal.

There's no reason any of the proposed methods so far have to be
exclusive: nsfs.c has a lot of flexibility.

> Starting with 4.8 we are also going to need to be able to retrieve 
> the user namespace owner of filesystems.  That will be an interesting
> mix.

This is per mount point, isn't it? so it can't be in /proc/fs/ and it
would have to be per local mount tree.  Yes, that is a bit nasty. 
 Sounds like we might need to unfold mount or mountinfo into something
that has one directory per entry?

James

> Eric
> 
> _______________________________________________
> Containers mailing list
> Containers@lists.linux-foundation.org
> https://lists.linuxfoundation.org/mailman/listinfo/containers
> 

[toc] | [prev] | [next] | [standalone]


#1439794 — Re: [CRIU] Introspecting userns relationships to other namespaces?

From"W. Trevor King" <wking@tremily.us>
Date2016-07-08 23:00 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSL2p-1mP-9@gated-at.bofh.it>
In reply to#1439525

[Multipart message — attachments visible in raw view] — view raw

On Fri, Jul 08, 2016 at 01:38:19PM -0700, Andrew Vagin wrote:
> What do you think about the idea to mount nsfs and be able to look up
> any alive namespace by inum:
>
>   $ tree .
>   .
>   ├── mnt{inum}
>   │   └── user -> ../user{inum}
>   ├── pid{inum}
>   │   ├── pid{inum}
>   │   │   └── user -> ../../user{inum}/user{inum}
>   │   └── user -> ../user{inum}
>   └── user{inum}
>       └── user{inum}
>
> https://lkml.org/lkml/2016/7/8/59
>
> I think it solves all requirements which were mentioned in this thread.

It may need an additional entry per directory for the bit you setns.
Maybe ‘handle’?

  $ tree .
  .
  ├── mnt{inum}
  │   ├── handle -> mnt:[{inum}]
  │   └── user -> ../user{inum}
  …

but that's not a major revision.

> On Fri, Jul 08, 2016 at 07:35:33AM -0700, James Bottomley wrote:
> > On Fri, 2016-07-08 at 02:44 -0500, Eric W. Biederman wrote:
> > > Starting with 4.8 we are also going to need to be able to
> > > retrieve the user namespace owner of filesystems.  That will be
> > > an interesting mix.
> >
> > This is per mount point, isn't it? so it can't be in /proc/fs/ and
> > it would have to be per local mount tree.  Yes, that is a bit
> > nasty.  Sounds like we might need to unfold mount or mountinfo
> > into something that has one directory per entry?
>
> If we will be able to look up namespaces in nsfs by inum, we can
> print an userns inum in mountinfo.

With the tree view you can find a namespace by inum (if it's one of
your descendants), but it's not going to be particularly efficient
(you'll have to walk the tree).  Folks that need to do that quickly
can index the tree (which would be fairly straightforward if the nsfs
mount supports inotify), but it would be nice to have a more elegant
solution for this use-case.

Cheers,
Trevor

-- 
This email may be signed or encrypted with GnuPG (http://www.gnupg.org).
For more information, see http://en.wikipedia.org/wiki/Pretty_Good_Privacy

[toc] | [prev] | [next] | [standalone]


#1439826 — Re: [CRIU] Introspecting userns relationships to other namespaces?

FromJames Bottomley <James.Bottomley@Hansenpartnership.com>
Date2016-07-09 00:30 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSMrw-2wN-21@gated-at.bofh.it>
In reply to#1439525
On July 8, 2016 1:38:19 PM PDT, Andrew Vagin <avagin@virtuozzo.com> wrote:
>On Fri, Jul 08, 2016 at 07:35:33AM -0700, James Bottomley wrote:
>> On Fri, 2016-07-08 at 02:44 -0500, Eric W. Biederman wrote:
>> > Andrew Vagin <avagin@virtuozzo.com> writes:
>> > 
>> > > On Wed, Jul 06, 2016 at 10:46:33AM -0500, Eric W. Biederman
>wrote:
>> > > > "Serge E. Hallyn" <serge@hallyn.com> writes:
>> > > > 
>> > > > > On Wed, Jul 06, 2016 at 10:41:48AM +0200, Michael Kerrisk
>(man
>> > > > > -pages) wrote:
>> > > > > > [Rats! Doing now what I should have down to start with.
>> > > > > > Looping some
>> > > > > > lists and CRIU and other possibly relevant people into this
>> > > > > > conversation]
>> > > > > > 
>> > > > > > Hi Eric,
>> > > > > > 
>> > > > > > On 5 July 2016 at 23:47, Eric W. Biederman <
>> > > > > > ebiederm@xmission.com> wrote:
>> > > > > > > "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
>> > > > > > > writes:
>> > > > > > > 
>> > > > > > > > Hi Eric,
>> > > > > > > > 
>> > > > > > > > I have a question. Is there any way currently to
>discover 
>> > > > > > > > which user namespace a particular nonuser namespace is 
>> > > > > > > > governed by? Maybe I am missing something, but there
>does 
>> > > > > > > > not seem to be a way to do this. Also, can one discover
>
>> > > > > > > > which userns is the parent of a given userns? Again, I 
>> > > > > > > > can't see a way to do this.
>> > > > > > > > 
>> > > > > > > > The point here is introspecting so that a process might
>
>> > > > > > > > determine what its capabilities are when operating on 
>> > > > > > > > some resource governed by a (nonuser) namespace.
>> > > > > > > 
>> > > > > > > To the best of my knowledge that there is not an
>interface 
>> > > > > > > to get that information.  It would be good to have such
>an 
>> > > > > > > interface for no other reason than the CRIU folks are
>going 
>> > > > > > > to need it at some point.  I am a bit surprised they have
>> > > > > > > not complained yet.
>> > > > > 
>> > > > > I don't think they need it.  They do in fact have what they 
>> > > > > need.  Assume you have tasks T1, T2, T1_1 and T2_1;  T1 and
>T2 
>> > > > > are in init_user_ns;  T1 spawned T1_1 in a new userns;  T2 
>> > > > > spawned T2_1 which setns()d to T1_1's ns. There's some
>> > > > > {handwave} uid mapping, does not matter.
>> > > > > 
>> > > > > At restart, it doesn't matter which task originally created
>the 
>> > > > > new userns. criu knows T1_1 and T2_1 are in the same userns; 
>> > > > >  it creates the userns, sets up the mapping, and T1_1 and
>T2_1
>> > > > > setns() to it.
>> > > > 
>> > > > Given that the simple cases are so easy it probably doesn't 
>> > > > matter in that sense.
>> > > > 
>> > > > However we now have the case where user namespaces own pid 
>> > > > namespaces, and uts namespaces, and network namespaces, and ipc
>
>> > > > namespaces, and filesystems.  Throw in some mount propagation
>and 
>> > > > use of setns and things could get confusing.   It is something 
>> > > > that will need to be figured out if CRIU is going to properly 
>> > > > checkpoint containers containing containers containing
>containers 
>> > > > containing containers.
>> > > 
>> > > It isn't a joke:). We have a few requests to support CR of 
>> > > containers with Docker containers inside. And we are going to
>start 
>> > > this task in a near future, so we would like to have interface to
>
>> > > get dependencies between namespaces too.
>> > > 
>> > > BTW: CRIU already supports nested mount namespaces, because
>systemd
>> > > creates them for services.
>> > 
>> > The tricky part about this and what messes up James proposed plan
>is
>> > that the interface needs to be something that returns a namespace 
>> > file descriptor.  So we can't print something out in a simple text
>> > file.
>> 
>> I actually described two problems: the first was how we get the
>> information in the first place.  Currently the owning or parent
>user_ns
>> is tucked inside an opaque structure.  I think we need to move that
>to
>> ns_common where it would be the owning userns for all non-user
>> namespaces and the parent for the userns.
>
>I'm agree with this.
>
>> 
>> Once we actually have the information, we can also add a set of proc
>> links, say either
>> 
>> /proc/<pid>/ns/X-userns
>> 
>> Which might be a bit messy since it doubles the number of files, or
>> perhaps in a simple directory.
>
>In this case we will need to enter into each namespace to build a full
>chain of dependencies.
>
>It's tricky, because if we enter into a child userns, we can't to enter
>into a parent userns from the same process, so to get the next branch,
>we will need to create a new process.
>
>				    process A
>					|
>init_user_ns->child_user_ns_1->child_userns_2
>
>fork() -> B
>  B: setns(/proc/A/ns/userns-parent)
>readlink(/proc/B/ns/userns)
>
>fork() -> C
>  C: setns(/proc/B/ns/userns-parent)
>readlink(/proc/C/ns/userns)
>
>
>> 
>> > Well I suppose we could print an device number and inode number
>pair.
>> > But then someone would still have to scour processes looking for a 
>> > user namespace so that is likely less than ideal.
>> 
>> There's no reason any of the proposed methods so far have to be
>> exclusive: nsfs.c has a lot of flexibility.
>
>
>What do you think about the idea to mount nsfs and be able to look up
>any alive namespace by inum:

I think I like it.  It will give us a way to enter any extant namespace.  It will work for Eric's fs namespaces as well.  Perhaps a /process/ns/<inum>
 Directory?

James

>  $ tree .
>  .
>  ├── mnt{inum}
>  │   └── user -> ../user{inum}
>  ├── pid{inum}
>  │   ├── pid{inum}
>  │   │   └── user -> ../../user{inum}/user{inum}
>  │   └── user -> ../user{inum}
>  └── user{inum}
>      └── user{inum}
>
>https://lkml.org/lkml/2016/7/8/59
>
>I think it solves all requirements which were mentioned in this thread.
>
>> 
>> > Starting with 4.8 we are also going to need to be able to retrieve 
>> > the user namespace owner of filesystems.  That will be an
>interesting
>> > mix.
>> 
>> This is per mount point, isn't it? so it can't be in /proc/fs/ and it
>> would have to be per local mount tree.  Yes, that is a bit nasty. 
>>  Sounds like we might need to unfold mount or mountinfo into
>something
>> that has one directory per entry?
>
>If we will be able to look up namespaces in nsfs by inum, we can print
>an userns inum in mountinfo.
>
>> 
>> James
>> 
>> > Eric
>> > 
>> > _______________________________________________
>> > Containers mailing list
>> > Containers@lists.linux-foundation.org
>> > https://lists.linuxfoundation.org/mailman/listinfo/containers
>> > 
>> 
>_______________________________________________
>Containers mailing list
>Containers@lists.linux-foundation.org
>https://lists.linuxfoundation.org/mailman/listinfo/containers


-- 
Sent from my Android device with K-9 Mail. Please excuse my brevity.

[toc] | [prev] | [next] | [standalone]


#1439829 — Re: [CRIU] Introspecting userns relationships to other namespaces?

FromJames Bottomley <James.Bottomley@Hansenpartnership.com>
Date2016-07-09 00:30 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSMrw-2wN-9@gated-at.bofh.it>
In reply to#1439525
On July 8, 2016 1:38:19 PM PDT, Andrew Vagin <avagin@virtuozzo.com> wrote:
>On Fri, Jul 08, 2016 at 07:35:33AM -0700, James Bottomley wrote:
>> On Fri, 2016-07-08 at 02:44 -0500, Eric W. Biederman wrote:
>> > Andrew Vagin <avagin@virtuozzo.com> writes:
>> > 
>> > > On Wed, Jul 06, 2016 at 10:46:33AM -0500, Eric W. Biederman
>wrote:
>> > > > "Serge E. Hallyn" <serge@hallyn.com> writes:
>> > > > 
>> > > > > On Wed, Jul 06, 2016 at 10:41:48AM +0200, Michael Kerrisk
>(man
>> > > > > -pages) wrote:
>> > > > > > [Rats! Doing now what I should have down to start with.
>> > > > > > Looping some
>> > > > > > lists and CRIU and other possibly relevant people into this
>> > > > > > conversation]
>> > > > > > 
>> > > > > > Hi Eric,
>> > > > > > 
>> > > > > > On 5 July 2016 at 23:47, Eric W. Biederman <
>> > > > > > ebiederm@xmission.com> wrote:
>> > > > > > > "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
>> > > > > > > writes:
>> > > > > > > 
>> > > > > > > > Hi Eric,
>> > > > > > > > 
>> > > > > > > > I have a question. Is there any way currently to
>discover 
>> > > > > > > > which user namespace a particular nonuser namespace is 
>> > > > > > > > governed by? Maybe I am missing something, but there
>does 
>> > > > > > > > not seem to be a way to do this. Also, can one discover
>
>> > > > > > > > which userns is the parent of a given userns? Again, I 
>> > > > > > > > can't see a way to do this.
>> > > > > > > > 
>> > > > > > > > The point here is introspecting so that a process might
>
>> > > > > > > > determine what its capabilities are when operating on 
>> > > > > > > > some resource governed by a (nonuser) namespace.
>> > > > > > > 
>> > > > > > > To the best of my knowledge that there is not an
>interface 
>> > > > > > > to get that information.  It would be good to have such
>an 
>> > > > > > > interface for no other reason than the CRIU folks are
>going 
>> > > > > > > to need it at some point.  I am a bit surprised they have
>> > > > > > > not complained yet.
>> > > > > 
>> > > > > I don't think they need it.  They do in fact have what they 
>> > > > > need.  Assume you have tasks T1, T2, T1_1 and T2_1;  T1 and
>T2 
>> > > > > are in init_user_ns;  T1 spawned T1_1 in a new userns;  T2 
>> > > > > spawned T2_1 which setns()d to T1_1's ns. There's some
>> > > > > {handwave} uid mapping, does not matter.
>> > > > > 
>> > > > > At restart, it doesn't matter which task originally created
>the 
>> > > > > new userns. criu knows T1_1 and T2_1 are in the same userns; 
>> > > > >  it creates the userns, sets up the mapping, and T1_1 and
>T2_1
>> > > > > setns() to it.
>> > > > 
>> > > > Given that the simple cases are so easy it probably doesn't 
>> > > > matter in that sense.
>> > > > 
>> > > > However we now have the case where user namespaces own pid 
>> > > > namespaces, and uts namespaces, and network namespaces, and ipc
>
>> > > > namespaces, and filesystems.  Throw in some mount propagation
>and 
>> > > > use of setns and things could get confusing.   It is something 
>> > > > that will need to be figured out if CRIU is going to properly 
>> > > > checkpoint containers containing containers containing
>containers 
>> > > > containing containers.
>> > > 
>> > > It isn't a joke:). We have a few requests to support CR of 
>> > > containers with Docker containers inside. And we are going to
>start 
>> > > this task in a near future, so we would like to have interface to
>
>> > > get dependencies between namespaces too.
>> > > 
>> > > BTW: CRIU already supports nested mount namespaces, because
>systemd
>> > > creates them for services.
>> > 
>> > The tricky part about this and what messes up James proposed plan
>is
>> > that the interface needs to be something that returns a namespace 
>> > file descriptor.  So we can't print something out in a simple text
>> > file.
>> 
>> I actually described two problems: the first was how we get the
>> information in the first place.  Currently the owning or parent
>user_ns
>> is tucked inside an opaque structure.  I think we need to move that
>to
>> ns_common where it would be the owning userns for all non-user
>> namespaces and the parent for the userns.
>
>I'm agree with this.
>
>> 
>> Once we actually have the information, we can also add a set of proc
>> links, say either
>> 
>> /proc/<pid>/ns/X-userns
>> 
>> Which might be a bit messy since it doubles the number of files, or
>> perhaps in a simple directory.
>
>In this case we will need to enter into each namespace to build a full
>chain of dependencies.
>
>It's tricky, because if we enter into a child userns, we can't to enter
>into a parent userns from the same process, so to get the next branch,
>we will need to create a new process.
>
>				    process A
>					|
>init_user_ns->child_user_ns_1->child_userns_2
>
>fork() -> B
>  B: setns(/proc/A/ns/userns-parent)
>readlink(/proc/B/ns/userns)
>
>fork() -> C
>  C: setns(/proc/B/ns/userns-parent)
>readlink(/proc/C/ns/userns)
>
>
>> 
>> > Well I suppose we could print an device number and inode number
>pair.
>> > But then someone would still have to scour processes looking for a 
>> > user namespace so that is likely less than ideal.
>> 
>> There's no reason any of the proposed methods so far have to be
>> exclusive: nsfs.c has a lot of flexibility.
>
>
>What do you think about the idea to mount nsfs and be able to look up
>any alive namespace by inum:

I think I like it.  It will give us a way to enter any extant namespace.  It will work for Eric's fs namespaces as well.  Perhaps a /process/ns/<inum>
 Directory?

James

>  $ tree .
>  .
>  ├── mnt{inum}
>  │   └── user -> ../user{inum}
>  ├── pid{inum}
>  │   ├── pid{inum}
>  │   │   └── user -> ../../user{inum}/user{inum}
>  │   └── user -> ../user{inum}
>  └── user{inum}
>      └── user{inum}
>
>https://lkml.org/lkml/2016/7/8/59
>
>I think it solves all requirements which were mentioned in this thread.
>
>> 
>> > Starting with 4.8 we are also going to need to be able to retrieve 
>> > the user namespace owner of filesystems.  That will be an
>interesting
>> > mix.
>> 
>> This is per mount point, isn't it? so it can't be in /proc/fs/ and it
>> would have to be per local mount tree.  Yes, that is a bit nasty. 
>>  Sounds like we might need to unfold mount or mountinfo into
>something
>> that has one directory per entry?
>
>If we will be able to look up namespaces in nsfs by inum, we can print
>an userns inum in mountinfo.
>
>> 
>> James
>> 
>> > Eric
>> > 
>> > _______________________________________________
>> > Containers mailing list
>> > Containers@lists.linux-foundation.org
>> > https://lists.linuxfoundation.org/mailman/listinfo/containers
>> > 
>> 
>_______________________________________________
>Containers mailing list
>Containers@lists.linux-foundation.org
>https://lists.linuxfoundation.org/mailman/listinfo/containers


-- 
Sent from my Android device with K-9 Mail. Please excuse my brevity.

[toc] | [prev] | [next] | [standalone]


#1439853 — Re: [CRIU] Introspecting userns relationships to other namespaces?

Fromebiederm@xmission.com (Eric W. Biederman)
Date2016-07-09 02:10 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSO0h-3Dw-3@gated-at.bofh.it>
In reply to#1439829
James Bottomley <James.Bottomley@Hansenpartnership.com> writes:

> On July 8, 2016 1:38:19 PM PDT, Andrew Vagin <avagin@virtuozzo.com> wrote:

>>What do you think about the idea to mount nsfs and be able to look up
>>any alive namespace by inum:
>
> I think I like it.  It will give us a way to enter any extant
> namespace.  It will work for Eric's fs namespaces as well.  Perhaps a
> /process/ns/<inum> Directory?

*Shivers*

That makes it very easy to bypass any existing controls that exist for
getting at namespaces.  It is true that everything of that kind is
directory based but still.

Plus I think it would serve as information leak to information outside
of the container.

An operation to get a user namespace file descriptor from some kernel
object sounds reasonably sane.

A great big list of things sounds about as scary as it can get.  This is
not the time to be making it easier to escape from containers.

Eric

[toc] | [prev] | [next] | [standalone]


#1439854 — Re: [CRIU] Introspecting userns relationships to other namespaces?

FromJames Bottomley <James.Bottomley@HansenPartnership.com>
Date2016-07-09 02:20 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSO9X-3Hn-3@gated-at.bofh.it>
In reply to#1439853
On Fri, 2016-07-08 at 18:52 -0500, Eric W. Biederman wrote:
> James Bottomley <James.Bottomley@Hansenpartnership.com> writes:
> 
> > On July 8, 2016 1:38:19 PM PDT, Andrew Vagin <avagin@virtuozzo.com>
> > wrote:
> 
> > > What do you think about the idea to mount nsfs and be able to 
> > > look up any alive namespace by inum:
> > 
> > I think I like it.  It will give us a way to enter any extant
> > namespace.  It will work for Eric's fs namespaces as well.  Perhaps 
> > a /process/ns/<inum> Directory?

As you understood, I meant /proc/ns/<inum> (damn mobile phone
completions).

> *Shivers*
> 
> That makes it very easy to bypass any existing controls that exist 
> for getting at namespaces.  It is true that everything of that kind 
> is directory based but still.
> 
> Plus I think it would serve as information leak to information 
> outside of the container.
> 
> An operation to get a user namespace file descriptor from some kernel
> object sounds reasonably sane.
> 
> A great big list of things sounds about as scary as it can get.  This 
> is not the time to be making it easier to escape from containers.

To be honest, I think this argument is rubbish.  If we're afraid of
giving out a list of all the namespaces, it means we're afraid there's
some security bug and we're trying to obscure it by making the list
hard to get.  All we've done is allayed fears about the bug but the
hackers still know the portals to get through.

If such a bug exists, it will be possible to exploit it by simply
reconstructing the information from the individual process directories,
so obscurity doesn't protect us and all it does is give us a false
sense of security.   If such a bug doesn't exist, then all the security
mechanisms currently in place (like no re-entry to prior namespace)
should protect us and we can give out the list.

Let's deal with the world as we'd like it to be (no obscure namespace
bugs) and accept the consequences and the responsibility for fixing
them if we turn out to be slightly incorrect.  We'll end up in a far
better place than security by obscurity would land us.

James

[toc] | [prev] | [next] | [standalone]


#1439881 — Re: [CRIU] Introspecting userns relationships to other namespaces?

Fromebiederm@xmission.com (Eric W. Biederman)
Date2016-07-09 05:20 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSQY9-5CR-1@gated-at.bofh.it>
In reply to#1439854
James Bottomley <James.Bottomley@HansenPartnership.com> writes:

> On Fri, 2016-07-08 at 18:52 -0500, Eric W. Biederman wrote:
>> James Bottomley <James.Bottomley@Hansenpartnership.com> writes:
>> 
>> > On July 8, 2016 1:38:19 PM PDT, Andrew Vagin <avagin@virtuozzo.com>
>> > wrote:
>> 
>> > > What do you think about the idea to mount nsfs and be able to 
>> > > look up any alive namespace by inum:
>> > 
>> > I think I like it.  It will give us a way to enter any extant
>> > namespace.  It will work for Eric's fs namespaces as well.  Perhaps 
>> > a /process/ns/<inum> Directory?
>
> As you understood, I meant /proc/ns/<inum> (damn mobile phone
> completions).
>
>> *Shivers*
>> 
>> That makes it very easy to bypass any existing controls that exist 
>> for getting at namespaces.  It is true that everything of that kind 
>> is directory based but still.
>> 
>> Plus I think it would serve as information leak to information 
>> outside of the container.
>> 
>> An operation to get a user namespace file descriptor from some kernel
>> object sounds reasonably sane.
>> 
>> A great big list of things sounds about as scary as it can get.  This 
>> is not the time to be making it easier to escape from containers.
>
> To be honest, I think this argument is rubbish.  If we're afraid of
> giving out a list of all the namespaces, it means we're afraid there's
> some security bug and we're trying to obscure it by making the list
> hard to get.  All we've done is allayed fears about the bug but the
> hackers still know the portals to get through.
>
> If such a bug exists, it will be possible to exploit it by simply
> reconstructing the information from the individual process directories,
> so obscurity doesn't protect us and all it does is give us a false
> sense of security.   If such a bug doesn't exist, then all the security
> mechanisms currently in place (like no re-entry to prior namespace)
> should protect us and we can give out the list.
>
> Let's deal with the world as we'd like it to be (no obscure namespace
> bugs) and accept the consequences and the responsibility for fixing
> them if we turn out to be slightly incorrect.  We'll end up in a far
> better place than security by obscurity would land us.

No.  That is not the fear.  The permission checks on /proc/self/ns/xxx
are different than if the namespace is bind mounted somewhere.

That was done deliberately and with a reasonable amount of forethought.
You are asking to throw those permission checks out.   The answer is no.

Furthermore there is a much clearer reason not to go with a list of all
namespaces. A list of all namespaces breaks CRIU.  As you have described
it the list will change depending upon which machine you restore a
checkpoint on.  I honestly don't know what kind of havoc that will cause
but it is certainly something we won't be able to checkpoint no matter
how hard we try.

A global list of namespaces especially of the kind that you can open
and get a handle to the namespace is just not appropriate.

I know inode numbers comes darn close to names but they aren't really
names and if it comes to it we can figure out how to preserve an
applications view of it all across a checkpoint/restart.  So far it
hasn't proven necessary to preserve any inode numbers across
checkpoint/restart but again it is theoretically possible if it becomes
necessary.

Throwing away checkpoint/restart support for the sake of
checkpoint/restart is a no-go.

Containers fundamentally imply you don't have global visibility,
and that is a good thing.

Eric

[toc] | [prev] | [next] | [standalone]


#1439937 — Re: [CRIU] Introspecting userns relationships to other namespaces?

FromJames Bottomley <James.Bottomley@Hansenpartnership.com>
Date2016-07-09 12:40 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSXPX-1Kv-3@gated-at.bofh.it>
In reply to#1439881
On July 9, 2016 4:26:28 PM GMT+09:00, Andrew Vagin <avagin@virtuozzo.com> wrote:
>On Fri, Jul 08, 2016 at 10:05:18PM -0500, Eric W. Biederman wrote:
>> James Bottomley <James.Bottomley@HansenPartnership.com> writes:
>> 
>> > On Fri, 2016-07-08 at 18:52 -0500, Eric W. Biederman wrote:
>> >> James Bottomley <James.Bottomley@Hansenpartnership.com> writes:
>> >> 
>> >> > On July 8, 2016 1:38:19 PM PDT, Andrew Vagin
><avagin@virtuozzo.com>
>> >> > wrote:
>> >> 
>> >> > > What do you think about the idea to mount nsfs and be able to 
>> >> > > look up any alive namespace by inum:
>> >> > 
>> >> > I think I like it.  It will give us a way to enter any extant
>> >> > namespace.  It will work for Eric's fs namespaces as well. 
>Perhaps 
>> >> > a /process/ns/<inum> Directory?
>> >
>> > As you understood, I meant /proc/ns/<inum> (damn mobile phone
>> > completions).
>> >
>> >> *Shivers*
>> >> 
>> >> That makes it very easy to bypass any existing controls that exist
>
>> >> for getting at namespaces.  It is true that everything of that
>kind 
>> >> is directory based but still.
>> >> 
>> >> Plus I think it would serve as information leak to information 
>> >> outside of the container.
>> >> 
>> >> An operation to get a user namespace file descriptor from some
>kernel
>> >> object sounds reasonably sane.
>> >> 
>> >> A great big list of things sounds about as scary as it can get. 
>This 
>> >> is not the time to be making it easier to escape from containers.
>> >
>> > To be honest, I think this argument is rubbish.  If we're afraid of
>> > giving out a list of all the namespaces, it means we're afraid
>there's
>> > some security bug and we're trying to obscure it by making the list
>> > hard to get.  All we've done is allayed fears about the bug but the
>> > hackers still know the portals to get through.
>> >
>> > If such a bug exists, it will be possible to exploit it by simply
>> > reconstructing the information from the individual process
>directories,
>> > so obscurity doesn't protect us and all it does is give us a false
>> > sense of security.   If such a bug doesn't exist, then all the
>security
>> > mechanisms currently in place (like no re-entry to prior namespace)
>> > should protect us and we can give out the list.
>> >
>> > Let's deal with the world as we'd like it to be (no obscure
>namespace
>> > bugs) and accept the consequences and the responsibility for fixing
>> > them if we turn out to be slightly incorrect.  We'll end up in a
>far
>> > better place than security by obscurity would land us.
>> 
>> No.  That is not the fear.  The permission checks on
>/proc/self/ns/xxx
>> are different than if the namespace is bind mounted somewhere.
>> 
>> That was done deliberately and with a reasonable amount of
>forethought.
>> You are asking to throw those permission checks out.   The answer is
>no.
>> 
>> Furthermore there is a much clearer reason not to go with a list of
>all
>> namespaces. A list of all namespaces breaks CRIU.  As you have
>described
>> it the list will change depending upon which machine you restore a
>> checkpoint on.  I honestly don't know what kind of havoc that will
>cause
>> but it is certainly something we won't be able to checkpoint no
>matter
>> how hard we try.
>
>It's right. I hadn't thought about this.

Me neither.  Sorry for the prior outburst.

I think this means we're back to exposing owning userns in the /proc /<pid >/ns directory. 

>> 
>> A global list of namespaces especially of the kind that you can open
>> and get a handle to the namespace is just not appropriate.
>> 
>> I know inode numbers comes darn close to names but they aren't really
>> names and if it comes to it we can figure out how to preserve an
>> applications view of it all across a checkpoint/restart.  So far it
>> hasn't proven necessary to preserve any inode numbers across
>> checkpoint/restart but again it is theoretically possible if it
>becomes
>> necessary.
>> 
>> Throwing away checkpoint/restart support for the sake of
>> checkpoint/restart is a no-go.
>> 
>> Containers fundamentally imply you don't have global visibility,
>> and that is a good thing.
>
>All these thoughts about security make me thinking that kcmp is what we
>should use here. It's maybe something like this:
>
>kcmp(pid1, pid2, KCMP_NS_USERNS, fd1, fd2)
>
>- to check if userns of the fd1 namepsace is equal to the fd2 userns
>
>kcmp(pid1, pid2, KCMP_NS_PARENT, fd1, fd2)
>
>- to check if a parent namespace of the fd1 pidns is equal to fd pidns.
>
>fd1 and fd2 is file descriptors to namespace files.
>
>So if we want to build a hierarchy, we need to collect all namespaces
>and then enumerate them to check dependencies with help of kcmp.

Sure, but we need a method for opening the filehandles first .. .

James 

>> 
>> Eric


-- 
Sent from my Android device with K-9 Mail. Please excuse my brevity.

[toc] | [prev] | [next] | [standalone]


#1439938 — Re: [CRIU] Introspecting userns relationships to other namespaces?

FromJames Bottomley <James.Bottomley@Hansenpartnership.com>
Date2016-07-09 12:40 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSXPX-1Kv-9@gated-at.bofh.it>
In reply to#1439881
On July 9, 2016 4:26:28 PM GMT+09:00, Andrew Vagin <avagin@virtuozzo.com> wrote:
>On Fri, Jul 08, 2016 at 10:05:18PM -0500, Eric W. Biederman wrote:
>> James Bottomley <James.Bottomley@HansenPartnership.com> writes:
>> 
>> > On Fri, 2016-07-08 at 18:52 -0500, Eric W. Biederman wrote:
>> >> James Bottomley <James.Bottomley@Hansenpartnership.com> writes:
>> >> 
>> >> > On July 8, 2016 1:38:19 PM PDT, Andrew Vagin
><avagin@virtuozzo.com>
>> >> > wrote:
>> >> 
>> >> > > What do you think about the idea to mount nsfs and be able to 
>> >> > > look up any alive namespace by inum:
>> >> > 
>> >> > I think I like it.  It will give us a way to enter any extant
>> >> > namespace.  It will work for Eric's fs namespaces as well. 
>Perhaps 
>> >> > a /process/ns/<inum> Directory?
>> >
>> > As you understood, I meant /proc/ns/<inum> (damn mobile phone
>> > completions).
>> >
>> >> *Shivers*
>> >> 
>> >> That makes it very easy to bypass any existing controls that exist
>
>> >> for getting at namespaces.  It is true that everything of that
>kind 
>> >> is directory based but still.
>> >> 
>> >> Plus I think it would serve as information leak to information 
>> >> outside of the container.
>> >> 
>> >> An operation to get a user namespace file descriptor from some
>kernel
>> >> object sounds reasonably sane.
>> >> 
>> >> A great big list of things sounds about as scary as it can get. 
>This 
>> >> is not the time to be making it easier to escape from containers.
>> >
>> > To be honest, I think this argument is rubbish.  If we're afraid of
>> > giving out a list of all the namespaces, it means we're afraid
>there's
>> > some security bug and we're trying to obscure it by making the list
>> > hard to get.  All we've done is allayed fears about the bug but the
>> > hackers still know the portals to get through.
>> >
>> > If such a bug exists, it will be possible to exploit it by simply
>> > reconstructing the information from the individual process
>directories,
>> > so obscurity doesn't protect us and all it does is give us a false
>> > sense of security.   If such a bug doesn't exist, then all the
>security
>> > mechanisms currently in place (like no re-entry to prior namespace)
>> > should protect us and we can give out the list.
>> >
>> > Let's deal with the world as we'd like it to be (no obscure
>namespace
>> > bugs) and accept the consequences and the responsibility for fixing
>> > them if we turn out to be slightly incorrect.  We'll end up in a
>far
>> > better place than security by obscurity would land us.
>> 
>> No.  That is not the fear.  The permission checks on
>/proc/self/ns/xxx
>> are different than if the namespace is bind mounted somewhere.
>> 
>> That was done deliberately and with a reasonable amount of
>forethought.
>> You are asking to throw those permission checks out.   The answer is
>no.
>> 
>> Furthermore there is a much clearer reason not to go with a list of
>all
>> namespaces. A list of all namespaces breaks CRIU.  As you have
>described
>> it the list will change depending upon which machine you restore a
>> checkpoint on.  I honestly don't know what kind of havoc that will
>cause
>> but it is certainly something we won't be able to checkpoint no
>matter
>> how hard we try.
>
>It's right. I hadn't thought about this.

Me neither.  Sorry for the prior outburst.

I think this means we're back to exposing owning userns in the /proc /<pid >/ns directory. 

>> 
>> A global list of namespaces especially of the kind that you can open
>> and get a handle to the namespace is just not appropriate.
>> 
>> I know inode numbers comes darn close to names but they aren't really
>> names and if it comes to it we can figure out how to preserve an
>> applications view of it all across a checkpoint/restart.  So far it
>> hasn't proven necessary to preserve any inode numbers across
>> checkpoint/restart but again it is theoretically possible if it
>becomes
>> necessary.
>> 
>> Throwing away checkpoint/restart support for the sake of
>> checkpoint/restart is a no-go.
>> 
>> Containers fundamentally imply you don't have global visibility,
>> and that is a good thing.
>
>All these thoughts about security make me thinking that kcmp is what we
>should use here. It's maybe something like this:
>
>kcmp(pid1, pid2, KCMP_NS_USERNS, fd1, fd2)
>
>- to check if userns of the fd1 namepsace is equal to the fd2 userns
>
>kcmp(pid1, pid2, KCMP_NS_PARENT, fd1, fd2)
>
>- to check if a parent namespace of the fd1 pidns is equal to fd pidns.
>
>fd1 and fd2 is file descriptors to namespace files.
>
>So if we want to build a hierarchy, we need to collect all namespaces
>and then enumerate them to check dependencies with help of kcmp.

Sure, but we need a method for opening the filehandles first .. .

James 

>> 
>> Eric


-- 
Sent from my Android device with K-9 Mail. Please excuse my brevity.

[toc] | [prev] | [next] | [standalone]


#1439978 — Re: [CRIU] Introspecting userns relationships to other namespaces?

Fromebiederm@xmission.com (Eric W. Biederman)
Date2016-07-09 20:30 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rT5aO-6vN-3@gated-at.bofh.it>
In reply to#1439881
Andrew Vagin <avagin@virtuozzo.com> writes:

> All these thoughts about security make me thinking that kcmp is what we
> should use here. It's maybe something like this:
>
> kcmp(pid1, pid2, KCMP_NS_USERNS, fd1, fd2)
>
> - to check if userns of the fd1 namepsace is equal to the fd2 userns
>
> kcmp(pid1, pid2, KCMP_NS_PARENT, fd1, fd2)
>
> - to check if a parent namespace of the fd1 pidns is equal to fd pidns.
>
> fd1 and fd2 is file descriptors to namespace files.
>
> So if we want to build a hierarchy, we need to collect all namespaces
> and then enumerate them to check dependencies with help of kcmp.

That is certainly one way to go.

There is a funny case where we would want to compare a user namespace
file descriptor to a parent user namespace file descriptor.


Grumble, Grumble.  I think this may actually a case for creating ioctls
for these two cases.  Now that random nsfs file descriptors are bind
mountable the original reason for using proc files is not as pressing.

One ioctl for the user namespace that owns a file descriptor.
One ioctl for the parent namespace of a namespace file descriptor.

We also need some way to get a command file descriptor for a file system
super block.  Al Viro has a pet project for cleaning up the mount API
and this might be the idea excuse to start looking at that.

(In principle we might be able to run commands through the namespace
 file descriptor and using an ioctl feels dirty.  But an ioctl that
 only uses the fd and request argument does not suffer from the same
 problems that ioctls that have to pass additional arguments suffer
 from.)

Eric

[toc] | [prev] | [next] | [standalone]


#1439979 — Re: [CRIU] Introspecting userns relationships to other namespaces?

Fromebiederm@xmission.com (Eric W. Biederman)
Date2016-07-09 20:50 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rT5ua-6CL-19@gated-at.bofh.it>
In reply to#1439978
ebiederm@xmission.com (Eric W. Biederman) writes:

> Andrew Vagin <avagin@virtuozzo.com> writes:
>
>> All these thoughts about security make me thinking that kcmp is what we
>> should use here. It's maybe something like this:
>>
>> kcmp(pid1, pid2, KCMP_NS_USERNS, fd1, fd2)
>>
>> - to check if userns of the fd1 namepsace is equal to the fd2 userns
>>
>> kcmp(pid1, pid2, KCMP_NS_PARENT, fd1, fd2)
>>
>> - to check if a parent namespace of the fd1 pidns is equal to fd pidns.
>>
>> fd1 and fd2 is file descriptors to namespace files.
>>
>> So if we want to build a hierarchy, we need to collect all namespaces
>> and then enumerate them to check dependencies with help of kcmp.
>
> That is certainly one way to go.
>
> There is a funny case where we would want to compare a user namespace
> file descriptor to a parent user namespace file descriptor.
>
>
> Grumble, Grumble.  I think this may actually a case for creating ioctls
> for these two cases.  Now that random nsfs file descriptors are bind
> mountable the original reason for using proc files is not as pressing.
>
> One ioctl for the user namespace that owns a file descriptor.
> One ioctl for the parent namespace of a namespace file descriptor.
>
> We also need some way to get a command file descriptor for a file system
> super block.  Al Viro has a pet project for cleaning up the mount API
> and this might be the idea excuse to start looking at that.
>
> (In principle we might be able to run commands through the namespace
>  file descriptor and using an ioctl feels dirty.  But an ioctl that
>  only uses the fd and request argument does not suffer from the same
>  problems that ioctls that have to pass additional arguments suffer
>  from.)

Of course it should be an error perhaps -EINVAL to get a user
namespace owner or parent namespace that is outside of a processes
current user namespace or pid namespace.  That way thing stay bounded
within the current namespaces the process is in.  Which prevents any
leak possibilities, and keeps CRIU working.

Eric

[toc] | [prev] | [next] | [standalone]


#1438314

From"Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
Date2016-07-07 10:20 +0200
Message-ID<rScHp-4lY-59@gated-at.bofh.it>
In reply to#1437759
Hi Serge,

On 6 July 2016 at 16:13, Serge E. Hallyn <serge@hallyn.com> wrote:
> On Wed, Jul 06, 2016 at 10:41:48AM +0200, Michael Kerrisk (man-pages) wrote:
>> [Rats! Doing now what I should have down to start with. Looping some
>> lists and CRIU and other possibly relevant people into this
>> conversation]
>>
>> Hi Eric,
>>
>> On 5 July 2016 at 23:47, Eric W. Biederman <ebiederm@xmission.com> wrote:
>> > "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> writes:
>> >
>> >> Hi Eric,
>> >>
>> >> I have a question. Is there any way currently to discover which
>> >> user namespace a particular nonuser namespace is governed by?
>> >> Maybe I am missing something, but there does not seem to be a
>> >> way to do this. Also, can one discover which userns is the
>> >> parent of a given userns? Again, I can't see a way to do this.
>> >>
>> >> The point here is introspecting so that a process might determine
>> >> what its capabilities are when operating on some resource governed
>> >> by a (nonuser) namespace.
>> >
>> > To the best of my knowledge that there is not an interface to get that
>> > information.  It would be good to have such an interface for no other
>> > reason than the CRIU folks are going to need it at some point.  I am a
>> > bit surprised they have not complained yet.
>
> I don't think they need it.  They do in fact have what they need.  Assume
> you have tasks T1, T2, T1_1 and T2_1;  T1 and T2 are in init_user_ns;  T1
> spawned T1_1 in a new userns;  T2 spawned T2_1 which setns()d to T1_1's ns.
> There's some {handwave} uid mapping, does not matter.
>
> At restart, it doesn't matter which task originally created the new userns.
> criu knows T1_1 and T2_1 are in the same userns;  it creates the userns, sets
> up the mapping, and T1_1 and T2_1 setns() to it.

I'm missing something here. How does the parental relationships
between the user namespaces get reconstructed? Those relationships
will govern what capabilities a process will have in various user
namespaces.

Cheers,

Michael


-- 
Michael Kerrisk
Linux man-pages maintainer; http://www.kernel.org/doc/man-pages/
Linux/UNIX System Programming Training: http://man7.org/training/

[toc] | [prev] | [next] | [standalone]


#1438629

From"Serge E. Hallyn" <serge@hallyn.com>
Date2016-07-07 15:40 +0200
Message-ID<rShH4-7v8-43@gated-at.bofh.it>
In reply to#1438314
Quoting Michael Kerrisk (man-pages) (mtk.manpages@gmail.com):
> Hi Serge,
> 
> On 6 July 2016 at 16:13, Serge E. Hallyn <serge@hallyn.com> wrote:
> > On Wed, Jul 06, 2016 at 10:41:48AM +0200, Michael Kerrisk (man-pages) wrote:
> >> [Rats! Doing now what I should have down to start with. Looping some
> >> lists and CRIU and other possibly relevant people into this
> >> conversation]
> >>
> >> Hi Eric,
> >>
> >> On 5 July 2016 at 23:47, Eric W. Biederman <ebiederm@xmission.com> wrote:
> >> > "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> writes:
> >> >
> >> >> Hi Eric,
> >> >>
> >> >> I have a question. Is there any way currently to discover which
> >> >> user namespace a particular nonuser namespace is governed by?
> >> >> Maybe I am missing something, but there does not seem to be a
> >> >> way to do this. Also, can one discover which userns is the
> >> >> parent of a given userns? Again, I can't see a way to do this.
> >> >>
> >> >> The point here is introspecting so that a process might determine
> >> >> what its capabilities are when operating on some resource governed
> >> >> by a (nonuser) namespace.
> >> >
> >> > To the best of my knowledge that there is not an interface to get that
> >> > information.  It would be good to have such an interface for no other
> >> > reason than the CRIU folks are going to need it at some point.  I am a
> >> > bit surprised they have not complained yet.
> >
> > I don't think they need it.  They do in fact have what they need.  Assume
> > you have tasks T1, T2, T1_1 and T2_1;  T1 and T2 are in init_user_ns;  T1
> > spawned T1_1 in a new userns;  T2 spawned T2_1 which setns()d to T1_1's ns.
> > There's some {handwave} uid mapping, does not matter.
> >
> > At restart, it doesn't matter which task originally created the new userns.
> > criu knows T1_1 and T2_1 are in the same userns;  it creates the userns, sets
> > up the mapping, and T1_1 and T2_1 setns() to it.
> 
> I'm missing something here. How does the parental relationships
> between the user namespaces get reconstructed? Those relationships
> will govern what capabilities a process will have in various user
> namespaces.

Hm.  Probably best-effort based on the process hierarchy.  So yeah you
could probably get a tree into a state that would be wrongly recreated.
Create a new netns, bind mount it, exit;  Have another task create a
new user_ns, bind mount it, exit;  Third task setns()s first to the new
netns then to the new user_ns.  I suspect criu will recreate that
wrongly.

[toc] | [prev] | [next] | [standalone]


#1438666

FromJames Bottomley <James.Bottomley@HansenPartnership.com>
Date2016-07-07 17:10 +0200
Message-ID<rSj6a-cB-11@gated-at.bofh.it>
In reply to#1438629
On Thu, 2016-07-07 at 08:36 -0500, Serge E. Hallyn wrote:
> Quoting Michael Kerrisk (man-pages) (mtk.manpages@gmail.com):
> > Hi Serge,
> > 
> > On 6 July 2016 at 16:13, Serge E. Hallyn <serge@hallyn.com> wrote:
> > > On Wed, Jul 06, 2016 at 10:41:48AM +0200, Michael Kerrisk (man
> > > -pages) wrote:
> > > > [Rats! Doing now what I should have down to start with. Looping 
> > > > some lists and CRIU and other possibly relevant people into 
> > > > this conversation]
> > > > 
> > > > Hi Eric,
> > > > 
> > > > On 5 July 2016 at 23:47, Eric W. Biederman <
> > > > ebiederm@xmission.com> wrote:
> > > > > "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
> > > > > writes:
> > > > > 
> > > > > > Hi Eric,
> > > > > > 
> > > > > > I have a question. Is there any way currently to discover 
> > > > > > which user namespace a particular nonuser namespace is 
> > > > > > governed by? Maybe I am missing something, but there does 
> > > > > > not seem to be a way to do this. Also, can one discover 
> > > > > > which userns is the parent of a given userns? Again, I 
> > > > > > can't see a way to do this.
> > > > > > 
> > > > > > The point here is introspecting so that a process might 
> > > > > > determine what its capabilities are when operating on some 
> > > > > > resource governed by a (nonuser) namespace.
> > > > > 
> > > > > To the best of my knowledge that there is not an interface to 
> > > > > get that information.  It would be good to have such an 
> > > > > interface for no other reason than the CRIU folks are going 
> > > > > to need it at some point.  I am a bit surprised they have not
> > > > > complained yet.
> > > 
> > > I don't think they need it.  They do in fact have what they need.
> > >   Assume you have tasks T1, T2, T1_1 and T2_1;  T1 and T2 are in
> > > init_user_ns;  T1 spawned T1_1 in a new userns;  T2 spawned T2_1 
> > > which setns()d to T1_1's ns. There's some {handwave} uid mapping,
> > > does not matter.
> > > 
> > > At restart, it doesn't matter which task originally created the 
> > > new userns. criu knows T1_1 and T2_1 are in the same userns;  it 
> > > creates the userns, sets up the mapping, and T1_1 and T2_1
> > > setns() to it.
> > 
> > I'm missing something here. How does the parental relationships
> > between the user namespaces get reconstructed? Those relationships
> > will govern what capabilities a process will have in various user
> > namespaces.

Actually, you get the parent namespace from the process tree by
tracking the user namespaces of the parent pids.  Currently non-root
users can't bind the namespace, so the only way to keep a new user_ns
around if you're not root is to keep the process around, so for
multiply nested user namespaces you can usually build the user_ns
hierarchy by looking at the process hierarchy.  Conversely, if the
process is reparented to init, chances are that the user_ns is also
parented to init_user_ns.

> Hm.  Probably best-effort based on the process hierarchy.  So yeah
> you could probably get a tree into a state that would be wrongly
> recreated. Create a new netns, bind mount it, exit;  Have another 
> task create a new user_ns, bind mount it, exit;  Third task setns()s 
> first to the new netns then to the new user_ns.  I suspect criu will 
> recreate that wrongly.

This is a bit pathological, and you have to be root to do it: so root
can set up a nesting hierarchy, bind it and destroy the pids but I know
of no current orchestration system which does this.

Actually, I have to back pedal a bit: the way I currently set up
architecture emulation containers does precisely this: I set up the
namespaces unprivileged with child mount namespaces, but then I ask
root to bind the userns and kill the process that created it so I have
a permanent handle to enter the namespace by, so I suspect that when
our current orchestration systems get more sophisticated, they might
eventually want to do something like this as well.

In theory, we could get nsfs to show this information as an option
(just add a show_options entry to the superblock ops), but the problem
is that although each namespace has a parent user_ns, there's no way to
get it without digging in the namespace specific structure.  Probably
we should restructure to move it into ns_common, then we could display
it (and enforce all namespaces having owning user_ns) but it would be a
reasonably large (but mechanical) change.

James

[toc] | [prev] | [next] | [standalone]


#1438789

From"Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
Date2016-07-07 20:30 +0200
Message-ID<rSmdH-275-7@gated-at.bofh.it>
In reply to#1438666
On 7 July 2016 at 17:01, James Bottomley
<James.Bottomley@hansenpartnership.com> wrote:
> On Thu, 2016-07-07 at 08:36 -0500, Serge E. Hallyn wrote:
>> Quoting Michael Kerrisk (man-pages) (mtk.manpages@gmail.com):
>> > Hi Serge,
>> >
>> > On 6 July 2016 at 16:13, Serge E. Hallyn <serge@hallyn.com> wrote:
>> > > On Wed, Jul 06, 2016 at 10:41:48AM +0200, Michael Kerrisk (man
>> > > -pages) wrote:
>> > > > [Rats! Doing now what I should have down to start with. Looping
>> > > > some lists and CRIU and other possibly relevant people into
>> > > > this conversation]
>> > > >
>> > > > Hi Eric,
>> > > >
>> > > > On 5 July 2016 at 23:47, Eric W. Biederman <
>> > > > ebiederm@xmission.com> wrote:
>> > > > > "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
>> > > > > writes:
>> > > > >
>> > > > > > Hi Eric,
>> > > > > >
>> > > > > > I have a question. Is there any way currently to discover
>> > > > > > which user namespace a particular nonuser namespace is
>> > > > > > governed by? Maybe I am missing something, but there does
>> > > > > > not seem to be a way to do this. Also, can one discover
>> > > > > > which userns is the parent of a given userns? Again, I
>> > > > > > can't see a way to do this.
>> > > > > >
>> > > > > > The point here is introspecting so that a process might
>> > > > > > determine what its capabilities are when operating on some
>> > > > > > resource governed by a (nonuser) namespace.
>> > > > >
>> > > > > To the best of my knowledge that there is not an interface to
>> > > > > get that information.  It would be good to have such an
>> > > > > interface for no other reason than the CRIU folks are going
>> > > > > to need it at some point.  I am a bit surprised they have not
>> > > > > complained yet.
>> > >
>> > > I don't think they need it.  They do in fact have what they need.
>> > >   Assume you have tasks T1, T2, T1_1 and T2_1;  T1 and T2 are in
>> > > init_user_ns;  T1 spawned T1_1 in a new userns;  T2 spawned T2_1
>> > > which setns()d to T1_1's ns. There's some {handwave} uid mapping,
>> > > does not matter.
>> > >
>> > > At restart, it doesn't matter which task originally created the
>> > > new userns. criu knows T1_1 and T2_1 are in the same userns;  it
>> > > creates the userns, sets up the mapping, and T1_1 and T2_1
>> > > setns() to it.
>> >
>> > I'm missing something here. How does the parental relationships
>> > between the user namespaces get reconstructed? Those relationships
>> > will govern what capabilities a process will have in various user
>> > namespaces.
>
> Actually, you get the parent namespace from the process tree by
> tracking the user namespaces of the parent pids.   Currently non-root
> users can't bind the namespace, so the only way to keep a new user_ns
> around if you're not root is to keep the process around, so for
> multiply nested user namespaces you can usually build the user_ns
> hierarchy by looking at the process hierarchy.  Conversely, if the
> process is reparented to init, chances are that the user_ns is also
> parented to init_user_ns.

Yes, but "chances are" == this isn't robust.  PR_SET_CHILD_SUBREAPER
further complicates things.

By the way, is that really what happens? Do child user namespaces get
reparented to the grandparent ns if the parent ns disappears (i.e.,
ceases to have any members and no bind mounts)? I hadn't thought about
that scenario before. It may be worth documenting in
user_namespaces(7).

>> Hm.  Probably best-effort based on the process hierarchy.  So yeah
>> you could probably get a tree into a state that would be wrongly
>> recreated. Create a new netns, bind mount it, exit;  Have another
>> task create a new user_ns, bind mount it, exit;  Third task setns()s
>> first to the new netns then to the new user_ns.  I suspect criu will
>> recreate that wrongly.
>
> This is a bit pathological, and you have to be root to do it: so root
> can set up a nesting hierarchy, bind it and destroy the pids but I know
> of no current orchestration system which does this.
>
> Actually, I have to back pedal a bit: the way I currently set up
> architecture emulation containers does precisely this: I set up the
> namespaces unprivileged with child mount namespaces, but then I ask
> root to bind the userns and kill the process that created it so I have
> a permanent handle to enter the namespace by, so I suspect that when
> our current orchestration systems get more sophisticated, they might
> eventually want to do something like this as well.
>
> In theory, we could get nsfs to show this information as an option
> (just add a show_options entry to the superblock ops), but the problem
> is that although each namespace has a parent user_ns, there's no way to
> get it without digging in the namespace specific structure.  Probably
> we should restructure to move it into ns_common, then we could display
> it (and enforce all namespaces having owning user_ns) but it would be a

I'm missing something here. Is it not already the case that all
namespaces have an owning user_ns?

Cheers,

Michael

> reasonably large (but mechanical) change.
>
> James
>



-- 
Michael Kerrisk
Linux man-pages maintainer; http://www.kernel.org/doc/man-pages/
Linux/UNIX System Programming Training: http://man7.org/training/

[toc] | [prev] | [next] | [standalone]


#1438790

From"Serge E. Hallyn" <serge@hallyn.com>
Date2016-07-07 20:30 +0200
Message-ID<rSmdH-275-5@gated-at.bofh.it>
In reply to#1438789
Quoting Michael Kerrisk (man-pages) (mtk.manpages@gmail.com):
> On 7 July 2016 at 17:01, James Bottomley
> <James.Bottomley@hansenpartnership.com> wrote:
> > On Thu, 2016-07-07 at 08:36 -0500, Serge E. Hallyn wrote:
> >> Quoting Michael Kerrisk (man-pages) (mtk.manpages@gmail.com):
> >> > Hi Serge,
> >> >
> >> > On 6 July 2016 at 16:13, Serge E. Hallyn <serge@hallyn.com> wrote:
> >> > > On Wed, Jul 06, 2016 at 10:41:48AM +0200, Michael Kerrisk (man
> >> > > -pages) wrote:
> >> > > > [Rats! Doing now what I should have down to start with. Looping
> >> > > > some lists and CRIU and other possibly relevant people into
> >> > > > this conversation]
> >> > > >
> >> > > > Hi Eric,
> >> > > >
> >> > > > On 5 July 2016 at 23:47, Eric W. Biederman <
> >> > > > ebiederm@xmission.com> wrote:
> >> > > > > "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
> >> > > > > writes:
> >> > > > >
> >> > > > > > Hi Eric,
> >> > > > > >
> >> > > > > > I have a question. Is there any way currently to discover
> >> > > > > > which user namespace a particular nonuser namespace is
> >> > > > > > governed by? Maybe I am missing something, but there does
> >> > > > > > not seem to be a way to do this. Also, can one discover
> >> > > > > > which userns is the parent of a given userns? Again, I
> >> > > > > > can't see a way to do this.
> >> > > > > >
> >> > > > > > The point here is introspecting so that a process might
> >> > > > > > determine what its capabilities are when operating on some
> >> > > > > > resource governed by a (nonuser) namespace.
> >> > > > >
> >> > > > > To the best of my knowledge that there is not an interface to
> >> > > > > get that information.  It would be good to have such an
> >> > > > > interface for no other reason than the CRIU folks are going
> >> > > > > to need it at some point.  I am a bit surprised they have not
> >> > > > > complained yet.
> >> > >
> >> > > I don't think they need it.  They do in fact have what they need.
> >> > >   Assume you have tasks T1, T2, T1_1 and T2_1;  T1 and T2 are in
> >> > > init_user_ns;  T1 spawned T1_1 in a new userns;  T2 spawned T2_1
> >> > > which setns()d to T1_1's ns. There's some {handwave} uid mapping,
> >> > > does not matter.
> >> > >
> >> > > At restart, it doesn't matter which task originally created the
> >> > > new userns. criu knows T1_1 and T2_1 are in the same userns;  it
> >> > > creates the userns, sets up the mapping, and T1_1 and T2_1
> >> > > setns() to it.
> >> >
> >> > I'm missing something here. How does the parental relationships
> >> > between the user namespaces get reconstructed? Those relationships
> >> > will govern what capabilities a process will have in various user
> >> > namespaces.
> >
> > Actually, you get the parent namespace from the process tree by
> > tracking the user namespaces of the parent pids.   Currently non-root
> > users can't bind the namespace, so the only way to keep a new user_ns
> > around if you're not root is to keep the process around, so for
> > multiply nested user namespaces you can usually build the user_ns
> > hierarchy by looking at the process hierarchy.  Conversely, if the
> > process is reparented to init, chances are that the user_ns is also
> > parented to init_user_ns.
> 
> Yes, but "chances are" == this isn't robust.  PR_SET_CHILD_SUBREAPER
> further complicates things.
> 
> By the way, is that really what happens? Do child user namespaces get
> reparented to the grandparent ns if the parent ns disappears (i.e.,

The parent ns cannot disappear.  The child ns pins the creator's cred,
which pins the parent user_ns.

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.kernel


csiph-web