Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1437548 > unrolled thread

Re: Introspecting userns relationships to other namespaces?

Started by"Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
First post2016-07-06 10:50 +0200
Last post2016-07-09 05:30 +0200
Articles 13 on this page of 33 — 7 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: Introspecting userns relationships to other namespaces? "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-06 10:50 +0200
    Re: Introspecting userns relationships to other namespaces? "Serge E. Hallyn" <serge@hallyn.com> - 2016-07-06 16:20 +0200
      Re: Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-06 18:00 +0200
        Re: [CRIU] Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-08 10:00 +0200
          Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-08 16:40 +0200
            Re: [CRIU] Introspecting userns relationships to other namespaces? "W. Trevor King" <wking@tremily.us> - 2016-07-08 23:00 +0200
            Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@Hansenpartnership.com> - 2016-07-09 00:30 +0200
            Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@Hansenpartnership.com> - 2016-07-09 00:30 +0200
              Re: [CRIU] Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-09 02:10 +0200
                Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-09 02:20 +0200
                  Re: [CRIU] Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-09 05:20 +0200
                    Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@Hansenpartnership.com> - 2016-07-09 12:40 +0200
                    Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@Hansenpartnership.com> - 2016-07-09 12:40 +0200
                    Re: [CRIU] Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-09 20:30 +0200
                      Re: [CRIU] Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-09 20:50 +0200
      Re: Introspecting userns relationships to other namespaces? "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-07 10:20 +0200
        Re: Introspecting userns relationships to other namespaces? "Serge E. Hallyn" <serge@hallyn.com> - 2016-07-07 15:40 +0200
          Re: Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-07 17:10 +0200
            Re: Introspecting userns relationships to other namespaces? "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-07 20:30 +0200
              Re: Introspecting userns relationships to other namespaces? "Serge E. Hallyn" <serge@hallyn.com> - 2016-07-07 20:30 +0200
              Re: Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-07 21:20 +0200
                Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-08 05:30 +0200
                  Re: Introspecting userns relationships to other namespaces? "W. Trevor King" <wking@tremily.us> - 2016-07-08 07:40 +0200
                    Re: Introspecting userns relationships to other namespaces? "W. Trevor King" <wking@tremily.us> - 2016-07-08 08:20 +0200
                    Re: Introspecting userns relationships to other namespaces? "W. Trevor King" <wking@tremily.us> - 2016-07-08 09:20 +0200
                  Re: [CRIU] Introspecting userns relationships to other namespaces? Andrei Vagin <avagin@gmail.com> - 2016-07-08 07:50 +0200
                    Re: [CRIU] Introspecting userns relationships to other namespaces? Andrei Vagin <avagin@gmail.com> - 2016-07-08 07:50 +0200
                    Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-08 08:20 +0200
                  Re: [CRIU] Introspecting userns relationships to other namespaces? "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-08 13:20 +0200
                Re: [CRIU] Introspecting userns relationships to other namespaces? James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-08 05:30 +0200
                Re: Introspecting userns relationships to other namespaces? "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-08 13:20 +0200
            Re: Introspecting userns relationships to other namespaces? "W. Trevor King" <wking@tremily.us> - 2016-07-09 05:20 +0200
              Re: Introspecting userns relationships to other namespaces? ebiederm@xmission.com (Eric W. Biederman) - 2016-07-09 05:30 +0200

Page 2 of 2 — ← Prev page 1 [2]


#1438906

FromJames Bottomley <James.Bottomley@HansenPartnership.com>
Date2016-07-07 21:20 +0200
Message-ID<rSn05-2Gq-1@gated-at.bofh.it>
In reply to#1438789
On Thu, 2016-07-07 at 20:21 +0200, Michael Kerrisk (man-pages) wrote:
> On 7 July 2016 at 17:01, James Bottomley
> <James.Bottomley@hansenpartnership.com> wrote:
[Serge already answered the parenting issue]
> > On Thu, 2016-07-07 at 08:36 -0500, Serge E. Hallyn wrote:
> > > Hm.  Probably best-effort based on the process hierarchy.  So 
> > > yeah you could probably get a tree into a state that would be 
> > > wrongly recreated. Create a new netns, bind mount it, exit;  Have 
> > > another task create a new user_ns, bind mount it, exit;  Third 
> > > task setns()s first to the new netns then to the new user_ns.  I 
> > > suspect criu will recreate that wrongly.
> > 
> > This is a bit pathological, and you have to be root to do it: so 
> > root can set up a nesting hierarchy, bind it and destroy the pids 
> > but I know of no current orchestration system which does this.
> > 
> > Actually, I have to back pedal a bit: the way I currently set up
> > architecture emulation containers does precisely this: I set up the
> > namespaces unprivileged with child mount namespaces, but then I ask
> > root to bind the userns and kill the process that created it so I 
> > have a permanent handle to enter the namespace by, so I suspect 
> > that when our current orchestration systems get more sophisticated, 
> > they might eventually want to do something like this as well.
> > 
> > In theory, we could get nsfs to show this information as an option
> > (just add a show_options entry to the superblock ops), but the 
> > problem is that although each namespace has a parent user_ns, 
> > there's no way to get it without digging in the namespace specific 
> > structure.  Probably we should restructure to move it into 
> > ns_common, then we could display it (and enforce all namespaces 
> > having owning user_ns) but it would be a
> 
> I'm missing something here. Is it not already the case that all
> namespaces have an owning user_ns?

Um, yes, I don't believe I said they don't.  The problem I thought you
were having is that there's no way of seeing what it is.

nsfs is the Namespace fileystem where bound namespaces appear to a cat
of /proc/self/mounts.  It can display any information that's in
ns_common (the common core of namespaces) but the owning user_ns
pointer currently isn't in this structure.  Every user namespace has a
pointer to it, but they're all privately embedded in the individual
namespace specific structures.  What I was proposing was that since
every current namespace has a pointer somewhere to the owning user
namespace, we could abstract this out into ns_common so it's now
accessible to be displayed by nsfs, probably as a mount option.

James

[toc] | [prev] | [next] | [standalone]


#1439083 — Re: [CRIU] Introspecting userns relationships to other namespaces?

FromJames Bottomley <James.Bottomley@HansenPartnership.com>
Date2016-07-08 05:30 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSuEh-7zI-1@gated-at.bofh.it>
In reply to#1438906
On Thu, 2016-07-07 at 20:00 -0700, Andrew Vagin wrote:
> On Thu, Jul 07, 2016 at 07:16:18PM -0700, Andrew Vagin wrote:
> > On Thu, Jul 07, 2016 at 12:17:35PM -0700, James Bottomley wrote:
> > > On Thu, 2016-07-07 at 20:21 +0200, Michael Kerrisk (man-pages)
> > > wrote:
> > > > On 7 July 2016 at 17:01, James Bottomley
> > > > <James.Bottomley@hansenpartnership.com> wrote:
> > > [Serge already answered the parenting issue]
> > > > > On Thu, 2016-07-07 at 08:36 -0500, Serge E. Hallyn wrote:
> > > > > > Hm.  Probably best-effort based on the process hierarchy. 
> > > > > >  So 
> > > > > > yeah you could probably get a tree into a state that would
> > > > > > be 
> > > > > > wrongly recreated. Create a new netns, bind mount it, exit;
> > > > > >   Have 
> > > > > > another task create a new user_ns, bind mount it, exit; 
> > > > > >  Third 
> > > > > > task setns()s first to the new netns then to the new
> > > > > > user_ns.  I 
> > > > > > suspect criu will recreate that wrongly.
> > > > > 
> > > > > This is a bit pathological, and you have to be root to do it:
> > > > > so 
> > > > > root can set up a nesting hierarchy, bind it and destroy the
> > > > > pids 
> > > > > but I know of no current orchestration system which does
> > > > > this.
> > > > > 
> > > > > Actually, I have to back pedal a bit: the way I currently set
> > > > > up
> > > > > architecture emulation containers does precisely this: I set
> > > > > up the
> > > > > namespaces unprivileged with child mount namespaces, but then
> > > > > I ask
> > > > > root to bind the userns and kill the process that created it
> > > > > so I 
> > > > > have a permanent handle to enter the namespace by, so I
> > > > > suspect 
> > > > > that when our current orchestration systems get more
> > > > > sophisticated, 
> > > > > they might eventually want to do something like this as well.
> > > > > 
> > > > > In theory, we could get nsfs to show this information as an
> > > > > option
> > > > > (just add a show_options entry to the superblock ops), but
> > > > > the 
> > > > > problem is that although each namespace has a parent user_ns,
> > > > > there's no way to get it without digging in the namespace
> > > > > specific 
> > > > > structure.  Probably we should restructure to move it into 
> > > > > ns_common, then we could display it (and enforce all
> > > > > namespaces 
> > > > > having owning user_ns) but it would be a
> > > > 
> > > > I'm missing something here. Is it not already the case that all
> > > > namespaces have an owning user_ns?
> > > 
> > > Um, yes, I don't believe I said they don't.  The problem I
> > > thought you
> > > were having is that there's no way of seeing what it is.
> > > 
> > > nsfs is the Namespace fileystem where bound namespaces appear to
> > > a cat
> > > of /proc/self/mounts.  It can display any information that's in
> > > ns_common (the common core of namespaces) but the owning user_ns
> > > pointer currently isn't in this structure.  Every user namespace
> > > has a
> > > pointer to it, but they're all privately embedded in the
> > > individual
> > > namespace specific structures.  What I was proposing was that
> > > since
> > > every current namespace has a pointer somewhere to the owning
> > > user
> > > namespace, we could abstract this out into ns_common so it's now
> > > accessible to be displayed by nsfs, probably as a mount option.
> > 
> > James, I am not sure that I understood you correctly. We have one
> > file system for all namespace files, how we can show per-file
> > properties
> > in mount options. I think we can show all required information in
> > fdinfo. We open a namespaces file (/proc/pid/ns/N) and then read
> > /proc/pid/fdinfo/X for it.
> 
> Here is a proof-of-concept patch.
> 
> How it works:
> 
> In [1]: import os
> 
> In [2]: fd = os.open("/proc/self/ns/pid", os.O_RDONLY)
> 
> In [3]: print open("/proc/self/fdinfo/%d" % fd).read()
> pos:	0
> flags:	0100000
> mnt_id:	2
> userns: 4026531837
> 
> In [4]: print "/proc/self/ns/user -> %s" %
> os.readlink("/proc/self/ns/user")
> /proc/self/ns/user -> user:[4026531837]

can't you just do

readlink /proc/self/ns/user | sed 's/.*\[\(.*\)\]/\1/'

?

But what Michael was asking about was the parent user_ns of all the
other namespaces ... I don't think there's any way we can get that out
of any information in /proc/self/

James

[toc] | [prev] | [next] | [standalone]


#1439114

From"W. Trevor King" <wking@tremily.us>
Date2016-07-08 07:40 +0200
Message-ID<rSwG5-xv-3@gated-at.bofh.it>
In reply to#1439083

[Multipart message — attachments visible in raw view] — view raw

On Thu, Jul 07, 2016 at 08:26:47PM -0700, James Bottomley wrote:
> On Thu, 2016-07-07 at 20:00 -0700, Andrew Vagin wrote:
> > On Thu, Jul 07, 2016 at 07:16:18PM -0700, Andrew Vagin wrote:
> > > I think we can show all required information in fdinfo. We open
> > > a namespaces file (/proc/pid/ns/N) and then read
> > > /proc/pid/fdinfo/X for it.
> > 
> > Here is a proof-of-concept patch.
> > …
> > In [2]: fd = os.open("/proc/self/ns/pid", os.O_RDONLY)
> > 
> > In [3]: print open("/proc/self/fdinfo/%d" % fd).read()
> > pos:	0
> > flags:	0100000
> > mnt_id:	2
> > userns: 4026531837
> > 
> > In [4]: print "/proc/self/ns/user -> %s" %
> > os.readlink("/proc/self/ns/user")
> > /proc/self/ns/user -> user:[4026531837]
> 
> can't you just do
> 
> readlink /proc/self/ns/user | sed 's/.*\[\(.*\)\]/\1/'

With Andrew's fdinfo approach you know the user namespace owning
/proc/self/ns/pid is 4026531837.  That happens to be
/proc/self/ns/user in this case, but doesn't have to be in general.

> But what Michael was asking about was the parent user_ns of all the
> other namespaces ... I don't think there's any way we can get that
> out of any information in /proc/self/

If fdinfo only shows immediate parents, you'd need to walk the tree to
get back to the root.  And at each layer of the PID namespace tree
there will be another user-namespace parent branching off).  With a
tree like:

  Namespace         | Parent       | Owning userns
 -------------------+--------------+-------------------
  Root userns       | -            | -
  Root PID ns       | -            | Root userns
  Child userns      | Root usens   | Root userns
  Child PID ns      | Root PID ns  | Root userns
  Grandchild userns | Child userns | Child userns
  Grandchild PID ns | Child PID ns | Grandchild userns

Walking from the granchild PID namespace would give you:

  Grandchild PID ns
  |-- Child PID ns
  |   |-- Root PID ns
  |   `-- Root userns 
  `-- Granchild userns
      `-- Child userns
          `-- Root userns

If you only put one level in fdinfo, you're stuck if one of the
namespaces involved has neither bind mounts nor a PID to give you
handle on it [1].  And if you want to put that whole ancestor tree in
fdinfo, you have to come up with some way to handle the two-parent
branching.

I'm also not sure how exposing nsfs information [2] would handle
namespaces that had neither a surviving bind mount nor a direct
process.

If all the information is available (possible after a mechanical patch
[3] makes it more accessible), then it seems easier to put it in a
separate /proc or /sys file.  There was a stab at this for PID
namespaces in [4] (the same series that landed NStgid, etc.) with
additional background and alternative approaches in [5].  There were
problems with that patch (and it was trying to do more by also listing
a process's ID in each PID namespace), but the “let's put the whole
tree in a new file” approach seems sound to me.

Cheers,
Trevor

[1]: http://thread.gmane.org/gmane.linux.kernel.containers/30456/focus=20536
     Subject: Re: Introspecting userns relationships to other namespaces?
     Date: Thu, 7 Jul 2016 13:24:42 -0500
     Message-ID: <20160707182442.GA6402@mail.hallyn.com>
[2]: http://thread.gmane.org/gmane.linux.kernel.containers/30456/focus=30499
     Subject: Re: [CRIU] Introspecting userns relationships to other namespaces?
     Date: Thu, 07 Jul 2016 20:20:05 -0700
     Message-ID: <1467948005.2322.84.camel@HansenPartnership.com>
[3]: http://thread.gmane.org/gmane.linux.kernel.containers/30456/focus=20537
     Subject: Re: Introspecting userns relationships to other namespaces?
     Message-ID: <1467903712.2347.16.camel@HansenPartnership.com>
     Date: Thu, 07 Jul 2016 08:01:52 -0700
[4]: http://thread.gmane.org/gmane.linux.kernel.containers/28925/focus=28928
     Subject: [resend][PATCH v9 1/3] procfs: show hierarchy of pid namespace
     Date: Tue, 23 Dec 2014 18:20:37 +0800
     Message-ID: <1419330039-29207-2-git-send-email-chenhanxiao@cn.fujitsu.com>
[5]: http://thread.gmane.org/gmane.linux.kernel.containers/28105
     Subject: [RFC]Pid conversion between pid namespace
     Date: Thu, 3 Jul 2014 12:18:33 +0000
     Message-ID: <5871495633F38949900D2BF2DC04883E55C374@G08CNEXMBPEKD02.g08.fujitsu.local>

-- 
This email may be signed or encrypted with GnuPG (http://www.gnupg.org).
For more information, see http://en.wikipedia.org/wiki/Pretty_Good_Privacy

[toc] | [prev] | [next] | [standalone]


#1439125

From"W. Trevor King" <wking@tremily.us>
Date2016-07-08 08:20 +0200
Message-ID<rSxiO-106-15@gated-at.bofh.it>
In reply to#1439114

[Multipart message — attachments visible in raw view] — view raw

On Thu, Jul 07, 2016 at 10:26:50PM -0700, W. Trevor King wrote:
> And if you want to put that whole ancestor tree in fdinfo, you have
> to come up with some way to handle the two-parent branching.

Going towards the roots is nice, because you know a given namespace
will only have two parents, but it leaks information about the system
into the container.  It's probably better to follow the NStgid,
etc. example and only walk toward the leaves.  So a (privileged?)
process in the root namespace could see the whole tree, while a
process in non-root namespaces could only see their namespaces and
descendants.  In situations where you were part of a namespace that
belonged to an external user namespace (e.g. you nsenter a child user
namespace but are still in the root PID namespace), you'd want an
“unknown” entry for the parent you couldn't see.

Cheers,
Trevor

-- 
This email may be signed or encrypted with GnuPG (http://www.gnupg.org).
For more information, see http://en.wikipedia.org/wiki/Pretty_Good_Privacy

[toc] | [prev] | [next] | [standalone]


#1439146

From"W. Trevor King" <wking@tremily.us>
Date2016-07-08 09:20 +0200
Message-ID<rSyeR-1z5-3@gated-at.bofh.it>
In reply to#1439114

[Multipart message — attachments visible in raw view] — view raw

On Thu, Jul 07, 2016 at 11:54:54PM -0700, Andrew Vagin wrote:
> On Thu, Jul 07, 2016 at 10:26:50PM -0700, W. Trevor King wrote:
> > On Thu, Jul 07, 2016 at 08:26:47PM -0700, James Bottomley wrote:
> > > On Thu, 2016-07-07 at 20:00 -0700, Andrew Vagin wrote:
> > > > On Thu, Jul 07, 2016 at 07:16:18PM -0700, Andrew Vagin wrote:
> > > > > I think we can show all required information in fdinfo. We open
> > > > > a namespaces file (/proc/pid/ns/N) and then read
> > > > > /proc/pid/fdinfo/X for it.
> > > > 
> > > > Here is a proof-of-concept patch.
> > > > …
> > > > In [2]: fd = os.open("/proc/self/ns/pid", os.O_RDONLY)
> > > > 
> > > > In [3]: print open("/proc/self/fdinfo/%d" % fd).read()
> > > > pos:	0
> > > > flags:	0100000
> > > > mnt_id:	2
> > > > userns: 4026531837
> > > > 
> > > > In [4]: print "/proc/self/ns/user -> %s" %
> > > > os.readlink("/proc/self/ns/user")
> > > > /proc/self/ns/user -> user:[4026531837]
> > > 
> > > can't you just do
> > > 
> > > readlink /proc/self/ns/user | sed 's/.*\[\(.*\)\]/\1/'
> > …
> > If you only put one level in fdinfo, you're stuck if one of the
> > namespaces involved has neither bind mounts nor a PID to give you
> > handle on it [1].  And if you want to put that whole ancestor tree in
> > fdinfo, you have to come up with some way to handle the two-parent
> > branching.
> 
> I think it's a bad idea to draw a tree in fdinfo. Why do we want to know
> this hierarchy? Probably we will want to access these namespaces (setns),
> in this case we need to have a way to open them.
> 
> Maybe we need to extend functionality of the nsfs filesystem
> (somethink like /proc/PID for namespaces)?

A similar idea came up during the PID-translation brainstorming [1],
but I'm not sure if anything ever came of that.  Once you're dealing
with a separate pseudo-filesystem, it seems easier to decouple it from
proc and just make a mountable namespace-hierarchy filesystem (like we
have mountable cgroup hierarchy filesystems).  That also gets you an
opt-in playground while the details of the nsfs filesystem view are
worked out.  Are you imagining something like:

  $ tree .
  .
  ├── mnt{inum}
  │   └── user -> ../user{inum}
  ├── pid{inum}
  │   ├── pid{inum}
  │   │   └── user -> ../../user{inum}/user{inum}
  │   └── user -> ../user{inum}
  └── user{inum}
      └── user{inum}

Cheers,
Trevor

[1]: http://thread.gmane.org/gmane.linux.kernel.containers/28105/focus=28164
     Subject: RE: [RFC]Pid conversion between pid namespace
     Date: Fri, 25 Jul 2014 10:01:45 +0000
     Message-ID: <5871495633F38949900D2BF2DC04883E56C7A2@G08CNEXMBPEKD02.g08.fujitsu.local>

-- 
This email may be signed or encrypted with GnuPG (http://www.gnupg.org).
For more information, see http://en.wikipedia.org/wiki/Pretty_Good_Privacy

[toc] | [prev] | [next] | [standalone]


#1439115 — Re: [CRIU] Introspecting userns relationships to other namespaces?

FromAndrei Vagin <avagin@gmail.com>
Date2016-07-08 07:50 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSwPL-B6-1@gated-at.bofh.it>
In reply to#1439083
On Thu, Jul 7, 2016 at 8:26 PM, James Bottomley
<James.Bottomley@hansenpartnership.com> wrote:
> On Thu, 2016-07-07 at 20:00 -0700, Andrew Vagin wrote:
>> On Thu, Jul 07, 2016 at 07:16:18PM -0700, Andrew Vagin wrote:
>> > On Thu, Jul 07, 2016 at 12:17:35PM -0700, James Bottomley wrote:
>> > > On Thu, 2016-07-07 at 20:21 +0200, Michael Kerrisk (man-pages)
>> > > wrote:
>> > > > On 7 July 2016 at 17:01, James Bottomley
>> > > > <James.Bottomley@hansenpartnership.com> wrote:
>> > > [Serge already answered the parenting issue]
>> > > > > On Thu, 2016-07-07 at 08:36 -0500, Serge E. Hallyn wrote:
>> > > > > > Hm.  Probably best-effort based on the process hierarchy.
>> > > > > >  So
>> > > > > > yeah you could probably get a tree into a state that would
>> > > > > > be
>> > > > > > wrongly recreated. Create a new netns, bind mount it, exit;
>> > > > > >   Have
>> > > > > > another task create a new user_ns, bind mount it, exit;
>> > > > > >  Third
>> > > > > > task setns()s first to the new netns then to the new
>> > > > > > user_ns.  I
>> > > > > > suspect criu will recreate that wrongly.
>> > > > >
>> > > > > This is a bit pathological, and you have to be root to do it:
>> > > > > so
>> > > > > root can set up a nesting hierarchy, bind it and destroy the
>> > > > > pids
>> > > > > but I know of no current orchestration system which does
>> > > > > this.
>> > > > >
>> > > > > Actually, I have to back pedal a bit: the way I currently set
>> > > > > up
>> > > > > architecture emulation containers does precisely this: I set
>> > > > > up the
>> > > > > namespaces unprivileged with child mount namespaces, but then
>> > > > > I ask
>> > > > > root to bind the userns and kill the process that created it
>> > > > > so I
>> > > > > have a permanent handle to enter the namespace by, so I
>> > > > > suspect
>> > > > > that when our current orchestration systems get more
>> > > > > sophisticated,
>> > > > > they might eventually want to do something like this as well.
>> > > > >
>> > > > > In theory, we could get nsfs to show this information as an
>> > > > > option
>> > > > > (just add a show_options entry to the superblock ops), but
>> > > > > the
>> > > > > problem is that although each namespace has a parent user_ns,
>> > > > > there's no way to get it without digging in the namespace
>> > > > > specific
>> > > > > structure.  Probably we should restructure to move it into
>> > > > > ns_common, then we could display it (and enforce all
>> > > > > namespaces
>> > > > > having owning user_ns) but it would be a
>> > > >
>> > > > I'm missing something here. Is it not already the case that all
>> > > > namespaces have an owning user_ns?
>> > >
>> > > Um, yes, I don't believe I said they don't.  The problem I
>> > > thought you
>> > > were having is that there's no way of seeing what it is.
>> > >
>> > > nsfs is the Namespace fileystem where bound namespaces appear to
>> > > a cat
>> > > of /proc/self/mounts.  It can display any information that's in
>> > > ns_common (the common core of namespaces) but the owning user_ns
>> > > pointer currently isn't in this structure.  Every user namespace
>> > > has a
>> > > pointer to it, but they're all privately embedded in the
>> > > individual
>> > > namespace specific structures.  What I was proposing was that
>> > > since
>> > > every current namespace has a pointer somewhere to the owning
>> > > user
>> > > namespace, we could abstract this out into ns_common so it's now
>> > > accessible to be displayed by nsfs, probably as a mount option.
>> >
>> > James, I am not sure that I understood you correctly. We have one
>> > file system for all namespace files, how we can show per-file
>> > properties
>> > in mount options. I think we can show all required information in
>> > fdinfo. We open a namespaces file (/proc/pid/ns/N) and then read
>> > /proc/pid/fdinfo/X for it.
>>
>> Here is a proof-of-concept patch.
>>
>> How it works:
>>
>> In [1]: import os
>>
>> In [2]: fd = os.open("/proc/self/ns/pid", os.O_RDONLY)
>>
>> In [3]: print open("/proc/self/fdinfo/%d" % fd).read()
>> pos:  0
>> flags:        0100000
>> mnt_id:       2
>> userns: 4026531837
>>
>> In [4]: print "/proc/self/ns/user -> %s" %
>> os.readlink("/proc/self/ns/user")
>> /proc/self/ns/user -> user:[4026531837]
>
> can't you just do
>
> readlink /proc/self/ns/user | sed 's/.*\[\(.*\)\]/\1/'

We can get fdinfo for any ns file. I used /proc/self/ns/pid as an example.

Look at another example:

[root@fc22-vm ~]# cat /proc/self/mountinfo | grep pid_ns_file
115 38 0:3 pid:[4026532306] /tmp/pid_ns_file rw shared:67 - nsfs nsfs rw

In [4]: print open("/proc/self/fdinfo/5").read()
pos: 0
flags: 0100000
mnt_id: 115
userns: 4026532305


In [5]: os.readlink("/proc/self/ns/user")
Out[5]: 'user:[4026531837]'


>
> ?
>
> But what Michael was asking about was the parent user_ns of all the
> other namespaces ... I don't think there's any way we can get that out
> of any information in /proc/self/
>
> James
>
>
> _______________________________________________
> Containers mailing list
> Containers@lists.linux-foundation.org
> https://lists.linuxfoundation.org/mailman/listinfo/containers

[toc] | [prev] | [next] | [standalone]


#1439116 — Re: [CRIU] Introspecting userns relationships to other namespaces?

FromAndrei Vagin <avagin@gmail.com>
Date2016-07-08 07:50 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSwPM-B6-7@gated-at.bofh.it>
In reply to#1439115
On Thu, Jul 7, 2016 at 10:41 PM, Andrei Vagin <avagin@gmail.com> wrote:
> On Thu, Jul 7, 2016 at 8:26 PM, James Bottomley
> <James.Bottomley@hansenpartnership.com> wrote:
>> On Thu, 2016-07-07 at 20:00 -0700, Andrew Vagin wrote:
>>> On Thu, Jul 07, 2016 at 07:16:18PM -0700, Andrew Vagin wrote:
>>> > On Thu, Jul 07, 2016 at 12:17:35PM -0700, James Bottomley wrote:
>>> > > On Thu, 2016-07-07 at 20:21 +0200, Michael Kerrisk (man-pages)
>>> > > wrote:
>>> > > > On 7 July 2016 at 17:01, James Bottomley
>>> > > > <James.Bottomley@hansenpartnership.com> wrote:
>>> > > [Serge already answered the parenting issue]
>>> > > > > On Thu, 2016-07-07 at 08:36 -0500, Serge E. Hallyn wrote:
>>> > > > > > Hm.  Probably best-effort based on the process hierarchy.
>>> > > > > >  So
>>> > > > > > yeah you could probably get a tree into a state that would
>>> > > > > > be
>>> > > > > > wrongly recreated. Create a new netns, bind mount it, exit;
>>> > > > > >   Have
>>> > > > > > another task create a new user_ns, bind mount it, exit;
>>> > > > > >  Third
>>> > > > > > task setns()s first to the new netns then to the new
>>> > > > > > user_ns.  I
>>> > > > > > suspect criu will recreate that wrongly.
>>> > > > >
>>> > > > > This is a bit pathological, and you have to be root to do it:
>>> > > > > so
>>> > > > > root can set up a nesting hierarchy, bind it and destroy the
>>> > > > > pids
>>> > > > > but I know of no current orchestration system which does
>>> > > > > this.
>>> > > > >
>>> > > > > Actually, I have to back pedal a bit: the way I currently set
>>> > > > > up
>>> > > > > architecture emulation containers does precisely this: I set
>>> > > > > up the
>>> > > > > namespaces unprivileged with child mount namespaces, but then
>>> > > > > I ask
>>> > > > > root to bind the userns and kill the process that created it
>>> > > > > so I
>>> > > > > have a permanent handle to enter the namespace by, so I
>>> > > > > suspect
>>> > > > > that when our current orchestration systems get more
>>> > > > > sophisticated,
>>> > > > > they might eventually want to do something like this as well.
>>> > > > >
>>> > > > > In theory, we could get nsfs to show this information as an
>>> > > > > option
>>> > > > > (just add a show_options entry to the superblock ops), but
>>> > > > > the
>>> > > > > problem is that although each namespace has a parent user_ns,
>>> > > > > there's no way to get it without digging in the namespace
>>> > > > > specific
>>> > > > > structure.  Probably we should restructure to move it into
>>> > > > > ns_common, then we could display it (and enforce all
>>> > > > > namespaces
>>> > > > > having owning user_ns) but it would be a
>>> > > >
>>> > > > I'm missing something here. Is it not already the case that all
>>> > > > namespaces have an owning user_ns?
>>> > >
>>> > > Um, yes, I don't believe I said they don't.  The problem I
>>> > > thought you
>>> > > were having is that there's no way of seeing what it is.
>>> > >
>>> > > nsfs is the Namespace fileystem where bound namespaces appear to
>>> > > a cat
>>> > > of /proc/self/mounts.  It can display any information that's in
>>> > > ns_common (the common core of namespaces) but the owning user_ns
>>> > > pointer currently isn't in this structure.  Every user namespace
>>> > > has a
>>> > > pointer to it, but they're all privately embedded in the
>>> > > individual
>>> > > namespace specific structures.  What I was proposing was that
>>> > > since
>>> > > every current namespace has a pointer somewhere to the owning
>>> > > user
>>> > > namespace, we could abstract this out into ns_common so it's now
>>> > > accessible to be displayed by nsfs, probably as a mount option.
>>> >
>>> > James, I am not sure that I understood you correctly. We have one
>>> > file system for all namespace files, how we can show per-file
>>> > properties
>>> > in mount options. I think we can show all required information in
>>> > fdinfo. We open a namespaces file (/proc/pid/ns/N) and then read
>>> > /proc/pid/fdinfo/X for it.
>>>
>>> Here is a proof-of-concept patch.
>>>
>>> How it works:
>>>
>>> In [1]: import os
>>>
>>> In [2]: fd = os.open("/proc/self/ns/pid", os.O_RDONLY)
>>>
>>> In [3]: print open("/proc/self/fdinfo/%d" % fd).read()
>>> pos:  0
>>> flags:        0100000
>>> mnt_id:       2
>>> userns: 4026531837
>>>
>>> In [4]: print "/proc/self/ns/user -> %s" %
>>> os.readlink("/proc/self/ns/user")
>>> /proc/self/ns/user -> user:[4026531837]
>>
>> can't you just do
>>
>> readlink /proc/self/ns/user | sed 's/.*\[\(.*\)\]/\1/'
>
> We can get fdinfo for any ns file. I used /proc/self/ns/pid as an example.
>
> Look at another example:
>
> [root@fc22-vm ~]# cat /proc/self/mountinfo | grep pid_ns_file
> 115 38 0:3 pid:[4026532306] /tmp/pid_ns_file rw shared:67 - nsfs nsfs rw
>

Sorry, I forgot to say that fd is a file descriptor for /tmp/pid_ns_file

In [2]  : fd = os.open("/tmp/pid_ns_file", os.O_RDONLY)
In [3]  : fd
Out[4]: 5

> In [4]: print open("/proc/self/fdinfo/5").read()
> pos: 0
> flags: 0100000
> mnt_id: 115
> userns: 4026532305
>
>
> In [5]: os.readlink("/proc/self/ns/user")
> Out[5]: 'user:[4026531837]'
>
>
>>
>> ?
>>
>> But what Michael was asking about was the parent user_ns of all the
>> other namespaces ... I don't think there's any way we can get that out
>> of any information in /proc/self/
>>
>> James
>>
>>
>> _______________________________________________
>> Containers mailing list
>> Containers@lists.linux-foundation.org
>> https://lists.linuxfoundation.org/mailman/listinfo/containers

[toc] | [prev] | [next] | [standalone]


#1439122 — Re: [CRIU] Introspecting userns relationships to other namespaces?

FromJames Bottomley <James.Bottomley@HansenPartnership.com>
Date2016-07-08 08:20 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSxiN-106-5@gated-at.bofh.it>
In reply to#1439115
On Thu, 2016-07-07 at 22:41 -0700, Andrei Vagin wrote:
> On Thu, Jul 7, 2016 at 8:26 PM, James Bottomley
> <James.Bottomley@hansenpartnership.com> wrote:
> > On Thu, 2016-07-07 at 20:00 -0700, Andrew Vagin wrote:
> > > On Thu, Jul 07, 2016 at 07:16:18PM -0700, Andrew Vagin wrote:
> > > > On Thu, Jul 07, 2016 at 12:17:35PM -0700, James Bottomley
> > > > wrote:
> > > > > On Thu, 2016-07-07 at 20:21 +0200, Michael Kerrisk (man
> > > > > -pages)
> > > > > wrote:
> > > > > > On 7 July 2016 at 17:01, James Bottomley
> > > > > > <James.Bottomley@hansenpartnership.com> wrote:
> > > > > [Serge already answered the parenting issue]
> > > > > > > On Thu, 2016-07-07 at 08:36 -0500, Serge E. Hallyn wrote:
> > > > > > > > Hm.  Probably best-effort based on the process
> > > > > > > > hierarchy.
> > > > > > > >  So
> > > > > > > > yeah you could probably get a tree into a state that
> > > > > > > > would
> > > > > > > > be
> > > > > > > > wrongly recreated. Create a new netns, bind mount it,
> > > > > > > > exit;
> > > > > > > >   Have
> > > > > > > > another task create a new user_ns, bind mount it, exit;
> > > > > > > >  Third
> > > > > > > > task setns()s first to the new netns then to the new
> > > > > > > > user_ns.  I
> > > > > > > > suspect criu will recreate that wrongly.
> > > > > > > 
> > > > > > > This is a bit pathological, and you have to be root to do
> > > > > > > it:
> > > > > > > so
> > > > > > > root can set up a nesting hierarchy, bind it and destroy
> > > > > > > the
> > > > > > > pids
> > > > > > > but I know of no current orchestration system which does
> > > > > > > this.
> > > > > > > 
> > > > > > > Actually, I have to back pedal a bit: the way I currently
> > > > > > > set
> > > > > > > up
> > > > > > > architecture emulation containers does precisely this: I
> > > > > > > set
> > > > > > > up the
> > > > > > > namespaces unprivileged with child mount namespaces, but
> > > > > > > then
> > > > > > > I ask
> > > > > > > root to bind the userns and kill the process that created
> > > > > > > it
> > > > > > > so I
> > > > > > > have a permanent handle to enter the namespace by, so I
> > > > > > > suspect
> > > > > > > that when our current orchestration systems get more
> > > > > > > sophisticated,
> > > > > > > they might eventually want to do something like this as
> > > > > > > well.
> > > > > > > 
> > > > > > > In theory, we could get nsfs to show this information as
> > > > > > > an
> > > > > > > option
> > > > > > > (just add a show_options entry to the superblock ops),
> > > > > > > but
> > > > > > > the
> > > > > > > problem is that although each namespace has a parent
> > > > > > > user_ns,
> > > > > > > there's no way to get it without digging in the namespace
> > > > > > > specific
> > > > > > > structure.  Probably we should restructure to move it
> > > > > > > into
> > > > > > > ns_common, then we could display it (and enforce all
> > > > > > > namespaces
> > > > > > > having owning user_ns) but it would be a
> > > > > > 
> > > > > > I'm missing something here. Is it not already the case that
> > > > > > all
> > > > > > namespaces have an owning user_ns?
> > > > > 
> > > > > Um, yes, I don't believe I said they don't.  The problem I
> > > > > thought you
> > > > > were having is that there's no way of seeing what it is.
> > > > > 
> > > > > nsfs is the Namespace fileystem where bound namespaces appear
> > > > > to
> > > > > a cat
> > > > > of /proc/self/mounts.  It can display any information that's
> > > > > in
> > > > > ns_common (the common core of namespaces) but the owning
> > > > > user_ns
> > > > > pointer currently isn't in this structure.  Every user
> > > > > namespace
> > > > > has a
> > > > > pointer to it, but they're all privately embedded in the
> > > > > individual
> > > > > namespace specific structures.  What I was proposing was that
> > > > > since
> > > > > every current namespace has a pointer somewhere to the owning
> > > > > user
> > > > > namespace, we could abstract this out into ns_common so it's
> > > > > now
> > > > > accessible to be displayed by nsfs, probably as a mount
> > > > > option.
> > > > 
> > > > James, I am not sure that I understood you correctly. We have
> > > > one
> > > > file system for all namespace files, how we can show per-file
> > > > properties
> > > > in mount options. I think we can show all required information
> > > > in
> > > > fdinfo. We open a namespaces file (/proc/pid/ns/N) and then
> > > > read
> > > > /proc/pid/fdinfo/X for it.
> > > 
> > > Here is a proof-of-concept patch.
> > > 
> > > How it works:
> > > 
> > > In [1]: import os
> > > 
> > > In [2]: fd = os.open("/proc/self/ns/pid", os.O_RDONLY)
> > > 
> > > In [3]: print open("/proc/self/fdinfo/%d" % fd).read()
> > > pos:  0
> > > flags:        0100000
> > > mnt_id:       2
> > > userns: 4026531837
> > > 
> > > In [4]: print "/proc/self/ns/user -> %s" %
> > > os.readlink("/proc/self/ns/user")
> > > /proc/self/ns/user -> user:[4026531837]
> > 
> > can't you just do
> > 
> > readlink /proc/self/ns/user | sed 's/.*\[\(.*\)\]/\1/'
> 
> We can get fdinfo for any ns file. I used /proc/self/ns/pid as an
> example.
> 
> Look at another example:
> 
> [root@fc22-vm ~]# cat /proc/self/mountinfo | grep pid_ns_file
> 115 38 0:3 pid:[4026532306] /tmp/pid_ns_file rw shared:67 - nsfs nsfs
> rw
> 
> In [4]: print open("/proc/self/fdinfo/5").read()
> pos: 0
> flags: 0100000
> mnt_id: 115
> userns: 4026532305

OK, I'm missing where this is coming from specifically.  There would
have to be a show_fdinfo() somewhere that did this and I'm not finding
it in linux-next.

James

[toc] | [prev] | [next] | [standalone]


#1439305 — Re: [CRIU] Introspecting userns relationships to other namespaces?

From"Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
Date2016-07-08 13:20 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSBZ8-42k-17@gated-at.bofh.it>
In reply to#1439083
On 07/08/2016 05:26 AM, James Bottomley wrote:
> On Thu, 2016-07-07 at 20:00 -0700, Andrew Vagin wrote:
>> On Thu, Jul 07, 2016 at 07:16:18PM -0700, Andrew Vagin wrote:
>>> On Thu, Jul 07, 2016 at 12:17:35PM -0700, James Bottomley wrote:
>>>> On Thu, 2016-07-07 at 20:21 +0200, Michael Kerrisk (man-pages)
>>>> wrote:
>>>>> On 7 July 2016 at 17:01, James Bottomley
>>>>> <James.Bottomley@hansenpartnership.com> wrote:
>>>> [Serge already answered the parenting issue]
>>>>>> On Thu, 2016-07-07 at 08:36 -0500, Serge E. Hallyn wrote:
>>>>>>> Hm.  Probably best-effort based on the process hierarchy.
>>>>>>>  So
>>>>>>> yeah you could probably get a tree into a state that would
>>>>>>> be
>>>>>>> wrongly recreated. Create a new netns, bind mount it, exit;
>>>>>>>   Have
>>>>>>> another task create a new user_ns, bind mount it, exit;
>>>>>>>  Third
>>>>>>> task setns()s first to the new netns then to the new
>>>>>>> user_ns.  I
>>>>>>> suspect criu will recreate that wrongly.
>>>>>>
>>>>>> This is a bit pathological, and you have to be root to do it:
>>>>>> so
>>>>>> root can set up a nesting hierarchy, bind it and destroy the
>>>>>> pids
>>>>>> but I know of no current orchestration system which does
>>>>>> this.
>>>>>>
>>>>>> Actually, I have to back pedal a bit: the way I currently set
>>>>>> up
>>>>>> architecture emulation containers does precisely this: I set
>>>>>> up the
>>>>>> namespaces unprivileged with child mount namespaces, but then
>>>>>> I ask
>>>>>> root to bind the userns and kill the process that created it
>>>>>> so I
>>>>>> have a permanent handle to enter the namespace by, so I
>>>>>> suspect
>>>>>> that when our current orchestration systems get more
>>>>>> sophisticated,
>>>>>> they might eventually want to do something like this as well.
>>>>>>
>>>>>> In theory, we could get nsfs to show this information as an
>>>>>> option
>>>>>> (just add a show_options entry to the superblock ops), but
>>>>>> the
>>>>>> problem is that although each namespace has a parent user_ns,
>>>>>> there's no way to get it without digging in the namespace
>>>>>> specific
>>>>>> structure.  Probably we should restructure to move it into
>>>>>> ns_common, then we could display it (and enforce all
>>>>>> namespaces
>>>>>> having owning user_ns) but it would be a
>>>>>
>>>>> I'm missing something here. Is it not already the case that all
>>>>> namespaces have an owning user_ns?
>>>>
>>>> Um, yes, I don't believe I said they don't.  The problem I
>>>> thought you
>>>> were having is that there's no way of seeing what it is.
>>>>
>>>> nsfs is the Namespace fileystem where bound namespaces appear to
>>>> a cat
>>>> of /proc/self/mounts.  It can display any information that's in
>>>> ns_common (the common core of namespaces) but the owning user_ns
>>>> pointer currently isn't in this structure.  Every user namespace
>>>> has a
>>>> pointer to it, but they're all privately embedded in the
>>>> individual
>>>> namespace specific structures.  What I was proposing was that
>>>> since
>>>> every current namespace has a pointer somewhere to the owning
>>>> user
>>>> namespace, we could abstract this out into ns_common so it's now
>>>> accessible to be displayed by nsfs, probably as a mount option.
>>>
>>> James, I am not sure that I understood you correctly. We have one
>>> file system for all namespace files, how we can show per-file
>>> properties
>>> in mount options. I think we can show all required information in
>>> fdinfo. We open a namespaces file (/proc/pid/ns/N) and then read
>>> /proc/pid/fdinfo/X for it.
>>
>> Here is a proof-of-concept patch.
>>
>> How it works:
>>
>> In [1]: import os
>>
>> In [2]: fd = os.open("/proc/self/ns/pid", os.O_RDONLY)
>>
>> In [3]: print open("/proc/self/fdinfo/%d" % fd).read()
>> pos:	0
>> flags:	0100000
>> mnt_id:	2
>> userns: 4026531837
>>
>> In [4]: print "/proc/self/ns/user -> %s" %
>> os.readlink("/proc/self/ns/user")
>> /proc/self/ns/user -> user:[4026531837]
>
> can't you just do
>
> readlink /proc/self/ns/user | sed 's/.*\[\(.*\)\]/\1/'
>
> ?
>
> But what Michael was asking about was the parent user_ns of all the
> other namespaces ...

Just to reiterate, what I'm interested in is the introspection use
case (but there's clearly several other interesting use cases here).
The idea is to be able to answer these questions

1. For each userns, what is the parent of that userns?

2. For each non-user namespace, what is the owning userns?

This enables us to understand the userns hierarchy, which
matters in terms of answering the question: what capabilities
does process X have in namespace Y?
    
Cheers,

Michael

-- 
Michael Kerrisk
Linux man-pages maintainer; http://www.kernel.org/doc/man-pages/
Linux/UNIX System Programming Training: http://man7.org/training/

[toc] | [prev] | [next] | [standalone]


#1439087 — Re: [CRIU] Introspecting userns relationships to other namespaces?

FromJames Bottomley <James.Bottomley@HansenPartnership.com>
Date2016-07-08 05:30 +0200
SubjectRe: [CRIU] Introspecting userns relationships to other namespaces?
Message-ID<rSuEh-7zI-15@gated-at.bofh.it>
In reply to#1438906
On Thu, 2016-07-07 at 19:16 -0700, Andrew Vagin wrote:
> On Thu, Jul 07, 2016 at 12:17:35PM -0700, James Bottomley wrote:
> > On Thu, 2016-07-07 at 20:21 +0200, Michael Kerrisk (man-pages)
> > wrote:
> > > On 7 July 2016 at 17:01, James Bottomley
> > > <James.Bottomley@hansenpartnership.com> wrote:
> > [Serge already answered the parenting issue]
> > > > On Thu, 2016-07-07 at 08:36 -0500, Serge E. Hallyn wrote:
> > > > > Hm.  Probably best-effort based on the process hierarchy.  So
> > > > > yeah you could probably get a tree into a state that would be
> > > > > wrongly recreated. Create a new netns, bind mount it, exit; 
> > > > >  Have 
> > > > > another task create a new user_ns, bind mount it, exit; 
> > > > >  Third 
> > > > > task setns()s first to the new netns then to the new user_ns.
> > > > >   I 
> > > > > suspect criu will recreate that wrongly.
> > > > 
> > > > This is a bit pathological, and you have to be root to do it:
> > > > so 
> > > > root can set up a nesting hierarchy, bind it and destroy the
> > > > pids 
> > > > but I know of no current orchestration system which does this.
> > > > 
> > > > Actually, I have to back pedal a bit: the way I currently set
> > > > up
> > > > architecture emulation containers does precisely this: I set up
> > > > the
> > > > namespaces unprivileged with child mount namespaces, but then I
> > > > ask
> > > > root to bind the userns and kill the process that created it so
> > > > I 
> > > > have a permanent handle to enter the namespace by, so I suspect
> > > > that when our current orchestration systems get more
> > > > sophisticated, 
> > > > they might eventually want to do something like this as well.
> > > > 
> > > > In theory, we could get nsfs to show this information as an
> > > > option
> > > > (just add a show_options entry to the superblock ops), but the 
> > > > problem is that although each namespace has a parent user_ns, 
> > > > there's no way to get it without digging in the namespace
> > > > specific 
> > > > structure.  Probably we should restructure to move it into 
> > > > ns_common, then we could display it (and enforce all namespaces
> > > > having owning user_ns) but it would be a
> > > 
> > > I'm missing something here. Is it not already the case that all
> > > namespaces have an owning user_ns?
> > 
> > Um, yes, I don't believe I said they don't.  The problem I thought
> > you
> > were having is that there's no way of seeing what it is.
> > 
> > nsfs is the Namespace fileystem where bound namespaces appear to a
> > cat
> > of /proc/self/mounts.  It can display any information that's in
> > ns_common (the common core of namespaces) but the owning user_ns
> > pointer currently isn't in this structure.  Every user namespace
> > has a
> > pointer to it, but they're all privately embedded in the individual
> > namespace specific structures.  What I was proposing was that since
> > every current namespace has a pointer somewhere to the owning user
> > namespace, we could abstract this out into ns_common so it's now
> > accessible to be displayed by nsfs, probably as a mount option.
> 
> James, I am not sure that I understood you correctly. We have one
> file system for all namespace files, how we can show per-file 
> properties in mount options.

We have two ways of getting information.  For a namespace that only
exists as a bind mount we only have what the mount/mountinfo shows, so
you see something like this:

jejb@jarvis:~> mount|grep nsfs
nsfs on /run/build-container/userns type nsfs (rw)
nsfs on /run/build-container/ppc64 type nsfs (rw)

the (rw) are the mount options.  We could add the ability to add other
mount options to this via the superblock .show_options callback.  We
could make it show the type and parent user namespace.

>  I think we can show all required information in fdinfo. We open a
> namespaces file (/proc/pid/ns/N) and then read /proc/pid/fdinfo/X for
> it.

Not if we don't have an extant process in the namespace, we can't use
these files because they don't exist, plus fdinfo on the
/proc/<pid>/ns/X doesn't tell you what the parent user_ns of X is
(again, we could add this information somewhere ... not sure where
yet).

James

[toc] | [prev] | [next] | [standalone]


#1439303

From"Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com>
Date2016-07-08 13:20 +0200
Message-ID<rSBZ8-42k-7@gated-at.bofh.it>
In reply to#1438906
On 07/07/2016 09:17 PM, James Bottomley wrote:
> On Thu, 2016-07-07 at 20:21 +0200, Michael Kerrisk (man-pages) wrote:
>> On 7 July 2016 at 17:01, James Bottomley
>> <James.Bottomley@hansenpartnership.com> wrote:
> [Serge already answered the parenting issue]
>>> On Thu, 2016-07-07 at 08:36 -0500, Serge E. Hallyn wrote:
>>>> Hm.  Probably best-effort based on the process hierarchy.  So
>>>> yeah you could probably get a tree into a state that would be
>>>> wrongly recreated. Create a new netns, bind mount it, exit;  Have
>>>> another task create a new user_ns, bind mount it, exit;  Third
>>>> task setns()s first to the new netns then to the new user_ns.  I
>>>> suspect criu will recreate that wrongly.
>>>
>>> This is a bit pathological, and you have to be root to do it: so
>>> root can set up a nesting hierarchy, bind it and destroy the pids
>>> but I know of no current orchestration system which does this.
>>>
>>> Actually, I have to back pedal a bit: the way I currently set up
>>> architecture emulation containers does precisely this: I set up the
>>> namespaces unprivileged with child mount namespaces, but then I ask
>>> root to bind the userns and kill the process that created it so I
>>> have a permanent handle to enter the namespace by, so I suspect
>>> that when our current orchestration systems get more sophisticated,
>>> they might eventually want to do something like this as well.
>>>
>>> In theory, we could get nsfs to show this information as an option
>>> (just add a show_options entry to the superblock ops), but the
>>> problem is that although each namespace has a parent user_ns,
>>> there's no way to get it without digging in the namespace specific
>>> structure.  Probably we should restructure to move it into
>>> ns_common, then we could display it (and enforce all namespaces
>>> having owning user_ns) but it would be a
>>
>> I'm missing something here. Is it not already the case that all
>> namespaces have an owning user_ns?
>
> Um, yes, I don't believe I said they don't.  The problem I thought you
> were having is that there's no way of seeing what it is.

Your words "and enforce all namespaces having owning user_ns" were
what left me puzzled--it sounded to me that the implication was
that this is not "enforced" right now.

Cheers,

Michael

-- 
Michael Kerrisk
Linux man-pages maintainer; http://www.kernel.org/doc/man-pages/
Linux/UNIX System Programming Training: http://man7.org/training/

[toc] | [prev] | [next] | [standalone]


#1439882

From"W. Trevor King" <wking@tremily.us>
Date2016-07-09 05:20 +0200
Message-ID<rSQY9-5CR-3@gated-at.bofh.it>
In reply to#1438666

[Multipart message — attachments visible in raw view] — view raw

On Thu, Jul 07, 2016 at 08:01:52AM -0700, James Bottomley wrote:
> In theory, we could get nsfs to show this information as an option
> (just add a show_options entry to the superblock ops), but the
> problem is that although each namespace has a parent user_ns,
> there's no way to get it without digging in the namespace specific
> structure.  Probably we should restructure to move it into
> ns_common, then we could display it (and enforce all namespaces
> having owning user_ns) but it would be a reasonably large (but
> mechanical) change.

It sounds like everyone is either positive or or neutral on this
groundwork, even if we haven't decided if/how to expose the
information to userspace.  I'm happy to work up a patch while the rest
of the discussion continues.  I'm also happy to let someone else work
up the patch, if anyone else is chomping at the bit ;).

Cheers,
Trevor

-- 
This email may be signed or encrypted with GnuPG (http://www.gnupg.org).
For more information, see http://en.wikipedia.org/wiki/Pretty_Good_Privacy

[toc] | [prev] | [next] | [standalone]


#1439886

Fromebiederm@xmission.com (Eric W. Biederman)
Date2016-07-09 05:30 +0200
Message-ID<rSR7P-5HH-1@gated-at.bofh.it>
In reply to#1439882
"W. Trevor King" <wking@tremily.us> writes:

> On Thu, Jul 07, 2016 at 08:01:52AM -0700, James Bottomley wrote:
>> In theory, we could get nsfs to show this information as an option
>> (just add a show_options entry to the superblock ops), but the
>> problem is that although each namespace has a parent user_ns,
>> there's no way to get it without digging in the namespace specific
>> structure.  Probably we should restructure to move it into
>> ns_common, then we could display it (and enforce all namespaces
>> having owning user_ns) but it would be a reasonably large (but
>> mechanical) change.
>
> It sounds like everyone is either positive or or neutral on this
> groundwork, even if we haven't decided if/how to expose the
> information to userspace.  I'm happy to work up a patch while the rest
> of the discussion continues.  I'm also happy to let someone else work
> up the patch, if anyone else is chomping at the bit ;).

I am dubious on moving all of the user namespace members into ns_common.

I would happy to be proved wrong but I suspect in the cases where we
actually use that user namespace the code will become uglier.  Making
the ordinary uses uglier to make a rare corner case nicer is the wrong
trade off.

But feel free to try it is certainly worth doing if it doesn't make the
code that uses the user namespaces uglier.

Eric

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web