Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1443665 > unrolled thread
| Started by | Andrey Vagin <avagin@openvz.org> |
|---|---|
| First post | 2016-07-14 20:30 +0200 |
| Last post | 2016-07-24 07:10 +0200 |
| Articles | 9 on this page of 29 — 6 participants |
Back to article view | Back to linux.kernel
[PATCH 0/5 RFC] Add an interface to discover relationships between namespaces Andrey Vagin <avagin@openvz.org> - 2016-07-14 20:30 +0200
[PATCH 2/5] kernel: add a helper to get an owning user namespace for a namespace Andrey Vagin <avagin@openvz.org> - 2016-07-14 20:30 +0200
Re: [PATCH 2/5] kernel: add a helper to get an owning user namespace for a namespace "W. Trevor King" <wking@tremily.us> - 2016-07-14 21:10 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces Andrey Vagin <avagin@openvz.org> - 2016-07-15 00:10 +0200
[PATCH 5/5] tools/testing: add a test to check nsfs ioctl-s Andrey Vagin <avagin@openvz.org> - 2016-07-15 04:20 +0200
[PATCH 4/5] nsfs: add ioctl to get a parent namespace Andrey Vagin <avagin@openvz.org> - 2016-07-15 04:20 +0200
Re: [PATCH 4/5] nsfs: add ioctl to get a parent namespace ebiederm@xmission.com (Eric W. Biederman) - 2016-07-24 07:30 +0200
[PATCH 3/5] nsfs: add ioctl to get an owning user namespace for ns file descriptor Andrey Vagin <avagin@openvz.org> - 2016-07-15 04:20 +0200
[PATCH 2/5] kernel: add a helper to get an owning user namespace for a namespace Andrey Vagin <avagin@openvz.org> - 2016-07-15 04:20 +0200
Re: [PATCH 2/5] kernel: add a helper to get an owning user namespace for a namespace ebiederm@xmission.com (Eric W. Biederman) - 2016-07-24 07:20 +0200
Re: [PATCH 2/5] kernel: add a helper to get an owning user namespace for a namespace ebiederm@xmission.com (Eric W. Biederman) - 2016-07-24 16:50 +0200
Re: [PATCH 2/5] kernel: add a helper to get an owning user namespace for a namespace "W. Trevor King" <wking@tremily.us> - 2016-07-24 19:10 +0200
Re: [PATCH 2/5] kernel: add a helper to get an owning user namespace for a namespace "W. Trevor King" <wking@tremily.us> - 2016-07-24 19:00 +0200
Re: [PATCH 1/5] namespaces: move user_ns into ns_common ebiederm@xmission.com (Eric W. Biederman) - 2016-07-24 07:20 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces ebiederm@xmission.com (Eric W. Biederman) - 2016-07-24 07:30 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-21 16:50 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces Andrey Vagin <avagin@openvz.org> - 2016-07-22 20:30 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-25 13:50 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces ebiederm@xmission.com (Eric W. Biederman) - 2016-07-25 15:40 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> - 2016-07-25 16:50 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces "Serge E. Hallyn" <serge@hallyn.com> - 2016-07-25 17:00 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces ebiederm@xmission.com (Eric W. Biederman) - 2016-07-25 17:40 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces ebiederm@xmission.com (Eric W. Biederman) - 2016-07-25 17:20 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces "W. Trevor King" <wking@tremily.us> - 2016-07-23 23:20 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-07-23 23:40 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces "W. Trevor King" <wking@tremily.us> - 2016-07-24 00:10 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces ebiederm@xmission.com (Eric W. Biederman) - 2016-07-24 00:20 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces "W. Trevor King" <wking@tremily.us> - 2016-07-24 00:40 +0200
Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces ebiederm@xmission.com (Eric W. Biederman) - 2016-07-24 07:10 +0200
Page 2 of 2 — ← Prev page 1 [2]
| From | "Serge E. Hallyn" <serge@hallyn.com> |
|---|---|
| Date | 2016-07-25 17:00 +0200 |
| Subject | Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces |
| Message-ID | <rYPwm-5P2-3@gated-at.bofh.it> |
| In reply to | #1449546 |
Quoting Michael Kerrisk (man-pages) (mtk.manpages@gmail.com): > Hi Eric, > > On 07/25/2016 03:18 PM, Eric W. Biederman wrote: > >"Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> writes: > > > >>Hi Andrey, > >> > >>On 07/22/2016 08:25 PM, Andrey Vagin wrote: > >>>On Thu, Jul 21, 2016 at 11:48 PM, Michael Kerrisk (man-pages) > >>><mtk.manpages@gmail.com> wrote: > >>>>Hi Andrey, > >>>> > >>>> > >>>>On 07/21/2016 11:06 PM, Andrew Vagin wrote: > >>>>> > >>>>>On Thu, Jul 21, 2016 at 04:41:12PM +0200, Michael Kerrisk (man-pages) > >>>>>wrote: > >>>>>> > >>>>>>Hi Andrey, > >>>>>> > >>>>>>On 07/14/2016 08:20 PM, Andrey Vagin wrote: > >>>>> > >>>>> > >>>>><snip> > >>>>> > >>>>>> > >>>>>>Could you add here an of the API in detail: what do these FDs refer to, > >>>>>>and how do you use them to solve the use case? And could you you add > >>>>>>that info to the commit messages please. > >>>>> > >>>>> > >>>>>Hi Michael, > >>>>> > >>>>>A patch for man-pages is attached. It adds the following text to > >>>>>namespaces(7). > >>>>> > >>>>>Since Linux 4.X, the following ioctl(2) calls are supported for names‐ > >>>>>pace file descriptors. The correct syntax is: > >>>>> > >>>>> fd = ioctl(ns_fd, ioctl_type); > >>>>> > >>>>>where ioctl_type is one of the following: > >>>>> > >>>>>NS_GET_USERNS > >>>>> Returns a file descriptor that refers to an owning user names‐ > >>>>> pace. > >>>>> > >>>>>NS_GET_PARENT > >>>>> Returns a file descriptor that refers to a parent namespace. > >>>>> This ioctl(2) can be used for pid and user namespaces. For user > >>>>> namespaces, NS_GET_PARENT and NS_GET_USERNS have the same mean‐ > >>>>> ing. > >> > >>For each of the above, I think it is worth mentioning that the > >>close-on-exec flag is set for the returned file descriptor. > > > >Hmm. That is an odd default. > > Why do you say that? It's pretty common as the default for various > APIs that create new FDs these days. (There's of course a strong argument > that the original UNIX default was a design blunder...) > > >>>>> > >>>>>In addition to generic ioctl(2) errors, the following specific ones can > >>>>>occur: > >>>>> > >>>>>EINVAL NS_GET_PARENT was called for a nonhierarchical namespace. > >>>>> > >>>>>EPERM The requested namespace is outside of the current namespace > >>>>> scope. > >> > >>Perhaps add "and the caller does not have CAP_SYS_ADMIN" in the initial > >>user namespace"? > > > >Having looked at that bit of code I don't think capabilities really > >have a role to play. > > Yes, I caught up with that now. I await to see how this plays out > in the next patch version. Thanks - that had caught my eye but I hadn't had time to look into the justification for this. Hiding this kind of thing indeed seems wrong to me, unless there is a really good justification for it, i.e. a way to use that info in an exploit.
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2016-07-25 17:40 +0200 |
| Message-ID | <rYQ94-6hx-23@gated-at.bofh.it> |
| In reply to | #1449552 |
"Serge E. Hallyn" <serge@hallyn.com> writes: > Quoting Michael Kerrisk (man-pages) (mtk.manpages@gmail.com): >> Hi Eric, >> >> On 07/25/2016 03:18 PM, Eric W. Biederman wrote: >> >"Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> writes: >> > >> >>Hi Andrey, >> >> >> >>On 07/22/2016 08:25 PM, Andrey Vagin wrote: >> >>Perhaps add "and the caller does not have CAP_SYS_ADMIN" in the initial >> >>user namespace"? >> > >> >Having looked at that bit of code I don't think capabilities really >> >have a role to play. >> >> Yes, I caught up with that now. I await to see how this plays out >> in the next patch version. > > Thanks - that had caught my eye but I hadn't had time to look into the > justification for this. Hiding this kind of thing indeed seems wrong to > me, unless there is a really good justification for it, i.e. a way > to use that info in an exploit. To avoid breaking checkpoint/restart we need to limit information to the namespaces the caller is a member of for the user and pid namespaces. This roughly duplicates the parentage checks in ns_capable. Conceptually this is the same as limiting .. in a chroot environment. Eric
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2016-07-25 17:20 +0200 |
| Message-ID | <rYPPH-6aR-13@gated-at.bofh.it> |
| In reply to | #1449546 |
"Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> writes:
> Hi Eric,
>
> On 07/25/2016 03:18 PM, Eric W. Biederman wrote:
>> "Michael Kerrisk (man-pages)" <mtk.manpages@gmail.com> writes:
>>
>>> Hi Andrey,
>>>
>>> On 07/22/2016 08:25 PM, Andrey Vagin wrote:
>>>> On Thu, Jul 21, 2016 at 11:48 PM, Michael Kerrisk (man-pages)
>>>> <mtk.manpages@gmail.com> wrote:
>>>>> Hi Andrey,
>>>>>
>>>>>
>>>>> On 07/21/2016 11:06 PM, Andrew Vagin wrote:
>>>>>>
[snip]
>>>>>> where ioctl_type is one of the following:
>>>>>>
>>>>>> NS_GET_USERNS
>>>>>> Returns a file descriptor that refers to an owning user names‐
>>>>>> pace.
>>>>>>
>>>>>> NS_GET_PARENT
>>>>>> Returns a file descriptor that refers to a parent namespace.
>>>>>> This ioctl(2) can be used for pid and user namespaces. For user
>>>>>> namespaces, NS_GET_PARENT and NS_GET_USERNS have the same mean‐
>>>>>> ing.
>>>
>>> For each of the above, I think it is worth mentioning that the
>>> close-on-exec flag is set for the returned file descriptor.
>>
>> Hmm. That is an odd default.
>
> Why do you say that? It's pretty common as the default for various
> APIs that create new FDs these days. (There's of course a strong argument
> that the original UNIX default was a design blunder...)
Interesting. I haven't kept up on that, but it seems reasonable.
[snip]
>>> So, from my point of view, the important piece that was missing from
>>> your commit message was the note to use readlink("/proc/self/fd/%d")
>>> on the returned FDs. I think that detail needs to be part of the
>>> commit message (and also the man page text). I think it even be
>>> helpful to include the above program as part of the commit message:
>>> it helps people more quickly grasp the API.
>>
>> Please, please make the standard way to compare these things fstat.
>> That is much less magic than a symlink, and a little more future proof.
>> Possibly even kcmp.
>
> As in fstat() to get the st_ino field, right?
Both the st_ino and st_dev fields.
The most likely change to support checkpoint/restart in the future is to
preserve st_ino across migrations and instantiate a different instance
of nsfs to hold the inode numbers from the previous machine.
We would need to handle the preservation carefully or else there is
a chance that two namespace file descriptors (collected from different
sources) with different st_dev and st_ino fields may actuall refer to
the same object.
Which is a long way of saying we have the st_dev field please use it,
it may matter at some point.
Eric
[toc] | [prev] | [next] | [standalone]
| From | "W. Trevor King" <wking@tremily.us> |
|---|---|
| Date | 2016-07-23 23:20 +0200 |
| Subject | Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces |
| Message-ID | <rYcv0-7Gq-9@gated-at.bofh.it> |
| In reply to | #1443665 |
[Multipart message — attachments visible in raw view] — view raw
On Thu, Jul 14, 2016 at 11:20:14AM -0700, Andrey Vagin wrote:
> Pid and user namepaces are hierarchical. There is no way to discover
> parent-child relationships too.
It bothers me that network namespaces are not hierarchical too ;).
namespaces(7) and clone(2) both have:
When a network namespace is freed (i.e., when the last process in
the namespace terminates), its physical network devices are moved
back to the initial network namespace (not to the parent of the
process).
So the initial network namespace (the head of net_namespace_list?) is
special [1]. To understand how physical network devices will be
handled, it seems like we want to treat network devices as a depth-1
tree, with all non-initial net namespaces as children of the initial
net namespace. Can we extend this series' NS_GET_PARENT to return:
* EPERM for an unprivileged caller (like this series currently does
for PID namespaces),
* ENOENT when called on net_namespace_list, and
* net_namespace_list when called on any other net namespace.
If that sounds reasonable, I'm happy to stumble my way through a patch
;).
And one benefit of the net_namespace_list approach is that it will be
really easy to walk children if we ever add a parent → children lookup
service to mirror this series' child → parent service.
Cheers,
Trevor
[1]: The commit message for 2b035b39 (net: Batch network namespace
destruction, 2009-11-29) opens with:
It is fairly common to kill several network namespaces at once.
Either because they are nested one inside the other or…
which I'm having trouble understanding if network namespaces aren't
hierarchical (and they don't seem to be, except for the initial
network namespace being special). Maybe nested network namespaces
were on the table at one point but never materialized?
net->list looks like a reference to that namespace's entry in
net_namespace_list, and I didn't see anything else that looked like
a reference to a parent or list of children.
--
This email may be signed or encrypted with GnuPG (http://www.gnupg.org).
For more information, see http://en.wikipedia.org/wiki/Pretty_Good_Privacy
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2016-07-23 23:40 +0200 |
| Subject | Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces |
| Message-ID | <rYcOm-7MO-13@gated-at.bofh.it> |
| In reply to | #1449006 |
[Multipart message — attachments visible in raw view] — view raw
On Sat, 2016-07-23 at 14:14 -0700, W. Trevor King wrote: > On Thu, Jul 14, 2016 at 11:20:14AM -0700, Andrey Vagin wrote: > > Pid and user namepaces are hierarchical. There is no way to > > discover parent-child relationships too. > > It bothers me that network namespaces are not hierarchical too ;). Well, there's a reason for that: mapping namespaces need to be be hierarchical because the mapping may be remapped; The initial point for creating a new namespace is the mapped endpoint of the old one. Label based namespaces don't really have any need to be. > namespaces(7) and clone(2) both have: > > When a network namespace is freed (i.e., when the last process in > the namespace terminates), its physical network devices are moved > back to the initial network namespace (not to the parent of the > process). > > So the initial network namespace (the head of net_namespace_list?) is > special [1]. To understand how physical network devices will be > handled, it seems like we want to treat network devices as a depth-1 > tree, with all non-initial net namespaces as children of the initial > net namespace. Can we extend this series' NS_GET_PARENT to return: > > * EPERM for an unprivileged caller (like this series currently does > for PID namespaces), > * ENOENT when called on net_namespace_list, and > * net_namespace_list when called on any other net namespace. What's the practical application of this? independent net namespaces are managed by the ip netns command. It pins them by a bind mount in a flat fashion; if we make them hierarchical the tool would probably need updating to reflect this, so we're going to need a reason to give the network people. Just having the interfaces not go back to root when you do an ip netns delete doesn't seem very compelling. James
[toc] | [prev] | [next] | [standalone]
| From | "W. Trevor King" <wking@tremily.us> |
|---|---|
| Date | 2016-07-24 00:10 +0200 |
| Subject | Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces |
| Message-ID | <rYdho-8bU-29@gated-at.bofh.it> |
| In reply to | #1449008 |
[Multipart message — attachments visible in raw view] — view raw
On Sat, Jul 23, 2016 at 02:38:56PM -0700, James Bottomley wrote: > On Sat, 2016-07-23 at 14:14 -0700, W. Trevor King wrote: > > namespaces(7) and clone(2) both have: > > > > When a network namespace is freed (i.e., when the last process > > in the namespace terminates), its physical network devices are > > moved back to the initial network namespace (not to the parent > > of the process). > > > > So the initial network namespace (the head of net_namespace_list?) > > is special [1]. To understand how physical network devices will > > be handled, it seems like we want to treat network devices as a > > depth-1 tree, with all non-initial net namespaces as children of > > the initial net namespace. Can we extend this series' > > NS_GET_PARENT to return: > > > > * EPERM for an unprivileged caller (like this series currently does > > for PID namespaces), > > * ENOENT when called on net_namespace_list, and > > * net_namespace_list when called on any other net namespace. > > What's the practical application of this? independent net > namespaces are managed by the ip netns command. It pins them by a > bind mount in a flat fashion; if we make them hierarchical the tool > would probably need updating to reflect this, so we're going to need > a reason to give the network people. Just having the interfaces not > go back to root when you do an ip netns delete doesn't seem very > compelling. I'm not suggesting we add support for deeper nesting, I'm suggesting we use NS_GET_PARENT to allow sufficiently privileged users to determine if a given net namespace is the initial net namespace. You could do this already with something like: 1. Create a new net namespace. 2. Add a physical network device to that namespace. 3. Delete that namespace. 4. See if the physical network device shows up in your initial-net-namespace candidate. 5. Delete the physical network device (hopefully it ended up somewhere you can find it ;). But using an NS_GET_PARENT call seems much safer and easier. Cheers, Trevor -- This email may be signed or encrypted with GnuPG (http://www.gnupg.org). For more information, see http://en.wikipedia.org/wiki/Pretty_Good_Privacy
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2016-07-24 00:20 +0200 |
| Message-ID | <rYdr3-8f1-5@gated-at.bofh.it> |
| In reply to | #1449011 |
"W. Trevor King" <wking@tremily.us> writes: 2> On Sat, Jul 23, 2016 at 02:38:56PM -0700, James Bottomley wrote: >> On Sat, 2016-07-23 at 14:14 -0700, W. Trevor King wrote: >> > namespaces(7) and clone(2) both have: >> > >> > When a network namespace is freed (i.e., when the last process >> > in the namespace terminates), its physical network devices are >> > moved back to the initial network namespace (not to the parent >> > of the process). >> > >> > So the initial network namespace (the head of net_namespace_list?) >> > is special [1]. To understand how physical network devices will >> > be handled, it seems like we want to treat network devices as a >> > depth-1 tree, with all non-initial net namespaces as children of >> > the initial net namespace. Can we extend this series' >> > NS_GET_PARENT to return: >> > >> > * EPERM for an unprivileged caller (like this series currently does >> > for PID namespaces), >> > * ENOENT when called on net_namespace_list, and >> > * net_namespace_list when called on any other net namespace. >> >> What's the practical application of this? independent net >> namespaces are managed by the ip netns command. It pins them by a >> bind mount in a flat fashion; if we make them hierarchical the tool >> would probably need updating to reflect this, so we're going to need >> a reason to give the network people. Just having the interfaces not >> go back to root when you do an ip netns delete doesn't seem very >> compelling. > > I'm not suggesting we add support for deeper nesting, I'm suggesting > we use NS_GET_PARENT to allow sufficiently privileged users to > determine if a given net namespace is the initial net namespace. You > could do this already with something like: > > 1. Create a new net namespace. > 2. Add a physical network device to that namespace. > 3. Delete that namespace. > 4. See if the physical network device shows up in your > initial-net-namespace candidate. > 5. Delete the physical network device (hopefully it ended up somewhere > you can find it ;). > > But using an NS_GET_PARENT call seems much safer and easier. Have you had the problem in practice where you can't tell which network namespace is the initial network namespace. This all seems like a theoretical problem rather than a real one. Eric
[toc] | [prev] | [next] | [standalone]
| From | "W. Trevor King" <wking@tremily.us> |
|---|---|
| Date | 2016-07-24 00:40 +0200 |
| Subject | Re: [PATCH 0/5 RFC] Add an interface to discover relationships between namespaces |
| Message-ID | <rYdKp-8kR-7@gated-at.bofh.it> |
| In reply to | #1449012 |
[Multipart message — attachments visible in raw view] — view raw
On Sat, Jul 23, 2016 at 04:56:44PM -0500, Eric W. Biederman wrote: > "W. Trevor King" <wking@tremily.us> writes: > > On Sat, Jul 23, 2016 at 02:38:56PM -0700, James Bottomley wrote: > >> On Sat, 2016-07-23 at 14:14 -0700, W. Trevor King wrote: > >> > namespaces(7) and clone(2) both have: > >> > > >> > When a network namespace is freed (i.e., when the last > >> > process in the namespace terminates), its physical network > >> > devices are moved back to the initial network namespace (not > >> > to the parent of the process). > >> > > >> > So the initial network namespace (the head of > >> > net_namespace_list?) is special [1]. To understand how > >> > physical network devices will be handled, it seems like we want > >> > to treat network devices as a depth-1 tree, with all > >> > non-initial net namespaces as children of the initial net > >> > namespace. Can we extend this series' NS_GET_PARENT to return: > >> > > >> > * EPERM for an unprivileged caller (like this series currently > >> > does for PID namespaces), > >> > * ENOENT when called on net_namespace_list, and > >> > * net_namespace_list when called on any other net namespace. > >> > >> What's the practical application of this? independent net > >> namespaces are managed by the ip netns command. It pins them by > >> a bind mount in a flat fashion; if we make them hierarchical the > >> tool would probably need updating to reflect this, so we're going > >> to need a reason to give the network people. Just having the > >> interfaces not go back to root when you do an ip netns delete > >> doesn't seem very compelling. > > > > I'm not suggesting we add support for deeper nesting, I'm suggesting > > we use NS_GET_PARENT to allow sufficiently privileged users to > > determine if a given net namespace is the initial net namespace. You > > could do this already with something like: > > > > 1. Create a new net namespace. > > 2. Add a physical network device to that namespace. > > 3. Delete that namespace. > > 4. See if the physical network device shows up in your > > initial-net-namespace candidate. > > 5. Delete the physical network device (hopefully it ended up > > somewhere you can find it ;). > > > > But using an NS_GET_PARENT call seems much safer and easier. > > Have you had the problem in practice where you can't tell which > network namespace is the initial network namespace. This all seems > like a theoretical problem rather than a real one. I haven't had any practical problems here, I'm just trying to wrap my head around namespace-relationship discovery. The special physical network device handling seems a lot like init re-parenting (with no PR_SET_CHILD_SUBREAPER analog in a 1-deep namespace tree), so calling the initial network namespace a parent (and all the other namespaces its direct children) seems natural enough. If that doesn't sound convincing, I'm happy to punt this idea until someone runs into a practical problem ;). Cheers, Trevor -- This email may be signed or encrypted with GnuPG (http://www.gnupg.org). For more information, see http://en.wikipedia.org/wiki/Pretty_Good_Privacy
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2016-07-24 07:10 +0200 |
| Message-ID | <rYjPP-3JA-1@gated-at.bofh.it> |
| In reply to | #1449013 |
"W. Trevor King" <wking@tremily.us> writes: > On Sat, Jul 23, 2016 at 04:56:44PM -0500, Eric W. Biederman wrote: >> "W. Trevor King" <wking@tremily.us> writes: >> > On Sat, Jul 23, 2016 at 02:38:56PM -0700, James Bottomley wrote: >> >> On Sat, 2016-07-23 at 14:14 -0700, W. Trevor King wrote: >> >> > namespaces(7) and clone(2) both have: >> >> > >> >> > When a network namespace is freed (i.e., when the last >> >> > process in the namespace terminates), its physical network >> >> > devices are moved back to the initial network namespace (not >> >> > to the parent of the process). >> >> > >> >> > So the initial network namespace (the head of >> >> > net_namespace_list?) is special [1]. To understand how >> >> > physical network devices will be handled, it seems like we want >> >> > to treat network devices as a depth-1 tree, with all >> >> > non-initial net namespaces as children of the initial net >> >> > namespace. Can we extend this series' NS_GET_PARENT to return: >> >> > >> >> > * EPERM for an unprivileged caller (like this series currently >> >> > does for PID namespaces), >> >> > * ENOENT when called on net_namespace_list, and >> >> > * net_namespace_list when called on any other net namespace. >> >> >> >> What's the practical application of this? independent net >> >> namespaces are managed by the ip netns command. It pins them by >> >> a bind mount in a flat fashion; if we make them hierarchical the >> >> tool would probably need updating to reflect this, so we're going >> >> to need a reason to give the network people. Just having the >> >> interfaces not go back to root when you do an ip netns delete >> >> doesn't seem very compelling. >> > >> > I'm not suggesting we add support for deeper nesting, I'm suggesting >> > we use NS_GET_PARENT to allow sufficiently privileged users to >> > determine if a given net namespace is the initial net namespace. You >> > could do this already with something like: >> > >> > 1. Create a new net namespace. >> > 2. Add a physical network device to that namespace. >> > 3. Delete that namespace. >> > 4. See if the physical network device shows up in your >> > initial-net-namespace candidate. >> > 5. Delete the physical network device (hopefully it ended up >> > somewhere you can find it ;). >> > >> > But using an NS_GET_PARENT call seems much safer and easier. >> >> Have you had the problem in practice where you can't tell which >> network namespace is the initial network namespace. This all seems >> like a theoretical problem rather than a real one. > > I haven't had any practical problems here, I'm just trying to wrap my > head around namespace-relationship discovery. The special physical > network device handling seems a lot like init re-parenting (with no > PR_SET_CHILD_SUBREAPER analog in a 1-deep namespace tree), so calling > the initial network namespace a parent (and all the other namespaces > its direct children) seems natural enough. If that doesn't sound > convincing, I'm happy to punt this idea until someone runs into a > practical problem ;). Then let's punt this until someone runs into a practical problem. For scaling and for sanity it is desirable to keep the connections between namespaces to a minimum. Further the initial instances of a namespace always tend to be a little bit special. Eric
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web