Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1647121 > unrolled thread
| Started by | David Howells <dhowells@redhat.com> |
|---|---|
| First post | 2017-05-22 18:30 +0200 |
| Last post | 2017-05-23 17:40 +0200 |
| Articles | 20 on this page of 37 — 9 participants |
Back to article view | Back to linux.kernel
[RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-22 18:30 +0200
[PATCH 8/9] Honour CONTAINER_NEW_EMPTY_FS_NS David Howells <dhowells@redhat.com> - 2017-05-22 18:30 +0200
[PATCH 3/9] Provide /proc/containers David Howells <dhowells@redhat.com> - 2017-05-22 18:30 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-05-22 19:00 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Aleksa Sarai <asarai@suse.de> - 2017-05-22 19:20 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 17:00 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-23 17:10 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 17:20 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-23 17:30 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-05-23 17:50 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 18:40 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-24 10:40 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Ian Kent <raven@themaw.net> - 2017-05-24 11:20 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Jessica Frazelle <me@jessfraz.com> - 2017-05-22 19:30 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Jeff Layton <jlayton@redhat.com> - 2017-05-22 20:40 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-05-22 21:30 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Jeff Layton <jlayton@redhat.com> - 2017-05-23 00:20 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Ian Kent <raven@themaw.net> - 2017-05-23 12:40 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Ian Kent <raven@themaw.net> - 2017-05-23 11:40 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 16:00 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-05-23 17:10 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-23 17:40 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Jessica Frazelle <me@jessfraz.com> - 2017-05-22 19:20 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 17:20 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-22 21:20 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Jeff Layton <jlayton@redhat.com> - 2017-05-23 00:30 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-23 15:10 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Jeff Layton <jlayton@redhat.com> - 2017-05-23 16:30 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Djalal Harouni <tixxdz@gmail.com> - 2017-05-23 16:40 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Colin Walters <walters@verbum.org> - 2017-05-23 17:00 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Colin Walters <walters@verbum.org> - 2017-05-23 17:40 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 17:40 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Jeff Layton <jlayton@redhat.com> - 2017-05-23 17:40 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Djalal Harouni <tixxdz@gmail.com> - 2017-05-23 16:30 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 18:20 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects Ian Kent <raven@themaw.net> - 2017-05-23 12:20 +0200
Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-23 17:40 +0200
Page 1 of 2 [1] 2 Next page →
| From | David Howells <dhowells@redhat.com> |
|---|---|
| Date | 2017-05-22 18:30 +0200 |
| Subject | [RFC][PATCH 0/9] Make containers kernel objects |
| Message-ID | <tJYnv-80c-3@gated-at.bofh.it> |
Here are a set of patches to define a container object for the kernel and
to provide some methods to create and manipulate them.
The reason I think this is necessary is that the kernel has no idea how to
direct upcalls to what userspace considers to be a container - current
Linux practice appears to make a "container" just an arbitrarily chosen
junction of namespaces, control groups and files, which may be changed
individually within the "container".
The kernel upcall mechanism then needs to decide which set of namespaces,
etc., it must exec the appropriate upcall program. Examples of this
include:
(1) The DNS resolver. The DNS cache in the kernel should probably be
per-network namespace, but in userspace the program, its libraries and
its config data are associated with a mount tree and a user namespace
and it gets run in a particular pid namespace.
(2) NFS ID mapper. The NFS ID mapping cache should also probably be
per-network namespace.
(3) nfsdcltrack. A way for NFSD to access stable storage for tracking
of persistent state. Again, network-namespace dependent, but also
perhaps mount-namespace dependent.
(4) General request-key upcalls. Not particularly namespace dependent,
apart from keyrings being somewhat governed by the user namespace and
the upcall being configured by the mount namespace.
These patches are built on top of the mount context patchset so that
namespaces can be properly propagated over submounts/automounts.
These patches implement a container object that holds the following things:
(1) Namespaces.
(2) A root directory.
(3) A set of processes, including a designated 'init' process.
(4) The creator's credentials, including ownership.
(5) A place to hang security for the container, allowing policies to be
set per-container.
I also want to add:
(6) Control groups.
(7) A per-container keyring that can be added to from outside of the
container, even once the container is live, for the provision of
filesystem authentication/encryption keys in advance of the container
being started.
You can get a list of containers by examining /proc/containers - but I'm
not sure how much value this gets you. Note that the container in which
you are running is called "<current>" and you can only see other containers
that were started from within yours. Containers are therefore effectively
hierarchical and an init_container is set up when the system boots.
Some management operations are provided:
(1) int fd = container_create(const char *name, unsigned int flags);
Create a container of the given name and return a handle to it as a
file descriptor. flags indicates what namespaces should be inherited
from the caller and what should be replaced new. It is possible to
set up a container with a null root filesystem that can be mounted
later.
(2) int fsfd = fsopen(const char *fsname, int container_fd,
unsigned int flags);
Prepare a mount context inside the container. This uses all the
containers namespaces instead of the caller's.
(3) fsmount(int fsfd, int dfd, const char *path, unsigned int at_flags,
unsigned int flags);
Mount a prepared superblock. dfd can be given container_fd to use the
container to which it refers as the root of the pathwalk.
If path is "/" and at_flags is AT_FSMOUNT_CONTAINER_ROOT, then this
will attempt to mount the root of the container and create a mount
namespace for it. The container must've been created with
CONTAINER_NEW_EMPTY_FS_NS.
(4) pid_t pid = fork_into_container(int container_fd);
Create the init process in a container. The process uses that
container's namespaces instead of the caller's.
(5) int sfd = container_socket(int container_fd,
int domain, int type, int protocol);
Create a socket inside a container. The socket gets the container's
namespaces. This allows netlink operations to be called within that
container to set it up from outside (at least in theory).
(6) mkdirat(int dfd, ...);
mknodat(int dfd, ...);
openat(int dfd, ...);
Supplying a container fd as dfd makes the pathwalk happen relative to
the root of the container. Note that the path must be *relative*.
And some need to be/could be added:
(7) Directly set a container's namespaces to allow cross-container
sharing.
(8) Adjust the control group membership of a container.
(9) Add a key inside a container keyring.
(10) Kill/suspend/freeze/reboot container, both from inside and out.
(11) Set container's root dir.
(12) Set the container's security policy.
(13) Allow overlayfs to access filesystems outside of the container in
which it is being created.
Kernel upcalls are invoked in the root of the container that incurs them
rather than in the init namespace context. There's still some awkwardness
here if you, say, share a network namespace between containers. Either the
upcall binaries and configuration must be duplicated between sharing
containers or a container must be elected as the one in which such upcalls
will be done.
Some further thoughts:
(*) Should there be an AT_IN_CONTAINER flag to provide to syscalls that
take a container in lieu of AT_FDCWD or a directory fd? The problem
is that such as mkdirat() and openat() don't have an at_flags
argument.
(*) Should there be a container hierarchy at all? It seems that this is
only really necessary for /proc/containers. Do we want to allow
containers-within-containers?
(*) Should each container automatically have its own pid namespace such
that its 'init' process always appears as pid 1?
(*) Does this allow kernel upcalls to be accounted against the correct
control group?
(*) Should each container have a 'list' of accessible device numbers such
that certain device files can be made usable within a container? And
can devtmpfs/udev be made to show the correct file set for each
container?
The patches can be found here also:
http://git.kernel.org/cgit/linux/kernel/git/dhowells/linux-fs.git/log/?h=container
Note that this is dependent on the mount-context branch.
David
---
David Howells (9):
containers: Rename linux/container.h to linux/container_dev.h
Implement containers as kernel objects
Provide /proc/containers
Allow processes to be forked and upcalled into a container
Open a socket inside a container
Allow fs syscall dfd arguments to take a container fd
Make fsopen() able to initiate mounting into a container
Honour CONTAINER_NEW_EMPTY_FS_NS
Sample program for driving container objects
arch/x86/entry/syscalls/syscall_32.tbl | 3
arch/x86/entry/syscalls/syscall_64.tbl | 3
drivers/acpi/container.c | 2
drivers/base/container.c | 2
fs/fsopen.c | 33 +-
fs/libfs.c | 3
fs/namei.c | 52 ++-
fs/namespace.c | 108 +++++-
fs/nfs/namespace.c | 2
fs/nfs/nfs4namespace.c | 4
fs/proc/root.c | 13 +
fs/sb_config.c | 29 +-
include/linux/container.h | 91 ++++-
include/linux/container_dev.h | 25 +
include/linux/cred.h | 3
include/linux/init_task.h | 4
include/linux/kmod.h | 1
include/linux/lsm_hooks.h | 25 +
include/linux/mount.h | 5
include/linux/nsproxy.h | 7
include/linux/pid.h | 5
include/linux/proc_ns.h | 3
include/linux/sb_config.h | 5
include/linux/sched.h | 3
include/linux/sched/task.h | 4
include/linux/security.h | 20 +
include/linux/syscalls.h | 6
include/uapi/linux/container.h | 28 ++
include/uapi/linux/fcntl.h | 2
include/uapi/linux/magic.h | 1
init/Kconfig | 7
init/main.c | 4
kernel/Makefile | 2
kernel/container.c | 576 ++++++++++++++++++++++++++++++++
kernel/cred.c | 45 ++-
kernel/exit.c | 1
kernel/fork.c | 117 ++++++-
kernel/kmod.c | 13 +
kernel/kthread.c | 3
kernel/namespaces.h | 15 +
kernel/nsproxy.c | 34 +-
kernel/pid.c | 4
kernel/sys_ni.c | 5
net/socket.c | 37 ++
samples/containers/test-container.c | 162 +++++++++
security/security.c | 18 +
security/selinux/hooks.c | 5
47 files changed, 1408 insertions(+), 132 deletions(-)
create mode 100644 include/linux/container_dev.h
create mode 100644 include/uapi/linux/container.h
create mode 100644 kernel/container.c
create mode 100644 kernel/namespaces.h
create mode 100644 samples/containers/test-container.c
[toc] | [next] | [standalone]
| From | David Howells <dhowells@redhat.com> |
|---|---|
| Date | 2017-05-22 18:30 +0200 |
| Subject | [PATCH 8/9] Honour CONTAINER_NEW_EMPTY_FS_NS |
| Message-ID | <tJYnx-80c-63@gated-at.bofh.it> |
| In reply to | #1647121 |
Allow a container to be created with an empty mount namespace, as specified
by passing CONTAINER_NEW_EMPTY_FS_NS to container_create(), and allow a
root filesystem to be mounted into the container:
cfd = container_create("foo", CONTAINER_NEW_EMPTY_FS_NS);
fd = fsopen("ext3", cfd, 0);
write(fd, "o foo");
...
fsmount(fd, -1, "/", AT_FSMOUNT_CONTAINER_ROOT, 0);
close(fd);
fd = fsopen("proc", cfd, 0);
fsmount(fd, cfd, "/proc", 0, 0);
close(fd);
---
fs/namespace.c | 84 ++++++++++++++++++++++++++++++++++++--------
include/linux/mount.h | 3 +-
include/uapi/linux/fcntl.h | 2 +
kernel/container.c | 6 +++
kernel/fork.c | 5 ++-
security/selinux/hooks.c | 2 +
6 files changed, 85 insertions(+), 17 deletions(-)
diff --git a/fs/namespace.c b/fs/namespace.c
index 9ca8b9f49f80..a365a7cba3ad 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2458,6 +2458,38 @@ static int do_add_mount(struct mount *newmnt, struct path *path, int mnt_flags,
}
static bool mount_too_revealing(struct vfsmount *mnt, int *new_mnt_flags);
+static struct mnt_namespace *create_mnt_ns(struct vfsmount *m);
+
+/*
+ * Create a mount namespace for a container and set the root mount in it.
+ */
+static int set_container_root(struct sb_config *sc, struct vfsmount *mnt)
+{
+ struct container *container = sc->container;
+ struct mnt_namespace *mnt_ns;
+ int ret = -EBUSY;
+
+ mnt_ns = create_mnt_ns(mnt);
+ if (IS_ERR(mnt_ns))
+ return PTR_ERR(mnt_ns);
+
+ spin_lock(&container->lock);
+ if (!container->ns->mnt_ns) {
+ container->ns->mnt_ns = mnt_ns;
+ write_seqcount_begin(&container->seq);
+ container->root.mnt = mnt;
+ container->root.dentry = mnt->mnt_root;
+ write_seqcount_end(&container->seq);
+ path_get(&container->root);
+ mnt_ns = NULL;
+ ret = 0;
+ }
+ spin_unlock(&container->lock);
+
+ if (ret < 0)
+ put_mnt_ns(mnt_ns);
+ return ret;
+}
/*
* Create a new mount using a superblock configuration and request it
@@ -2479,8 +2511,12 @@ static int do_new_mount_sc(struct sb_config *sc, struct path *mountpoint,
goto err_mnt;
}
- ret = do_add_mount(real_mount(mnt), mountpoint, mnt_flags,
- sc->container ? sc->container->ns->mnt_ns : NULL);
+ if (mnt_flags & MNT_CONTAINER_ROOT)
+ ret = set_container_root(sc, mnt);
+ else
+ ret = do_add_mount(real_mount(mnt), mountpoint, mnt_flags,
+ sc->container ? sc->container->ns->mnt_ns : NULL);
+
if (ret < 0) {
errorf("VFS: Failed to add mount");
goto err_mnt;
@@ -3262,10 +3298,17 @@ SYSCALL_DEFINE5(fsmount, int, fs_fd, int, dfd, const char __user *, dir_name,
struct fd f;
unsigned int lookup_flags, mnt_flags = 0;
long ret;
+ char buf[2];
if ((at_flags & ~(AT_SYMLINK_NOFOLLOW | AT_NO_AUTOMOUNT |
- AT_EMPTY_PATH)) != 0)
+ AT_EMPTY_PATH | AT_FSMOUNT_CONTAINER_ROOT)) != 0)
return -EINVAL;
+ if (at_flags & AT_FSMOUNT_CONTAINER_ROOT) {
+ if (strncpy_from_user(buf, dir_name, 2) < 0)
+ return -EFAULT;
+ if (buf[0] != '/' || buf[1] != '\0')
+ return -EINVAL;
+ }
if (flags & ~(MS_RDONLY | MS_NOSUID | MS_NODEV | MS_NOEXEC |
MS_NOATIME | MS_NODIRATIME | MS_RELATIME | MS_STRICTATIME))
@@ -3317,18 +3360,29 @@ SYSCALL_DEFINE5(fsmount, int, fs_fd, int, dfd, const char __user *, dir_name,
if (ret < 0)
goto err_fsfd;
- /* Find the mountpoint. A container can be specified in dfd. */
- lookup_flags = LOOKUP_FOLLOW | LOOKUP_AUTOMOUNT;
- if (at_flags & AT_SYMLINK_NOFOLLOW)
- lookup_flags &= ~LOOKUP_FOLLOW;
- if (at_flags & AT_NO_AUTOMOUNT)
- lookup_flags &= ~LOOKUP_AUTOMOUNT;
- if (at_flags & AT_EMPTY_PATH)
- lookup_flags |= LOOKUP_EMPTY;
- ret = user_path_at(dfd, dir_name, lookup_flags, &mountpoint);
- if (ret < 0) {
- errorf("VFS: Mountpoint lookup failed");
- goto err_fsfd;
+ if (at_flags & AT_FSMOUNT_CONTAINER_ROOT) {
+ /* We're mounting the root of the container that was specified
+ * to sys_fsopen(). The dir_name should be specified as "/"
+ * and dfd is ignored.
+ */
+ mountpoint.mnt = NULL;
+ mountpoint.dentry = NULL;
+ mnt_flags |= MNT_CONTAINER_ROOT;
+ } else {
+ /* Find the mountpoint. A container can be specified in dfd. */
+ lookup_flags = LOOKUP_FOLLOW | LOOKUP_AUTOMOUNT;
+
+ if (at_flags & AT_SYMLINK_NOFOLLOW)
+ lookup_flags &= ~LOOKUP_FOLLOW;
+ if (at_flags & AT_NO_AUTOMOUNT)
+ lookup_flags &= ~LOOKUP_AUTOMOUNT;
+ if (at_flags & AT_EMPTY_PATH)
+ lookup_flags |= LOOKUP_EMPTY;
+ ret = user_path_at(dfd, dir_name, lookup_flags, &mountpoint);
+ if (ret < 0) {
+ errorf("VFS: Mountpoint lookup failed");
+ goto err_fsfd;
+ }
}
ret = security_sb_mountpoint(sc, &mountpoint);
diff --git a/include/linux/mount.h b/include/linux/mount.h
index 265e9aa2ab0b..480c6b4061e0 100644
--- a/include/linux/mount.h
+++ b/include/linux/mount.h
@@ -51,7 +51,8 @@ struct sb_config;
#define MNT_INTERNAL_FLAGS (MNT_SHARED | MNT_WRITE_HOLD | MNT_INTERNAL | \
MNT_DOOMED | MNT_SYNC_UMOUNT | MNT_MARKED)
-#define MNT_INTERNAL 0x4000
+#define MNT_INTERNAL 0x4000
+#define MNT_CONTAINER_ROOT 0x8000 /* Mounting a container root */
#define MNT_LOCK_ATIME 0x040000
#define MNT_LOCK_NOEXEC 0x080000
diff --git a/include/uapi/linux/fcntl.h b/include/uapi/linux/fcntl.h
index 813afd6eee71..747af8704bbf 100644
--- a/include/uapi/linux/fcntl.h
+++ b/include/uapi/linux/fcntl.h
@@ -68,5 +68,7 @@
#define AT_STATX_FORCE_SYNC 0x2000 /* - Force the attributes to be sync'd with the server */
#define AT_STATX_DONT_SYNC 0x4000 /* - Don't sync attributes with the server */
+#define AT_FSMOUNT_CONTAINER_ROOT 0x2000
+
#endif /* _UAPI_LINUX_FCNTL_H */
diff --git a/kernel/container.c b/kernel/container.c
index 5ebbf548f01a..68276603d255 100644
--- a/kernel/container.c
+++ b/kernel/container.c
@@ -23,6 +23,7 @@
#include <linux/printk.h>
#include <linux/security.h>
#include <linux/proc_fs.h>
+#include <linux/mnt_namespace.h>
#include "namespaces.h"
struct container init_container = {
@@ -500,6 +501,11 @@ static struct container *create_container(const char *name, unsigned int flags)
fs->root.mnt = NULL;
fs->root.dentry = NULL;
+ if (flags & CONTAINER_NEW_EMPTY_FS_NS) {
+ put_mnt_ns(ns->mnt_ns);
+ ns->mnt_ns = NULL;
+ }
+
ret = security_container_alloc(c, flags);
if (ret < 0)
goto err_fs;
diff --git a/kernel/fork.c b/kernel/fork.c
index 68cd7367fcd5..e5111d4bcc1c 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -2169,7 +2169,10 @@ SYSCALL_DEFINE1(fork_into_container, int, containerfd)
if (is_container_file(f.file)) {
struct container *c = f.file->private_data;
- ret = _do_fork(SIGCHLD, 0, 0, NULL, NULL, 0, c);
+ if (!c->ns->mnt_ns)
+ ret = -ENOENT;
+ else
+ ret = _do_fork(SIGCHLD, 0, 0, NULL, NULL, 0, c);
}
fdput(f);
return ret;
diff --git a/security/selinux/hooks.c b/security/selinux/hooks.c
index 23bdbb0c2de5..f6b994b15a4d 100644
--- a/security/selinux/hooks.c
+++ b/security/selinux/hooks.c
@@ -2975,6 +2975,8 @@ static int selinux_sb_mountpoint(struct sb_config *sc, struct path *mountpoint)
const struct cred *cred = current_cred();
int ret;
+ if (!mountpoint->mnt)
+ return 0; /* This is the root in an empty namespace */
ret = path_has_perm(cred, mountpoint, FILE__MOUNTON);
if (ret < 0)
errorf("SELinux: Mount on mountpoint not permitted");
[toc] | [prev] | [next] | [standalone]
| From | David Howells <dhowells@redhat.com> |
|---|---|
| Date | 2017-05-22 18:30 +0200 |
| Subject | [PATCH 3/9] Provide /proc/containers |
| Message-ID | <tJYnx-80c-65@gated-at.bofh.it> |
| In reply to | #1647121 |
Provide /proc/containers to view the current container and all the
containers created within it:
# ./foo-container
NAME USE FL OWNER GROUP
<current> 141 01 0 0
foo-test 1 04 0 0
I'm not sure whether this is really desirable, though.
Signed-off-by: David Howells <dhowells@redhat.com>
---
kernel/container.c | 104 ++++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 104 insertions(+)
diff --git a/kernel/container.c b/kernel/container.c
index eef1566835eb..d5849c07a76b 100644
--- a/kernel/container.c
+++ b/kernel/container.c
@@ -22,6 +22,7 @@
#include <linux/syscalls.h>
#include <linux/printk.h>
#include <linux/security.h>
+#include <linux/proc_fs.h>
#include "namespaces.h"
struct container init_container = {
@@ -70,6 +71,108 @@ void put_container(struct container *c)
}
}
+static void *container_proc_start(struct seq_file *m, loff_t *_pos)
+{
+ struct container *c = m->private;
+ struct list_head *p;
+ loff_t pos = *_pos;
+
+ spin_lock(&c->lock);
+
+ if (pos <= 1) {
+ *_pos = 1;
+ return (void *)1UL; /* Banner on first line */
+ }
+
+ if (pos == 2)
+ return m->private; /* Current container on second line */
+
+ /* Subordinate containers thereafter */
+ p = c->children.next;
+ pos--;
+ for (pos--; pos > 0 && p != &c->children; pos--) {
+ p = p->next;
+ }
+
+ if (p == &c->children)
+ return NULL;
+ return container_of(p, struct container, child_link);
+}
+
+static void *container_proc_next(struct seq_file *m, void *v, loff_t *_pos)
+{
+ struct container *c = m->private, *vc = v;
+ struct list_head *p;
+ loff_t pos = *_pos;
+
+ pos++;
+ *_pos = pos;
+ if (pos == 2)
+ return c; /* Current container on second line */
+
+ if (pos == 3)
+ p = &c->children;
+ else
+ p = &vc->child_link;
+ p = p->next;
+ if (p == &c->children)
+ return NULL;
+ return container_of(p, struct container, child_link);
+}
+
+static void container_proc_stop(struct seq_file *m, void *v)
+{
+ struct container *c = m->private;
+
+ spin_unlock(&c->lock);
+}
+
+static int container_proc_show(struct seq_file *m, void *v)
+{
+ struct user_namespace *uns = current_user_ns();
+ struct container *c = v;
+ const char *name;
+
+ if (v == (void *)1UL) {
+ seq_puts(m, "NAME USE FL OWNER GROUP\n");
+ return 0;
+ }
+
+ name = (c == m->private) ? "<current>" : c->name;
+ seq_printf(m, "%-24s %3u %02lx %0d %5d\n",
+ name, refcount_read(&c->usage), c->flags,
+ from_kuid_munged(uns, c->cred->uid),
+ from_kgid_munged(uns, c->cred->gid));
+
+ return 0;
+}
+
+static const struct seq_operations container_proc_ops = {
+ .start = container_proc_start,
+ .next = container_proc_next,
+ .stop = container_proc_stop,
+ .show = container_proc_show,
+};
+
+static int container_proc_open(struct inode *inode, struct file *file)
+{
+ struct seq_file *m;
+ int ret = seq_open(file, &container_proc_ops);
+
+ if (ret == 0) {
+ m = file->private_data;
+ m->private = current->container;
+ }
+ return ret;
+}
+
+static const struct file_operations container_proc_fops = {
+ .open = container_proc_open,
+ .read = seq_read,
+ .llseek = seq_lseek,
+ .release = seq_release,
+};
+
/*
* Allow the user to poll for the container dying.
*/
@@ -230,6 +333,7 @@ static int __init init_container_fs(void)
panic("Cannot mount containerfs: %ld\n",
PTR_ERR(containerfs_mnt));
+ proc_create("containers", 0, NULL, &container_proc_fops);
return 0;
}
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2017-05-22 19:00 +0200 |
| Message-ID | <tJYQy-8az-29@gated-at.bofh.it> |
| In reply to | #1647121 |
[Added missing cc to containers list] On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote: > Here are a set of patches to define a container object for the kernel > and to provide some methods to create and manipulate them. > > The reason I think this is necessary is that the kernel has no idea > how to direct upcalls to what userspace considers to be a container - > current Linux practice appears to make a "container" just an > arbitrarily chosen junction of namespaces, control groups and files, > which may be changed individually within the "container". This sounds like a step in the wrong direction: the strength of the current container interfaces in Linux is that people who set up containers don't have to agree what they look like. So I can set up a user namespace without a mount namespace or an architecture emulation container with only a mount namespace. But ignoring my fun foibles with containers and to give a concrete example in terms of a popular orchestration system: in kubernetes, where certain namespaces are shared across pods, do you imagine the kernel's view of the "container" to be the pod or what kubernetes thinks of as the container? This is important, because half the examples you give below are network related and usually pods share a network namespace. > The kernel upcall mechanism then needs to decide which set of > namespaces, etc., it must exec the appropriate upcall program. > Examples of this include: > > (1) The DNS resolver. The DNS cache in the kernel should probably > be per-network namespace, but in userspace the program, its > libraries and its config data are associated with a mount tree and a > user namespace and it gets run in a particular pid namespace. All persistent (written to fs data) has to be mount ns associated; there are no ifs, ands and buts to that. I agree this implies that if you want to run a separate network namespace, you either take DNS from the parent (a lot of containers do) or you set up a daemon to run within the mount namespace. I agree the latter is a slightly fiddly operation you have to get right, but that's why we have orchestration systems. What is it we could do with the above that we cannot do today? > (2) NFS ID mapper. The NFS ID mapping cache should also probably be > per-network namespace. I think this is a view but not the only one: Right at the moment, NFS ID mapping is used as the one of the ways we can get the user namespace ID mapping writes to file problems fixed ... that makes it a property of the mount namespace for a lot of containers. There are many other instances where they do exactly as you say, but what I'm saying is that we don't want to lose the flexibility we currently have. > (3) nfsdcltrack. A way for NFSD to access stable storage for > tracking of persistent state. Again, network-namespace dependent, > but also perhaps mount-namespace dependent. So again, given we can set this up to work today, this sounds like more a restriction that will bite us than an enhancement that gives us extra features. > (4) General request-key upcalls. Not particularly namespace > dependent, apart from keyrings being somewhat governed by the user > namespace and the upcall being configured by the mount namespace. All mount namespaces have an owning user namespace, so the data relations are already there in the kernel, is the problem simply finding them? > These patches are built on top of the mount context patchset so that > namespaces can be properly propagated over submounts/automounts. I'll stop here ... you get the idea that I think this is imposing a set of restrictions that will come back to bite us later. If this is just for the sake of figuring out how to get keyring upcalls to work, then I'm sure we can come up with something. James
[toc] | [prev] | [next] | [standalone]
| From | Aleksa Sarai <asarai@suse.de> |
|---|---|
| Date | 2017-05-22 19:20 +0200 |
| Message-ID | <tJZ9U-8vX-7@gated-at.bofh.it> |
| In reply to | #1647163 |
>> The reason I think this is necessary is that the kernel has no idea >> how to direct upcalls to what userspace considers to be a container - >> current Linux practice appears to make a "container" just an >> arbitrarily chosen junction of namespaces, control groups and files, >> which may be changed individually within the "container". Just want to point out that if the kernel APIs for containers massively change, then the OCI will have to completely rework how we describe containers (and so will all existing runtimes). Not to mention that while I don't like how hard it is (from a runtime perspective) to actually set up a container securely, there are undoubtedly benefits to having namespaces split out. The network namespace being separate means that in certain contexts you actually don't want to create a new network namespace when creating a container. I had some ideas about how you could implement bridging in userspace (as an unprivileged user, for rootless containers) but if you can't join namespaces individually then such a setup is not practically possible. -- Aleksa Sarai Software Engineer (Containers) SUSE Linux GmbH https://www.cyphar.com/
[toc] | [prev] | [next] | [standalone]
| From | David Howells <dhowells@redhat.com> |
|---|---|
| Date | 2017-05-23 17:00 +0200 |
| Message-ID | <tKjrY-4z9-23@gated-at.bofh.it> |
| In reply to | #1647203 |
Aleksa Sarai <asarai@suse.de> wrote:
> >> The reason I think this is necessary is that the kernel has no idea
> >> how to direct upcalls to what userspace considers to be a container -
> >> current Linux practice appears to make a "container" just an
> >> arbitrarily chosen junction of namespaces, control groups and files,
> >> which may be changed individually within the "container".
>
> Just want to point out that if the kernel APIs for containers massively
> change, then the OCI will have to completely rework how we describe containers
> (and so will all existing runtimes).
>
> Not to mention that while I don't like how hard it is (from a runtime
> perspective) to actually set up a container securely, there are undoubtedly
> benefits to having namespaces split out. The network namespace being separate
> means that in certain contexts you actually don't want to create a new network
> namespace when creating a container.
Yep, I quite agree.
However, certain things need to be made per-net namespace that *aren't*. DNS
results, for instance.
As an example, I could set up a client machine with two ethernet ports, set up
two DNS+NFS servers, each of which think they're called "foo.bar" and attach
each server to a different port on the client machine. Then I could create a
pair of containers on the client machine and route the network in each
container to a different port. Now there's a problem because the names of the
cached DNS records for each port overlap.
Further, the NFS idmapper needs to be able to direct its calls to the
appropriate network.
> I had some ideas about how you could implement bridging in userspace (as an
> unprivileged user, for rootless containers) but if you can't join namespaces
> individually then such a setup is not practically possible.
I'm not proposing to take away the ability to arbitrarily set the namespaces
in a container. I haven't implemented it yet, but it was on the to-do list:
(7) Directly set a container's namespaces to allow cross-container
sharing.
David
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2017-05-23 17:10 +0200 |
| Message-ID | <tKjBD-4S4-3@gated-at.bofh.it> |
| In reply to | #1648161 |
David Howells <dhowells@redhat.com> writes: > Aleksa Sarai <asarai@suse.de> wrote: > >> >> The reason I think this is necessary is that the kernel has no idea >> >> how to direct upcalls to what userspace considers to be a container - >> >> current Linux practice appears to make a "container" just an >> >> arbitrarily chosen junction of namespaces, control groups and files, >> >> which may be changed individually within the "container". >> >> Just want to point out that if the kernel APIs for containers massively >> change, then the OCI will have to completely rework how we describe containers >> (and so will all existing runtimes). >> >> Not to mention that while I don't like how hard it is (from a runtime >> perspective) to actually set up a container securely, there are undoubtedly >> benefits to having namespaces split out. The network namespace being separate >> means that in certain contexts you actually don't want to create a new network >> namespace when creating a container. > > Yep, I quite agree. > > However, certain things need to be made per-net namespace that *aren't*. DNS > results, for instance. > > As an example, I could set up a client machine with two ethernet ports, set up > two DNS+NFS servers, each of which think they're called "foo.bar" and attach > each server to a different port on the client machine. Then I could create a > pair of containers on the client machine and route the network in each > container to a different port. Now there's a problem because the names of the > cached DNS records for each port overlap. Please look at ip netns add. It does solve this in userspace rather simply. > Further, the NFS idmapper needs to be able to direct its calls to the > appropriate network. Eric
[toc] | [prev] | [next] | [standalone]
| From | David Howells <dhowells@redhat.com> |
|---|---|
| Date | 2017-05-23 17:20 +0200 |
| Message-ID | <tKjLl-4VL-33@gated-at.bofh.it> |
| In reply to | #1648164 |
Eric W. Biederman <ebiederm@xmission.com> wrote: > > As an example, I could set up a client machine with two ethernet ports, > > set up two DNS+NFS servers, each of which think they're called "foo.bar" > > and attach each server to a different port on the client machine. Then I > > could create a pair of containers on the client machine and route the > > network in each container to a different port. Now there's a problem > > because the names of the cached DNS records for each port overlap. > > Please look at ip netns add. warthog>man ip | grep setns warthog1> > It does solve this in userspace rather simply. Ummm... How? The kernel DNS resolver is not namespace aware. David
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2017-05-23 17:30 +0200 |
| Message-ID | <tKjV0-51B-25@gated-at.bofh.it> |
| In reply to | #1648182 |
David Howells <dhowells@redhat.com> writes: > Eric W. Biederman <ebiederm@xmission.com> wrote: > >> > As an example, I could set up a client machine with two ethernet ports, >> > set up two DNS+NFS servers, each of which think they're called "foo.bar" >> > and attach each server to a different port on the client machine. Then I >> > could create a pair of containers on the client machine and route the >> > network in each container to a different port. Now there's a problem >> > because the names of the cached DNS records for each port overlap. >> >> Please look at ip netns add. > > warthog>man ip | grep setns > warthog1> Not setns netns >> It does solve this in userspace rather simply. > > Ummm... How? The kernel DNS resolver is not namespace aware. But it works fine if called in the proper context and we have a defacto standard for where to put all of the files (the tricky part) if you are dealing with multiple network namespaces simultaneously. Eric
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2017-05-23 17:50 +0200 |
| Message-ID | <tKkel-58v-11@gated-at.bofh.it> |
| In reply to | #1648194 |
On Tue, 2017-05-23 at 10:17 -0500, Eric W. Biederman wrote: > David Howells <dhowells@redhat.com> writes: > > Eric W. Biederman <ebiederm@xmission.com> wrote: > > > It does solve this in userspace rather simply. > > > > Ummm... How? The kernel DNS resolver is not namespace aware. > > But it works fine if called in the proper context and we have a > defacto standard for where to put all of the files (the tricky part) > if you are dealing with multiple network namespaces simultaneously. I think you're missing each other's points slightly. What David is pointing out is that the kernel has a DNS cache (net/dns_resolver/) it can do name to IP translations, but isn't namespaced. Once it has one entry all containers would see it if they cause a lookup to go through the kernel cache, so going through the cache you can't have a name resolving to different IP addresses on a per container basis. I think Eric's point is that if you need the same DNS names resolving to different IP addresses on a per container basis, you can do this in userspace today but you have to disable the in-kernel DNS cache. James
[toc] | [prev] | [next] | [standalone]
| From | David Howells <dhowells@redhat.com> |
|---|---|
| Date | 2017-05-23 18:40 +0200 |
| Message-ID | <tKl0L-5Ft-25@gated-at.bofh.it> |
| In reply to | #1648214 |
James Bottomley <James.Bottomley@HansenPartnership.com> wrote: > What David is pointing out is that the kernel has a DNS cache > (net/dns_resolver/) it can do name to IP translations, but isn't > namespaced. Once it has one entry all containers would see it if they > cause a lookup to go through the kernel cache, so going through the > cache you can't have a name resolving to different IP addresses on a > per container basis. Yes - and the transport to userspace, the request_key() upcall, isn't namespaced either. Namespacing it isn't entirely simple since we have to set the right mount namespace (for execve, config, etc.), plus any other relevant namespaces (such as network) - which is dependent on key type. I can't record the mount namespace in the network namespace because that would create a dependency loop: mnt_ns -> mnt -> sb -> net_ns -> mnt_ns > I think Eric's point is that if you need the same DNS names resolving > to different IP addresses on a per container basis, you can do this in > userspace today but you have to disable the in-kernel DNS cache. You could disable the in-kernel dns resolver in your config, but then you don't get referrals in NFS. Also, CIFS, AFS and other filesystems would be affected. If you're fine with the restrictions, then there is no problem. David
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2017-05-24 10:40 +0200 |
| Message-ID | <tKzZL-8mT-1@gated-at.bofh.it> |
| In reply to | #1648257 |
David Howells <dhowells@redhat.com> writes: > James Bottomley <James.Bottomley@HansenPartnership.com> wrote: > >> What David is pointing out is that the kernel has a DNS cache >> (net/dns_resolver/) it can do name to IP translations, but isn't >> namespaced. Once it has one entry all containers would see it if they >> cause a lookup to go through the kernel cache, so going through the >> cache you can't have a name resolving to different IP addresses on a >> per container basis. > > Yes - and the transport to userspace, the request_key() upcall, isn't > namespaced either. Namespacing it isn't entirely simple since we have to set > the right mount namespace (for execve, config, etc.), plus any other relevant > namespaces (such as network) - which is dependent on key type. > > I can't record the mount namespace in the network namespace because that would > create a dependency loop: > > mnt_ns -> mnt -> sb -> net_ns -> mnt_ns I have already given a concrete suggest on how this might be untangled. So I won't repeat it here. >> I think Eric's point is that if you need the same DNS names resolving >> to different IP addresses on a per container basis, you can do this in >> userspace today but you have to disable the in-kernel DNS cache. > > You could disable the in-kernel dns resolver in your config, but then you > don't get referrals in NFS. Also, CIFS, AFS and other filesystems would be > affected. If you're fine with the restrictions, then there is no > problem. I haven't been arguing that at all. I was only pointing out that this issue is not an issue with DNS. Userspace handles this all fine. The issue is exclusively with this request_key api and generally user mode upcalls. I have no problem seeing that there is an issue with the kernel code. I am well aware of the problem. Unfortunately the people who cared enough to start addressing this have not been able to write kernel code that fixes this. My personal experience when I tried to use the request_key api at the beginning of this was it was too hard to test. There was no room for goofing up as at that time it was impossible to invalidate a cached reply from userspace if you happened to know it was wrong. Which meant that if something incorrect was cached it required rebooting the kernel. I have a lot of sympathy with the view that the best way to do some of this is with socket activations or perhaps something with rpc portmapper. Where something like inetd is used to start the user space component on-demand. I won't call that a solution to this case but I do think it makes a good example to compare with. When you need run something in a clean context having that something only need to worry about the contents of the data it is receiving and not about it's environment as suid applications do is a nice simplification. The entire user mode helper paradigm removes from user space the freedom to specify what context it's code should run in. In a world where everything is global that is fine. But in a world with containers where not everything is global it becomes a royal pain. And I am very very sympathetic to solving this. The only solution that I know would work is to capture the context at some point in a process and then to use that process to fork user mode helpers. So far no one has even bothered to seriously try the one solution that is guaranteed to work because it takes a lot of changes to kernel code. I believe the last effort snagged on what a pain it is to refactor the user mode helper infrastructure. I don't see in your code any of that work. I am glad to see that you also see the problem. At least when it comes to the request_key api. What I am hoping to see is someone who has the will to dig in and understand all of the interactions and refactor the kernel to solve the problem. This is not a case where our user space interfaces are preventing a solution to this problem (as your patchset implies). This is a case where things need to be refactored kernel side to solve this. So far this attempt is just another in the bazillion or so bad half-assed attempts to solve this problem I have seen over the years. Eric
[toc] | [prev] | [next] | [standalone]
| From | Ian Kent <raven@themaw.net> |
|---|---|
| Date | 2017-05-24 11:20 +0200 |
| Message-ID | <tKACt-pV-5@gated-at.bofh.it> |
| In reply to | #1649269 |
On Wed, 2017-05-24 at 03:26 -0500, Eric W. Biederman wrote: > > So far no one has even bothered to seriously try the one solution that > is guaranteed to work because it takes a lot of changes to kernel code. > I believe the last effort snagged on what a pain it is to refactor the > user mode helper infrastructure. Yes, that's mostly true in my case although I wouldn't say I haven't looked at it seriously but equally I haven't got anything towards it yet either, sorry. I'm likely going to revisit this based on a couple of approaches. One is just what you describe and I had already been looking at this some time ago. It seems to me that adding a work queue type that starts and retains a process until the work queue is destroyed (similar to the way the work queue sub system starts a fail over thread for use under resource exhaustion) would be a sensible way to do it. This doesn't mean I think it's a good idea for reasons I've outlined in the past but the approach does warrant the effort to work out if it can be used without problems. And there's also the request key infrastructure which, as it is now, gets in the road of verifying results, *sigh*. Ian
[toc] | [prev] | [next] | [standalone]
| From | Jessica Frazelle <me@jessfraz.com> |
|---|---|
| Date | 2017-05-22 19:30 +0200 |
| Message-ID | <tJZjz-7i-1@gated-at.bofh.it> |
| In reply to | #1647163 |
I had replied but not to the thread with the containers mailing list. See https://marc.info/?l=linux-cgroups&m=149547317006676&w=2 On Mon, May 22, 2017 at 5:53 PM, James Bottomley <James.Bottomley@hansenpartnership.com> wrote: > [Added missing cc to containers list] > On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote: >> Here are a set of patches to define a container object for the kernel >> and to provide some methods to create and manipulate them. >> >> The reason I think this is necessary is that the kernel has no idea >> how to direct upcalls to what userspace considers to be a container - >> current Linux practice appears to make a "container" just an >> arbitrarily chosen junction of namespaces, control groups and files, >> which may be changed individually within the "container". > > This sounds like a step in the wrong direction: the strength of the > current container interfaces in Linux is that people who set up > containers don't have to agree what they look like. So I can set up a > user namespace without a mount namespace or an architecture emulation > container with only a mount namespace. > > But ignoring my fun foibles with containers and to give a concrete > example in terms of a popular orchestration system: in kubernetes, > where certain namespaces are shared across pods, do you imagine the > kernel's view of the "container" to be the pod or what kubernetes > thinks of as the container? This is important, because half the > examples you give below are network related and usually pods share a > network namespace. I am glad you pointed this out because I was trying to make the same point, various definitions of containers differ and who is to say whether the various container runtimes (runc, rkt, systemd-nspawn) or consumers of containers (kubernetes) won't modify their definition in the future. How will this scale as new LSMs like Landlock or new namespaces are added in the future will they be included in the container kernel object as well... Seems like a lot more maintenance for something that is really just making the keyring namespace-aware... unless there are other things I missed. > >> The kernel upcall mechanism then needs to decide which set of >> namespaces, etc., it must exec the appropriate upcall program. >> Examples of this include: >> >> (1) The DNS resolver. The DNS cache in the kernel should probably >> be per-network namespace, but in userspace the program, its >> libraries and its config data are associated with a mount tree and a >> user namespace and it gets run in a particular pid namespace. > > All persistent (written to fs data) has to be mount ns associated; > there are no ifs, ands and buts to that. I agree this implies that if > you want to run a separate network namespace, you either take DNS from > the parent (a lot of containers do) or you set up a daemon to run > within the mount namespace. I agree the latter is a slightly fiddly > operation you have to get right, but that's why we have orchestration > systems. > > What is it we could do with the above that we cannot do today? > >> (2) NFS ID mapper. The NFS ID mapping cache should also probably be >> per-network namespace. > > I think this is a view but not the only one: Right at the moment, NFS > ID mapping is used as the one of the ways we can get the user namespace > ID mapping writes to file problems fixed ... that makes it a property > of the mount namespace for a lot of containers. There are many other > instances where they do exactly as you say, but what I'm saying is that > we don't want to lose the flexibility we currently have. > >> (3) nfsdcltrack. A way for NFSD to access stable storage for >> tracking of persistent state. Again, network-namespace dependent, >> but also perhaps mount-namespace dependent. > > So again, given we can set this up to work today, this sounds like more > a restriction that will bite us than an enhancement that gives us extra > features. > >> (4) General request-key upcalls. Not particularly namespace >> dependent, apart from keyrings being somewhat governed by the user >> namespace and the upcall being configured by the mount namespace. > > All mount namespaces have an owning user namespace, so the data > relations are already there in the kernel, is the problem simply > finding them? > >> These patches are built on top of the mount context patchset so that >> namespaces can be properly propagated over submounts/automounts. > > I'll stop here ... you get the idea that I think this is imposing a set > of restrictions that will come back to bite us later. If this is just > for the sake of figuring out how to get keyring upcalls to work, then > I'm sure we can come up with something. > > James > > -- > To unsubscribe from this list: send the line "unsubscribe cgroups" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html -- Jessie Frazelle 4096R / D4C4 DD60 0D66 F65A 8EFC 511E 18F3 685C 0022 BFF3 pgp.mit.edu
[toc] | [prev] | [next] | [standalone]
| From | Jeff Layton <jlayton@redhat.com> |
|---|---|
| Date | 2017-05-22 20:40 +0200 |
| Message-ID | <tK0pj-KA-9@gated-at.bofh.it> |
| In reply to | #1647163 |
On Mon, 2017-05-22 at 09:53 -0700, James Bottomley wrote: > [Added missing cc to containers list] > On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote: > > Here are a set of patches to define a container object for the kernel > > and to provide some methods to create and manipulate them. > > > > The reason I think this is necessary is that the kernel has no idea > > how to direct upcalls to what userspace considers to be a container - > > current Linux practice appears to make a "container" just an > > arbitrarily chosen junction of namespaces, control groups and files, > > which may be changed individually within the "container". > > This sounds like a step in the wrong direction: the strength of the > current container interfaces in Linux is that people who set up > containers don't have to agree what they look like. So I can set up a > user namespace without a mount namespace or an architecture emulation > container with only a mount namespace. > Does this really mandate what they look like though? AFAICT, you can still spawn disconnected namespaces to your heart's content. What this does is provide a container for several different namespaces so that the kernel can actually be aware of the association between them. The way you populate the different namespaces looks to be pretty flexible. > But ignoring my fun foibles with containers and to give a concrete > example in terms of a popular orchestration system: in kubernetes, > where certain namespaces are shared across pods, do you imagine the > kernel's view of the "container" to be the pod or what kubernetes > thinks of as the container? This is important, because half the > examples you give below are network related and usually pods share a > network namespace. > > > The kernel upcall mechanism then needs to decide which set of > > namespaces, etc., it must exec the appropriate upcall program. > > Examples of this include: > > > > (1) The DNS resolver. The DNS cache in the kernel should probably > > be per-network namespace, but in userspace the program, its > > libraries and its config data are associated with a mount tree and a > > user namespace and it gets run in a particular pid namespace. > > All persistent (written to fs data) has to be mount ns associated; > there are no ifs, ands and buts to that. I agree this implies that if > you want to run a separate network namespace, you either take DNS from > the parent (a lot of containers do) or you set up a daemon to run > within the mount namespace. I agree the latter is a slightly fiddly > operation you have to get right, but that's why we have orchestration > systems. > > What is it we could do with the above that we cannot do today? > Spawn a task directly from the kernel, already set up in the correct namespaces, a'la call_usermodehelper. So far there is no way to do that, and it is something we'd very much desire. Ian Kent has made several passes at it recently. > > (2) NFS ID mapper. The NFS ID mapping cache should also probably be > > per-network namespace. > > I think this is a view but not the only one: Right at the moment, NFS > ID mapping is used as the one of the ways we can get the user namespace > ID mapping writes to file problems fixed ... that makes it a property > of the mount namespace for a lot of containers. There are many other > instances where they do exactly as you say, but what I'm saying is that > we don't want to lose the flexibility we currently have. > > > (3) nfsdcltrack. A way for NFSD to access stable storage for > > tracking of persistent state. Again, network-namespace dependent, > > but also perhaps mount-namespace dependent. Definitely mount-namespace dependent. > > So again, given we can set this up to work today, this sounds like more > a restriction that will bite us than an enhancement that gives us extra > features. > How do you set this up to work today? AFAIK, if you want to run knfsd in a container today, you're out of luck for any non-trivial configuration. The main reason is that most of knfsd is namespace-ized in the network namespace, but there is no clear way to associate that with a mount namespace, which is what we need to do this properly inside a container. I think David's patches would get us there. > > (4) General request-key upcalls. Not particularly namespace > > dependent, apart from keyrings being somewhat governed by the user > > namespace and the upcall being configured by the mount namespace. > > All mount namespaces have an owning user namespace, so the data > relations are already there in the kernel, is the problem simply > finding them? > > > These patches are built on top of the mount context patchset so that > > namespaces can be properly propagated over submounts/automounts. > > I'll stop here ... you get the idea that I think this is imposing a set > of restrictions that will come back to bite us later. If this is just > for the sake of figuring out how to get keyring upcalls to work, then > I'm sure we can come up with something. > -- Jeff Layton <jlayton@redhat.com>
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2017-05-22 21:30 +0200 |
| Message-ID | <tK1bI-1hH-17@gated-at.bofh.it> |
| In reply to | #1647255 |
On Mon, 2017-05-22 at 14:34 -0400, Jeff Layton wrote: > On Mon, 2017-05-22 at 09:53 -0700, James Bottomley wrote: > > [Added missing cc to containers list] > > On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote: > > > Here are a set of patches to define a container object for the > > > kernel and to provide some methods to create and manipulate them. > > > > > > The reason I think this is necessary is that the kernel has no > > > idea how to direct upcalls to what userspace considers to be a > > > container - current Linux practice appears to make a "container" > > > just an arbitrarily chosen junction of namespaces, control groups > > > and files, which may be changed individually within the > > > "container". > > > > This sounds like a step in the wrong direction: the strength of the > > current container interfaces in Linux is that people who set up > > containers don't have to agree what they look like. So I can set > > up a user namespace without a mount namespace or an architecture > > emulation container with only a mount namespace. > > > > Does this really mandate what they look like though? AFAICT, you can > still spawn disconnected namespaces to your heart's content. What > this does is provide a container for several different namespaces so > that the kernel can actually be aware of the association between > them. Yes, because it imposes a view of what is in a container. As the several replies have pointed out (and indeed as I pointed out below for kubernetes), this isn't something the orchestration systems would find usable. > The way you populate the different namespaces looks to be pretty > flexible. OK, but look at it another way: If we provides a container API no actual consumer of container technologies wants to use just because we think it makes certain tasks easy, is it really a good API? Containers are multi-layered and complex. If you're not ready for this as a user, then you should use an orchestration system that prevents you from screwing up. > > But ignoring my fun foibles with containers and to give a concrete > > example in terms of a popular orchestration system: in kubernetes, > > where certain namespaces are shared across pods, do you imagine the > > kernel's view of the "container" to be the pod or what kubernetes > > thinks of as the container? This is important, because half the > > examples you give below are network related and usually pods share > > a network namespace. > > > > > The kernel upcall mechanism then needs to decide which set of > > > namespaces, etc., it must exec the appropriate upcall program. > > > Examples of this include: > > > > > > (1) The DNS resolver. The DNS cache in the kernel should > > > probably be per-network namespace, but in userspace the program, > > > its libraries and its config data are associated with a mount > > > tree and a user namespace and it gets run in a particular pid > > > namespace. > > > > All persistent (written to fs data) has to be mount ns associated; > > there are no ifs, ands and buts to that. I agree this implies that > > if you want to run a separate network namespace, you either take > > DNS from the parent (a lot of containers do) or you set up a daemon > > to run within the mount namespace. I agree the latter is a > > slightly fiddly operation you have to get right, but that's why we > > have orchestration systems. > > > > What is it we could do with the above that we cannot do today? > > > > Spawn a task directly from the kernel, already set up in the correct > namespaces, a'la call_usermodehelper. So far there is no way to do > that, Today the usermode helper has to be namespace aware. We spawn it into the root namespace and it jumps into the correct namespace/cgroup combination and re-executes itself or simply performs the requisite task on behalf of the container. Is this simple, no; does it work, yes, provided the host OS is aware of what the container orchestration system wants it to do. > and it is something we'd very much desire. Ian Kent has made several > passes at it recently. Well, every time we try to remove some of the complexity from userspace, we end up wrapping around the axle of what exactly we're trying to achieve, yes. > > > (2) NFS ID mapper. The NFS ID mapping cache should also > > > probably be per-network namespace. > > > > I think this is a view but not the only one: Right at the moment, > > NFS ID mapping is used as the one of the ways we can get the user > > namespace ID mapping writes to file problems fixed ... that makes > > it a property of the mount namespace for a lot of containers. > > There are many other instances where they do exactly as you say, > > but what I'm saying is that we don't want to lose the flexibility > > we currently have. > > > > > (3) nfsdcltrack. A way for NFSD to access stable storage for > > > tracking of persistent state. Again, network-namespace > > > dependent, but also perhaps mount-namespace dependent. > > Definitely mount-namespace dependent. > > > > > So again, given we can set this up to work today, this sounds like > > more a restriction that will bite us than an enhancement that gives > > us extra features. > > > > How do you set this up to work today? Well, as above, it spawns into the root, you jump it to where it should be and re-execute or simply handle in the host. > AFAIK, if you want to run knfsd in a container today, you're out of > luck for any non-trivial configuration. Well "running knfsd in a container" is actually different from having a containerised nfs export. My understanding was that thanks to the work of Stas Kinsbursky, the latter has mostly worked since the 3.9 kernel for v3 and below. I assume the current issue is that there's a problem with v4? James > The main reason is that most of knfsd is namespace-ized in the > network namespace, but there is no clear way to associate that with a > mount namespace, which is what we need to do this properly inside a > container. I think David's patches would get us there. > > > > (4) General request-key upcalls. Not particularly namespace > > > dependent, apart from keyrings being somewhat governed by the > > > user namespace and the upcall being configured by the mount > > > namespace. > > > > All mount namespaces have an owning user namespace, so the data > > relations are already there in the kernel, is the problem simply > > finding them? > > > > > These patches are built on top of the mount context patchset so > > > that namespaces can be properly propagated over > > > submounts/automounts. > > > > I'll stop here ... you get the idea that I think this is imposing a > > set of restrictions that will come back to bite us later. If this > > is just for the sake of figuring out how to get keyring upcalls to > > work, then I'm sure we can come up with something. > > >
[toc] | [prev] | [next] | [standalone]
| From | Jeff Layton <jlayton@redhat.com> |
|---|---|
| Date | 2017-05-23 00:20 +0200 |
| Message-ID | <tK3Qd-30F-1@gated-at.bofh.it> |
| In reply to | #1647295 |
On Mon, 2017-05-22 at 12:21 -0700, James Bottomley wrote: > On Mon, 2017-05-22 at 14:34 -0400, Jeff Layton wrote: > > On Mon, 2017-05-22 at 09:53 -0700, James Bottomley wrote: > > > [Added missing cc to containers list] > > > On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote: > > > > Here are a set of patches to define a container object for the > > > > kernel and to provide some methods to create and manipulate them. > > > > > > > > The reason I think this is necessary is that the kernel has no > > > > idea how to direct upcalls to what userspace considers to be a > > > > container - current Linux practice appears to make a "container" > > > > just an arbitrarily chosen junction of namespaces, control groups > > > > and files, which may be changed individually within the > > > > "container". > > > > > > This sounds like a step in the wrong direction: the strength of the > > > current container interfaces in Linux is that people who set up > > > containers don't have to agree what they look like. So I can set > > > up a user namespace without a mount namespace or an architecture > > > emulation container with only a mount namespace. > > > > > > > Does this really mandate what they look like though? AFAICT, you can > > still spawn disconnected namespaces to your heart's content. What > > this does is provide a container for several different namespaces so > > that the kernel can actually be aware of the association between > > them. > > Yes, because it imposes a view of what is in a container. As the > several replies have pointed out (and indeed as I pointed out below for > kubernetes), this isn't something the orchestration systems would find > usable. > > > The way you populate the different namespaces looks to be pretty > > flexible. > > OK, but look at it another way: If we provides a container API no > actual consumer of container technologies wants to use just because we > think it makes certain tasks easy, is it really a good API? > > Containers are multi-layered and complex. If you're not ready for this > as a user, then you should use an orchestration system that prevents > you from screwing up. > > > > But ignoring my fun foibles with containers and to give a concrete > > > example in terms of a popular orchestration system: in kubernetes, > > > where certain namespaces are shared across pods, do you imagine the > > > kernel's view of the "container" to be the pod or what kubernetes > > > thinks of as the container? This is important, because half the > > > examples you give below are network related and usually pods share > > > a network namespace. > > > > > > > The kernel upcall mechanism then needs to decide which set of > > > > namespaces, etc., it must exec the appropriate upcall program. > > > > Examples of this include: > > > > > > > > (1) The DNS resolver. The DNS cache in the kernel should > > > > probably be per-network namespace, but in userspace the program, > > > > its libraries and its config data are associated with a mount > > > > tree and a user namespace and it gets run in a particular pid > > > > namespace. > > > > > > All persistent (written to fs data) has to be mount ns associated; > > > there are no ifs, ands and buts to that. I agree this implies that > > > if you want to run a separate network namespace, you either take > > > DNS from the parent (a lot of containers do) or you set up a daemon > > > to run within the mount namespace. I agree the latter is a > > > slightly fiddly operation you have to get right, but that's why we > > > have orchestration systems. > > > > > > What is it we could do with the above that we cannot do today? > > > > > > > Spawn a task directly from the kernel, already set up in the correct > > namespaces, a'la call_usermodehelper. So far there is no way to do > > that, > > Today the usermode helper has to be namespace aware. We spawn it into > the root namespace and it jumps into the correct namespace/cgroup > combination and re-executes itself or simply performs the requisite > task on behalf of the container. Is this simple, no; does it work, > yes, provided the host OS is aware of what the container orchestration > system wants it to do. > > > and it is something we'd very much desire. Ian Kent has made several > > passes at it recently. > > Well, every time we try to remove some of the complexity from > userspace, we end up wrapping around the axle of what exactly we're > trying to achieve, yes. > > > > > (2) NFS ID mapper. The NFS ID mapping cache should also > > > > probably be per-network namespace. > > > > > > I think this is a view but not the only one: Right at the moment, > > > NFS ID mapping is used as the one of the ways we can get the user > > > namespace ID mapping writes to file problems fixed ... that makes > > > it a property of the mount namespace for a lot of containers. > > > There are many other instances where they do exactly as you say, > > > but what I'm saying is that we don't want to lose the flexibility > > > we currently have. > > > > > > > (3) nfsdcltrack. A way for NFSD to access stable storage for > > > > tracking of persistent state. Again, network-namespace > > > > dependent, but also perhaps mount-namespace dependent. > > > > Definitely mount-namespace dependent. > > > > > > > > So again, given we can set this up to work today, this sounds like > > > more a restriction that will bite us than an enhancement that gives > > > us extra features. > > > > > > > How do you set this up to work today? > > Well, as above, it spawns into the root, you jump it to where it should > be and re-execute or simply handle in the host. > > > AFAIK, if you want to run knfsd in a container today, you're out of > > luck for any non-trivial configuration. > > Well "running knfsd in a container" is actually different from having a > containerised nfs export. My understanding was that thanks to the work > of Stas Kinsbursky, the latter has mostly worked since the 3.9 kernel > for v3 and below. I assume the current issue is that there's a problem > with v4? > Yes -- v3 mostly works because the equivalent state-tracking (rpc.statd) is run as a long-running daemon. nfsdcltrack uses call_usermodehelper, so for that you need to be able to determine what mount namespace to run the thing in. All we really know in knfsd when we want to do an upcall is the net namespace. We could really use a way to associate the two and spawn the thing in the correct container (or pass it enough info for it to setns() into the right ones). In principle, we could just ensure that we do all of this sort of thing with long-running daemons that are started whenever the container starts. But...having to run daemons full-time for infrequently-used services sort of sucks and requires it to be setup. UMH helpers just get run as long as the binary is in the right place. I've also been reading over Eric suggestion, and that seems like it might work as well though. > > The main reason is that most of knfsd is namespace-ized in the > > network namespace, but there is no clear way to associate that with a > > mount namespace, which is what we need to do this properly inside a > > container. I think David's patches would get us there. > > > > > > (4) General request-key upcalls. Not particularly namespace > > > > dependent, apart from keyrings being somewhat governed by the > > > > user namespace and the upcall being configured by the mount > > > > namespace. > > > > > > All mount namespaces have an owning user namespace, so the data > > > relations are already there in the kernel, is the problem simply > > > finding them? > > > > > > > These patches are built on top of the mount context patchset so > > > > that namespaces can be properly propagated over > > > > submounts/automounts. > > > > > > I'll stop here ... you get the idea that I think this is imposing a > > > set of restrictions that will come back to bite us later. If this > > > is just for the sake of figuring out how to get keyring upcalls to > > > work, then I'm sure we can come up with something. > > > > > -- Jeff Layton <jlayton@redhat.com>
[toc] | [prev] | [next] | [standalone]
| From | Ian Kent <raven@themaw.net> |
|---|---|
| Date | 2017-05-23 12:40 +0200 |
| Message-ID | <tKfol-1W6-5@gated-at.bofh.it> |
| In reply to | #1647295 |
On Mon, 2017-05-22 at 12:21 -0700, James Bottomley wrote: > > > > > (3) nfsdcltrack. A way for NFSD to access stable storage for > > > > tracking of persistent state. Again, network-namespace > > > > dependent, but also perhaps mount-namespace dependent. > > > > Definitely mount-namespace dependent. > > > > > > > > So again, given we can set this up to work today, this sounds like > > > more a restriction that will bite us than an enhancement that gives > > > us extra features. > > > > > > > How do you set this up to work today? > > Well, as above, it spawns into the root, you jump it to where it should > be and re-execute or simply handle in the host. > > > AFAIK, if you want to run knfsd in a container today, you're out of > > luck for any non-trivial configuration. > > Well "running knfsd in a container" is actually different from having a > containerised nfs export. My understanding was that thanks to the work > of Stas Kinsbursky, the latter has mostly worked since the 3.9 kernel > for v3 and below. I assume the current issue is that there's a problem > with v4? Oh, ok, I thought that, say, a docker (NFS) volumes-from a container to another container didn't work for any version of NFS. Certainly didn't work last time I tried, it was a while ago though. Ian
[toc] | [prev] | [next] | [standalone]
| From | Ian Kent <raven@themaw.net> |
|---|---|
| Date | 2017-05-23 11:40 +0200 |
| Message-ID | <tKesi-1hY-29@gated-at.bofh.it> |
| In reply to | #1647163 |
On Mon, 2017-05-22 at 09:53 -0700, James Bottomley wrote: > [Added missing cc to containers list] > On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote: > > Here are a set of patches to define a container object for the kernel > > and to provide some methods to create and manipulate them. > > > > The reason I think this is necessary is that the kernel has no idea > > how to direct upcalls to what userspace considers to be a container - > > current Linux practice appears to make a "container" just an > > arbitrarily chosen junction of namespaces, control groups and files, > > which may be changed individually within the "container". > > This sounds like a step in the wrong direction: the strength of the > current container interfaces in Linux is that people who set up > containers don't have to agree what they look like. So I can set up a > user namespace without a mount namespace or an architecture emulation > container with only a mount namespace. > > But ignoring my fun foibles with containers and to give a concrete > example in terms of a popular orchestration system: in kubernetes, > where certain namespaces are shared across pods, do you imagine the > kernel's view of the "container" to be the pod or what kubernetes > thinks of as the container? This is important, because half the > examples you give below are network related and usually pods share a > network namespace. > > > The kernel upcall mechanism then needs to decide which set of > > namespaces, etc., it must exec the appropriate upcall program. > > Examples of this include: > > > > (1) The DNS resolver. The DNS cache in the kernel should probably > > be per-network namespace, but in userspace the program, its > > libraries and its config data are associated with a mount tree and a > > user namespace and it gets run in a particular pid namespace. > > All persistent (written to fs data) has to be mount ns associated; > there are no ifs, ands and buts to that. I agree this implies that if > you want to run a separate network namespace, you either take DNS from > the parent (a lot of containers do) or you set up a daemon to run > within the mount namespace. I agree the latter is a slightly fiddly > operation you have to get right, but that's why we have orchestration > systems. > > What is it we could do with the above that we cannot do today? > > > (2) NFS ID mapper. The NFS ID mapping cache should also probably be > > per-network namespace. > > I think this is a view but not the only one: Right at the moment, NFS > ID mapping is used as the one of the ways we can get the user namespace > ID mapping writes to file problems fixed ... that makes it a property > of the mount namespace for a lot of containers. There are many other > instances where they do exactly as you say, but what I'm saying is that > we don't want to lose the flexibility we currently have. > > > (3) nfsdcltrack. A way for NFSD to access stable storage for > > tracking of persistent state. Again, network-namespace dependent, > > but also perhaps mount-namespace dependent. > > So again, given we can set this up to work today, this sounds like more > a restriction that will bite us than an enhancement that gives us extra > features. > > > (4) General request-key upcalls. Not particularly namespace > > dependent, apart from keyrings being somewhat governed by the user > > namespace and the upcall being configured by the mount namespace. > > All mount namespaces have an owning user namespace, so the data > relations are already there in the kernel, is the problem simply > finding them? > > > These patches are built on top of the mount context patchset so that > > namespaces can be properly propagated over submounts/automounts. > > I'll stop here ... you get the idea that I think this is imposing a set > of restrictions that will come back to bite us later. If this is just > for the sake of figuring out how to get keyring upcalls to work, then > I'm sure we can come up with something. You talk about a number of things I'm simply not aware of so I can't answer your questions. But your points do sound like issues that need to be covered. I think you mentioned user space used NFS ID mapper works fine. I wonder, could you give more detail on that please. Perhaps nsenter(1) is being used, I tried that as a possible usermode helper solution and it probably did "work" in the sense of in container execution but no-one liked it, it seems kernel folk expect to do things, well, in kernel. Not only that there were other problems, probably request key sub system not being namespace aware, or id caching within nfs or somewhere else, and there was a question of not being able to cater for user namespace usage. Anyway I do have a different view from my own experiences. First there are a number of subsystems involved in creating a process from within a container that has the container environment and, AFAICS (from the usermode helper experience), it needs to be done from outside the container. For example sub systems that need to be handled properly are the namespaces (and the pid namespace in particular is tricky), credentials and cgroups, to name those that come immediately to mind. I just couldn't get all that right after a number of tries. From this the problem that occurred to me is that we have a comprehensive namespace implementation within the kernel but no container implementation to help the binding together of the various sub systems for container use cases in a way that satisfies container (or a process within an existing container) creation. The risk is that, as time passes, problems like usermode helper will be solved in different places in different ways, not necessarily satisfactorily and potentially hard to find when there are bugs and even harder to maintain than the implementation here. At least the interface here provides a "goto" place to define and maintain the procedures required to do these things. Yes, it would require change but change happens and the first pass may not be on a par with what is currently done from a simplicity POV. But why not see it as solving a kernel development problem and focus on what needs to be done to make it on par with (and perhaps simpler) to cover the current usage .... Eric Biederman's comment is attractive indeed but I don't see how that solves problem of having a containers sub system implementation which I think is needed, if for no other reason than as a definition of correct usage and localization of that usage for easier (right, nothing is easy) bug resolution. Ian
[toc] | [prev] | [next] | [standalone]
| From | David Howells <dhowells@redhat.com> |
|---|---|
| Date | 2017-05-23 16:00 +0200 |
| Message-ID | <tKivU-3WY-19@gated-at.bofh.it> |
| In reply to | #1647163 |
James Bottomley <James.Bottomley@HansenPartnership.com> wrote: > This sounds like a step in the wrong direction: the strength of the > current container interfaces in Linux is that people who set up > containers don't have to agree what they look like. It may be a strength, but it is also a problem. > So I can set up a user namespace without a mount namespace or an > architecture emulation container with only a mount namespace. (I presume you mean with only the mount namespace separate) Yep. You can do that with this too. > But ignoring my fun foibles with containers and to give a concrete > example in terms of a popular orchestration system: in kubernetes, > where certain namespaces are shared across pods, do you imagine the > kernel's view of the "container" to be the pod or what kubernetes > thinks of as the container? Why not both? If the net_ns is created in the pod container, then probably network-related upcalls should be directed there. Unless instructed otherwise, upon creation a container object will inherit the caller's namespaces. > This is important, because half the examples you give below are network > related and usually pods share a network namespace. Yeah - I'm more familiar with upcalls made by NFS, AFS and keyrings. > > (1) The DNS resolver. ... > > All persistent (written to fs data) has to be mount ns associated; > there are no ifs, ands and buts to that. I agree this implies that if > you want to run a separate network namespace, you either take DNS from > the parent (a lot of containers do) My intention is to make the DNS cache per-network namespace within the kernel. Currently, there's only one and it's shared between all namespaces, but that won't work if you end up with two net namespaces that direct the same addresses to different things. > or you set up a daemon to run within the mount namespace. That's not currently an option: the DNS service upcalls only, and /sbin/request-key is invoked in the init_ns. This performs the network accesses in the wrong network namespace. > I agree the latter is a slightly fiddly operation you have to get right, but > that's why we have orchestration systems. An orchestration system can use this. This is not a replacement for Kubernetes or Docker or whatever. > What is it we could do with the above that we cannot do today? Upcall into an appropriate set of namespaces and keep the results separate by network namespace. > > (2) NFS ID mapper. The NFS ID mapping cache should also probably be > > per-network namespace. > > I think this is a view but not the only one: Right at the moment, NFS > ID mapping is used as the one of the ways we can get the user namespace > ID mapping writes to file problems fixed ... that makes it a property > of the mount namespace for a lot of containers. In some ways it's really a property of the server, and two different servers may appear in two separate network namespaces with the same apparent name and address. It's not a property of the mount namespace because mount namespaces share superblocks, and this is done at the superblock level. Possibly it should be done on the vfsmount, as a filter on the interaction between userspace and kernel. > There are many other instances where they do exactly as you say, but what > I'm saying is that we don't want to lose the flexibility we currently have. You don't really lose any flexibility; if anything, you gain it. (Note that in case your objection is that I haven't yet implemented the ability to set namespaces arbitrarily in a namespace, that's on list of things to do that I included, as is adjusting the control groups.) > All mount namespaces have an owning user namespace, so the data > relations are already there in the kernel, is the problem simply > finding them? The superblocks used by the vfsmounts in a mount namespace aren't all necessarily in the same user_ns, so none of: sb->s_user_ns == current_user_ns() sb->s_user_ns == current->ns->mnt_ns->user_ns current->ns->mnt_ns->user_ns == current_user_ns() need hold true that I can see. > > These patches are built on top of the mount context patchset so that > > namespaces can be properly propagated over submounts/automounts. > > I'll stop here ... you get the idea that I think this is imposing a set > of restrictions that will come back to bite us later. What restrictions am I imposing? > If this is just for the sake of figuring out how to get keyring upcalls to > work, then I'm sure we can come up with something. No, it's not just for that, though, admittedly, all of the upcall mechanisms I outlined use request_key() at the core. Really, a container is an anchor for the resources you need to make an upcall, but it can also be used to anchor other things. One thing I've been asked for by a number of people is a per-container keyring for the provision of authentication keys, fs decryption keys and other things - but there's no actual container off which this can be hung. Another thing that could be useful is a list of what device files a container may access, so that we can allow limited mounting by the container root user within the container. Now these could be made into their own namespaces or added to one that already exists - perhaps the mount namespace being the most logical. David
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web