Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1647121 > unrolled thread

[RFC][PATCH 0/9] Make containers kernel objects

Started byDavid Howells <dhowells@redhat.com>
First post2017-05-22 18:30 +0200
Last post2017-05-23 17:40 +0200
Articles 20 on this page of 37 — 9 participants

Back to article view | Back to linux.kernel


Contents

  [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-22 18:30 +0200
    [PATCH 8/9] Honour CONTAINER_NEW_EMPTY_FS_NS David Howells <dhowells@redhat.com> - 2017-05-22 18:30 +0200
    [PATCH 3/9] Provide /proc/containers David Howells <dhowells@redhat.com> - 2017-05-22 18:30 +0200
    Re: [RFC][PATCH 0/9] Make containers kernel objects James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-05-22 19:00 +0200
      Re: [RFC][PATCH 0/9] Make containers kernel objects Aleksa Sarai <asarai@suse.de> - 2017-05-22 19:20 +0200
        Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 17:00 +0200
          Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-23 17:10 +0200
            Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 17:20 +0200
              Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-23 17:30 +0200
                Re: [RFC][PATCH 0/9] Make containers kernel objects James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-05-23 17:50 +0200
                  Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 18:40 +0200
                    Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-24 10:40 +0200
                      Re: [RFC][PATCH 0/9] Make containers kernel objects Ian Kent <raven@themaw.net> - 2017-05-24 11:20 +0200
      Re: [RFC][PATCH 0/9] Make containers kernel objects Jessica Frazelle <me@jessfraz.com> - 2017-05-22 19:30 +0200
      Re: [RFC][PATCH 0/9] Make containers kernel objects Jeff Layton <jlayton@redhat.com> - 2017-05-22 20:40 +0200
        Re: [RFC][PATCH 0/9] Make containers kernel objects James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-05-22 21:30 +0200
          Re: [RFC][PATCH 0/9] Make containers kernel objects Jeff Layton <jlayton@redhat.com> - 2017-05-23 00:20 +0200
          Re: [RFC][PATCH 0/9] Make containers kernel objects Ian Kent <raven@themaw.net> - 2017-05-23 12:40 +0200
      Re: [RFC][PATCH 0/9] Make containers kernel objects Ian Kent <raven@themaw.net> - 2017-05-23 11:40 +0200
      Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 16:00 +0200
        Re: [RFC][PATCH 0/9] Make containers kernel objects James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-05-23 17:10 +0200
        Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-23 17:40 +0200
    Re: [RFC][PATCH 0/9] Make containers kernel objects Jessica Frazelle <me@jessfraz.com> - 2017-05-22 19:20 +0200
      Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 17:20 +0200
    Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-22 21:20 +0200
      Re: [RFC][PATCH 0/9] Make containers kernel objects Jeff Layton <jlayton@redhat.com> - 2017-05-23 00:30 +0200
        Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-23 15:10 +0200
          Re: [RFC][PATCH 0/9] Make containers kernel objects Jeff Layton <jlayton@redhat.com> - 2017-05-23 16:30 +0200
          Re: [RFC][PATCH 0/9] Make containers kernel objects Djalal Harouni <tixxdz@gmail.com> - 2017-05-23 16:40 +0200
            Re: [RFC][PATCH 0/9] Make containers kernel objects Colin Walters <walters@verbum.org> - 2017-05-23 17:00 +0200
              Re: [RFC][PATCH 0/9] Make containers kernel objects Colin Walters <walters@verbum.org> - 2017-05-23 17:40 +0200
              Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 17:40 +0200
              Re: [RFC][PATCH 0/9] Make containers kernel objects Jeff Layton <jlayton@redhat.com> - 2017-05-23 17:40 +0200
        Re: [RFC][PATCH 0/9] Make containers kernel objects Djalal Harouni <tixxdz@gmail.com> - 2017-05-23 16:30 +0200
      Re: [RFC][PATCH 0/9] Make containers kernel objects David Howells <dhowells@redhat.com> - 2017-05-23 18:20 +0200
    Re: [RFC][PATCH 0/9] Make containers kernel objects Ian Kent <raven@themaw.net> - 2017-05-23 12:20 +0200
    Re: [RFC][PATCH 0/9] Make containers kernel objects ebiederm@xmission.com (Eric W. Biederman) - 2017-05-23 17:40 +0200

Page 1 of 2  [1] 2  Next page →


#1647121 — [RFC][PATCH 0/9] Make containers kernel objects

FromDavid Howells <dhowells@redhat.com>
Date2017-05-22 18:30 +0200
Subject[RFC][PATCH 0/9] Make containers kernel objects
Message-ID<tJYnv-80c-3@gated-at.bofh.it>
Here are a set of patches to define a container object for the kernel and
to provide some methods to create and manipulate them.

The reason I think this is necessary is that the kernel has no idea how to
direct upcalls to what userspace considers to be a container - current
Linux practice appears to make a "container" just an arbitrarily chosen
junction of namespaces, control groups and files, which may be changed
individually within the "container".

The kernel upcall mechanism then needs to decide which set of namespaces,
etc., it must exec the appropriate upcall program.  Examples of this
include:

 (1) The DNS resolver.  The DNS cache in the kernel should probably be
     per-network namespace, but in userspace the program, its libraries and
     its config data are associated with a mount tree and a user namespace
     and it gets run in a particular pid namespace.

 (2) NFS ID mapper.  The NFS ID mapping cache should also probably be
     per-network namespace.

 (3) nfsdcltrack.  A way for NFSD to access stable storage for tracking
     of persistent state.  Again, network-namespace dependent, but also
     perhaps mount-namespace dependent.

 (4) General request-key upcalls.  Not particularly namespace dependent,
     apart from keyrings being somewhat governed by the user namespace and
     the upcall being configured by the mount namespace.

These patches are built on top of the mount context patchset so that
namespaces can be properly propagated over submounts/automounts.

These patches implement a container object that holds the following things:

 (1) Namespaces.

 (2) A root directory.

 (3) A set of processes, including a designated 'init' process.

 (4) The creator's credentials, including ownership.

 (5) A place to hang security for the container, allowing policies to be
     set per-container.

I also want to add:

 (6) Control groups.

 (7) A per-container keyring that can be added to from outside of the
     container, even once the container is live, for the provision of
     filesystem authentication/encryption keys in advance of the container
     being started.

You can get a list of containers by examining /proc/containers - but I'm
not sure how much value this gets you.  Note that the container in which
you are running is called "<current>" and you can only see other containers
that were started from within yours.  Containers are therefore effectively
hierarchical and an init_container is set up when the system boots.


Some management operations are provided:

 (1) int fd = container_create(const char *name, unsigned int flags);

     Create a container of the given name and return a handle to it as a
     file descriptor.  flags indicates what namespaces should be inherited
     from the caller and what should be replaced new.  It is possible to
     set up a container with a null root filesystem that can be mounted
     later.

 (2) int fsfd = fsopen(const char *fsname, int container_fd,
		       unsigned int flags);

     Prepare a mount context inside the container.  This uses all the
     containers namespaces instead of the caller's.

 (3) fsmount(int fsfd, int dfd, const char *path, unsigned int at_flags,
	     unsigned int flags);

     Mount a prepared superblock.  dfd can be given container_fd to use the
     container to which it refers as the root of the pathwalk.

     If path is "/" and at_flags is AT_FSMOUNT_CONTAINER_ROOT, then this
     will attempt to mount the root of the container and create a mount
     namespace for it.  The container must've been created with
     CONTAINER_NEW_EMPTY_FS_NS.

 (4) pid_t pid = fork_into_container(int container_fd);

     Create the init process in a container.  The process uses that
     container's namespaces instead of the caller's.

 (5) int sfd = container_socket(int container_fd,
				int domain, int type, int protocol);

     Create a socket inside a container.  The socket gets the container's
     namespaces.  This allows netlink operations to be called within that
     container to set it up from outside (at least in theory).

 (6) mkdirat(int dfd, ...);
     mknodat(int dfd, ...);
     openat(int dfd, ...);

     Supplying a container fd as dfd makes the pathwalk happen relative to
     the root of the container.  Note that the path must be *relative*.

And some need to be/could be added:

 (7) Directly set a container's namespaces to allow cross-container
     sharing.

 (8) Adjust the control group membership of a container.

 (9) Add a key inside a container keyring.

(10) Kill/suspend/freeze/reboot container, both from inside and out.

(11) Set container's root dir.

(12) Set the container's security policy.

(13) Allow overlayfs to access filesystems outside of the container in
     which it is being created.


Kernel upcalls are invoked in the root of the container that incurs them
rather than in the init namespace context.  There's still some awkwardness
here if you, say, share a network namespace between containers.  Either the
upcall binaries and configuration must be duplicated between sharing
containers or a container must be elected as the one in which such upcalls
will be done.


Some further thoughts:

 (*) Should there be an AT_IN_CONTAINER flag to provide to syscalls that
     take a container in lieu of AT_FDCWD or a directory fd?  The problem
     is that such as mkdirat() and openat() don't have an at_flags
     argument.

 (*) Should there be a container hierarchy at all?  It seems that this is
     only really necessary for /proc/containers.  Do we want to allow
     containers-within-containers?

 (*) Should each container automatically have its own pid namespace such
     that its 'init' process always appears as pid 1?

 (*) Does this allow kernel upcalls to be accounted against the correct
     control group?

 (*) Should each container have a 'list' of accessible device numbers such
     that certain device files can be made usable within a container?  And
     can devtmpfs/udev be made to show the correct file set for each
     container?


The patches can be found here also:

	http://git.kernel.org/cgit/linux/kernel/git/dhowells/linux-fs.git/log/?h=container

Note that this is dependent on the mount-context branch.

David
---
David Howells (9):
      containers: Rename linux/container.h to linux/container_dev.h
      Implement containers as kernel objects
      Provide /proc/containers
      Allow processes to be forked and upcalled into a container
      Open a socket inside a container
      Allow fs syscall dfd arguments to take a container fd
      Make fsopen() able to initiate mounting into a container
      Honour CONTAINER_NEW_EMPTY_FS_NS
      Sample program for driving container objects


 arch/x86/entry/syscalls/syscall_32.tbl |    3 
 arch/x86/entry/syscalls/syscall_64.tbl |    3 
 drivers/acpi/container.c               |    2 
 drivers/base/container.c               |    2 
 fs/fsopen.c                            |   33 +-
 fs/libfs.c                             |    3 
 fs/namei.c                             |   52 ++-
 fs/namespace.c                         |  108 +++++-
 fs/nfs/namespace.c                     |    2 
 fs/nfs/nfs4namespace.c                 |    4 
 fs/proc/root.c                         |   13 +
 fs/sb_config.c                         |   29 +-
 include/linux/container.h              |   91 ++++-
 include/linux/container_dev.h          |   25 +
 include/linux/cred.h                   |    3 
 include/linux/init_task.h              |    4 
 include/linux/kmod.h                   |    1 
 include/linux/lsm_hooks.h              |   25 +
 include/linux/mount.h                  |    5 
 include/linux/nsproxy.h                |    7 
 include/linux/pid.h                    |    5 
 include/linux/proc_ns.h                |    3 
 include/linux/sb_config.h              |    5 
 include/linux/sched.h                  |    3 
 include/linux/sched/task.h             |    4 
 include/linux/security.h               |   20 +
 include/linux/syscalls.h               |    6 
 include/uapi/linux/container.h         |   28 ++
 include/uapi/linux/fcntl.h             |    2 
 include/uapi/linux/magic.h             |    1 
 init/Kconfig                           |    7 
 init/main.c                            |    4 
 kernel/Makefile                        |    2 
 kernel/container.c                     |  576 ++++++++++++++++++++++++++++++++
 kernel/cred.c                          |   45 ++-
 kernel/exit.c                          |    1 
 kernel/fork.c                          |  117 ++++++-
 kernel/kmod.c                          |   13 +
 kernel/kthread.c                       |    3 
 kernel/namespaces.h                    |   15 +
 kernel/nsproxy.c                       |   34 +-
 kernel/pid.c                           |    4 
 kernel/sys_ni.c                        |    5 
 net/socket.c                           |   37 ++
 samples/containers/test-container.c    |  162 +++++++++
 security/security.c                    |   18 +
 security/selinux/hooks.c               |    5 
 47 files changed, 1408 insertions(+), 132 deletions(-)
 create mode 100644 include/linux/container_dev.h
 create mode 100644 include/uapi/linux/container.h
 create mode 100644 kernel/container.c
 create mode 100644 kernel/namespaces.h
 create mode 100644 samples/containers/test-container.c

[toc] | [next] | [standalone]


#1647125 — [PATCH 8/9] Honour CONTAINER_NEW_EMPTY_FS_NS

FromDavid Howells <dhowells@redhat.com>
Date2017-05-22 18:30 +0200
Subject[PATCH 8/9] Honour CONTAINER_NEW_EMPTY_FS_NS
Message-ID<tJYnx-80c-63@gated-at.bofh.it>
In reply to#1647121
Allow a container to be created with an empty mount namespace, as specified
by passing CONTAINER_NEW_EMPTY_FS_NS to container_create(), and allow a
root filesystem to be mounted into the container:

	cfd = container_create("foo", CONTAINER_NEW_EMPTY_FS_NS);
	fd = fsopen("ext3", cfd, 0);
	write(fd, "o foo");
	...
	fsmount(fd, -1, "/", AT_FSMOUNT_CONTAINER_ROOT, 0);
	close(fd);
	fd = fsopen("proc", cfd, 0);
	fsmount(fd, cfd, "/proc", 0, 0);
	close(fd);
---

 fs/namespace.c             |   84 ++++++++++++++++++++++++++++++++++++--------
 include/linux/mount.h      |    3 +-
 include/uapi/linux/fcntl.h |    2 +
 kernel/container.c         |    6 +++
 kernel/fork.c              |    5 ++-
 security/selinux/hooks.c   |    2 +
 6 files changed, 85 insertions(+), 17 deletions(-)

diff --git a/fs/namespace.c b/fs/namespace.c
index 9ca8b9f49f80..a365a7cba3ad 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2458,6 +2458,38 @@ static int do_add_mount(struct mount *newmnt, struct path *path, int mnt_flags,
 }
 
 static bool mount_too_revealing(struct vfsmount *mnt, int *new_mnt_flags);
+static struct mnt_namespace *create_mnt_ns(struct vfsmount *m);
+
+/*
+ * Create a mount namespace for a container and set the root mount in it.
+ */
+static int set_container_root(struct sb_config *sc, struct vfsmount *mnt)
+{
+	struct container *container = sc->container;
+	struct mnt_namespace *mnt_ns;
+	int ret = -EBUSY;
+
+	mnt_ns = create_mnt_ns(mnt);
+	if (IS_ERR(mnt_ns))
+		return PTR_ERR(mnt_ns);
+
+	spin_lock(&container->lock);
+	if (!container->ns->mnt_ns) {
+		container->ns->mnt_ns = mnt_ns;
+		write_seqcount_begin(&container->seq);
+		container->root.mnt = mnt;
+		container->root.dentry = mnt->mnt_root;
+		write_seqcount_end(&container->seq);
+		path_get(&container->root);
+		mnt_ns = NULL;
+		ret = 0;
+	}
+	spin_unlock(&container->lock);
+
+	if (ret < 0)
+		put_mnt_ns(mnt_ns);
+	return ret;
+}
 
 /*
  * Create a new mount using a superblock configuration and request it
@@ -2479,8 +2511,12 @@ static int do_new_mount_sc(struct sb_config *sc, struct path *mountpoint,
 		goto err_mnt;
 	}
 
-	ret = do_add_mount(real_mount(mnt), mountpoint, mnt_flags,
-			   sc->container ? sc->container->ns->mnt_ns : NULL);
+	if (mnt_flags & MNT_CONTAINER_ROOT)
+		ret = set_container_root(sc, mnt);
+	else
+		ret = do_add_mount(real_mount(mnt), mountpoint, mnt_flags,
+				   sc->container ? sc->container->ns->mnt_ns : NULL);
+
 	if (ret < 0) {
 		errorf("VFS: Failed to add mount");
 		goto err_mnt;
@@ -3262,10 +3298,17 @@ SYSCALL_DEFINE5(fsmount, int, fs_fd, int, dfd, const char __user *, dir_name,
 	struct fd f;
 	unsigned int lookup_flags, mnt_flags = 0;
 	long ret;
+	char buf[2];
 
 	if ((at_flags & ~(AT_SYMLINK_NOFOLLOW | AT_NO_AUTOMOUNT |
-			  AT_EMPTY_PATH)) != 0)
+			  AT_EMPTY_PATH | AT_FSMOUNT_CONTAINER_ROOT)) != 0)
 		return -EINVAL;
+	if (at_flags & AT_FSMOUNT_CONTAINER_ROOT) {
+		if (strncpy_from_user(buf, dir_name, 2) < 0)
+			return -EFAULT;
+		if (buf[0] != '/' || buf[1] != '\0')
+			return -EINVAL;
+	}
 
 	if (flags & ~(MS_RDONLY | MS_NOSUID | MS_NODEV | MS_NOEXEC |
 		      MS_NOATIME | MS_NODIRATIME | MS_RELATIME | MS_STRICTATIME))
@@ -3317,18 +3360,29 @@ SYSCALL_DEFINE5(fsmount, int, fs_fd, int, dfd, const char __user *, dir_name,
 	if (ret < 0)
 		goto err_fsfd;
 
-	/* Find the mountpoint.  A container can be specified in dfd. */
-	lookup_flags = LOOKUP_FOLLOW | LOOKUP_AUTOMOUNT;
-	if (at_flags & AT_SYMLINK_NOFOLLOW)
-		lookup_flags &= ~LOOKUP_FOLLOW;
-	if (at_flags & AT_NO_AUTOMOUNT)
-		lookup_flags &= ~LOOKUP_AUTOMOUNT;
-	if (at_flags & AT_EMPTY_PATH)
-		lookup_flags |= LOOKUP_EMPTY;
-	ret = user_path_at(dfd, dir_name, lookup_flags, &mountpoint);
-	if (ret < 0) {
-		errorf("VFS: Mountpoint lookup failed");
-		goto err_fsfd;
+	if (at_flags & AT_FSMOUNT_CONTAINER_ROOT) {
+		/* We're mounting the root of the container that was specified
+		 * to sys_fsopen().  The dir_name should be specified as "/"
+		 * and dfd is ignored.
+		 */
+		mountpoint.mnt = NULL;
+		mountpoint.dentry = NULL;
+		mnt_flags |= MNT_CONTAINER_ROOT;
+	} else {
+		/* Find the mountpoint.  A container can be specified in dfd. */
+		lookup_flags = LOOKUP_FOLLOW | LOOKUP_AUTOMOUNT;
+
+		if (at_flags & AT_SYMLINK_NOFOLLOW)
+			lookup_flags &= ~LOOKUP_FOLLOW;
+		if (at_flags & AT_NO_AUTOMOUNT)
+			lookup_flags &= ~LOOKUP_AUTOMOUNT;
+		if (at_flags & AT_EMPTY_PATH)
+			lookup_flags |= LOOKUP_EMPTY;
+		ret = user_path_at(dfd, dir_name, lookup_flags, &mountpoint);
+		if (ret < 0) {
+			errorf("VFS: Mountpoint lookup failed");
+			goto err_fsfd;
+		}
 	}
 
 	ret = security_sb_mountpoint(sc, &mountpoint);
diff --git a/include/linux/mount.h b/include/linux/mount.h
index 265e9aa2ab0b..480c6b4061e0 100644
--- a/include/linux/mount.h
+++ b/include/linux/mount.h
@@ -51,7 +51,8 @@ struct sb_config;
 #define MNT_INTERNAL_FLAGS (MNT_SHARED | MNT_WRITE_HOLD | MNT_INTERNAL | \
 			    MNT_DOOMED | MNT_SYNC_UMOUNT | MNT_MARKED)
 
-#define MNT_INTERNAL	0x4000
+#define MNT_INTERNAL		0x4000
+#define MNT_CONTAINER_ROOT	0x8000		/* Mounting a container root */
 
 #define MNT_LOCK_ATIME		0x040000
 #define MNT_LOCK_NOEXEC		0x080000
diff --git a/include/uapi/linux/fcntl.h b/include/uapi/linux/fcntl.h
index 813afd6eee71..747af8704bbf 100644
--- a/include/uapi/linux/fcntl.h
+++ b/include/uapi/linux/fcntl.h
@@ -68,5 +68,7 @@
 #define AT_STATX_FORCE_SYNC	0x2000	/* - Force the attributes to be sync'd with the server */
 #define AT_STATX_DONT_SYNC	0x4000	/* - Don't sync attributes with the server */
 
+#define AT_FSMOUNT_CONTAINER_ROOT	0x2000
+
 
 #endif /* _UAPI_LINUX_FCNTL_H */
diff --git a/kernel/container.c b/kernel/container.c
index 5ebbf548f01a..68276603d255 100644
--- a/kernel/container.c
+++ b/kernel/container.c
@@ -23,6 +23,7 @@
 #include <linux/printk.h>
 #include <linux/security.h>
 #include <linux/proc_fs.h>
+#include <linux/mnt_namespace.h>
 #include "namespaces.h"
 
 struct container init_container = {
@@ -500,6 +501,11 @@ static struct container *create_container(const char *name, unsigned int flags)
 	fs->root.mnt = NULL;
 	fs->root.dentry = NULL;
 
+	if (flags & CONTAINER_NEW_EMPTY_FS_NS) {
+		put_mnt_ns(ns->mnt_ns);
+		ns->mnt_ns = NULL;
+	}
+
 	ret = security_container_alloc(c, flags);
 	if (ret < 0)
 		goto err_fs;
diff --git a/kernel/fork.c b/kernel/fork.c
index 68cd7367fcd5..e5111d4bcc1c 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -2169,7 +2169,10 @@ SYSCALL_DEFINE1(fork_into_container, int, containerfd)
 	if (is_container_file(f.file)) {
 		struct container *c = f.file->private_data;
 
-		ret = _do_fork(SIGCHLD, 0, 0, NULL, NULL, 0, c);
+		if (!c->ns->mnt_ns)
+			ret = -ENOENT;
+		else
+			ret = _do_fork(SIGCHLD, 0, 0, NULL, NULL, 0, c);
 	}
 	fdput(f);
 	return ret;
diff --git a/security/selinux/hooks.c b/security/selinux/hooks.c
index 23bdbb0c2de5..f6b994b15a4d 100644
--- a/security/selinux/hooks.c
+++ b/security/selinux/hooks.c
@@ -2975,6 +2975,8 @@ static int selinux_sb_mountpoint(struct sb_config *sc, struct path *mountpoint)
 	const struct cred *cred = current_cred();
 	int ret;
 
+	if (!mountpoint->mnt)
+		return 0; /* This is the root in an empty namespace */
 	ret = path_has_perm(cred, mountpoint, FILE__MOUNTON);
 	if (ret < 0)
 		errorf("SELinux: Mount on mountpoint not permitted");

[toc] | [prev] | [next] | [standalone]


#1647126 — [PATCH 3/9] Provide /proc/containers

FromDavid Howells <dhowells@redhat.com>
Date2017-05-22 18:30 +0200
Subject[PATCH 3/9] Provide /proc/containers
Message-ID<tJYnx-80c-65@gated-at.bofh.it>
In reply to#1647121
Provide /proc/containers to view the current container and all the
containers created within it:

	# ./foo-container
	NAME                     USE FL OWNER GROUP
	<current>                141 01 0     0
	foo-test                   1 04 0     0

I'm not sure whether this is really desirable, though.

Signed-off-by: David Howells <dhowells@redhat.com>
---

 kernel/container.c |  104 ++++++++++++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 104 insertions(+)

diff --git a/kernel/container.c b/kernel/container.c
index eef1566835eb..d5849c07a76b 100644
--- a/kernel/container.c
+++ b/kernel/container.c
@@ -22,6 +22,7 @@
 #include <linux/syscalls.h>
 #include <linux/printk.h>
 #include <linux/security.h>
+#include <linux/proc_fs.h>
 #include "namespaces.h"
 
 struct container init_container = {
@@ -70,6 +71,108 @@ void put_container(struct container *c)
 	}
 }
 
+static void *container_proc_start(struct seq_file *m, loff_t *_pos)
+{
+	struct container *c = m->private;
+	struct list_head *p;
+	loff_t pos = *_pos;
+
+	spin_lock(&c->lock);
+
+	if (pos <= 1) {
+		*_pos = 1;
+		return (void *)1UL; /* Banner on first line */
+	}
+
+	if (pos == 2)
+		return m->private; /* Current container on second line */
+
+	/* Subordinate containers thereafter */
+	p = c->children.next;
+	pos--;
+	for (pos--; pos > 0 && p != &c->children; pos--) {
+		p = p->next;
+	}
+
+	if (p == &c->children)
+		return NULL;
+	return container_of(p, struct container, child_link);
+}
+
+static void *container_proc_next(struct seq_file *m, void *v, loff_t *_pos)
+{
+	struct container *c = m->private, *vc = v;
+	struct list_head *p;
+	loff_t pos = *_pos;
+
+	pos++;
+	*_pos = pos;
+	if (pos == 2)
+		return c; /* Current container on second line */
+
+	if (pos == 3)
+		p = &c->children;
+	else
+		p = &vc->child_link;
+	p = p->next;
+	if (p == &c->children)
+		return NULL;
+	return container_of(p, struct container, child_link);
+}
+
+static void container_proc_stop(struct seq_file *m, void *v)
+{
+	struct container *c = m->private;
+
+	spin_unlock(&c->lock);
+}
+
+static int container_proc_show(struct seq_file *m, void *v)
+{
+	struct user_namespace *uns = current_user_ns();
+	struct container *c = v;
+	const char *name;
+
+	if (v == (void *)1UL) {
+		seq_puts(m, "NAME                     USE FL OWNER GROUP\n");
+		return 0;
+	}
+
+	name = (c == m->private) ? "<current>" : c->name;
+	seq_printf(m, "%-24s %3u %02lx %0d %5d\n",
+		   name, refcount_read(&c->usage), c->flags,
+		   from_kuid_munged(uns, c->cred->uid),
+		   from_kgid_munged(uns, c->cred->gid));
+
+	return 0;
+}
+
+static const struct seq_operations container_proc_ops = {
+	.start	= container_proc_start,
+	.next	= container_proc_next,
+	.stop	= container_proc_stop,
+	.show	= container_proc_show,
+};
+
+static int container_proc_open(struct inode *inode, struct file *file)
+{
+	struct seq_file *m;
+	int ret = seq_open(file, &container_proc_ops);
+
+	if (ret == 0) {
+		m = file->private_data;
+		m->private = current->container;
+	}
+	return ret;
+}
+
+static const struct file_operations container_proc_fops = {
+	.open		= container_proc_open,
+	.read		= seq_read,
+	.llseek		= seq_lseek,
+	.release	= seq_release,
+};
+
 /*
  * Allow the user to poll for the container dying.
  */
@@ -230,6 +333,7 @@ static int __init init_container_fs(void)
 		panic("Cannot mount containerfs: %ld\n",
 		      PTR_ERR(containerfs_mnt));
 
+	proc_create("containers", 0, NULL, &container_proc_fops);
 	return 0;
 }
 

[toc] | [prev] | [next] | [standalone]


#1647163

FromJames Bottomley <James.Bottomley@HansenPartnership.com>
Date2017-05-22 19:00 +0200
Message-ID<tJYQy-8az-29@gated-at.bofh.it>
In reply to#1647121
[Added missing cc to containers list]
On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote:
> Here are a set of patches to define a container object for the kernel 
> and to provide some methods to create and manipulate them.
> 
> The reason I think this is necessary is that the kernel has no idea 
> how to direct upcalls to what userspace considers to be a container -
> current Linux practice appears to make a "container" just an 
> arbitrarily chosen junction of namespaces, control groups and files, 
> which may be changed individually within the "container".

This sounds like a step in the wrong direction: the strength of the
current container interfaces in Linux is that people who set up
containers don't have to agree what they look like.  So I can set up a
user namespace without a mount namespace or an architecture emulation
container with only a mount namespace.

But ignoring my fun foibles with containers and to give a concrete
example in terms of a popular orchestration system: in kubernetes,
where certain namespaces are shared across pods, do you imagine the
kernel's view of the "container" to be the pod or what kubernetes
thinks of as the container?  This is important, because half the
examples you give below are network related and usually pods share a
network namespace.

> The kernel upcall mechanism then needs to decide which set of 
> namespaces, etc., it must exec the appropriate upcall program. 
>  Examples of this include:
> 
>  (1) The DNS resolver.  The DNS cache in the kernel should probably 
> be per-network namespace, but in userspace the program, its
> libraries and its config data are associated with a mount tree and a 
> user namespace and it gets run in a particular pid namespace.

All persistent (written to fs data) has to be mount ns associated;
there are no ifs, ands and buts to that.  I agree this implies that if
you want to run a separate network namespace, you either take DNS from
the parent (a lot of containers do) or you set up a daemon to run
within the mount namespace.  I agree the latter is a slightly fiddly
operation you have to get right, but that's why we have orchestration
systems.

What is it we could do with the above that we cannot do today?

>  (2) NFS ID mapper.  The NFS ID mapping cache should also probably be
>      per-network namespace.

I think this is a view but not the only one:  Right at the moment, NFS
ID mapping is used as the one of the ways we can get the user namespace
ID mapping writes to file problems fixed ... that makes it a property
of the mount namespace for a lot of containers.  There are many other
instances where they do exactly as you say, but what I'm saying is that
we don't want to lose the flexibility we currently have.

>  (3) nfsdcltrack.  A way for NFSD to access stable storage for 
> tracking of persistent state.  Again, network-namespace dependent, 
> but also perhaps mount-namespace dependent.

So again, given we can set this up to work today, this sounds like more
a restriction that will bite us than an enhancement that gives us extra
features.

>  (4) General request-key upcalls.  Not particularly namespace 
> dependent, apart from keyrings being somewhat governed by the user
> namespace and the upcall being configured by the mount namespace.

All mount namespaces have an owning user namespace, so the data
relations are already there in the kernel, is the problem simply
finding them?

> These patches are built on top of the mount context patchset so that
> namespaces can be properly propagated over submounts/automounts.

I'll stop here ... you get the idea that I think this is imposing a set
of restrictions that will come back to bite us later.  If this is just
for the sake of figuring out how to get keyring upcalls to work, then
I'm sure we can come up with something.

James

[toc] | [prev] | [next] | [standalone]


#1647203

FromAleksa Sarai <asarai@suse.de>
Date2017-05-22 19:20 +0200
Message-ID<tJZ9U-8vX-7@gated-at.bofh.it>
In reply to#1647163
>> The reason I think this is necessary is that the kernel has no idea
>> how to direct upcalls to what userspace considers to be a container -
>> current Linux practice appears to make a "container" just an
>> arbitrarily chosen junction of namespaces, control groups and files,
>> which may be changed individually within the "container".

Just want to point out that if the kernel APIs for containers massively 
change, then the OCI will have to completely rework how we describe 
containers (and so will all existing runtimes).

Not to mention that while I don't like how hard it is (from a runtime 
perspective) to actually set up a container securely, there are 
undoubtedly benefits to having namespaces split out. The network 
namespace being separate means that in certain contexts you actually 
don't want to create a new network namespace when creating a container.

I had some ideas about how you could implement bridging in userspace (as 
an unprivileged user, for rootless containers) but if you can't join 
namespaces individually then such a setup is not practically possible.

-- 
Aleksa Sarai
Software Engineer (Containers)
SUSE Linux GmbH
https://www.cyphar.com/

[toc] | [prev] | [next] | [standalone]


#1648161

FromDavid Howells <dhowells@redhat.com>
Date2017-05-23 17:00 +0200
Message-ID<tKjrY-4z9-23@gated-at.bofh.it>
In reply to#1647203
Aleksa Sarai <asarai@suse.de> wrote:

> >> The reason I think this is necessary is that the kernel has no idea
> >> how to direct upcalls to what userspace considers to be a container -
> >> current Linux practice appears to make a "container" just an
> >> arbitrarily chosen junction of namespaces, control groups and files,
> >> which may be changed individually within the "container".
> 
> Just want to point out that if the kernel APIs for containers massively
> change, then the OCI will have to completely rework how we describe containers
> (and so will all existing runtimes).
> 
> Not to mention that while I don't like how hard it is (from a runtime
> perspective) to actually set up a container securely, there are undoubtedly
> benefits to having namespaces split out. The network namespace being separate
> means that in certain contexts you actually don't want to create a new network
> namespace when creating a container.

Yep, I quite agree.

However, certain things need to be made per-net namespace that *aren't*.  DNS
results, for instance.

As an example, I could set up a client machine with two ethernet ports, set up
two DNS+NFS servers, each of which think they're called "foo.bar" and attach
each server to a different port on the client machine.  Then I could create a
pair of containers on the client machine and route the network in each
container to a different port.  Now there's a problem because the names of the
cached DNS records for each port overlap.

Further, the NFS idmapper needs to be able to direct its calls to the
appropriate network.

> I had some ideas about how you could implement bridging in userspace (as an
> unprivileged user, for rootless containers) but if you can't join namespaces
> individually then such a setup is not practically possible.

I'm not proposing to take away the ability to arbitrarily set the namespaces
in a container.  I haven't implemented it yet, but it was on the to-do list:

 (7) Directly set a container's namespaces to allow cross-container
     sharing.

David

[toc] | [prev] | [next] | [standalone]


#1648164

Fromebiederm@xmission.com (Eric W. Biederman)
Date2017-05-23 17:10 +0200
Message-ID<tKjBD-4S4-3@gated-at.bofh.it>
In reply to#1648161
David Howells <dhowells@redhat.com> writes:

> Aleksa Sarai <asarai@suse.de> wrote:
>
>> >> The reason I think this is necessary is that the kernel has no idea
>> >> how to direct upcalls to what userspace considers to be a container -
>> >> current Linux practice appears to make a "container" just an
>> >> arbitrarily chosen junction of namespaces, control groups and files,
>> >> which may be changed individually within the "container".
>> 
>> Just want to point out that if the kernel APIs for containers massively
>> change, then the OCI will have to completely rework how we describe containers
>> (and so will all existing runtimes).
>> 
>> Not to mention that while I don't like how hard it is (from a runtime
>> perspective) to actually set up a container securely, there are undoubtedly
>> benefits to having namespaces split out. The network namespace being separate
>> means that in certain contexts you actually don't want to create a new network
>> namespace when creating a container.
>
> Yep, I quite agree.
>
> However, certain things need to be made per-net namespace that *aren't*.  DNS
> results, for instance.
>
> As an example, I could set up a client machine with two ethernet ports, set up
> two DNS+NFS servers, each of which think they're called "foo.bar" and attach
> each server to a different port on the client machine.  Then I could create a
> pair of containers on the client machine and route the network in each
> container to a different port.  Now there's a problem because the names of the
> cached DNS records for each port overlap.

Please look at ip netns add.  It does solve this in userspace rather
simply.

> Further, the NFS idmapper needs to be able to direct its calls to the
> appropriate network.

Eric

[toc] | [prev] | [next] | [standalone]


#1648182

FromDavid Howells <dhowells@redhat.com>
Date2017-05-23 17:20 +0200
Message-ID<tKjLl-4VL-33@gated-at.bofh.it>
In reply to#1648164
Eric W. Biederman <ebiederm@xmission.com> wrote:

> > As an example, I could set up a client machine with two ethernet ports,
> > set up two DNS+NFS servers, each of which think they're called "foo.bar"
> > and attach each server to a different port on the client machine.  Then I
> > could create a pair of containers on the client machine and route the
> > network in each container to a different port.  Now there's a problem
> > because the names of the cached DNS records for each port overlap.
> 
> Please look at ip netns add.

	warthog>man ip | grep setns
	warthog1>

> It does solve this in userspace rather simply.

Ummm...  How?  The kernel DNS resolver is not namespace aware.

David

[toc] | [prev] | [next] | [standalone]


#1648194

Fromebiederm@xmission.com (Eric W. Biederman)
Date2017-05-23 17:30 +0200
Message-ID<tKjV0-51B-25@gated-at.bofh.it>
In reply to#1648182
David Howells <dhowells@redhat.com> writes:

> Eric W. Biederman <ebiederm@xmission.com> wrote:
>
>> > As an example, I could set up a client machine with two ethernet ports,
>> > set up two DNS+NFS servers, each of which think they're called "foo.bar"
>> > and attach each server to a different port on the client machine.  Then I
>> > could create a pair of containers on the client machine and route the
>> > network in each container to a different port.  Now there's a problem
>> > because the names of the cached DNS records for each port overlap.
>> 
>> Please look at ip netns add.
>
> 	warthog>man ip | grep setns
> 	warthog1>

Not setns netns


>> It does solve this in userspace rather simply.
>
> Ummm...  How?  The kernel DNS resolver is not namespace aware.

But it works fine if called in the proper context and we have a defacto
standard for where to put all of the files (the tricky part) if you are
dealing with multiple network namespaces simultaneously.

Eric

[toc] | [prev] | [next] | [standalone]


#1648214

FromJames Bottomley <James.Bottomley@HansenPartnership.com>
Date2017-05-23 17:50 +0200
Message-ID<tKkel-58v-11@gated-at.bofh.it>
In reply to#1648194
On Tue, 2017-05-23 at 10:17 -0500, Eric W. Biederman wrote:
> David Howells <dhowells@redhat.com> writes:
> > Eric W. Biederman <ebiederm@xmission.com> wrote:
> > > It does solve this in userspace rather simply.
> > 
> > Ummm...  How?  The kernel DNS resolver is not namespace aware.
> 
> But it works fine if called in the proper context and we have a 
> defacto standard for where to put all of the files (the tricky part) 
> if you are dealing with multiple network namespaces simultaneously.

I think you're missing each other's points slightly.

What David is pointing out is that the kernel has a DNS cache
(net/dns_resolver/) it can do name to IP translations, but isn't
namespaced.  Once it has one entry all containers would see it if they
cause a lookup to go through the kernel cache, so going through the
cache you can't have a name resolving to different IP addresses on a
per container basis.

I think Eric's point is that if you need the same DNS names resolving
to different IP addresses on a per container basis, you can do this in
userspace today but you have to disable the in-kernel DNS cache.

James

[toc] | [prev] | [next] | [standalone]


#1648257

FromDavid Howells <dhowells@redhat.com>
Date2017-05-23 18:40 +0200
Message-ID<tKl0L-5Ft-25@gated-at.bofh.it>
In reply to#1648214
James Bottomley <James.Bottomley@HansenPartnership.com> wrote:

> What David is pointing out is that the kernel has a DNS cache
> (net/dns_resolver/) it can do name to IP translations, but isn't
> namespaced.  Once it has one entry all containers would see it if they
> cause a lookup to go through the kernel cache, so going through the
> cache you can't have a name resolving to different IP addresses on a
> per container basis.

Yes - and the transport to userspace, the request_key() upcall, isn't
namespaced either.  Namespacing it isn't entirely simple since we have to set
the right mount namespace (for execve, config, etc.), plus any other relevant
namespaces (such as network) - which is dependent on key type.

I can't record the mount namespace in the network namespace because that would
create a dependency loop:

	mnt_ns -> mnt -> sb -> net_ns -> mnt_ns

> I think Eric's point is that if you need the same DNS names resolving
> to different IP addresses on a per container basis, you can do this in
> userspace today but you have to disable the in-kernel DNS cache.

You could disable the in-kernel dns resolver in your config, but then you
don't get referrals in NFS.  Also, CIFS, AFS and other filesystems would be
affected.  If you're fine with the restrictions, then there is no problem.

David

[toc] | [prev] | [next] | [standalone]


#1649269

Fromebiederm@xmission.com (Eric W. Biederman)
Date2017-05-24 10:40 +0200
Message-ID<tKzZL-8mT-1@gated-at.bofh.it>
In reply to#1648257
David Howells <dhowells@redhat.com> writes:

> James Bottomley <James.Bottomley@HansenPartnership.com> wrote:
>
>> What David is pointing out is that the kernel has a DNS cache
>> (net/dns_resolver/) it can do name to IP translations, but isn't
>> namespaced.  Once it has one entry all containers would see it if they
>> cause a lookup to go through the kernel cache, so going through the
>> cache you can't have a name resolving to different IP addresses on a
>> per container basis.
>
> Yes - and the transport to userspace, the request_key() upcall, isn't
> namespaced either.  Namespacing it isn't entirely simple since we have to set
> the right mount namespace (for execve, config, etc.), plus any other relevant
> namespaces (such as network) - which is dependent on key type.
>
> I can't record the mount namespace in the network namespace because that would
> create a dependency loop:
>
> 	mnt_ns -> mnt -> sb -> net_ns -> mnt_ns

I have already given a concrete suggest on how this might be untangled.
So I won't repeat it here.

>> I think Eric's point is that if you need the same DNS names resolving
>> to different IP addresses on a per container basis, you can do this in
>> userspace today but you have to disable the in-kernel DNS cache.
>
> You could disable the in-kernel dns resolver in your config, but then you
> don't get referrals in NFS.  Also, CIFS, AFS and other filesystems would be
> affected.  If you're fine with the restrictions, then there is no
> problem.


I haven't been arguing that at all.  I was only pointing out that this
issue is not an issue with DNS.  Userspace handles this all fine.
The issue is exclusively with this request_key api and generally user
mode upcalls.

I have no problem seeing that there is an issue with the kernel code.
I am well aware of the problem.  Unfortunately the people who cared
enough to start addressing this have not been able to write kernel
code that fixes this.

My personal experience when I tried to use the request_key api at
the beginning of this was it was too hard to test.  There was no room
for goofing up as at that time it was impossible to invalidate a cached
reply from userspace if you happened to know it was wrong.  Which meant
that if something incorrect was cached it required rebooting the kernel.

I have a lot of sympathy with the view that the best way to do
some of this is with socket activations or perhaps something with rpc
portmapper.  Where something like inetd is used to start the user space
component on-demand.  I won't call that a solution to this case but I do
think it makes a good example to compare with.

When you need run something in a clean context having that something
only need to worry about the contents of the data it is receiving and
not about it's environment as suid applications do is a nice
simplification.

The entire user mode helper paradigm removes from user space the freedom
to specify what context it's code should run in.  In a world where
everything is global that is fine.  But in a world with containers where
not everything is global it becomes a royal pain.

And I am very very sympathetic to solving this.  The only solution that
I know would work is to capture the context at some point in a process
and then to use that process to fork user mode helpers.

So far no one has even bothered to seriously try the one solution that
is guaranteed to work because it takes a lot of changes to kernel code.
I believe the last effort snagged on what a pain it is to refactor the
user mode helper infrastructure.

I don't see in your code any of that work.

I am glad to see that you also see the problem.  At least when it comes
to the request_key api.

What I am hoping to see is someone who has the will to dig in and
understand all of the interactions and refactor the kernel to solve
the problem.

This is not a case where our user space interfaces are preventing a
solution to this problem (as your patchset implies).  This is a case
where things need to be refactored kernel side to solve this.

So far this attempt is just another in the bazillion or so bad
half-assed attempts to solve this problem I have seen over the years.

Eric

[toc] | [prev] | [next] | [standalone]


#1649363

FromIan Kent <raven@themaw.net>
Date2017-05-24 11:20 +0200
Message-ID<tKACt-pV-5@gated-at.bofh.it>
In reply to#1649269
On Wed, 2017-05-24 at 03:26 -0500, Eric W. Biederman wrote:
> 
> So far no one has even bothered to seriously try the one solution that
> is guaranteed to work because it takes a lot of changes to kernel code.
> I believe the last effort snagged on what a pain it is to refactor the
> user mode helper infrastructure.

Yes, that's mostly true in my case although I wouldn't say I haven't looked at
it seriously but equally I haven't got anything towards it yet either, sorry.

I'm likely going to revisit this based on a couple of approaches.

One is just what you describe and I had already been looking at this some time
ago. It seems to me that adding a work queue type that starts and retains a
process until the work queue is destroyed (similar to the way the work queue sub
system starts a fail over thread for use under resource exhaustion) would be a
sensible way to do it.

This doesn't mean I think it's a good idea for reasons I've outlined in the past
but the approach does warrant the effort to work out if it can be used without
problems.

And there's also the request key infrastructure which, as it is now, gets in the
road of verifying results, *sigh*.

Ian

[toc] | [prev] | [next] | [standalone]


#1647206

FromJessica Frazelle <me@jessfraz.com>
Date2017-05-22 19:30 +0200
Message-ID<tJZjz-7i-1@gated-at.bofh.it>
In reply to#1647163
I had replied but not to the thread with the containers mailing list.
See https://marc.info/?l=linux-cgroups&m=149547317006676&w=2

On Mon, May 22, 2017 at 5:53 PM, James Bottomley
<James.Bottomley@hansenpartnership.com> wrote:
> [Added missing cc to containers list]
> On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote:
>> Here are a set of patches to define a container object for the kernel
>> and to provide some methods to create and manipulate them.
>>
>> The reason I think this is necessary is that the kernel has no idea
>> how to direct upcalls to what userspace considers to be a container -
>> current Linux practice appears to make a "container" just an
>> arbitrarily chosen junction of namespaces, control groups and files,
>> which may be changed individually within the "container".
>
> This sounds like a step in the wrong direction: the strength of the
> current container interfaces in Linux is that people who set up
> containers don't have to agree what they look like.  So I can set up a
> user namespace without a mount namespace or an architecture emulation
> container with only a mount namespace.
>
> But ignoring my fun foibles with containers and to give a concrete
> example in terms of a popular orchestration system: in kubernetes,
> where certain namespaces are shared across pods, do you imagine the
> kernel's view of the "container" to be the pod or what kubernetes
> thinks of as the container?  This is important, because half the
> examples you give below are network related and usually pods share a
> network namespace.

I am glad you pointed this out because I was trying to make the same
point, various definitions of containers differ and who is to say
whether the various container runtimes (runc, rkt, systemd-nspawn) or
consumers of containers (kubernetes) won't modify their definition in
the future. How will this scale as new LSMs like Landlock or new
namespaces are added in the future will they be included in the
container kernel object as well...

Seems like a lot more maintenance for something that is really just
making the keyring namespace-aware... unless there are other things I
missed.

>
>> The kernel upcall mechanism then needs to decide which set of
>> namespaces, etc., it must exec the appropriate upcall program.
>>  Examples of this include:
>>
>>  (1) The DNS resolver.  The DNS cache in the kernel should probably
>> be per-network namespace, but in userspace the program, its
>> libraries and its config data are associated with a mount tree and a
>> user namespace and it gets run in a particular pid namespace.
>
> All persistent (written to fs data) has to be mount ns associated;
> there are no ifs, ands and buts to that.  I agree this implies that if
> you want to run a separate network namespace, you either take DNS from
> the parent (a lot of containers do) or you set up a daemon to run
> within the mount namespace.  I agree the latter is a slightly fiddly
> operation you have to get right, but that's why we have orchestration
> systems.
>
> What is it we could do with the above that we cannot do today?
>
>>  (2) NFS ID mapper.  The NFS ID mapping cache should also probably be
>>      per-network namespace.
>
> I think this is a view but not the only one:  Right at the moment, NFS
> ID mapping is used as the one of the ways we can get the user namespace
> ID mapping writes to file problems fixed ... that makes it a property
> of the mount namespace for a lot of containers.  There are many other
> instances where they do exactly as you say, but what I'm saying is that
> we don't want to lose the flexibility we currently have.
>
>>  (3) nfsdcltrack.  A way for NFSD to access stable storage for
>> tracking of persistent state.  Again, network-namespace dependent,
>> but also perhaps mount-namespace dependent.
>
> So again, given we can set this up to work today, this sounds like more
> a restriction that will bite us than an enhancement that gives us extra
> features.
>
>>  (4) General request-key upcalls.  Not particularly namespace
>> dependent, apart from keyrings being somewhat governed by the user
>> namespace and the upcall being configured by the mount namespace.
>
> All mount namespaces have an owning user namespace, so the data
> relations are already there in the kernel, is the problem simply
> finding them?
>
>> These patches are built on top of the mount context patchset so that
>> namespaces can be properly propagated over submounts/automounts.
>
> I'll stop here ... you get the idea that I think this is imposing a set
> of restrictions that will come back to bite us later.  If this is just
> for the sake of figuring out how to get keyring upcalls to work, then
> I'm sure we can come up with something.
>
> James
>
> --
> To unsubscribe from this list: send the line "unsubscribe cgroups" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html



-- 


Jessie Frazelle
4096R / D4C4 DD60 0D66 F65A 8EFC  511E 18F3 685C 0022 BFF3
pgp.mit.edu

[toc] | [prev] | [next] | [standalone]


#1647255

FromJeff Layton <jlayton@redhat.com>
Date2017-05-22 20:40 +0200
Message-ID<tK0pj-KA-9@gated-at.bofh.it>
In reply to#1647163
On Mon, 2017-05-22 at 09:53 -0700, James Bottomley wrote:
> [Added missing cc to containers list]
> On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote:
> > Here are a set of patches to define a container object for the kernel 
> > and to provide some methods to create and manipulate them.
> > 
> > The reason I think this is necessary is that the kernel has no idea 
> > how to direct upcalls to what userspace considers to be a container -
> > current Linux practice appears to make a "container" just an 
> > arbitrarily chosen junction of namespaces, control groups and files, 
> > which may be changed individually within the "container".
> 
> This sounds like a step in the wrong direction: the strength of the
> current container interfaces in Linux is that people who set up
> containers don't have to agree what they look like.  So I can set up a
> user namespace without a mount namespace or an architecture emulation
> container with only a mount namespace.
> 

Does this really mandate what they look like though? AFAICT, you can
still spawn disconnected namespaces to your heart's content. What this
does is provide a container for several different namespaces so that the
kernel can actually be aware of the association between them. The way
you populate the different namespaces looks to be pretty flexible.

> But ignoring my fun foibles with containers and to give a concrete
> example in terms of a popular orchestration system: in kubernetes,
> where certain namespaces are shared across pods, do you imagine the
> kernel's view of the "container" to be the pod or what kubernetes
> thinks of as the container?  This is important, because half the
> examples you give below are network related and usually pods share a
> network namespace.
> 
> > The kernel upcall mechanism then needs to decide which set of 
> > namespaces, etc., it must exec the appropriate upcall program. 
> >  Examples of this include:
> > 
> >  (1) The DNS resolver.  The DNS cache in the kernel should probably 
> > be per-network namespace, but in userspace the program, its
> > libraries and its config data are associated with a mount tree and a 
> > user namespace and it gets run in a particular pid namespace.
> 
> All persistent (written to fs data) has to be mount ns associated;
> there are no ifs, ands and buts to that.  I agree this implies that if
> you want to run a separate network namespace, you either take DNS from
> the parent (a lot of containers do) or you set up a daemon to run
> within the mount namespace.  I agree the latter is a slightly fiddly
> operation you have to get right, but that's why we have orchestration
> systems.
> 
> What is it we could do with the above that we cannot do today?
> 

Spawn a task directly from the kernel, already set up in the correct
namespaces, a'la call_usermodehelper. So far there is no way to do that,
and it is something we'd very much desire. Ian Kent has made several
passes at it recently.

> >  (2) NFS ID mapper.  The NFS ID mapping cache should also probably be
> >      per-network namespace.
> 
> I think this is a view but not the only one:  Right at the moment, NFS
> ID mapping is used as the one of the ways we can get the user namespace
> ID mapping writes to file problems fixed ... that makes it a property
> of the mount namespace for a lot of containers.  There are many other
> instances where they do exactly as you say, but what I'm saying is that
> we don't want to lose the flexibility we currently have.
> 
> >  (3) nfsdcltrack.  A way for NFSD to access stable storage for 
> > tracking of persistent state.  Again, network-namespace dependent, 
> > but also perhaps mount-namespace dependent.

Definitely mount-namespace dependent.

> 
> So again, given we can set this up to work today, this sounds like more
> a restriction that will bite us than an enhancement that gives us extra
> features.
> 

How do you set this up to work today?

AFAIK, if you want to run knfsd in a container today, you're out of luck
for any non-trivial configuration. The main reason is that most of knfsd
is namespace-ized in the network namespace, but there is no clear way to
associate that with a mount namespace, which is what we need to do this
properly inside a container. I think David's patches would get us there.

> >  (4) General request-key upcalls.  Not particularly namespace 
> > dependent, apart from keyrings being somewhat governed by the user
> > namespace and the upcall being configured by the mount namespace.
> 
> All mount namespaces have an owning user namespace, so the data
> relations are already there in the kernel, is the problem simply
> finding them?
> 
> > These patches are built on top of the mount context patchset so that
> > namespaces can be properly propagated over submounts/automounts.
> 
> I'll stop here ... you get the idea that I think this is imposing a set
> of restrictions that will come back to bite us later.  If this is just
> for the sake of figuring out how to get keyring upcalls to work, then
> I'm sure we can come up with something.
> 

-- 
Jeff Layton <jlayton@redhat.com>

[toc] | [prev] | [next] | [standalone]


#1647295

FromJames Bottomley <James.Bottomley@HansenPartnership.com>
Date2017-05-22 21:30 +0200
Message-ID<tK1bI-1hH-17@gated-at.bofh.it>
In reply to#1647255
On Mon, 2017-05-22 at 14:34 -0400, Jeff Layton wrote:
> On Mon, 2017-05-22 at 09:53 -0700, James Bottomley wrote:
> > [Added missing cc to containers list]
> > On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote:
> > > Here are a set of patches to define a container object for the 
> > > kernel and to provide some methods to create and manipulate them.
> > > 
> > > The reason I think this is necessary is that the kernel has no 
> > > idea how to direct upcalls to what userspace considers to be a
> > > container - current Linux practice appears to make a "container" 
> > > just an arbitrarily chosen junction of namespaces, control groups 
> > > and files, which may be changed individually within the
> > > "container".
> > 
> > This sounds like a step in the wrong direction: the strength of the
> > current container interfaces in Linux is that people who set up
> > containers don't have to agree what they look like.  So I can set 
> > up a user namespace without a mount namespace or an architecture
> > emulation container with only a mount namespace.
> > 
> 
> Does this really mandate what they look like though? AFAICT, you can
> still spawn disconnected namespaces to your heart's content. What 
> this does is provide a container for several different namespaces so 
> that the kernel can actually be aware of the association between 
> them.

Yes, because it imposes a view of what is in a container.  As the
several replies have pointed out (and indeed as I pointed out below for
kubernetes), this isn't something the orchestration systems would find
usable.

>  The way you populate the different namespaces looks to be pretty
> flexible.

OK, but look at it another way: If we provides a container API no
actual consumer of container technologies wants to use just because we
think it makes certain tasks easy, is it really a good API?

Containers are multi-layered and complex.  If you're not ready for this
as a user, then you should use an orchestration system that prevents
you from screwing up.

> > But ignoring my fun foibles with containers and to give a concrete
> > example in terms of a popular orchestration system: in kubernetes,
> > where certain namespaces are shared across pods, do you imagine the
> > kernel's view of the "container" to be the pod or what kubernetes
> > thinks of as the container?  This is important, because half the
> > examples you give below are network related and usually pods share 
> > a network namespace.
> > 
> > > The kernel upcall mechanism then needs to decide which set of 
> > > namespaces, etc., it must exec the appropriate upcall program. 
> > >  Examples of this include:
> > > 
> > >  (1) The DNS resolver.  The DNS cache in the kernel should 
> > > probably be per-network namespace, but in userspace the program, 
> > > its libraries and its config data are associated with a mount 
> > > tree and a user namespace and it gets run in a particular pid
> > > namespace.
> > 
> > All persistent (written to fs data) has to be mount ns associated;
> > there are no ifs, ands and buts to that.  I agree this implies that 
> > if you want to run a separate network namespace, you either take 
> > DNS from the parent (a lot of containers do) or you set up a daemon 
> > to run within the mount namespace.  I agree the latter is a 
> > slightly fiddly operation you have to get right, but that's why we 
> > have orchestration systems.
> > 
> > What is it we could do with the above that we cannot do today?
> > 
> 
> Spawn a task directly from the kernel, already set up in the correct
> namespaces, a'la call_usermodehelper. So far there is no way to do
> that,

Today the usermode helper has to be namespace aware.  We spawn it into
the root namespace and it jumps into the correct namespace/cgroup
combination and re-executes itself or simply performs the requisite
task on behalf of the container.  Is this simple, no; does it work,
yes, provided the host OS is aware of what the container orchestration
system wants it to do.

> and it is something we'd very much desire. Ian Kent has made several
> passes at it recently.

Well, every time we try to remove some of the complexity from
userspace, we end up wrapping around the axle of what exactly we're
trying to achieve, yes.

> > >  (2) NFS ID mapper.  The NFS ID mapping cache should also 
> > > probably be per-network namespace.
> > 
> > I think this is a view but not the only one:  Right at the moment, 
> > NFS ID mapping is used as the one of the ways we can get the user
> > namespace ID mapping writes to file problems fixed ... that makes 
> > it a property of the mount namespace for a lot of containers. 
> >  There are many other instances where they do exactly as you say, 
> > but what I'm saying is that we don't want to lose the flexibility
> > we currently have.
> > 
> > >  (3) nfsdcltrack.  A way for NFSD to access stable storage for 
> > > tracking of persistent state.  Again, network-namespace 
> > > dependent, but also perhaps mount-namespace dependent.
> 
> Definitely mount-namespace dependent.
> 
> > 
> > So again, given we can set this up to work today, this sounds like 
> > more a restriction that will bite us than an enhancement that gives 
> > us extra features.
> > 
> 
> How do you set this up to work today?

Well, as above, it spawns into the root, you jump it to where it should
be and re-execute or simply handle in the host. 

> AFAIK, if you want to run knfsd in a container today, you're out of 
> luck for any non-trivial configuration.

Well "running knfsd in a container" is actually different from having a
containerised nfs export.  My understanding was that thanks to the work
of Stas Kinsbursky, the latter has mostly worked since the 3.9 kernel
for v3 and below.  I assume the current issue is that there's a problem
with v4?

James

>  The main reason is that most of knfsd is namespace-ized in the
> network namespace, but there is no clear way to associate that with a
> mount namespace, which is what we need to do this properly inside a
> container. I think David's patches would get us there.
>
> > >  (4) General request-key upcalls.  Not particularly namespace 
> > > dependent, apart from keyrings being somewhat governed by the 
> > > user namespace and the upcall being configured by the mount
> > > namespace.
> > 
> > All mount namespaces have an owning user namespace, so the data
> > relations are already there in the kernel, is the problem simply
> > finding them?
> > 
> > > These patches are built on top of the mount context patchset so 
> > > that namespaces can be properly propagated over
> > > submounts/automounts.
> > 
> > I'll stop here ... you get the idea that I think this is imposing a 
> > set of restrictions that will come back to bite us later.  If this 
> > is just for the sake of figuring out how to get keyring upcalls to 
> > work, then I'm sure we can come up with something.
> > 
> 

[toc] | [prev] | [next] | [standalone]


#1647437

FromJeff Layton <jlayton@redhat.com>
Date2017-05-23 00:20 +0200
Message-ID<tK3Qd-30F-1@gated-at.bofh.it>
In reply to#1647295
On Mon, 2017-05-22 at 12:21 -0700, James Bottomley wrote:
> On Mon, 2017-05-22 at 14:34 -0400, Jeff Layton wrote:
> > On Mon, 2017-05-22 at 09:53 -0700, James Bottomley wrote:
> > > [Added missing cc to containers list]
> > > On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote:
> > > > Here are a set of patches to define a container object for the 
> > > > kernel and to provide some methods to create and manipulate them.
> > > > 
> > > > The reason I think this is necessary is that the kernel has no 
> > > > idea how to direct upcalls to what userspace considers to be a
> > > > container - current Linux practice appears to make a "container" 
> > > > just an arbitrarily chosen junction of namespaces, control groups 
> > > > and files, which may be changed individually within the
> > > > "container".
> > > 
> > > This sounds like a step in the wrong direction: the strength of the
> > > current container interfaces in Linux is that people who set up
> > > containers don't have to agree what they look like.  So I can set 
> > > up a user namespace without a mount namespace or an architecture
> > > emulation container with only a mount namespace.
> > > 
> > 
> > Does this really mandate what they look like though? AFAICT, you can
> > still spawn disconnected namespaces to your heart's content. What 
> > this does is provide a container for several different namespaces so 
> > that the kernel can actually be aware of the association between 
> > them.
> 
> Yes, because it imposes a view of what is in a container.  As the
> several replies have pointed out (and indeed as I pointed out below for
> kubernetes), this isn't something the orchestration systems would find
> usable.
> 
> >  The way you populate the different namespaces looks to be pretty
> > flexible.
> 
> OK, but look at it another way: If we provides a container API no
> actual consumer of container technologies wants to use just because we
> think it makes certain tasks easy, is it really a good API?
> 
> Containers are multi-layered and complex.  If you're not ready for this
> as a user, then you should use an orchestration system that prevents
> you from screwing up.
> 
> > > But ignoring my fun foibles with containers and to give a concrete
> > > example in terms of a popular orchestration system: in kubernetes,
> > > where certain namespaces are shared across pods, do you imagine the
> > > kernel's view of the "container" to be the pod or what kubernetes
> > > thinks of as the container?  This is important, because half the
> > > examples you give below are network related and usually pods share 
> > > a network namespace.
> > > 
> > > > The kernel upcall mechanism then needs to decide which set of 
> > > > namespaces, etc., it must exec the appropriate upcall program. 
> > > >  Examples of this include:
> > > > 
> > > >  (1) The DNS resolver.  The DNS cache in the kernel should 
> > > > probably be per-network namespace, but in userspace the program, 
> > > > its libraries and its config data are associated with a mount 
> > > > tree and a user namespace and it gets run in a particular pid
> > > > namespace.
> > > 
> > > All persistent (written to fs data) has to be mount ns associated;
> > > there are no ifs, ands and buts to that.  I agree this implies that 
> > > if you want to run a separate network namespace, you either take 
> > > DNS from the parent (a lot of containers do) or you set up a daemon 
> > > to run within the mount namespace.  I agree the latter is a 
> > > slightly fiddly operation you have to get right, but that's why we 
> > > have orchestration systems.
> > > 
> > > What is it we could do with the above that we cannot do today?
> > > 
> > 
> > Spawn a task directly from the kernel, already set up in the correct
> > namespaces, a'la call_usermodehelper. So far there is no way to do
> > that,
> 
> Today the usermode helper has to be namespace aware.  We spawn it into
> the root namespace and it jumps into the correct namespace/cgroup
> combination and re-executes itself or simply performs the requisite
> task on behalf of the container.  Is this simple, no; does it work,
> yes, provided the host OS is aware of what the container orchestration
> system wants it to do.
> 
> > and it is something we'd very much desire. Ian Kent has made several
> > passes at it recently.
> 
> Well, every time we try to remove some of the complexity from
> userspace, we end up wrapping around the axle of what exactly we're
> trying to achieve, yes.
> 
> > > >  (2) NFS ID mapper.  The NFS ID mapping cache should also 
> > > > probably be per-network namespace.
> > > 
> > > I think this is a view but not the only one:  Right at the moment, 
> > > NFS ID mapping is used as the one of the ways we can get the user
> > > namespace ID mapping writes to file problems fixed ... that makes 
> > > it a property of the mount namespace for a lot of containers. 
> > >  There are many other instances where they do exactly as you say, 
> > > but what I'm saying is that we don't want to lose the flexibility
> > > we currently have.
> > > 
> > > >  (3) nfsdcltrack.  A way for NFSD to access stable storage for 
> > > > tracking of persistent state.  Again, network-namespace 
> > > > dependent, but also perhaps mount-namespace dependent.
> > 
> > Definitely mount-namespace dependent.
> > 
> > > 
> > > So again, given we can set this up to work today, this sounds like 
> > > more a restriction that will bite us than an enhancement that gives 
> > > us extra features.
> > > 
> > 
> > How do you set this up to work today?
> 
> Well, as above, it spawns into the root, you jump it to where it should
> be and re-execute or simply handle in the host. 
> 
> > AFAIK, if you want to run knfsd in a container today, you're out of 
> > luck for any non-trivial configuration.
> 
> Well "running knfsd in a container" is actually different from having a
> containerised nfs export.  My understanding was that thanks to the work
> of Stas Kinsbursky, the latter has mostly worked since the 3.9 kernel
> for v3 and below.  I assume the current issue is that there's a problem
> with v4?
> 

Yes -- v3 mostly works because the equivalent state-tracking (rpc.statd)
is run as a long-running daemon.

nfsdcltrack uses call_usermodehelper, so for that you need to be able to
determine what mount namespace to run the thing in. All we really know
in knfsd when we want to do an upcall is the net namespace. We could
really use a way to associate the two and spawn the thing in the correct
container (or pass it enough info for it to setns() into the right
ones).

In principle, we could just ensure that we do all of this sort of thing
with long-running daemons that are started whenever the container
starts. But...having to run daemons full-time for infrequently-used
services sort of sucks and requires it to be setup. UMH helpers just get
run as long as the binary is in the right place.

I've also been reading over Eric suggestion, and that seems like it
might work as well though.

> >  The main reason is that most of knfsd is namespace-ized in the
> > network namespace, but there is no clear way to associate that with a
> > mount namespace, which is what we need to do this properly inside a
> > container. I think David's patches would get us there.
> > 
> > > >  (4) General request-key upcalls.  Not particularly namespace 
> > > > dependent, apart from keyrings being somewhat governed by the 
> > > > user namespace and the upcall being configured by the mount
> > > > namespace.
> > > 
> > > All mount namespaces have an owning user namespace, so the data
> > > relations are already there in the kernel, is the problem simply
> > > finding them?
> > > 
> > > > These patches are built on top of the mount context patchset so 
> > > > that namespaces can be properly propagated over
> > > > submounts/automounts.
> > > 
> > > I'll stop here ... you get the idea that I think this is imposing a 
> > > set of restrictions that will come back to bite us later.  If this 
> > > is just for the sake of figuring out how to get keyring upcalls to 
> > > work, then I'm sure we can come up with something.
> > > 
> 
> 

-- 
Jeff Layton <jlayton@redhat.com>

[toc] | [prev] | [next] | [standalone]


#1647924

FromIan Kent <raven@themaw.net>
Date2017-05-23 12:40 +0200
Message-ID<tKfol-1W6-5@gated-at.bofh.it>
In reply to#1647295
On Mon, 2017-05-22 at 12:21 -0700, James Bottomley wrote:
> 
> > > >  (3) nfsdcltrack.  A way for NFSD to access stable storage for 
> > > > tracking of persistent state.  Again, network-namespace 
> > > > dependent, but also perhaps mount-namespace dependent.
> > 
> > Definitely mount-namespace dependent.
> > 
> > > 
> > > So again, given we can set this up to work today, this sounds like 
> > > more a restriction that will bite us than an enhancement that gives 
> > > us extra features.
> > > 
> > 
> > How do you set this up to work today?
> 
> Well, as above, it spawns into the root, you jump it to where it should
> be and re-execute or simply handle in the host. 
> 
> > AFAIK, if you want to run knfsd in a container today, you're out of 
> > luck for any non-trivial configuration.
> 
> Well "running knfsd in a container" is actually different from having a
> containerised nfs export.  My understanding was that thanks to the work
> of Stas Kinsbursky, the latter has mostly worked since the 3.9 kernel
> for v3 and below.  I assume the current issue is that there's a problem
> with v4?

Oh, ok, I thought that, say, a docker (NFS) volumes-from a container to another
container didn't work for any version of NFS.

Certainly didn't work last time I tried, it was a while ago though.

Ian

[toc] | [prev] | [next] | [standalone]


#1647893

FromIan Kent <raven@themaw.net>
Date2017-05-23 11:40 +0200
Message-ID<tKesi-1hY-29@gated-at.bofh.it>
In reply to#1647163
On Mon, 2017-05-22 at 09:53 -0700, James Bottomley wrote:
> [Added missing cc to containers list]
> On Mon, 2017-05-22 at 17:22 +0100, David Howells wrote:
> > Here are a set of patches to define a container object for the kernel 
> > and to provide some methods to create and manipulate them.
> > 
> > The reason I think this is necessary is that the kernel has no idea 
> > how to direct upcalls to what userspace considers to be a container -
> > current Linux practice appears to make a "container" just an 
> > arbitrarily chosen junction of namespaces, control groups and files, 
> > which may be changed individually within the "container".
> 
> This sounds like a step in the wrong direction: the strength of the
> current container interfaces in Linux is that people who set up
> containers don't have to agree what they look like.  So I can set up a
> user namespace without a mount namespace or an architecture emulation
> container with only a mount namespace.
> 
> But ignoring my fun foibles with containers and to give a concrete
> example in terms of a popular orchestration system: in kubernetes,
> where certain namespaces are shared across pods, do you imagine the
> kernel's view of the "container" to be the pod or what kubernetes
> thinks of as the container?  This is important, because half the
> examples you give below are network related and usually pods share a
> network namespace.
> 
> > The kernel upcall mechanism then needs to decide which set of 
> > namespaces, etc., it must exec the appropriate upcall program. 
> >  Examples of this include:
> > 
> >  (1) The DNS resolver.  The DNS cache in the kernel should probably 
> > be per-network namespace, but in userspace the program, its
> > libraries and its config data are associated with a mount tree and a 
> > user namespace and it gets run in a particular pid namespace.
> 
> All persistent (written to fs data) has to be mount ns associated;
> there are no ifs, ands and buts to that.  I agree this implies that if
> you want to run a separate network namespace, you either take DNS from
> the parent (a lot of containers do) or you set up a daemon to run
> within the mount namespace.  I agree the latter is a slightly fiddly
> operation you have to get right, but that's why we have orchestration
> systems.
> 
> What is it we could do with the above that we cannot do today?
> 
> >  (2) NFS ID mapper.  The NFS ID mapping cache should also probably be
> >      per-network namespace.
> 
> I think this is a view but not the only one:  Right at the moment, NFS
> ID mapping is used as the one of the ways we can get the user namespace
> ID mapping writes to file problems fixed ... that makes it a property
> of the mount namespace for a lot of containers.  There are many other
> instances where they do exactly as you say, but what I'm saying is that
> we don't want to lose the flexibility we currently have.
> 
> >  (3) nfsdcltrack.  A way for NFSD to access stable storage for 
> > tracking of persistent state.  Again, network-namespace dependent, 
> > but also perhaps mount-namespace dependent.
> 
> So again, given we can set this up to work today, this sounds like more
> a restriction that will bite us than an enhancement that gives us extra
> features.
> 
> >  (4) General request-key upcalls.  Not particularly namespace 
> > dependent, apart from keyrings being somewhat governed by the user
> > namespace and the upcall being configured by the mount namespace.
> 
> All mount namespaces have an owning user namespace, so the data
> relations are already there in the kernel, is the problem simply
> finding them?
> 
> > These patches are built on top of the mount context patchset so that
> > namespaces can be properly propagated over submounts/automounts.
> 
> I'll stop here ... you get the idea that I think this is imposing a set
> of restrictions that will come back to bite us later.  If this is just
> for the sake of figuring out how to get keyring upcalls to work, then
> I'm sure we can come up with something.

You talk about a number of things I'm simply not aware of so I can't answer your
questions. But your points do sound like issues that need to be covered.

I think you mentioned user space used NFS ID mapper works fine.
I wonder, could you give more detail on that please.

Perhaps nsenter(1) is being used, I tried that as a possible usermode helper
solution and it probably did "work" in the sense of in container execution but
no-one liked it, it seems kernel folk expect to do things, well, in kernel.

Not only that there were other problems, probably request key sub system not
being namespace aware, or id caching within nfs or somewhere else, and there was
a question of not being able to cater for user namespace usage.

Anyway I do have a different view from my own experiences.

First there are a number of subsystems involved in creating a process from
within a container that has the container environment and, AFAICS (from the
usermode helper experience), it needs to be done from outside the container. For
example sub systems that need to be handled properly are the namespaces (and the
pid namespace in particular is tricky), credentials and cgroups, to name those
that come immediately to mind. I just couldn't get all that right after a number
of tries.

From this the problem that occurred to me is that we have a comprehensive
namespace implementation within the kernel but no container implementation to
help the binding together of the various sub systems for container use cases in
a way that satisfies container (or a process within an existing container)
creation.

The risk is that, as time passes, problems like usermode helper will be solved
in different places in different ways, not necessarily satisfactorily and
potentially hard to find when there are bugs and even harder to maintain than
the implementation here.

At least the interface here provides a "goto" place to define and maintain the
procedures required to do these things.

Yes, it would require change but change happens and the first pass may not be on
a par with what is currently done from a simplicity POV.

But why not see it as solving a kernel development problem and focus on what
needs to be done to make it on par with (and perhaps simpler) to cover the
current usage ....

Eric Biederman's comment is attractive indeed but I don't see how that solves
problem of having a containers sub system implementation which I think is
needed, if for no other reason than as a definition of correct usage and
localization of that usage for easier (right, nothing is easy) bug resolution.

Ian

[toc] | [prev] | [next] | [standalone]


#1648086

FromDavid Howells <dhowells@redhat.com>
Date2017-05-23 16:00 +0200
Message-ID<tKivU-3WY-19@gated-at.bofh.it>
In reply to#1647163
James Bottomley <James.Bottomley@HansenPartnership.com> wrote:

> This sounds like a step in the wrong direction: the strength of the
> current container interfaces in Linux is that people who set up
> containers don't have to agree what they look like.

It may be a strength, but it is also a problem.

> So I can set up a user namespace without a mount namespace or an
> architecture emulation container with only a mount namespace.

(I presume you mean with only the mount namespace separate)

Yep.  You can do that with this too.

> But ignoring my fun foibles with containers and to give a concrete
> example in terms of a popular orchestration system: in kubernetes,
> where certain namespaces are shared across pods, do you imagine the
> kernel's view of the "container" to be the pod or what kubernetes
> thinks of as the container?

Why not both?  If the net_ns is created in the pod container, then probably
network-related upcalls should be directed there.  Unless instructed
otherwise, upon creation a container object will inherit the caller's
namespaces.

> This is important, because half the examples you give below are network
> related and usually pods share a network namespace.

Yeah - I'm more familiar with upcalls made by NFS, AFS and keyrings.

> >  (1) The DNS resolver.  ...
> 
> All persistent (written to fs data) has to be mount ns associated;
> there are no ifs, ands and buts to that.  I agree this implies that if
> you want to run a separate network namespace, you either take DNS from
> the parent (a lot of containers do)

My intention is to make the DNS cache per-network namespace within the kernel.
Currently, there's only one and it's shared between all namespaces, but that
won't work if you end up with two net namespaces that direct the same
addresses to different things.

> or you set up a daemon to run within the mount namespace.

That's not currently an option: the DNS service upcalls only, and
/sbin/request-key is invoked in the init_ns.  This performs the network
accesses in the wrong network namespace.

> I agree the latter is a slightly fiddly operation you have to get right, but
> that's why we have orchestration systems.

An orchestration system can use this.  This is not a replacement for
Kubernetes or Docker or whatever.

> What is it we could do with the above that we cannot do today?

Upcall into an appropriate set of namespaces and keep the results separate by
network namespace.

> >  (2) NFS ID mapper.  The NFS ID mapping cache should also probably be
> >      per-network namespace.
> 
> I think this is a view but not the only one:  Right at the moment, NFS
> ID mapping is used as the one of the ways we can get the user namespace
> ID mapping writes to file problems fixed ... that makes it a property
> of the mount namespace for a lot of containers.

In some ways it's really a property of the server, and two different servers
may appear in two separate network namespaces with the same apparent name and
address.

It's not a property of the mount namespace because mount namespaces share
superblocks, and this is done at the superblock level.

Possibly it should be done on the vfsmount, as a filter on the interaction
between userspace and kernel.

> There are many other instances where they do exactly as you say, but what
> I'm saying is that we don't want to lose the flexibility we currently have.

You don't really lose any flexibility; if anything, you gain it.

(Note that in case your objection is that I haven't yet implemented the
ability to set namespaces arbitrarily in a namespace, that's on list of things
to do that I included, as is adjusting the control groups.)

> All mount namespaces have an owning user namespace, so the data
> relations are already there in the kernel, is the problem simply
> finding them?

The superblocks used by the vfsmounts in a mount namespace aren't all
necessarily in the same user_ns, so none of:

	sb->s_user_ns == current_user_ns()
	sb->s_user_ns == current->ns->mnt_ns->user_ns
	current->ns->mnt_ns->user_ns == current_user_ns()

need hold true that I can see.

> > These patches are built on top of the mount context patchset so that
> > namespaces can be properly propagated over submounts/automounts.
> 
> I'll stop here ... you get the idea that I think this is imposing a set
> of restrictions that will come back to bite us later.

What restrictions am I imposing?

> If this is just for the sake of figuring out how to get keyring upcalls to
> work, then I'm sure we can come up with something.

No, it's not just for that, though, admittedly, all of the upcall mechanisms I
outlined use request_key() at the core.


Really, a container is an anchor for the resources you need to make an upcall,
but it can also be used to anchor other things.

One thing I've been asked for by a number of people is a per-container keyring
for the provision of authentication keys, fs decryption keys and other things
- but there's no actual container off which this can be hung.

Another thing that could be useful is a list of what device files a container
may access, so that we can allow limited mounting by the container root user
within the container.

Now these could be made into their own namespaces or added to one that already
exists - perhaps the mount namespace being the most logical.

David

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.kernel


csiph-web