Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1315327 > unrolled thread
| Started by | Kees Cook <keescook@chromium.org> |
|---|---|
| First post | 2016-01-22 23:40 +0100 |
| Last post | 2016-01-25 20:00 +0100 |
| Articles | 20 on this page of 38 — 10 participants |
Back to article view | Back to linux.kernel
[PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Kees Cook <keescook@chromium.org> - 2016-01-22 23:40 +0100
[PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin Kees Cook <keescook@chromium.org> - 2016-01-22 23:40 +0100
Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin ebiederm@xmission.com (Eric W. Biederman) - 2016-01-23 04:30 +0100
Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin Jann Horn <jann@thejh.net> - 2016-01-23 23:30 +0100
Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin ebiederm@xmission.com (Eric W. Biederman) - 2016-01-24 02:40 +0100
Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin Al Viro <viro@ZenIV.linux.org.uk> - 2016-01-24 02:50 +0100
Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin Jann Horn <jann@thejh.net> - 2016-01-24 03:00 +0100
Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin ebiederm@xmission.com (Eric W. Biederman) - 2016-01-24 07:20 +0100
Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin Jann Horn <jann@thejh.net> - 2016-01-24 07:40 +0100
Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin ebiederm@xmission.com (Eric W. Biederman) - 2016-01-24 08:00 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Richard Weinberger <richard@nod.at> - 2016-01-22 23:50 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled ebiederm@xmission.com (Eric W. Biederman) - 2016-01-23 04:20 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Kees Cook <keescook@chromium.org> - 2016-01-24 22:00 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Serge Hallyn <serge.hallyn@ubuntu.com> - 2016-01-26 08:40 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Andy Lutomirski <luto@amacapital.net> - 2016-01-24 23:30 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Kees Cook <keescook@chromium.org> - 2016-01-25 20:00 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled ebiederm@xmission.com (Eric W. Biederman) - 2016-01-25 21:00 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Kees Cook <keescook@chromium.org> - 2016-01-25 23:40 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Andy Lutomirski <luto@amacapital.net> - 2016-01-26 00:40 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Daniel Micay <danielmicay@gmail.com> - 2016-01-26 03:30 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled ebiederm@xmission.com (Eric W. Biederman) - 2016-01-26 06:10 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Josh Boyer <jwboyer@fedoraproject.org> - 2016-01-26 15:40 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled "Austin S. Hemmelgarn" <ahferroin7@gmail.com> - 2016-01-26 15:50 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Josh Boyer <jwboyer@fedoraproject.org> - 2016-01-26 16:00 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Serge Hallyn <serge.hallyn@ubuntu.com> - 2016-01-26 18:30 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Josh Boyer <jwboyer@fedoraproject.org> - 2016-01-26 21:00 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled "Austin S. Hemmelgarn" <ahferroin7@gmail.com> - 2016-01-26 21:20 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Serge Hallyn <serge.hallyn@ubuntu.com> - 2016-01-26 18:20 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled "Austin S. Hemmelgarn" <ahferroin7@gmail.com> - 2016-01-26 19:20 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Andy Lutomirski <luto@amacapital.net> - 2016-01-26 19:30 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled "Austin S. Hemmelgarn" <ahferroin7@gmail.com> - 2016-01-26 19:50 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Kees Cook <keescook@chromium.org> - 2016-01-27 00:20 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Kees Cook <keescook@chromium.org> - 2016-01-27 00:20 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled ebiederm@xmission.com (Eric W. Biederman) - 2016-01-27 11:50 +0100
Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled "Austin S. Hemmelgarn" <ahferroin7@gmail.com> - 2016-01-27 13:40 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Kees Cook <keescook@chromium.org> - 2016-01-26 17:40 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Kees Cook <keescook@chromium.org> - 2016-01-25 20:00 +0100
Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled Andy Lutomirski <luto@amacapital.net> - 2016-01-25 20:00 +0100
Page 1 of 2 [1] 2 Next page →
| From | Kees Cook <keescook@chromium.org> |
|---|---|
| Date | 2016-01-22 23:40 +0100 |
| Subject | [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled |
| Message-ID | <qTSx5-20U-27@gated-at.bofh.it> |
There continues to be unexpected side-effects and security exposures via CLONE_NEWUSER. For many end-users running distro kernels with CONFIG_USER_NS enabled, there is no way to disable this feature when desired. As such, this creates a sysctl to restrict CLONE_NEWUSER so admins not running containers or Chrome can avoid the risks of this feature. -Kees
[toc] | [next] | [standalone]
| From | Kees Cook <keescook@chromium.org> |
|---|---|
| Date | 2016-01-22 23:40 +0100 |
| Subject | [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin |
| Message-ID | <qTSx6-20U-41@gated-at.bofh.it> |
| In reply to | #1315327 |
Several sysctls expect a state where the highest value (in extra2) is
locked once set for that boot. Yama does this, and kptr_restrict should
be doing it. This extracts Yama's logic and adds it to the existing
proc_dointvec_minmax_sysadmin, taking care to avoid the simple boolean
states (which do not get locked). Since Yama wants to be checking a
different capability, we build wrappers for both cases (CAP_SYS_ADMIN
and CAP_SYS_PTRACE).
Signed-off-by: Kees Cook <keescook@chromium.org>
---
Documentation/sysctl/kernel.txt | 4 +++-
include/linux/sysctl.h | 18 ++++++++++++++++++
kernel/sysctl.c | 34 +++++++++++++++++++++-------------
security/yama/yama_lsm.c | 18 +-----------------
4 files changed, 43 insertions(+), 31 deletions(-)
diff --git a/Documentation/sysctl/kernel.txt b/Documentation/sysctl/kernel.txt
index 73c6b1ef0e84..bbfc5e339a3d 100644
--- a/Documentation/sysctl/kernel.txt
+++ b/Documentation/sysctl/kernel.txt
@@ -385,7 +385,9 @@ to protect against uses of %pK in dmesg(8) if leaking kernel pointer
values to unprivileged users is a concern.
When kptr_restrict is set to (2), kernel pointers printed using
-%pK will be replaced with 0's regardless of privileges.
+%pK will be replaced with 0's regardless of privileges, and the value
+will be locked at "2", so that the root user cannot remove this
+restriction.
==============================================================
diff --git a/include/linux/sysctl.h b/include/linux/sysctl.h
index fa7bc29925c9..f8f0b991fe3e 100644
--- a/include/linux/sysctl.h
+++ b/include/linux/sysctl.h
@@ -23,6 +23,7 @@
#include <linux/list.h>
#include <linux/rcupdate.h>
+#include <linux/capability.h>
#include <linux/wait.h>
#include <linux/rbtree.h>
#include <uapi/linux/sysctl.h>
@@ -55,6 +56,23 @@ extern int proc_doulongvec_ms_jiffies_minmax(struct ctl_table *table, int,
void __user *, size_t *, loff_t *);
extern int proc_do_large_bitmap(struct ctl_table *, int,
void __user *, size_t *, loff_t *);
+extern int proc_dointvec_minmax_cap(int cap, struct ctl_table *table,
+ int write, void __user *buffer,
+ size_t *lenp, loff_t *ppos);
+static inline int proc_dointvec_minmax_cap_sysadmin(struct ctl_table *table,
+ int write, void __user *buffer,
+ size_t *lenp, loff_t *ppos)
+{
+ return proc_dointvec_minmax_cap(CAP_SYS_ADMIN, table, write, buffer,
+ lenp, ppos);
+}
+static inline int proc_dointvec_minmax_cap_ptrace(struct ctl_table *table,
+ int write, void __user *buffer,
+ size_t *lenp, loff_t *ppos)
+{
+ return proc_dointvec_minmax_cap(CAP_SYS_PTRACE, table, write, buffer,
+ lenp, ppos);
+}
/*
* Register a set of sysctl names by calling register_sysctl_table
diff --git a/kernel/sysctl.c b/kernel/sysctl.c
index c810f8afdb7f..fc8899dd636d 100644
--- a/kernel/sysctl.c
+++ b/kernel/sysctl.c
@@ -181,11 +181,6 @@ static int proc_taint(struct ctl_table *table, int write,
void __user *buffer, size_t *lenp, loff_t *ppos);
#endif
-#ifdef CONFIG_PRINTK
-static int proc_dointvec_minmax_sysadmin(struct ctl_table *table, int write,
- void __user *buffer, size_t *lenp, loff_t *ppos);
-#endif
-
static int proc_dointvec_minmax_coredump(struct ctl_table *table, int write,
void __user *buffer, size_t *lenp, loff_t *ppos);
#ifdef CONFIG_COREDUMP
@@ -803,7 +798,7 @@ static struct ctl_table kern_table[] = {
.data = &dmesg_restrict,
.maxlen = sizeof(int),
.mode = 0644,
- .proc_handler = proc_dointvec_minmax_sysadmin,
+ .proc_handler = proc_dointvec_minmax_cap_sysadmin,
.extra1 = &zero,
.extra2 = &one,
},
@@ -812,7 +807,7 @@ static struct ctl_table kern_table[] = {
.data = &kptr_restrict,
.maxlen = sizeof(int),
.mode = 0644,
- .proc_handler = proc_dointvec_minmax_sysadmin,
+ .proc_handler = proc_dointvec_minmax_cap_sysadmin,
.extra1 = &zero,
.extra2 = &two,
},
@@ -2217,16 +2212,29 @@ static int proc_taint(struct ctl_table *table, int write,
return err;
}
-#ifdef CONFIG_PRINTK
-static int proc_dointvec_minmax_sysadmin(struct ctl_table *table, int write,
- void __user *buffer, size_t *lenp, loff_t *ppos)
+int proc_dointvec_minmax_cap(int cap, struct ctl_table *table, int write,
+ void __user *buffer, size_t *lenp, loff_t *ppos)
{
- if (write && !capable(CAP_SYS_ADMIN))
+ struct ctl_table table_copy;
+ int value;
+
+ /* Require init capabilities to make changes. */
+ if (write && !capable(cap))
return -EPERM;
- return proc_dointvec_minmax(table, write, buffer, lenp, ppos);
+ /*
+ * To deal with const sysctl tables, we make a copy to perform
+ * the locking. When data is >1 and ==extra2, lock extra1 to
+ * extra2 to stop the value from being changed any further at
+ * runtime.
+ */
+ table_copy = *table;
+ value = *(int *)table_copy.data;
+ if (value > 1 && value == *(int *)table_copy.extra2)
+ table_copy.extra1 = table_copy.extra2;
+
+ return proc_dointvec_minmax(&table_copy, write, buffer, lenp, ppos);
}
-#endif
struct do_proc_dointvec_minmax_conv_param {
int *min;
diff --git a/security/yama/yama_lsm.c b/security/yama/yama_lsm.c
index d3c19c970a06..3215afd08fbd 100644
--- a/security/yama/yama_lsm.c
+++ b/security/yama/yama_lsm.c
@@ -354,22 +354,6 @@ static struct security_hook_list yama_hooks[] = {
};
#ifdef CONFIG_SYSCTL
-static int yama_dointvec_minmax(struct ctl_table *table, int write,
- void __user *buffer, size_t *lenp, loff_t *ppos)
-{
- struct ctl_table table_copy;
-
- if (write && !capable(CAP_SYS_PTRACE))
- return -EPERM;
-
- /* Lock the max value if it ever gets set. */
- table_copy = *table;
- if (*(int *)table_copy.data == *(int *)table_copy.extra2)
- table_copy.extra1 = table_copy.extra2;
-
- return proc_dointvec_minmax(&table_copy, write, buffer, lenp, ppos);
-}
-
static int zero;
static int max_scope = YAMA_SCOPE_NO_ATTACH;
@@ -385,7 +369,7 @@ static struct ctl_table yama_sysctl_table[] = {
.data = &ptrace_scope,
.maxlen = sizeof(int),
.mode = 0644,
- .proc_handler = yama_dointvec_minmax,
+ .proc_handler = proc_dointvec_minmax_cap_ptrace,
.extra1 = &zero,
.extra2 = &max_scope,
},
--
2.6.3
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2016-01-23 04:30 +0100 |
| Subject | Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin |
| Message-ID | <qTX3I-5bz-21@gated-at.bofh.it> |
| In reply to | #1315329 |
Kees Cook <keescook@chromium.org> writes:
> Several sysctls expect a state where the highest value (in extra2) is
> locked once set for that boot. Yama does this, and kptr_restrict should
> be doing it. This extracts Yama's logic and adds it to the existing
> proc_dointvec_minmax_sysadmin, taking care to avoid the simple boolean
> states (which do not get locked). Since Yama wants to be checking a
> different capability, we build wrappers for both cases (CAP_SYS_ADMIN
> and CAP_SYS_PTRACE).
Sigh this sysctl appears susceptible to known attacks.
In my quick skim I believe this sysctl implementation that checks
capabilities is susceptible to attacks where the already open file
descriptor is set as stdout on a setuid root application.
Can we come up with an interface that isn't exploitable by an
application that will act as a setuid cat?
Eric
> -#ifdef CONFIG_PRINTK
> -static int proc_dointvec_minmax_sysadmin(struct ctl_table *table, int write,
> - void __user *buffer, size_t *lenp, loff_t *ppos)
> +int proc_dointvec_minmax_cap(int cap, struct ctl_table *table, int write,
> + void __user *buffer, size_t *lenp, loff_t *ppos)
> {
> - if (write && !capable(CAP_SYS_ADMIN))
> + struct ctl_table table_copy;
> + int value;
> +
> + /* Require init capabilities to make changes. */
> + if (write && !capable(cap))
> return -EPERM;
>
> - return proc_dointvec_minmax(table, write, buffer, lenp, ppos);
> + /*
> + * To deal with const sysctl tables, we make a copy to perform
> + * the locking. When data is >1 and ==extra2, lock extra1 to
> + * extra2 to stop the value from being changed any further at
> + * runtime.
> + */
> + table_copy = *table;
> + value = *(int *)table_copy.data;
> + if (value > 1 && value == *(int *)table_copy.extra2)
> + table_copy.extra1 = table_copy.extra2;
> +
> + return proc_dointvec_minmax(&table_copy, write, buffer, lenp, ppos);
> }
> -#endif
[toc] | [prev] | [next] | [standalone]
| From | Jann Horn <jann@thejh.net> |
|---|---|
| Date | 2016-01-23 23:30 +0100 |
| Subject | Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin |
| Message-ID | <qUeQV-3GP-1@gated-at.bofh.it> |
| In reply to | #1315500 |
[Multipart message — attachments visible in raw view] — view raw
On Fri, Jan 22, 2016 at 09:10:07PM -0600, Eric W. Biederman wrote: > Kees Cook <keescook@chromium.org> writes: > > > Several sysctls expect a state where the highest value (in extra2) is > > locked once set for that boot. Yama does this, and kptr_restrict should > > be doing it. This extracts Yama's logic and adds it to the existing > > proc_dointvec_minmax_sysadmin, taking care to avoid the simple boolean > > states (which do not get locked). Since Yama wants to be checking a > > different capability, we build wrappers for both cases (CAP_SYS_ADMIN > > and CAP_SYS_PTRACE). > > Sigh this sysctl appears susceptible to known attacks. > > In my quick skim I believe this sysctl implementation that checks > capabilities is susceptible to attacks where the already open file > descriptor is set as stdout on a setuid root application. > > Can we come up with an interface that isn't exploitable by an > application that will act as a setuid cat? Adding the struct file * to the parameters of all proc_handler functions would work, right? (Or just filp->f_cred? That would be less generic.) A quick grep says that's just about 160 functions that'll need to be changed. :/
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2016-01-24 02:40 +0100 |
| Subject | Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin |
| Message-ID | <qUhON-5Q5-9@gated-at.bofh.it> |
| In reply to | #1315768 |
Jann Horn <jann@thejh.net> writes: > On Fri, Jan 22, 2016 at 09:10:07PM -0600, Eric W. Biederman wrote: >> Kees Cook <keescook@chromium.org> writes: >> >> > Several sysctls expect a state where the highest value (in extra2) is >> > locked once set for that boot. Yama does this, and kptr_restrict should >> > be doing it. This extracts Yama's logic and adds it to the existing >> > proc_dointvec_minmax_sysadmin, taking care to avoid the simple boolean >> > states (which do not get locked). Since Yama wants to be checking a >> > different capability, we build wrappers for both cases (CAP_SYS_ADMIN >> > and CAP_SYS_PTRACE). >> >> Sigh this sysctl appears susceptible to known attacks. >> >> In my quick skim I believe this sysctl implementation that checks >> capabilities is susceptible to attacks where the already open file >> descriptor is set as stdout on a setuid root application. >> >> Can we come up with an interface that isn't exploitable by an >> application that will act as a setuid cat? > > Adding the struct file * to the parameters of all proc_handler > functions would work, right? (Or just filp->f_cred? That would be > less generic.) > > A quick grep says that's just about 160 functions that'll need to > be changed. :/ Yep. That is about the size of it. file * used to be passed to the sysctl methods but it was removed several years ago because no one was using it. Eric
[toc] | [prev] | [next] | [standalone]
| From | Al Viro <viro@ZenIV.linux.org.uk> |
|---|---|
| Date | 2016-01-24 02:50 +0100 |
| Subject | Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin |
| Message-ID | <qUhYt-5To-1@gated-at.bofh.it> |
| In reply to | #1315803 |
On Sat, Jan 23, 2016 at 07:20:17PM -0600, Eric W. Biederman wrote: > Yep. That is about the size of it. file * used to be passed to the > sysctl methods but it was removed several years ago because no one was > using it. Generally cred would be better... Alternatively we could eat one more pointer in task_struct and stash a reference to that sucker there, rather than adding an explicit argument (again, with cred instead of file). Not sure...
[toc] | [prev] | [next] | [standalone]
| From | Jann Horn <jann@thejh.net> |
|---|---|
| Date | 2016-01-24 03:00 +0100 |
| Subject | Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin |
| Message-ID | <qUi8a-5WE-5@gated-at.bofh.it> |
| In reply to | #1315804 |
[Multipart message — attachments visible in raw view] — view raw
On Sun, Jan 24, 2016 at 01:43:42AM +0000, Al Viro wrote: > On Sat, Jan 23, 2016 at 07:20:17PM -0600, Eric W. Biederman wrote: > > > Yep. That is about the size of it. file * used to be passed to the > > sysctl methods but it was removed several years ago because no one was > > using it. > > Generally cred would be better... > Alternatively we could eat one more > pointer in task_struct and stash a reference to that sucker there, rather > than adding an explicit argument (again, with cred instead of file). > Not sure... I think it makes sense to do this the same way as the rest of the VFS code here (which passes the creds down through an argument). And adding the arguments everywhere doesn't really mean more work - either way, someone should probably go through all of those sysctl handlers and fix them up to use the file creds.
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2016-01-24 07:20 +0100 |
| Subject | Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin |
| Message-ID | <qUmbL-Hp-11@gated-at.bofh.it> |
| In reply to | #1315806 |
Jann Horn <jann@thejh.net> writes: > On Sun, Jan 24, 2016 at 01:43:42AM +0000, Al Viro wrote: >> On Sat, Jan 23, 2016 at 07:20:17PM -0600, Eric W. Biederman wrote: >> >> > Yep. That is about the size of it. file * used to be passed to the >> > sysctl methods but it was removed several years ago because no one was >> > using it. >> >> Generally cred would be better... > >> Alternatively we could eat one more >> pointer in task_struct and stash a reference to that sucker there, rather >> than adding an explicit argument (again, with cred instead of file). >> Not sure... > > I think it makes sense to do this the same way as the rest of the VFS code > here (which passes the creds down through an argument). > > And adding the arguments everywhere doesn't really mean more work - either > way, someone should probably go through all of those sysctl handlers and > fix them up to use the file creds. Not all of them need it. It might be worth figuring out the necessary rigamarole to hook into sysctl_perm the way the networking code does and have that require the capability at open time. The advantage is that open time is when it is actually appropraite to check permissions. I could be wrong but I doubt there is enough madness with the handful of sysctl users that call capable to require the checks to happen on write and not on open. Eric
[toc] | [prev] | [next] | [standalone]
| From | Jann Horn <jann@thejh.net> |
|---|---|
| Date | 2016-01-24 07:40 +0100 |
| Subject | Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin |
| Message-ID | <qUmv8-PM-3@gated-at.bofh.it> |
| In reply to | #1315832 |
[Multipart message — attachments visible in raw view] — view raw
On Sun, Jan 24, 2016 at 12:02:41AM -0600, Eric W. Biederman wrote: > Jann Horn <jann@thejh.net> writes: > > > On Sun, Jan 24, 2016 at 01:43:42AM +0000, Al Viro wrote: > >> On Sat, Jan 23, 2016 at 07:20:17PM -0600, Eric W. Biederman wrote: > >> > >> > Yep. That is about the size of it. file * used to be passed to the > >> > sysctl methods but it was removed several years ago because no one was > >> > using it. > >> > >> Generally cred would be better... > > > >> Alternatively we could eat one more > >> pointer in task_struct and stash a reference to that sucker there, rather > >> than adding an explicit argument (again, with cred instead of file). > >> Not sure... > > > > I think it makes sense to do this the same way as the rest of the VFS code > > here (which passes the creds down through an argument). > > > > And adding the arguments everywhere doesn't really mean more work - either > > way, someone should probably go through all of those sysctl handlers and > > fix them up to use the file creds. > > Not all of them need it. It might be worth figuring out the necessary > rigamarole to hook into sysctl_perm the way the networking code does and > have that require the capability at open time. > > The advantage is that open time is when it is actually appropraite to > check permissions. I could be wrong but I doubt there is enough madness > with the handful of sysctl users that call capable to require the checks > to happen on write and not on open. That would work - if all sysctls know whether a capability will be needed for writing later on and don't decide it based on the written data. Is that always true? Looking through some of the sysctl handlers, I found proc_do_uts_string and pid_ns_ctl_handler, which operate on a namespace looked up through current at write time. I think that's buggy and ought to be done using the file opener creds and on the file opener's namespaces, but where can those be stored?
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2016-01-24 08:00 +0100 |
| Subject | Re: [kernel-hardening] Re: [PATCH 1/2] sysctl: expand use of proc_dointvec_minmax_sysadmin |
| Message-ID | <qUmOu-Wv-7@gated-at.bofh.it> |
| In reply to | #1315834 |
Jann Horn <jann@thejh.net> writes: > On Sun, Jan 24, 2016 at 12:02:41AM -0600, Eric W. Biederman wrote: >> Jann Horn <jann@thejh.net> writes: >> >> > On Sun, Jan 24, 2016 at 01:43:42AM +0000, Al Viro wrote: >> >> On Sat, Jan 23, 2016 at 07:20:17PM -0600, Eric W. Biederman wrote: >> >> >> >> > Yep. That is about the size of it. file * used to be passed to the >> >> > sysctl methods but it was removed several years ago because no one was >> >> > using it. >> >> >> >> Generally cred would be better... >> > >> >> Alternatively we could eat one more >> >> pointer in task_struct and stash a reference to that sucker there, rather >> >> than adding an explicit argument (again, with cred instead of file). >> >> Not sure... >> > >> > I think it makes sense to do this the same way as the rest of the VFS code >> > here (which passes the creds down through an argument). >> > >> > And adding the arguments everywhere doesn't really mean more work - either >> > way, someone should probably go through all of those sysctl handlers and >> > fix them up to use the file creds. >> >> Not all of them need it. It might be worth figuring out the necessary >> rigamarole to hook into sysctl_perm the way the networking code does and >> have that require the capability at open time. >> >> The advantage is that open time is when it is actually appropraite to >> check permissions. I could be wrong but I doubt there is enough madness >> with the handful of sysctl users that call capable to require the checks >> to happen on write and not on open. > > That would work - if all sysctls know whether a capability will be needed > for writing later on and don't decide it based on the written data. Is that > always true? That is certainly the common case to stick a pointer to static data in struct sysctl_table. > Looking through some of the sysctl handlers, I found proc_do_uts_string and > pid_ns_ctl_handler, which operate on a namespace looked up through current > at write time. I think that's buggy and ought to be done using the file > opener creds and on the file opener's namespaces, but where can those be > stored? Interesting point. I had not though about it from that angle. From all other angles it has just been something that would be nice to fix but I hadn't seen the need. We have all of the infrastructure needed to register sysctls per namespace and the networking stack uses it. That would not be hard to use for uts, pid and ipc namespaces as well. That would remove any race. Eric
[toc] | [prev] | [next] | [standalone]
| From | Richard Weinberger <richard@nod.at> |
|---|---|
| Date | 2016-01-22 23:50 +0100 |
| Message-ID | <qTSGJ-24Y-3@gated-at.bofh.it> |
| In reply to | #1315327 |
Am 22.01.2016 um 23:39 schrieb Kees Cook: > There continues to be unexpected side-effects and security exposures > via CLONE_NEWUSER. For many end-users running distro kernels with > CONFIG_USER_NS enabled, there is no way to disable this feature when > desired. As such, this creates a sysctl to restrict CLONE_NEWUSER so > admins not running containers or Chrome can avoid the risks of this > feature. Last time such a patch came up I was not thrilled because hiding a scary feature behind a knob IMHO doesn't make it any better nor helps finding issues. But as userns is still a source of a lot of issues and distros enable it by default a knob for the admin seems to be a good idea by now. ;-\ Thanks, //richard
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2016-01-23 04:20 +0100 |
| Message-ID | <qTWU1-56B-1@gated-at.bofh.it> |
| In reply to | #1315327 |
Kees Cook <keescook@chromium.org> writes: > There continues to be unexpected side-effects and security exposures > via CLONE_NEWUSER. For many end-users running distro kernels with > CONFIG_USER_NS enabled, there is no way to disable this feature when > desired. As such, this creates a sysctl to restrict CLONE_NEWUSER so > admins not running containers or Chrome can avoid the risks of this > feature. I don't actually think there do continue to be unexpected side-effects and security exposures with CLONE_NEWUSER. It takes a while for all of the fixes to trickle out to distros. At most what I have seen recently are problems with other kernel interfaces being amplified with user namespaces. AKA the current mess with devpts, and the unexpected issues with bind mounts in mount namespaces. I have a couple of concerns with a sysctl. 1) As user namespaces settle out this sysctl has the potential to decrease the security of the system overall as sandboxing features of the kernel will not be available to unprivileged applications. Web browsing with chrome will be less safe for example. 2) I strongly suspect the granularity of a sysctl is wrong for access to user namespaces on a production system. In general I suspect what we want is something like seccomp. I believe all of the relevant bits are in registers. I actually thought that was enough for seccomp. Does seccomp not work for some reason? 3) A sysctl breeds a false sense of security in thinking that if a security issue is discovered you can just flip a switch, disable all new user namespaces and you won't be vulnerable. In fact most of the issues in the past have only required being in a user namespace to trigger. Which means any containers or user namespaces that already exist could be used to exploit any new found issue. Which means that a I don't think a sysctl will give the desired level of protection. In my analysis of the issues to date I don't know of anything short of a reboot that would meaninfully remove the threat. 4) With applications like docker coming on-line I don't think a restriction to processes with capabilities is actually meaninful for restricting access to user namespaces. So I have concerns about both efficacy and usability with the proposed sysctl. So to keep this productive. Please tell me about the threat model you envision, and how you envision knobs in the kernel being used to counter those threats. Eric
[toc] | [prev] | [next] | [standalone]
| From | Kees Cook <keescook@chromium.org> |
|---|---|
| Date | 2016-01-24 22:00 +0100 |
| Message-ID | <qUzVo-1M4-21@gated-at.bofh.it> |
| In reply to | #1315491 |
On Fri, Jan 22, 2016 at 7:02 PM, Eric W. Biederman <ebiederm@xmission.com> wrote: > Kees Cook <keescook@chromium.org> writes: > >> There continues to be unexpected side-effects and security exposures >> via CLONE_NEWUSER. For many end-users running distro kernels with >> CONFIG_USER_NS enabled, there is no way to disable this feature when >> desired. As such, this creates a sysctl to restrict CLONE_NEWUSER so >> admins not running containers or Chrome can avoid the risks of this >> feature. > > I don't actually think there do continue to be unexpected side-effects > and security exposures with CLONE_NEWUSER. It takes a while for all of > the fixes to trickle out to distros. At most what I have seen recently > are problems with other kernel interfaces being amplified with user > namespaces. AKA the current mess with devpts, and the unexpected > issues with bind mounts in mount namespaces. Access to CLONE_NEWUSER has lead to a lot of security issues over the last 3 years. There has to be a way to avoid this for people that have no interest in containers. For admins running servers where there are no containers (which is still a giant number of systems -- containers are popular but not ubiquitous), the sysctl makes perfect sense. > I have a couple of concerns with a sysctl. > > 1) As user namespaces settle out this sysctl has the potential to > decrease the security of the system overall as sandboxing > features of the kernel will not be available to unprivileged > applications. > > Web browsing with chrome will be less safe for example. I don't propose this for Desktops. > 2) I strongly suspect the granularity of a sysctl is wrong for access to > user namespaces on a production system. > > In general I suspect what we want is something like seccomp. I > believe all of the relevant bits are in registers. I actually > thought that was enough for seccomp. Does seccomp not work for > some reason? Setting a global seccomp filter on init is not possible with any inits yet, and for some architectures it would push all processes onto the slow path. It's an extraordinarily big hammer for wanting to turn off a single area of the kernel with a long history of problems. Also, seccomp is arguably a program author's policy tool, not a system policy tool. We could offer this sysctl as an LSM too, but that's even messier. This is a trivial change to user namespaces and provides a large protection to people that aren't interested in the risks of running containers. > 3) A sysctl breeds a false sense of security in thinking that if a > security issue is discovered you can just flip a switch, disable > all new user namespaces and you won't be vulnerable. > > In fact most of the issues in the past have only required being in > a user namespace to trigger. Which means any containers or user > namespaces that already exist could be used to exploit any new > found issue. Which means that a I don't think a sysctl will give > the desired level of protection. > > In my analysis of the issues to date I don't know of anything > short of a reboot that would meaninfully remove the threat. Any admin that decides to just turn off CLONE_NEWUSER in the middle of still using it is insane. I don't think this breeds any false sense of security as most sysctls are set at boot time. > 4) With applications like docker coming on-line I don't think a > restriction to processes with capabilities is actually meaninful > for restricting access to user namespaces. Admins who are currently using containers are already exposed to so much attack surface. This is not for them, it's for people that don't use containers. > So I have concerns about both efficacy and usability with the proposed > sysctl. Two distros already have this sysctl because it was so strongly requested by their users. This needs to be upstream so we can manage the effects correctly. > So to keep this productive. Please tell me about the threat model > you envision, and how you envision knobs in the kernel being used to > counter those threats. The threat model I envision is post-intrusion escalation of privileges on systems that run distro kernels and do not use containers. I envision the sysctl being used at boot time to kill the entire class of current and future vulnerabilities exposed by CLONE_NEWUSER. Just like the sysctls used to turn off modules at boot or turn off kexec at boot. As Linux developers I feel we have an obligation to provide our end users with run-time choices (not just compile-time choices), since most of our users are using kernels built by someone else. Given the repeated problems with module auto-loading, we provided a way to disable module loading. Given the physical-memory-rewriting exposure of kexec, we provides a way to disable kexec. Given the conflict between hibernation and kASLR, we provided a way to choose one at runtime. Here, we're looking back on three years of vulnerabilities around CLONE_NEWUSER with no end in sight, and we have an obligation to help the end users that don't want to be exposed to this any more. Note I'm not suggesting we stop trying to fix the problems we find with user namespaces, but we need to provide a way to disable them. Having this sysctl is vastly superior to telling people how to rewrite their kernel memory at boot time to disable syscalls: https://outflux.net/blog/archives/2013/12/10/live-patching-the-kernel/ -Kees -- Kees Cook Chrome OS & Brillo Security
[toc] | [prev] | [next] | [standalone]
| From | Serge Hallyn <serge.hallyn@ubuntu.com> |
|---|---|
| Date | 2016-01-26 08:40 +0100 |
| Subject | Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled |
| Message-ID | <qV6oi-KI-7@gated-at.bofh.it> |
| In reply to | #1315977 |
Quoting Kees Cook (keescook@chromium.org): > On Fri, Jan 22, 2016 at 7:02 PM, Eric W. Biederman > > So I have concerns about both efficacy and usability with the proposed > > sysctl. > > Two distros already have this sysctl because it was so strongly > requested by their users. This needs to be upstream so we can manage > the effects correctly. Which two distros? Was it in fact requested by their users? My opinion remains that long-term this is a bad thing. If we're going to have this upstream, it should be clearly marked so as to be easily removable at some point down the road. Userspace that cannot count on a feature (in the best case) won't use it or (much worse) will fall back to broken behavior in one case or the other.
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-01-24 23:30 +0100 |
| Message-ID | <qUBku-2UA-13@gated-at.bofh.it> |
| In reply to | #1315491 |
On Fri, Jan 22, 2016 at 7:02 PM, Eric W. Biederman <ebiederm@xmission.com> wrote: > Kees Cook <keescook@chromium.org> writes: > >> There continues to be unexpected side-effects and security exposures >> via CLONE_NEWUSER. For many end-users running distro kernels with >> CONFIG_USER_NS enabled, there is no way to disable this feature when >> desired. As such, this creates a sysctl to restrict CLONE_NEWUSER so >> admins not running containers or Chrome can avoid the risks of this >> feature. > > I don't actually think there do continue to be unexpected side-effects > and security exposures with CLONE_NEWUSER. It takes a while for all of > the fixes to trickle out to distros. At most what I have seen recently > are problems with other kernel interfaces being amplified with user > namespaces. AKA the current mess with devpts, and the unexpected > issues with bind mounts in mount namespaces. > > > So to keep this productive. Please tell me about the threat model > you envision, and how you envision knobs in the kernel being used to > counter those threats. I consider the ability to use CLONE_NEWUSER to acquire CAP_NET_ADMIN over /any/ network namespace and to thus access the network configuration API to be a huge risk. For example, unprivileged users can program iptables. I'll eat my hat if there are no privilege escalations in there. (They can't request module loading, but still.) --Andy
[toc] | [prev] | [next] | [standalone]
| From | Kees Cook <keescook@chromium.org> |
|---|---|
| Date | 2016-01-25 20:00 +0100 |
| Message-ID | <qUUwO-8gR-19@gated-at.bofh.it> |
| In reply to | #1316061 |
On Mon, Jan 25, 2016 at 10:53 AM, Andy Lutomirski <luto@amacapital.net> wrote: > On Mon, Jan 25, 2016 at 10:51 AM, Kees Cook <keescook@chromium.org> wrote: >> On Sun, Jan 24, 2016 at 2:22 PM, Andy Lutomirski <luto@amacapital.net> wrote: >>> On Fri, Jan 22, 2016 at 7:02 PM, Eric W. Biederman >>> <ebiederm@xmission.com> wrote: >>>> Kees Cook <keescook@chromium.org> writes: >>>> >>>>> There continues to be unexpected side-effects and security exposures >>>>> via CLONE_NEWUSER. For many end-users running distro kernels with >>>>> CONFIG_USER_NS enabled, there is no way to disable this feature when >>>>> desired. As such, this creates a sysctl to restrict CLONE_NEWUSER so >>>>> admins not running containers or Chrome can avoid the risks of this >>>>> feature. >>>> >>>> I don't actually think there do continue to be unexpected side-effects >>>> and security exposures with CLONE_NEWUSER. It takes a while for all of >>>> the fixes to trickle out to distros. At most what I have seen recently >>>> are problems with other kernel interfaces being amplified with user >>>> namespaces. AKA the current mess with devpts, and the unexpected >>>> issues with bind mounts in mount namespaces. >>>> >>> >>>> >>>> So to keep this productive. Please tell me about the threat model >>>> you envision, and how you envision knobs in the kernel being used to >>>> counter those threats. >>> >>> I consider the ability to use CLONE_NEWUSER to acquire CAP_NET_ADMIN >>> over /any/ network namespace and to thus access the network >>> configuration API to be a huge risk. For example, unprivileged users >>> can program iptables. I'll eat my hat if there are no privilege >>> escalations in there. (They can't request module loading, but still.) >> >> Should I consider this an Ack for the patch? :) > > Only if you explain why you need the CAP_SYS_ADMIN check. :) Hm? In the sysctl write? Because otherwise a non-cap root user could turn "1" to "0". The restriction on CLONE_NEWUSER checks caps, not uid, so the uid must be protected by cap checks. The DAC permissions on sysctls for cap-based restrictions make no sense -- they need to be doing cap checks not DAC checks. It's the same logic for why dmesg_restrict and kptr_restrict use the same cap check. > IOW, I think you could change that one line of code and have a less > weird version of the patch that would work just fine. Well, I don't know about less weird, but it would leave a unneeded hole in the permission checks. -Kees -- Kees Cook Chrome OS & Brillo Security
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2016-01-25 21:00 +0100 |
| Message-ID | <qUVsT-t7-19@gated-at.bofh.it> |
| In reply to | #1317202 |
Kees Cook <keescook@chromium.org> writes: > > Well, I don't know about less weird, but it would leave a unneeded > hole in the permission checks. To be clear the current patch has my: Nacked-by: "Eric W. Biederman" <ebiederm@xmission.com> The code is buggy, and poorly thought through. Your lack of interest in fixing the bugs in your patch is distressing. So broken code, not willing to fix. No. We are not merging this sysctl. Eric
[toc] | [prev] | [next] | [standalone]
| From | Kees Cook <keescook@chromium.org> |
|---|---|
| Date | 2016-01-25 23:40 +0100 |
| Message-ID | <qUXXI-2jh-13@gated-at.bofh.it> |
| In reply to | #1317263 |
On Mon, Jan 25, 2016 at 11:33 AM, Eric W. Biederman <ebiederm@xmission.com> wrote: > Kees Cook <keescook@chromium.org> writes: >> >> Well, I don't know about less weird, but it would leave a unneeded >> hole in the permission checks. > > To be clear the current patch has my: > > Nacked-by: "Eric W. Biederman" <ebiederm@xmission.com> > > The code is buggy, and poorly thought through. Your lack of interest in > fixing the bugs in your patch is distressing. I'm not sure where you see me having a "lack of interest". The existing cap-checking sysctls have a corner-case bug, which is orthogonal to this change. > So broken code, not willing to fix. No. We are not merging this sysctl. I think you're jumping to conclusions. :) This feature is already implemented by two distros, and likely wanted by others. We cannot ignore that. The sysctl default doesn't change the existing behavior, so this doesn't get in your way at all. Can you please respond to my earlier email where I rebutted each of your arguments against it? Just saying "no" and putting words in my mouth isn't very productive. Andy, given your interest in this feature, and my explanation of the CAP_SYSADMIN check, what are your thoughts? -Kees -- Kees Cook Chrome OS & Brillo Security
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-01-26 00:40 +0100 |
| Message-ID | <qUYTM-2XX-11@gated-at.bofh.it> |
| In reply to | #1317377 |
On Mon, Jan 25, 2016 at 2:34 PM, Kees Cook <keescook@chromium.org> wrote: > On Mon, Jan 25, 2016 at 11:33 AM, Eric W. Biederman > <ebiederm@xmission.com> wrote: >> Kees Cook <keescook@chromium.org> writes: >>> >>> Well, I don't know about less weird, but it would leave a unneeded >>> hole in the permission checks. >> >> To be clear the current patch has my: >> >> Nacked-by: "Eric W. Biederman" <ebiederm@xmission.com> >> >> The code is buggy, and poorly thought through. Your lack of interest in >> fixing the bugs in your patch is distressing. > > I'm not sure where you see me having a "lack of interest". The > existing cap-checking sysctls have a corner-case bug, which is > orthogonal to this change. > >> So broken code, not willing to fix. No. We are not merging this sysctl. > > I think you're jumping to conclusions. :) > > This feature is already implemented by two distros, and likely wanted > by others. We cannot ignore that. The sysctl default doesn't change > the existing behavior, so this doesn't get in your way at all. Can you > please respond to my earlier email where I rebutted each of your > arguments against it? Just saying "no" and putting words in my mouth > isn't very productive. > > Andy, given your interest in this feature, and my explanation of the > CAP_SYSADMIN check, what are your thoughts? > I think the sysctl sucks, but I don't have a better idea, so I think it should go in. There's clearly demand. A better change would be to have an option to tighten up what namespaces can be manipulated in which manner. If anyone wants to step up and do that, it sounds useful. Denying CAP_NET_ADMIN in an unprivileged user ns would address one piece of attack surface. There are probably others. *However*, I think that trying to protect against a hypothetical attacker with uid == global root who has procfs access but doesn't have CAP_SYS_ADMIN isn't important enough to be worth introducing yet another bad capable() call. Whoever wants to tilt at the windmill of fixng sysctl permissions can address all of them, and then maybe this sysctl would be worth tightening up. --Andy
[toc] | [prev] | [next] | [standalone]
| From | Daniel Micay <danielmicay@gmail.com> |
|---|---|
| Date | 2016-01-26 03:30 +0100 |
| Subject | Re: [kernel-hardening] Re: [PATCH 0/2] sysctl: allow CLONE_NEWUSER to be disabled |
| Message-ID | <qV1yi-4Nn-9@gated-at.bofh.it> |
| In reply to | #1317377 |
[Multipart message — attachments visible in raw view] — view raw
> This feature is already implemented by two distros, and likely wanted > by others. We cannot ignore that. Date point: Arch Linux won't be enabling CONFIG_USERNS until there's a way to disable unprivileged user namespaces. The kernel maintainers are unwilling to carry long-term out-of-tree patches. https://github.com/sandstorm-io/sandstorm/blob/d270755b1b55e5be6c96df2cce7c914f35f0d2a2/install.sh#L464-L474
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web