Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1438576 > unrolled thread
| Started by | Dave Hansen <dave@sr71.net> |
|---|---|
| First post | 2016-07-07 14:50 +0200 |
| Last post | 2016-07-08 12:20 +0200 |
| Articles | 4 — 2 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
[PATCH 2/9] mm: implement new pkey_mprotect() system call Dave Hansen <dave@sr71.net> - 2016-07-07 14:50 +0200
Re: [PATCH 2/9] mm: implement new pkey_mprotect() system call Mel Gorman <mgorman@techsingularity.net> - 2016-07-07 16:50 +0200
Re: [PATCH 2/9] mm: implement new pkey_mprotect() system call Dave Hansen <dave@sr71.net> - 2016-07-07 19:00 +0200
Re: [PATCH 2/9] mm: implement new pkey_mprotect() system call Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 12:20 +0200
| From | Dave Hansen <dave@sr71.net> |
|---|---|
| Date | 2016-07-07 14:50 +0200 |
| Subject | [PATCH 2/9] mm: implement new pkey_mprotect() system call |
| Message-ID | <rSgUF-6Xx-27@gated-at.bofh.it> |
From: Dave Hansen <dave.hansen@linux.intel.com>
pkey_mprotect() is just like mprotect, except it also takes a
protection key as an argument. On systems that do not support
protection keys, it still works, but requires that key=0.
Otherwise it does exactly what mprotect does.
I expect it to get used like this, if you want to guarantee that
any mapping you create can *never* be accessed without the right
protection keys set up.
int real_prot = PROT_READ|PROT_WRITE;
pkey = pkey_alloc(0, PKEY_DENY_ACCESS);
ptr = mmap(NULL, PAGE_SIZE, PROT_NONE, MAP_ANONYMOUS|MAP_PRIVATE, -1, 0);
ret = pkey_mprotect(ptr, PAGE_SIZE, real_prot, pkey);
This way, there is *no* window where the mapping is accessible
since it was always either PROT_NONE or had a protection key set
that denied all access.
We settled on 'unsigned long' for the type of the key here. We
only need 4 bits on x86 today, but I figured that other
architectures might need some more space.
Semantically, we have a bit of a problem if we combine this
syscall with our previously-introduced execute-only support:
What do we do when we mix execute-only pkey use with
pkey_mprotect() use? For instance:
pkey_mprotect(ptr, PAGE_SIZE, PROT_WRITE, 6); // set pkey=6
mprotect(ptr, PAGE_SIZE, PROT_EXEC); // set pkey=X_ONLY_PKEY?
mprotect(ptr, PAGE_SIZE, PROT_WRITE); // is pkey=6 again?
To solve that, we make the plain-mprotect()-initiated execute-only
support only apply to VMAs that have the default protection key (0)
set on them.
Proposed semantics:
1. protection key 0 is special and represents the default,
"unassigned" protection key. It is always allocated.
2. mprotect() never affects a mapping's pkey_mprotect()-assigned
protection key. A protection key of 0 (even if set explicitly)
represents an unassigned protection key.
2a. mprotect(PROT_EXEC) on a mapping with an assigned protection
key may or may not result in a mapping with execute-only
properties. pkey_mprotect() plus pkey_set() on all threads
should be used to _guarantee_ execute-only semantics if this
is not a strong enough semantic.
3. mprotect(PROT_EXEC) may result in an "execute-only" mapping. The
kernel will internally attempt to allocate and dedicate a
protection key for the purpose of execute-only mappings. This
may not be possible in cases where there are no free protection
keys available. It can also happen, of course, in situations
where there is no hardware support for protection keys.
Signed-off-by: Dave Hansen <dave.hansen@linux.intel.com>
Cc: linux-api@vger.kernel.org
Cc: linux-arch@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: x86@kernel.org
Cc: torvalds@linux-foundation.org
Cc: akpm@linux-foundation.org
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: mgorman@techsingularity.net
Cc: hughd@google.com
Cc: viro@zeniv.linux.org.uk
---
b/arch/x86/include/asm/mmu_context.h | 15 ++++++++++-----
b/arch/x86/include/asm/pkeys.h | 11 +++++++++--
b/arch/x86/kernel/fpu/xstate.c | 15 ++++++++++++++-
b/arch/x86/mm/pkeys.c | 2 +-
b/mm/mprotect.c | 27 +++++++++++++++++++++++----
5 files changed, 57 insertions(+), 13 deletions(-)
diff -puN arch/x86/include/asm/mmu_context.h~pkeys-110-syscalls-mprotect_pkey arch/x86/include/asm/mmu_context.h
--- a/arch/x86/include/asm/mmu_context.h~pkeys-110-syscalls-mprotect_pkey 2016-07-07 05:46:59.974764757 -0700
+++ b/arch/x86/include/asm/mmu_context.h 2016-07-07 05:46:59.986765301 -0700
@@ -4,6 +4,7 @@
#include <asm/desc.h>
#include <linux/atomic.h>
#include <linux/mm_types.h>
+#include <linux/pkeys.h>
#include <trace/events/tlb.h>
@@ -195,16 +196,20 @@ static inline void arch_unmap(struct mm_
mpx_notify_unmap(mm, vma, start, end);
}
+#ifdef CONFIG_X86_INTEL_MEMORY_PROTECTION_KEYS
static inline int vma_pkey(struct vm_area_struct *vma)
{
- u16 pkey = 0;
-#ifdef CONFIG_X86_INTEL_MEMORY_PROTECTION_KEYS
unsigned long vma_pkey_mask = VM_PKEY_BIT0 | VM_PKEY_BIT1 |
VM_PKEY_BIT2 | VM_PKEY_BIT3;
- pkey = (vma->vm_flags & vma_pkey_mask) >> VM_PKEY_SHIFT;
-#endif
- return pkey;
+
+ return (vma->vm_flags & vma_pkey_mask) >> VM_PKEY_SHIFT;
+}
+#else
+static inline int vma_pkey(struct vm_area_struct *vma)
+{
+ return 0;
}
+#endif
static inline bool __pkru_allows_pkey(u16 pkey, bool write)
{
diff -puN arch/x86/include/asm/pkeys.h~pkeys-110-syscalls-mprotect_pkey arch/x86/include/asm/pkeys.h
--- a/arch/x86/include/asm/pkeys.h~pkeys-110-syscalls-mprotect_pkey 2016-07-07 05:46:59.976764847 -0700
+++ b/arch/x86/include/asm/pkeys.h 2016-07-07 05:46:59.986765301 -0700
@@ -1,7 +1,12 @@
#ifndef _ASM_X86_PKEYS_H
#define _ASM_X86_PKEYS_H
-#define arch_max_pkey() (boot_cpu_has(X86_FEATURE_OSPKE) ? 16 : 1)
+#define PKEY_DEDICATED_EXECUTE_ONLY 15
+/*
+ * Consider the PKEY_DEDICATED_EXECUTE_ONLY key unavailable.
+ */
+#define arch_max_pkey() (boot_cpu_has(X86_FEATURE_OSPKE) ? \
+ PKEY_DEDICATED_EXECUTE_ONLY : 1)
extern int arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
unsigned long init_val);
@@ -10,7 +15,6 @@ extern int arch_set_user_pkey_access(str
* Try to dedicate one of the protection keys to be used as an
* execute-only protection key.
*/
-#define PKEY_DEDICATED_EXECUTE_ONLY 15
extern int __execute_only_pkey(struct mm_struct *mm);
static inline int execute_only_pkey(struct mm_struct *mm)
{
@@ -31,4 +35,7 @@ static inline int arch_override_mprotect
return __arch_override_mprotect_pkey(vma, prot, pkey);
}
+extern int __arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
+ unsigned long init_val);
+
#endif /*_ASM_X86_PKEYS_H */
diff -puN arch/x86/kernel/fpu/xstate.c~pkeys-110-syscalls-mprotect_pkey arch/x86/kernel/fpu/xstate.c
--- a/arch/x86/kernel/fpu/xstate.c~pkeys-110-syscalls-mprotect_pkey 2016-07-07 05:46:59.977764893 -0700
+++ b/arch/x86/kernel/fpu/xstate.c 2016-07-07 05:46:59.987765346 -0700
@@ -889,7 +889,7 @@ out:
* not modfiy PKRU *itself* here, only the XSAVE state that will
* be restored in to PKRU when we return back to userspace.
*/
-int arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
+int __arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
unsigned long init_val)
{
struct xregs_state *xsave = &tsk->thread.fpu.state.xsave;
@@ -948,3 +948,16 @@ int arch_set_user_pkey_access(struct tas
return 0;
}
+
+/*
+ * When setting a userspace-provided value, we need to ensure
+ * that it is valid. The __ version can get used by
+ * kernel-internal uses like the execute-only support.
+ */
+int arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
+ unsigned long init_val)
+{
+ if (!validate_pkey(pkey))
+ return -EINVAL;
+ return __arch_set_user_pkey_access(tsk, pkey, init_val);
+}
diff -puN arch/x86/mm/pkeys.c~pkeys-110-syscalls-mprotect_pkey arch/x86/mm/pkeys.c
--- a/arch/x86/mm/pkeys.c~pkeys-110-syscalls-mprotect_pkey 2016-07-07 05:46:59.980765029 -0700
+++ b/arch/x86/mm/pkeys.c 2016-07-07 05:46:59.987765346 -0700
@@ -38,7 +38,7 @@ int __execute_only_pkey(struct mm_struct
return PKEY_DEDICATED_EXECUTE_ONLY;
}
preempt_enable();
- ret = arch_set_user_pkey_access(current, PKEY_DEDICATED_EXECUTE_ONLY,
+ ret = __arch_set_user_pkey_access(current, PKEY_DEDICATED_EXECUTE_ONLY,
PKEY_DISABLE_ACCESS);
/*
* If the PKRU-set operation failed somehow, just return
diff -puN mm/mprotect.c~pkeys-110-syscalls-mprotect_pkey mm/mprotect.c
--- a/mm/mprotect.c~pkeys-110-syscalls-mprotect_pkey 2016-07-07 05:46:59.982765119 -0700
+++ b/mm/mprotect.c 2016-07-07 05:46:59.987765346 -0700
@@ -352,8 +352,11 @@ fail:
return error;
}
-SYSCALL_DEFINE3(mprotect, unsigned long, start, size_t, len,
- unsigned long, prot)
+/*
+ * pkey==-1 when doing a legacy mprotect()
+ */
+static int do_mprotect_pkey(unsigned long start, size_t len,
+ unsigned long prot, int pkey)
{
unsigned long nstart, end, tmp, reqprot;
struct vm_area_struct *vma, *prev;
@@ -409,7 +412,7 @@ SYSCALL_DEFINE3(mprotect, unsigned long,
for (nstart = start ; ; ) {
unsigned long newflags;
- int pkey = arch_override_mprotect_pkey(vma, prot, -1);
+ int new_vma_pkey;
/* Here we know that vma->vm_start <= nstart < vma->vm_end. */
@@ -417,7 +420,8 @@ SYSCALL_DEFINE3(mprotect, unsigned long,
if (rier && (vma->vm_flags & VM_MAYEXEC))
prot |= PROT_EXEC;
- newflags = calc_vm_prot_bits(prot, pkey);
+ new_vma_pkey = arch_override_mprotect_pkey(vma, prot, pkey);
+ newflags = calc_vm_prot_bits(prot, new_vma_pkey);
newflags |= (vma->vm_flags & ~(VM_READ | VM_WRITE | VM_EXEC));
/* newflags >> 4 shift VM_MAY% in place of VM_% */
@@ -454,3 +458,18 @@ out:
up_write(¤t->mm->mmap_sem);
return error;
}
+
+SYSCALL_DEFINE3(mprotect, unsigned long, start, size_t, len,
+ unsigned long, prot)
+{
+ return do_mprotect_pkey(start, len, prot, -1);
+}
+
+SYSCALL_DEFINE4(pkey_mprotect, unsigned long, start, size_t, len,
+ unsigned long, prot, int, pkey)
+{
+ if (!validate_pkey(pkey))
+ return -EINVAL;
+
+ return do_mprotect_pkey(start, len, prot, pkey);
+}
_
[toc] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2016-07-07 16:50 +0200 |
| Message-ID | <rSiMN-8hI-23@gated-at.bofh.it> |
| In reply to | #1438576 |
On Thu, Jul 07, 2016 at 05:47:22AM -0700, Dave Hansen wrote:
> diff -puN arch/x86/include/asm/mmu_context.h~pkeys-110-syscalls-mprotect_pkey arch/x86/include/asm/mmu_context.h
> --- a/arch/x86/include/asm/mmu_context.h~pkeys-110-syscalls-mprotect_pkey 2016-07-07 05:46:59.974764757 -0700
> +++ b/arch/x86/include/asm/mmu_context.h 2016-07-07 05:46:59.986765301 -0700
> @@ -4,6 +4,7 @@
> #include <asm/desc.h>
> #include <linux/atomic.h>
> #include <linux/mm_types.h>
> +#include <linux/pkeys.h>
>
> #include <trace/events/tlb.h>
>
> @@ -195,16 +196,20 @@ static inline void arch_unmap(struct mm_
> mpx_notify_unmap(mm, vma, start, end);
> }
>
> +#ifdef CONFIG_X86_INTEL_MEMORY_PROTECTION_KEYS
> static inline int vma_pkey(struct vm_area_struct *vma)
> {
> - u16 pkey = 0;
> -#ifdef CONFIG_X86_INTEL_MEMORY_PROTECTION_KEYS
> unsigned long vma_pkey_mask = VM_PKEY_BIT0 | VM_PKEY_BIT1 |
> VM_PKEY_BIT2 | VM_PKEY_BIT3;
> - pkey = (vma->vm_flags & vma_pkey_mask) >> VM_PKEY_SHIFT;
> -#endif
> - return pkey;
> +
> + return (vma->vm_flags & vma_pkey_mask) >> VM_PKEY_SHIFT;
> +}
> +#else
> +static inline int vma_pkey(struct vm_area_struct *vma)
> +{
> + return 0;
> }
> +#endif
>
> static inline bool __pkru_allows_pkey(u16 pkey, bool write)
> {
Looks like MASK could have been statically defined and be a simple shift
and mask known at compile time. Minor though.
> diff -puN arch/x86/include/asm/pkeys.h~pkeys-110-syscalls-mprotect_pkey arch/x86/include/asm/pkeys.h
> --- a/arch/x86/include/asm/pkeys.h~pkeys-110-syscalls-mprotect_pkey 2016-07-07 05:46:59.976764847 -0700
> +++ b/arch/x86/include/asm/pkeys.h 2016-07-07 05:46:59.986765301 -0700
> @@ -1,7 +1,12 @@
> #ifndef _ASM_X86_PKEYS_H
> #define _ASM_X86_PKEYS_H
>
> -#define arch_max_pkey() (boot_cpu_has(X86_FEATURE_OSPKE) ? 16 : 1)
> +#define PKEY_DEDICATED_EXECUTE_ONLY 15
> +/*
> + * Consider the PKEY_DEDICATED_EXECUTE_ONLY key unavailable.
> + */
> +#define arch_max_pkey() (boot_cpu_has(X86_FEATURE_OSPKE) ? \
> + PKEY_DEDICATED_EXECUTE_ONLY : 1)
>
> extern int arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
> unsigned long init_val);
> @@ -10,7 +15,6 @@ extern int arch_set_user_pkey_access(str
> * Try to dedicate one of the protection keys to be used as an
> * execute-only protection key.
> */
> -#define PKEY_DEDICATED_EXECUTE_ONLY 15
> extern int __execute_only_pkey(struct mm_struct *mm);
> static inline int execute_only_pkey(struct mm_struct *mm)
> {
> @@ -31,4 +35,7 @@ static inline int arch_override_mprotect
> return __arch_override_mprotect_pkey(vma, prot, pkey);
> }
>
> +extern int __arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
> + unsigned long init_val);
> +
> #endif /*_ASM_X86_PKEYS_H */
> diff -puN arch/x86/kernel/fpu/xstate.c~pkeys-110-syscalls-mprotect_pkey arch/x86/kernel/fpu/xstate.c
> --- a/arch/x86/kernel/fpu/xstate.c~pkeys-110-syscalls-mprotect_pkey 2016-07-07 05:46:59.977764893 -0700
> +++ b/arch/x86/kernel/fpu/xstate.c 2016-07-07 05:46:59.987765346 -0700
> @@ -889,7 +889,7 @@ out:
> * not modfiy PKRU *itself* here, only the XSAVE state that will
> * be restored in to PKRU when we return back to userspace.
> */
> -int arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
> +int __arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
> unsigned long init_val)
> {
In the changelog, you state that the key is unsigned long yet here and
in the documentation you linked, it's int. Minimally, it's surprising
that the key is signed.
> struct xregs_state *xsave = &tsk->thread.fpu.state.xsave;
> @@ -948,3 +948,16 @@ int arch_set_user_pkey_access(struct tas
>
> return 0;
> }
> +
> +/*
> + * When setting a userspace-provided value, we need to ensure
> + * that it is valid. The __ version can get used by
> + * kernel-internal uses like the execute-only support.
> + */
> +int arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
> + unsigned long init_val)
> +{
> + if (!validate_pkey(pkey))
> + return -EINVAL;
> + return __arch_set_user_pkey_access(tsk, pkey, init_val);
> +}
There appears to be a subtle bug fixed for validate_key. It appears
there wasn't protection of the dedicated key before but nothing could
reach it.
The arch_max_pkey and PKEY_DEDICATE_EXECUTE_ONLY interaction is subtle
but I can't find a problem with it either.
That aside, the validate_pkey check looks weak. It might be a number
that works but no guarantee it's an allocated key or initialised
properly. At this point, garbage can be handed into the system call
potentially but maybe that gets fixed later.
> diff -puN arch/x86/mm/pkeys.c~pkeys-110-syscalls-mprotect_pkey arch/x86/mm/pkeys.c
> --- a/arch/x86/mm/pkeys.c~pkeys-110-syscalls-mprotect_pkey 2016-07-07 05:46:59.980765029 -0700
> +++ b/arch/x86/mm/pkeys.c 2016-07-07 05:46:59.987765346 -0700
> @@ -38,7 +38,7 @@ int __execute_only_pkey(struct mm_struct
> return PKEY_DEDICATED_EXECUTE_ONLY;
> }
> preempt_enable();
> - ret = arch_set_user_pkey_access(current, PKEY_DEDICATED_EXECUTE_ONLY,
> + ret = __arch_set_user_pkey_access(current, PKEY_DEDICATED_EXECUTE_ONLY,
> PKEY_DISABLE_ACCESS);
> /*
> * If the PKRU-set operation failed somehow, just return
> diff -puN mm/mprotect.c~pkeys-110-syscalls-mprotect_pkey mm/mprotect.c
> --- a/mm/mprotect.c~pkeys-110-syscalls-mprotect_pkey 2016-07-07 05:46:59.982765119 -0700
> +++ b/mm/mprotect.c 2016-07-07 05:46:59.987765346 -0700
> @@ -352,8 +352,11 @@ fail:
> return error;
> }
>
> -SYSCALL_DEFINE3(mprotect, unsigned long, start, size_t, len,
> - unsigned long, prot)
> +/*
> + * pkey==-1 when doing a legacy mprotect()
> + */
> +static int do_mprotect_pkey(unsigned long start, size_t len,
> + unsigned long prot, int pkey)
> {
> unsigned long nstart, end, tmp, reqprot;
> struct vm_area_struct *vma, *prev;
> @@ -409,7 +412,7 @@ SYSCALL_DEFINE3(mprotect, unsigned long,
>
> for (nstart = start ; ; ) {
> unsigned long newflags;
> - int pkey = arch_override_mprotect_pkey(vma, prot, -1);
> + int new_vma_pkey;
>
> /* Here we know that vma->vm_start <= nstart < vma->vm_end. */
>
> @@ -417,7 +420,8 @@ SYSCALL_DEFINE3(mprotect, unsigned long,
> if (rier && (vma->vm_flags & VM_MAYEXEC))
> prot |= PROT_EXEC;
>
> - newflags = calc_vm_prot_bits(prot, pkey);
> + new_vma_pkey = arch_override_mprotect_pkey(vma, prot, pkey);
> + newflags = calc_vm_prot_bits(prot, new_vma_pkey);
> newflags |= (vma->vm_flags & ~(VM_READ | VM_WRITE | VM_EXEC));
>
On CPUs that do not support the feature, arch_override_mprotect_pkey
returns 0 and the normal protections are used. It's not clear how an
application is meant to detect if the operation succeeded or not. What
if the application relies on pkeys to be working?
Should pkey_mprotect return ENOSYS if the CPU does not support the
requested feature or is that handled somewhere else?
> /* newflags >> 4 shift VM_MAY% in place of VM_% */
> @@ -454,3 +458,18 @@ out:
> up_write(¤t->mm->mmap_sem);
> return error;
> }
> +
> +SYSCALL_DEFINE3(mprotect, unsigned long, start, size_t, len,
> + unsigned long, prot)
> +{
> + return do_mprotect_pkey(start, len, prot, -1);
> +}
> +
> +SYSCALL_DEFINE4(pkey_mprotect, unsigned long, start, size_t, len,
> + unsigned long, prot, int, pkey)
> +{
> + if (!validate_pkey(pkey))
> + return -EINVAL;
> +
> + return do_mprotect_pkey(start, len, prot, pkey);
> +}
> _
--
Mel Gorman
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Dave Hansen <dave@sr71.net> |
|---|---|
| Date | 2016-07-07 19:00 +0200 |
| Message-ID | <rSkOC-16v-17@gated-at.bofh.it> |
| In reply to | #1438658 |
On 07/07/2016 07:40 AM, Mel Gorman wrote:
> On Thu, Jul 07, 2016 at 05:47:22AM -0700, Dave Hansen wrote:
>> +#ifdef CONFIG_X86_INTEL_MEMORY_PROTECTION_KEYS
>> static inline int vma_pkey(struct vm_area_struct *vma)
>> {
>> - u16 pkey = 0;
>> -#ifdef CONFIG_X86_INTEL_MEMORY_PROTECTION_KEYS
>> unsigned long vma_pkey_mask = VM_PKEY_BIT0 | VM_PKEY_BIT1 |
>> VM_PKEY_BIT2 | VM_PKEY_BIT3;
>> - pkey = (vma->vm_flags & vma_pkey_mask) >> VM_PKEY_SHIFT;
>> -#endif
>> - return pkey;
>> +
>> + return (vma->vm_flags & vma_pkey_mask) >> VM_PKEY_SHIFT;
>> +}
>> +#else
>> +static inline int vma_pkey(struct vm_area_struct *vma)
>> +{
>> + return 0;
>> }
>> +#endif
>>
>> static inline bool __pkru_allows_pkey(u16 pkey, bool write)
>> {
>
> Looks like MASK could have been statically defined and be a simple shift
> and mask known at compile time. Minor though.
The VM_PKEY_BIT*'s are only ever defined as masks and not bit numbers.
So, if you want to use a mask, you end up doing something like:
unsigned long mask = (NR_PKEYS-1) << ffz(~VM_PKEY_BIT0);
Which ends up with the same thing, but I think ends up being pretty on
par for ugliness.
...
>> +/*
>> + * When setting a userspace-provided value, we need to ensure
>> + * that it is valid. The __ version can get used by
>> + * kernel-internal uses like the execute-only support.
>> + */
>> +int arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
>> + unsigned long init_val)
>> +{
>> + if (!validate_pkey(pkey))
>> + return -EINVAL;
>> + return __arch_set_user_pkey_access(tsk, pkey, init_val);
>> +}
>
> There appears to be a subtle bug fixed for validate_key. It appears
> there wasn't protection of the dedicated key before but nothing could
> reach it.
Right. There was no user interface that took a key and we trusted that
the kernel knew what it was doing.
> The arch_max_pkey and PKEY_DEDICATE_EXECUTE_ONLY interaction is subtle
> but I can't find a problem with it either.
>
> That aside, the validate_pkey check looks weak. It might be a number
> that works but no guarantee it's an allocated key or initialised
> properly. At this point, garbage can be handed into the system call
> potentially but maybe that gets fixed later.
It's called in three paths:
1. by the kernel when setting up execute-only support
2. by pkey_alloc() on the pkey we just allocated
3. by pkey_set() on a pkey we just checked was allocated
So, it isn't broken, but it's also not clear at all why it is safe and
what validate_pkey() is actually validating.
But, that said, this does make me realize that with
pkey_alloc()/pkey_free(), this is probably redundant. We verify that
the key is allocated, and we only allow valid keys to be allocated.
IOW, I think I can remove validate_pkey(), but only if we keep pkey_alloc().
...
>> - newflags = calc_vm_prot_bits(prot, pkey);
>> + new_vma_pkey = arch_override_mprotect_pkey(vma, prot, pkey);
>> + newflags = calc_vm_prot_bits(prot, new_vma_pkey);
>> newflags |= (vma->vm_flags & ~(VM_READ | VM_WRITE | VM_EXEC));
>>
>
> On CPUs that do not support the feature, arch_override_mprotect_pkey
> returns 0 and the normal protections are used. It's not clear how an
> application is meant to detect if the operation succeeded or not. What
> if the application relies on pkeys to be working?
It actually shows up as -ENOSPC from pkey_alloc(). This sounds goofy,
but it teaches programs something very important: they always have to
look for ENOSPC, and must always be prepared to function without
protection keys. A library might have stolen all the keys, or an
LD_PRELOAD, so an app can never be sure what is available.
If we teach them to check for ENOSPC from day one, they'll never be
surprised.
I've tried to spell this out a bit more clearly in the manpages. I'll
also add it to the changelog.
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2016-07-08 12:20 +0200 |
| Message-ID | <rSB34-3r2-29@gated-at.bofh.it> |
| In reply to | #1438728 |
On Thu, Jul 07, 2016 at 09:51:52AM -0700, Dave Hansen wrote:
> > Looks like MASK could have been statically defined and be a simple shift
> > and mask known at compile time. Minor though.
>
> The VM_PKEY_BIT*'s are only ever defined as masks and not bit numbers.
> So, if you want to use a mask, you end up doing something like:
>
> unsigned long mask = (NR_PKEYS-1) << ffz(~VM_PKEY_BIT0);
>
> Which ends up with the same thing, but I think ends up being pretty on
> par for ugliness.
>
Fair enough.
> >> +/*
> >> + * When setting a userspace-provided value, we need to ensure
> >> + * that it is valid. The __ version can get used by
> >> + * kernel-internal uses like the execute-only support.
> >> + */
> >> +int arch_set_user_pkey_access(struct task_struct *tsk, int pkey,
> >> + unsigned long init_val)
> >> +{
> >> + if (!validate_pkey(pkey))
> >> + return -EINVAL;
> >> + return __arch_set_user_pkey_access(tsk, pkey, init_val);
> >> +}
> >
> > There appears to be a subtle bug fixed for validate_key. It appears
> > there wasn't protection of the dedicated key before but nothing could
> > reach it.
>
> Right. There was no user interface that took a key and we trusted that
> the kernel knew what it was doing.
>
Ok. I was fairly sure that was the thinking behind it but wanted to be suire.
> > The arch_max_pkey and PKEY_DEDICATE_EXECUTE_ONLY interaction is subtle
> > but I can't find a problem with it either.
> >
> > That aside, the validate_pkey check looks weak. It might be a number
> > that works but no guarantee it's an allocated key or initialised
> > properly. At this point, garbage can be handed into the system call
> > potentially but maybe that gets fixed later.
>
> It's called in three paths:
> 1. by the kernel when setting up execute-only support
> 2. by pkey_alloc() on the pkey we just allocated
> 3. by pkey_set() on a pkey we just checked was allocated
>
> So, it isn't broken, but it's also not clear at all why it is safe and
> what validate_pkey() is actually validating.
>
> But, that said, this does make me realize that with
> pkey_alloc()/pkey_free(), this is probably redundant. We verify that
> the key is allocated, and we only allow valid keys to be allocated.
>
> IOW, I think I can remove validate_pkey(), but only if we keep pkey_alloc().
>
Ok, it's not a major problem. I simply worried that the protection of
key slots is pretty weak as it can be interfered with from userspace.
On the other hand, the kernel never interprets the information so it's
unlikely to cause a security problem. Applications can still shoot
themselves in the foot but hopefully the developers are aware that the
protection they get with keys is not absolute.
> ...
> >> - newflags = calc_vm_prot_bits(prot, pkey);
> >> + new_vma_pkey = arch_override_mprotect_pkey(vma, prot, pkey);
> >> + newflags = calc_vm_prot_bits(prot, new_vma_pkey);
> >> newflags |= (vma->vm_flags & ~(VM_READ | VM_WRITE | VM_EXEC));
> >>
> >
> > On CPUs that do not support the feature, arch_override_mprotect_pkey
> > returns 0 and the normal protections are used. It's not clear how an
> > application is meant to detect if the operation succeeded or not. What
> > if the application relies on pkeys to be working?
>
> It actually shows up as -ENOSPC from pkey_alloc(). This sounds goofy,
> but it teaches programs something very important: they always have to
> look for ENOSPC, and must always be prepared to function without
> protection keys.
Ok, that makes sense. I don't think it's goofy. Sure, they cannot detect
the CPU support directly from the interface but it's close enough.
--
Mel Gorman
SUSE Labs
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web