Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1289855 > unrolled thread

[PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

Started byTony Luck <tony.luck@intel.com>
First post2015-12-11 20:40 +0100
Last post2015-12-15 20:30 +0100
Articles 20 on this page of 25 — 7 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks Tony Luck <tony.luck@intel.com> - 2015-12-11 20:40 +0100
    Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Andy Lutomirski <luto@amacapital.net> - 2015-12-11 21:10 +0100
      RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks "Luck, Tony" <tony.luck@intel.com> - 2015-12-11 22:20 +0100
        Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Andy Lutomirski <luto@amacapital.net> - 2015-12-11 23:00 +0100
          RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks "Luck, Tony" <tony.luck@intel.com> - 2015-12-11 23:20 +0100
            Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Dan Williams <dan.j.williams@intel.com> - 2015-12-11 23:30 +0100
              Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Andy Lutomirski <luto@amacapital.net> - 2015-12-11 23:30 +0100
                Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Andy Lutomirski <luto@amacapital.net> - 2015-12-11 23:40 +0100
                  RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks "Luck, Tony" <tony.luck@intel.com> - 2015-12-11 23:50 +0100
                    Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Andy Lutomirski <luto@amacapital.net> - 2015-12-12 00:00 +0100
                      Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Ingo Molnar <mingo@kernel.org> - 2015-12-14 09:40 +0100
                        Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks "Luck, Tony" <tony.luck@intel.com> - 2015-12-14 20:50 +0100
                          Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Andy Lutomirski <luto@amacapital.net> - 2015-12-14 21:20 +0100
                RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks "Luck, Tony" <tony.luck@intel.com> - 2015-12-11 23:40 +0100
    Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Borislav Petkov <bp@alien8.de> - 2015-12-15 14:20 +0100
      Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Dan Williams <dan.j.williams@intel.com> - 2015-12-15 18:50 +0100
        RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks "Luck, Tony" <tony.luck@intel.com> - 2015-12-15 19:00 +0100
          Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Borislav Petkov <bp@alien8.de> - 2015-12-15 19:30 +0100
          Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Dan Williams <dan.j.williams@intel.com> - 2015-12-15 19:30 +0100
            Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Dan Williams <dan.j.williams@intel.com> - 2015-12-15 19:40 +0100
              Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Borislav Petkov <bp@alien8.de> - 2015-12-15 19:40 +0100
                Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Borislav Petkov <bp@alien8.de> - 2015-12-15 20:30 +0100
                  RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks "Elliott, Robert (Persistent Memory)" <elliott@hpe.com> - 2015-12-15 21:30 +0100
                    Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks Borislav Petkov <bp@alien8.de> - 2015-12-21 18:40 +0100
                RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover  from machine checks "Elliott, Robert (Persistent Memory)" <elliott@hpe.com> - 2015-12-15 20:30 +0100

Page 1 of 2  [1] 2  Next page →


#1289855 — [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromTony Luck <tony.luck@intel.com>
Date2015-12-11 20:40 +0100
Subject[PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qEBHQ-ir-13@gated-at.bofh.it>
Using __copy_user_nocache() as inspiration create a memory copy
routine for use by kernel code with annotations to allow for
recovery from machine checks.

Notes:
1) Unlike the original we make no attempt to copy all the bytes
   up to the faulting address. The original achieves that by
   re-executing the failing part as a byte-by-byte copy,
   which will take another page fault. We don't want to have
   a second machine check!
2) Likewise the return value for the original indicates exactly
   how many bytes were not copied. Instead we provide the physical
   address of the fault (thanks to help from do_machine_check()
3) Provide helpful macros to decode the return value.

Signed-off-by: Tony Luck <tony.luck@intel.com>
---
 arch/x86/include/asm/uaccess_64.h |  5 +++
 arch/x86/kernel/x8664_ksyms_64.c  |  2 +
 arch/x86/lib/copy_user_64.S       | 91 +++++++++++++++++++++++++++++++++++++++
 3 files changed, 98 insertions(+)

diff --git a/arch/x86/include/asm/uaccess_64.h b/arch/x86/include/asm/uaccess_64.h
index f2f9b39b274a..779cb0e77ecc 100644
--- a/arch/x86/include/asm/uaccess_64.h
+++ b/arch/x86/include/asm/uaccess_64.h
@@ -216,6 +216,11 @@ __copy_to_user_inatomic(void __user *dst, const void *src, unsigned size)
 extern long __copy_user_nocache(void *dst, const void __user *src,
 				unsigned size, int zerorest);
 
+extern u64 mcsafe_memcpy(void *dst, const void __user *src,
+				unsigned size);
+#define COPY_HAD_MCHECK(ret)	((ret) & BIT(63))
+#define	COPY_MCHECK_PADDR(ret)	((ret) & ~BIT(63))
+
 static inline int
 __copy_from_user_nocache(void *dst, const void __user *src, unsigned size)
 {
diff --git a/arch/x86/kernel/x8664_ksyms_64.c b/arch/x86/kernel/x8664_ksyms_64.c
index a0695be19864..ec988c92c055 100644
--- a/arch/x86/kernel/x8664_ksyms_64.c
+++ b/arch/x86/kernel/x8664_ksyms_64.c
@@ -37,6 +37,8 @@ EXPORT_SYMBOL(__copy_user_nocache);
 EXPORT_SYMBOL(_copy_from_user);
 EXPORT_SYMBOL(_copy_to_user);
 
+EXPORT_SYMBOL(mcsafe_memcpy);
+
 EXPORT_SYMBOL(copy_page);
 EXPORT_SYMBOL(clear_page);
 
diff --git a/arch/x86/lib/copy_user_64.S b/arch/x86/lib/copy_user_64.S
index 982ce34f4a9b..ffce93cbc9a5 100644
--- a/arch/x86/lib/copy_user_64.S
+++ b/arch/x86/lib/copy_user_64.S
@@ -319,3 +319,94 @@ ENTRY(__copy_user_nocache)
 	_ASM_EXTABLE(21b,50b)
 	_ASM_EXTABLE(22b,50b)
 ENDPROC(__copy_user_nocache)
+
+/*
+ * mcsafe_memcpy - Uncached memory copy with machine check exception handling
+ * Note that we only catch machine checks when reading the source addresses.
+ * Writes to target are posted and don't generate machine checks.
+ * This will force destination/source out of cache for more performance.
+ */
+ENTRY(mcsafe_memcpy)
+	cmpl $8,%edx
+	jb 20f		/* less then 8 bytes, go to byte copy loop */
+
+	/* check for bad alignment of destination */
+	movl %edi,%ecx
+	andl $7,%ecx
+	jz 102f				/* already aligned */
+	subl $8,%ecx
+	negl %ecx
+	subl %ecx,%edx
+0:	movb (%rsi),%al
+	movb %al,(%rdi)
+	incq %rsi
+	incq %rdi
+	decl %ecx
+	jnz 100b
+102:
+	movl %edx,%ecx
+	andl $63,%edx
+	shrl $6,%ecx
+	jz 17f
+1:	movq (%rsi),%r8
+2:	movq 1*8(%rsi),%r9
+3:	movq 2*8(%rsi),%r10
+4:	movq 3*8(%rsi),%r11
+	movnti %r8,(%rdi)
+	movnti %r9,1*8(%rdi)
+	movnti %r10,2*8(%rdi)
+	movnti %r11,3*8(%rdi)
+9:	movq 4*8(%rsi),%r8
+10:	movq 5*8(%rsi),%r9
+11:	movq 6*8(%rsi),%r10
+12:	movq 7*8(%rsi),%r11
+	movnti %r8,4*8(%rdi)
+	movnti %r9,5*8(%rdi)
+	movnti %r10,6*8(%rdi)
+	movnti %r11,7*8(%rdi)
+	leaq 64(%rsi),%rsi
+	leaq 64(%rdi),%rdi
+	decl %ecx
+	jnz 1b
+17:	movl %edx,%ecx
+	andl $7,%edx
+	shrl $3,%ecx
+	jz 20f
+18:	movq (%rsi),%r8
+	movnti %r8,(%rdi)
+	leaq 8(%rsi),%rsi
+	leaq 8(%rdi),%rdi
+	decl %ecx
+	jnz 18b
+20:	andl %edx,%edx
+	jz 23f
+	movl %edx,%ecx
+21:	movb (%rsi),%al
+	movb %al,(%rdi)
+	incq %rsi
+	incq %rdi
+	decl %ecx
+	jnz 21b
+23:	xorl %eax,%eax
+	sfence
+	ret
+
+	.section .fixup,"ax"
+30:
+	sfence
+	/* do_machine_check() sets %eax return value */
+	ret
+	.previous
+
+	_ASM_MCEXTABLE(0b,30b)
+	_ASM_MCEXTABLE(1b,30b)
+	_ASM_MCEXTABLE(2b,30b)
+	_ASM_MCEXTABLE(3b,30b)
+	_ASM_MCEXTABLE(4b,30b)
+	_ASM_MCEXTABLE(9b,30b)
+	_ASM_MCEXTABLE(10b,30b)
+	_ASM_MCEXTABLE(11b,30b)
+	_ASM_MCEXTABLE(12b,30b)
+	_ASM_MCEXTABLE(18b,30b)
+	_ASM_MCEXTABLE(21b,30b)
+ENDPROC(mcsafe_memcpy)
-- 
2.1.4

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1289876 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromAndy Lutomirski <luto@amacapital.net>
Date2015-12-11 21:10 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qECaR-IL-7@gated-at.bofh.it>
In reply to#1289855
On Thu, Dec 10, 2015 at 4:21 PM, Tony Luck <tony.luck@intel.com> wrote:
> Using __copy_user_nocache() as inspiration create a memory copy
> routine for use by kernel code with annotations to allow for
> recovery from machine checks.
>
> Notes:
> 1) Unlike the original we make no attempt to copy all the bytes
>    up to the faulting address. The original achieves that by
>    re-executing the failing part as a byte-by-byte copy,
>    which will take another page fault. We don't want to have
>    a second machine check!
> 2) Likewise the return value for the original indicates exactly
>    how many bytes were not copied. Instead we provide the physical
>    address of the fault (thanks to help from do_machine_check()
> 3) Provide helpful macros to decode the return value.
>
> Signed-off-by: Tony Luck <tony.luck@intel.com>
> ---
>  arch/x86/include/asm/uaccess_64.h |  5 +++
>  arch/x86/kernel/x8664_ksyms_64.c  |  2 +
>  arch/x86/lib/copy_user_64.S       | 91 +++++++++++++++++++++++++++++++++++++++
>  3 files changed, 98 insertions(+)
>
> diff --git a/arch/x86/include/asm/uaccess_64.h b/arch/x86/include/asm/uaccess_64.h
> index f2f9b39b274a..779cb0e77ecc 100644
> --- a/arch/x86/include/asm/uaccess_64.h
> +++ b/arch/x86/include/asm/uaccess_64.h
> @@ -216,6 +216,11 @@ __copy_to_user_inatomic(void __user *dst, const void *src, unsigned size)
>  extern long __copy_user_nocache(void *dst, const void __user *src,
>                                 unsigned size, int zerorest);
>
> +extern u64 mcsafe_memcpy(void *dst, const void __user *src,
> +                               unsigned size);
> +#define COPY_HAD_MCHECK(ret)   ((ret) & BIT(63))
> +#define        COPY_MCHECK_PADDR(ret)  ((ret) & ~BIT(63))
> +
>  static inline int
>  __copy_from_user_nocache(void *dst, const void __user *src, unsigned size)
>  {
> diff --git a/arch/x86/kernel/x8664_ksyms_64.c b/arch/x86/kernel/x8664_ksyms_64.c
> index a0695be19864..ec988c92c055 100644
> --- a/arch/x86/kernel/x8664_ksyms_64.c
> +++ b/arch/x86/kernel/x8664_ksyms_64.c
> @@ -37,6 +37,8 @@ EXPORT_SYMBOL(__copy_user_nocache);
>  EXPORT_SYMBOL(_copy_from_user);
>  EXPORT_SYMBOL(_copy_to_user);
>
> +EXPORT_SYMBOL(mcsafe_memcpy);
> +
>  EXPORT_SYMBOL(copy_page);
>  EXPORT_SYMBOL(clear_page);
>
> diff --git a/arch/x86/lib/copy_user_64.S b/arch/x86/lib/copy_user_64.S
> index 982ce34f4a9b..ffce93cbc9a5 100644
> --- a/arch/x86/lib/copy_user_64.S
> +++ b/arch/x86/lib/copy_user_64.S
> @@ -319,3 +319,94 @@ ENTRY(__copy_user_nocache)
>         _ASM_EXTABLE(21b,50b)
>         _ASM_EXTABLE(22b,50b)
>  ENDPROC(__copy_user_nocache)
> +
> +/*
> + * mcsafe_memcpy - Uncached memory copy with machine check exception handling
> + * Note that we only catch machine checks when reading the source addresses.
> + * Writes to target are posted and don't generate machine checks.
> + * This will force destination/source out of cache for more performance.
> + */
> +ENTRY(mcsafe_memcpy)
> +       cmpl $8,%edx
> +       jb 20f          /* less then 8 bytes, go to byte copy loop */
> +
> +       /* check for bad alignment of destination */
> +       movl %edi,%ecx
> +       andl $7,%ecx
> +       jz 102f                         /* already aligned */
> +       subl $8,%ecx
> +       negl %ecx
> +       subl %ecx,%edx
> +0:     movb (%rsi),%al
> +       movb %al,(%rdi)
> +       incq %rsi
> +       incq %rdi
> +       decl %ecx
> +       jnz 100b
> +102:
> +       movl %edx,%ecx
> +       andl $63,%edx
> +       shrl $6,%ecx
> +       jz 17f
> +1:     movq (%rsi),%r8
> +2:     movq 1*8(%rsi),%r9
> +3:     movq 2*8(%rsi),%r10
> +4:     movq 3*8(%rsi),%r11
> +       movnti %r8,(%rdi)
> +       movnti %r9,1*8(%rdi)
> +       movnti %r10,2*8(%rdi)
> +       movnti %r11,3*8(%rdi)
> +9:     movq 4*8(%rsi),%r8
> +10:    movq 5*8(%rsi),%r9
> +11:    movq 6*8(%rsi),%r10
> +12:    movq 7*8(%rsi),%r11
> +       movnti %r8,4*8(%rdi)
> +       movnti %r9,5*8(%rdi)
> +       movnti %r10,6*8(%rdi)
> +       movnti %r11,7*8(%rdi)
> +       leaq 64(%rsi),%rsi
> +       leaq 64(%rdi),%rdi
> +       decl %ecx
> +       jnz 1b
> +17:    movl %edx,%ecx
> +       andl $7,%edx
> +       shrl $3,%ecx
> +       jz 20f
> +18:    movq (%rsi),%r8
> +       movnti %r8,(%rdi)
> +       leaq 8(%rsi),%rsi
> +       leaq 8(%rdi),%rdi
> +       decl %ecx
> +       jnz 18b
> +20:    andl %edx,%edx
> +       jz 23f
> +       movl %edx,%ecx
> +21:    movb (%rsi),%al
> +       movb %al,(%rdi)
> +       incq %rsi
> +       incq %rdi
> +       decl %ecx
> +       jnz 21b
> +23:    xorl %eax,%eax
> +       sfence
> +       ret
> +
> +       .section .fixup,"ax"
> +30:
> +       sfence
> +       /* do_machine_check() sets %eax return value */
> +       ret
> +       .previous
> +
> +       _ASM_MCEXTABLE(0b,30b)
> +       _ASM_MCEXTABLE(1b,30b)
> +       _ASM_MCEXTABLE(2b,30b)
> +       _ASM_MCEXTABLE(3b,30b)
> +       _ASM_MCEXTABLE(4b,30b)
> +       _ASM_MCEXTABLE(9b,30b)
> +       _ASM_MCEXTABLE(10b,30b)
> +       _ASM_MCEXTABLE(11b,30b)
> +       _ASM_MCEXTABLE(12b,30b)
> +       _ASM_MCEXTABLE(18b,30b)
> +       _ASM_MCEXTABLE(21b,30b)
> +ENDPROC(mcsafe_memcpy)

I still don't get the BIT(63) thing.  Can you explain it?

--Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1289907 — RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

From"Luck, Tony" <tony.luck@intel.com>
Date2015-12-11 22:20 +0100
SubjectRE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qEDgB-1n0-11@gated-at.bofh.it>
In reply to#1289876
PiBJIHN0aWxsIGRvbid0IGdldCB0aGUgQklUKDYzKSB0aGluZy4gIENhbiB5b3UgZXhwbGFpbiBp
dD8NCg0KSXQgd2lsbCBiZSBtb3JlIG9idmlvdXMgd2hlbiBJIGdldCBhcm91bmQgdG8gd3JpdGlu
ZyBjb3B5X2Zyb21fdXNlcigpLg0KDQpUaGVuIHdlIHdpbGwgaGF2ZSBhIGZ1bmN0aW9uIHRoYXQg
Y2FuIHRha2UgcGFnZSBmYXVsdHMgaWYgdGhlcmUgYXJlIHBhZ2VzDQp0aGF0IGFyZSBub3QgcHJl
c2VudC4gIElmIHRoZSBwYWdlIGZhdWx0cyBjYW4ndCBiZSBmaXhlZCB3ZSBoYXZlIGEgLUVGQVVM
VA0KY29uZGl0aW9uLiBXZSBjYW4gYWxzbyB0YWtlIG1hY2hpbmUgY2hlY2tzIGlmIHdlIHJlYWRz
IGZyb20gYSBsb2NhdGlvbiB3aXRoIGFuDQp1bmNvcnJlY3RlZCBlcnJvci4NCg0KV2UgbmVlZCB0
byBkaXN0aW5ndWlzaCB0aGVzZSB0d28gY2FzZXMgYmVjYXVzZSB0aGUgYWN0aW9uIHdlIHRha2Ug
aXMNCmRpZmZlcmVudC4gRm9yIHRoZSB1bnJlc29sdmVkIHBhZ2UgZmF1bHQgd2UgYWxyZWFkeSBo
YXZlIHRoZSBBQkkgdGhhdCB0aGUNCmNvcHlfdG8vZnJvbV91c2VyKCkgZnVuY3Rpb25zIHJldHVy
biB6ZXJvIGZvciBzdWNjZXNzLCBhbmQgYSBub24temVybw0KcmV0dXJuIGlzIHRoZSBudW1iZXIg
b2Ygbm90LWNvcGllZCBieXRlcy4NCg0KU28gZm9yIG15IG5ldyBjYXNlIEknbSBzZXR0aW5nIGJp
dDYzIC4uLiB0aGlzIGlzIG5ldmVyIGdvaW5nIHRvIGJlIHNldCBmb3INCmEgZmFpbGVkIHBhZ2Ug
ZmF1bHQuDQoNCmNvcHlfZnJvbV91c2VyKCkgY29uY2VwdHVhbGx5IHdpbGwgbG9vayBsaWtlIHRo
aXM6DQoNCmludCBjb3B5X2Zyb21fdXNlcih2b2lkICp0bywgdm9pZCAqZnJvbSwgdW5zaWduZWQg
bG9uZyBuKQ0Kew0KCXU2NCByZXQgPSBtY3NhZmVfbWVtY3B5KHRvLCBmcm9tLCBuKTsNCg0KCWlm
IChDT1BZX0hBRF9NQ0hFQ0socikpIHsNCgkJaWYgKG1lbW9yeV9mYWlsdXJlKENPUFlfTUNIRUNL
X1BBRERSKHJldCkgPj4gUEFHRV9TSVpFLCAuLi4pKQ0KCQkJZm9yY2Vfc2lnKFNJR0JVUywgY3Vy
cmVudCk7DQoJCXJldHVybiBzb21ldGhpbmc7DQoJfSBlbHNlDQoJCXJldHVybiByZXQ7DQp9DQoN
Ci1Ub255DQo=
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1289937 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromAndy Lutomirski <luto@amacapital.net>
Date2015-12-11 23:00 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qEDTk-1BH-5@gated-at.bofh.it>
In reply to#1289907
On Fri, Dec 11, 2015 at 1:19 PM, Luck, Tony <tony.luck@intel.com> wrote:
>> I still don't get the BIT(63) thing.  Can you explain it?
>
> It will be more obvious when I get around to writing copy_from_user().
>
> Then we will have a function that can take page faults if there are pages
> that are not present.  If the page faults can't be fixed we have a -EFAULT
> condition. We can also take machine checks if we reads from a location with an
> uncorrected error.
>
> We need to distinguish these two cases because the action we take is
> different. For the unresolved page fault we already have the ABI that the
> copy_to/from_user() functions return zero for success, and a non-zero
> return is the number of not-copied bytes.

I'm missing something, though.  The normal fixup_exception path
doesn't touch rax at all.  The memory_failure path does.  But couldn't
you distinguish them by just pointing the exception handlers at
different landing pads?

Also, would it be more straightforward if the mcexception landing pad
looked up the va -> pa mapping by itself?  Or is that somehow not
reliable?

--Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1289948 — RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

From"Luck, Tony" <tony.luck@intel.com>
Date2015-12-11 23:20 +0100
SubjectRE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qEEcF-1XA-11@gated-at.bofh.it>
In reply to#1289937
PiBJJ20gbWlzc2luZyBzb21ldGhpbmcsIHRob3VnaC4gIFRoZSBub3JtYWwgZml4dXBfZXhjZXB0
aW9uIHBhdGgNCj4gZG9lc24ndCB0b3VjaCByYXggYXQgYWxsLiAgVGhlIG1lbW9yeV9mYWlsdXJl
IHBhdGggZG9lcy4gIEJ1dCBjb3VsZG4ndA0KPiB5b3UgZGlzdGluZ3Vpc2ggdGhlbSBieSBqdXN0
IHBvaW50aW5nIHRoZSBleGNlcHRpb24gaGFuZGxlcnMgYXQNCj4gZGlmZmVyZW50IGxhbmRpbmcg
cGFkcz8NCg0KUGVyaGFwcyBJJ20ganVzdCB0cnlpbmcgdG8gdGFrZSBhIHNob3J0IGN1dCB0byBh
dm9pZCB3cml0aW5nDQpzb21lIGNsZXZlciBmaXh1cCBjb2RlIGZvciB0aGUgdGFyZ2V0IGlwIHRo
YXQgZ29lcyBpbnRvIHRoZQ0KZXhjZXB0aW9uIHRhYmxlLg0KDQpGb3IgX19jb3B5X3VzZXJfbm9j
YWNoZSgpIHdlIGhhdmUgZm91ciBwb3NzaWJsZSB0YXJnZXRzDQpmb3IgZml4dXAgZGVwZW5kaW5n
IG9uIHdoZXJlIHdlIHdlcmUgaW4gdGhlIGZ1bmN0aW9uLg0KDQogICAgICAgIC5zZWN0aW9uIC5m
aXh1cCwiYXgiDQozMDogICAgIHNobGwgJDYsJWVjeA0KICAgICAgICBhZGRsICVlY3gsJWVkeA0K
ICAgICAgICBqbXAgNjBmDQo0MDogICAgIGxlYSAoJXJkeCwlcmN4LDgpLCVyZHgNCiAgICAgICAg
am1wIDYwZg0KNTA6ICAgICBtb3ZsICVlY3gsJWVkeA0KNjA6ICAgICBzZmVuY2UNCiAgICAgICAg
am1wIGNvcHlfdXNlcl9oYW5kbGVfdGFpbA0KICAgICAgICAucHJldmlvdXMNCg0KTm90ZSB0aGF0
IHRoaXMgY29kZSBhbHNvIHRha2VzIGEgc2hvcnRjdXQNCmJ5IGp1bXBpbmcgdG8gY29weV91c2Vy
X2hhbmRsZV90YWlsKCkgdG8NCmZpbmlzaCB1cCB0aGUgY29weSBhIGJ5dGUgYXQgYSB0aW1lIC4u
LiBhbmQNCnJ1bm5pbmcgYmFjayBpbnRvIHRoZSBzYW1lIHBhZ2UgZmF1bHQgYSAybmQNCnRpbWUg
dG8gbWFrZSBzdXJlIHRoZSBieXRlIGNvdW50IGlzIGV4YWN0bHkNCnJpZ2h0Lg0KDQpJIHJlYWxs
eSwgcmVhbGx5LCBkb24ndCB3YW50IHRvIHJ1biBiYWNrIGludG8NCnRoZSBwb2lzb24gYWdhaW4u
ICBJdCB3b3VsZCBwcm9iYWJseSB3b3JrLCBidXQNCmJlY2F1c2UgY3VycmVudCBnZW5lcmF0aW9u
IEludGVsIGNwdXMgYnJvYWRjYXN0IG1hY2hpbmUNCmNoZWNrcyB0byBldmVyeSBsb2dpY2FsIGNw
dSwgaXQgaXMgYSBsb3Qgb2Ygb3ZlcmhlYWQsDQphbmQgcG90ZW50aWFsbHkgcmlza3kuDQoNCj4g
QWxzbywgd291bGQgaXQgYmUgbW9yZSBzdHJhaWdodGZvcndhcmQgaWYgdGhlIG1jZXhjZXB0aW9u
IGxhbmRpbmcgcGFkDQo+IGxvb2tlZCB1cCB0aGUgdmEgLT4gcGEgbWFwcGluZyBieSBpdHNlbGY/
ICBPciBpcyB0aGF0IHNvbWVob3cgbm90DQo+IHJlbGlhYmxlPw0KDQpJZiB3ZSBkaWQgZ2V0IGFs
bCB0aGUgYWJvdmUgcmlnaHQsIHRoZW4gd2UgY291bGQgaGF2ZQ0KdGFyZ2V0IHVzZSB2aXJ0X3Rv
X3BoeXMoKSB0byBjb252ZXJ0IHRvIHBoeXNpY2FsIC4uLg0KSSBkb24ndCBzZWUgdGhhdCB0aGlz
IHBhcnQgd291bGQgYmUgYSBwcm9ibGVtLg0KDQotVG9ueQ0KDQoNCg0KDQo=
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1289960 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromDan Williams <dan.j.williams@intel.com>
Date2015-12-11 23:30 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qEEmm-21X-11@gated-at.bofh.it>
In reply to#1289948
On Fri, Dec 11, 2015 at 2:17 PM, Luck, Tony <tony.luck@intel.com> wrote:
>> Also, would it be more straightforward if the mcexception landing pad
>> looked up the va -> pa mapping by itself?  Or is that somehow not
>> reliable?
>
> If we did get all the above right, then we could have
> target use virt_to_phys() to convert to physical ...
> I don't see that this part would be a problem.

virt_to_phys() implies a linear address.  In the case of the use in
the pmem driver we'll be using an ioremap()'d address off somewherein
vmalloc space.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1289965 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromAndy Lutomirski <luto@amacapital.net>
Date2015-12-11 23:30 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qEEmm-21X-25@gated-at.bofh.it>
In reply to#1289960
On Fri, Dec 11, 2015 at 2:20 PM, Dan Williams <dan.j.williams@intel.com> wrote:
> On Fri, Dec 11, 2015 at 2:17 PM, Luck, Tony <tony.luck@intel.com> wrote:
>>> Also, would it be more straightforward if the mcexception landing pad
>>> looked up the va -> pa mapping by itself?  Or is that somehow not
>>> reliable?
>>
>> If we did get all the above right, then we could have
>> target use virt_to_phys() to convert to physical ...
>> I don't see that this part would be a problem.
>
> virt_to_phys() implies a linear address.  In the case of the use in
> the pmem driver we'll be using an ioremap()'d address off somewherein
> vmalloc space.

There's always slow_virt_to_phys.

Note that I don't fundamentally object to passing the pa to the fixup
handler.  I just think we should try to disentangle that from figuring
out what exactly the failure was.

Also, are there really PCOMMIT-capable CPUs that still forcibly
broadcast MCE?  If, so, that's unfortunate.

--Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1289969 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromAndy Lutomirski <luto@amacapital.net>
Date2015-12-11 23:40 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qEEw1-261-1@gated-at.bofh.it>
In reply to#1289965
On Fri, Dec 11, 2015 at 2:35 PM, Luck, Tony <tony.luck@intel.com> wrote:
>> Also, are there really PCOMMIT-capable CPUs that still forcibly
>> broadcast MCE?  If, so, that's unfortunate.
>
> PCOMMIT and LMCE arrive together ... though BIOS is in the decision
> path to enable LMCE, so it is possible that some systems could still
> broadcast if the BIOS writer decides to not allow local.

I really wish Intel would stop doing that.

>
> But a machine check safe copy_from_user() would be useful
> current generation cpus that broadcast all the time.

Fair enough.

--Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1289979 — RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

From"Luck, Tony" <tony.luck@intel.com>
Date2015-12-11 23:50 +0100
SubjectRE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qEEFI-29G-23@gated-at.bofh.it>
In reply to#1289969
Pj4gQnV0IGEgbWFjaGluZSBjaGVjayBzYWZlIGNvcHlfZnJvbV91c2VyKCkgd291bGQgYmUgdXNl
ZnVsDQo+PiBjdXJyZW50IGdlbmVyYXRpb24gY3B1cyB0aGF0IGJyb2FkY2FzdCBhbGwgdGhlIHRp
bWUuDQo+DQo+IEZhaXIgZW5vdWdoLg0KDQpUaGFua3MgZm9yIHNwZW5kaW5nIHRoZSB0aW1lIHRv
IGxvb2sgYXQgdGhpcy4gIENvYXhpbmcgbWUgdG8gcmUtd3JpdGUgdGhlDQp0YWlsIG9mIGRvX21h
Y2hpbmVfY2hlY2soKSBoYXMgbWFkZSB0aGF0IGNvZGUgbXVjaCBiZXR0ZXIuIFRvbyBtYW55DQp5
ZWFycyBvZiBvbmUgcGF0Y2ggb24gdG9wIG9mIGFub3RoZXIgd2l0aG91dCBsb29raW5nIGF0IHRo
ZSB3aG9sZSBjb250ZXh0Lg0KDQpDb2dpdGF0ZSBvbiB0aGlzIHNlcmllcyBvdmVyIHRoZSB3ZWVr
ZW5kIGFuZCBzZWUgaWYgeW91IGNhbiBnaXZlIG1lDQphbiBBY2tlZC1ieSBvciBSZXZpZXdlZC1i
eSAoSSdsbCBiZSBhZGRpbmcgYSAjZGVmaW5lIGZvciBCSVQoNjMpKS4NCg0KLVRvbnkNCg==
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1289983 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromAndy Lutomirski <luto@amacapital.net>
Date2015-12-12 00:00 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qEEPo-2db-3@gated-at.bofh.it>
In reply to#1289979
On Fri, Dec 11, 2015 at 2:45 PM, Luck, Tony <tony.luck@intel.com> wrote:
>>> But a machine check safe copy_from_user() would be useful
>>> current generation cpus that broadcast all the time.
>>
>> Fair enough.
>
> Thanks for spending the time to look at this.  Coaxing me to re-write the
> tail of do_machine_check() has made that code much better. Too many
> years of one patch on top of another without looking at the whole context.
>
> Cogitate on this series over the weekend and see if you can give me
> an Acked-by or Reviewed-by (I'll be adding a #define for BIT(63)).

I can't review the MCE decoding part, because I don't understand it
nearly well enough.  The interaction with the core fault handling
looks fine, modulo any need to bikeshed on the macro naming (which
I'll refrain from doing).

I still think it would be better if you get rid of BIT(63) and use a
pair of landing pads, though.  They could be as simple as:

.Lpage_fault_goes_here:
    xorq %rax, %rax
    jmp .Lbad

.Lmce_goes_here:
    /* set high bit of rax or whatever */
    /* fall through */

.Lbad:
    /* deal with it */

That way the magic is isolated to the function that needs the magic.

Also, at least renaming the macro to EXTABLE_MC_PA_IN_AX might be
nice.  It'll keep future users honest.  Maybe some day there'll be a
PA_IN_AX flag, and, heck, maybe some day there'll be ways to get info
for non-MCE faults delivered through fixup_exception.

--Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1290963 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromIngo Molnar <mingo@kernel.org>
Date2015-12-14 09:40 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qFwPN-3DO-43@gated-at.bofh.it>
In reply to#1289983
* Andy Lutomirski <luto@amacapital.net> wrote:

> I still think it would be better if you get rid of BIT(63) and use a
> pair of landing pads, though.  They could be as simple as:
> 
> .Lpage_fault_goes_here:
>     xorq %rax, %rax
>     jmp .Lbad
> 
> .Lmce_goes_here:
>     /* set high bit of rax or whatever */
>     /* fall through */
> 
> .Lbad:
>     /* deal with it */
> 
> That way the magic is isolated to the function that needs the magic.

Seconded - this is the usual pattern we use in all assembly functions.

Thanks,

	Ingo
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1291509 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

From"Luck, Tony" <tony.luck@intel.com>
Date2015-12-14 20:50 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qFHia-26U-5@gated-at.bofh.it>
In reply to#1290963
On Mon, Dec 14, 2015 at 09:36:25AM +0100, Ingo Molnar wrote:
> >     /* deal with it */
> > 
> > That way the magic is isolated to the function that needs the magic.
> 
> Seconded - this is the usual pattern we use in all assembly functions.

Ok - you want me to write some x86 assembly code (you may regret that).

Initial question ... here's the fixup for __copy_user_nocache()

		.section .fixup,"ax"
	30:     shll $6,%ecx
		addl %ecx,%edx
		jmp 60f
	40:     lea (%rdx,%rcx,8),%rdx
		jmp 60f
	50:     movl %ecx,%edx
	60:     sfence
		jmp copy_user_handle_tail
		.previous

Are %ecx and %rcx synonyms for the same register? Is there some
super subtle reason we use the 'r' names in the "40" fixup, but
the 'e' names everywhere else in this code (and the 'e' names in
the body of the original function)?

-Tony
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1291526 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromAndy Lutomirski <luto@amacapital.net>
Date2015-12-14 21:20 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qFHLc-2w9-17@gated-at.bofh.it>
In reply to#1291509
On Mon, Dec 14, 2015 at 11:46 AM, Luck, Tony <tony.luck@intel.com> wrote:
> On Mon, Dec 14, 2015 at 09:36:25AM +0100, Ingo Molnar wrote:
>> >     /* deal with it */
>> >
>> > That way the magic is isolated to the function that needs the magic.
>>
>> Seconded - this is the usual pattern we use in all assembly functions.
>
> Ok - you want me to write some x86 assembly code (you may regret that).
>

All you have to do is erase all of the ia64 asm knowledge from your
brain and repurpose 1% of that space for x86 asm.  You'll be a
world-class expert!

> Initial question ... here's the fixup for __copy_user_nocache()
>
>                 .section .fixup,"ax"
>         30:     shll $6,%ecx
>                 addl %ecx,%edx
>                 jmp 60f
>         40:     lea (%rdx,%rcx,8),%rdx
>                 jmp 60f
>         50:     movl %ecx,%edx
>         60:     sfence
>                 jmp copy_user_handle_tail
>                 .previous
>
> Are %ecx and %rcx synonyms for the same register? Is there some
> super subtle reason we use the 'r' names in the "40" fixup, but
> the 'e' names everywhere else in this code (and the 'e' names in
> the body of the original function)?

rcx is a 64-bit register.  ecx is the low 32 bits of it.  If you read
from ecx, you get the low 32 bits, but if you write to ecx, you zero
the high bits as a side-effect.

--Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1289972 — RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

From"Luck, Tony" <tony.luck@intel.com>
Date2015-12-11 23:40 +0100
SubjectRE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qEEw1-261-3@gated-at.bofh.it>
In reply to#1289965
PiBBbHNvLCBhcmUgdGhlcmUgcmVhbGx5IFBDT01NSVQtY2FwYWJsZSBDUFVzIHRoYXQgc3RpbGwg
Zm9yY2libHkNCj4gYnJvYWRjYXN0IE1DRT8gIElmLCBzbywgdGhhdCdzIHVuZm9ydHVuYXRlLg0K
DQpQQ09NTUlUIGFuZCBMTUNFIGFycml2ZSB0b2dldGhlciAuLi4gdGhvdWdoIEJJT1MgaXMgaW4g
dGhlIGRlY2lzaW9uDQpwYXRoIHRvIGVuYWJsZSBMTUNFLCBzbyBpdCBpcyBwb3NzaWJsZSB0aGF0
IHNvbWUgc3lzdGVtcyBjb3VsZCBzdGlsbA0KYnJvYWRjYXN0IGlmIHRoZSBCSU9TIHdyaXRlciBk
ZWNpZGVzIHRvIG5vdCBhbGxvdyBsb2NhbC4NCg0KQnV0IGEgbWFjaGluZSBjaGVjayBzYWZlIGNv
cHlfZnJvbV91c2VyKCkgd291bGQgYmUgdXNlZnVsDQpjdXJyZW50IGdlbmVyYXRpb24gY3B1cyB0
aGF0IGJyb2FkY2FzdCBhbGwgdGhlIHRpbWUuDQoNCi1Ub255DQo=
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1292168 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromBorislav Petkov <bp@alien8.de>
Date2015-12-15 14:20 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qFXGi-4AF-13@gated-at.bofh.it>
In reply to#1289855
On Thu, Dec 10, 2015 at 04:21:50PM -0800, Tony Luck wrote:
> Using __copy_user_nocache() as inspiration create a memory copy
> routine for use by kernel code with annotations to allow for
> recovery from machine checks.
> 
> Notes:
> 1) Unlike the original we make no attempt to copy all the bytes
>    up to the faulting address. The original achieves that by
>    re-executing the failing part as a byte-by-byte copy,
>    which will take another page fault. We don't want to have
>    a second machine check!
> 2) Likewise the return value for the original indicates exactly
>    how many bytes were not copied. Instead we provide the physical
>    address of the fault (thanks to help from do_machine_check()
> 3) Provide helpful macros to decode the return value.
> 
> Signed-off-by: Tony Luck <tony.luck@intel.com>
> ---
>  arch/x86/include/asm/uaccess_64.h |  5 +++
>  arch/x86/kernel/x8664_ksyms_64.c  |  2 +
>  arch/x86/lib/copy_user_64.S       | 91 +++++++++++++++++++++++++++++++++++++++
>  3 files changed, 98 insertions(+)

...

> + * mcsafe_memcpy - Uncached memory copy with machine check exception handling
> + * Note that we only catch machine checks when reading the source addresses.
> + * Writes to target are posted and don't generate machine checks.
> + * This will force destination/source out of cache for more performance.

... and the non-temporal version is the optimal one even though we're
defaulting to copy_user_enhanced_fast_string for memcpy on modern Intel
CPUs...?

Btw, it should be also inside an ifdef if we're going to ifdef
CONFIG_MCE_KERNEL_RECOVERY everywhere else.

-- 
Regards/Gruss,
    Boris.

ECO tip #101: Trim your mails when you reply.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1292407 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromDan Williams <dan.j.williams@intel.com>
Date2015-12-15 18:50 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qG1Tz-7mv-5@gated-at.bofh.it>
In reply to#1292168
On Tue, Dec 15, 2015 at 5:11 AM, Borislav Petkov <bp@alien8.de> wrote:
> On Thu, Dec 10, 2015 at 04:21:50PM -0800, Tony Luck wrote:
>> Using __copy_user_nocache() as inspiration create a memory copy
>> routine for use by kernel code with annotations to allow for
>> recovery from machine checks.
>>
>> Notes:
>> 1) Unlike the original we make no attempt to copy all the bytes
>>    up to the faulting address. The original achieves that by
>>    re-executing the failing part as a byte-by-byte copy,
>>    which will take another page fault. We don't want to have
>>    a second machine check!
>> 2) Likewise the return value for the original indicates exactly
>>    how many bytes were not copied. Instead we provide the physical
>>    address of the fault (thanks to help from do_machine_check()
>> 3) Provide helpful macros to decode the return value.
>>
>> Signed-off-by: Tony Luck <tony.luck@intel.com>
>> ---
>>  arch/x86/include/asm/uaccess_64.h |  5 +++
>>  arch/x86/kernel/x8664_ksyms_64.c  |  2 +
>>  arch/x86/lib/copy_user_64.S       | 91 +++++++++++++++++++++++++++++++++++++++
>>  3 files changed, 98 insertions(+)
>
> ...
>
>> + * mcsafe_memcpy - Uncached memory copy with machine check exception handling
>> + * Note that we only catch machine checks when reading the source addresses.
>> + * Writes to target are posted and don't generate machine checks.
>> + * This will force destination/source out of cache for more performance.
>
> ... and the non-temporal version is the optimal one even though we're
> defaulting to copy_user_enhanced_fast_string for memcpy on modern Intel
> CPUs...?

At least the pmem driver use case does not want caching of the
source-buffer since that is the raw "disk" media.  I.e. in
pmem_do_bvec() we'd use this to implement memcpy_from_pmem().
However, caching the destination-buffer may prove beneficial since
that data is likely to be consumed immediately by the thread that
submitted the i/o.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1292417 — RE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

From"Luck, Tony" <tony.luck@intel.com>
Date2015-12-15 19:00 +0100
SubjectRE: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qG23g-7pS-19@gated-at.bofh.it>
In reply to#1292407
Pj4gLi4uIGFuZCB0aGUgbm9uLXRlbXBvcmFsIHZlcnNpb24gaXMgdGhlIG9wdGltYWwgb25lIGV2
ZW4gdGhvdWdoIHdlJ3JlDQo+PiBkZWZhdWx0aW5nIHRvIGNvcHlfdXNlcl9lbmhhbmNlZF9mYXN0
X3N0cmluZyBmb3IgbWVtY3B5IG9uIG1vZGVybiBJbnRlbA0KPj4gQ1BVcy4uLj8NCg0KTXkgY3Vy
cmVudCBnZW5lcmF0aW9uIGNwdSBoYXMgYSBiaXQgb2YgYW4gaXNzdWUgd2l0aCByZWNvdmVyaW5n
IGZyb20gYQ0KbWFjaGluZSBjaGVjayBpbiBhICJyZXAgbW92IiAuLi4gc28gSSdtIHdvcmtpbmcg
d2l0aCBhIHZlcnNpb24gb2YgbWVtY3B5DQp0aGF0IHVucm9sbHMgaW50byBpbmRpdmlkdWFsIG1v
diBpbnN0cnVjdGlvbnMgZm9yIG5vdy4NCg0KPiBBdCBsZWFzdCB0aGUgcG1lbSBkcml2ZXIgdXNl
IGNhc2UgZG9lcyBub3Qgd2FudCBjYWNoaW5nIG9mIHRoZQ0KPiBzb3VyY2UtYnVmZmVyIHNpbmNl
IHRoYXQgaXMgdGhlIHJhdyAiZGlzayIgbWVkaWEuICBJLmUuIGluDQo+IHBtZW1fZG9fYnZlYygp
IHdlJ2QgdXNlIHRoaXMgdG8gaW1wbGVtZW50IG1lbWNweV9mcm9tX3BtZW0oKS4NCj4gSG93ZXZl
ciwgY2FjaGluZyB0aGUgZGVzdGluYXRpb24tYnVmZmVyIG1heSBwcm92ZSBiZW5lZmljaWFsIHNp
bmNlDQo+IHRoYXQgZGF0YSBpcyBsaWtlbHkgdG8gYmUgY29uc3VtZWQgaW1tZWRpYXRlbHkgYnkg
dGhlIHRocmVhZCB0aGF0DQo+IHN1Ym1pdHRlZCB0aGUgaS9vLg0KDQpJIGNhbiBkcm9wIHRoZSAi
bnRpIiBmcm9tIHRoZSBkZXN0aW5hdGlvbiBtb3Zlcy4gIERvZXMgIm50aSIgd29yaw0Kb24gdGhl
IGxvYWQgZnJvbSBzb3VyY2UgYWRkcmVzcyBzaWRlIHRvIGF2b2lkIGNhY2hlIGFsbG9jYXRpb24/
DQoNCk9uIGFub3RoZXIgdG9waWMgcmFpc2VkIGJ5IEJvcmlzIC4uLiBpcyB0aGVyZSBzb21lIENP
TkZJR19QTUVNKg0KdGhhdCBJIHNob3VsZCB1c2UgYXMgYSBkZXBlbmRlbmN5IHRvIGVuYWJsZSBh
bGwgdGhpcz8NCg0KLVRvbnkNCg==
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1292434 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromBorislav Petkov <bp@alien8.de>
Date2015-12-15 19:30 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qG2wi-7PJ-3@gated-at.bofh.it>
In reply to#1292417
On Tue, Dec 15, 2015 at 05:53:31PM +0000, Luck, Tony wrote:
> My current generation cpu has a bit of an issue with recovering from a
> machine check in a "rep mov" ... so I'm working with a version of memcpy
> that unrolls into individual mov instructions for now.

Ah.

> I can drop the "nti" from the destination moves.  Does "nti" work
> on the load from source address side to avoid cache allocation?

I don't think so:

+1:     movq (%rsi),%r8
+2:     movq 1*8(%rsi),%r9
+3:     movq 2*8(%rsi),%r10
+4:     movq 3*8(%rsi),%r11
...

You need to load the data into registers first because MOVNTI needs them
there as it does reg -> mem movement. That first load from memory into
registers with a normal MOV will pull the data into the cache.

Perhaps the first thing to try would be to see what slowdown normal MOVs
bring and if not really noticeable, use those instead.

> On another topic raised by Boris ... is there some CONFIG_PMEM*
> that I should use as a dependency to enable all this?

I found CONFIG_LIBNVDIMM only today:

drivers/nvdimm/Kconfig:1:menuconfig LIBNVDIMM
drivers/nvdimm/Kconfig:2:       tristate "NVDIMM (Non-Volatile Memory Device) Support"

-- 
Regards/Gruss,
    Boris.

ECO tip #101: Trim your mails when you reply.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1292438 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromDan Williams <dan.j.williams@intel.com>
Date2015-12-15 19:30 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qG2wi-7PJ-5@gated-at.bofh.it>
In reply to#1292417
On Tue, Dec 15, 2015 at 9:53 AM, Luck, Tony <tony.luck@intel.com> wrote:
>>> ... and the non-temporal version is the optimal one even though we're
>>> defaulting to copy_user_enhanced_fast_string for memcpy on modern Intel
>>> CPUs...?
>
> My current generation cpu has a bit of an issue with recovering from a
> machine check in a "rep mov" ... so I'm working with a version of memcpy
> that unrolls into individual mov instructions for now.
>
>> At least the pmem driver use case does not want caching of the
>> source-buffer since that is the raw "disk" media.  I.e. in
>> pmem_do_bvec() we'd use this to implement memcpy_from_pmem().
>> However, caching the destination-buffer may prove beneficial since
>> that data is likely to be consumed immediately by the thread that
>> submitted the i/o.
>
> I can drop the "nti" from the destination moves.  Does "nti" work
> on the load from source address side to avoid cache allocation?

My mistake, I don't think we have an uncached load capability, only store.

> On another topic raised by Boris ... is there some CONFIG_PMEM*
> that I should use as a dependency to enable all this?

I'd rather make this a "select ARCH_MCSAFE_MEMCPY".  Since it's not a
hard dependency and the details will be hidden behind
memcpy_from_pmem().  Specifically, the details will be handled by a
new arch_memcpy_from_pmem() in arch/x86/include/asm/pmem.h to
supplement the existing arch_memcpy_to_pmem().
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1292448 — Re: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks

FromDan Williams <dan.j.williams@intel.com>
Date2015-12-15 19:40 +0100
SubjectRe: [PATCHV2 3/3] x86, ras: Add mcsafe_memcpy() function to recover from machine checks
Message-ID<qG2FX-7U3-1@gated-at.bofh.it>
In reply to#1292438
On Tue, Dec 15, 2015 at 10:27 AM, Dan Williams <dan.j.williams@intel.com> wrote:
> On Tue, Dec 15, 2015 at 9:53 AM, Luck, Tony <tony.luck@intel.com> wrote:
>>>> ... and the non-temporal version is the optimal one even though we're
>>>> defaulting to copy_user_enhanced_fast_string for memcpy on modern Intel
>>>> CPUs...?
>>
>> My current generation cpu has a bit of an issue with recovering from a
>> machine check in a "rep mov" ... so I'm working with a version of memcpy
>> that unrolls into individual mov instructions for now.
>>
>>> At least the pmem driver use case does not want caching of the
>>> source-buffer since that is the raw "disk" media.  I.e. in
>>> pmem_do_bvec() we'd use this to implement memcpy_from_pmem().
>>> However, caching the destination-buffer may prove beneficial since
>>> that data is likely to be consumed immediately by the thread that
>>> submitted the i/o.
>>
>> I can drop the "nti" from the destination moves.  Does "nti" work
>> on the load from source address side to avoid cache allocation?
>
> My mistake, I don't think we have an uncached load capability, only store.

Correction we have MOVNTDQA, but that requires saving the fpu state
and marking the memory as WC, i.e. probably not worth it.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.kernel


csiph-web