Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1547445 > unrolled thread
| Started by | "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> |
|---|---|
| First post | 2016-12-27 03:00 +0100 |
| Last post | 2017-01-05 18:00 +0100 |
| Articles | 20 on this page of 52 — 9 participants |
Back to article view | Back to linux.kernel
[PATCHv2 00/29] 5-level paging "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:00 +0100
[PATCHv2 17/29] x86/asm: remove __VIRTUAL_MASK_SHIFT==47 assert "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:00 +0100
[PATCHv2 03/29] asm-generic: introduce __ARCH_USE_5LEVEL_HACK "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:00 +0100
[RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:00 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2016-12-27 03:20 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-12-27 03:40 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2016-12-27 04:30 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-01-02 10:20 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Carlos O'Donell <carlos@redhat.com> - 2016-12-29 04:00 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2016-12-31 03:10 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-01-02 09:40 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "H.J. Lu" <hjl.tools@gmail.com> - 2017-01-13 21:20 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Arnd Bergmann <arnd@arndb.de> - 2017-01-02 09:50 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2017-01-03 07:10 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Arnd Bergmann <arnd@arndb.de> - 2017-01-03 14:30 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2017-01-03 19:40 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2017-01-03 23:10 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Arnd Bergmann <arnd@arndb.de> - 2017-01-04 15:00 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Arnd Bergmann <arnd@arndb.de> - 2017-01-03 23:20 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-01-03 17:10 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2017-01-03 19:30 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-01-04 15:30 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2017-01-05 19:20 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Dave Hansen <dave.hansen@intel.com> - 2017-01-05 20:20 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2017-01-05 20:40 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Dave Hansen <dave.hansen@intel.com> - 2017-01-05 20:40 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2017-01-05 21:20 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Dave Hansen <dave.hansen@intel.com> - 2017-01-05 21:50 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2017-01-05 22:30 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Dave Hansen <dave.hansen@intel.com> - 2017-01-06 00:20 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-01-11 15:30 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2017-01-11 19:10 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-01-11 19:40 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Dave Hansen <dave.hansen@intel.com> - 2017-01-11 20:00 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andy Lutomirski <luto@amacapital.net> - 2017-01-11 20:30 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Linus Torvalds <torvalds@linux-foundation.org> - 2017-01-11 20:40 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Andi Kleen <ak@linux.intel.com> - 2017-01-11 22:50 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-01-11 20:40 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Linus Torvalds <torvalds@linux-foundation.org> - 2017-01-11 20:40 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR Dave Hansen <dave.hansen@intel.com> - 2017-01-11 19:30 +0100
Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-01-05 21:20 +0100
[PATCHv2 19/29] x86/paravirt: make paravirt code support 5-level paging "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:00 +0100
[PATCHv2 27/29] x86/mm: add support for 5-level paging for KASLR "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:00 +0100
[PATCHv2 12/29] x86/mm: add support of p4d_t in vmalloc_fault() "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:00 +0100
[PATCHv2 18/29] x86/mm: define virtual memory map for 5-level paging "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:00 +0100
[PATCHv2 04/29] arch, mm: convert all architectures to use 5level-fixup.h "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:10 +0100
[PATCHv2 02/29] asm-generic: introduce 5level-fixup.h "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:10 +0100
[PATCHv2 14/29] x86/kexec: support p4d_t "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:10 +0100
[PATCHv2 08/29] x86: basic changes into headers for 5-level paging "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:10 +0100
[PATCHv2 05/29] asm-generic: introduce <asm-generic/pgtable-nop4d.h> "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:10 +0100
[PATCHv2 15/29] x86: convert the rest of the code to support p4d_t "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-12-27 03:10 +0100
Re: [PATCHv2 00/29] 5-level paging "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-01-05 18:00 +0100
Page 1 of 3 [1] 2 3 Next page →
| From | "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> |
|---|---|
| Date | 2016-12-27 03:00 +0100 |
| Subject | [PATCHv2 00/29] 5-level paging |
| Message-ID | <sSPdv-45l-3@gated-at.bofh.it> |
Here is v2 of 5-level paging patchset.
Please consider applying first 7 patches.
Main x86 portion requires more work (mostly Xen), but any feedback is
welcome.
I've also included a patch proposal to address compatibility issue with
wide addresses that some software has.
Ingo wants to to see opt-out, but I believe opt-in is required to not
break userspace. That's not purely theoretical, I see one crash due to
the issue.
See description of patchset split below.
== Overview ==
x86-64 is currently limited to 256 TiB of virtual address space and 64 TiB
of physical address space. We are already bumping into this limit: some
vendors offers servers with 64 TiB of memory today.
To overcome the limitation upcoming hardware will introduce support for
5-level paging[1]. It is a straight-forward extension of the current page
table structure adding one more layer of translation.
It bumps the limits to 128 PiB of virtual address space and 4 PiB of
physical address space. This "ought to be enough for anybody" ©.
== Patches ==
The patchset is build on top of v4.10-rc1.
Current QEMU upstream git supports 5-level paging. Use "-cpu qemu64,+la57"
to enable it.
Patch 1:
Detect la57 feature for /proc/cpuinfo. It's trivial and can be
applied now.
Patches 2-7:
Brings 5-level paging to generic code and convert all
architectures to it using <asm-generic/5level-fixup.h>
I think this should be ready for -mm tree. Please consider
applying.
Patches 8-15:
Convert x86 to properly folded p4d layer using
<asm-generic/pgtable-nop4d.h>.
Patches 16-28:
Enabling of real 5-level paging.
CONFIG_X86_5LEVEL=y will enable new paging mode.
Patch 29:
Extends rlimit interface to set/get maximum virtual address that
the application can map.
This aims to address compatibility issue. Only supports x86 for
now.
I also patched dash to add the rlimit support. It was handy for
testing. Let me know if anybody wants to play with it.
Comments are welcome.
Git:
git://git.kernel.org/pub/scm/linux/kernel/git/kas/linux.git la57/v2
== TODO ==
There is still work to do:
- CONFIG_XEN is broken.
Paravirt Xen MMU support hasn't yet adjusted to work with 5-level
paging. It's legacy feature, not sure if we really need to support it
with new paging, but it blocks Xen drivers too.
I haven't got around to setup testing environment for XEN, so left it
broken for now.
I would appreciate help with the code.
- Boot-time switch between 4- and 5-level paging.
We assume that distributions will be keen to avoid returning to the
i386 days where we shipped one kernel binary for each page table
layout.
As page table format is the same for 4- and 5-level paging it should
be possible to have single kernel binary and switch between them at
boot-time without too much hassle.
For now I only implemented compile-time switch.
This will implemented with separate patchset.
== Changelong ==
v2:
- Rebased to v4.10-rc1;
- RLIMIT_VADDR proposal;
- Fix virtual map and update documentation;
- Fix few build errors;
- Rework cpuid helpers in boot code;
- Fix espfix code to work with 5-level pages;
[1] https://software.intel.com/sites/default/files/managed/2b/80/5-level_paging_white_paper.pdf
Kirill A. Shutemov (29):
x86/cpufeature: Add 5-level paging detecton
asm-generic: introduce 5level-fixup.h
asm-generic: introduce __ARCH_USE_5LEVEL_HACK
arch, mm: convert all architectures to use 5level-fixup.h
asm-generic: introduce <asm-generic/pgtable-nop4d.h>
mm: convert generic code to 5-level paging
mm: introduce __p4d_alloc()
x86: basic changes into headers for 5-level paging
x86: trivial portion of 5-level paging conversion
x86/gup: add 5-level paging support
x86/ident_map: add 5-level paging support
x86/mm: add support of p4d_t in vmalloc_fault()
x86/power: support p4d_t in hibernate code
x86/kexec: support p4d_t
x86: convert the rest of the code to support p4d_t
x86: detect 5-level paging support
x86/asm: remove __VIRTUAL_MASK_SHIFT==47 assert
x86/mm: define virtual memory map for 5-level paging
x86/paravirt: make paravirt code support 5-level paging
x86/mm: basic defines/helpers for CONFIG_X86_5LEVEL
x86/dump_pagetables: support 5-level paging
x86/mm: extend kasan to support 5-level paging
x86/espfix: support 5-level paging
x86/mm: add support of additional page table level during early boot
x86/mm: add sync_global_pgds() for configuration with 5-level paging
x86/mm: make kernel_physical_mapping_init() support 5-level paging
x86/mm: add support for 5-level paging for KASLR
x86: enable 5-level paging support
mm, x86: introduce RLIMIT_VADDR
Documentation/x86/x86_64/mm.txt | 33 ++-
arch/arc/include/asm/hugepage.h | 1 +
arch/arc/include/asm/pgtable.h | 1 +
arch/arm/include/asm/pgtable.h | 1 +
arch/arm64/include/asm/pgtable-types.h | 4 +
arch/avr32/include/asm/pgtable-2level.h | 1 +
arch/cris/include/asm/pgtable.h | 1 +
arch/frv/include/asm/pgtable.h | 1 +
arch/h8300/include/asm/pgtable.h | 1 +
arch/hexagon/include/asm/pgtable.h | 1 +
arch/ia64/include/asm/pgtable.h | 2 +
arch/metag/include/asm/pgtable.h | 1 +
arch/mips/include/asm/pgtable-32.h | 1 +
arch/mips/include/asm/pgtable-64.h | 1 +
arch/mn10300/include/asm/page.h | 1 +
arch/nios2/include/asm/pgtable.h | 1 +
arch/openrisc/include/asm/pgtable.h | 1 +
arch/powerpc/include/asm/book3s/32/pgtable.h | 1 +
arch/powerpc/include/asm/book3s/64/pgtable.h | 2 +
arch/powerpc/include/asm/nohash/32/pgtable.h | 1 +
arch/powerpc/include/asm/nohash/64/pgtable-4k.h | 3 +
arch/powerpc/include/asm/nohash/64/pgtable-64k.h | 1 +
arch/s390/include/asm/pgtable.h | 1 +
arch/score/include/asm/pgtable.h | 1 +
arch/sh/include/asm/pgtable-2level.h | 1 +
arch/sh/include/asm/pgtable-3level.h | 1 +
arch/sparc/include/asm/pgtable_64.h | 1 +
arch/tile/include/asm/pgtable_32.h | 1 +
arch/tile/include/asm/pgtable_64.h | 1 +
arch/um/include/asm/pgtable-2level.h | 1 +
arch/um/include/asm/pgtable-3level.h | 1 +
arch/unicore32/include/asm/pgtable.h | 1 +
arch/x86/Kconfig | 7 +
arch/x86/boot/compressed/head_64.S | 23 ++-
arch/x86/boot/cpucheck.c | 9 +
arch/x86/boot/cpuflags.c | 12 +-
arch/x86/entry/entry_64.S | 7 +-
arch/x86/include/asm/cpufeatures.h | 3 +-
arch/x86/include/asm/disabled-features.h | 8 +-
arch/x86/include/asm/elf.h | 2 +-
arch/x86/include/asm/kasan.h | 9 +-
arch/x86/include/asm/kexec.h | 1 +
arch/x86/include/asm/page_64_types.h | 10 +
arch/x86/include/asm/paravirt.h | 64 +++++-
arch/x86/include/asm/paravirt_types.h | 17 +-
arch/x86/include/asm/pgalloc.h | 36 +++-
arch/x86/include/asm/pgtable-2level_types.h | 1 +
arch/x86/include/asm/pgtable-3level_types.h | 1 +
arch/x86/include/asm/pgtable.h | 91 ++++++++-
arch/x86/include/asm/pgtable_64.h | 29 ++-
arch/x86/include/asm/pgtable_64_types.h | 27 +++
arch/x86/include/asm/pgtable_types.h | 42 +++-
arch/x86/include/asm/processor.h | 17 +-
arch/x86/include/asm/required-features.h | 8 +-
arch/x86/include/asm/sparsemem.h | 9 +-
arch/x86/include/uapi/asm/processor-flags.h | 2 +
arch/x86/kernel/espfix_64.c | 12 +-
arch/x86/kernel/head64.c | 40 +++-
arch/x86/kernel/head_64.S | 58 ++++--
arch/x86/kernel/machine_kexec_32.c | 4 +-
arch/x86/kernel/machine_kexec_64.c | 14 +-
arch/x86/kernel/paravirt.c | 13 +-
arch/x86/kernel/sys_x86_64.c | 6 +-
arch/x86/kernel/tboot.c | 6 +-
arch/x86/kernel/vm86_32.c | 6 +-
arch/x86/mm/dump_pagetables.c | 51 ++++-
arch/x86/mm/fault.c | 57 +++++-
arch/x86/mm/gup.c | 33 ++-
arch/x86/mm/hugetlbpage.c | 8 +-
arch/x86/mm/ident_map.c | 42 +++-
arch/x86/mm/init_32.c | 22 +-
arch/x86/mm/init_64.c | 248 ++++++++++++++++++++---
arch/x86/mm/ioremap.c | 3 +-
arch/x86/mm/kasan_init_64.c | 42 +++-
arch/x86/mm/kaslr.c | 82 ++++++--
arch/x86/mm/mmap.c | 4 +-
arch/x86/mm/pageattr.c | 56 +++--
arch/x86/mm/pgtable.c | 38 +++-
arch/x86/mm/pgtable_32.c | 8 +-
arch/x86/platform/efi/efi_64.c | 21 +-
arch/x86/power/hibernate_32.c | 7 +-
arch/x86/power/hibernate_64.c | 35 ++--
arch/x86/realmode/init.c | 2 +-
arch/x86/xen/Kconfig | 1 +
arch/xtensa/include/asm/pgtable.h | 1 +
drivers/misc/sgi-gru/grufault.c | 9 +-
fs/binfmt_aout.c | 2 -
fs/binfmt_elf.c | 10 +-
fs/hugetlbfs/inode.c | 6 +-
fs/proc/base.c | 1 +
fs/userfaultfd.c | 6 +-
include/asm-generic/4level-fixup.h | 3 +-
include/asm-generic/5level-fixup.h | 41 ++++
include/asm-generic/pgtable-nop4d-hack.h | 62 ++++++
include/asm-generic/pgtable-nop4d.h | 56 +++++
include/asm-generic/pgtable-nopud.h | 48 +++--
include/asm-generic/pgtable.h | 48 ++++-
include/asm-generic/resource.h | 4 +
include/asm-generic/tlb.h | 14 +-
include/linux/hugetlb.h | 5 +-
include/linux/kasan.h | 1 +
include/linux/mm.h | 34 +++-
include/linux/sched.h | 5 +
include/uapi/asm-generic/resource.h | 3 +-
kernel/events/uprobes.c | 5 +-
kernel/sys.c | 6 +-
lib/ioremap.c | 39 +++-
mm/gup.c | 46 ++++-
mm/huge_memory.c | 7 +-
mm/hugetlb.c | 29 ++-
mm/kasan/kasan_init.c | 35 +++-
mm/memory.c | 230 +++++++++++++++++----
mm/mlock.c | 1 +
mm/mmap.c | 20 +-
mm/mprotect.c | 26 ++-
mm/mremap.c | 16 +-
mm/nommu.c | 2 +-
mm/pagewalk.c | 32 ++-
mm/pgtable-generic.c | 6 +
mm/rmap.c | 13 +-
mm/shmem.c | 8 +-
mm/sparse-vmemmap.c | 22 +-
mm/swapfile.c | 26 ++-
mm/userfaultfd.c | 23 ++-
mm/vmalloc.c | 81 ++++++--
125 files changed, 2047 insertions(+), 410 deletions(-)
create mode 100644 include/asm-generic/5level-fixup.h
create mode 100644 include/asm-generic/pgtable-nop4d-hack.h
create mode 100644 include/asm-generic/pgtable-nop4d.h
--
2.11.0
[toc] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> |
|---|---|
| Date | 2016-12-27 03:00 +0100 |
| Subject | [PATCHv2 17/29] x86/asm: remove __VIRTUAL_MASK_SHIFT==47 assert |
| Message-ID | <sSPdB-45l-45@gated-at.bofh.it> |
| In reply to | #1547445 |
We don't need it anymore. 17be0aec74fb ("x86/asm/entry/64: Implement
better check for canonical addresses") made canonical address check
generic wrt. address width.
Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
---
arch/x86/entry/entry_64.S | 7 ++-----
1 file changed, 2 insertions(+), 5 deletions(-)
diff --git a/arch/x86/entry/entry_64.S b/arch/x86/entry/entry_64.S
index 5b219707c2f2..a6b84c309685 100644
--- a/arch/x86/entry/entry_64.S
+++ b/arch/x86/entry/entry_64.S
@@ -264,12 +264,9 @@ return_from_SYSCALL_64:
*
* If width of "canonical tail" ever becomes variable, this will need
* to be updated to remain correct on both old and new CPUs.
+ *
+ * Change top 16 bits to be the sign-extension of 47th bit
*/
- .ifne __VIRTUAL_MASK_SHIFT - 47
- .error "virtual address width changed -- SYSRET checks need update"
- .endif
-
- /* Change top 16 bits to be the sign-extension of 47th bit */
shl $(64 - (__VIRTUAL_MASK_SHIFT+1)), %rcx
sar $(64 - (__VIRTUAL_MASK_SHIFT+1)), %rcx
--
2.11.0
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> |
|---|---|
| Date | 2016-12-27 03:00 +0100 |
| Subject | [PATCHv2 03/29] asm-generic: introduce __ARCH_USE_5LEVEL_HACK |
| Message-ID | <sSPdB-45l-47@gated-at.bofh.it> |
| In reply to | #1547445 |
We are going to introduce <asm-generic/pgtable-nop4d.h> to provide
abstraction for properly (in opposite to 5level-fixup.h hack) folded
p4d level. The new header will be included from pgtable-nopud.h.
If an architecture uses <asm-generic/nop*d.h>, we cannot use
5level-fixup.h directly to quickly convert the architecture to 5-level
paging as it would conflict with pgtable-nop4d.h.
With this patch an architecture can define __ARCH_USE_5LEVEL_HACK before
inclusion <asm-genenric/nop*d.h> to 5level-fixup.h.
Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
---
include/asm-generic/pgtable-nop4d-hack.h | 62 ++++++++++++++++++++++++++++++++
include/asm-generic/pgtable-nopud.h | 5 +++
2 files changed, 67 insertions(+)
create mode 100644 include/asm-generic/pgtable-nop4d-hack.h
diff --git a/include/asm-generic/pgtable-nop4d-hack.h b/include/asm-generic/pgtable-nop4d-hack.h
new file mode 100644
index 000000000000..752fb7511750
--- /dev/null
+++ b/include/asm-generic/pgtable-nop4d-hack.h
@@ -0,0 +1,62 @@
+#ifndef _PGTABLE_NOP4D_HACK_H
+#define _PGTABLE_NOP4D_HACK_H
+
+#ifndef __ASSEMBLY__
+#include <asm-generic/5level-fixup.h>
+
+#define __PAGETABLE_PUD_FOLDED
+
+/*
+ * Having the pud type consist of a pgd gets the size right, and allows
+ * us to conceptually access the pgd entry that this pud is folded into
+ * without casting.
+ */
+typedef struct { pgd_t pgd; } pud_t;
+
+#define PUD_SHIFT PGDIR_SHIFT
+#define PTRS_PER_PUD 1
+#define PUD_SIZE (1UL << PUD_SHIFT)
+#define PUD_MASK (~(PUD_SIZE-1))
+
+/*
+ * The "pgd_xxx()" functions here are trivial for a folded two-level
+ * setup: the pud is never bad, and a pud always exists (as it's folded
+ * into the pgd entry)
+ */
+static inline int pgd_none(pgd_t pgd) { return 0; }
+static inline int pgd_bad(pgd_t pgd) { return 0; }
+static inline int pgd_present(pgd_t pgd) { return 1; }
+static inline void pgd_clear(pgd_t *pgd) { }
+#define pud_ERROR(pud) (pgd_ERROR((pud).pgd))
+
+#define pgd_populate(mm, pgd, pud) do { } while (0)
+/*
+ * (puds are folded into pgds so this doesn't get actually called,
+ * but the define is needed for a generic inline function.)
+ */
+#define set_pgd(pgdptr, pgdval) set_pud((pud_t *)(pgdptr), (pud_t) { pgdval })
+
+static inline pud_t *pud_offset(pgd_t *pgd, unsigned long address)
+{
+ return (pud_t *)pgd;
+}
+
+#define pud_val(x) (pgd_val((x).pgd))
+#define __pud(x) ((pud_t) { __pgd(x) })
+
+#define pgd_page(pgd) (pud_page((pud_t){ pgd }))
+#define pgd_page_vaddr(pgd) (pud_page_vaddr((pud_t){ pgd }))
+
+/*
+ * allocating and freeing a pud is trivial: the 1-entry pud is
+ * inside the pgd, so has no extra memory associated with it.
+ */
+#define pud_alloc_one(mm, address) NULL
+#define pud_free(mm, x) do { } while (0)
+#define __pud_free_tlb(tlb, x, a) do { } while (0)
+
+#undef pud_addr_end
+#define pud_addr_end(addr, end) (end)
+
+#endif /* __ASSEMBLY__ */
+#endif /* _PGTABLE_NOP4D_HACK_H */
diff --git a/include/asm-generic/pgtable-nopud.h b/include/asm-generic/pgtable-nopud.h
index 810431d8351b..5e49430a30a4 100644
--- a/include/asm-generic/pgtable-nopud.h
+++ b/include/asm-generic/pgtable-nopud.h
@@ -3,6 +3,10 @@
#ifndef __ASSEMBLY__
+#ifdef __ARCH_USE_5LEVEL_HACK
+#include <asm-generic/pgtable-nop4d-hack.h>
+#else
+
#define __PAGETABLE_PUD_FOLDED
/*
@@ -58,4 +62,5 @@ static inline pud_t * pud_offset(pgd_t * pgd, unsigned long address)
#define pud_addr_end(addr, end) (end)
#endif /* __ASSEMBLY__ */
+#endif /* !__ARCH_USE_5LEVEL_HACK */
#endif /* _PGTABLE_NOPUD_H */
--
2.11.0
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> |
|---|---|
| Date | 2016-12-27 03:00 +0100 |
| Subject | [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sSPdB-45l-27@gated-at.bofh.it> |
| In reply to | #1547445 |
This patch introduces new rlimit resource to manage maximum virtual
address available to userspace to map.
On x86, 5-level paging enables 56-bit userspace virtual address space.
Not all user space is ready to handle wide addresses. It's known that
at least some JIT compilers use high bit in pointers to encode their
information. It collides with valid pointers with 5-level paging and
leads to crashes.
The patch aims to address this compatibility issue.
MM would use min(RLIMIT_VADDR, TASK_SIZE) as upper limit of virtual
address available to map by userspace.
The default hard limit will be RLIM_INFINITY, which basically means that
TASK_SIZE limits available address space.
The soft limit will also be RLIM_INFINITY everywhere, but the machine
with 5-level paging enabled. In this case, soft limit would be
(1UL << 47) - PAGE_SIZE. It’s current x86-64 TASK_SIZE_MAX with 4-level
paging which known to be safe
New rlimit resource would follow usual semantics with regards to
inheritance: preserved on fork(2) and exec(2). This has potential to
break application if limits set too wide or too narrow, but this is not
uncommon for other resources (consider RLIMIT_DATA or RLIMIT_AS).
As with other resources you can set the limit lower than current usage.
It would affect only future virtual address space allocations.
Use-cases for new rlimit:
- Bumping the soft limit to RLIM_INFINITY, allows current process all
its children to use addresses above 47-bits.
- Bumping the soft limit to RLIM_INFINITY after fork(2), but before
exec(2) allows the child to use addresses above 47-bits.
- Lowering the hard limit to 47-bits would prevent current process all
its children to use addresses above 47-bits, unless a process has
CAP_SYS_RESOURCES.
- It’s also can be handy to lower hard or soft limit to arbitrary
address. User-mode emulation in QEMU may lower the limit to 32-bit
to emulate 32-bit machine on 64-bit host.
TODO:
- port to non-x86;
Not-yet-signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Cc: linux-api@vger.kernel.org
---
arch/x86/include/asm/elf.h | 2 +-
arch/x86/include/asm/processor.h | 17 ++++++++++++-----
arch/x86/kernel/sys_x86_64.c | 6 +++---
arch/x86/mm/hugetlbpage.c | 8 ++++----
arch/x86/mm/mmap.c | 4 ++--
fs/binfmt_aout.c | 2 --
fs/binfmt_elf.c | 10 +++++-----
fs/hugetlbfs/inode.c | 6 +++---
fs/proc/base.c | 1 +
include/asm-generic/resource.h | 4 ++++
include/linux/sched.h | 5 +++++
include/uapi/asm-generic/resource.h | 3 ++-
kernel/events/uprobes.c | 5 +++--
kernel/sys.c | 6 +++---
mm/mmap.c | 20 +++++++++++---------
mm/mremap.c | 3 ++-
mm/nommu.c | 2 +-
mm/shmem.c | 8 ++++----
18 files changed, 66 insertions(+), 46 deletions(-)
diff --git a/arch/x86/include/asm/elf.h b/arch/x86/include/asm/elf.h
index e7f155c3045e..5ce6f2b2b105 100644
--- a/arch/x86/include/asm/elf.h
+++ b/arch/x86/include/asm/elf.h
@@ -250,7 +250,7 @@ extern int force_personality32;
the loader. We need to make sure that it is out of the way of the program
that it will "exec", and that there is sufficient room for the brk. */
-#define ELF_ET_DYN_BASE (TASK_SIZE / 3 * 2)
+#define ELF_ET_DYN_BASE (mmap_max_addr() / 3 * 2)
/* This yields a mask that user programs can use to figure out what
instruction set this CPU supports. This could be done in user space,
diff --git a/arch/x86/include/asm/processor.h b/arch/x86/include/asm/processor.h
index eaf100508c36..e02917126859 100644
--- a/arch/x86/include/asm/processor.h
+++ b/arch/x86/include/asm/processor.h
@@ -770,8 +770,8 @@ static inline void spin_lock_prefetch(const void *x)
*/
#define TASK_SIZE PAGE_OFFSET
#define TASK_SIZE_MAX TASK_SIZE
-#define STACK_TOP TASK_SIZE
-#define STACK_TOP_MAX STACK_TOP
+#define STACK_TOP mmap_max_addr()
+#define STACK_TOP_MAX TASK_SIZE
#define INIT_THREAD { \
.sp0 = TOP_OF_INIT_STACK, \
@@ -809,7 +809,14 @@ static inline void spin_lock_prefetch(const void *x)
* particular problem by preventing anything from being mapped
* at the maximum canonical address.
*/
-#define TASK_SIZE_MAX ((1UL << 47) - PAGE_SIZE)
+#define TASK_SIZE_MAX ((1UL << __VIRTUAL_MASK_SHIFT) - PAGE_SIZE)
+
+/*
+ * Default limit on maximum virtual address. This is required for
+ * compatibility with applications that assumes 47-bit VA.
+ * The limit be overrided with setrlimit(2).
+ */
+#define USER_VADDR_LIM ((1UL << 47) - PAGE_SIZE)
/* This decides where the kernel will search for a free chunk of vm
* space during mmap's.
@@ -822,7 +829,7 @@ static inline void spin_lock_prefetch(const void *x)
#define TASK_SIZE_OF(child) ((test_tsk_thread_flag(child, TIF_ADDR32)) ? \
IA32_PAGE_OFFSET : TASK_SIZE_MAX)
-#define STACK_TOP TASK_SIZE
+#define STACK_TOP mmap_max_addr()
#define STACK_TOP_MAX TASK_SIZE_MAX
#define INIT_THREAD { \
@@ -844,7 +851,7 @@ extern void start_thread(struct pt_regs *regs, unsigned long new_ip,
* This decides where the kernel will search for a free chunk of vm
* space during mmap's.
*/
-#define TASK_UNMAPPED_BASE (PAGE_ALIGN(TASK_SIZE / 3))
+#define TASK_UNMAPPED_BASE (PAGE_ALIGN(mmap_max_addr() / 3))
#define KSTK_EIP(task) (task_pt_regs(task)->ip)
diff --git a/arch/x86/kernel/sys_x86_64.c b/arch/x86/kernel/sys_x86_64.c
index a55ed63b9f91..e31f5b0c5468 100644
--- a/arch/x86/kernel/sys_x86_64.c
+++ b/arch/x86/kernel/sys_x86_64.c
@@ -115,7 +115,7 @@ static void find_start_end(unsigned long flags, unsigned long *begin,
}
} else {
*begin = current->mm->mmap_legacy_base;
- *end = TASK_SIZE;
+ *end = mmap_max_addr();
}
}
@@ -168,7 +168,7 @@ arch_get_unmapped_area_topdown(struct file *filp, const unsigned long addr0,
struct vm_unmapped_area_info info;
/* requested length too big for entire address space */
- if (len > TASK_SIZE)
+ if (len > mmap_max_addr())
return -ENOMEM;
if (flags & MAP_FIXED)
@@ -182,7 +182,7 @@ arch_get_unmapped_area_topdown(struct file *filp, const unsigned long addr0,
if (addr) {
addr = PAGE_ALIGN(addr);
vma = find_vma(mm, addr);
- if (TASK_SIZE - len >= addr &&
+ if (mmap_max_addr() - len >= addr &&
(!vma || addr + len <= vma->vm_start))
return addr;
}
diff --git a/arch/x86/mm/hugetlbpage.c b/arch/x86/mm/hugetlbpage.c
index 2ae8584b44c7..b55b04b82097 100644
--- a/arch/x86/mm/hugetlbpage.c
+++ b/arch/x86/mm/hugetlbpage.c
@@ -82,7 +82,7 @@ static unsigned long hugetlb_get_unmapped_area_bottomup(struct file *file,
info.flags = 0;
info.length = len;
info.low_limit = current->mm->mmap_legacy_base;
- info.high_limit = TASK_SIZE;
+ info.high_limit = mmap_max_addr();
info.align_mask = PAGE_MASK & ~huge_page_mask(h);
info.align_offset = 0;
return vm_unmapped_area(&info);
@@ -114,7 +114,7 @@ static unsigned long hugetlb_get_unmapped_area_topdown(struct file *file,
VM_BUG_ON(addr != -ENOMEM);
info.flags = 0;
info.low_limit = TASK_UNMAPPED_BASE;
- info.high_limit = TASK_SIZE;
+ info.high_limit = mmap_max_addr();
addr = vm_unmapped_area(&info);
}
@@ -131,7 +131,7 @@ hugetlb_get_unmapped_area(struct file *file, unsigned long addr,
if (len & ~huge_page_mask(h))
return -EINVAL;
- if (len > TASK_SIZE)
+ if (len > mmap_max_addr())
return -ENOMEM;
if (flags & MAP_FIXED) {
@@ -143,7 +143,7 @@ hugetlb_get_unmapped_area(struct file *file, unsigned long addr,
if (addr) {
addr = ALIGN(addr, huge_page_size(h));
vma = find_vma(mm, addr);
- if (TASK_SIZE - len >= addr &&
+ if (mmap_max_addr() - len >= addr &&
(!vma || addr + len <= vma->vm_start))
return addr;
}
diff --git a/arch/x86/mm/mmap.c b/arch/x86/mm/mmap.c
index d2dc0438d654..c22f0b802576 100644
--- a/arch/x86/mm/mmap.c
+++ b/arch/x86/mm/mmap.c
@@ -52,7 +52,7 @@ static unsigned long stack_maxrandom_size(void)
* Leave an at least ~128 MB hole with possible stack randomization.
*/
#define MIN_GAP (128*1024*1024UL + stack_maxrandom_size())
-#define MAX_GAP (TASK_SIZE/6*5)
+#define MAX_GAP (mmap_max_addr()/6*5)
static int mmap_is_legacy(void)
{
@@ -90,7 +90,7 @@ static unsigned long mmap_base(unsigned long rnd)
else if (gap > MAX_GAP)
gap = MAX_GAP;
- return PAGE_ALIGN(TASK_SIZE - gap - rnd);
+ return PAGE_ALIGN(mmap_max_addr() - gap - rnd);
}
/*
diff --git a/fs/binfmt_aout.c b/fs/binfmt_aout.c
index 2a59139f520b..7a7f6dba6b00 100644
--- a/fs/binfmt_aout.c
+++ b/fs/binfmt_aout.c
@@ -121,8 +121,6 @@ static struct linux_binfmt aout_format = {
.min_coredump = PAGE_SIZE
};
-#define BAD_ADDR(x) ((unsigned long)(x) >= TASK_SIZE)
-
static int set_brk(unsigned long start, unsigned long end)
{
start = PAGE_ALIGN(start);
diff --git a/fs/binfmt_elf.c b/fs/binfmt_elf.c
index 29a02daf08a9..1f8034aed298 100644
--- a/fs/binfmt_elf.c
+++ b/fs/binfmt_elf.c
@@ -89,7 +89,7 @@ static struct linux_binfmt elf_format = {
.min_coredump = ELF_EXEC_PAGESIZE,
};
-#define BAD_ADDR(x) ((unsigned long)(x) >= TASK_SIZE)
+#define BAD_ADDR(x) ((unsigned long)(x) >= mmap_max_addr())
static int set_brk(unsigned long start, unsigned long end)
{
@@ -587,8 +587,8 @@ static unsigned long load_elf_interp(struct elfhdr *interp_elf_ex,
k = load_addr + eppnt->p_vaddr;
if (BAD_ADDR(k) ||
eppnt->p_filesz > eppnt->p_memsz ||
- eppnt->p_memsz > TASK_SIZE ||
- TASK_SIZE - eppnt->p_memsz < k) {
+ eppnt->p_memsz > mmap_max_addr() ||
+ mmap_max_addr() - eppnt->p_memsz < k) {
error = -ENOMEM;
goto out;
}
@@ -960,8 +960,8 @@ static int load_elf_binary(struct linux_binprm *bprm)
* <= p_memsz so it is only necessary to check p_memsz.
*/
if (BAD_ADDR(k) || elf_ppnt->p_filesz > elf_ppnt->p_memsz ||
- elf_ppnt->p_memsz > TASK_SIZE ||
- TASK_SIZE - elf_ppnt->p_memsz < k) {
+ elf_ppnt->p_memsz > mmap_max_addr() ||
+ mmap_max_addr() - elf_ppnt->p_memsz < k) {
/* set_brk can never work. Avoid overflows. */
retval = -EINVAL;
goto out_free_dentry;
diff --git a/fs/hugetlbfs/inode.c b/fs/hugetlbfs/inode.c
index 54de77e78775..e132e93b85fb 100644
--- a/fs/hugetlbfs/inode.c
+++ b/fs/hugetlbfs/inode.c
@@ -178,7 +178,7 @@ hugetlb_get_unmapped_area(struct file *file, unsigned long addr,
if (len & ~huge_page_mask(h))
return -EINVAL;
- if (len > TASK_SIZE)
+ if (len > mmap_max_addr())
return -ENOMEM;
if (flags & MAP_FIXED) {
@@ -190,7 +190,7 @@ hugetlb_get_unmapped_area(struct file *file, unsigned long addr,
if (addr) {
addr = ALIGN(addr, huge_page_size(h));
vma = find_vma(mm, addr);
- if (TASK_SIZE - len >= addr &&
+ if (mmap_max_addr() - len >= addr &&
(!vma || addr + len <= vma->vm_start))
return addr;
}
@@ -198,7 +198,7 @@ hugetlb_get_unmapped_area(struct file *file, unsigned long addr,
info.flags = 0;
info.length = len;
info.low_limit = TASK_UNMAPPED_BASE;
- info.high_limit = TASK_SIZE;
+ info.high_limit = mmap_max_addr();
info.align_mask = PAGE_MASK & ~huge_page_mask(h);
info.align_offset = 0;
return vm_unmapped_area(&info);
diff --git a/fs/proc/base.c b/fs/proc/base.c
index 8e7e61b28f31..b91247cd171d 100644
--- a/fs/proc/base.c
+++ b/fs/proc/base.c
@@ -594,6 +594,7 @@ static const struct limit_names lnames[RLIM_NLIMITS] = {
[RLIMIT_NICE] = {"Max nice priority", NULL},
[RLIMIT_RTPRIO] = {"Max realtime priority", NULL},
[RLIMIT_RTTIME] = {"Max realtime timeout", "us"},
+ [RLIMIT_VADDR] = {"Max virtual address", NULL},
};
/* Display limits for a process */
diff --git a/include/asm-generic/resource.h b/include/asm-generic/resource.h
index 5e752b959054..d24c978103e5 100644
--- a/include/asm-generic/resource.h
+++ b/include/asm-generic/resource.h
@@ -3,6 +3,9 @@
#include <uapi/asm-generic/resource.h>
+#ifndef USER_VADDR_LIM
+#define USER_VADDR_LIM RLIM_INFINITY
+#endif
/*
* boot-time rlimit defaults for the init task:
@@ -25,6 +28,7 @@
[RLIMIT_NICE] = { 0, 0 }, \
[RLIMIT_RTPRIO] = { 0, 0 }, \
[RLIMIT_RTTIME] = { RLIM_INFINITY, RLIM_INFINITY }, \
+ [RLIMIT_VADDR] = { USER_VADDR_LIM, RLIM_INFINITY }, \
}
#endif
diff --git a/include/linux/sched.h b/include/linux/sched.h
index 4d1905245c7a..f0f23afe0838 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -3661,4 +3661,9 @@ void cpufreq_add_update_util_hook(int cpu, struct update_util_data *data,
void cpufreq_remove_update_util_hook(int cpu);
#endif /* CONFIG_CPU_FREQ */
+static inline unsigned long mmap_max_addr(void)
+{
+ return min(TASK_SIZE, rlimit(RLIMIT_VADDR));
+}
+
#endif
diff --git a/include/uapi/asm-generic/resource.h b/include/uapi/asm-generic/resource.h
index c6d10af50123..7843ed0ed7a7 100644
--- a/include/uapi/asm-generic/resource.h
+++ b/include/uapi/asm-generic/resource.h
@@ -45,7 +45,8 @@
0-39 for nice level 19 .. -20 */
#define RLIMIT_RTPRIO 14 /* maximum realtime priority */
#define RLIMIT_RTTIME 15 /* timeout for RT tasks in us */
-#define RLIM_NLIMITS 16
+#define RLIMIT_VADDR 16 /* maximum virtual address */
+#define RLIM_NLIMITS 17
/*
* SuS says limits have to be unsigned.
diff --git a/kernel/events/uprobes.c b/kernel/events/uprobes.c
index d416f3baf392..651f571a1a79 100644
--- a/kernel/events/uprobes.c
+++ b/kernel/events/uprobes.c
@@ -1142,8 +1142,9 @@ static int xol_add_vma(struct mm_struct *mm, struct xol_area *area)
if (!area->vaddr) {
/* Try to map as high as possible, this is only a hint. */
- area->vaddr = get_unmapped_area(NULL, TASK_SIZE - PAGE_SIZE,
- PAGE_SIZE, 0, 0);
+ area->vaddr = get_unmapped_area(NULL,
+ mmap_max_addr() - PAGE_SIZE,
+ PAGE_SIZE, 0, 0);
if (area->vaddr & ~PAGE_MASK) {
ret = area->vaddr;
goto fail;
diff --git a/kernel/sys.c b/kernel/sys.c
index 842914ef7de4..a5ee7f23beda 100644
--- a/kernel/sys.c
+++ b/kernel/sys.c
@@ -1718,7 +1718,7 @@ static int prctl_set_mm_exe_file(struct mm_struct *mm, unsigned int fd)
*/
static int validate_prctl_map(struct prctl_mm_map *prctl_map)
{
- unsigned long mmap_max_addr = TASK_SIZE;
+ unsigned long max_addr = mmap_max_addr();
struct mm_struct *mm = current->mm;
int error = -EINVAL, i;
@@ -1743,7 +1743,7 @@ static int validate_prctl_map(struct prctl_mm_map *prctl_map)
for (i = 0; i < ARRAY_SIZE(offsets); i++) {
u64 val = *(u64 *)((char *)prctl_map + offsets[i]);
- if ((unsigned long)val >= mmap_max_addr ||
+ if ((unsigned long)val >= max_addr ||
(unsigned long)val < mmap_min_addr)
goto out;
}
@@ -1949,7 +1949,7 @@ static int prctl_set_mm(int opt, unsigned long addr,
if (opt == PR_SET_MM_AUXV)
return prctl_set_auxv(mm, addr, arg4);
- if (addr >= TASK_SIZE || addr < mmap_min_addr)
+ if (addr >= mmap_max_addr() || addr < mmap_min_addr)
return -EINVAL;
error = -EINVAL;
diff --git a/mm/mmap.c b/mm/mmap.c
index dc4291dcc99b..a3384f23359e 100644
--- a/mm/mmap.c
+++ b/mm/mmap.c
@@ -1966,7 +1966,7 @@ arch_get_unmapped_area(struct file *filp, unsigned long addr,
struct vm_area_struct *vma;
struct vm_unmapped_area_info info;
- if (len > TASK_SIZE - mmap_min_addr)
+ if (len > mmap_max_addr() - mmap_min_addr)
return -ENOMEM;
if (flags & MAP_FIXED)
@@ -1975,15 +1975,16 @@ arch_get_unmapped_area(struct file *filp, unsigned long addr,
if (addr) {
addr = PAGE_ALIGN(addr);
vma = find_vma(mm, addr);
- if (TASK_SIZE - len >= addr && addr >= mmap_min_addr &&
- (!vma || addr + len <= vma->vm_start))
+ if (mmap_max_addr() - len >= addr &&
+ addr >= mmap_min_addr &&
+ (!vma || addr + len <= vma->vm_start))
return addr;
}
info.flags = 0;
info.length = len;
info.low_limit = mm->mmap_base;
- info.high_limit = TASK_SIZE;
+ info.high_limit = mmap_max_addr();
info.align_mask = 0;
return vm_unmapped_area(&info);
}
@@ -2005,7 +2006,7 @@ arch_get_unmapped_area_topdown(struct file *filp, const unsigned long addr0,
struct vm_unmapped_area_info info;
/* requested length too big for entire address space */
- if (len > TASK_SIZE - mmap_min_addr)
+ if (len > mmap_max_addr() - mmap_min_addr)
return -ENOMEM;
if (flags & MAP_FIXED)
@@ -2015,7 +2016,8 @@ arch_get_unmapped_area_topdown(struct file *filp, const unsigned long addr0,
if (addr) {
addr = PAGE_ALIGN(addr);
vma = find_vma(mm, addr);
- if (TASK_SIZE - len >= addr && addr >= mmap_min_addr &&
+ if (mmap_max_addr() - len >= addr &&
+ addr >= mmap_min_addr &&
(!vma || addr + len <= vma->vm_start))
return addr;
}
@@ -2037,7 +2039,7 @@ arch_get_unmapped_area_topdown(struct file *filp, const unsigned long addr0,
VM_BUG_ON(addr != -ENOMEM);
info.flags = 0;
info.low_limit = TASK_UNMAPPED_BASE;
- info.high_limit = TASK_SIZE;
+ info.high_limit = mmap_max_addr();
addr = vm_unmapped_area(&info);
}
@@ -2057,7 +2059,7 @@ get_unmapped_area(struct file *file, unsigned long addr, unsigned long len,
return error;
/* Careful about overflows.. */
- if (len > TASK_SIZE)
+ if (len > mmap_max_addr())
return -ENOMEM;
get_area = current->mm->get_unmapped_area;
@@ -2078,7 +2080,7 @@ get_unmapped_area(struct file *file, unsigned long addr, unsigned long len,
if (IS_ERR_VALUE(addr))
return addr;
- if (addr > TASK_SIZE - len)
+ if (addr > mmap_max_addr() - len)
return -ENOMEM;
if (offset_in_page(addr))
return -EINVAL;
diff --git a/mm/mremap.c b/mm/mremap.c
index 2b3bfcd51c75..a8b4fba3dce6 100644
--- a/mm/mremap.c
+++ b/mm/mremap.c
@@ -433,7 +433,8 @@ static unsigned long mremap_to(unsigned long addr, unsigned long old_len,
if (offset_in_page(new_addr))
goto out;
- if (new_len > TASK_SIZE || new_addr > TASK_SIZE - new_len)
+ if (new_len > mmap_max_addr() ||
+ new_addr > mmap_max_addr() - new_len)
goto out;
/* Ensure the old/new locations do not overlap */
diff --git a/mm/nommu.c b/mm/nommu.c
index 24f9f5f39145..6043b8b82083 100644
--- a/mm/nommu.c
+++ b/mm/nommu.c
@@ -905,7 +905,7 @@ static int validate_mmap_request(struct file *file,
/* Careful about overflows.. */
rlen = PAGE_ALIGN(len);
- if (!rlen || rlen > TASK_SIZE)
+ if (!rlen || rlen > mmap_max_addr())
return -ENOMEM;
/* offset overflow? */
diff --git a/mm/shmem.c b/mm/shmem.c
index bb53285a1d99..3c9be716083f 100644
--- a/mm/shmem.c
+++ b/mm/shmem.c
@@ -1976,7 +1976,7 @@ unsigned long shmem_get_unmapped_area(struct file *file,
unsigned long inflated_addr;
unsigned long inflated_offset;
- if (len > TASK_SIZE)
+ if (len > mmap_max_addr())
return -ENOMEM;
get_area = current->mm->get_unmapped_area;
@@ -1988,7 +1988,7 @@ unsigned long shmem_get_unmapped_area(struct file *file,
return addr;
if (addr & ~PAGE_MASK)
return addr;
- if (addr > TASK_SIZE - len)
+ if (addr > mmap_max_addr() - len)
return addr;
if (shmem_huge == SHMEM_HUGE_DENY)
@@ -2031,7 +2031,7 @@ unsigned long shmem_get_unmapped_area(struct file *file,
return addr;
inflated_len = len + HPAGE_PMD_SIZE - PAGE_SIZE;
- if (inflated_len > TASK_SIZE)
+ if (inflated_len > mmap_max_addr())
return addr;
if (inflated_len < len)
return addr;
@@ -2047,7 +2047,7 @@ unsigned long shmem_get_unmapped_area(struct file *file,
if (inflated_offset > offset)
inflated_addr += HPAGE_PMD_SIZE;
- if (inflated_addr > TASK_SIZE - len)
+ if (inflated_addr > mmap_max_addr() - len)
return addr;
return inflated_addr;
}
--
2.11.0
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-12-27 03:20 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sSPwR-4qF-9@gated-at.bofh.it> |
| In reply to | #1547448 |
On Mon, Dec 26, 2016 at 5:54 PM, Kirill A. Shutemov <kirill.shutemov@linux.intel.com> wrote: > This patch introduces new rlimit resource to manage maximum virtual > address available to userspace to map. > > On x86, 5-level paging enables 56-bit userspace virtual address space. > Not all user space is ready to handle wide addresses. It's known that > at least some JIT compilers use high bit in pointers to encode their > information. It collides with valid pointers with 5-level paging and > leads to crashes. > > The patch aims to address this compatibility issue. > > MM would use min(RLIMIT_VADDR, TASK_SIZE) as upper limit of virtual > address available to map by userspace. > > The default hard limit will be RLIM_INFINITY, which basically means that > TASK_SIZE limits available address space. > > The soft limit will also be RLIM_INFINITY everywhere, but the machine > with 5-level paging enabled. In this case, soft limit would be > (1UL << 47) - PAGE_SIZE. It’s current x86-64 TASK_SIZE_MAX with 4-level > paging which known to be safe > > New rlimit resource would follow usual semantics with regards to > inheritance: preserved on fork(2) and exec(2). This has potential to > break application if limits set too wide or too narrow, but this is not > uncommon for other resources (consider RLIMIT_DATA or RLIMIT_AS). > > As with other resources you can set the limit lower than current usage. > It would affect only future virtual address space allocations. > > Use-cases for new rlimit: > > - Bumping the soft limit to RLIM_INFINITY, allows current process all > its children to use addresses above 47-bits. > > - Bumping the soft limit to RLIM_INFINITY after fork(2), but before > exec(2) allows the child to use addresses above 47-bits. > > - Lowering the hard limit to 47-bits would prevent current process all > its children to use addresses above 47-bits, unless a process has > CAP_SYS_RESOURCES. > > - It’s also can be handy to lower hard or soft limit to arbitrary > address. User-mode emulation in QEMU may lower the limit to 32-bit > to emulate 32-bit machine on 64-bit host. I tend to think that this should be a personality or an ELF flag, not an rlimit. That way setuid works right.
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2016-12-27 03:40 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sSPQd-4wp-3@gated-at.bofh.it> |
| In reply to | #1547464 |
On Mon, Dec 26, 2016 at 06:06:01PM -0800, Andy Lutomirski wrote: > On Mon, Dec 26, 2016 at 5:54 PM, Kirill A. Shutemov > <kirill.shutemov@linux.intel.com> wrote: > > This patch introduces new rlimit resource to manage maximum virtual > > address available to userspace to map. > > > > On x86, 5-level paging enables 56-bit userspace virtual address space. > > Not all user space is ready to handle wide addresses. It's known that > > at least some JIT compilers use high bit in pointers to encode their > > information. It collides with valid pointers with 5-level paging and > > leads to crashes. > > > > The patch aims to address this compatibility issue. > > > > MM would use min(RLIMIT_VADDR, TASK_SIZE) as upper limit of virtual > > address available to map by userspace. > > > > The default hard limit will be RLIM_INFINITY, which basically means that > > TASK_SIZE limits available address space. > > > > The soft limit will also be RLIM_INFINITY everywhere, but the machine > > with 5-level paging enabled. In this case, soft limit would be > > (1UL << 47) - PAGE_SIZE. It’s current x86-64 TASK_SIZE_MAX with 4-level > > paging which known to be safe > > > > New rlimit resource would follow usual semantics with regards to > > inheritance: preserved on fork(2) and exec(2). This has potential to > > break application if limits set too wide or too narrow, but this is not > > uncommon for other resources (consider RLIMIT_DATA or RLIMIT_AS). > > > > As with other resources you can set the limit lower than current usage. > > It would affect only future virtual address space allocations. > > > > Use-cases for new rlimit: > > > > - Bumping the soft limit to RLIM_INFINITY, allows current process all > > its children to use addresses above 47-bits. > > > > - Bumping the soft limit to RLIM_INFINITY after fork(2), but before > > exec(2) allows the child to use addresses above 47-bits. > > > > - Lowering the hard limit to 47-bits would prevent current process all > > its children to use addresses above 47-bits, unless a process has > > CAP_SYS_RESOURCES. > > > > - It’s also can be handy to lower hard or soft limit to arbitrary > > address. User-mode emulation in QEMU may lower the limit to 32-bit > > to emulate 32-bit machine on 64-bit host. > > I tend to think that this should be a personality or an ELF flag, not > an rlimit. My plan was to implement ELF flag on top. Basically, ELF flag would mean that we bump soft limit to hard limit on exec. > That way setuid works right. Um.. I probably miss background here. -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-12-27 04:30 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sSQCB-596-1@gated-at.bofh.it> |
| In reply to | #1547466 |
On Mon, Dec 26, 2016 at 6:24 PM, Kirill A. Shutemov <kirill@shutemov.name> wrote: > On Mon, Dec 26, 2016 at 06:06:01PM -0800, Andy Lutomirski wrote: >> On Mon, Dec 26, 2016 at 5:54 PM, Kirill A. Shutemov >> <kirill.shutemov@linux.intel.com> wrote: >> > This patch introduces new rlimit resource to manage maximum virtual >> > address available to userspace to map. >> > >> > On x86, 5-level paging enables 56-bit userspace virtual address space. >> > Not all user space is ready to handle wide addresses. It's known that >> > at least some JIT compilers use high bit in pointers to encode their >> > information. It collides with valid pointers with 5-level paging and >> > leads to crashes. >> > >> > The patch aims to address this compatibility issue. >> > >> > MM would use min(RLIMIT_VADDR, TASK_SIZE) as upper limit of virtual >> > address available to map by userspace. >> > >> > The default hard limit will be RLIM_INFINITY, which basically means that >> > TASK_SIZE limits available address space. >> > >> > The soft limit will also be RLIM_INFINITY everywhere, but the machine >> > with 5-level paging enabled. In this case, soft limit would be >> > (1UL << 47) - PAGE_SIZE. It’s current x86-64 TASK_SIZE_MAX with 4-level >> > paging which known to be safe >> > >> > New rlimit resource would follow usual semantics with regards to >> > inheritance: preserved on fork(2) and exec(2). This has potential to >> > break application if limits set too wide or too narrow, but this is not >> > uncommon for other resources (consider RLIMIT_DATA or RLIMIT_AS). >> > >> > As with other resources you can set the limit lower than current usage. >> > It would affect only future virtual address space allocations. >> > >> > Use-cases for new rlimit: >> > >> > - Bumping the soft limit to RLIM_INFINITY, allows current process all >> > its children to use addresses above 47-bits. >> > >> > - Bumping the soft limit to RLIM_INFINITY after fork(2), but before >> > exec(2) allows the child to use addresses above 47-bits. >> > >> > - Lowering the hard limit to 47-bits would prevent current process all >> > its children to use addresses above 47-bits, unless a process has >> > CAP_SYS_RESOURCES. >> > >> > - It’s also can be handy to lower hard or soft limit to arbitrary >> > address. User-mode emulation in QEMU may lower the limit to 32-bit >> > to emulate 32-bit machine on 64-bit host. >> >> I tend to think that this should be a personality or an ELF flag, not >> an rlimit. > > My plan was to implement ELF flag on top. Basically, ELF flag would mean > that we bump soft limit to hard limit on exec. > >> That way setuid works right. > > Um.. I probably miss background here. > If a setuid program depends on the lower limit, then a malicious program shouldn't be able to cause it to run with the higher limit. The personality code should already get this case right because personalities are reset when setuid happens. --Andy
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2017-01-02 10:20 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sV6WB-4G8-5@gated-at.bofh.it> |
| In reply to | #1547477 |
On Mon, Dec 26, 2016 at 07:22:03PM -0800, Andy Lutomirski wrote: > On Mon, Dec 26, 2016 at 6:24 PM, Kirill A. Shutemov > <kirill@shutemov.name> wrote: > > On Mon, Dec 26, 2016 at 06:06:01PM -0800, Andy Lutomirski wrote: > >> On Mon, Dec 26, 2016 at 5:54 PM, Kirill A. Shutemov > >> <kirill.shutemov@linux.intel.com> wrote: > >> > This patch introduces new rlimit resource to manage maximum virtual > >> > address available to userspace to map. > >> > > >> > On x86, 5-level paging enables 56-bit userspace virtual address space. > >> > Not all user space is ready to handle wide addresses. It's known that > >> > at least some JIT compilers use high bit in pointers to encode their > >> > information. It collides with valid pointers with 5-level paging and > >> > leads to crashes. > >> > > >> > The patch aims to address this compatibility issue. > >> > > >> > MM would use min(RLIMIT_VADDR, TASK_SIZE) as upper limit of virtual > >> > address available to map by userspace. > >> > > >> > The default hard limit will be RLIM_INFINITY, which basically means that > >> > TASK_SIZE limits available address space. > >> > > >> > The soft limit will also be RLIM_INFINITY everywhere, but the machine > >> > with 5-level paging enabled. In this case, soft limit would be > >> > (1UL << 47) - PAGE_SIZE. It’s current x86-64 TASK_SIZE_MAX with 4-level > >> > paging which known to be safe > >> > > >> > New rlimit resource would follow usual semantics with regards to > >> > inheritance: preserved on fork(2) and exec(2). This has potential to > >> > break application if limits set too wide or too narrow, but this is not > >> > uncommon for other resources (consider RLIMIT_DATA or RLIMIT_AS). > >> > > >> > As with other resources you can set the limit lower than current usage. > >> > It would affect only future virtual address space allocations. > >> > > >> > Use-cases for new rlimit: > >> > > >> > - Bumping the soft limit to RLIM_INFINITY, allows current process all > >> > its children to use addresses above 47-bits. > >> > > >> > - Bumping the soft limit to RLIM_INFINITY after fork(2), but before > >> > exec(2) allows the child to use addresses above 47-bits. > >> > > >> > - Lowering the hard limit to 47-bits would prevent current process all > >> > its children to use addresses above 47-bits, unless a process has > >> > CAP_SYS_RESOURCES. > >> > > >> > - It’s also can be handy to lower hard or soft limit to arbitrary > >> > address. User-mode emulation in QEMU may lower the limit to 32-bit > >> > to emulate 32-bit machine on 64-bit host. > >> > >> I tend to think that this should be a personality or an ELF flag, not > >> an rlimit. > > > > My plan was to implement ELF flag on top. Basically, ELF flag would mean > > that we bump soft limit to hard limit on exec. > > > >> That way setuid works right. > > > > Um.. I probably miss background here. > > > > If a setuid program depends on the lower limit, then a malicious > program shouldn't be able to cause it to run with the higher limit. > The personality code should already get this case right because > personalities are reset when setuid happens. It would be nice to have more fine-grained control than binary personality flag gives. It would cover more use-cases. Well, we could reset the limit on exec of setuid binary too. That's not ideal, but... -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Carlos O'Donell <carlos@redhat.com> |
|---|---|
| Date | 2016-12-29 04:00 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sTz6G-yi-9@gated-at.bofh.it> |
| In reply to | #1547466 |
On 12/26/2016 09:24 PM, Kirill A. Shutemov wrote: > On Mon, Dec 26, 2016 at 06:06:01PM -0800, Andy Lutomirski wrote: >> On Mon, Dec 26, 2016 at 5:54 PM, Kirill A. Shutemov >> <kirill.shutemov@linux.intel.com> wrote: >>> This patch introduces new rlimit resource to manage maximum virtual >>> address available to userspace to map. >>> >>> On x86, 5-level paging enables 56-bit userspace virtual address space. >>> Not all user space is ready to handle wide addresses. It's known that >>> at least some JIT compilers use high bit in pointers to encode their >>> information. It collides with valid pointers with 5-level paging and >>> leads to crashes. >>> >>> The patch aims to address this compatibility issue. >>> >>> MM would use min(RLIMIT_VADDR, TASK_SIZE) as upper limit of virtual >>> address available to map by userspace. >>> >>> The default hard limit will be RLIM_INFINITY, which basically means that >>> TASK_SIZE limits available address space. >>> >>> The soft limit will also be RLIM_INFINITY everywhere, but the machine >>> with 5-level paging enabled. In this case, soft limit would be >>> (1UL << 47) - PAGE_SIZE. It’s current x86-64 TASK_SIZE_MAX with 4-level >>> paging which known to be safe >>> >>> New rlimit resource would follow usual semantics with regards to >>> inheritance: preserved on fork(2) and exec(2). This has potential to >>> break application if limits set too wide or too narrow, but this is not >>> uncommon for other resources (consider RLIMIT_DATA or RLIMIT_AS). >>> >>> As with other resources you can set the limit lower than current usage. >>> It would affect only future virtual address space allocations. >>> >>> Use-cases for new rlimit: >>> >>> - Bumping the soft limit to RLIM_INFINITY, allows current process all >>> its children to use addresses above 47-bits. >>> >>> - Bumping the soft limit to RLIM_INFINITY after fork(2), but before >>> exec(2) allows the child to use addresses above 47-bits. >>> >>> - Lowering the hard limit to 47-bits would prevent current process all >>> its children to use addresses above 47-bits, unless a process has >>> CAP_SYS_RESOURCES. >>> >>> - It’s also can be handy to lower hard or soft limit to arbitrary >>> address. User-mode emulation in QEMU may lower the limit to 32-bit >>> to emulate 32-bit machine on 64-bit host. >> >> I tend to think that this should be a personality or an ELF flag, not >> an rlimit. > > My plan was to implement ELF flag on top. Basically, ELF flag would mean > that we bump soft limit to hard limit on exec. Could you clarify what you mean by an "ELF flag?" -- Cheers, Carlos.
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-12-31 03:10 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sUhho-57Y-5@gated-at.bofh.it> |
| In reply to | #1548065 |
On Wed, Dec 28, 2016 at 6:53 PM, Carlos O'Donell <carlos@redhat.com> wrote: > On 12/26/2016 09:24 PM, Kirill A. Shutemov wrote: >> On Mon, Dec 26, 2016 at 06:06:01PM -0800, Andy Lutomirski wrote: >>> On Mon, Dec 26, 2016 at 5:54 PM, Kirill A. Shutemov >>> <kirill.shutemov@linux.intel.com> wrote: >>>> This patch introduces new rlimit resource to manage maximum virtual >>>> address available to userspace to map. >>>> >>>> On x86, 5-level paging enables 56-bit userspace virtual address space. >>>> Not all user space is ready to handle wide addresses. It's known that >>>> at least some JIT compilers use high bit in pointers to encode their >>>> information. It collides with valid pointers with 5-level paging and >>>> leads to crashes. >>>> >>>> The patch aims to address this compatibility issue. >>>> >>>> MM would use min(RLIMIT_VADDR, TASK_SIZE) as upper limit of virtual >>>> address available to map by userspace. >>>> >>>> The default hard limit will be RLIM_INFINITY, which basically means that >>>> TASK_SIZE limits available address space. >>>> >>>> The soft limit will also be RLIM_INFINITY everywhere, but the machine >>>> with 5-level paging enabled. In this case, soft limit would be >>>> (1UL << 47) - PAGE_SIZE. It’s current x86-64 TASK_SIZE_MAX with 4-level >>>> paging which known to be safe >>>> >>>> New rlimit resource would follow usual semantics with regards to >>>> inheritance: preserved on fork(2) and exec(2). This has potential to >>>> break application if limits set too wide or too narrow, but this is not >>>> uncommon for other resources (consider RLIMIT_DATA or RLIMIT_AS). >>>> >>>> As with other resources you can set the limit lower than current usage. >>>> It would affect only future virtual address space allocations. >>>> >>>> Use-cases for new rlimit: >>>> >>>> - Bumping the soft limit to RLIM_INFINITY, allows current process all >>>> its children to use addresses above 47-bits. >>>> >>>> - Bumping the soft limit to RLIM_INFINITY after fork(2), but before >>>> exec(2) allows the child to use addresses above 47-bits. >>>> >>>> - Lowering the hard limit to 47-bits would prevent current process all >>>> its children to use addresses above 47-bits, unless a process has >>>> CAP_SYS_RESOURCES. >>>> >>>> - It’s also can be handy to lower hard or soft limit to arbitrary >>>> address. User-mode emulation in QEMU may lower the limit to 32-bit >>>> to emulate 32-bit machine on 64-bit host. >>> >>> I tend to think that this should be a personality or an ELF flag, not >>> an rlimit. >> >> My plan was to implement ELF flag on top. Basically, ELF flag would mean >> that we bump soft limit to hard limit on exec. > > Could you clarify what you mean by an "ELF flag?" Some way to mark a binary as supporting a larger address space. I don't have a precise solution in mind, but an ELF note might be a good way to go here. --Andy
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2017-01-02 09:40 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sV6jU-4bE-1@gated-at.bofh.it> |
| In reply to | #1548784 |
On Fri, Dec 30, 2016 at 06:08:27PM -0800, Andy Lutomirski wrote: > On Wed, Dec 28, 2016 at 6:53 PM, Carlos O'Donell <carlos@redhat.com> wrote: > > On 12/26/2016 09:24 PM, Kirill A. Shutemov wrote: > >> On Mon, Dec 26, 2016 at 06:06:01PM -0800, Andy Lutomirski wrote: > >>> On Mon, Dec 26, 2016 at 5:54 PM, Kirill A. Shutemov > >>> <kirill.shutemov@linux.intel.com> wrote: > >>>> This patch introduces new rlimit resource to manage maximum virtual > >>>> address available to userspace to map. > >>>> > >>>> On x86, 5-level paging enables 56-bit userspace virtual address space. > >>>> Not all user space is ready to handle wide addresses. It's known that > >>>> at least some JIT compilers use high bit in pointers to encode their > >>>> information. It collides with valid pointers with 5-level paging and > >>>> leads to crashes. > >>>> > >>>> The patch aims to address this compatibility issue. > >>>> > >>>> MM would use min(RLIMIT_VADDR, TASK_SIZE) as upper limit of virtual > >>>> address available to map by userspace. > >>>> > >>>> The default hard limit will be RLIM_INFINITY, which basically means that > >>>> TASK_SIZE limits available address space. > >>>> > >>>> The soft limit will also be RLIM_INFINITY everywhere, but the machine > >>>> with 5-level paging enabled. In this case, soft limit would be > >>>> (1UL << 47) - PAGE_SIZE. It’s current x86-64 TASK_SIZE_MAX with 4-level > >>>> paging which known to be safe > >>>> > >>>> New rlimit resource would follow usual semantics with regards to > >>>> inheritance: preserved on fork(2) and exec(2). This has potential to > >>>> break application if limits set too wide or too narrow, but this is not > >>>> uncommon for other resources (consider RLIMIT_DATA or RLIMIT_AS). > >>>> > >>>> As with other resources you can set the limit lower than current usage. > >>>> It would affect only future virtual address space allocations. > >>>> > >>>> Use-cases for new rlimit: > >>>> > >>>> - Bumping the soft limit to RLIM_INFINITY, allows current process all > >>>> its children to use addresses above 47-bits. > >>>> > >>>> - Bumping the soft limit to RLIM_INFINITY after fork(2), but before > >>>> exec(2) allows the child to use addresses above 47-bits. > >>>> > >>>> - Lowering the hard limit to 47-bits would prevent current process all > >>>> its children to use addresses above 47-bits, unless a process has > >>>> CAP_SYS_RESOURCES. > >>>> > >>>> - It’s also can be handy to lower hard or soft limit to arbitrary > >>>> address. User-mode emulation in QEMU may lower the limit to 32-bit > >>>> to emulate 32-bit machine on 64-bit host. > >>> > >>> I tend to think that this should be a personality or an ELF flag, not > >>> an rlimit. > >> > >> My plan was to implement ELF flag on top. Basically, ELF flag would mean > >> that we bump soft limit to hard limit on exec. > > > > Could you clarify what you mean by an "ELF flag?" > > Some way to mark a binary as supporting a larger address space. I > don't have a precise solution in mind, but an ELF note might be a good > way to go here. + H.J. There's discussion of proposal of "Program Properties"[1]. It seems fits the purpose. [1] https://sourceware.org/ml/gnu-gabi/2016-q4/msg00000.html -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | "H.J. Lu" <hjl.tools@gmail.com> |
|---|---|
| Date | 2017-01-13 21:20 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sZgul-5h6-5@gated-at.bofh.it> |
| In reply to | #1549023 |
On Mon, Jan 2, 2017 at 12:35 AM, Kirill A. Shutemov <kirill@shutemov.name> wrote: > On Fri, Dec 30, 2016 at 06:08:27PM -0800, Andy Lutomirski wrote: >> On Wed, Dec 28, 2016 at 6:53 PM, Carlos O'Donell <carlos@redhat.com> wrote: >> > On 12/26/2016 09:24 PM, Kirill A. Shutemov wrote: >> >> On Mon, Dec 26, 2016 at 06:06:01PM -0800, Andy Lutomirski wrote: >> >>> On Mon, Dec 26, 2016 at 5:54 PM, Kirill A. Shutemov >> >>> <kirill.shutemov@linux.intel.com> wrote: >> >>>> This patch introduces new rlimit resource to manage maximum virtual >> >>>> address available to userspace to map. >> >>>> >> >>>> On x86, 5-level paging enables 56-bit userspace virtual address space. >> >>>> Not all user space is ready to handle wide addresses. It's known that >> >>>> at least some JIT compilers use high bit in pointers to encode their >> >>>> information. It collides with valid pointers with 5-level paging and >> >>>> leads to crashes. >> >>>> >> >>>> The patch aims to address this compatibility issue. >> >>>> >> >>>> MM would use min(RLIMIT_VADDR, TASK_SIZE) as upper limit of virtual >> >>>> address available to map by userspace. >> >>>> >> >>>> The default hard limit will be RLIM_INFINITY, which basically means that >> >>>> TASK_SIZE limits available address space. >> >>>> >> >>>> The soft limit will also be RLIM_INFINITY everywhere, but the machine >> >>>> with 5-level paging enabled. In this case, soft limit would be >> >>>> (1UL << 47) - PAGE_SIZE. It’s current x86-64 TASK_SIZE_MAX with 4-level >> >>>> paging which known to be safe >> >>>> >> >>>> New rlimit resource would follow usual semantics with regards to >> >>>> inheritance: preserved on fork(2) and exec(2). This has potential to >> >>>> break application if limits set too wide or too narrow, but this is not >> >>>> uncommon for other resources (consider RLIMIT_DATA or RLIMIT_AS). >> >>>> >> >>>> As with other resources you can set the limit lower than current usage. >> >>>> It would affect only future virtual address space allocations. >> >>>> >> >>>> Use-cases for new rlimit: >> >>>> >> >>>> - Bumping the soft limit to RLIM_INFINITY, allows current process all >> >>>> its children to use addresses above 47-bits. >> >>>> >> >>>> - Bumping the soft limit to RLIM_INFINITY after fork(2), but before >> >>>> exec(2) allows the child to use addresses above 47-bits. >> >>>> >> >>>> - Lowering the hard limit to 47-bits would prevent current process all >> >>>> its children to use addresses above 47-bits, unless a process has >> >>>> CAP_SYS_RESOURCES. >> >>>> >> >>>> - It’s also can be handy to lower hard or soft limit to arbitrary >> >>>> address. User-mode emulation in QEMU may lower the limit to 32-bit >> >>>> to emulate 32-bit machine on 64-bit host. >> >>> >> >>> I tend to think that this should be a personality or an ELF flag, not >> >>> an rlimit. >> >> >> >> My plan was to implement ELF flag on top. Basically, ELF flag would mean >> >> that we bump soft limit to hard limit on exec. >> > >> > Could you clarify what you mean by an "ELF flag?" >> >> Some way to mark a binary as supporting a larger address space. I >> don't have a precise solution in mind, but an ELF note might be a good >> way to go here. > > + H.J. > > There's discussion of proposal of "Program Properties"[1]. It seems fits > the purpose. > > [1] https://sourceware.org/ml/gnu-gabi/2016-q4/msg00000.html > > -- > Kirill A. Shutemov There is another proposal: https://fedoraproject.org/wiki/Toolchain/Watermark#Markup_for_ELF_objects which covers much more than mine. -- H.J.
[toc] | [prev] | [next] | [standalone]
| From | Arnd Bergmann <arnd@arndb.de> |
|---|---|
| Date | 2017-01-02 09:50 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sV6tz-4f8-9@gated-at.bofh.it> |
| In reply to | #1547448 |
On Tuesday, December 27, 2016 4:54:13 AM CET Kirill A. Shutemov wrote: > This patch introduces new rlimit resource to manage maximum virtual > address available to userspace to map. > > On x86, 5-level paging enables 56-bit userspace virtual address space. > Not all user space is ready to handle wide addresses. It's known that > at least some JIT compilers use high bit in pointers to encode their > information. It collides with valid pointers with 5-level paging and > leads to crashes. > > The patch aims to address this compatibility issue. > > MM would use min(RLIMIT_VADDR, TASK_SIZE) as upper limit of virtual > address available to map by userspace. > > The default hard limit will be RLIM_INFINITY, which basically means that > TASK_SIZE limits available address space. > > The soft limit will also be RLIM_INFINITY everywhere, but the machine > with 5-level paging enabled. In this case, soft limit would be > (1UL << 47) - PAGE_SIZE. It’s current x86-64 TASK_SIZE_MAX with 4-level > paging which known to be safe > > New rlimit resource would follow usual semantics with regards to > inheritance: preserved on fork(2) and exec(2). This has potential to > break application if limits set too wide or too narrow, but this is not > uncommon for other resources (consider RLIMIT_DATA or RLIMIT_AS). > > As with other resources you can set the limit lower than current usage. > It would affect only future virtual address space allocations. > > Use-cases for new rlimit: > > - Bumping the soft limit to RLIM_INFINITY, allows current process all > its children to use addresses above 47-bits. > > - Bumping the soft limit to RLIM_INFINITY after fork(2), but before > exec(2) allows the child to use addresses above 47-bits. > > - Lowering the hard limit to 47-bits would prevent current process all > its children to use addresses above 47-bits, unless a process has > CAP_SYS_RESOURCES. > > - It’s also can be handy to lower hard or soft limit to arbitrary > address. User-mode emulation in QEMU may lower the limit to 32-bit > to emulate 32-bit machine on 64-bit host. > > TODO: > - port to non-x86; > > Not-yet-signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> > Cc: linux-api@vger.kernel.org This seems to nicely address the same problem on arm64, which has run into the same issue due to the various page table formats that can currently be chosen at compile time. I don't see how this interacts with the existing PER_LINUX32/PER_LINUX32_3GB personality flags, but I assume you have either already thought of that, or we can come up with a good way to define what happens when conflicting settings are applied. The two reasonable ways I can think of are to either use the minimum of the two limits, or to make the personality syscall set the soft rlimit and use whatever limit was last set. Arnd
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2017-01-03 07:10 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sVqsh-24D-3@gated-at.bofh.it> |
| In reply to | #1549031 |
On Mon, Jan 2, 2017 at 12:44 AM, Arnd Bergmann <arnd@arndb.de> wrote: > On Tuesday, December 27, 2016 4:54:13 AM CET Kirill A. Shutemov wrote: >> As with other resources you can set the limit lower than current usage. >> It would affect only future virtual address space allocations. I still don't buy all these use cases: >> >> Use-cases for new rlimit: >> >> - Bumping the soft limit to RLIM_INFINITY, allows current process all >> its children to use addresses above 47-bits. OK, I get this, but only as a workaround for programs that make assumptions about the address space and don't use some mechanism (to be designed?) to work correctly in spite of a larger address space. >> >> - Bumping the soft limit to RLIM_INFINITY after fork(2), but before >> exec(2) allows the child to use addresses above 47-bits. Ditto. >> >> - Lowering the hard limit to 47-bits would prevent current process all >> its children to use addresses above 47-bits, unless a process has >> CAP_SYS_RESOURCES. I've tried and I can't imagine any reason to do this. >> >> - It’s also can be handy to lower hard or soft limit to arbitrary >> address. User-mode emulation in QEMU may lower the limit to 32-bit >> to emulate 32-bit machine on 64-bit host. I don't understand. QEMU user-mode emulation intercepts all syscalls. What QEMU would *actually* want is a way to say "allocate me some memory with the high N bits clear". mmap-via-int80 on x86 should be fixed to do this, but a new syscall with an explicit parameter would work, as would a prctl changing the current limit. >> >> TODO: >> - port to non-x86; >> >> Not-yet-signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> >> Cc: linux-api@vger.kernel.org > > This seems to nicely address the same problem on arm64, which has > run into the same issue due to the various page table formats > that can currently be chosen at compile time. On further reflection, I think this has very little to do with paging formats except insofar as paging formats make us notice the problem. The issue is that user code wants to be able to assume an upper limit on an address, and it gets an upper limit right now that depends on architecture due to paging formats. But someone really might want to write a *portable* 64-bit program that allocates memory with the high 16 bits clear. So let's add such a mechanism directly. As a thought experiment, what if x86_64 simply never allocated "high" (above 2^47-1) addresses unless a new mmap-with-explicit-limit syscall were used? Old glibc would continue working. Old VMs would work. New programs that want to use ginormous mappings would have to use the new syscall. This would be totally stateless and would have no issues with CRIU. If necessary, we could also have a prctl that changes a "personality-like" limit that is in effect when the old mmap was used. I say "personality-like" because it would reset under exactly the same conditions that personality resets itself. Thoughts?
[toc] | [prev] | [next] | [standalone]
| From | Arnd Bergmann <arnd@arndb.de> |
|---|---|
| Date | 2017-01-03 14:30 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sVxk6-6WU-17@gated-at.bofh.it> |
| In reply to | #1549550 |
On Monday, January 2, 2017 10:08:28 PM CET Andy Lutomirski wrote:
>
> > This seems to nicely address the same problem on arm64, which has
> > run into the same issue due to the various page table formats
> > that can currently be chosen at compile time.
>
> On further reflection, I think this has very little to do with paging
> formats except insofar as paging formats make us notice the problem.
> The issue is that user code wants to be able to assume an upper limit
> on an address, and it gets an upper limit right now that depends on
> architecture due to paging formats. But someone really might want to
> write a *portable* 64-bit program that allocates memory with the high
> 16 bits clear. So let's add such a mechanism directly.
>
> As a thought experiment, what if x86_64 simply never allocated "high"
> (above 2^47-1) addresses unless a new mmap-with-explicit-limit syscall
> were used? Old glibc would continue working. Old VMs would work.
> New programs that want to use ginormous mappings would have to use the
> new syscall. This would be totally stateless and would have no issues
> with CRIU.
I can see this working well for the 47-bit addressing default, but
what about applications that actually rely on 39-bit addressing
(I'd have to double-check, but I think this was the limit that
people were most interested in for arm64)?
39 bits seems a little small to make that the default for everyone
who doesn't pass the extra flag. Having to pass another flag to
limit the addresses introduces other problems (e.g. mmap from
library call that doesn't pass that flag).
> If necessary, we could also have a prctl that changes a
> "personality-like" limit that is in effect when the old mmap was used.
> I say "personality-like" because it would reset under exactly the same
> conditions that personality resets itself.
For "personality-like", it would still have to interact
with the existing PER_LINUX32 and PER_LINUX32_3GB flags that
do the exact same thing, so actually using personality might
be better.
We still have a few bits in the personality arguments, and
we could combine them with the existing ADDR_LIMIT_3GB
and ADDR_LIMIT_32BIT flags that are mutually exclusive by
definition, such as
ADDR_LIMIT_32BIT = 0x0800000, /* existing */
ADDR_LIMIT_3GB = 0x8000000, /* existing */
ADDR_LIMIT_39BIT = 0x0010000, /* next free bit */
ADDR_LIMIT_42BIT = 0x8010000,
ADDR_LIMIT_47BIT = 0x0810000,
ADDR_LIMIT_48BIT = 0x8810000,
This would probably take only one or two personality bits for the
limits that are interesting in practice.
Arnd
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2017-01-03 19:40 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sVCa7-1SP-63@gated-at.bofh.it> |
| In reply to | #1549753 |
On Tue, Jan 3, 2017 at 5:18 AM, Arnd Bergmann <arnd@arndb.de> wrote: > On Monday, January 2, 2017 10:08:28 PM CET Andy Lutomirski wrote: >> >> > This seems to nicely address the same problem on arm64, which has >> > run into the same issue due to the various page table formats >> > that can currently be chosen at compile time. >> >> On further reflection, I think this has very little to do with paging >> formats except insofar as paging formats make us notice the problem. >> The issue is that user code wants to be able to assume an upper limit >> on an address, and it gets an upper limit right now that depends on >> architecture due to paging formats. But someone really might want to >> write a *portable* 64-bit program that allocates memory with the high >> 16 bits clear. So let's add such a mechanism directly. >> >> As a thought experiment, what if x86_64 simply never allocated "high" >> (above 2^47-1) addresses unless a new mmap-with-explicit-limit syscall >> were used? Old glibc would continue working. Old VMs would work. >> New programs that want to use ginormous mappings would have to use the >> new syscall. This would be totally stateless and would have no issues >> with CRIU. > > I can see this working well for the 47-bit addressing default, but > what about applications that actually rely on 39-bit addressing > (I'd have to double-check, but I think this was the limit that > people were most interested in for arm64)? > > 39 bits seems a little small to make that the default for everyone > who doesn't pass the extra flag. Having to pass another flag to > limit the addresses introduces other problems (e.g. mmap from > library call that doesn't pass that flag). That's a fair point. Maybe my straw man isn't so good. > >> If necessary, we could also have a prctl that changes a >> "personality-like" limit that is in effect when the old mmap was used. >> I say "personality-like" because it would reset under exactly the same >> conditions that personality resets itself. > > For "personality-like", it would still have to interact > with the existing PER_LINUX32 and PER_LINUX32_3GB flags that > do the exact same thing, so actually using personality might > be better. > > We still have a few bits in the personality arguments, and > we could combine them with the existing ADDR_LIMIT_3GB > and ADDR_LIMIT_32BIT flags that are mutually exclusive by > definition, such as > > ADDR_LIMIT_32BIT = 0x0800000, /* existing */ > ADDR_LIMIT_3GB = 0x8000000, /* existing */ > ADDR_LIMIT_39BIT = 0x0010000, /* next free bit */ > ADDR_LIMIT_42BIT = 0x8010000, > ADDR_LIMIT_47BIT = 0x0810000, > ADDR_LIMIT_48BIT = 0x8810000, > > This would probably take only one or two personality bits for the > limits that are interesting in practice. Hmm. What if we approached this a bit differently? We could add a single new personality bit ADDR_LIMIT_EXPLICIT. Setting this bit cause PER_LINUX32_3GB etc to be automatically cleared. When ADDR_LIMIT_EXPLICIT is in effect, prctl can set a 64-bit numeric limit. If ADDR_LIMIT_EXPLICIT is cleared, the prctl value stops being settable and reading it via prctl returns whatever is implied by the other personality bits. --Andy
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2017-01-03 23:10 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sVFrk-48v-31@gated-at.bofh.it> |
| In reply to | #1550082 |
On Tue, Jan 3, 2017 at 2:07 PM, Arnd Bergmann <arnd@arndb.de> wrote: > On Tuesday, January 3, 2017 10:29:33 AM CET Andy Lutomirski wrote: >> >> Hmm. What if we approached this a bit differently? We could add a >> single new personality bit ADDR_LIMIT_EXPLICIT. Setting this bit >> cause PER_LINUX32_3GB etc to be automatically cleared. > > Both the ADDR_LIMIT_32BIT and ADDR_LIMIT_3GB flags I guess? Yes. > >> When >> ADDR_LIMIT_EXPLICIT is in effect, prctl can set a 64-bit numeric >> limit. If ADDR_LIMIT_EXPLICIT is cleared, the prctl value stops being >> settable and reading it via prctl returns whatever is implied by the >> other personality bits. > > I don't see anything wrong with it, but I'm a bit confused now > what this would be good for, compared to using just prctl. > > Is this about setuid clearing the personality but not the prctl, > or something else? It's to avid ambiguity as to what happens if you set ADDR_LIMIT_32BIT and use the prctl. ISTM it would be nice for the semantics to be fully defined in all cases. --Andy
[toc] | [prev] | [next] | [standalone]
| From | Arnd Bergmann <arnd@arndb.de> |
|---|---|
| Date | 2017-01-04 15:00 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sVUgF-5oR-9@gated-at.bofh.it> |
| In reply to | #1550248 |
On Tuesday, January 3, 2017 2:09:16 PM CET Andy Lutomirski wrote: > > > >> When > >> ADDR_LIMIT_EXPLICIT is in effect, prctl can set a 64-bit numeric > >> limit. If ADDR_LIMIT_EXPLICIT is cleared, the prctl value stops being > >> settable and reading it via prctl returns whatever is implied by the > >> other personality bits. > > > > I don't see anything wrong with it, but I'm a bit confused now > > what this would be good for, compared to using just prctl. > > > > Is this about setuid clearing the personality but not the prctl, > > or something else? > > It's to avid ambiguity as to what happens if you set ADDR_LIMIT_32BIT > and use the prctl. ISTM it would be nice for the semantics to be > fully defined in all cases. > Ok, got it. Arnd
[toc] | [prev] | [next] | [standalone]
| From | Arnd Bergmann <arnd@arndb.de> |
|---|---|
| Date | 2017-01-03 23:20 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sVFrk-48v-33@gated-at.bofh.it> |
| In reply to | #1550082 |
On Tuesday, January 3, 2017 10:29:33 AM CET Andy Lutomirski wrote: > > Hmm. What if we approached this a bit differently? We could add a > single new personality bit ADDR_LIMIT_EXPLICIT. Setting this bit > cause PER_LINUX32_3GB etc to be automatically cleared. Both the ADDR_LIMIT_32BIT and ADDR_LIMIT_3GB flags I guess? > When > ADDR_LIMIT_EXPLICIT is in effect, prctl can set a 64-bit numeric > limit. If ADDR_LIMIT_EXPLICIT is cleared, the prctl value stops being > settable and reading it via prctl returns whatever is implied by the > other personality bits. I don't see anything wrong with it, but I'm a bit confused now what this would be good for, compared to using just prctl. Is this about setuid clearing the personality but not the prctl, or something else? Arnd
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2017-01-03 17:10 +0100 |
| Subject | Re: [RFC, PATCHv2 29/29] mm, x86: introduce RLIMIT_VADDR |
| Message-ID | <sVzOW-jT-35@gated-at.bofh.it> |
| In reply to | #1549550 |
On Mon, Jan 02, 2017 at 10:08:28PM -0800, Andy Lutomirski wrote: > On Mon, Jan 2, 2017 at 12:44 AM, Arnd Bergmann <arnd@arndb.de> wrote: > > On Tuesday, December 27, 2016 4:54:13 AM CET Kirill A. Shutemov wrote: > >> As with other resources you can set the limit lower than current usage. > >> It would affect only future virtual address space allocations. > > I still don't buy all these use cases: > > >> > >> Use-cases for new rlimit: > >> > >> - Bumping the soft limit to RLIM_INFINITY, allows current process all > >> its children to use addresses above 47-bits. > > OK, I get this, but only as a workaround for programs that make > assumptions about the address space and don't use some mechanism (to > be designed?) to work correctly in spite of a larger address space. I guess you've misread the case. It's opt-in for large adrress space, not other way around. I believe 47-bit VA by default is right way to go to make the transition without breaking userspace. > >> - Bumping the soft limit to RLIM_INFINITY after fork(2), but before > >> exec(2) allows the child to use addresses above 47-bits. > > Ditto. > > >> > >> - Lowering the hard limit to 47-bits would prevent current process all > >> its children to use addresses above 47-bits, unless a process has > >> CAP_SYS_RESOURCES. > > I've tried and I can't imagine any reason to do this. That's just if something went wrong and we want to stop an application from use addresses above 47-bit. > >> - It’s also can be handy to lower hard or soft limit to arbitrary > >> address. User-mode emulation in QEMU may lower the limit to 32-bit > >> to emulate 32-bit machine on 64-bit host. > > I don't understand. QEMU user-mode emulation intercepts all syscalls. > What QEMU would *actually* want is a way to say "allocate me some > memory with the high N bits clear". mmap-via-int80 on x86 should be > fixed to do this, but a new syscall with an explicit parameter would > work, as would a prctl changing the current limit. Look at mess in mmap_find_vma(). QEmu has to guess where is free virtual memory. That's unnessesary complex. prctl would work for this too. new-mmap would *not*: there are more ways to allocate vitual address space: shmat(), mremap(). Changing all of them just for this is stupid. > >> > >> TODO: > >> - port to non-x86; > >> > >> Not-yet-signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> > >> Cc: linux-api@vger.kernel.org > > > > This seems to nicely address the same problem on arm64, which has > > run into the same issue due to the various page table formats > > that can currently be chosen at compile time. > > On further reflection, I think this has very little to do with paging > formats except insofar as paging formats make us notice the problem. > The issue is that user code wants to be able to assume an upper limit > on an address, and it gets an upper limit right now that depends on > architecture due to paging formats. But someone really might want to > write a *portable* 64-bit program that allocates memory with the high > 16 bits clear. So let's add such a mechanism directly. > > As a thought experiment, what if x86_64 simply never allocated "high" > (above 2^47-1) addresses unless a new mmap-with-explicit-limit syscall > were used? Old glibc would continue working. Old VMs would work. > New programs that want to use ginormous mappings would have to use the > new syscall. This would be totally stateless and would have no issues > with CRIU. Except, we need more than mmap as I mentioned. And what about stack? I'm not sure that everybody would be happy with stack in the middle of address space. > If necessary, we could also have a prctl that changes a > "personality-like" limit that is in effect when the old mmap was used. > I say "personality-like" because it would reset under exactly the same > conditions that personality resets itself. > > Thoughts? > > -- > To unsubscribe, send a message with 'unsubscribe linux-mm' in > the body to majordomo@kvack.org. For more info on Linux MM, > see: http://www.linux-mm.org/ . > Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a> -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | linux.kernel
csiph-web