Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1501116 > unrolled thread
| Started by | Boris Ostrovsky <boris.ostrovsky@oracle.com> |
|---|---|
| First post | 2016-10-14 20:20 +0200 |
| Last post | 2016-10-18 17:50 +0200 |
| Articles | 14 on this page of 34 — 6 participants |
Back to article view | Back to linux.kernel
[PATCH 0/8] PVH v2 support Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 20:20 +0200
[PATCH 7/8] xen/pvh: PVH guests always have PV devices Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 20:20 +0200
Re: [Xen-devel] [PATCH 7/8] xen/pvh: PVH guests always have PV devices Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-14 21:30 +0200
Re: [PATCH 7/8] xen/pvh: PVH guests always have PV devices Juergen Gross <jgross@suse.com> - 2016-10-18 18:00 +0200
[PATCH 5/8] xen/pvh: Prevent PVH guests from using PIC, RTC and IOAPIC Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 20:20 +0200
Re: [Xen-devel] [PATCH 5/8] xen/pvh: Prevent PVH guests from using PIC, RTC and IOAPIC Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-14 21:20 +0200
Re: [Xen-devel] [PATCH 5/8] xen/pvh: Prevent PVH guests from using PIC, RTC and IOAPIC Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 21:40 +0200
Re: [PATCH 5/8] xen/pvh: Prevent PVH guests from using PIC, RTC and IOAPIC Roger Pau Monné <roger.pau@citrix.com> - 2016-10-26 12:50 +0200
Re: [PATCH 5/8] xen/pvh: Prevent PVH guests from using PIC, RTC and IOAPIC Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-26 16:50 +0200
Re: [PATCH 5/8] xen/pvh: Prevent PVH guests from using PIC, RTC and IOAPIC Roger Pau Monné <roger.pau@citrix.com> - 2016-10-26 17:20 +0200
Re: [PATCH 5/8] xen/pvh: Prevent PVH guests from using PIC, RTC and IOAPIC Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-26 18:10 +0200
[PATCH 8/8] xen/pvh: Enable CPU hotplug Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 20:20 +0200
Re: [Xen-devel] [PATCH 8/8] xen/pvh: Enable CPU hotplug Andrew Cooper <andrew.cooper3@citrix.com> - 2016-10-14 20:50 +0200
Re: [Xen-devel] [PATCH 8/8] xen/pvh: Enable CPU hotplug Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 21:10 +0200
[PATCH 4/8] xen/pvh: Bootstrap PVH guest Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 20:20 +0200
Re: [Xen-devel] [PATCH 4/8] xen/pvh: Bootstrap PVH guest Andrew Cooper <andrew.cooper3@citrix.com> - 2016-10-14 20:50 +0200
Re: [Xen-devel] [PATCH 4/8] xen/pvh: Bootstrap PVH guest Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 21:00 +0200
Re: [Xen-devel] [PATCH 4/8] xen/pvh: Bootstrap PVH guest Andrew Cooper <andrew.cooper3@citrix.com> - 2016-10-14 21:20 +0200
Re: [Xen-devel] [PATCH 4/8] xen/pvh: Bootstrap PVH guest Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-14 21:20 +0200
Re: [Xen-devel] [PATCH 4/8] xen/pvh: Bootstrap PVH guest Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 21:40 +0200
[PATCH 2/8] x86/head: Refactor 32-bit pgtable setup Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 20:20 +0200
Re: [PATCH 2/8] x86/head: Refactor 32-bit pgtable setup hpa@zytor.com - 2016-10-14 20:40 +0200
Re: [PATCH 2/8] x86/head: Refactor 32-bit pgtable setup Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 20:50 +0200
Re: [PATCH 2/8] x86/head: Refactor 32-bit pgtable setup hpa@zytor.com - 2016-10-14 21:10 +0200
Re: [PATCH 2/8] x86/head: Refactor 32-bit pgtable setup Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 21:30 +0200
[PATCH 3/8] xen/pvh: Import PVH-related Xen public interfaces Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 20:20 +0200
Re: [Xen-devel] [PATCH 3/8] xen/pvh: Import PVH-related Xen public interfaces Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-14 20:40 +0200
Re: [PATCH 3/8] xen/pvh: Import PVH-related Xen public interfaces Juergen Gross <jgross@suse.com> - 2016-10-21 13:00 +0200
[PATCH 1/8] xen/x86: Remove PVH support Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-14 20:20 +0200
Re: [Xen-devel] [PATCH 1/8] xen/x86: Remove PVH support Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-14 20:40 +0200
Re: [PATCH 1/8] xen/x86: Remove PVH support Juergen Gross <jgross@suse.com> - 2016-10-18 15:50 +0200
Re: [PATCH 1/8] xen/x86: Remove PVH support Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-18 16:50 +0200
Re: [PATCH 1/8] xen/x86: Remove PVH support Juergen Gross <jgross@suse.com> - 2016-10-18 17:40 +0200
Re: [PATCH 1/8] xen/x86: Remove PVH support Boris Ostrovsky <boris.ostrovsky@oracle.com> - 2016-10-18 17:50 +0200
Page 2 of 2 — ← Prev page 1 [2]
| From | Boris Ostrovsky <boris.ostrovsky@oracle.com> |
|---|---|
| Date | 2016-10-14 20:20 +0200 |
| Subject | [PATCH 2/8] x86/head: Refactor 32-bit pgtable setup |
| Message-ID | <ssffk-4Pd-35@gated-at.bofh.it> |
| In reply to | #1501116 |
From: Matt Fleming <matt@codeblueprint.co.uk> The new Xen PVH entry point requires page tables to be setup by the kernel since it is entered with paging disabled. Pull the common code out of head_32.S and into pgtable_32.S so that setup_pgtable_32 can be invoked from both the new Xen entry point and the existing startup_32 code. Cc: Boris Ostrovsky <boris.ostrovsky@oracle.com> Cc: Thomas Gleixner <tglx@linutronix.de> Cc: Ingo Molnar <mingo@redhat.com> Cc: "H. Peter Anvin" <hpa@zytor.com> Cc: x86@kernel.org Signed-off-by: Matt Fleming <matt@codeblueprint.co.uk> --- arch/x86/Makefile | 2 + arch/x86/kernel/Makefile | 2 + arch/x86/kernel/head_32.S | 168 +------------------------------------ arch/x86/kernel/pgtable_32.S | 196 +++++++++++++++++++++++++++++++++++++++++++ 4 files changed, 201 insertions(+), 167 deletions(-) create mode 100644 arch/x86/kernel/pgtable_32.S diff --git a/arch/x86/Makefile b/arch/x86/Makefile index 2d44933..67cc771 100644 --- a/arch/x86/Makefile +++ b/arch/x86/Makefile @@ -204,6 +204,8 @@ head-y += arch/x86/kernel/head$(BITS).o head-y += arch/x86/kernel/ebda.o head-y += arch/x86/kernel/platform-quirks.o +head-$(CONFIG_X86_32) += arch/x86/kernel/pgtable_32.o + libs-y += arch/x86/lib/ # See arch/x86/Kbuild for content of core part of the kernel diff --git a/arch/x86/kernel/Makefile b/arch/x86/kernel/Makefile index 4dd5d50..eae85a5 100644 --- a/arch/x86/kernel/Makefile +++ b/arch/x86/kernel/Makefile @@ -8,6 +8,8 @@ extra-y += ebda.o extra-y += platform-quirks.o extra-y += vmlinux.lds +extra-$(CONFIG_X86_32) += pgtable_32.o + CPPFLAGS_vmlinux.lds += -U$(UTS_MACHINE) ifdef CONFIG_FUNCTION_TRACER diff --git a/arch/x86/kernel/head_32.S b/arch/x86/kernel/head_32.S index 5f40126..0db066e 100644 --- a/arch/x86/kernel/head_32.S +++ b/arch/x86/kernel/head_32.S @@ -41,51 +41,6 @@ #define X86_VENDOR_ID new_cpu_data+CPUINFO_x86_vendor_id /* - * This is how much memory in addition to the memory covered up to - * and including _end we need mapped initially. - * We need: - * (KERNEL_IMAGE_SIZE/4096) / 1024 pages (worst case, non PAE) - * (KERNEL_IMAGE_SIZE/4096) / 512 + 4 pages (worst case for PAE) - * - * Modulo rounding, each megabyte assigned here requires a kilobyte of - * memory, which is currently unreclaimed. - * - * This should be a multiple of a page. - * - * KERNEL_IMAGE_SIZE should be greater than pa(_end) - * and small than max_low_pfn, otherwise will waste some page table entries - */ - -#if PTRS_PER_PMD > 1 -#define PAGE_TABLE_SIZE(pages) (((pages) / PTRS_PER_PMD) + PTRS_PER_PGD) -#else -#define PAGE_TABLE_SIZE(pages) ((pages) / PTRS_PER_PGD) -#endif - -/* - * Number of possible pages in the lowmem region. - * - * We shift 2 by 31 instead of 1 by 32 to the left in order to avoid a - * gas warning about overflowing shift count when gas has been compiled - * with only a host target support using a 32-bit type for internal - * representation. - */ -LOWMEM_PAGES = (((2<<31) - __PAGE_OFFSET) >> PAGE_SHIFT) - -/* Enough space to fit pagetables for the low memory linear map */ -MAPPING_BEYOND_END = PAGE_TABLE_SIZE(LOWMEM_PAGES) << PAGE_SHIFT - -/* - * Worst-case size of the kernel mapping we need to make: - * a relocatable kernel can live anywhere in lowmem, so we need to be able - * to map all of lowmem. - */ -KERNEL_PAGES = LOWMEM_PAGES - -INIT_MAP_SIZE = PAGE_TABLE_SIZE(KERNEL_PAGES) * PAGE_SIZE -RESERVE_BRK(pagetables, INIT_MAP_SIZE) - -/* * 32-bit kernel entrypoint; only used by the boot CPU. On entry, * %esi points to the real-mode code as a 32-bit pointer. * CS and DS must be 4 GB flat segments, but we don't depend on @@ -157,92 +112,7 @@ ENTRY(startup_32) call load_ucode_bsp #endif -/* - * Initialize page tables. This creates a PDE and a set of page - * tables, which are located immediately beyond __brk_base. The variable - * _brk_end is set up to point to the first "safe" location. - * Mappings are created both at virtual address 0 (identity mapping) - * and PAGE_OFFSET for up to _end. - */ -#ifdef CONFIG_X86_PAE - - /* - * In PAE mode initial_page_table is statically defined to contain - * enough entries to cover the VMSPLIT option (that is the top 1, 2 or 3 - * entries). The identity mapping is handled by pointing two PGD entries - * to the first kernel PMD. - * - * Note the upper half of each PMD or PTE are always zero at this stage. - */ - -#define KPMDS (((-__PAGE_OFFSET) >> 30) & 3) /* Number of kernel PMDs */ - - xorl %ebx,%ebx /* %ebx is kept at zero */ - - movl $pa(__brk_base), %edi - movl $pa(initial_pg_pmd), %edx - movl $PTE_IDENT_ATTR, %eax -10: - leal PDE_IDENT_ATTR(%edi),%ecx /* Create PMD entry */ - movl %ecx,(%edx) /* Store PMD entry */ - /* Upper half already zero */ - addl $8,%edx - movl $512,%ecx -11: - stosl - xchgl %eax,%ebx - stosl - xchgl %eax,%ebx - addl $0x1000,%eax - loop 11b - - /* - * End condition: we must map up to the end + MAPPING_BEYOND_END. - */ - movl $pa(_end) + MAPPING_BEYOND_END + PTE_IDENT_ATTR, %ebp - cmpl %ebp,%eax - jb 10b -1: - addl $__PAGE_OFFSET, %edi - movl %edi, pa(_brk_end) - shrl $12, %eax - movl %eax, pa(max_pfn_mapped) - - /* Do early initialization of the fixmap area */ - movl $pa(initial_pg_fixmap)+PDE_IDENT_ATTR,%eax - movl %eax,pa(initial_pg_pmd+0x1000*KPMDS-8) -#else /* Not PAE */ - -page_pde_offset = (__PAGE_OFFSET >> 20); - - movl $pa(__brk_base), %edi - movl $pa(initial_page_table), %edx - movl $PTE_IDENT_ATTR, %eax -10: - leal PDE_IDENT_ATTR(%edi),%ecx /* Create PDE entry */ - movl %ecx,(%edx) /* Store identity PDE entry */ - movl %ecx,page_pde_offset(%edx) /* Store kernel PDE entry */ - addl $4,%edx - movl $1024, %ecx -11: - stosl - addl $0x1000,%eax - loop 11b - /* - * End condition: we must map up to the end + MAPPING_BEYOND_END. - */ - movl $pa(_end) + MAPPING_BEYOND_END + PTE_IDENT_ATTR, %ebp - cmpl %ebp,%eax - jb 10b - addl $__PAGE_OFFSET, %edi - movl %edi, pa(_brk_end) - shrl $12, %eax - movl %eax, pa(max_pfn_mapped) - - /* Do early initialization of the fixmap area */ - movl $pa(initial_pg_fixmap)+PDE_IDENT_ATTR,%eax - movl %eax,pa(initial_page_table+0xffc) -#endif + call setup_pgtable_32 #ifdef CONFIG_PARAVIRT /* This is can only trip for a broken bootloader... */ @@ -660,47 +530,11 @@ ENTRY(setup_once_ref) */ __PAGE_ALIGNED_BSS .align PAGE_SIZE -#ifdef CONFIG_X86_PAE -initial_pg_pmd: - .fill 1024*KPMDS,4,0 -#else -ENTRY(initial_page_table) - .fill 1024,4,0 -#endif -initial_pg_fixmap: - .fill 1024,4,0 ENTRY(empty_zero_page) .fill 4096,1,0 ENTRY(swapper_pg_dir) .fill 1024,4,0 -/* - * This starts the data section. - */ -#ifdef CONFIG_X86_PAE -__PAGE_ALIGNED_DATA - /* Page-aligned for the benefit of paravirt? */ - .align PAGE_SIZE -ENTRY(initial_page_table) - .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 /* low identity map */ -# if KPMDS == 3 - .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 - .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x1000),0 - .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x2000),0 -# elif KPMDS == 2 - .long 0,0 - .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 - .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x1000),0 -# elif KPMDS == 1 - .long 0,0 - .long 0,0 - .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 -# else -# error "Kernel PMDs should be 1, 2 or 3" -# endif - .align PAGE_SIZE /* needs to be page-sized too */ -#endif - .data .balign 4 ENTRY(initial_stack) diff --git a/arch/x86/kernel/pgtable_32.S b/arch/x86/kernel/pgtable_32.S new file mode 100644 index 0000000..aded718 --- /dev/null +++ b/arch/x86/kernel/pgtable_32.S @@ -0,0 +1,196 @@ +#include <linux/threads.h> +#include <linux/init.h> +#include <linux/linkage.h> +#include <asm/segment.h> +#include <asm/page_types.h> +#include <asm/pgtable_types.h> +#include <asm/cache.h> +#include <asm/thread_info.h> +#include <asm/asm-offsets.h> +#include <asm/setup.h> +#include <asm/processor-flags.h> +#include <asm/msr-index.h> +#include <asm/cpufeatures.h> +#include <asm/percpu.h> +#include <asm/nops.h> +#include <asm/bootparam.h> + +/* Physical address */ +#define pa(X) ((X) - __PAGE_OFFSET) + +/* + * This is how much memory in addition to the memory covered up to + * and including _end we need mapped initially. + * We need: + * (KERNEL_IMAGE_SIZE/4096) / 1024 pages (worst case, non PAE) + * (KERNEL_IMAGE_SIZE/4096) / 512 + 4 pages (worst case for PAE) + * + * Modulo rounding, each megabyte assigned here requires a kilobyte of + * memory, which is currently unreclaimed. + * + * This should be a multiple of a page. + * + * KERNEL_IMAGE_SIZE should be greater than pa(_end) + * and small than max_low_pfn, otherwise will waste some page table entries + */ + +#if PTRS_PER_PMD > 1 +#define PAGE_TABLE_SIZE(pages) (((pages) / PTRS_PER_PMD) + PTRS_PER_PGD) +#else +#define PAGE_TABLE_SIZE(pages) ((pages) / PTRS_PER_PGD) +#endif + +/* + * Number of possible pages in the lowmem region. + * + * We shift 2 by 31 instead of 1 by 32 to the left in order to avoid a + * gas warning about overflowing shift count when gas has been compiled + * with only a host target support using a 32-bit type for internal + * representation. + */ +LOWMEM_PAGES = (((2<<31) - __PAGE_OFFSET) >> PAGE_SHIFT) + +/* Enough space to fit pagetables for the low memory linear map */ +MAPPING_BEYOND_END = PAGE_TABLE_SIZE(LOWMEM_PAGES) << PAGE_SHIFT + +/* + * Worst-case size of the kernel mapping we need to make: + * a relocatable kernel can live anywhere in lowmem, so we need to be able + * to map all of lowmem. + */ +KERNEL_PAGES = LOWMEM_PAGES + +INIT_MAP_SIZE = PAGE_TABLE_SIZE(KERNEL_PAGES) * PAGE_SIZE +RESERVE_BRK(pagetables, INIT_MAP_SIZE) + +/* + * Initialize page tables. This creates a PDE and a set of page + * tables, which are located immediately beyond __brk_base. The variable + * _brk_end is set up to point to the first "safe" location. + * Mappings are created both at virtual address 0 (identity mapping) + * and PAGE_OFFSET for up to _end. + */ + .text +ENTRY(setup_pgtable_32) +#ifdef CONFIG_X86_PAE + /* + * In PAE mode initial_page_table is statically defined to contain + * enough entries to cover the VMSPLIT option (that is the top 1, 2 or 3 + * entries). The identity mapping is handled by pointing two PGD entries + * to the first kernel PMD. + * + * Note the upper half of each PMD or PTE are always zero at this stage. + */ + +#define KPMDS (((-__PAGE_OFFSET) >> 30) & 3) /* Number of kernel PMDs */ + + xorl %ebx,%ebx /* %ebx is kept at zero */ + + movl $pa(__brk_base), %edi + movl $pa(initial_pg_pmd), %edx + movl $PTE_IDENT_ATTR, %eax +10: + leal PDE_IDENT_ATTR(%edi),%ecx /* Create PMD entry */ + movl %ecx,(%edx) /* Store PMD entry */ + /* Upper half already zero */ + addl $8,%edx + movl $512,%ecx +11: + stosl + xchgl %eax,%ebx + stosl + xchgl %eax,%ebx + addl $0x1000,%eax + loop 11b + + /* + * End condition: we must map up to the end + MAPPING_BEYOND_END. + */ + movl $pa(_end) + MAPPING_BEYOND_END + PTE_IDENT_ATTR, %ebp + cmpl %ebp,%eax + jb 10b +1: + addl $__PAGE_OFFSET, %edi + movl %edi, pa(_brk_end) + shrl $12, %eax + movl %eax, pa(max_pfn_mapped) + + /* Do early initialization of the fixmap area */ + movl $pa(initial_pg_fixmap)+PDE_IDENT_ATTR,%eax + movl %eax,pa(initial_pg_pmd+0x1000*KPMDS-8) +#else /* Not PAE */ + +page_pde_offset = (__PAGE_OFFSET >> 20); + + movl $pa(__brk_base), %edi + movl $pa(initial_page_table), %edx + movl $PTE_IDENT_ATTR, %eax +10: + leal PDE_IDENT_ATTR(%edi),%ecx /* Create PDE entry */ + movl %ecx,(%edx) /* Store identity PDE entry */ + movl %ecx,page_pde_offset(%edx) /* Store kernel PDE entry */ + addl $4,%edx + movl $1024, %ecx +11: + stosl + addl $0x1000,%eax + loop 11b + /* + * End condition: we must map up to the end + MAPPING_BEYOND_END. + */ + movl $pa(_end) + MAPPING_BEYOND_END + PTE_IDENT_ATTR, %ebp + cmpl %ebp,%eax + jb 10b + addl $__PAGE_OFFSET, %edi + movl %edi, pa(_brk_end) + shrl $12, %eax + movl %eax, pa(max_pfn_mapped) + + /* Do early initialization of the fixmap area */ + movl $pa(initial_pg_fixmap)+PDE_IDENT_ATTR,%eax + movl %eax,pa(initial_page_table+0xffc) +#endif + ret +ENDPROC(setup_pgtable_32) + +/* + * BSS section + */ +__PAGE_ALIGNED_BSS + .align PAGE_SIZE +#ifdef CONFIG_X86_PAE +initial_pg_pmd: + .fill 1024*KPMDS,4,0 +#else +ENTRY(initial_page_table) + .fill 1024,4,0 +#endif +initial_pg_fixmap: + .fill 1024,4,0 + +/* + * This starts the data section. + */ +#ifdef CONFIG_X86_PAE +__PAGE_ALIGNED_DATA + /* Page-aligned for the benefit of paravirt? */ + .align PAGE_SIZE +ENTRY(initial_page_table) + .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 /* low identity map */ +# if KPMDS == 3 + .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 + .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x1000),0 + .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x2000),0 +# elif KPMDS == 2 + .long 0,0 + .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 + .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x1000),0 +# elif KPMDS == 1 + .long 0,0 + .long 0,0 + .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 +# else +# error "Kernel PMDs should be 1, 2 or 3" +# endif + .align PAGE_SIZE /* needs to be page-sized too */ +#endif -- 1.8.3.1
[toc] | [prev] | [next] | [standalone]
| From | hpa@zytor.com |
|---|---|
| Date | 2016-10-14 20:40 +0200 |
| Subject | Re: [PATCH 2/8] x86/head: Refactor 32-bit pgtable setup |
| Message-ID | <ssfyG-4WS-11@gated-at.bofh.it> |
| In reply to | #1501121 |
On October 14, 2016 11:05:12 AM PDT, Boris Ostrovsky <boris.ostrovsky@oracle.com> wrote: >From: Matt Fleming <matt@codeblueprint.co.uk> > >The new Xen PVH entry point requires page tables to be setup by the >kernel since it is entered with paging disabled. > >Pull the common code out of head_32.S and into pgtable_32.S so that >setup_pgtable_32 can be invoked from both the new Xen entry point and >the existing startup_32 code. > >Cc: Boris Ostrovsky <boris.ostrovsky@oracle.com> >Cc: Thomas Gleixner <tglx@linutronix.de> >Cc: Ingo Molnar <mingo@redhat.com> >Cc: "H. Peter Anvin" <hpa@zytor.com> >Cc: x86@kernel.org >Signed-off-by: Matt Fleming <matt@codeblueprint.co.uk> >--- > arch/x86/Makefile | 2 + > arch/x86/kernel/Makefile | 2 + >arch/x86/kernel/head_32.S | 168 >+------------------------------------ >arch/x86/kernel/pgtable_32.S | 196 >+++++++++++++++++++++++++++++++++++++++++++ > 4 files changed, 201 insertions(+), 167 deletions(-) > create mode 100644 arch/x86/kernel/pgtable_32.S > >diff --git a/arch/x86/Makefile b/arch/x86/Makefile >index 2d44933..67cc771 100644 >--- a/arch/x86/Makefile >+++ b/arch/x86/Makefile >@@ -204,6 +204,8 @@ head-y += arch/x86/kernel/head$(BITS).o > head-y += arch/x86/kernel/ebda.o > head-y += arch/x86/kernel/platform-quirks.o > >+head-$(CONFIG_X86_32) += arch/x86/kernel/pgtable_32.o >+ > libs-y += arch/x86/lib/ > > # See arch/x86/Kbuild for content of core part of the kernel >diff --git a/arch/x86/kernel/Makefile b/arch/x86/kernel/Makefile >index 4dd5d50..eae85a5 100644 >--- a/arch/x86/kernel/Makefile >+++ b/arch/x86/kernel/Makefile >@@ -8,6 +8,8 @@ extra-y += ebda.o > extra-y += platform-quirks.o > extra-y += vmlinux.lds > >+extra-$(CONFIG_X86_32) += pgtable_32.o >+ > CPPFLAGS_vmlinux.lds += -U$(UTS_MACHINE) > > ifdef CONFIG_FUNCTION_TRACER >diff --git a/arch/x86/kernel/head_32.S b/arch/x86/kernel/head_32.S >index 5f40126..0db066e 100644 >--- a/arch/x86/kernel/head_32.S >+++ b/arch/x86/kernel/head_32.S >@@ -41,51 +41,6 @@ > #define X86_VENDOR_ID new_cpu_data+CPUINFO_x86_vendor_id > > /* >- * This is how much memory in addition to the memory covered up to >- * and including _end we need mapped initially. >- * We need: >- * (KERNEL_IMAGE_SIZE/4096) / 1024 pages (worst case, non PAE) >- * (KERNEL_IMAGE_SIZE/4096) / 512 + 4 pages (worst case for PAE) >- * >- * Modulo rounding, each megabyte assigned here requires a kilobyte of >- * memory, which is currently unreclaimed. >- * >- * This should be a multiple of a page. >- * >- * KERNEL_IMAGE_SIZE should be greater than pa(_end) >- * and small than max_low_pfn, otherwise will waste some page table >entries >- */ >- >-#if PTRS_PER_PMD > 1 >-#define PAGE_TABLE_SIZE(pages) (((pages) / PTRS_PER_PMD) + >PTRS_PER_PGD) >-#else >-#define PAGE_TABLE_SIZE(pages) ((pages) / PTRS_PER_PGD) >-#endif >- >-/* >- * Number of possible pages in the lowmem region. >- * >- * We shift 2 by 31 instead of 1 by 32 to the left in order to avoid a >- * gas warning about overflowing shift count when gas has been >compiled >- * with only a host target support using a 32-bit type for internal >- * representation. >- */ >-LOWMEM_PAGES = (((2<<31) - __PAGE_OFFSET) >> PAGE_SHIFT) >- >-/* Enough space to fit pagetables for the low memory linear map */ >-MAPPING_BEYOND_END = PAGE_TABLE_SIZE(LOWMEM_PAGES) << PAGE_SHIFT >- >-/* >- * Worst-case size of the kernel mapping we need to make: >- * a relocatable kernel can live anywhere in lowmem, so we need to be >able >- * to map all of lowmem. >- */ >-KERNEL_PAGES = LOWMEM_PAGES >- >-INIT_MAP_SIZE = PAGE_TABLE_SIZE(KERNEL_PAGES) * PAGE_SIZE >-RESERVE_BRK(pagetables, INIT_MAP_SIZE) >- >-/* > * 32-bit kernel entrypoint; only used by the boot CPU. On entry, > * %esi points to the real-mode code as a 32-bit pointer. > * CS and DS must be 4 GB flat segments, but we don't depend on >@@ -157,92 +112,7 @@ ENTRY(startup_32) > call load_ucode_bsp > #endif > >-/* >- * Initialize page tables. This creates a PDE and a set of page >- * tables, which are located immediately beyond __brk_base. The >variable >- * _brk_end is set up to point to the first "safe" location. >- * Mappings are created both at virtual address 0 (identity mapping) >- * and PAGE_OFFSET for up to _end. >- */ >-#ifdef CONFIG_X86_PAE >- >- /* >- * In PAE mode initial_page_table is statically defined to contain >- * enough entries to cover the VMSPLIT option (that is the top 1, 2 >or 3 >- * entries). The identity mapping is handled by pointing two PGD >entries >- * to the first kernel PMD. >- * >- * Note the upper half of each PMD or PTE are always zero at this >stage. >- */ >- >-#define KPMDS (((-__PAGE_OFFSET) >> 30) & 3) /* Number of kernel PMDs >*/ >- >- xorl %ebx,%ebx /* %ebx is kept at zero */ >- >- movl $pa(__brk_base), %edi >- movl $pa(initial_pg_pmd), %edx >- movl $PTE_IDENT_ATTR, %eax >-10: >- leal PDE_IDENT_ATTR(%edi),%ecx /* Create PMD entry */ >- movl %ecx,(%edx) /* Store PMD entry */ >- /* Upper half already zero */ >- addl $8,%edx >- movl $512,%ecx >-11: >- stosl >- xchgl %eax,%ebx >- stosl >- xchgl %eax,%ebx >- addl $0x1000,%eax >- loop 11b >- >- /* >- * End condition: we must map up to the end + MAPPING_BEYOND_END. >- */ >- movl $pa(_end) + MAPPING_BEYOND_END + PTE_IDENT_ATTR, %ebp >- cmpl %ebp,%eax >- jb 10b >-1: >- addl $__PAGE_OFFSET, %edi >- movl %edi, pa(_brk_end) >- shrl $12, %eax >- movl %eax, pa(max_pfn_mapped) >- >- /* Do early initialization of the fixmap area */ >- movl $pa(initial_pg_fixmap)+PDE_IDENT_ATTR,%eax >- movl %eax,pa(initial_pg_pmd+0x1000*KPMDS-8) >-#else /* Not PAE */ >- >-page_pde_offset = (__PAGE_OFFSET >> 20); >- >- movl $pa(__brk_base), %edi >- movl $pa(initial_page_table), %edx >- movl $PTE_IDENT_ATTR, %eax >-10: >- leal PDE_IDENT_ATTR(%edi),%ecx /* Create PDE entry */ >- movl %ecx,(%edx) /* Store identity PDE entry */ >- movl %ecx,page_pde_offset(%edx) /* Store kernel PDE entry */ >- addl $4,%edx >- movl $1024, %ecx >-11: >- stosl >- addl $0x1000,%eax >- loop 11b >- /* >- * End condition: we must map up to the end + MAPPING_BEYOND_END. >- */ >- movl $pa(_end) + MAPPING_BEYOND_END + PTE_IDENT_ATTR, %ebp >- cmpl %ebp,%eax >- jb 10b >- addl $__PAGE_OFFSET, %edi >- movl %edi, pa(_brk_end) >- shrl $12, %eax >- movl %eax, pa(max_pfn_mapped) >- >- /* Do early initialization of the fixmap area */ >- movl $pa(initial_pg_fixmap)+PDE_IDENT_ATTR,%eax >- movl %eax,pa(initial_page_table+0xffc) >-#endif >+ call setup_pgtable_32 > > #ifdef CONFIG_PARAVIRT > /* This is can only trip for a broken bootloader... */ >@@ -660,47 +530,11 @@ ENTRY(setup_once_ref) > */ > __PAGE_ALIGNED_BSS > .align PAGE_SIZE >-#ifdef CONFIG_X86_PAE >-initial_pg_pmd: >- .fill 1024*KPMDS,4,0 >-#else >-ENTRY(initial_page_table) >- .fill 1024,4,0 >-#endif >-initial_pg_fixmap: >- .fill 1024,4,0 > ENTRY(empty_zero_page) > .fill 4096,1,0 > ENTRY(swapper_pg_dir) > .fill 1024,4,0 > >-/* >- * This starts the data section. >- */ >-#ifdef CONFIG_X86_PAE >-__PAGE_ALIGNED_DATA >- /* Page-aligned for the benefit of paravirt? */ >- .align PAGE_SIZE >-ENTRY(initial_page_table) >- .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 /* low identity map */ >-# if KPMDS == 3 >- .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 >- .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x1000),0 >- .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x2000),0 >-# elif KPMDS == 2 >- .long 0,0 >- .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 >- .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x1000),0 >-# elif KPMDS == 1 >- .long 0,0 >- .long 0,0 >- .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 >-# else >-# error "Kernel PMDs should be 1, 2 or 3" >-# endif >- .align PAGE_SIZE /* needs to be page-sized too */ >-#endif >- > .data > .balign 4 > ENTRY(initial_stack) >diff --git a/arch/x86/kernel/pgtable_32.S >b/arch/x86/kernel/pgtable_32.S >new file mode 100644 >index 0000000..aded718 >--- /dev/null >+++ b/arch/x86/kernel/pgtable_32.S >@@ -0,0 +1,196 @@ >+#include <linux/threads.h> >+#include <linux/init.h> >+#include <linux/linkage.h> >+#include <asm/segment.h> >+#include <asm/page_types.h> >+#include <asm/pgtable_types.h> >+#include <asm/cache.h> >+#include <asm/thread_info.h> >+#include <asm/asm-offsets.h> >+#include <asm/setup.h> >+#include <asm/processor-flags.h> >+#include <asm/msr-index.h> >+#include <asm/cpufeatures.h> >+#include <asm/percpu.h> >+#include <asm/nops.h> >+#include <asm/bootparam.h> >+ >+/* Physical address */ >+#define pa(X) ((X) - __PAGE_OFFSET) >+ >+/* >+ * This is how much memory in addition to the memory covered up to >+ * and including _end we need mapped initially. >+ * We need: >+ * (KERNEL_IMAGE_SIZE/4096) / 1024 pages (worst case, non PAE) >+ * (KERNEL_IMAGE_SIZE/4096) / 512 + 4 pages (worst case for PAE) >+ * >+ * Modulo rounding, each megabyte assigned here requires a kilobyte of >+ * memory, which is currently unreclaimed. >+ * >+ * This should be a multiple of a page. >+ * >+ * KERNEL_IMAGE_SIZE should be greater than pa(_end) >+ * and small than max_low_pfn, otherwise will waste some page table >entries >+ */ >+ >+#if PTRS_PER_PMD > 1 >+#define PAGE_TABLE_SIZE(pages) (((pages) / PTRS_PER_PMD) + >PTRS_PER_PGD) >+#else >+#define PAGE_TABLE_SIZE(pages) ((pages) / PTRS_PER_PGD) >+#endif >+ >+/* >+ * Number of possible pages in the lowmem region. >+ * >+ * We shift 2 by 31 instead of 1 by 32 to the left in order to avoid a >+ * gas warning about overflowing shift count when gas has been >compiled >+ * with only a host target support using a 32-bit type for internal >+ * representation. >+ */ >+LOWMEM_PAGES = (((2<<31) - __PAGE_OFFSET) >> PAGE_SHIFT) >+ >+/* Enough space to fit pagetables for the low memory linear map */ >+MAPPING_BEYOND_END = PAGE_TABLE_SIZE(LOWMEM_PAGES) << PAGE_SHIFT >+ >+/* >+ * Worst-case size of the kernel mapping we need to make: >+ * a relocatable kernel can live anywhere in lowmem, so we need to be >able >+ * to map all of lowmem. >+ */ >+KERNEL_PAGES = LOWMEM_PAGES >+ >+INIT_MAP_SIZE = PAGE_TABLE_SIZE(KERNEL_PAGES) * PAGE_SIZE >+RESERVE_BRK(pagetables, INIT_MAP_SIZE) >+ >+/* >+ * Initialize page tables. This creates a PDE and a set of page >+ * tables, which are located immediately beyond __brk_base. The >variable >+ * _brk_end is set up to point to the first "safe" location. >+ * Mappings are created both at virtual address 0 (identity mapping) >+ * and PAGE_OFFSET for up to _end. >+ */ >+ .text >+ENTRY(setup_pgtable_32) >+#ifdef CONFIG_X86_PAE >+ /* >+ * In PAE mode initial_page_table is statically defined to contain >+ * enough entries to cover the VMSPLIT option (that is the top 1, 2 >or 3 >+ * entries). The identity mapping is handled by pointing two PGD >entries >+ * to the first kernel PMD. >+ * >+ * Note the upper half of each PMD or PTE are always zero at this >stage. >+ */ >+ >+#define KPMDS (((-__PAGE_OFFSET) >> 30) & 3) /* Number of kernel PMDs >*/ >+ >+ xorl %ebx,%ebx /* %ebx is kept at zero */ >+ >+ movl $pa(__brk_base), %edi >+ movl $pa(initial_pg_pmd), %edx >+ movl $PTE_IDENT_ATTR, %eax >+10: >+ leal PDE_IDENT_ATTR(%edi),%ecx /* Create PMD entry */ >+ movl %ecx,(%edx) /* Store PMD entry */ >+ /* Upper half already zero */ >+ addl $8,%edx >+ movl $512,%ecx >+11: >+ stosl >+ xchgl %eax,%ebx >+ stosl >+ xchgl %eax,%ebx >+ addl $0x1000,%eax >+ loop 11b >+ >+ /* >+ * End condition: we must map up to the end + MAPPING_BEYOND_END. >+ */ >+ movl $pa(_end) + MAPPING_BEYOND_END + PTE_IDENT_ATTR, %ebp >+ cmpl %ebp,%eax >+ jb 10b >+1: >+ addl $__PAGE_OFFSET, %edi >+ movl %edi, pa(_brk_end) >+ shrl $12, %eax >+ movl %eax, pa(max_pfn_mapped) >+ >+ /* Do early initialization of the fixmap area */ >+ movl $pa(initial_pg_fixmap)+PDE_IDENT_ATTR,%eax >+ movl %eax,pa(initial_pg_pmd+0x1000*KPMDS-8) >+#else /* Not PAE */ >+ >+page_pde_offset = (__PAGE_OFFSET >> 20); >+ >+ movl $pa(__brk_base), %edi >+ movl $pa(initial_page_table), %edx >+ movl $PTE_IDENT_ATTR, %eax >+10: >+ leal PDE_IDENT_ATTR(%edi),%ecx /* Create PDE entry */ >+ movl %ecx,(%edx) /* Store identity PDE entry */ >+ movl %ecx,page_pde_offset(%edx) /* Store kernel PDE entry */ >+ addl $4,%edx >+ movl $1024, %ecx >+11: >+ stosl >+ addl $0x1000,%eax >+ loop 11b >+ /* >+ * End condition: we must map up to the end + MAPPING_BEYOND_END. >+ */ >+ movl $pa(_end) + MAPPING_BEYOND_END + PTE_IDENT_ATTR, %ebp >+ cmpl %ebp,%eax >+ jb 10b >+ addl $__PAGE_OFFSET, %edi >+ movl %edi, pa(_brk_end) >+ shrl $12, %eax >+ movl %eax, pa(max_pfn_mapped) >+ >+ /* Do early initialization of the fixmap area */ >+ movl $pa(initial_pg_fixmap)+PDE_IDENT_ATTR,%eax >+ movl %eax,pa(initial_page_table+0xffc) >+#endif >+ ret >+ENDPROC(setup_pgtable_32) >+ >+/* >+ * BSS section >+ */ >+__PAGE_ALIGNED_BSS >+ .align PAGE_SIZE >+#ifdef CONFIG_X86_PAE >+initial_pg_pmd: >+ .fill 1024*KPMDS,4,0 >+#else >+ENTRY(initial_page_table) >+ .fill 1024,4,0 >+#endif >+initial_pg_fixmap: >+ .fill 1024,4,0 >+ >+/* >+ * This starts the data section. >+ */ >+#ifdef CONFIG_X86_PAE >+__PAGE_ALIGNED_DATA >+ /* Page-aligned for the benefit of paravirt? */ >+ .align PAGE_SIZE >+ENTRY(initial_page_table) >+ .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 /* low identity map */ >+# if KPMDS == 3 >+ .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 >+ .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x1000),0 >+ .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x2000),0 >+# elif KPMDS == 2 >+ .long 0,0 >+ .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 >+ .long pa(initial_pg_pmd+PGD_IDENT_ATTR+0x1000),0 >+# elif KPMDS == 1 >+ .long 0,0 >+ .long 0,0 >+ .long pa(initial_pg_pmd+PGD_IDENT_ATTR),0 >+# else >+# error "Kernel PMDs should be 1, 2 or 3" >+# endif >+ .align PAGE_SIZE /* needs to be page-sized too */ >+#endif And why does it need a separate entry point as opposed to the plain one? -- Sent from my Android device with K-9 Mail. Please excuse my brevity.
[toc] | [prev] | [next] | [standalone]
| From | Boris Ostrovsky <boris.ostrovsky@oracle.com> |
|---|---|
| Date | 2016-10-14 20:50 +0200 |
| Subject | Re: [PATCH 2/8] x86/head: Refactor 32-bit pgtable setup |
| Message-ID | <ssfIm-53c-23@gated-at.bofh.it> |
| In reply to | #1501138 |
On 10/14/2016 02:31 PM, hpa@zytor.com wrote:
> On October 14, 2016 11:05:12 AM PDT, Boris Ostrovsky <boris.ostrovsky@oracle.com> wrote:
>> From: Matt Fleming <matt@codeblueprint.co.uk>
>>
>> The new Xen PVH entry point requires page tables to be setup by the
>> kernel since it is entered with paging disabled.
>>
>> Pull the common code out of head_32.S and into pgtable_32.S so that
>> setup_pgtable_32 can be invoked from both the new Xen entry point and
>> the existing startup_32 code.
>>
> And why does it need a separate entry point as opposed to the plain one?
One reason is that we need to prepare boot_params before jumping to
startup_{32|64}.
When the guest is loaded (always in 32-bit mode) the only thing we have
is a pointer to Xen-specific datastructure. The early PVH code will
prepare zeropage based on that structure and then jump to regular
startup_*() code.
-boris
[toc] | [prev] | [next] | [standalone]
| From | hpa@zytor.com |
|---|---|
| Date | 2016-10-14 21:10 +0200 |
| Subject | Re: [PATCH 2/8] x86/head: Refactor 32-bit pgtable setup |
| Message-ID | <ssg1H-5pD-1@gated-at.bofh.it> |
| In reply to | #1501145 |
On October 14, 2016 11:44:18 AM PDT, Boris Ostrovsky <boris.ostrovsky@oracle.com> wrote:
>On 10/14/2016 02:31 PM, hpa@zytor.com wrote:
>> On October 14, 2016 11:05:12 AM PDT, Boris Ostrovsky
><boris.ostrovsky@oracle.com> wrote:
>>> From: Matt Fleming <matt@codeblueprint.co.uk>
>>>
>>> The new Xen PVH entry point requires page tables to be setup by the
>>> kernel since it is entered with paging disabled.
>>>
>>> Pull the common code out of head_32.S and into pgtable_32.S so that
>>> setup_pgtable_32 can be invoked from both the new Xen entry point
>and
>>> the existing startup_32 code.
>>>
>> And why does it need a separate entry point as opposed to the plain
>one?
>
>One reason is that we need to prepare boot_params before jumping to
>startup_{32|64}.
>
>When the guest is loaded (always in 32-bit mode) the only thing we have
>is a pointer to Xen-specific datastructure. The early PVH code will
>prepare zeropage based on that structure and then jump to regular
>startup_*() code.
>
>-boris
And why not just resume execution at start_32 then?
--
Sent from my Android device with K-9 Mail. Please excuse my brevity.
[toc] | [prev] | [next] | [standalone]
| From | Boris Ostrovsky <boris.ostrovsky@oracle.com> |
|---|---|
| Date | 2016-10-14 21:30 +0200 |
| Subject | Re: [PATCH 2/8] x86/head: Refactor 32-bit pgtable setup |
| Message-ID | <ssgl4-5wE-49@gated-at.bofh.it> |
| In reply to | #1501154 |
On 10/14/2016 03:04 PM, hpa@zytor.com wrote:
> On October 14, 2016 11:44:18 AM PDT, Boris Ostrovsky <boris.ostrovsky@oracle.com> wrote:
>> On 10/14/2016 02:31 PM, hpa@zytor.com wrote:
>>> On October 14, 2016 11:05:12 AM PDT, Boris Ostrovsky
>> <boris.ostrovsky@oracle.com> wrote:
>>>> From: Matt Fleming <matt@codeblueprint.co.uk>
>>>>
>>>> The new Xen PVH entry point requires page tables to be setup by the
>>>> kernel since it is entered with paging disabled.
>>>>
>>>> Pull the common code out of head_32.S and into pgtable_32.S so that
>>>> setup_pgtable_32 can be invoked from both the new Xen entry point
>> and
>>>> the existing startup_32 code.
>>>>
>>> And why does it need a separate entry point as opposed to the plain
>> one?
>>
>> One reason is that we need to prepare boot_params before jumping to
>> startup_{32|64}.
>>
>> When the guest is loaded (always in 32-bit mode) the only thing we have
>> is a pointer to Xen-specific datastructure. The early PVH code will
>> prepare zeropage based on that structure and then jump to regular
>> startup_*() code.
>>
>> -boris
> And why not just resume execution at start_32 then?
I am not sure what start_32 is.
If you meant startup_32 then that's exactly what we do (for 32-bit
guests) once zeropage is set up.
-boris
-boris
[toc] | [prev] | [next] | [standalone]
| From | Boris Ostrovsky <boris.ostrovsky@oracle.com> |
|---|---|
| Date | 2016-10-14 20:20 +0200 |
| Subject | [PATCH 3/8] xen/pvh: Import PVH-related Xen public interfaces |
| Message-ID | <ssffk-4Pd-37@gated-at.bofh.it> |
| In reply to | #1501116 |
Signed-off-by: Boris Ostrovsky <boris.ostrovsky@oracle.com>
---
include/xen/interface/elfnote.h | 12 ++-
include/xen/interface/hvm/hvm_vcpu.h | 143 +++++++++++++++++++++++++++++++++
include/xen/interface/hvm/start_info.h | 98 ++++++++++++++++++++++
3 files changed, 252 insertions(+), 1 deletion(-)
create mode 100644 include/xen/interface/hvm/hvm_vcpu.h
create mode 100644 include/xen/interface/hvm/start_info.h
diff --git a/include/xen/interface/elfnote.h b/include/xen/interface/elfnote.h
index f90b034..9e9f9bf 100644
--- a/include/xen/interface/elfnote.h
+++ b/include/xen/interface/elfnote.h
@@ -193,9 +193,19 @@
#define XEN_ELFNOTE_SUPPORTED_FEATURES 17
/*
+ * Physical entry point into the kernel.
+ *
+ * 32bit entry point into the kernel. When requested to launch the
+ * guest kernel in a HVM container, Xen will use this entry point to
+ * launch the guest in 32bit protected mode with paging disabled.
+ * Ignored otherwise.
+ */
+#define XEN_ELFNOTE_PHYS32_ENTRY 18
+
+/*
* The number of the highest elfnote defined.
*/
-#define XEN_ELFNOTE_MAX XEN_ELFNOTE_SUPPORTED_FEATURES
+#define XEN_ELFNOTE_MAX XEN_ELFNOTE_PHYS32_ENTRY
#endif /* __XEN_PUBLIC_ELFNOTE_H__ */
diff --git a/include/xen/interface/hvm/hvm_vcpu.h b/include/xen/interface/hvm/hvm_vcpu.h
new file mode 100644
index 0000000..32ca83e
--- /dev/null
+++ b/include/xen/interface/hvm/hvm_vcpu.h
@@ -0,0 +1,143 @@
+/*
+ * Permission is hereby granted, free of charge, to any person obtaining a copy
+ * of this software and associated documentation files (the "Software"), to
+ * deal in the Software without restriction, including without limitation the
+ * rights to use, copy, modify, merge, publish, distribute, sublicense, and/or
+ * sell copies of the Software, and to permit persons to whom the Software is
+ * furnished to do so, subject to the following conditions:
+ *
+ * The above copyright notice and this permission notice shall be included in
+ * all copies or substantial portions of the Software.
+ *
+ * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
+ * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
+ * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
+ * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
+ * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING
+ * FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER
+ * DEALINGS IN THE SOFTWARE.
+ *
+ * Copyright (c) 2015, Roger Pau Monne <roger.pau@citrix.com>
+ */
+
+#ifndef __XEN_PUBLIC_HVM_HVM_VCPU_H__
+#define __XEN_PUBLIC_HVM_HVM_VCPU_H__
+
+#include "../xen.h"
+
+struct vcpu_hvm_x86_32 {
+ uint32_t eax;
+ uint32_t ecx;
+ uint32_t edx;
+ uint32_t ebx;
+ uint32_t esp;
+ uint32_t ebp;
+ uint32_t esi;
+ uint32_t edi;
+ uint32_t eip;
+ uint32_t eflags;
+
+ uint32_t cr0;
+ uint32_t cr3;
+ uint32_t cr4;
+
+ uint32_t pad1;
+
+ /*
+ * EFER should only be used to set the NXE bit (if required)
+ * when starting a vCPU in 32bit mode with paging enabled or
+ * to set the LME/LMA bits in order to start the vCPU in
+ * compatibility mode.
+ */
+ uint64_t efer;
+
+ uint32_t cs_base;
+ uint32_t ds_base;
+ uint32_t ss_base;
+ uint32_t es_base;
+ uint32_t tr_base;
+ uint32_t cs_limit;
+ uint32_t ds_limit;
+ uint32_t ss_limit;
+ uint32_t es_limit;
+ uint32_t tr_limit;
+ uint16_t cs_ar;
+ uint16_t ds_ar;
+ uint16_t ss_ar;
+ uint16_t es_ar;
+ uint16_t tr_ar;
+
+ uint16_t pad2[3];
+};
+
+/*
+ * The layout of the _ar fields of the segment registers is the
+ * following:
+ *
+ * Bits [0,3]: type (bits 40-43).
+ * Bit 4: s (descriptor type, bit 44).
+ * Bit [5,6]: dpl (descriptor privilege level, bits 45-46).
+ * Bit 7: p (segment-present, bit 47).
+ * Bit 8: avl (available for system software, bit 52).
+ * Bit 9: l (64-bit code segment, bit 53).
+ * Bit 10: db (meaning depends on the segment, bit 54).
+ * Bit 11: g (granularity, bit 55)
+ * Bits [12,15]: unused, must be blank.
+ *
+ * A more complete description of the meaning of this fields can be
+ * obtained from the Intel SDM, Volume 3, section 3.4.5.
+ */
+
+struct vcpu_hvm_x86_64 {
+ uint64_t rax;
+ uint64_t rcx;
+ uint64_t rdx;
+ uint64_t rbx;
+ uint64_t rsp;
+ uint64_t rbp;
+ uint64_t rsi;
+ uint64_t rdi;
+ uint64_t rip;
+ uint64_t rflags;
+
+ uint64_t cr0;
+ uint64_t cr3;
+ uint64_t cr4;
+ uint64_t efer;
+
+ /*
+ * Using VCPU_HVM_MODE_64B implies that the vCPU is launched
+ * directly in long mode, so the cached parts of the segment
+ * registers get set to match that environment.
+ *
+ * If the user wants to launch the vCPU in compatibility mode
+ * the 32-bit structure should be used instead.
+ */
+};
+
+struct vcpu_hvm_context {
+#define VCPU_HVM_MODE_32B 0 /* 32bit fields of the structure will be used. */
+#define VCPU_HVM_MODE_64B 1 /* 64bit fields of the structure will be used. */
+ uint32_t mode;
+
+ uint32_t pad;
+
+ /* CPU registers. */
+ union {
+ struct vcpu_hvm_x86_32 x86_32;
+ struct vcpu_hvm_x86_64 x86_64;
+ } cpu_regs;
+};
+typedef struct vcpu_hvm_context vcpu_hvm_context_t;
+
+#endif /* __XEN_PUBLIC_HVM_HVM_VCPU_H__ */
+
+/*
+ * Local variables:
+ * mode: C
+ * c-file-style: "BSD"
+ * c-basic-offset: 4
+ * tab-width: 4
+ * indent-tabs-mode: nil
+ * End:
+ */
diff --git a/include/xen/interface/hvm/start_info.h b/include/xen/interface/hvm/start_info.h
new file mode 100644
index 0000000..6484159
--- /dev/null
+++ b/include/xen/interface/hvm/start_info.h
@@ -0,0 +1,98 @@
+/*
+ * Permission is hereby granted, free of charge, to any person obtaining a copy
+ * of this software and associated documentation files (the "Software"), to
+ * deal in the Software without restriction, including without limitation the
+ * rights to use, copy, modify, merge, publish, distribute, sublicense, and/or
+ * sell copies of the Software, and to permit persons to whom the Software is
+ * furnished to do so, subject to the following conditions:
+ *
+ * The above copyright notice and this permission notice shall be included in
+ * all copies or substantial portions of the Software.
+ *
+ * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
+ * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
+ * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
+ * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
+ * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING
+ * FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER
+ * DEALINGS IN THE SOFTWARE.
+ *
+ * Copyright (c) 2016, Citrix Systems, Inc.
+ */
+
+#ifndef __XEN_PUBLIC_ARCH_X86_HVM_START_INFO_H__
+#define __XEN_PUBLIC_ARCH_X86_HVM_START_INFO_H__
+
+/*
+ * Start of day structure passed to PVH guests and to HVM guests in %ebx.
+ *
+ * NOTE: nothing will be loaded at physical address 0, so a 0 value in any
+ * of the address fields should be treated as not present.
+ *
+ * 0 +----------------+
+ * | magic | Contains the magic value XEN_HVM_START_MAGIC_VALUE
+ * | | ("xEn3" with the 0x80 bit of the "E" set).
+ * 4 +----------------+
+ * | version | Version of this structure. Current version is 0. New
+ * | | versions are guaranteed to be backwards-compatible.
+ * 8 +----------------+
+ * | flags | SIF_xxx flags.
+ * 12 +----------------+
+ * | nr_modules | Number of modules passed to the kernel.
+ * 16 +----------------+
+ * | modlist_paddr | Physical address of an array of modules
+ * | | (layout of the structure below).
+ * 24 +----------------+
+ * | cmdline_paddr | Physical address of the command line,
+ * | | a zero-terminated ASCII string.
+ * 32 +----------------+
+ * | rsdp_paddr | Physical address of the RSDP ACPI data structure.
+ * 40 +----------------+
+ *
+ * The layout of each entry in the module structure is the following:
+ *
+ * 0 +----------------+
+ * | paddr | Physical address of the module.
+ * 8 +----------------+
+ * | size | Size of the module in bytes.
+ * 16 +----------------+
+ * | cmdline_paddr | Physical address of the command line,
+ * | | a zero-terminated ASCII string.
+ * 24 +----------------+
+ * | reserved |
+ * 32 +----------------+
+ *
+ * The address and sizes are always a 64bit little endian unsigned integer.
+ *
+ * NB: Xen on x86 will always try to place all the data below the 4GiB
+ * boundary.
+ */
+#define XEN_HVM_START_MAGIC_VALUE 0x336ec578
+
+/*
+ * C representation of the x86/HVM start info layout.
+ *
+ * The canonical definition of this layout is above, this is just a way to
+ * represent the layout described there using C types.
+ */
+struct hvm_start_info {
+ uint32_t magic; /* Contains the magic value 0x336ec578 */
+ /* ("xEn3" with the 0x80 bit of the "E" set).*/
+ uint32_t version; /* Version of this structure. */
+ uint32_t flags; /* SIF_xxx flags. */
+ uint32_t nr_modules; /* Number of modules passed to the kernel. */
+ uint64_t modlist_paddr; /* Physical address of an array of */
+ /* hvm_modlist_entry. */
+ uint64_t cmdline_paddr; /* Physical address of the command line. */
+ uint64_t rsdp_paddr; /* Physical address of the RSDP ACPI data */
+ /* structure. */
+};
+
+struct hvm_modlist_entry {
+ uint64_t paddr; /* Physical address of the module. */
+ uint64_t size; /* Size of the module in bytes. */
+ uint64_t cmdline_paddr; /* Physical address of the command line. */
+ uint64_t reserved;
+};
+
+#endif /* __XEN_PUBLIC_ARCH_X86_HVM_START_INFO_H__ */
--
1.8.3.1
[toc] | [prev] | [next] | [standalone]
| From | Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> |
|---|---|
| Date | 2016-10-14 20:40 +0200 |
| Subject | Re: [Xen-devel] [PATCH 3/8] xen/pvh: Import PVH-related Xen public interfaces |
| Message-ID | <ssfyG-4WS-27@gated-at.bofh.it> |
| In reply to | #1501122 |
On Fri, Oct 14, 2016 at 02:05:13PM -0400, Boris Ostrovsky wrote: > Signed-off-by: Boris Ostrovsky <boris.ostrovsky@oracle.com> Reviewed-by: Konrad Rzeszutek Wilk <konrad.wilk@oracle.com>
[toc] | [prev] | [next] | [standalone]
| From | Juergen Gross <jgross@suse.com> |
|---|---|
| Date | 2016-10-21 13:00 +0200 |
| Subject | Re: [PATCH 3/8] xen/pvh: Import PVH-related Xen public interfaces |
| Message-ID | <suFIl-4q7-1@gated-at.bofh.it> |
| In reply to | #1501122 |
On 14/10/16 20:05, Boris Ostrovsky wrote: > Signed-off-by: Boris Ostrovsky <boris.ostrovsky@oracle.com> Reviewed-by: Juergen Gross <jgross@suse.com> Juergen
[toc] | [prev] | [next] | [standalone]
| From | Boris Ostrovsky <boris.ostrovsky@oracle.com> |
|---|---|
| Date | 2016-10-14 20:20 +0200 |
| Subject | [PATCH 1/8] xen/x86: Remove PVH support |
| Message-ID | <ssffk-4Pd-39@gated-at.bofh.it> |
| In reply to | #1501116 |
We are replacing existing PVH guests with new implementation.
Signed-off-by: Boris Ostrovsky <boris.ostrovsky@oracle.com>
---
arch/x86/xen/enlighten.c | 140 ++++++---------------------------------
arch/x86/xen/mmu.c | 21 +-----
arch/x86/xen/setup.c | 37 +----------
arch/x86/xen/smp.c | 78 ++++++++--------------
arch/x86/xen/smp.h | 8 ---
arch/x86/xen/xen-head.S | 62 ++---------------
arch/x86/xen/xen-ops.h | 1 -
drivers/xen/events/events_base.c | 1 -
include/xen/xen.h | 13 +---
9 files changed, 54 insertions(+), 307 deletions(-)
diff --git a/arch/x86/xen/enlighten.c b/arch/x86/xen/enlighten.c
index c0fdd57..dc4ed0c 100644
--- a/arch/x86/xen/enlighten.c
+++ b/arch/x86/xen/enlighten.c
@@ -1149,10 +1149,11 @@ void xen_setup_vcpu_info_placement(void)
xen_vcpu_setup(cpu);
}
- /* xen_vcpu_setup managed to place the vcpu_info within the
- * percpu area for all cpus, so make use of it. Note that for
- * PVH we want to use native IRQ mechanism. */
- if (have_vcpu_info_placement && !xen_pvh_domain()) {
+ /*
+ * xen_vcpu_setup managed to place the vcpu_info within the
+ * percpu area for all cpus, so make use of it.
+ */
+ if (have_vcpu_info_placement) {
pv_irq_ops.save_fl = __PV_IS_CALLEE_SAVE(xen_save_fl_direct);
pv_irq_ops.restore_fl = __PV_IS_CALLEE_SAVE(xen_restore_fl_direct);
pv_irq_ops.irq_disable = __PV_IS_CALLEE_SAVE(xen_irq_disable_direct);
@@ -1426,49 +1427,9 @@ static void __init xen_boot_params_init_edd(void)
* Set up the GDT and segment registers for -fstack-protector. Until
* we do this, we have to be careful not to call any stack-protected
* function, which is most of the kernel.
- *
- * Note, that it is __ref because the only caller of this after init
- * is PVH which is not going to use xen_load_gdt_boot or other
- * __init functions.
*/
-static void __ref xen_setup_gdt(int cpu)
+static void xen_setup_gdt(int cpu)
{
- if (xen_feature(XENFEAT_auto_translated_physmap)) {
-#ifdef CONFIG_X86_64
- unsigned long dummy;
-
- load_percpu_segment(cpu); /* We need to access per-cpu area */
- switch_to_new_gdt(cpu); /* GDT and GS set */
-
- /* We are switching of the Xen provided GDT to our HVM mode
- * GDT. The new GDT has __KERNEL_CS with CS.L = 1
- * and we are jumping to reload it.
- */
- asm volatile ("pushq %0\n"
- "leaq 1f(%%rip),%0\n"
- "pushq %0\n"
- "lretq\n"
- "1:\n"
- : "=&r" (dummy) : "0" (__KERNEL_CS));
-
- /*
- * While not needed, we also set the %es, %ds, and %fs
- * to zero. We don't care about %ss as it is NULL.
- * Strictly speaking this is not needed as Xen zeros those
- * out (and also MSR_FS_BASE, MSR_GS_BASE, MSR_KERNEL_GS_BASE)
- *
- * Linux zeros them in cpu_init() and in secondary_startup_64
- * (for BSP).
- */
- loadsegment(es, 0);
- loadsegment(ds, 0);
- loadsegment(fs, 0);
-#else
- /* PVH: TODO Implement. */
- BUG();
-#endif
- return; /* PVH does not need any PV GDT ops. */
- }
pv_cpu_ops.write_gdt_entry = xen_write_gdt_entry_boot;
pv_cpu_ops.load_gdt = xen_load_gdt_boot;
@@ -1479,59 +1440,6 @@ static void __ref xen_setup_gdt(int cpu)
pv_cpu_ops.load_gdt = xen_load_gdt;
}
-#ifdef CONFIG_XEN_PVH
-/*
- * A PV guest starts with default flags that are not set for PVH, set them
- * here asap.
- */
-static void xen_pvh_set_cr_flags(int cpu)
-{
-
- /* Some of these are setup in 'secondary_startup_64'. The others:
- * X86_CR0_TS, X86_CR0_PE, X86_CR0_ET are set by Xen for HVM guests
- * (which PVH shared codepaths), while X86_CR0_PG is for PVH. */
- write_cr0(read_cr0() | X86_CR0_MP | X86_CR0_NE | X86_CR0_WP | X86_CR0_AM);
-
- if (!cpu)
- return;
- /*
- * For BSP, PSE PGE are set in probe_page_size_mask(), for APs
- * set them here. For all, OSFXSR OSXMMEXCPT are set in fpu__init_cpu().
- */
- if (boot_cpu_has(X86_FEATURE_PSE))
- cr4_set_bits_and_update_boot(X86_CR4_PSE);
-
- if (boot_cpu_has(X86_FEATURE_PGE))
- cr4_set_bits_and_update_boot(X86_CR4_PGE);
-}
-
-/*
- * Note, that it is ref - because the only caller of this after init
- * is PVH which is not going to use xen_load_gdt_boot or other
- * __init functions.
- */
-void __ref xen_pvh_secondary_vcpu_init(int cpu)
-{
- xen_setup_gdt(cpu);
- xen_pvh_set_cr_flags(cpu);
-}
-
-static void __init xen_pvh_early_guest_init(void)
-{
- if (!xen_feature(XENFEAT_auto_translated_physmap))
- return;
-
- BUG_ON(!xen_feature(XENFEAT_hvm_callback_vector));
-
- xen_pvh_early_cpu_init(0, false);
- xen_pvh_set_cr_flags(0);
-
-#ifdef CONFIG_X86_32
- BUG(); /* PVH: Implement proper support. */
-#endif
-}
-#endif /* CONFIG_XEN_PVH */
-
static void __init xen_dom0_set_legacy_features(void)
{
x86_platform.legacy.rtc = 1;
@@ -1568,24 +1476,17 @@ asmlinkage __visible void __init xen_start_kernel(void)
xen_domain_type = XEN_PV_DOMAIN;
xen_setup_features();
-#ifdef CONFIG_XEN_PVH
- xen_pvh_early_guest_init();
-#endif
+
xen_setup_machphys_mapping();
/* Install Xen paravirt ops */
pv_info = xen_info;
pv_init_ops = xen_init_ops;
- if (!xen_pvh_domain()) {
- pv_cpu_ops = xen_cpu_ops;
+ pv_cpu_ops = xen_cpu_ops;
- x86_platform.get_nmi_reason = xen_get_nmi_reason;
- }
+ x86_platform.get_nmi_reason = xen_get_nmi_reason;
- if (xen_feature(XENFEAT_auto_translated_physmap))
- x86_init.resources.memory_setup = xen_auto_xlated_memory_setup;
- else
- x86_init.resources.memory_setup = xen_memory_setup;
+ x86_init.resources.memory_setup = xen_memory_setup;
x86_init.oem.arch_setup = xen_arch_setup;
x86_init.oem.banner = xen_banner;
@@ -1678,18 +1579,15 @@ asmlinkage __visible void __init xen_start_kernel(void)
/* set the limit of our address space */
xen_reserve_top();
- /* PVH: runs at default kernel iopl of 0 */
- if (!xen_pvh_domain()) {
- /*
- * We used to do this in xen_arch_setup, but that is too late
- * on AMD were early_cpu_init (run before ->arch_setup()) calls
- * early_amd_init which pokes 0xcf8 port.
- */
- set_iopl.iopl = 1;
- rc = HYPERVISOR_physdev_op(PHYSDEVOP_set_iopl, &set_iopl);
- if (rc != 0)
- xen_raw_printk("physdev_op failed %d\n", rc);
- }
+ /*
+ * We used to do this in xen_arch_setup, but that is too late
+ * on AMD were early_cpu_init (run before ->arch_setup()) calls
+ * early_amd_init which pokes 0xcf8 port.
+ */
+ set_iopl.iopl = 1;
+ rc = HYPERVISOR_physdev_op(PHYSDEVOP_set_iopl, &set_iopl);
+ if (rc != 0)
+ xen_raw_printk("physdev_op failed %d\n", rc);
#ifdef CONFIG_X86_32
/* set up basic CPUID stuff */
diff --git a/arch/x86/xen/mmu.c b/arch/x86/xen/mmu.c
index 7d5afdb..f6740b5 100644
--- a/arch/x86/xen/mmu.c
+++ b/arch/x86/xen/mmu.c
@@ -1792,10 +1792,6 @@ static void __init set_page_prot_flags(void *addr, pgprot_t prot,
unsigned long pfn = __pa(addr) >> PAGE_SHIFT;
pte_t pte = pfn_pte(pfn, prot);
- /* For PVH no need to set R/O or R/W to pin them or unpin them. */
- if (xen_feature(XENFEAT_auto_translated_physmap))
- return;
-
if (HYPERVISOR_update_va_mapping((unsigned long)addr, pte, flags))
BUG();
}
@@ -1902,8 +1898,7 @@ static void __init check_pt_base(unsigned long *pt_base, unsigned long *pt_end,
* level2_ident_pgt, and level2_kernel_pgt. This means that only the
* kernel has a physical mapping to start with - but that's enough to
* get __va working. We need to fill in the rest of the physical
- * mapping once some sort of allocator has been set up. NOTE: for
- * PVH, the page tables are native.
+ * mapping once some sort of allocator has been set up.
*/
void __init xen_setup_kernel_pagetable(pgd_t *pgd, unsigned long max_pfn)
{
@@ -2812,16 +2807,6 @@ static int do_remap_gfn(struct vm_area_struct *vma,
BUG_ON(!((vma->vm_flags & (VM_PFNMAP | VM_IO)) == (VM_PFNMAP | VM_IO)));
- if (xen_feature(XENFEAT_auto_translated_physmap)) {
-#ifdef CONFIG_XEN_PVH
- /* We need to update the local page tables and the xen HAP */
- return xen_xlate_remap_gfn_array(vma, addr, gfn, nr, err_ptr,
- prot, domid, pages);
-#else
- return -EINVAL;
-#endif
- }
-
rmd.mfn = gfn;
rmd.prot = prot;
/* We use the err_ptr to indicate if there we are doing a contiguous
@@ -2915,10 +2900,6 @@ int xen_unmap_domain_gfn_range(struct vm_area_struct *vma,
if (!pages || !xen_feature(XENFEAT_auto_translated_physmap))
return 0;
-#ifdef CONFIG_XEN_PVH
- return xen_xlate_unmap_gfn_range(vma, numpgs, pages);
-#else
return -EINVAL;
-#endif
}
EXPORT_SYMBOL_GPL(xen_unmap_domain_gfn_range);
diff --git a/arch/x86/xen/setup.c b/arch/x86/xen/setup.c
index f8960fc..6999016 100644
--- a/arch/x86/xen/setup.c
+++ b/arch/x86/xen/setup.c
@@ -915,39 +915,6 @@ char * __init xen_memory_setup(void)
}
/*
- * Machine specific memory setup for auto-translated guests.
- */
-char * __init xen_auto_xlated_memory_setup(void)
-{
- struct xen_memory_map memmap;
- int i;
- int rc;
-
- memmap.nr_entries = E820MAX;
- set_xen_guest_handle(memmap.buffer, xen_e820_map);
-
- rc = HYPERVISOR_memory_op(XENMEM_memory_map, &memmap);
- if (rc < 0)
- panic("No memory map (%d)\n", rc);
-
- xen_e820_map_entries = memmap.nr_entries;
-
- sanitize_e820_map(xen_e820_map, ARRAY_SIZE(xen_e820_map),
- &xen_e820_map_entries);
-
- for (i = 0; i < xen_e820_map_entries; i++)
- e820_add_region(xen_e820_map[i].addr, xen_e820_map[i].size,
- xen_e820_map[i].type);
-
- /* Remove p2m info, it is not needed. */
- xen_start_info->mfn_list = 0;
- xen_start_info->first_p2m_pfn = 0;
- xen_start_info->nr_p2m_frames = 0;
-
- return "Xen";
-}
-
-/*
* Set the bit indicating "nosegneg" library variants should be used.
* We only need to bother in pure 32-bit mode; compat 32-bit processes
* can have un-truncated segments, so wrapping around is allowed.
@@ -1032,8 +999,8 @@ void __init xen_pvmmu_arch_setup(void)
void __init xen_arch_setup(void)
{
xen_panic_handler_init();
- if (!xen_feature(XENFEAT_auto_translated_physmap))
- xen_pvmmu_arch_setup();
+
+ xen_pvmmu_arch_setup();
#ifdef CONFIG_ACPI
if (!(xen_start_info->flags & SIF_INITDOMAIN)) {
diff --git a/arch/x86/xen/smp.c b/arch/x86/xen/smp.c
index 9fa27ce..914e320 100644
--- a/arch/x86/xen/smp.c
+++ b/arch/x86/xen/smp.c
@@ -105,18 +105,8 @@ static void cpu_bringup(void)
local_irq_enable();
}
-/*
- * Note: cpu parameter is only relevant for PVH. The reason for passing it
- * is we can't do smp_processor_id until the percpu segments are loaded, for
- * which we need the cpu number! So we pass it in rdi as first parameter.
- */
-asmlinkage __visible void cpu_bringup_and_idle(int cpu)
+asmlinkage __visible void cpu_bringup_and_idle(void)
{
-#ifdef CONFIG_XEN_PVH
- if (xen_feature(XENFEAT_auto_translated_physmap) &&
- xen_feature(XENFEAT_supervisor_mode_kernel))
- xen_pvh_secondary_vcpu_init(cpu);
-#endif
cpu_bringup();
cpu_startup_entry(CPUHP_AP_ONLINE_IDLE);
}
@@ -410,61 +400,47 @@ static void __init xen_smp_prepare_cpus(unsigned int max_cpus)
gdt = get_cpu_gdt_table(cpu);
#ifdef CONFIG_X86_32
- /* Note: PVH is not yet supported on x86_32. */
ctxt->user_regs.fs = __KERNEL_PERCPU;
ctxt->user_regs.gs = __KERNEL_STACK_CANARY;
#endif
memset(&ctxt->fpu_ctxt, 0, sizeof(ctxt->fpu_ctxt));
- if (!xen_feature(XENFEAT_auto_translated_physmap)) {
- ctxt->user_regs.eip = (unsigned long)cpu_bringup_and_idle;
- ctxt->flags = VGCF_IN_KERNEL;
- ctxt->user_regs.eflags = 0x1000; /* IOPL_RING1 */
- ctxt->user_regs.ds = __USER_DS;
- ctxt->user_regs.es = __USER_DS;
- ctxt->user_regs.ss = __KERNEL_DS;
+ ctxt->user_regs.eip = (unsigned long)cpu_bringup_and_idle;
+ ctxt->flags = VGCF_IN_KERNEL;
+ ctxt->user_regs.eflags = 0x1000; /* IOPL_RING1 */
+ ctxt->user_regs.ds = __USER_DS;
+ ctxt->user_regs.es = __USER_DS;
+ ctxt->user_regs.ss = __KERNEL_DS;
- xen_copy_trap_info(ctxt->trap_ctxt);
+ xen_copy_trap_info(ctxt->trap_ctxt);
- ctxt->ldt_ents = 0;
+ ctxt->ldt_ents = 0;
- BUG_ON((unsigned long)gdt & ~PAGE_MASK);
+ BUG_ON((unsigned long)gdt & ~PAGE_MASK);
- gdt_mfn = arbitrary_virt_to_mfn(gdt);
- make_lowmem_page_readonly(gdt);
- make_lowmem_page_readonly(mfn_to_virt(gdt_mfn));
+ gdt_mfn = arbitrary_virt_to_mfn(gdt);
+ make_lowmem_page_readonly(gdt);
+ make_lowmem_page_readonly(mfn_to_virt(gdt_mfn));
- ctxt->gdt_frames[0] = gdt_mfn;
- ctxt->gdt_ents = GDT_ENTRIES;
+ ctxt->gdt_frames[0] = gdt_mfn;
+ ctxt->gdt_ents = GDT_ENTRIES;
- ctxt->kernel_ss = __KERNEL_DS;
- ctxt->kernel_sp = idle->thread.sp0;
+ ctxt->kernel_ss = __KERNEL_DS;
+ ctxt->kernel_sp = idle->thread.sp0;
#ifdef CONFIG_X86_32
- ctxt->event_callback_cs = __KERNEL_CS;
- ctxt->failsafe_callback_cs = __KERNEL_CS;
+ ctxt->event_callback_cs = __KERNEL_CS;
+ ctxt->failsafe_callback_cs = __KERNEL_CS;
#else
- ctxt->gs_base_kernel = per_cpu_offset(cpu);
-#endif
- ctxt->event_callback_eip =
- (unsigned long)xen_hypervisor_callback;
- ctxt->failsafe_callback_eip =
- (unsigned long)xen_failsafe_callback;
- ctxt->user_regs.cs = __KERNEL_CS;
- per_cpu(xen_cr3, cpu) = __pa(swapper_pg_dir);
- }
-#ifdef CONFIG_XEN_PVH
- else {
- /*
- * The vcpu comes on kernel page tables which have the NX pte
- * bit set. This means before DS/SS is touched, NX in
- * EFER must be set. Hence the following assembly glue code.
- */
- ctxt->user_regs.eip = (unsigned long)xen_pvh_early_cpu_init;
- ctxt->user_regs.rdi = cpu;
- ctxt->user_regs.rsi = true; /* entry == true */
- }
+ ctxt->gs_base_kernel = per_cpu_offset(cpu);
#endif
+ ctxt->event_callback_eip =
+ (unsigned long)xen_hypervisor_callback;
+ ctxt->failsafe_callback_eip =
+ (unsigned long)xen_failsafe_callback;
+ ctxt->user_regs.cs = __KERNEL_CS;
+ per_cpu(xen_cr3, cpu) = __pa(swapper_pg_dir);
+
ctxt->user_regs.esp = idle->thread.sp0 - sizeof(struct pt_regs);
ctxt->ctrlreg[3] = xen_pfn_to_cr3(virt_to_gfn(swapper_pg_dir));
if (HYPERVISOR_vcpu_op(VCPUOP_initialise, xen_vcpu_nr(cpu), ctxt))
diff --git a/arch/x86/xen/smp.h b/arch/x86/xen/smp.h
index c5c16dc..9beef33 100644
--- a/arch/x86/xen/smp.h
+++ b/arch/x86/xen/smp.h
@@ -21,12 +21,4 @@ static inline int xen_smp_intr_init(unsigned int cpu)
static inline void xen_smp_intr_free(unsigned int cpu) {}
#endif /* CONFIG_SMP */
-#ifdef CONFIG_XEN_PVH
-extern void xen_pvh_early_cpu_init(int cpu, bool entry);
-#else
-static inline void xen_pvh_early_cpu_init(int cpu, bool entry)
-{
-}
-#endif
-
#endif
diff --git a/arch/x86/xen/xen-head.S b/arch/x86/xen/xen-head.S
index 7f8d8ab..37794e4 100644
--- a/arch/x86/xen/xen-head.S
+++ b/arch/x86/xen/xen-head.S
@@ -16,25 +16,6 @@
#include <xen/interface/xen-mca.h>
#include <asm/xen/interface.h>
-#ifdef CONFIG_XEN_PVH
-#define PVH_FEATURES_STR "|writable_descriptor_tables|auto_translated_physmap|supervisor_mode_kernel"
-/* Note the lack of 'hvm_callback_vector'. Older hypervisor will
- * balk at this being part of XEN_ELFNOTE_FEATURES, so we put it in
- * XEN_ELFNOTE_SUPPORTED_FEATURES which older hypervisors will ignore.
- */
-#define PVH_FEATURES ((1 << XENFEAT_writable_page_tables) | \
- (1 << XENFEAT_auto_translated_physmap) | \
- (1 << XENFEAT_supervisor_mode_kernel) | \
- (1 << XENFEAT_hvm_callback_vector))
-/* The XENFEAT_writable_page_tables is not stricly necessary as we set that
- * up regardless whether this CONFIG option is enabled or not, but it
- * clarifies what the right flags need to be.
- */
-#else
-#define PVH_FEATURES_STR ""
-#define PVH_FEATURES (0)
-#endif
-
__INIT
ENTRY(startup_xen)
cld
@@ -54,41 +35,6 @@ ENTRY(startup_xen)
__FINIT
-#ifdef CONFIG_XEN_PVH
-/*
- * xen_pvh_early_cpu_init() - early PVH VCPU initialization
- * @cpu: this cpu number (%rdi)
- * @entry: true if this is a secondary vcpu coming up on this entry
- * point, false if this is the boot CPU being initialized for
- * the first time (%rsi)
- *
- * Note: This is called as a function on the boot CPU, and is the entry point
- * on the secondary CPU.
- */
-ENTRY(xen_pvh_early_cpu_init)
- mov %rsi, %r11
-
- /* Gather features to see if NX implemented. */
- mov $0x80000001, %eax
- cpuid
- mov %edx, %esi
-
- mov $MSR_EFER, %ecx
- rdmsr
- bts $_EFER_SCE, %eax
-
- bt $20, %esi
- jnc 1f /* No NX, skip setting it */
- bts $_EFER_NX, %eax
-1: wrmsr
-#ifdef CONFIG_SMP
- cmp $0, %r11b
- jne cpu_bringup_and_idle
-#endif
- ret
-
-#endif /* CONFIG_XEN_PVH */
-
.pushsection .text
.balign PAGE_SIZE
ENTRY(hypercall_page)
@@ -114,10 +60,10 @@ ENTRY(hypercall_page)
#endif
ELFNOTE(Xen, XEN_ELFNOTE_ENTRY, _ASM_PTR startup_xen)
ELFNOTE(Xen, XEN_ELFNOTE_HYPERCALL_PAGE, _ASM_PTR hypercall_page)
- ELFNOTE(Xen, XEN_ELFNOTE_FEATURES, .ascii "!writable_page_tables|pae_pgdir_above_4gb"; .asciz PVH_FEATURES_STR)
- ELFNOTE(Xen, XEN_ELFNOTE_SUPPORTED_FEATURES, .long (PVH_FEATURES) |
- (1 << XENFEAT_writable_page_tables) |
- (1 << XENFEAT_dom0))
+ ELFNOTE(Xen, XEN_ELFNOTE_FEATURES,
+ .ascii "!writable_page_tables|pae_pgdir_above_4gb")
+ ELFNOTE(Xen, XEN_ELFNOTE_SUPPORTED_FEATURES,
+ .long (1 << XENFEAT_writable_page_tables) | (1 << XENFEAT_dom0))
ELFNOTE(Xen, XEN_ELFNOTE_PAE_MODE, .asciz "yes")
ELFNOTE(Xen, XEN_ELFNOTE_LOADER, .asciz "generic")
ELFNOTE(Xen, XEN_ELFNOTE_L1_MFN_VALID,
diff --git a/arch/x86/xen/xen-ops.h b/arch/x86/xen/xen-ops.h
index 3cbce3b..3688ecc 100644
--- a/arch/x86/xen/xen-ops.h
+++ b/arch/x86/xen/xen-ops.h
@@ -146,5 +146,4 @@ static inline void __init xen_efi_init(void)
extern int xen_panic_handler_init(void);
-void xen_pvh_secondary_vcpu_init(int cpu);
#endif /* XEN_OPS_H */
diff --git a/drivers/xen/events/events_base.c b/drivers/xen/events/events_base.c
index 9ecfcdc..6a0aa5c 100644
--- a/drivers/xen/events/events_base.c
+++ b/drivers/xen/events/events_base.c
@@ -1706,7 +1706,6 @@ void __init xen_init_IRQ(void)
pirq_eoi_map = (void *)__get_free_page(GFP_KERNEL|__GFP_ZERO);
eoi_gmfn.gmfn = virt_to_gfn(pirq_eoi_map);
rc = HYPERVISOR_physdev_op(PHYSDEVOP_pirq_eoi_gmfn_v2, &eoi_gmfn);
- /* TODO: No PVH support for PIRQ EOI */
if (rc != 0) {
free_page((unsigned long) pirq_eoi_map);
pirq_eoi_map = NULL;
diff --git a/include/xen/xen.h b/include/xen/xen.h
index f0f0252..d0f9684 100644
--- a/include/xen/xen.h
+++ b/include/xen/xen.h
@@ -29,17 +29,6 @@ enum xen_domain_type {
#define xen_initial_domain() (0)
#endif /* CONFIG_XEN_DOM0 */
-#ifdef CONFIG_XEN_PVH
-/* This functionality exists only for x86. The XEN_PVHVM support exists
- * only in x86 world - hence on ARM it will be always disabled.
- * N.B. ARM guests are neither PV nor HVM nor PVHVM.
- * It's a bit like PVH but is different also (it's further towards the H
- * end of the spectrum than even PVH).
- */
-#include <xen/features.h>
-#define xen_pvh_domain() (xen_pv_domain() && \
- xen_feature(XENFEAT_auto_translated_physmap))
-#else
#define xen_pvh_domain() (0)
-#endif
+
#endif /* _XEN_XEN_H */
--
1.8.3.1
[toc] | [prev] | [next] | [standalone]
| From | Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> |
|---|---|
| Date | 2016-10-14 20:40 +0200 |
| Subject | Re: [Xen-devel] [PATCH 1/8] xen/x86: Remove PVH support |
| Message-ID | <ssfyG-4WS-5@gated-at.bofh.it> |
| In reply to | #1501123 |
On Fri, Oct 14, 2016 at 02:05:11PM -0400, Boris Ostrovsky wrote: > We are replacing existing PVH guests with new implementation. > > Signed-off-by: Boris Ostrovsky <boris.ostrovsky@oracle.com> Reviewed-by: Konrad Rzeszutek Wilk <konrad.wilk@oracle.com>
[toc] | [prev] | [next] | [standalone]
| From | Juergen Gross <jgross@suse.com> |
|---|---|
| Date | 2016-10-18 15:50 +0200 |
| Subject | Re: [PATCH 1/8] xen/x86: Remove PVH support |
| Message-ID | <stCWd-2lM-7@gated-at.bofh.it> |
| In reply to | #1501123 |
On 14/10/16 20:05, Boris Ostrovsky wrote:
> We are replacing existing PVH guests with new implementation.
>
> Signed-off-by: Boris Ostrovsky <boris.ostrovsky@oracle.com>
Reviewed-by: Juergen Gross <jgross@suse.com>
with the following addressed:
> diff --git a/include/xen/xen.h b/include/xen/xen.h
> index f0f0252..d0f9684 100644
> --- a/include/xen/xen.h
> +++ b/include/xen/xen.h
> @@ -29,17 +29,6 @@ enum xen_domain_type {
> #define xen_initial_domain() (0)
> #endif /* CONFIG_XEN_DOM0 */
>
> -#ifdef CONFIG_XEN_PVH
> -/* This functionality exists only for x86. The XEN_PVHVM support exists
> - * only in x86 world - hence on ARM it will be always disabled.
> - * N.B. ARM guests are neither PV nor HVM nor PVHVM.
> - * It's a bit like PVH but is different also (it's further towards the H
> - * end of the spectrum than even PVH).
> - */
> -#include <xen/features.h>
> -#define xen_pvh_domain() (xen_pv_domain() && \
> - xen_feature(XENFEAT_auto_translated_physmap))
> -#else
> #define xen_pvh_domain() (0)
Any reason you don't remove this, too (together with its last user in
arch/x86/xen/grant-table.c) ?
Juergen
[toc] | [prev] | [next] | [standalone]
| From | Boris Ostrovsky <boris.ostrovsky@oracle.com> |
|---|---|
| Date | 2016-10-18 16:50 +0200 |
| Subject | Re: [PATCH 1/8] xen/x86: Remove PVH support |
| Message-ID | <stDSi-31t-13@gated-at.bofh.it> |
| In reply to | #1502998 |
On 10/18/2016 09:46 AM, Juergen Gross wrote:
> On 14/10/16 20:05, Boris Ostrovsky wrote:
>> We are replacing existing PVH guests with new implementation.
>>
>> Signed-off-by: Boris Ostrovsky <boris.ostrovsky@oracle.com>
> Reviewed-by: Juergen Gross <jgross@suse.com>
>
> with the following addressed:
>
>> diff --git a/include/xen/xen.h b/include/xen/xen.h
>> index f0f0252..d0f9684 100644
>> --- a/include/xen/xen.h
>> +++ b/include/xen/xen.h
>> @@ -29,17 +29,6 @@ enum xen_domain_type {
>> #define xen_initial_domain() (0)
>> #endif /* CONFIG_XEN_DOM0 */
>>
>> -#ifdef CONFIG_XEN_PVH
>> -/* This functionality exists only for x86. The XEN_PVHVM support exists
>> - * only in x86 world - hence on ARM it will be always disabled.
>> - * N.B. ARM guests are neither PV nor HVM nor PVHVM.
>> - * It's a bit like PVH but is different also (it's further towards the H
>> - * end of the spectrum than even PVH).
>> - */
>> -#include <xen/features.h>
>> -#define xen_pvh_domain() (xen_pv_domain() && \
>> - xen_feature(XENFEAT_auto_translated_physmap))
>> -#else
>> #define xen_pvh_domain() (0)
> Any reason you don't remove this, too (together with its last user in
> arch/x86/xen/grant-table.c) ?
grant-table.c is in fact one of the reasons: we will be using that code
for PVHv2 again so I kept it to avoid unnecessary code churn.
Also, we want to have a nop definition of xen_pvh_domain() for
!CONFIG_XEN_PVH.
-boris
[toc] | [prev] | [next] | [standalone]
| From | Juergen Gross <jgross@suse.com> |
|---|---|
| Date | 2016-10-18 17:40 +0200 |
| Subject | Re: [PATCH 1/8] xen/x86: Remove PVH support |
| Message-ID | <stEEF-3BE-9@gated-at.bofh.it> |
| In reply to | #1503065 |
On 18/10/16 16:45, Boris Ostrovsky wrote:
> On 10/18/2016 09:46 AM, Juergen Gross wrote:
>> On 14/10/16 20:05, Boris Ostrovsky wrote:
>>> We are replacing existing PVH guests with new implementation.
>>>
>>> Signed-off-by: Boris Ostrovsky <boris.ostrovsky@oracle.com>
>> Reviewed-by: Juergen Gross <jgross@suse.com>
>>
>> with the following addressed:
>>
>>> diff --git a/include/xen/xen.h b/include/xen/xen.h
>>> index f0f0252..d0f9684 100644
>>> --- a/include/xen/xen.h
>>> +++ b/include/xen/xen.h
>>> @@ -29,17 +29,6 @@ enum xen_domain_type {
>>> #define xen_initial_domain() (0)
>>> #endif /* CONFIG_XEN_DOM0 */
>>>
>>> -#ifdef CONFIG_XEN_PVH
>>> -/* This functionality exists only for x86. The XEN_PVHVM support exists
>>> - * only in x86 world - hence on ARM it will be always disabled.
>>> - * N.B. ARM guests are neither PV nor HVM nor PVHVM.
>>> - * It's a bit like PVH but is different also (it's further towards the H
>>> - * end of the spectrum than even PVH).
>>> - */
>>> -#include <xen/features.h>
>>> -#define xen_pvh_domain() (xen_pv_domain() && \
>>> - xen_feature(XENFEAT_auto_translated_physmap))
>>> -#else
>>> #define xen_pvh_domain() (0)
>> Any reason you don't remove this, too (together with its last user in
>> arch/x86/xen/grant-table.c) ?
>
> grant-table.c is in fact one of the reasons: we will be using that code
> for PVHv2 again so I kept it to avoid unnecessary code churn.
>
> Also, we want to have a nop definition of xen_pvh_domain() for
> !CONFIG_XEN_PVH.
Okay, could you mention this in the commit message, please?
Juergen
[toc] | [prev] | [next] | [standalone]
| From | Boris Ostrovsky <boris.ostrovsky@oracle.com> |
|---|---|
| Date | 2016-10-18 17:50 +0200 |
| Subject | Re: [PATCH 1/8] xen/x86: Remove PVH support |
| Message-ID | <stEOm-3Fs-1@gated-at.bofh.it> |
| In reply to | #1503111 |
On 10/18/2016 11:33 AM, Juergen Gross wrote:
> On 18/10/16 16:45, Boris Ostrovsky wrote:
>> On 10/18/2016 09:46 AM, Juergen Gross wrote:
>>> On 14/10/16 20:05, Boris Ostrovsky wrote:
>>>> We are replacing existing PVH guests with new implementation.
>>>>
>>>> Signed-off-by: Boris Ostrovsky <boris.ostrovsky@oracle.com>
>>> Reviewed-by: Juergen Gross <jgross@suse.com>
>>>
>>> with the following addressed:
>>>
>>>> diff --git a/include/xen/xen.h b/include/xen/xen.h
>>>> index f0f0252..d0f9684 100644
>>>> --- a/include/xen/xen.h
>>>> +++ b/include/xen/xen.h
>>>> @@ -29,17 +29,6 @@ enum xen_domain_type {
>>>> #define xen_initial_domain() (0)
>>>> #endif /* CONFIG_XEN_DOM0 */
>>>>
>>>> -#ifdef CONFIG_XEN_PVH
>>>> -/* This functionality exists only for x86. The XEN_PVHVM support exists
>>>> - * only in x86 world - hence on ARM it will be always disabled.
>>>> - * N.B. ARM guests are neither PV nor HVM nor PVHVM.
>>>> - * It's a bit like PVH but is different also (it's further towards the H
>>>> - * end of the spectrum than even PVH).
>>>> - */
>>>> -#include <xen/features.h>
>>>> -#define xen_pvh_domain() (xen_pv_domain() && \
>>>> - xen_feature(XENFEAT_auto_translated_physmap))
>>>> -#else
>>>> #define xen_pvh_domain() (0)
>>> Any reason you don't remove this, too (together with its last user in
>>> arch/x86/xen/grant-table.c) ?
>> grant-table.c is in fact one of the reasons: we will be using that code
>> for PVHv2 again so I kept it to avoid unnecessary code churn.
>>
>> Also, we want to have a nop definition of xen_pvh_domain() for
>> !CONFIG_XEN_PVH.
> Okay, could you mention this in the commit message, please?
Will do.
-boris
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web