Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1372156 > unrolled thread

HVMLite / PVHv2 - using x86 EFI boot entry

Started by"Luis R. Rodriguez" <mcgrof@kernel.org>
First post2016-04-06 04:50 +0200
Last post2016-04-09 19:10 +0200
Articles 20 on this page of 56 — 11 participants

Back to article view | Back to linux.kernel


Contents

  HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-06 04:50 +0200
    Re: HVMLite / PVHv2 - using x86 EFI boot entry David Vrabel <david.vrabel@citrix.com> - 2016-04-06 11:50 +0200
      Re: HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-08 22:50 +0200
        Re: HVMLite / PVHv2 - using x86 EFI boot entry Juergen Gross <jgross@suse.com> - 2016-04-11 07:20 +0200
          Re: HVMLite / PVHv2 - using x86 EFI boot entry Andy Lutomirski <luto@amacapital.net> - 2016-04-12 23:10 +0200
            Re: HVMLite / PVHv2 - using x86 EFI boot entry Roger Pau Monné <roger.pau@citrix.com> - 2016-04-13 11:10 +0200
              Re: HVMLite / PVHv2 - using x86 EFI boot entry Matt Fleming <matt@codeblueprint.co.uk> - 2016-04-13 12:20 +0200
                Re: HVMLite / PVHv2 - using x86 EFI boot entry Matt Fleming <matt@codeblueprint.co.uk> - 2016-04-13 12:50 +0200
                Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry George Dunlap <george.dunlap@citrix.com> - 2016-04-13 13:20 +0200
                Re: HVMLite / PVHv2 - using x86 EFI boot entry Roger Pau Monné <roger.pau@citrix.com> - 2016-04-13 14:00 +0200
          Re: HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-13 20:40 +0200
            Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-13 22:50 +0200
              Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-14 00:30 +0200
                Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-04-14 03:10 +0200
                  Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-14 20:50 +0200
                    Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-04-14 22:00 +0200
                      Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-14 23:00 +0200
                        Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-04-15 04:10 +0200
                        Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Julien Grall <julien.grall@arm.com> - 2016-04-15 12:10 +0200
                          Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-15 17:00 +0200
    Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry George Dunlap <george.dunlap@citrix.com> - 2016-04-06 13:10 +0200
      Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Matt Fleming <matt@codeblueprint.co.uk> - 2016-04-06 17:10 +0200
        Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-09 00:00 +0200
        Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Roger Pau Monné <roger.pau@citrix.com> - 2016-04-13 12:10 +0200
          Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Matt Fleming <matt@codeblueprint.co.uk> - 2016-04-13 12:30 +0200
      Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@suse.com> - 2016-04-07 21:00 +0200
        Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry George Dunlap <george.dunlap@citrix.com> - 2016-04-08 16:20 +0200
          Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-09 00:00 +0200
            Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-13 00:20 +0200
              Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry George Dunlap <george.dunlap@citrix.com> - 2016-04-13 12:10 +0200
                Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-13 21:00 +0200
                  Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry George Dunlap <george.dunlap@citrix.com> - 2016-04-14 11:50 +0200
                    Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-14 22:00 +0200
              Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Roger Pau Monné <roger.pau@citrix.com> - 2016-04-13 12:30 +0200
                Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-13 21:20 +0200
            Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Roger Pau Monné <roger.pau@citrix.com> - 2016-04-13 12:00 +0200
              Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-13 21:00 +0200
                Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-13 21:20 +0200
                  Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-13 22:10 +0200
                    Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-13 22:40 +0200
                  Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry George Dunlap <george.dunlap@citrix.com> - 2016-04-14 12:20 +0200
        Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry George Dunlap <george.dunlap@citrix.com> - 2016-04-13 18:00 +0200
          Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-13 22:00 +0200
            Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry George Dunlap <george.dunlap@citrix.com> - 2016-04-14 12:00 +0200
              Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-14 21:50 +0200
                Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-04-14 22:50 +0200
                  Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-14 23:20 +0200
                    Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-04-15 04:20 +0200
                Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry Juergen Gross <jgross@suse.com> - 2016-04-15 08:00 +0200
                  Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-15 17:30 +0200
                Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry George Dunlap <george.dunlap@citrix.com> - 2016-04-15 12:00 +0200
                  Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-15 17:40 +0200
                    Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry George Dunlap <george.dunlap@citrix.com> - 2016-04-15 18:10 +0200
    Re: HVMLite / PVHv2 - using x86 EFI boot entry Daniel Kiper <daniel.kiper@oracle.com> - 2016-04-06 13:20 +0200
      Re: HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-07 21:20 +0200
      Re: HVMLite / PVHv2 - using x86 EFI boot entry "Luis R. Rodriguez" <mcgrof@kernel.org> - 2016-04-09 19:10 +0200

Page 1 of 3  [1] 2 3  Next page →


#1372156 — HVMLite / PVHv2 - using x86 EFI boot entry

From"Luis R. Rodriguez" <mcgrof@kernel.org>
Date2016-04-06 04:50 +0200
SubjectHVMLite / PVHv2 - using x86 EFI boot entry
Message-ID<rkLHz-5vj-5@gated-at.bofh.it>
Boris sent out the first HVMLite series of patches to add a new Xen guest type
February 1, 2016 [0]. We've been talking off list with a few folks now over
the prospect of instead of adding yet-another-boot-entry we instead fixate
HVMLite to use the x86 EFI boot entry. There's a series of reasons to consider
this, likewise there are reasons to question the effort required and if its
really needed. We'd like some more public review of this proposal, and see if
others can come up with other ideas, both in favor or against this proposal.

This in particular is also a good time to get x86 Linux folks to chime on on
the general design proposal of HVMLite design, given that outside of the boot
entry discussion it would seem including myself that we didn't get the memo
over the proposed architecture review [1]. At least on my behalf perhaps the
only sticking thorns of the design was the new boot entry, which came to me
as a surprise, and this thread addresses and the lack of addressing semantics 
for early boot (which we may seem to need to address; some of this is being
addressing in parallels through other work). The HVMLite document talks about
using ACPI_FADT_NO_VGA -- we don't use this yet upstream but I have some pending
changes which should make it easy to integrate its use on HVMLite. Perhaps
there are others that may have some other points they may want to raise now...

A huge summary of the discussion over EFI boot option for HVMLite is now on a
wiki [2], below I'll just provide the outline of the discussion. Consider this a
request for more public review, feel free to take any of the items below and
elaborate on it as you see fit.

Worth mentioning also is that this topic will be discussed at the 2016 Xen
Hackathon April 18-19 [3] at the ARM Cambridge, UK Headquarters so if you can
attend and this topic interests you, consider attending.

  * Linux x86 Xen EFI boot entry evaluation
  * Issues with boot x86 boot entries
    * Bypassing native startup_32() / startup_64()
    * Small x86 zero page stubs

  * Xen evolution and roadmap
    * About PVH
    * About HVMLite
    * Xen ARM solution

  * Why use EFI for HVMlite
    * EFI calling conventions are standardized
    * EFI entry generalizes what new HVMLite entry proposes
    * Further semantics may be needed
    * Match Xen ARM's clean solution
    * You don't need full EFI emulation
      * Minimal EFI stubs for guests
        * GetMemoryMap()
        * ExitBootServices()
      * EFI stubs which may be needed for guests
        * Exit()
        * Variable operation functions
      * EFI stubs not needed for guests
        * GetTime()/SetTime()
        * SetVirtualAddressMap()
        * ResetSystem()
      * dom0 EFI
      * domU EFI emulation possibilities
        * Xen implements its own EFI environment for guests
        * Xen uses Tianocore / OVMF
    * kexec needs a boot path as well

  * Points against using EFI
    * Legacy PV guests need to be supported
    * Nulling the claimed boot loader effect
    * startup_32 / startup_64 flexibility
  * Remaining questions

[0] http://lkml.kernel.org/r/1454341137-14110-3-git-send-email-boris.ostrovsky@oracle.com
[1] http://lists.xen.org/archives/html/xen-devel/2016-02/msg01609.html
[2] http://kernelnewbies.org/KernelProjects/x86-xen-efi
[3] http://wiki.xenproject.org/wiki/Hackathon/April2016

  Luis

[toc] | [next] | [standalone]


#1372339

FromDavid Vrabel <david.vrabel@citrix.com>
Date2016-04-06 11:50 +0200
Message-ID<rkSg1-1O6-7@gated-at.bofh.it>
In reply to#1372156
On 06/04/16 03:40, Luis R. Rodriguez wrote:
> 
>     * You don't need full EFI emulation

I think needing any EFI emulation inside Xen (which is where it would
need to be for dom0) is not suitable because of the increase in
hypervisor ABI.

I also still do not understand your objection to the current tiny stub.

David

[toc] | [prev] | [next] | [standalone]


#1374459

From"Luis R. Rodriguez" <mcgrof@kernel.org>
Date2016-04-08 22:50 +0200
Message-ID<rlLvR-133-9@gated-at.bofh.it>
In reply to#1372339
On Wed, Apr 06, 2016 at 10:40:08AM +0100, David Vrabel wrote:
> On 06/04/16 03:40, Luis R. Rodriguez wrote:
> > 
> >     * You don't need full EFI emulation
> 
> I think needing any EFI emulation inside Xen (which is where it would
> need to be for dom0) is not suitable because of the increase in
> hypervisor ABI.

Is this because of timing on architecture / design of HVMLite, or
a general position that the complexity to deal with EFI emulation
is too much for Xen's taste ?

ARM already went the EFI entry way for domU -- it went the OVMF route,
would such a possibility be possible for x86 domU HVMLite ? If not why
not, I mean it would seem to make sense to at least mimic the same type
of early boot environment, and perhaps there are some lessons to be
learned from that effort too.

Are there some lessons to be learned with ARM's effort? What are they?
If that could be re-done again with any type of cleaner path, what
could that be that could help the x86 side ?

Although emulating EFI may require work, some folks have pointed out
that the amount of work may not be that much. If that is done can
we instead rely on the same code to replace OVMF to support both
Xen ARM and Xen HVMLite on x86 ? What would be the pros / cons of
this ?

> I also still do not understand your objection to the current tiny stub.

Its more of a hypothetical -- can an EFI entry be used instead given
it already does exactly what the new small entry does ? Its also rather
odd to add a new entry without evaluating fully a possible alternative
that would provide the same exact mechanism.

A full technical unbiased evaluation of the different approaches is what I'd
hope we could strive to achieve through discussion and peer review, thinking
and prioritizing ultimately what is best to minimize the impact on Linux
and also help take advantage of the best features possible through both
means. Thinking long term, not immediate short term.

  Luis

[toc] | [prev] | [next] | [standalone]


#1375496

FromJuergen Gross <jgross@suse.com>
Date2016-04-11 07:20 +0200
Message-ID<rmCqt-vD-5@gated-at.bofh.it>
In reply to#1374459
On 08/04/16 22:40, Luis R. Rodriguez wrote:
> On Wed, Apr 06, 2016 at 10:40:08AM +0100, David Vrabel wrote:
>> On 06/04/16 03:40, Luis R. Rodriguez wrote:
>>>
>>>     * You don't need full EFI emulation
>>
>> I think needing any EFI emulation inside Xen (which is where it would
>> need to be for dom0) is not suitable because of the increase in
>> hypervisor ABI.
> 
> Is this because of timing on architecture / design of HVMLite, or
> a general position that the complexity to deal with EFI emulation
> is too much for Xen's taste ?

The Xen hypervisor should be as small as possible. Adding an EFI
emulator will be adding quite some code. This should be done after a
very thorough evaluation only.

> ARM already went the EFI entry way for domU -- it went the OVMF route,
> would such a possibility be possible for x86 domU HVMLite ? If not why
> not, I mean it would seem to make sense to at least mimic the same type
> of early boot environment, and perhaps there are some lessons to be
> learned from that effort too.

The final solution must be appropriate for dom0, too. So don't try
to limit the discussion to domU. If dom0 isn't going to be acceptable
there will no need to discuss domU.

> Are there some lessons to be learned with ARM's effort? What are they?
> If that could be re-done again with any type of cleaner path, what
> could that be that could help the x86 side ?
> 
> Although emulating EFI may require work, some folks have pointed out
> that the amount of work may not be that much. If that is done can
> we instead rely on the same code to replace OVMF to support both
> Xen ARM and Xen HVMLite on x86 ? What would be the pros / cons of
> this ?
> 
>> I also still do not understand your objection to the current tiny stub.
> 
> Its more of a hypothetical -- can an EFI entry be used instead given
> it already does exactly what the new small entry does ? Its also rather
> odd to add a new entry without evaluating fully a possible alternative
> that would provide the same exact mechanism.

The interface isn't the new entry only. It should be evaluated how much
of the early EFI boot path would be common to the HVMlite one. What
would be gained by using the same entry but having two different boot
paths after it? You still need a way to distinguish between bare metal
EFI and HVMlite. And Xen needs a way to find out whether a kernel is
supporting HVMlite to boot it in the correct mode.

> A full technical unbiased evaluation of the different approaches is what I'd
> hope we could strive to achieve through discussion and peer review, thinking
> and prioritizing ultimately what is best to minimize the impact on Linux
> and also help take advantage of the best features possible through both
> means. Thinking long term, not immediate short term.

Sure.


Juergen

[toc] | [prev] | [next] | [standalone]


#1377238

FromAndy Lutomirski <luto@amacapital.net>
Date2016-04-12 23:10 +0200
Message-ID<rndJq-5Cv-67@gated-at.bofh.it>
In reply to#1375496
On Sun, Apr 10, 2016 at 10:12 PM, Juergen Gross <jgross@suse.com> wrote:
> On 08/04/16 22:40, Luis R. Rodriguez wrote:
>> On Wed, Apr 06, 2016 at 10:40:08AM +0100, David Vrabel wrote:
>>> On 06/04/16 03:40, Luis R. Rodriguez wrote:
>>>>
>>>>     * You don't need full EFI emulation
>>>
>>> I think needing any EFI emulation inside Xen (which is where it would
>>> need to be for dom0) is not suitable because of the increase in
>>> hypervisor ABI.
>>
>> Is this because of timing on architecture / design of HVMLite, or
>> a general position that the complexity to deal with EFI emulation
>> is too much for Xen's taste ?
>
> The Xen hypervisor should be as small as possible. Adding an EFI
> emulator will be adding quite some code. This should be done after a
> very thorough evaluation only.
>
>> ARM already went the EFI entry way for domU -- it went the OVMF route,
>> would such a possibility be possible for x86 domU HVMLite ? If not why
>> not, I mean it would seem to make sense to at least mimic the same type
>> of early boot environment, and perhaps there are some lessons to be
>> learned from that effort too.
>
> The final solution must be appropriate for dom0, too. So don't try
> to limit the discussion to domU. If dom0 isn't going to be acceptable
> there will no need to discuss domU.
>
>> Are there some lessons to be learned with ARM's effort? What are they?
>> If that could be re-done again with any type of cleaner path, what
>> could that be that could help the x86 side ?
>>
>> Although emulating EFI may require work, some folks have pointed out
>> that the amount of work may not be that much. If that is done can
>> we instead rely on the same code to replace OVMF to support both
>> Xen ARM and Xen HVMLite on x86 ? What would be the pros / cons of
>> this ?
>>
>>> I also still do not understand your objection to the current tiny stub.
>>
>> Its more of a hypothetical -- can an EFI entry be used instead given
>> it already does exactly what the new small entry does ? Its also rather
>> odd to add a new entry without evaluating fully a possible alternative
>> that would provide the same exact mechanism.
>
> The interface isn't the new entry only. It should be evaluated how much
> of the early EFI boot path would be common to the HVMlite one. What
> would be gained by using the same entry but having two different boot
> paths after it? You still need a way to distinguish between bare metal
> EFI and HVMlite. And Xen needs a way to find out whether a kernel is
> supporting HVMlite to boot it in the correct mode.
>
>> A full technical unbiased evaluation of the different approaches is what I'd
>> hope we could strive to achieve through discussion and peer review, thinking
>> and prioritizing ultimately what is best to minimize the impact on Linux
>> and also help take advantage of the best features possible through both
>> means. Thinking long term, not immediate short term.
>
> Sure.

FWIW, someone just pointed me to u-boot's EFI implementation.
u-boot's lib/efi_loader contains a tiny (<3k LOC, 10kB compiled) UEFI
implementation that's sufficient to boot a Linux EFI payload.

An argument against making Xen's default domU entry use UEFI is that
it might become unnecessarily awkward to do something like
chainloading to OVMF.   But maybe OVMF can be compiled as a UEFI
binary :)

--Andy

[toc] | [prev] | [next] | [standalone]


#1377670

FromRoger Pau Monné <roger.pau@citrix.com>
Date2016-04-13 11:10 +0200
Message-ID<rnoYa-7qk-5@gated-at.bofh.it>
In reply to#1377238
On Tue, Apr 12, 2016 at 02:02:52PM -0700, Andy Lutomirski wrote:
> On Sun, Apr 10, 2016 at 10:12 PM, Juergen Gross <jgross@suse.com> wrote:
> > On 08/04/16 22:40, Luis R. Rodriguez wrote:
> >> On Wed, Apr 06, 2016 at 10:40:08AM +0100, David Vrabel wrote:
> >>> On 06/04/16 03:40, Luis R. Rodriguez wrote:
> >>>>
> >>>>     * You don't need full EFI emulation
> >>>
> >>> I think needing any EFI emulation inside Xen (which is where it would
> >>> need to be for dom0) is not suitable because of the increase in
> >>> hypervisor ABI.
> >>
> >> Is this because of timing on architecture / design of HVMLite, or
> >> a general position that the complexity to deal with EFI emulation
> >> is too much for Xen's taste ?
> >
> > The Xen hypervisor should be as small as possible. Adding an EFI
> > emulator will be adding quite some code. This should be done after a
> > very thorough evaluation only.
> >
> >> ARM already went the EFI entry way for domU -- it went the OVMF route,
> >> would such a possibility be possible for x86 domU HVMLite ? If not why
> >> not, I mean it would seem to make sense to at least mimic the same type
> >> of early boot environment, and perhaps there are some lessons to be
> >> learned from that effort too.
> >
> > The final solution must be appropriate for dom0, too. So don't try
> > to limit the discussion to domU. If dom0 isn't going to be acceptable
> > there will no need to discuss domU.
> >
> >> Are there some lessons to be learned with ARM's effort? What are they?
> >> If that could be re-done again with any type of cleaner path, what
> >> could that be that could help the x86 side ?
> >>
> >> Although emulating EFI may require work, some folks have pointed out
> >> that the amount of work may not be that much. If that is done can
> >> we instead rely on the same code to replace OVMF to support both
> >> Xen ARM and Xen HVMLite on x86 ? What would be the pros / cons of
> >> this ?
> >>
> >>> I also still do not understand your objection to the current tiny stub.
> >>
> >> Its more of a hypothetical -- can an EFI entry be used instead given
> >> it already does exactly what the new small entry does ? Its also rather
> >> odd to add a new entry without evaluating fully a possible alternative
> >> that would provide the same exact mechanism.
> >
> > The interface isn't the new entry only. It should be evaluated how much
> > of the early EFI boot path would be common to the HVMlite one. What
> > would be gained by using the same entry but having two different boot
> > paths after it? You still need a way to distinguish between bare metal
> > EFI and HVMlite. And Xen needs a way to find out whether a kernel is
> > supporting HVMlite to boot it in the correct mode.
> >
> >> A full technical unbiased evaluation of the different approaches is what I'd
> >> hope we could strive to achieve through discussion and peer review, thinking
> >> and prioritizing ultimately what is best to minimize the impact on Linux
> >> and also help take advantage of the best features possible through both
> >> means. Thinking long term, not immediate short term.
> >
> > Sure.
> 
> FWIW, someone just pointed me to u-boot's EFI implementation.
> u-boot's lib/efi_loader contains a tiny (<3k LOC, 10kB compiled) UEFI
> implementation that's sufficient to boot a Linux EFI payload.

I guess this is a pretty minimal EFI implementation, is this something 
standard, or just an EFI implementation tailored to Linux needs? (ie: is 
there any standard EFI flag to signal this kind of minimal EFI environment?)
 
> An argument against making Xen's default domU entry use UEFI is that
> it might become unnecessarily awkward to do something like
> chainloading to OVMF.   But maybe OVMF can be compiled as a UEFI
> binary :)

With my FreeBSD committer hat:

The FreeBSD kernel doesn't contain an EFI entry point, it just contains one 
single entry point that's used for both legacy BIOS and EFI. Then the 
FreeBSD loader is the one that contains the different entry points. I would 
really like to avoid adding an EFI entry point and the PE header to the 
FreeBSD kernel. The current trampoline in FreeBSD to tie the Xen entry point 
into the native path contains 96 lines of assembly (half of them are 
actually comments) and 66 lines of C. I think adding an EFI entry point is 
going to add a lot more of code than this, and we would probably need 
changes to the build system in order to assembly the PE header and the ELF 
headers together.

IMHO, if we want to boot PVH using EFI the right solution is to use OVMF (or 
any other UEFI firmware) and port it so it's able to run as a PVH guest. I 
guess it should even be possible to use it for Dom0, although I think this 
is cumbersome.

Roger.

[toc] | [prev] | [next] | [standalone]


#1377735

FromMatt Fleming <matt@codeblueprint.co.uk>
Date2016-04-13 12:20 +0200
Message-ID<rnq3U-8a8-17@gated-at.bofh.it>
In reply to#1377670
On Wed, 13 Apr, at 11:02:02AM, Roger Pau Monné wrote:
> 
> With my FreeBSD committer hat:
> 
> The FreeBSD kernel doesn't contain an EFI entry point, it just contains one 
> single entry point that's used for both legacy BIOS and EFI. Then the 
> FreeBSD loader is the one that contains the different entry points. I would 
> really like to avoid adding an EFI entry point and the PE header to the 
> FreeBSD kernel. The current trampoline in FreeBSD to tie the Xen entry point 
> into the native path contains 96 lines of assembly (half of them are 
> actually comments) and 66 lines of C. I think adding an EFI entry point is 
> going to add a lot more of code than this, and we would probably need 
> changes to the build system in order to assembly the PE header and the ELF 
> headers together.
 
What does the boot flow look like for PVH2 on FreeBSD today?
Presumably it doesn't have the same entry point that Boris proposed
for Linux?

Does it go, Hypervisor -> FreeBSD loader -> FreeBSD kernel? Or are you
able to directly boot the kernel from the hypervisor and skip the
middle part by having secondary entry point for Xen marked by the ELF
note?

> IMHO, if we want to boot PVH using EFI the right solution is to use OVMF (or 
> any other UEFI firmware) and port it so it's able to run as a PVH guest. I 
> guess it should even be possible to use it for Dom0, although I think this 
> is cumbersome.

There are two levels of EFI boot entry features being discussed,

 1. Make the OS kernel a PE/COFF executable
 2. Provide some level of EFI service functionality

You can adopt 1. without 2, i.e. without actually providing any EFI
services at all, as long as the Xen hypervisor grows a PE/COFF loader
(since EFI firmware has to provide you one, for EFI platforms you
could use the LoadImage() service in the firmware, but for BIOS
platforms you'd need your own in Xen).

On Linux, this has the advantage of deferring the decompression of the
bzImage (x86 Linux kernel file format) to the stub on the front of the
bzImage. And while I realise that the toolstack already has support
for decompressing bzImages, given what Andrew has said about reducing
attack surface, having the guest perform the decompression should be a
win.

Of course, this is offset somewhat by the fact that you need to audit
the PE/COFF loader ;) But decompression in general is notoriously
vulnerable to security issues.

Using the in-kernel decompressor is how most (all?) Linux boot loaders
work today, so there's the added benefit of reducing the differences
between booting on Xen and booting bare metal. For example, you'd
probably be able to use CONFIG_RANDOMIZE_BASE (ASLR for kernel image)
for Xen if you use the kernel's decompressor. Xen would also get
future features in this area for free, and there is a tendency to push
boot features into the early stub.

For 1. we'd basically be using the PE/COFF file format with the EFI
ABI as an OS agnostic boot protocol, but not as a full firmware
runtime environment.

2. is also interesting, though I think less so than 1. I agree that
making OVMF work as a PVH guest is probably the right way to go, even
for Dom0, not least because you'd have a much cleaner/less buggy
implementation than what we see in the real world ;)

[toc] | [prev] | [next] | [standalone]


#1377754

FromMatt Fleming <matt@codeblueprint.co.uk>
Date2016-04-13 12:50 +0200
Message-ID<rnqwW-8nP-27@gated-at.bofh.it>
In reply to#1377735
On Wed, 13 Apr, at 11:15:15AM, Matt Fleming wrote:
> 
> For 1. we'd basically be using the PE/COFF file format with the EFI
> ABI as an OS agnostic boot protocol, but not as a full firmware
> runtime environment.

To add some balance to this proposal (since there's no such thing as a
free lunch) some of the disadvantages are,

The PE/COFF stub in Linux does assume that it is executing in native
cpu mode and does not perform any mode switching, i.e. from 32-bit
protected to long mode. This is due to the way that EFI works - by the
time the OS image entry point is jumped to on a 64-bit cpu we're
running in long mode with identity mapped page tables. To be fair,
when running Xen on EFI (bare metal) this would save you one cpu mode
switch when compared with the current HVMLite proposal.

I'm not aware of a direct equivalent for ELF notes in the PE/COFF
format. I'm still re-reading the spec to find something suitable.

[toc] | [prev] | [next] | [standalone]


#1377770 — Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry

FromGeorge Dunlap <george.dunlap@citrix.com>
Date2016-04-13 13:20 +0200
SubjectRe: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry
Message-ID<rnqZY-n3-11@gated-at.bofh.it>
In reply to#1377735
On Wed, Apr 13, 2016 at 11:15 AM, Matt Fleming <matt@codeblueprint.co.uk> wrote:
> For 1. we'd basically be using the PE/COFF file format with the EFI
> ABI as an OS agnostic boot protocol, but not as a full firmware
> runtime environment.

But we still have the issue here that the now the EFI entry point in
Linux has to figure out, "Am I running in a full firmware runtime
environment, or am I running under Xen?", and then change behavior
appropriately.  Then we get back to Juergen's comment:  "[The EFI
proposal] should be evaluated how much of the early EFI boot path
would be common to the HVMlite one. What would be gained by using the
same entry but having two different boot paths after it?"

> 2. is also interesting, though I think less so than 1. I agree that
> making OVMF work as a PVH guest is probably the right way to go, even
> for Dom0, not least because you'd have a much cleaner/less buggy
> implementation than what we see in the real world ;)

So rather than just add an extra entry point and a Xen-to-zero-page
stub, you're going to ask Xen on dom0 to import a full OVMF binary?
Or have the bootloader entries include xen, linux, the initrd, *and*
ovmf?  That seems a bit extreme. :-)

Keep in mind also that PVH needs to support not only the traditional
VM use-case (e.g., booting a full distro), but the small service VM
usecase (a la unikernels).  Booting a traditional distro as a domU via
OVMF -> EFI Linux makes sense; it reduces the distro's test burden,
and the OVMF doesn't add a lot to the memory or boot time compared to
the size and boot time of a full distro.  But booting tiny service
VMs, sometimes with not even any disk of their own (other than a
ramdisk), the extra cost of including OVMF in the guest address space
can be a non-negligible addition to the memory requirements and
boot-up time.

One of the reasons Xen on ARM prioritized getting EFI working for
domUs was that a representative from a certain distro vendor made it
absolutely clear that *their* distro would *only* support booting via
EFI on ARM.  But you can still, as I understand it, use uBoot with DT
to boot a lightweight domU if you want.

 -George

[toc] | [prev] | [next] | [standalone]


#1377851

FromRoger Pau Monné <roger.pau@citrix.com>
Date2016-04-13 14:00 +0200
Message-ID<rnrCF-Gd-7@gated-at.bofh.it>
In reply to#1377735
On Wed, Apr 13, 2016 at 11:15:15AM +0100, Matt Fleming wrote:
> On Wed, 13 Apr, at 11:02:02AM, Roger Pau Monné wrote:
> > 
> > With my FreeBSD committer hat:
> > 
> > The FreeBSD kernel doesn't contain an EFI entry point, it just contains one 
> > single entry point that's used for both legacy BIOS and EFI. Then the 
> > FreeBSD loader is the one that contains the different entry points. I would 
> > really like to avoid adding an EFI entry point and the PE header to the 
> > FreeBSD kernel. The current trampoline in FreeBSD to tie the Xen entry point 
> > into the native path contains 96 lines of assembly (half of them are 
> > actually comments) and 66 lines of C. I think adding an EFI entry point is 
> > going to add a lot more of code than this, and we would probably need 
> > changes to the build system in order to assembly the PE header and the ELF 
> > headers together.
>  
> What does the boot flow look like for PVH2 on FreeBSD today?
> Presumably it doesn't have the same entry point that Boris proposed
> for Linux?

Yes it does have something quite similar to the entry point that Boris 
proposed for Linux.
 
> Does it go, Hypervisor -> FreeBSD loader -> FreeBSD kernel? Or are you
> able to directly boot the kernel from the hypervisor and skip the
> middle part by having secondary entry point for Xen marked by the ELF
> note?

We skip the bootloader and Xen loads the FreeBSD kernel directly using the 
ELF note that contains the PVH entry point.

I certainly want to be able to run the FreeBSD loader inside of a PVH guest, 
but I plan to simply chainload it from OVMF, so it would look like:

Hypervisor -> OVMF -> FreeBSD EFI loader -> FreeBSD kernel

> > IMHO, if we want to boot PVH using EFI the right solution is to use OVMF (or 
> > any other UEFI firmware) and port it so it's able to run as a PVH guest. I 
> > guess it should even be possible to use it for Dom0, although I think this 
> > is cumbersome.
> 
> There are two levels of EFI boot entry features being discussed,
> 
>  1. Make the OS kernel a PE/COFF executable
>  2. Provide some level of EFI service functionality
> 
> You can adopt 1. without 2, i.e. without actually providing any EFI
> services at all, as long as the Xen hypervisor grows a PE/COFF loader
> (since EFI firmware has to provide you one, for EFI platforms you
> could use the LoadImage() service in the firmware, but for BIOS
> platforms you'd need your own in Xen).

We could use native LoadImage for Dom0 maybe if we are booted on an EFI 
platform, but for DomUs we certainly need to implement our own inside of 
Xen, at which point we could do the same and always use the one inside of 
Xen in order to avoid diverging paths.

TBH, I don't think this is the right solution. We would force every OS 
kernel that wants to be loaded using Xen to become a PE/COFF executable. 
This also includes unikernels like MirageOS, which will be forced to become 
a PE/COFF executable.

Is this header compatible with the ELF header? Con both co-exist in the 
same binary without issues?

> On Linux, this has the advantage of deferring the decompression of the
> bzImage (x86 Linux kernel file format) to the stub on the front of the
> bzImage. And while I realise that the toolstack already has support
> for decompressing bzImages, given what Andrew has said about reducing
> attack surface, having the guest perform the decompression should be a
> win.
> 
> Of course, this is offset somewhat by the fact that you need to audit
> the PE/COFF loader ;) But decompression in general is notoriously
> vulnerable to security issues.
> 
> Using the in-kernel decompressor is how most (all?) Linux boot loaders
> work today, so there's the added benefit of reducing the differences
> between booting on Xen and booting bare metal. For example, you'd
> probably be able to use CONFIG_RANDOMIZE_BASE (ASLR for kernel image)
> for Xen if you use the kernel's decompressor. Xen would also get
> future features in this area for free, and there is a tendency to push
> boot features into the early stub.

All the issues that you mention above are also solved by chainloading OVMF 
instead of directly loading the guest kernel, and it avoids adding a PE/COFF 
loader into Xen.

> For 1. we'd basically be using the PE/COFF file format with the EFI
> ABI as an OS agnostic boot protocol, but not as a full firmware
> runtime environment.

This also means that we will be adding PE/COFF headers to (uni)kernels, but 
we won't still implement full EFI support inside of them, so although it 
would seem like they are capable of being loaded by a native EFI loader, 
they would not.

This seems misleading, and I think it's going to cause grief amongst OS 
developers in general. The current proposed entry point is unique to Xen 
(it's only mentioned in Xen ELF notes), and is certainly not going to cause 
confusion at all.

Also, doesn't this (the fact that Xen will use the EFI entry point 
without a runtime environment) mean that there are going to be diverging 
paths inside of Linux EFI entry point anyway?

At which point, does it really matter that much if this divergence includes 
a new entry point or not?

> 2. is also interesting, though I think less so than 1. I agree that
> making OVMF work as a PVH guest is probably the right way to go, even
> for Dom0, not least because you'd have a much cleaner/less buggy
> implementation than what we see in the real world ;)

I think we all agree that this is not suitable.

Roger.

[toc] | [prev] | [next] | [standalone]


#1378169

From"Luis R. Rodriguez" <mcgrof@kernel.org>
Date2016-04-13 20:40 +0200
Message-ID<rnxRM-5zu-9@gated-at.bofh.it>
In reply to#1375496
On Mon, Apr 11, 2016 at 07:12:08AM +0200, Juergen Gross wrote:
> On 08/04/16 22:40, Luis R. Rodriguez wrote:
> > On Wed, Apr 06, 2016 at 10:40:08AM +0100, David Vrabel wrote:
> >> On 06/04/16 03:40, Luis R. Rodriguez wrote:
> >>>
> >>>     * You don't need full EFI emulation
> >>
> >> I think needing any EFI emulation inside Xen (which is where it would
> >> need to be for dom0) is not suitable because of the increase in
> >> hypervisor ABI.
> > 
> > Is this because of timing on architecture / design of HVMLite, or
> > a general position that the complexity to deal with EFI emulation
> > is too much for Xen's taste ?
> 
> The Xen hypervisor should be as small as possible. Adding an EFI
> emulator will be adding quite some code. This should be done after a
> very thorough evaluation only.

Sure.

> > ARM already went the EFI entry way for domU -- it went the OVMF route,
> > would such a possibility be possible for x86 domU HVMLite ? If not why
> > not, I mean it would seem to make sense to at least mimic the same type
> > of early boot environment, and perhaps there are some lessons to be
> > learned from that effort too.
> 
> The final solution must be appropriate for dom0, too. So don't try
> to limit the discussion to domU. If dom0 isn't going to be acceptable
> there will no need to discuss domU.

Understood. George noted that on ARM dom0 still uses the ARM native entry
point, it seems to accomplish this as it uses a device tree node. I'll
chime in on that in another thread.

> > Are there some lessons to be learned with ARM's effort? What are they?
> > If that could be re-done again with any type of cleaner path, what
> > could that be that could help the x86 side ?
> > 
> > Although emulating EFI may require work, some folks have pointed out
> > that the amount of work may not be that much. If that is done can
> > we instead rely on the same code to replace OVMF to support both
> > Xen ARM and Xen HVMLite on x86 ? What would be the pros / cons of
> > this ?
> > 
> >> I also still do not understand your objection to the current tiny stub.
> > 
> > Its more of a hypothetical -- can an EFI entry be used instead given
> > it already does exactly what the new small entry does ? Its also rather
> > odd to add a new entry without evaluating fully a possible alternative
> > that would provide the same exact mechanism.
> 
> The interface isn't the new entry only. It should be evaluated how much
> of the early EFI boot path would be common to the HVMlite one.

We also have other asm code which can be shared. I'll reply to Boris'
original e-mail with what I can identify as perhaps sharable. There is
obviously more as you allude.

> What would be gained by using the same entry but having two different boot
> paths after it?

Its a good question. In summary for me it would be the push for sharing more
code and the push for semantics on early boot to address differences
proactively, and ultimately it may enable us to help bring closer the old PV
boot path closer.

I'll elaborate on this but first let's clarify why a new entry is used for
HVMlite to start of with:

  1) Xen ABI has historically not wanted to set up the boot params for Linux
     guests, instead it insists on letting the Linux kernel Xen boot stubs fill
     that out for it. This sticking point means it has implicated a boot stub.
     The HVMLite boot entry tries to bring the boot entries paths closer as it
     leverages more of the HVM boot path philosophy to mimic the regular PC boot
     path.

     Is HVMLite supposed to support legacy PV guests as well BTW ?

     Reason I'm highlighting Xen ABI as a *reason* alone is that even with
     today's large discrepancy on the old PV boot path I believe we can
     bring together the boot paths closer together if the Xen ABI was slightly
     flexible about this, I've highlighted how I believe that is possible before,
     *iff* the Xen ABI would at the very least set 2 things only:

     a) Hypervisor type
     b) A custom data pointer

     This would enable a single boot entry on the guest to handle then:

	Pseudo code:

	startup_32()                         startup_64()
	       |                                  |
	       |                                  |
	       V                                  V
	pre_hypervisor_stub_32()        pre_hypervisor_stub_64()
	       |                                  |
	       |                                  |
	       V                                  V
	 [existing startup_32()]       [existing startup_64()]
	       |                                  |
	       |                                  |
	       V                                  V
	post_hypervisor_stub_32()       post_hypervisor_stub_64()

     
     If the Xen ABI was flexible about setting a hypervisor type and custom
     data pointer then we would haven handlers for it, and in it, it can
     do whatever it thinks is needed for its own guest types. It could
     also continue to set the zero page on its own as it sees fit.

     Again, note that if this is done it could also mean even bringing together
     the old PV boot path closer together... so this is not just a prospect
     for HVMLite but also for old PV guests.

  2) Because of 1) it has meant we have no formal semantics for early boot
     code is available and so severe differences can best be addressed also
     by yet another boot entry. This has meant often times not addressing
     or not knowing if we've addressed real differences between the different
     entries. Case in point, dead code [0]. How do we know we will not run
     certain code that should not run for the different entries ? Without
     *any* semantics later in boot code to distinguish where we came from
     and because we strive to build single kernels with different possible
     run time environments it means we have tons of code available to
     execute / run that we may not need.

     Because of the lack of semantics we may still have dead code prospects
     with the new HVMLite entry. How are we sure there is no differences ?

[0] http://www.do-not-panic.com/2015/12/avoiding-dead-code-pvops-not-silver-bullet.html

  3) Unikernel / other OS requirements: this is really tied to 2) but even if
     we tried to evolve the Xen ABI it would mean considering existing solutions
     out there. Things to consider as an example: FreeBSD doesn't have an EFI
     entry, unikernels want a simple boot entry.

With this in mind then, that I can think of:

Cons of using the same entry but having two different boot paths:

  * Pushes the Xen ABI, needs to make everyone happy, this is hard
  * Perhaps harder to implement

Gains of striving to use the same entry but having two different boot:

 * Helps to share more code easily
 * Reduce attack surface
 * Requires us to have semantics for early boot; this has a series of
   side benefits:
   - Means you should try to address differences explicitly rather than
     implicitly -- case in point Dead Code

> You still need a way to distinguish between bare metal
> EFI and HVMlite.

Great point! This is the semantics aspect. The new entry for HVMlite approach
deals with this by making the differences implicit by the new entry point.
My call for addressing this through a hypervisor type was to see if we can
get those semantics added explicitly so we can also later address dead
code concerns for the new HVMLite guest type.

Part of my own interest in an EFI entry here is that EFI could be used to help
expand on the semantics in an OS/agnostic form rather than pushing the x86 boot
protocol further. That seems to have its own set of drawbacks though.


> And Xen needs a way to find out whether a kernel is
> supporting HVMlite to boot it in the correct mode.

How was Xen going to find out if new kernels had HVMlite support with the
new entry ? An ELFNOTE() ? If an entry is shared could we note use an
ELFNOTE() also for this though too ?

  Luis

[toc] | [prev] | [next] | [standalone]


#1378286 — Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry

From"Luis R. Rodriguez" <mcgrof@kernel.org>
Date2016-04-13 22:50 +0200
SubjectRe: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry
Message-ID<rnzTA-78a-17@gated-at.bofh.it>
In reply to#1378169
On Wed, Apr 13, 2016 at 02:56:29PM -0400, Konrad Rzeszutek Wilk wrote:
> On Wed, Apr 13, 2016 at 08:29:51PM +0200, Luis R. Rodriguez wrote:
> > On Mon, Apr 11, 2016 at 07:12:08AM +0200, Juergen Gross wrote:
> > 
> > > What would be gained by using the same entry but having two different boot
> > > paths after it?
> > 
> > Its a good question. In summary for me it would be the push for sharing more
> > code and the push for semantics on early boot to address differences
> > proactively, and ultimately it may enable us to help bring closer the old PV
> > boot path closer.
> 
> But why? We want to kill PV (eventually).

Yeah yeah, but its still there, and we'll have to live with it for
at least minimum 5 years I hear. Part of my interest is to see to it
that this path gets less disruption and issues, and we also address
dead code issues which pvops simply folded under the rug. The dead code
concerns may exist still for hvmlite, so unless someone is willing
to make a bold claim there is none, its something to consider.

How we address semantics then is *very* important to me.

> > I'll elaborate on this but first let's clarify why a new entry is used for
> > HVMlite to start of with:
> > 
> >   1) Xen ABI has historically not wanted to set up the boot params for Linux
> >      guests, instead it insists on letting the Linux kernel Xen boot stubs fill
> >      that out for it. This sticking point means it has implicated a boot stub.
> 
> 
> Which is b/c it has to be OS agnostic. It has nothing to do 'not wanting'.

It can still be OS agnostic and pass on type and custom data pointer.

Would that be reasonable ?

> >      The HVMLite boot entry tries to bring the boot entries paths closer as it
> >      leverages more of the HVM boot path philosophy to mimic the regular PC boot
> >      path.
> > 
> >      Is HVMLite supposed to support legacy PV guests as well BTW ?
> 
> Gosh no.

Interesting.. and *everyone* is happy about this?

> >      Reason I'm highlighting Xen ABI as a *reason* alone is that even with
> >      today's large discrepancy on the old PV boot path I believe we can
> >      bring together the boot paths closer together if the Xen ABI was slightly
> >      flexible about this, I've highlighted how I believe that is possible before,
> 
> <runs away screaming>

Everyone has. If you need to support old PV guests for more than 5 years the
work I'm doing should help with that. I'm trying to leverage gains of the
work I'm doing for HVMLite, and part of this is trying to address semantics
proactively.

> >      *iff* the Xen ABI would at the very least set 2 things only:
> > 
> >      a) Hypervisor type
> >      b) A custom data pointer
> > 
> >      This would enable a single boot entry on the guest to handle then:
> > 
> > 	Pseudo code:
> > 
> > 	startup_32()                         startup_64()
> > 	       |                                  |
> > 	       |                                  |
> > 	       V                                  V
> > 	pre_hypervisor_stub_32()        pre_hypervisor_stub_64()
> > 	       |                                  |
> > 	       |                                  |
> > 	       V                                  V
> > 	 [existing startup_32()]       [existing startup_64()]
> > 	       |                                  |
> > 	       |                                  |
> > 	       V                                  V
> > 	post_hypervisor_stub_32()       post_hypervisor_stub_64()
> > 
> >      
> >      If the Xen ABI was flexible about setting a hypervisor type and custom
> >      data pointer then we would haven handlers for it, and in it, it can
> >      do whatever it thinks is needed for its own guest types. It could
> >      also continue to set the zero page on its own as it sees fit.
> > 
> >      Again, note that if this is done it could also mean even bringing together
> >      the old PV boot path closer together... so this is not just a prospect
> >      for HVMLite but also for old PV guests.
> > 
> >   2) Because of 1) it has meant we have no formal semantics for early boot
> >      code is available and so severe differences can best be addressed also
> >      by yet another boot entry. This has meant often times not addressing
> 
> There are semantics written for this new code: http://xenbits.xen.org/docs/unstable/misc/hvmlite.html

That only addressed semantics for early boot code implicitly through a new entry...

> All other ones related to low-level operations are described in Intel SDM.
> 
> 
> >      or not knowing if we've addressed real differences between the different
> >      entries. Case in point, dead code [0]. How do we know we will not run
> >      certain code that should not run for the different entries ? Without
> >      *any* semantics later in boot code to distinguish where we came from
> >      and because we strive to build single kernels with different possible
> >      run time environments it means we have tons of code available to
> >      execute / run that we may not need.
> 
> I am not following that. PVH aka HVMLite will pretty much erase the need for the
> pvops.

It does not mean there are no dead code concerns with HVMlite.

> > 
> >      Because of the lack of semantics we may still have dead code prospects
> >      with the new HVMLite entry. How are we sure there is no differences ?
> > 
> > [0] http://www.do-not-panic.com/2015/12/avoiding-dead-code-pvops-not-silver-bullet.html
> > 
> >   3) Unikernel / other OS requirements: this is really tied to 2) but even if
> >      we tried to evolve the Xen ABI it would mean considering existing solutions
> >      out there. Things to consider as an example: FreeBSD doesn't have an EFI
> >      entry, unikernels want a simple boot entry.
> > 
> > With this in mind then, that I can think of:
> > 
> > Cons of using the same entry but having two different boot paths:
> > 
> >   * Pushes the Xen ABI, needs to make everyone happy, this is hard
> >   * Perhaps harder to implement
> > 
> > Gains of striving to use the same entry but having two different boot:
> > 
> >  * Helps to share more code easily
> >  * Reduce attack surface
> >  * Requires us to have semantics for early boot; this has a series of
> >    side benefits:
> >    - Means you should try to address differences explicitly rather than
> >      implicitly -- case in point Dead Code
> > 
> > > You still need a way to distinguish between bare metal
> > > EFI and HVMlite.
> > 
> > Great point! This is the semantics aspect. The new entry for HVMlite approach
> > deals with this by making the differences implicit by the new entry point.
> > My call for addressing this through a hypervisor type was to see if we can
> > get those semantics added explicitly so we can also later address dead
> > code concerns for the new HVMLite guest type.
> 
> Right, they are..

There is huge merit to address a huge chunks of dead code concerns by sticking
more closer to the native booth paths, it doesn't mean you still have no
dead code concerns with HVMlite, nor that HVMLite has no platform quirks,
it does and part of some recent work is to pave a *clean* path for setting
these differences apart.

> > Part of my own interest in an EFI entry here is that EFI could be used to help
> > expand on the semantics in an OS/agnostic form rather than pushing the x86 boot
> > protocol further. That seems to have its own set of drawbacks though.
> > 
> > 
> > > And Xen needs a way to find out whether a kernel is
> > > supporting HVMlite to boot it in the correct mode.
> > 
> > How was Xen going to find out if new kernels had HVMlite support with the
> > new entry ? An ELFNOTE() ? If an entry is shared could we note use an
> 
> Yeah.
> > ELFNOTE() also for this though too ?
> 
> Not sure what you mean by 'shared'. But you can add multiple Elf PT_NOTEs.
> See the ELF document.

OK so even if we used a common/shared entry point we can address letting
Xen find out whether or not a kernel supports HVMlite.

  Luis

[toc] | [prev] | [next] | [standalone]


#1378319 — Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry

From"Luis R. Rodriguez" <mcgrof@kernel.org>
Date2016-04-14 00:30 +0200
SubjectRe: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry
Message-ID<rnBsm-8rm-11@gated-at.bofh.it>
In reply to#1378286
On Wed, Apr 13, 2016 at 05:08:01PM -0400, Konrad Rzeszutek Wilk wrote:
> On Wed, Apr 13, 2016 at 10:40:55PM +0200, Luis R. Rodriguez wrote:
> > On Wed, Apr 13, 2016 at 02:56:29PM -0400, Konrad Rzeszutek Wilk wrote:
> > > On Wed, Apr 13, 2016 at 08:29:51PM +0200, Luis R. Rodriguez wrote:
> > > > On Mon, Apr 11, 2016 at 07:12:08AM +0200, Juergen Gross wrote:
> > > > 
> > > > > What would be gained by using the same entry but having two different boot
> > > > > paths after it?
> > > > 
> > > > Its a good question. In summary for me it would be the push for sharing more
> > > > code and the push for semantics on early boot to address differences
> > > > proactively, and ultimately it may enable us to help bring closer the old PV
> > > > boot path closer.
> > > 
> > > But why? We want to kill PV (eventually).
> > 
> > Yeah yeah, but its still there, and we'll have to live with it for
> > at least minimum 5 years I hear. Part of my interest is to see to it
> > that this path gets less disruption and issues, and we also address
> > dead code issues which pvops simply folded under the rug. The dead code
> > concerns may exist still for hvmlite, so unless someone is willing
> > to make a bold claim there is none, its something to consider.
> 
> What is this dead code you speak of? Is it MTRR? Is early path code
> that PV misses (like KASL or other?)

Kasan is dead code to Xen. If you boot x86 Xen with Kasan enabled
Xen explodes. Quick question, will Kasan not explode with HVMLite ?

MTRR used to be dead code concern but since we have vetted most of that code
now we are pretty certain that code should never run now.

KASLR may be -- not sure as I  haven't vetted that, but from
what I have loosely heard maybe.

VGA code will be dead code for HVMlite for sure as the design doc
says it will not run VGA, the ACPI flag will be set but the check
for that is not yet on Linux. That means the VGA Linux code will
be there but we have no way to ensure it will not run nor that
anything will muck with it.

To be clear -- dead code concerns still exist even without
virtualization solutions, its just that with virtualization
this stuff comes up more and there has been no proactive
measures to address this. The question of semantics here is
to see to what extent we need earlier boot code annotations
to ensure we address semantics proactively.

> The entrace point in Linux "proper" is startup_32 or startup_64 - the same
> path that EFI uses.
> 
> If you were to draw this (very simplified):
> 
> a)- GRUB2 ---------------------\ (creates an bootparam structure)
>                                 \
>                                  +---- startup_32 or startup_64
> b) EFI -> Linux EFI stub -------/
>        (creates bootparm)      /
> c) GRUB2-EFI  -> Linux EFI----/
>                stub         /
> d) HVMLite ----------------/
>       (creates bootparm)

b) and d) might be able to share paths there...
d) still has its own entry, it does more than create boot params.

> (I am not sure about the c) - I would have to look in source to
> be source). There is also LILO in this, but I am not even sure if
> works anymore.
> 
> 
> What you have is that every entry point creates the bootparams
> and ends up calling startup_X. The startup_64 then hit the rest
> of the kernel. The startp_X code is the one that would setup
> the basic pagetables, segments, etc.

Sure.. a full diagram should include both sides and how when using
a custom entry one runs the risk of skipping a lot of code setup.
There is that and as others have pointed out how certain guests types
are assumed to not have certain peripherals, and we have no idea
to ensure certain old legacy code may not ever run or be accessed
by drivers.

> > How we address semantics then is *very* important to me.
> 
> Which semantics? How the CPU is going to be at startup_X ? Or
> how the CPU is going to be when EFI firmware invokes the EFI stub?
> Or when GRUB2 loads Linux?

What hypervisor kicked me and what guest type I am.

Let me elaborate more below.

> That (those bootloaders) is clearly defined. The URL I provided
> mentions the HVMLite one. The Documentation/x86/boot.c mentions
> what the semantics are to expected when providing an bootstrap
> (which is what HVMLitel stub code in Linux would write against -
> and what EFI stub code had been written against too).
> > 
> > > > I'll elaborate on this but first let's clarify why a new entry is used for
> > > > HVMlite to start of with:
> > > > 
> > > >   1) Xen ABI has historically not wanted to set up the boot params for Linux
> > > >      guests, instead it insists on letting the Linux kernel Xen boot stubs fill
> > > >      that out for it. This sticking point means it has implicated a boot stub.
> > > 
> > > 
> > > Which is b/c it has to be OS agnostic. It has nothing to do 'not wanting'.
> > 
> > It can still be OS agnostic and pass on type and custom data pointer.
> 
> Sure. It has that (it MUST otherwise how else would you pass data).
> It is documented as well http://xenbits.xen.org/docs/unstable/hypercall/x86_64/include,public,xen.h.html#incontents_startofday
> (see " Start of day structure passed to PVH guests in %ebx.")

The design doc begs for a custom OS entry point though.
If we had a single 'type' and 'custom data' passed to the kernel that
should suffice for the default Linux entry point to just pivot off
of that and do what it needs without more entry points. Once.

  Luis

[toc] | [prev] | [next] | [standalone]


#1378378 — Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry

FromKonrad Rzeszutek Wilk <konrad.wilk@oracle.com>
Date2016-04-14 03:10 +0200
SubjectRe: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry
Message-ID<rnDXc-1Tk-3@gated-at.bofh.it>
In reply to#1378319
On Thu, Apr 14, 2016 at 12:23:17AM +0200, Luis R. Rodriguez wrote:
> On Wed, Apr 13, 2016 at 05:08:01PM -0400, Konrad Rzeszutek Wilk wrote:
> > On Wed, Apr 13, 2016 at 10:40:55PM +0200, Luis R. Rodriguez wrote:
> > > On Wed, Apr 13, 2016 at 02:56:29PM -0400, Konrad Rzeszutek Wilk wrote:
> > > > On Wed, Apr 13, 2016 at 08:29:51PM +0200, Luis R. Rodriguez wrote:
> > > > > On Mon, Apr 11, 2016 at 07:12:08AM +0200, Juergen Gross wrote:
> > > > > 
> > > > > > What would be gained by using the same entry but having two different boot
> > > > > > paths after it?
> > > > > 
> > > > > Its a good question. In summary for me it would be the push for sharing more
> > > > > code and the push for semantics on early boot to address differences
> > > > > proactively, and ultimately it may enable us to help bring closer the old PV
> > > > > boot path closer.
> > > > 
> > > > But why? We want to kill PV (eventually).
> > > 
> > > Yeah yeah, but its still there, and we'll have to live with it for
> > > at least minimum 5 years I hear. Part of my interest is to see to it
> > > that this path gets less disruption and issues, and we also address
> > > dead code issues which pvops simply folded under the rug. The dead code
> > > concerns may exist still for hvmlite, so unless someone is willing
> > > to make a bold claim there is none, its something to consider.
> > 
> > What is this dead code you speak of? Is it MTRR? Is early path code
> > that PV misses (like KASL or other?)
> 
> Kasan is dead code to Xen. If you boot x86 Xen with Kasan enabled

For Xen PV guests,
> Xen explodes. Quick question, will Kasan not explode with HVMLite ?

.. but for HVMLite of Xen HVM guest Kasan will run.
> 
> MTRR used to be dead code concern but since we have vetted most of that code
> now we are pretty certain that code should never run now.
> 
> KASLR may be -- not sure as I  haven't vetted that, but from
> what I have loosely heard maybe.
> 
> VGA code will be dead code for HVMlite for sure as the design doc
> says it will not run VGA, the ACPI flag will be set but the check
> for that is not yet on Linux. That means the VGA Linux code will
> be there but we have no way to ensure it will not run nor that
> anything will muck with it.

<shrugs> The worst it will do is try to read non-existent registers.
The VGA code should be able to handle failures like that and
not initialize itself when the hardware is dead (or non-existent).
> 
> To be clear -- dead code concerns still exist even without
> virtualization solutions, its just that with virtualization
> this stuff comes up more and there has been no proactive
> measures to address this. The question of semantics here is
> to see to what extent we need earlier boot code annotations
> to ensure we address semantics proactively.

I think what you mean by dead code is another word for
hardware test coverage?
> 
> > The entrace point in Linux "proper" is startup_32 or startup_64 - the same
> > path that EFI uses.
> > 
> > If you were to draw this (very simplified):
> > 
> > a)- GRUB2 ---------------------\ (creates an bootparam structure)
> >                                 \
> >                                  +---- startup_32 or startup_64
> > b) EFI -> Linux EFI stub -------/
> >        (creates bootparm)      /
> > c) GRUB2-EFI  -> Linux EFI----/
> >                stub         /
> > d) HVMLite ----------------/
> >       (creates bootparm)
> 
> b) and d) might be able to share paths there...

No idea. You would have to look in the assembler code to
figure that out.

> d) still has its own entry, it does more than create boot params.

d) purpose is to create boot params. It may do more as nobody likes
to muck in assembler and make bootparams from within assembler.

> 
> > (I am not sure about the c) - I would have to look in source to
> > be source). There is also LILO in this, but I am not even sure if
> > works anymore.
> > 
> > 
> > What you have is that every entry point creates the bootparams
> > and ends up calling startup_X. The startup_64 then hit the rest
> > of the kernel. The startp_X code is the one that would setup
> > the basic pagetables, segments, etc.
> 
> Sure.. a full diagram should include both sides and how when using
> a custom entry one runs the risk of skipping a lot of code setup.

But it does not skip a lot of code setup. It starts exactly
at the same code startup that _all_ bootstraping code start at.

> There is that and as others have pointed out how certain guests types
> are assumed to not have certain peripherals, and we have no idea
> to ensure certain old legacy code may not ever run or be accessed
> by drivers.

Ok, but that is not at code setup. That is later - when device
drivers are initialized. This no different than booting on
some hardware with missing functionality. ACPI, PCI and PnP
PnP are set there to help OSes discover this.
> 
> > > How we address semantics then is *very* important to me.
> > 
> > Which semantics? How the CPU is going to be at startup_X ? Or
> > how the CPU is going to be when EFI firmware invokes the EFI stub?
> > Or when GRUB2 loads Linux?
> 
> What hypervisor kicked me and what guest type I am.

cpuid software flags have that - and that semantics has been 
there for eons.
> 
> Let me elaborate more below.
> 
> > That (those bootloaders) is clearly defined. The URL I provided
> > mentions the HVMLite one. The Documentation/x86/boot.c mentions
> > what the semantics are to expected when providing an bootstrap
> > (which is what HVMLitel stub code in Linux would write against -
> > and what EFI stub code had been written against too).
> > > 
> > > > > I'll elaborate on this but first let's clarify why a new entry is used for
> > > > > HVMlite to start of with:
> > > > > 
> > > > >   1) Xen ABI has historically not wanted to set up the boot params for Linux
> > > > >      guests, instead it insists on letting the Linux kernel Xen boot stubs fill
> > > > >      that out for it. This sticking point means it has implicated a boot stub.
> > > > 
> > > > 
> > > > Which is b/c it has to be OS agnostic. It has nothing to do 'not wanting'.
> > > 
> > > It can still be OS agnostic and pass on type and custom data pointer.
> > 
> > Sure. It has that (it MUST otherwise how else would you pass data).
> > It is documented as well http://xenbits.xen.org/docs/unstable/hypercall/x86_64/include,public,xen.h.html#incontents_startofday
> > (see " Start of day structure passed to PVH guests in %ebx.")
> 
> The design doc begs for a custom OS entry point though.

That is what the ELF Note has.
> If we had a single 'type' and 'custom data' passed to the kernel that
> should suffice for the default Linux entry point to just pivot off
> of that and do what it needs without more entry points. Once.

And what about ramdisk? What about multiple ramdisks?
What about command line? All of that is what bootparams
tries to unify on Linux. But 'bootparams' is unique to Linux,
it does not exist on FreeBSD. Hence some stub code to transplant
OS-agnostic simple data to OS-specific is neccessary.
> 
>   Luis

[toc] | [prev] | [next] | [standalone]


#1379210 — Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry

From"Luis R. Rodriguez" <mcgrof@kernel.org>
Date2016-04-14 20:50 +0200
SubjectRe: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry
Message-ID<rnUuZ-6rs-5@gated-at.bofh.it>
In reply to#1378378
On Wed, Apr 13, 2016 at 09:01:32PM -0400, Konrad Rzeszutek Wilk wrote:
> On Thu, Apr 14, 2016 at 12:23:17AM +0200, Luis R. Rodriguez wrote:
> > On Wed, Apr 13, 2016 at 05:08:01PM -0400, Konrad Rzeszutek Wilk wrote:
> > > On Wed, Apr 13, 2016 at 10:40:55PM +0200, Luis R. Rodriguez wrote:
> > > > On Wed, Apr 13, 2016 at 02:56:29PM -0400, Konrad Rzeszutek Wilk wrote:
> > > > > On Wed, Apr 13, 2016 at 08:29:51PM +0200, Luis R. Rodriguez wrote:
> > > > 
> > > > and we also want to address dead code issues which pvops simply folded
> > > > under the rug. The dead code concerns may exist still for hvmlite, so
> > > > unless someone is willing to make a bold claim there is none, its
> > > > something to consider.
> > > 
> > > What is this dead code you speak of?
> > 
> > Kasan is dead code to Xen. If you boot x86 Xen with Kasan enabled
> 
> For Xen PV guests,

That's right. For 5 years this will be a bomb. That went unnoticed
and I feel I have to pull hair now to try to get folks to fix this.

How many other issues will go in which will explode during this 5 year
time line? How can we proactively address a solution to this now so we
avoid this in future ?

Do you believe me now its a real issue?

Fortunately I have a proactive solution for pvops now in my pipeline that
should help avoid us having to blow more things up on Xen but also that should
cause no headaches on behalf of x86 developers. But, reason I have been so
engaged on HVMLite design review is I want to ensure we take the lessons
learned from pvops and avoid this for an architecture that will be def-facto
Xen on Linux 5 years from now.

Not bringing this up or addressing this now for HVMLite / PVH2 would simply be
silly, and since it wasn't addressed in pvops I obviously have to ensure I
convince enough people it was a real issue and ensure that we have enough
semantics available to address it.

Part of the semantics question, which has made my quest hard, was the use of
semantics for virtualization for code on early boot and later in boot has been
rather sloppy, so we have recently needed to address some of these gaps. Some
of these discussions have however been productive, as I'll explain to George
soon regarding his DT questions.  The discussion is not over though and we need
to ensure that if we need semantics for HVMLite we'll have them available in
*clean* way. One of the things where early semantics and design to address
these issue help in a proactive manner is to address a clean boot entry -- and
that's also why I've been so pedantic over review of the new HVMlite boot
entry.

> > Xen explodes. Quick question, will Kasan not explode with HVMLite ?
> 
> .. but for HVMLite of Xen HVM guest Kasan will run.

Are you sure? Should that mean that Xen HVM should be fine as well.  Does that
work? Are we sure?

> > MTRR used to be dead code concern but since we have vetted most of that code
> > now we are pretty certain that code should never run now.
> > 
> > KASLR may be -- not sure as I  haven't vetted that, but from
> > what I have loosely heard maybe.
> > 
> > VGA code will be dead code for HVMlite for sure as the design doc
> > says it will not run VGA, the ACPI flag will be set but the check
> > for that is not yet on Linux. That means the VGA Linux code will
> > be there but we have no way to ensure it will not run nor that
> > anything will muck with it.
> 
> <shrugs> The worst it will do is try to read non-existent registers.

Really ?

Is that your position on all other possible dead code that may have been
possible on old Xen PV guests as well ?

As I hinted, after thinking about this for a while I realized that dead code is
likely present on bare metal as well even without virtualization, specially if
you build large single kernels to support a wide array of features which only
late at run time can be determined. Virtualization and the pvops design just
makes this issue much more prominent. If there are other areas of code exposed
that actually may run, but we are not sure may run, I figured some other folks
with a bit more security conscience minds might even simply take the position
it may be a security risk to leave that code exposed. So to take a position
that 'the worst it will do is try to read non-existent registers' -- seems
rather shortsighted here.

Anyway for more details on thoughts on this refer to the this wiki:

http://kernelnewbies.org/KernelProjects/kernel-sandboxing

Since this is now getting off topic please send me your feedback on another
thread for the non-virtualization aspects of this if that interests you. My
point here was rather to highlight the importance of clear semantics due to
virtualization in light of possible dead code.

> The VGA code should be able to handle failures like that and
> not initialize itself when the hardware is dead (or non-existent).

That's right, its through ACPI_FADT_NO_VGA and since its part of the HVMLite
design doc we want HVMlite design to address ACPI_FADT_NO_VGA properly.  I've
paved the way for this to be done cleanly and easily now, but that code should
be in place before HVMLite code gets merged.

Does domU for old Xen PV also set ACPI_FADT_NO_VGA as well ?  Should it ?

> > To be clear -- dead code concerns still exist even without
> > virtualization solutions, its just that with virtualization
> > this stuff comes up more and there has been no proactive
> > measures to address this. The question of semantics here is
> > to see to what extent we need earlier boot code annotations
> > to ensure we address semantics proactively.
> 
> I think what you mean by dead code is another word for
> hardware test coverage?

No, no, its very different given that with virtualization the scope of possible
dead code is significant and at run time you are certain a huge portion of code
should *never ever* run. So for instance we know once we boot bare metal none
of the Xen stuff should ever run, likewise on Xen dom0 we know none of the KVM
/ bare-metal only stuff should never run, when on Xen domU, none of the Xen
domU-only stuff should ever run.

> > > The entrace point in Linux "proper" is startup_32 or startup_64 - the same
> > > path that EFI uses.
> > > 
> > > If you were to draw this (very simplified):
> > > 
> > > a)- GRUB2 ---------------------\ (creates an bootparam structure)
> > >                                 \
> > >                                  +---- startup_32 or startup_64
> > > b) EFI -> Linux EFI stub -------/
> > >        (creates bootparm)      /
> > > c) GRUB2-EFI  -> Linux EFI----/
> > >                stub         /
> > > d) HVMLite ----------------/
> > >       (creates bootparm)
> > 
> > b) and d) might be able to share paths there...
> 
> No idea. You would have to look in the assembler code to
> figure that out.

And that's a pain, I get it.

I spotted one place already -- will note to Boris. I think Matt may have more
ideas ;)

> > d) still has its own entry, it does more than create boot params.
> 
> d) purpose is to create boot params.

OK good to know that's the only thing we acknowledge it *should* do.

>  It may do more as nobody likes to muck in assembler and make bootparams from
>  within assembler.

OK -- it does do more and that's where we'd like to avoid duplication if
possible and yet-another-entry (TM).

> > > (I am not sure about the c) - I would have to look in source to
> > > be source). There is also LILO in this, but I am not even sure if
> > > works anymore.
> > > 
> > > 
> > > What you have is that every entry point creates the bootparams
> > > and ends up calling startup_X. The startup_64 then hit the rest
> > > of the kernel. The startp_X code is the one that would setup
> > > the basic pagetables, segments, etc.
> > 
> > Sure.. a full diagram should include both sides and how when using
> > a custom entry one runs the risk of skipping a lot of code setup.
> 
> But it does not skip a lot of code setup. It starts exactly
> at the same code startup that _all_ bootstraping code start at.

Its a fair point.

> > There is that and as others have pointed out how certain guests types
> > are assumed to not have certain peripherals, and we have no idea
> > to ensure certain old legacy code may not ever run or be accessed
> > by drivers.
> 
> Ok, but that is not at code setup. That is later - when device
> drivers are initialized. This no different than booting on
> some hardware with missing functionality. ACPI, PCI and PnP
> PnP are set there to help OSes discover this.

To a certain extent this is true, but there may things which are missing still.

We really have no idea what the full list of those things are.

It may be that things may have been running for ages without notice of an issue
or that only under certain situations will certain issues or bugs trigger a
failure. For instance, just yesterday I was Cc'd on a brand-spanking new legacy
conflict [0], caused by upstream commit 8c058b0b9c34d8c ("x86/irq: Probe for
PIC presence before allocating descs for legacy IRQs") merged on v4.4 where
some new code used nr_legacy_irqs() -- one proposed solution seems to be that
for Xen code NR_IRQS_LEGACY should be used instead is as it lacks PCI [1] and
another was to peg the legacy requirements as a quirk on the new x86 platform
legacy quirk stuff [2]. Are other uses of nr_legacy_irqs() correct ? Are
we sure ?

[0] http://lkml.kernel.org/r/570F90DF.1020508@oracle.com
[1] https://lkml.org/lkml/2016/4/14/532
[2] http://lkml.kernel.org/r/1460592286-300-1-git-send-email-mcgrof@kernel.org

> > > > How we address semantics then is *very* important to me.
> > > 
> > > Which semantics? How the CPU is going to be at startup_X ? Or
> > > how the CPU is going to be when EFI firmware invokes the EFI stub?
> > > Or when GRUB2 loads Linux?
> > 
> > What hypervisor kicked me and what guest type I am.
> 
> cpuid software flags have that - and that semantics has been 
> there for eons.

We cannot use cpuid early in asm code, I'm looking for something we
can even use on asm early in boot code, on x86 the best option we
have is the boot_params, but I've even have had issues with that
early in code, as I can only access it after load_idt() where I
described my effort to unify Xen PV and x86_64 init paths [3].

[3] http://lkml.kernel.org/r/CAB=NE6VTCRCazcNpCdJ7pN1eD3=x_fcGOdH37MzVpxkKEN5esw@mail.gmail.com

> > Let me elaborate more below.
> > 
> > > That (those bootloaders) is clearly defined. The URL I provided
> > > mentions the HVMLite one. The Documentation/x86/boot.c mentions
> > > what the semantics are to expected when providing an bootstrap
> > > (which is what HVMLitel stub code in Linux would write against -
> > > and what EFI stub code had been written against too).
> > > > 
> > > > > > I'll elaborate on this but first let's clarify why a new entry is used for
> > > > > > HVMlite to start of with:
> > > > > > 
> > > > > >   1) Xen ABI has historically not wanted to set up the boot params for Linux
> > > > > >      guests, instead it insists on letting the Linux kernel Xen boot stubs fill
> > > > > >      that out for it. This sticking point means it has implicated a boot stub.
> > > > > 
> > > > > 
> > > > > Which is b/c it has to be OS agnostic. It has nothing to do 'not wanting'.
> > > > 
> > > > It can still be OS agnostic and pass on type and custom data pointer.
> > > 
> > > Sure. It has that (it MUST otherwise how else would you pass data).
> > > It is documented as well http://xenbits.xen.org/docs/unstable/hypercall/x86_64/include,public,xen.h.html#incontents_startofday
> > > (see " Start of day structure passed to PVH guests in %ebx.")
> > 
> > The design doc begs for a custom OS entry point though.
> 
> That is what the ELF Note has.

Right, but I'm saying that its rather silly to be adding entry points if
all we want the code to do is copy the boot params for us. The design
doc requires a new entry, and likewise you'd need yet-another-entry if
HVMLite is thrown out the window and come back 5 years later after new
hardware solutions are in place and need to redesign HVMLite. Kind of
where we are with PVH today. Likewise if other paravirtualization
developers want to support Linux and want to copy your strategy they'd
add yet-another-entry-point as well.

This is dumb.

> > If we had a single 'type' and 'custom data' passed to the kernel that
> > should suffice for the default Linux entry point to just pivot off
> > of that and do what it needs without more entry points. Once.
> 
> And what about ramdisk? What about multiple ramdisks?
> What about command line? All of that is what bootparams
> tries to unify on Linux. But 'bootparams' is unique to Linux,
> it does not exist on FreeBSD. Hence some stub code to transplant
> OS-agnostic simple data to OS-specific is neccessary.

If we had a Xen ABI option where *all* that I'm asking is you pass
first:

  a) hypervisor type
  b) custom data pointer

We'd be able to avoid adding *any* entry point and just address
the requirements as I noted with pre / post stubs for the type.
This would require an x86 boot protocol bump, but all the issues
creeping up randomly I think that's worth putting on the table now.

And maybe we don't want it to be hypervisor specific, perhaps there are other
*needs* for custom pre-post startup_32()/startup_64() stubs.

To avoid extending boot_params further I figured perhaps we can look
at EFI as another option instead. If we are going to drop all legacy
PV support from the kernel (not the hypervisor) and require hardware
virtualization 5 years from now on the Linux kernel, it doesn't seem
to me far fetched to at the very least consider using an EFI entry
instead, specially since all it does is set boot params and we can
make re-use this for HVMLite too.

  Luis

[toc] | [prev] | [next] | [standalone]


#1379276 — Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry

FromKonrad Rzeszutek Wilk <konrad.wilk@oracle.com>
Date2016-04-14 22:00 +0200
SubjectRe: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry
Message-ID<rnVAJ-7pv-7@gated-at.bofh.it>
In reply to#1379210
On Thu, Apr 14, 2016 at 08:40:48PM +0200, Luis R. Rodriguez wrote:
> On Wed, Apr 13, 2016 at 09:01:32PM -0400, Konrad Rzeszutek Wilk wrote:
> > On Thu, Apr 14, 2016 at 12:23:17AM +0200, Luis R. Rodriguez wrote:
> > > On Wed, Apr 13, 2016 at 05:08:01PM -0400, Konrad Rzeszutek Wilk wrote:
> > > > On Wed, Apr 13, 2016 at 10:40:55PM +0200, Luis R. Rodriguez wrote:
> > > > > On Wed, Apr 13, 2016 at 02:56:29PM -0400, Konrad Rzeszutek Wilk wrote:
> > > > > > On Wed, Apr 13, 2016 at 08:29:51PM +0200, Luis R. Rodriguez wrote:
> > > > > 
> > > > > and we also want to address dead code issues which pvops simply folded
> > > > > under the rug. The dead code concerns may exist still for hvmlite, so
> > > > > unless someone is willing to make a bold claim there is none, its
> > > > > something to consider.
> > > > 
> > > > What is this dead code you speak of?
> > > 
> > > Kasan is dead code to Xen. If you boot x86 Xen with Kasan enabled
> > 
> > For Xen PV guests,
> 
> That's right. For 5 years this will be a bomb. That went unnoticed
> and I feel I have to pull hair now to try to get folks to fix this.

Sometimes you have to roll up your sleeves and do the work yourself.
> 
> How many other issues will go in which will explode during this 5 year
> time line? How can we proactively address a solution to this now so we
> avoid this in future ?
> 
> Do you believe me now its a real issue?

I never said otherwise. What I was confused was that you grouped
this with HVMLite - which would not have a problem like this.

> 
> Fortunately I have a proactive solution for pvops now in my pipeline that
> should help avoid us having to blow more things up on Xen but also that should
> cause no headaches on behalf of x86 developers. But, reason I have been so
> engaged on HVMLite design review is I want to ensure we take the lessons
> learned from pvops and avoid this for an architecture that will be def-facto
> Xen on Linux 5 years from now.
> 
> Not bringing this up or addressing this now for HVMLite / PVH2 would simply be
> silly, and since it wasn't addressed in pvops I obviously have to ensure I
> convince enough people it was a real issue and ensure that we have enough
> semantics available to address it.

Kasan came after pvops, so of course it was not addressed in pvops.

> 
> Part of the semantics question, which has made my quest hard, was the use of
> semantics for virtualization for code on early boot and later in boot has been
> rather sloppy, so we have recently needed to address some of these gaps. Some
> of these discussions have however been productive, as I'll explain to George
> soon regarding his DT questions.  The discussion is not over though and we need
> to ensure that if we need semantics for HVMLite we'll have them available in
> *clean* way. One of the things where early semantics and design to address
> these issue help in a proactive manner is to address a clean boot entry -- and
> that's also why I've been so pedantic over review of the new HVMlite boot
> entry.

I must have missed your review of the patches. Sorry!
> 
> > > Xen explodes. Quick question, will Kasan not explode with HVMLite ?
> > 
> > .. but for HVMLite of Xen HVM guest Kasan will run.
> 
> Are you sure? Should that mean that Xen HVM should be fine as well.  Does that
> work? Are we sure?

Yes, and yes.
> 
> > > MTRR used to be dead code concern but since we have vetted most of that code
> > > now we are pretty certain that code should never run now.
> > > 
> > > KASLR may be -- not sure as I  haven't vetted that, but from
> > > what I have loosely heard maybe.
> > > 
> > > VGA code will be dead code for HVMlite for sure as the design doc
> > > says it will not run VGA, the ACPI flag will be set but the check
> > > for that is not yet on Linux. That means the VGA Linux code will
> > > be there but we have no way to ensure it will not run nor that
> > > anything will muck with it.
> > 
> > <shrugs> The worst it will do is try to read non-existent registers.
> 
> Really ?
> 
> Is that your position on all other possible dead code that may have been
> possible on old Xen PV guests as well ?

This is not just with Xen - it with other device drivers that are being
invoked on baremetal and are not present in hardware anymore.
> 
> As I hinted, after thinking about this for a while I realized that dead code is
> likely present on bare metal as well even without virtualization, specially if

Yes!
> you build large single kernels to support a wide array of features which only
> late at run time can be determined. Virtualization and the pvops design just
> makes this issue much more prominent. If there are other areas of code exposed
> that actually may run, but we are not sure may run, I figured some other folks
> with a bit more security conscience minds might even simply take the position
> it may be a security risk to leave that code exposed. So to take a position
> that 'the worst it will do is try to read non-existent registers' -- seems
> rather shortsighted here.

Security conscious people trim their CONFIG.
>  
> Anyway for more details on thoughts on this refer to the this wiki:
> 
> http://kernelnewbies.org/KernelProjects/kernel-sandboxing
> 
> Since this is now getting off topic please send me your feedback on another
> thread for the non-virtualization aspects of this if that interests you. My
> point here was rather to highlight the importance of clear semantics due to
> virtualization in light of possible dead code.

Thank you.
> 
> > The VGA code should be able to handle failures like that and
> > not initialize itself when the hardware is dead (or non-existent).
> 
> That's right, its through ACPI_FADT_NO_VGA and since its part of the HVMLite
> design doc we want HVMlite design to address ACPI_FADT_NO_VGA properly.  I've
> paved the way for this to be done cleanly and easily now, but that code should
> be in place before HVMLite code gets merged.
> 
> Does domU for old Xen PV also set ACPI_FADT_NO_VGA as well ?  Should it ?

It does not. Not sure - it seems to have worked fine for the last ten
years?
> 
> > > To be clear -- dead code concerns still exist even without
> > > virtualization solutions, its just that with virtualization
> > > this stuff comes up more and there has been no proactive
> > > measures to address this. The question of semantics here is
> > > to see to what extent we need earlier boot code annotations
> > > to ensure we address semantics proactively.
> > 
> > I think what you mean by dead code is another word for
> > hardware test coverage?
> 
> No, no, its very different given that with virtualization the scope of possible
> dead code is significant and at run time you are certain a huge portion of code
> should *never ever* run. So for instance we know once we boot bare metal none
> of the Xen stuff should ever run, likewise on Xen dom0 we know none of the KVM
> / bare-metal only stuff should never run, when on Xen domU, none of the Xen

What is this 'bare metal only stuff' you speak of? On Xen dom0 most of
the baremetal code is running. In fact that is how the device drivers
work. Or are you talking about low level baremetal code? If so, then
PVH/HVMLite does that - it skips pvops so that it can run this
'low-level baremetal code'

> domU-only stuff should ever run.

You forgot KVM guest support on baremetal. That shouldn't run either.

> 
> > > > The entrace point in Linux "proper" is startup_32 or startup_64 - the same
> > > > path that EFI uses.
> > > > 
> > > > If you were to draw this (very simplified):
> > > > 
> > > > a)- GRUB2 ---------------------\ (creates an bootparam structure)
> > > >                                 \
> > > >                                  +---- startup_32 or startup_64
> > > > b) EFI -> Linux EFI stub -------/
> > > >        (creates bootparm)      /
> > > > c) GRUB2-EFI  -> Linux EFI----/
> > > >                stub         /
> > > > d) HVMLite ----------------/
> > > >       (creates bootparm)
> > > 
> > > b) and d) might be able to share paths there...
> > 
> > No idea. You would have to look in the assembler code to
> > figure that out.
> 
> And that's a pain, I get it.
> 
> I spotted one place already -- will note to Boris. I think Matt may have more
> ideas ;)
> 
> > > d) still has its own entry, it does more than create boot params.
> > 
> > d) purpose is to create boot params.
> 
> OK good to know that's the only thing we acknowledge it *should* do.

And b), c) purpose is for that too - amongts providing an mechanism
to call in EFI firmware.

And I realized that early baremetal boot option also ends up calling C during
its startup (see main in arch/x86/boot/main.c) amongst then switching
different modes.

> 
> >  It may do more as nobody likes to muck in assembler and make bootparams from
> >  within assembler.
> 
> OK -- it does do more and that's where we'd like to avoid duplication if
> possible and yet-another-entry (TM).

It does more? EFI stub entry does more than the GRUB2 entry.

If you have some patches to trim the code duplication within
those boot paths- please post it.
> 
> > > > (I am not sure about the c) - I would have to look in source to
> > > > be source). There is also LILO in this, but I am not even sure if
> > > > works anymore.
> > > > 
> > > > 
> > > > What you have is that every entry point creates the bootparams
> > > > and ends up calling startup_X. The startup_64 then hit the rest
> > > > of the kernel. The startp_X code is the one that would setup
> > > > the basic pagetables, segments, etc.
> > > 
> > > Sure.. a full diagram should include both sides and how when using
> > > a custom entry one runs the risk of skipping a lot of code setup.
> > 
> > But it does not skip a lot of code setup. It starts exactly
> > at the same code startup that _all_ bootstraping code start at.
> 
> Its a fair point.
> 
> > > There is that and as others have pointed out how certain guests types
> > > are assumed to not have certain peripherals, and we have no idea
> > > to ensure certain old legacy code may not ever run or be accessed
> > > by drivers.
> > 
> > Ok, but that is not at code setup. That is later - when device
> > drivers are initialized. This no different than booting on
> > some hardware with missing functionality. ACPI, PCI and PnP
> > PnP are set there to help OSes discover this.
> 
> To a certain extent this is true, but there may things which are missing still.

Like?
> 
> We really have no idea what the full list of those things are.

Ok, it sounds like you have some homework.
> 
> It may be that things may have been running for ages without notice of an issue
> or that only under certain situations will certain issues or bugs trigger a
> failure. For instance, just yesterday I was Cc'd on a brand-spanking new legacy
> conflict [0], caused by upstream commit 8c058b0b9c34d8c ("x86/irq: Probe for
> PIC presence before allocating descs for legacy IRQs") merged on v4.4 where
> some new code used nr_legacy_irqs() -- one proposed solution seems to be that
> for Xen code NR_IRQS_LEGACY should be used instead is as it lacks PCI [1] and
> another was to peg the legacy requirements as a quirk on the new x86 platform
> legacy quirk stuff [2]. Are other uses of nr_legacy_irqs() correct ? Are
> we sure ?

And how is this example related to 'early bootup' path?

It is not.

It is in fact related to PV codepaths - which PVH/HVMLite and HVM guests
do not exercise.
> 
> [0] http://lkml.kernel.org/r/570F90DF.1020508@oracle.com
> [1] https://lkml.org/lkml/2016/4/14/532
> [2] http://lkml.kernel.org/r/1460592286-300-1-git-send-email-mcgrof@kernel.org
> 
> > > > > How we address semantics then is *very* important to me.
> > > > 
> > > > Which semantics? How the CPU is going to be at startup_X ? Or
> > > > how the CPU is going to be when EFI firmware invokes the EFI stub?
> > > > Or when GRUB2 loads Linux?
> > > 
> > > What hypervisor kicked me and what guest type I am.
> > 
> > cpuid software flags have that - and that semantics has been 
> > there for eons.
> 
> We cannot use cpuid early in asm code, I'm looking for something we

?! Why!?
> can even use on asm early in boot code, on x86 the best option we
> have is the boot_params, but I've even have had issues with that
> early in code, as I can only access it after load_idt() where I
> described my effort to unify Xen PV and x86_64 init paths [3].

Well, Xen PV skips x86_64_start_kernel..
> 
> [3] http://lkml.kernel.org/r/CAB=NE6VTCRCazcNpCdJ7pN1eD3=x_fcGOdH37MzVpxkKEN5esw@mail.gmail.com
> 
> > > Let me elaborate more below.
> > > 
> > > > That (those bootloaders) is clearly defined. The URL I provided
> > > > mentions the HVMLite one. The Documentation/x86/boot.c mentions
> > > > what the semantics are to expected when providing an bootstrap
> > > > (which is what HVMLitel stub code in Linux would write against -
> > > > and what EFI stub code had been written against too).
> > > > > 
> > > > > > > I'll elaborate on this but first let's clarify why a new entry is used for
> > > > > > > HVMlite to start of with:
> > > > > > > 
> > > > > > >   1) Xen ABI has historically not wanted to set up the boot params for Linux
> > > > > > >      guests, instead it insists on letting the Linux kernel Xen boot stubs fill
> > > > > > >      that out for it. This sticking point means it has implicated a boot stub.
> > > > > > 
> > > > > > 
> > > > > > Which is b/c it has to be OS agnostic. It has nothing to do 'not wanting'.
> > > > > 
> > > > > It can still be OS agnostic and pass on type and custom data pointer.
> > > > 
> > > > Sure. It has that (it MUST otherwise how else would you pass data).
> > > > It is documented as well http://xenbits.xen.org/docs/unstable/hypercall/x86_64/include,public,xen.h.html#incontents_startofday
> > > > (see " Start of day structure passed to PVH guests in %ebx.")
> > > 
> > > The design doc begs for a custom OS entry point though.
> > 
> > That is what the ELF Note has.
> 
> Right, but I'm saying that its rather silly to be adding entry points if
> all we want the code to do is copy the boot params for us. The design
> doc requires a new entry, and likewise you'd need yet-another-entry if
> HVMLite is thrown out the window and come back 5 years later after new
> hardware solutions are in place and need to redesign HVMLite. Kind of

Why would you need to redesign HVMLite based on hardware solutions?
The entrace point and the CPU state are pretty well known - it is akin
to what GRUB2 bootloader path is (protected mode).
> where we are with PVH today. Likewise if other paravirtualization
> developers want to support Linux and want to copy your strategy they'd
> add yet-another-entry-point as well.
> 
> This is dumb.

You saying the EFI entry point is dumb? That instead the EFI
firmware should understand Linux bootparams and booted that?

> 
> > > If we had a single 'type' and 'custom data' passed to the kernel that
> > > should suffice for the default Linux entry point to just pivot off
> > > of that and do what it needs without more entry points. Once.
> > 
> > And what about ramdisk? What about multiple ramdisks?
> > What about command line? All of that is what bootparams
> > tries to unify on Linux. But 'bootparams' is unique to Linux,
> > it does not exist on FreeBSD. Hence some stub code to transplant
> > OS-agnostic simple data to OS-specific is neccessary.
> 
> If we had a Xen ABI option where *all* that I'm asking is you pass
> first:
> 
>   a) hypervisor type

Why can't you use cpuid.
>   b) custom data pointer

What is this custom data pointer you speak of?
> 
> We'd be able to avoid adding *any* entry point and just address
> the requirements as I noted with pre / post stubs for the type.

But you need some entry point to call into Linux. Are you
suggesting to use the existing ones? No, the existing one
wouldn't understand this.

> This would require an x86 boot protocol bump, but all the issues
> creeping up randomly I think that's worth putting on the table now.

Aaaah, so you are saying expand the bootparams. In other words
make Xen ABI call into Linux using the bootparams structure, similar
to how GRUB2 does it.

How is that OS agnostic?
> 
> And maybe we don't want it to be hypervisor specific, perhaps there are other
> *needs* for custom pre-post startup_32()/startup_64() stubs.

Multiboot?
> 
> To avoid extending boot_params further I figured perhaps we can look
> at EFI as another option instead. If we are going to drop all legacy

But EFI support is _huge_.
> PV support from the kernel (not the hypervisor) and require hardware
> virtualization 5 years from now on the Linux kernel, it doesn't seem
> to me far fetched to at the very least consider using an EFI entry
> instead, specially since all it does is set boot params and we can
> make re-use this for HVMLite too.

But to make that work you have to emulate EFI firmware in the
hypervisor. Is that work you are signing up for?
> 
>   Luis

[toc] | [prev] | [next] | [standalone]


#1379301 — Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry

From"Luis R. Rodriguez" <mcgrof@kernel.org>
Date2016-04-14 23:00 +0200
SubjectRe: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry
Message-ID<rnWwN-8aU-5@gated-at.bofh.it>
In reply to#1379276
On Thu, Apr 14, 2016 at 03:56:53PM -0400, Konrad Rzeszutek Wilk wrote:
> On Thu, Apr 14, 2016 at 08:40:48PM +0200, Luis R. Rodriguez wrote:
> > On Wed, Apr 13, 2016 at 09:01:32PM -0400, Konrad Rzeszutek Wilk wrote:
> > > On Thu, Apr 14, 2016 at 12:23:17AM +0200, Luis R. Rodriguez wrote:
> > > > VGA code will be dead code for HVMlite for sure as the design doc
> > > > says it will not run VGA, the ACPI flag will be set but the check
> > > > for that is not yet on Linux. That means the VGA Linux code will
> > > > be there but we have no way to ensure it will not run nor that
> > > > anything will muck with it.
> > > 
> > > <shrugs> The worst it will do is try to read non-existent registers.
> > 
> > Really ?
> > 
> > Is that your position on all other possible dead code that may have been
> > possible on old Xen PV guests as well ?
> 
> This is not just with Xen - it with other device drivers that are being
> invoked on baremetal and are not present in hardware anymore.

Indeed, however virtualization makes this issue much more prominent.

> > As I hinted, after thinking about this for a while I realized that dead code is
> > likely present on bare metal as well even without virtualization, specially if
> 
> Yes!
> > you build large single kernels to support a wide array of features which only
> > late at run time can be determined. Virtualization and the pvops design just
> > makes this issue much more prominent. If there are other areas of code exposed
> > that actually may run, but we are not sure may run, I figured some other folks
> > with a bit more security conscience minds might even simply take the position
> > it may be a security risk to leave that code exposed. So to take a position
> > that 'the worst it will do is try to read non-existent registers' -- seems
> > rather shortsighted here.
> 
> Security conscious people trim their CONFIG.

Not all Linux distributions want to do this, the more binaries the
higher the cost to test / vet.

> > Anyway for more details on thoughts on this refer to the this wiki:
> > 
> > http://kernelnewbies.org/KernelProjects/kernel-sandboxing
> > 
> > Since this is now getting off topic please send me your feedback on another
> > thread for the non-virtualization aspects of this if that interests you. My
> > point here was rather to highlight the importance of clear semantics due to
> > virtualization in light of possible dead code.
> 
> Thank you.
> > 
> > > The VGA code should be able to handle failures like that and
> > > not initialize itself when the hardware is dead (or non-existent).
> > 
> > That's right, its through ACPI_FADT_NO_VGA and since its part of the HVMLite
> > design doc we want HVMlite design to address ACPI_FADT_NO_VGA properly.  I've
> > paved the way for this to be done cleanly and easily now, but that code should
> > be in place before HVMLite code gets merged.
> > 
> > Does domU for old Xen PV also set ACPI_FADT_NO_VGA as well ?  Should it ?
> 
> It does not. Not sure - it seems to have worked fine for the last ten
> years?

Maybe HVMLite will need it enabled then too, just for bug parity.

> > > > To be clear -- dead code concerns still exist even without
> > > > virtualization solutions, its just that with virtualization
> > > > this stuff comes up more and there has been no proactive
> > > > measures to address this. The question of semantics here is
> > > > to see to what extent we need earlier boot code annotations
> > > > to ensure we address semantics proactively.
> > > 
> > > I think what you mean by dead code is another word for
> > > hardware test coverage?
> > 
> > No, no, its very different given that with virtualization the scope of possible
> > dead code is significant and at run time you are certain a huge portion of code
> > should *never ever* run. So for instance we know once we boot bare metal none
> > of the Xen stuff should ever run, likewise on Xen dom0 we know none of the KVM
> > / bare-metal only stuff should never run, when on Xen domU, none of the Xen
> 
> What is this 'bare metal only stuff' you speak of? On Xen dom0 most of
> the baremetal code is running.

A lot, not all. In the past folks added stubs (used to be paravirt_enabled()
checks) to some code, but we are simply not sure of other possible conflicts.
This is an known unknown if you will.

> In fact that is how the device drivers work. Or are you talking about low
> level baremetal code? If so, then PVH/HVMLite does that - it skips pvops so
> that it can run this 'low-level baremetal code'

Are you telling me that HVMLite has no dead code issues ?

> > domU-only stuff should ever run.
> 
> You forgot KVM guest support on baremetal. That shouldn't run either.

Glad you bring that up, yes, that is correct. I'm being just as cautious with
Xen as with KVM on their dead-code possible issues, however their dead code
conerns should be smaller given as you not the boot path.

It doesn't mean dead-cod concerns do not exist for KVM... or other
virtualization solutions.

> > > > > The entrace point in Linux "proper" is startup_32 or startup_64 - the same
> > > > > path that EFI uses.
> > > > > 
> > > > > If you were to draw this (very simplified):
> > > > > 
> > > > > a)- GRUB2 ---------------------\ (creates an bootparam structure)
> > > > >                                 \
> > > > >                                  +---- startup_32 or startup_64
> > > > > b) EFI -> Linux EFI stub -------/
> > > > >        (creates bootparm)      /
> > > > > c) GRUB2-EFI  -> Linux EFI----/
> > > > >                stub         /
> > > > > d) HVMLite ----------------/
> > > > >       (creates bootparm)
> > > > 
> > > > b) and d) might be able to share paths there...
> > > 
> > > No idea. You would have to look in the assembler code to
> > > figure that out.
> > 
> > And that's a pain, I get it.
> > 
> > I spotted one place already -- will note to Boris. I think Matt may have more
> > ideas ;)
> > 
> > > > d) still has its own entry, it does more than create boot params.
> > > 
> > > d) purpose is to create boot params.
> > 
> > OK good to know that's the only thing we acknowledge it *should* do.
> 
> And b), c) purpose is for that too - amongts providing an mechanism
> to call in EFI firmware.

Sure.

> And I realized that early baremetal boot option also ends up calling C during
> its startup (see main in arch/x86/boot/main.c) amongst then switching
> different modes.

Sure.

> > >  It may do more as nobody likes to muck in assembler and make bootparams from
> > >  within assembler.
> > 
> > OK -- it does do more and that's where we'd like to avoid duplication if
> > possible and yet-another-entry (TM).
> 
> It does more? EFI stub entry does more than the GRUB2 entry.
> 
> If you have some patches to trim the code duplication within
> those boot paths- please post it.

Sure.

> > > > > (I am not sure about the c) - I would have to look in source to
> > > > > be source). There is also LILO in this, but I am not even sure if
> > > > > works anymore.
> > > > > 
> > > > > 
> > > > > What you have is that every entry point creates the bootparams
> > > > > and ends up calling startup_X. The startup_64 then hit the rest
> > > > > of the kernel. The startp_X code is the one that would setup
> > > > > the basic pagetables, segments, etc.
> > > > 
> > > > Sure.. a full diagram should include both sides and how when using
> > > > a custom entry one runs the risk of skipping a lot of code setup.
> > > 
> > > But it does not skip a lot of code setup. It starts exactly
> > > at the same code startup that _all_ bootstraping code start at.
> > 
> > Its a fair point.
> > 
> > > > There is that and as others have pointed out how certain guests types
> > > > are assumed to not have certain peripherals, and we have no idea
> > > > to ensure certain old legacy code may not ever run or be accessed
> > > > by drivers.
> > > 
> > > Ok, but that is not at code setup. That is later - when device
> > > drivers are initialized. This no different than booting on
> > > some hardware with missing functionality. ACPI, PCI and PnP
> > > PnP are set there to help OSes discover this.
> > 
> > To a certain extent this is true, but there may things which are missing still.
> 
> Like?

That's the thing, I had a list of thing to look out for and then things
I ran across over code inspection. We need more work to be sure we're
really well covered.

Are you *sure* we have no dead code concerns with HVMLite ?
If there are dead code concerns are you sure there might not
be differences between KVM and HVMLite ? Should cpuid be used to
address differences ? Will that enable to distinguish between
hybrid versions of HVMLite ? Are we sure ?

> > We really have no idea what the full list of those things are.
> 
> Ok, it sounds like you have some homework.

We all do.

> > It may be that things may have been running for ages without notice of an issue
> > or that only under certain situations will certain issues or bugs trigger a
> > failure. For instance, just yesterday I was Cc'd on a brand-spanking new legacy
> > conflict [0], caused by upstream commit 8c058b0b9c34d8c ("x86/irq: Probe for
> > PIC presence before allocating descs for legacy IRQs") merged on v4.4 where
> > some new code used nr_legacy_irqs() -- one proposed solution seems to be that
> > for Xen code NR_IRQS_LEGACY should be used instead is as it lacks PCI [1] and
> > another was to peg the legacy requirements as a quirk on the new x86 platform
> > legacy quirk stuff [2]. Are other uses of nr_legacy_irqs() correct ? Are
> > we sure ?
> 
> And how is this example related to 'early bootup' path?
> 
> It is not.

For early boot code -- it is not. HVMLite is not merged, and PHV was never
completed.. so how are you sure we won't have any issues there ?

> It is in fact related to PV codepaths - which PVH/HVMLite and HVM guests
> do not exercise.

Agreed.

> > [0] http://lkml.kernel.org/r/570F90DF.1020508@oracle.com
> > [1] https://lkml.org/lkml/2016/4/14/532
> > [2] http://lkml.kernel.org/r/1460592286-300-1-git-send-email-mcgrof@kernel.org
> > 
> > > > > > How we address semantics then is *very* important to me.
> > > > > 
> > > > > Which semantics? How the CPU is going to be at startup_X ? Or
> > > > > how the CPU is going to be when EFI firmware invokes the EFI stub?
> > > > > Or when GRUB2 loads Linux?
> > > > 
> > > > What hypervisor kicked me and what guest type I am.
> > > 
> > > cpuid software flags have that - and that semantics has been 
> > > there for eons.
> > 
> > We cannot use cpuid early in asm code, I'm looking for something we
> 
> ?! Why!?

What existing code uses it? If there is nothing you are still certain
it should work ? Would that work for old PV guest as well BTW ?

> > can even use on asm early in boot code, on x86 the best option we
> > have is the boot_params, but I've even have had issues with that
> > early in code, as I can only access it after load_idt() where I
> > described my effort to unify Xen PV and x86_64 init paths [3].
> 
> Well, Xen PV skips x86_64_start_kernel..

Yes, and in doing so often times people skip adding Xen PV specific
code, as was the case with Kasan.

> > [3] http://lkml.kernel.org/r/CAB=NE6VTCRCazcNpCdJ7pN1eD3=x_fcGOdH37MzVpxkKEN5esw@mail.gmail.com
> > 
> > > > Let me elaborate more below.
> > > > 
> > > > > That (those bootloaders) is clearly defined. The URL I provided
> > > > > mentions the HVMLite one. The Documentation/x86/boot.c mentions
> > > > > what the semantics are to expected when providing an bootstrap
> > > > > (which is what HVMLitel stub code in Linux would write against -
> > > > > and what EFI stub code had been written against too).
> > > > > > 
> > > > > > > > I'll elaborate on this but first let's clarify why a new entry is used for
> > > > > > > > HVMlite to start of with:
> > > > > > > > 
> > > > > > > >   1) Xen ABI has historically not wanted to set up the boot params for Linux
> > > > > > > >      guests, instead it insists on letting the Linux kernel Xen boot stubs fill
> > > > > > > >      that out for it. This sticking point means it has implicated a boot stub.
> > > > > > > 
> > > > > > > 
> > > > > > > Which is b/c it has to be OS agnostic. It has nothing to do 'not wanting'.
> > > > > > 
> > > > > > It can still be OS agnostic and pass on type and custom data pointer.
> > > > > 
> > > > > Sure. It has that (it MUST otherwise how else would you pass data).
> > > > > It is documented as well http://xenbits.xen.org/docs/unstable/hypercall/x86_64/include,public,xen.h.html#incontents_startofday
> > > > > (see " Start of day structure passed to PVH guests in %ebx.")
> > > > 
> > > > The design doc begs for a custom OS entry point though.
> > > 
> > > That is what the ELF Note has.
> > 
> > Right, but I'm saying that its rather silly to be adding entry points if
> > all we want the code to do is copy the boot params for us. The design
> > doc requires a new entry, and likewise you'd need yet-another-entry if
> > HVMLite is thrown out the window and come back 5 years later after new
> > hardware solutions are in place and need to redesign HVMLite. Kind of
> 
> Why would you need to redesign HVMLite based on hardware solutions?

That's what happened to Xen PV, right ? Are we sure 5 years from now we won't
have any new hardware virtualization features that will just obsolete HVMLite?

> The entrace point and the CPU state are pretty well known - it is akin
> to what GRUB2 bootloader path is (protected mode).
> > where we are with PVH today. Likewise if other paravirtualization
> > developers want to support Linux and want to copy your strategy they'd
> > add yet-another-entry-point as well.
> > 
> > This is dumb.
> 
> You saying the EFI entry point is dumb? That instead the EFI
> firmware should understand Linux bootparams and booted that?

EFI is a standard. Xen is not. And since we are not talking about legacy
hardware in the future, EFI seems like a sensible option to consider for an
entry point. Specially given that it may mean that we can ultimately also help
unify more entry points on Linux in general. I'd prefer to consider using
EFI configuration tables instead of extending the x86 boot protocol.

> > > > If we had a single 'type' and 'custom data' passed to the kernel that
> > > > should suffice for the default Linux entry point to just pivot off
> > > > of that and do what it needs without more entry points. Once.
> > > 
> > > And what about ramdisk? What about multiple ramdisks?
> > > What about command line? All of that is what bootparams
> > > tries to unify on Linux. But 'bootparams' is unique to Linux,
> > > it does not exist on FreeBSD. Hence some stub code to transplant
> > > OS-agnostic simple data to OS-specific is neccessary.
> > 
> > If we had a Xen ABI option where *all* that I'm asking is you pass
> > first:
> > 
> >   a) hypervisor type
> 
> Why can't you use cpuid.

I'll evaluate that.

> >   b) custom data pointer
> 
> What is this custom data pointer you speak of?

For Xen this is the en_start_info, the structure that Xen stuffs in
a copy of its version of what we need to fill the boot_params.

> > We'd be able to avoid adding *any* entry point and just address
> > the requirements as I noted with pre / post stubs for the type.
> 
> But you need some entry point to call into Linux. Are you
> suggesting to use the existing ones? No, the existing one
> wouldn't understand this.

If we used the boot_parms, yes it would be possible.

> > This would require an x86 boot protocol bump, but all the issues
> > creeping up randomly I think that's worth putting on the table now.
> 
> Aaaah, so you are saying expand the bootparams. In other words
> make Xen ABI call into Linux using the bootparams structure, similar
> to how GRUB2 does it.
> 
> How is that OS agnostic?

That's an issue, I understand. EFI is OS agnostic though.

> > And maybe we don't want it to be hypervisor specific, perhaps there are other
> > *needs* for custom pre-post startup_32()/startup_64() stubs.
> 
> Multiboot?

Can you elaborate?

> > To avoid extending boot_params further I figured perhaps we can look
> > at EFI as another option instead. If we are going to drop all legacy
> 
> But EFI support is _huge_.

I get the sense now. Perhaps we should explore to what extent now really
at the Hackathon.

> > PV support from the kernel (not the hypervisor) and require hardware
> > virtualization 5 years from now on the Linux kernel, it doesn't seem
> > to me far fetched to at the very least consider using an EFI entry
> > instead, specially since all it does is set boot params and we can
> > make re-use this for HVMLite too.
> 
> But to make that work you have to emulate EFI firmware in the
> hypervisor. Is that work you are signing up for?

I'll do what is needed, as I have done before. If EFI is on the long
term roadmap for ARM perhaps there are a few birds to knock with one
stone here. If there is also interest to support other OSes through
EFI standard means this also should help make that easier.

  Luis

[toc] | [prev] | [next] | [standalone]


#1379417 — Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry

FromKonrad Rzeszutek Wilk <konrad.wilk@oracle.com>
Date2016-04-15 04:10 +0200
SubjectRe: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry
Message-ID<ro1mN-3Dy-3@gated-at.bofh.it>
In reply to#1379301
On Thu, Apr 14, 2016 at 10:56:19PM +0200, Luis R. Rodriguez wrote:
> On Thu, Apr 14, 2016 at 03:56:53PM -0400, Konrad Rzeszutek Wilk wrote:
> > On Thu, Apr 14, 2016 at 08:40:48PM +0200, Luis R. Rodriguez wrote:
> > > On Wed, Apr 13, 2016 at 09:01:32PM -0400, Konrad Rzeszutek Wilk wrote:
> > > > On Thu, Apr 14, 2016 at 12:23:17AM +0200, Luis R. Rodriguez wrote:
> > > > > VGA code will be dead code for HVMlite for sure as the design doc
> > > > > says it will not run VGA, the ACPI flag will be set but the check
> > > > > for that is not yet on Linux. That means the VGA Linux code will
> > > > > be there but we have no way to ensure it will not run nor that
> > > > > anything will muck with it.
> > > > 
> > > > <shrugs> The worst it will do is try to read non-existent registers.
> > > 
> > > Really ?
> > > 
> > > Is that your position on all other possible dead code that may have been
> > > possible on old Xen PV guests as well ?
> > 
> > This is not just with Xen - it with other device drivers that are being
> > invoked on baremetal and are not present in hardware anymore.
> 
> Indeed, however virtualization makes this issue much more prominent.

I suppose - as it only exposes a certain type of platform and nothing
else.
> 
> > > As I hinted, after thinking about this for a while I realized that dead code is
> > > likely present on bare metal as well even without virtualization, specially if
> > 
> > Yes!
> > > you build large single kernels to support a wide array of features which only
> > > late at run time can be determined. Virtualization and the pvops design just
> > > makes this issue much more prominent. If there are other areas of code exposed
> > > that actually may run, but we are not sure may run, I figured some other folks
> > > with a bit more security conscience minds might even simply take the position
> > > it may be a security risk to leave that code exposed. So to take a position
> > > that 'the worst it will do is try to read non-existent registers' -- seems
> > > rather shortsighted here.
> > 
> > Security conscious people trim their CONFIG.
> 
> Not all Linux distributions want to do this, the more binaries the
> higher the cost to test / vet.

OK, but Linux distributions have many goals - and are pulled in
different directions so they cannot always achieve the 'low footprint -
small amount of code to do inspection from security standpoint'

> 
> > > Anyway for more details on thoughts on this refer to the this wiki:
> > > 
> > > http://kernelnewbies.org/KernelProjects/kernel-sandboxing
> > > 
> > > Since this is now getting off topic please send me your feedback on another
> > > thread for the non-virtualization aspects of this if that interests you. My
> > > point here was rather to highlight the importance of clear semantics due to
> > > virtualization in light of possible dead code.
> > 
> > Thank you.
> > > 
> > > > The VGA code should be able to handle failures like that and
> > > > not initialize itself when the hardware is dead (or non-existent).
> > > 
> > > That's right, its through ACPI_FADT_NO_VGA and since its part of the HVMLite
> > > design doc we want HVMlite design to address ACPI_FADT_NO_VGA properly.  I've
> > > paved the way for this to be done cleanly and easily now, but that code should
> > > be in place before HVMLite code gets merged.
> > > 
> > > Does domU for old Xen PV also set ACPI_FADT_NO_VGA as well ?  Should it ?
> > 
> > It does not. Not sure - it seems to have worked fine for the last ten
> > years?
> 
> Maybe HVMLite will need it enabled then too, just for bug parity.

<shrugs> Sure.
> 
> > > > > To be clear -- dead code concerns still exist even without
> > > > > virtualization solutions, its just that with virtualization
> > > > > this stuff comes up more and there has been no proactive
> > > > > measures to address this. The question of semantics here is
> > > > > to see to what extent we need earlier boot code annotations
> > > > > to ensure we address semantics proactively.
> > > > 
> > > > I think what you mean by dead code is another word for
> > > > hardware test coverage?
> > > 
> > > No, no, its very different given that with virtualization the scope of possible
> > > dead code is significant and at run time you are certain a huge portion of code
> > > should *never ever* run. So for instance we know once we boot bare metal none
> > > of the Xen stuff should ever run, likewise on Xen dom0 we know none of the KVM
> > > / bare-metal only stuff should never run, when on Xen domU, none of the Xen
> > 
> > What is this 'bare metal only stuff' you speak of? On Xen dom0 most of
> > the baremetal code is running.
> 
> A lot, not all. In the past folks added stubs (used to be paravirt_enabled()
> checks) to some code, but we are simply not sure of other possible conflicts.
> This is an known unknown if you will.
> 
> > In fact that is how the device drivers work. Or are you talking about low
> > level baremetal code? If so, then PVH/HVMLite does that - it skips pvops so
> > that it can run this 'low-level baremetal code'
> 
> Are you telling me that HVMLite has no dead code issues ?

You said earlier that baremetal has dead code issue. Then by extensions
_any_ execution path has dead code issues.

..snip..
> > > > > There is that and as others have pointed out how certain guests types
> > > > > are assumed to not have certain peripherals, and we have no idea
> > > > > to ensure certain old legacy code may not ever run or be accessed
> > > > > by drivers.
> > > > 
> > > > Ok, but that is not at code setup. That is later - when device
> > > > drivers are initialized. This no different than booting on
> > > > some hardware with missing functionality. ACPI, PCI and PnP
> > > > PnP are set there to help OSes discover this.
> > > 
> > > To a certain extent this is true, but there may things which are missing still.
> > 
> > Like?
> 
> That's the thing, I had a list of thing to look out for and then things
> I ran across over code inspection. We need more work to be sure we're
> really well covered.
> 
> Are you *sure* we have no dead code concerns with HVMLite ?
> If there are dead code concerns are you sure there might not
> be differences between KVM and HVMLite ? Should cpuid be used to
> address differences ? Will that enable to distinguish between
> hybrid versions of HVMLite ? Are we sure ?

HVMLite CPU semantics will be the same as what a baremetal CPU
semantics are.

Platform wise it will be different - as in, instead of say
having a speaker (to emulated it) or RTC clock (again, another
thing to emulate), or say IDE controller (again, another
thing to emulate), or Realtek network card (again, another
thing to emulate) - it has none of those.

[Keep in mind 'another thing to emulate', means 'another
@$@() thing in QEMU that could be a security bug']

So it differs from an consumer x86 platform in that it has
none of the 'legacy' stuff. And it requires PV drivers to
function. And since it requires PV drivers to function
only OSes that have those can use this mode.
> 
> > > We really have no idea what the full list of those things are.
> > 
> > Ok, it sounds like you have some homework.
> 
> We all do.
> 
> > > It may be that things may have been running for ages without notice of an issue
> > > or that only under certain situations will certain issues or bugs trigger a
> > > failure. For instance, just yesterday I was Cc'd on a brand-spanking new legacy
> > > conflict [0], caused by upstream commit 8c058b0b9c34d8c ("x86/irq: Probe for
> > > PIC presence before allocating descs for legacy IRQs") merged on v4.4 where
> > > some new code used nr_legacy_irqs() -- one proposed solution seems to be that
> > > for Xen code NR_IRQS_LEGACY should be used instead is as it lacks PCI [1] and
> > > another was to peg the legacy requirements as a quirk on the new x86 platform
> > > legacy quirk stuff [2]. Are other uses of nr_legacy_irqs() correct ? Are
> > > we sure ?
> > 
> > And how is this example related to 'early bootup' path?
> > 
> > It is not.
> 
> For early boot code -- it is not. HVMLite is not merged, and PHV was never
> completed.. so how are you sure we won't have any issues there ?

If we did not have issues we would be out of jobs.

But this is a seperate topic - it is an issue about device drivers and
the assumptions they have. And those assumptions are not always
true (even with normal hardware).

> 
> > It is in fact related to PV codepaths - which PVH/HVMLite and HVM guests
> > do not exercise.
> 
> Agreed.
> 
> > > [0] http://lkml.kernel.org/r/570F90DF.1020508@oracle.com
> > > [1] https://lkml.org/lkml/2016/4/14/532
> > > [2] http://lkml.kernel.org/r/1460592286-300-1-git-send-email-mcgrof@kernel.org
> > > 
> > > > > > > How we address semantics then is *very* important to me.
> > > > > > 
> > > > > > Which semantics? How the CPU is going to be at startup_X ? Or
> > > > > > how the CPU is going to be when EFI firmware invokes the EFI stub?
> > > > > > Or when GRUB2 loads Linux?
> > > > > 
> > > > > What hypervisor kicked me and what guest type I am.
> > > > 
> > > > cpuid software flags have that - and that semantics has been 
> > > > there for eons.
> > > 
> > > We cannot use cpuid early in asm code, I'm looking for something we
> > 
> > ?! Why!?
> 
> What existing code uses it? If there is nothing you are still certain
> it should work ? Would that work for old PV guest as well BTW ?

Yeah. For HVM/HVMLite it traps to the hypervisor.

For old PV guests it is unwise to use it as it goes straight to
the hardware (as PV guests run in ring3 - they are considered
'userspace' and the Intel nor AMD do not trap on 'cpuid' in ring3
-unless  you run in an VMX container).
> 
> > > can even use on asm early in boot code, on x86 the best option we
> > > have is the boot_params, but I've even have had issues with that
> > > early in code, as I can only access it after load_idt() where I
> > > described my effort to unify Xen PV and x86_64 init paths [3].
> > 
> > Well, Xen PV skips x86_64_start_kernel..
> 
> Yes, and in doing so often times people skip adding Xen PV specific
> code, as was the case with Kasan.

Right. That is an existing problem Xen PV code has.

> 
> > > [3] http://lkml.kernel.org/r/CAB=NE6VTCRCazcNpCdJ7pN1eD3=x_fcGOdH37MzVpxkKEN5esw@mail.gmail.com
> > > 
> > > > > Let me elaborate more below.
> > > > > 
> > > > > > That (those bootloaders) is clearly defined. The URL I provided
> > > > > > mentions the HVMLite one. The Documentation/x86/boot.c mentions
> > > > > > what the semantics are to expected when providing an bootstrap
> > > > > > (which is what HVMLitel stub code in Linux would write against -
> > > > > > and what EFI stub code had been written against too).
> > > > > > > 
> > > > > > > > > I'll elaborate on this but first let's clarify why a new entry is used for
> > > > > > > > > HVMlite to start of with:
> > > > > > > > > 
> > > > > > > > >   1) Xen ABI has historically not wanted to set up the boot params for Linux
> > > > > > > > >      guests, instead it insists on letting the Linux kernel Xen boot stubs fill
> > > > > > > > >      that out for it. This sticking point means it has implicated a boot stub.
> > > > > > > > 
> > > > > > > > 
> > > > > > > > Which is b/c it has to be OS agnostic. It has nothing to do 'not wanting'.
> > > > > > > 
> > > > > > > It can still be OS agnostic and pass on type and custom data pointer.
> > > > > > 
> > > > > > Sure. It has that (it MUST otherwise how else would you pass data).
> > > > > > It is documented as well http://xenbits.xen.org/docs/unstable/hypercall/x86_64/include,public,xen.h.html#incontents_startofday
> > > > > > (see " Start of day structure passed to PVH guests in %ebx.")
> > > > > 
> > > > > The design doc begs for a custom OS entry point though.
> > > > 
> > > > That is what the ELF Note has.
> > > 
> > > Right, but I'm saying that its rather silly to be adding entry points if
> > > all we want the code to do is copy the boot params for us. The design
> > > doc requires a new entry, and likewise you'd need yet-another-entry if
> > > HVMLite is thrown out the window and come back 5 years later after new
> > > hardware solutions are in place and need to redesign HVMLite. Kind of
> > 
> > Why would you need to redesign HVMLite based on hardware solutions?
> 
> That's what happened to Xen PV, right ? Are we sure 5 years from now we won't
> have any new hardware virtualization features that will just obsolete HVMLite?

There were no hardware virtualization when Xen PV came about.

If there is hardware virtualization that obsoletes HVMLite that means
it would also obsolete KVM and HVM mode - as HVMLite runs in an VMX
container - the same type that KVM and Xen HVM guests run in.

> 
> > The entrace point and the CPU state are pretty well known - it is akin
> > to what GRUB2 bootloader path is (protected mode).
> > > where we are with PVH today. Likewise if other paravirtualization
> > > developers want to support Linux and want to copy your strategy they'd
> > > add yet-another-entry-point as well.
> > > 
> > > This is dumb.
> > 
> > You saying the EFI entry point is dumb? That instead the EFI
> > firmware should understand Linux bootparams and booted that?
> 
> EFI is a standard. Xen is not. And since we are not talking about legacy

And is a standard something that has to come out of a committee?

If so, then Linux bootparams is not a standard. Nor is LILO bootup
path.

> hardware in the future, EFI seems like a sensible option to consider for an
> entry point. Specially given that it may mean that we can ultimately also help
> unify more entry points on Linux in general. I'd prefer to consider using

<chokes>
I can just see that. On non-EFI hardware GRUB2/SYSLINUX would use the EFI entry
point and create an fake firmware.
> EFI configuration tables instead of extending the x86 boot protocol.

What is that? Are you talking about EFI runtime services? Take a look
at the EFI spec and see what you have to implement to emulate this.
> 
> > > > > If we had a single 'type' and 'custom data' passed to the kernel that
> > > > > should suffice for the default Linux entry point to just pivot off
> > > > > of that and do what it needs without more entry points. Once.
> > > > 
> > > > And what about ramdisk? What about multiple ramdisks?
> > > > What about command line? All of that is what bootparams
> > > > tries to unify on Linux. But 'bootparams' is unique to Linux,
> > > > it does not exist on FreeBSD. Hence some stub code to transplant
> > > > OS-agnostic simple data to OS-specific is neccessary.
> > > 
> > > If we had a Xen ABI option where *all* that I'm asking is you pass
> > > first:
> > > 
> > >   a) hypervisor type
> > 
> > Why can't you use cpuid.
> 
> I'll evaluate that.
> 
> > >   b) custom data pointer
> > 
> > What is this custom data pointer you speak of?
> 
> For Xen this is the en_start_info, the structure that Xen stuffs in
> a copy of its version of what we need to fill the boot_params.

Ok, but that is what we do in some way provide.

I am lost here. You seem to saying you want something that is
already there?

> 
> > > We'd be able to avoid adding *any* entry point and just address
> > > the requirements as I noted with pre / post stubs for the type.
> > 
> > But you need some entry point to call into Linux. Are you
> > suggesting to use the existing ones? No, the existing one
> > wouldn't understand this.
> 
> If we used the boot_parms, yes it would be possible.

...OS agnostic... they are not.

> 
> > > This would require an x86 boot protocol bump, but all the issues
> > > creeping up randomly I think that's worth putting on the table now.
> > 
> > Aaaah, so you are saying expand the bootparams. In other words
> > make Xen ABI call into Linux using the bootparams structure, similar
> > to how GRUB2 does it.
> > 
> > How is that OS agnostic?
> 
> That's an issue, I understand. EFI is OS agnostic though.
> 
> > > And maybe we don't want it to be hypervisor specific, perhaps there are other
> > > *needs* for custom pre-post startup_32()/startup_64() stubs.
> > 
> > Multiboot?
> 
> Can you elaborate?

Google Multiboot specification.
> 
> > > To avoid extending boot_params further I figured perhaps we can look
> > > at EFI as another option instead. If we are going to drop all legacy
> > 
> > But EFI support is _huge_.
> 
> I get the sense now. Perhaps we should explore to what extent now really
> at the Hackathon.

Print out the EFI spec and carry it on the plane. The plane will tilt
to one side when trying to take off.

> 
> > > PV support from the kernel (not the hypervisor) and require hardware
> > > virtualization 5 years from now on the Linux kernel, it doesn't seem
> > > to me far fetched to at the very least consider using an EFI entry
> > > instead, specially since all it does is set boot params and we can
> > > make re-use this for HVMLite too.
> > 
> > But to make that work you have to emulate EFI firmware in the
> > hypervisor. Is that work you are signing up for?
> 
> I'll do what is needed, as I have done before. If EFI is on the long
> term roadmap for ARM perhaps there are a few birds to knock with one
> stone here. If there is also interest to support other OSes through
> EFI standard means this also should help make that easier.
> 
>   Luis

[toc] | [prev] | [next] | [standalone]


#1379662 — Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry

FromJulien Grall <julien.grall@arm.com>
Date2016-04-15 12:10 +0200
SubjectRe: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry
Message-ID<ro8Rk-1jQ-9@gated-at.bofh.it>
In reply to#1379301
Hello Luis,

On 14/04/16 21:56, Luis R. Rodriguez wrote:
> On Thu, Apr 14, 2016 at 03:56:53PM -0400, Konrad Rzeszutek Wilk wrote:
>> On Thu, Apr 14, 2016 at 08:40:48PM +0200, Luis R. Rodriguez wrote:
>>> On Wed, Apr 13, 2016 at 09:01:32PM -0400, Konrad Rzeszutek Wilk wrote:
>>>> On Thu, Apr 14, 2016 at 12:23:17AM +0200, Luis R. Rodriguez wrote:
>>> PV support from the kernel (not the hypervisor) and require hardware
>>> virtualization 5 years from now on the Linux kernel, it doesn't seem
>>> to me far fetched to at the very least consider using an EFI entry
>>> instead, specially since all it does is set boot params and we can
>>> make re-use this for HVMLite too.
>>
>> But to make that work you have to emulate EFI firmware in the
>> hypervisor. Is that work you are signing up for?
>
> I'll do what is needed, as I have done before. If EFI is on the long
> term roadmap for ARM perhaps there are a few birds to knock with one
> stone here. If there is also interest to support other OSes through
> EFI standard means this also should help make that easier.

We already have a working solution for EFI on ARM which does not require 
to emulate the firmware in the hypervisor.

On ARM, the EFI stub is communicating with the kernel using device-tree 
[1]. Once the EFI stub has ended, the native path (i.e non-UEFI) will be 
executed normally and it won't be possible to use BootServices anymore.

For the guest, we provide a full support of EFI using OVMF. For DOM0, 
Xen will craft the UEFI system table and the UEFI memory map. The 
locations of those tables will be passed to DOM0 using a tiny 
device-tree [1] and the kernel will boot using the native path. The 
runtime services for DOM0 will be provided via hypercall.

The DOM0 approach has been discussed for a long time (see [3]) and I 
believe this is better than emulating UEFI firmware in Xen. We want to 
keep Xen on ARM tiny. Adding any sort of emulation will increase the 
attack surface and require more maintenance from our side.

Regards,

[1] Documentation/arm/uefi.txt in Linux.

[2] 
http://xenbits.xen.org/docs/unstable-staging/misc/arm/device-tree/guest.txt

[3] http://www.gossamer-threads.com/lists/xen/devel/397349

-- 
Julien Grall

[toc] | [prev] | [next] | [standalone]


#1379895 — Re: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry

From"Luis R. Rodriguez" <mcgrof@kernel.org>
Date2016-04-15 17:00 +0200
SubjectRe: [Xen-devel] HVMLite / PVHv2 - using x86 EFI boot entry
Message-ID<rodnY-4BO-19@gated-at.bofh.it>
In reply to#1379662
On Fri, Apr 15, 2016 at 3:06 AM, Julien Grall <julien.grall@arm.com> wrote:
> On 14/04/16 21:56, Luis R. Rodriguez wrote:
>> On Thu, Apr 14, 2016 at 03:56:53PM -0400, Konrad Rzeszutek Wilk wrote:
>>> But to make that work you have to emulate EFI firmware in the
>>> hypervisor. Is that work you are signing up for?
>>
>> I'll do what is needed, as I have done before. If EFI is on the long
>> term roadmap for ARM perhaps there are a few birds to knock with one
>> stone here. If there is also interest to support other OSes through
>> EFI standard means this also should help make that easier.
>
> We already have a working solution for EFI on ARM which does not require to
> emulate the firmware in the hypervisor.

I get that.

> On ARM, the EFI stub is communicating with the kernel using device-tree [1].
> Once the EFI stub has ended, the native path (i.e non-UEFI) will be executed
> normally and it won't be possible to use BootServices anymore.
>
> For the guest, we provide a full support of EFI using OVMF.

I get that as well, is this the long term solution ? That still
requires OVMF, will relying on OVMF always be what is used on Xen ARM
? Was it too much of a burden to require OVMF? Is the upstream OVMF
code pulled by Xen at build time on ARM, or just wget a binary ?

> For DOM0, Xen
> will craft the UEFI system table and the UEFI memory map. The locations of
> those tables will be passed to DOM0 using a tiny device-tree [1] and the
> kernel will boot using the native path. The runtime services for DOM0 will
> be provided via hypercall.

Thanks this helps!

> The DOM0 approach has been discussed for a long time (see [3]) and I believe
> this is better than emulating UEFI firmware in Xen. We want to keep Xen on
> ARM tiny. Adding any sort of emulation will increase the attack surface and
> require more maintenance from our side.

OK thanks, would re-using OVMF (note, DT perhaps may not be ideal for
x86 for the rest though) be a reasonable solution on x86 as an option
then?

  Luis

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | linux.kernel


csiph-web