Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1736595 > unrolled thread

[patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall

Started byMarcelo Tosatti <mtosatti@redhat.com>
First post2017-09-21 13:50 +0200
Last post2017-09-25 20:40 +0200
Articles 7 on this page of 27 — 6 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall Marcelo Tosatti <mtosatti@redhat.com> - 2017-09-21 13:50 +0200
    Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2017-09-21 15:40 +0200
      Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Peter Zijlstra <peterz@infradead.org> - 2017-09-21 16:10 +0200
        Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Marcelo Tosatti <mtosatti@redhat.com> - 2017-09-22 03:20 +0200
          Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Peter Zijlstra <peterz@infradead.org> - 2017-09-22 12:10 +0200
            Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Peter Zijlstra <peterz@infradead.org> - 2017-09-22 13:00 +0200
              Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Marcelo Tosatti <mtosatti@redhat.com> - 2017-09-22 14:40 +0200
                Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Peter Zijlstra <peterz@infradead.org> - 2017-09-22 15:00 +0200
                  Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Paolo Bonzini <pbonzini@redhat.com> - 2017-09-23 13:00 +0200
                    Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Peter Zijlstra <peterz@infradead.org> - 2017-09-23 15:50 +0200
                      Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Paolo Bonzini <pbonzini@redhat.com> - 2017-09-24 15:10 +0200
                        Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Marcelo Tosatti <mtosatti@redhat.com> - 2017-09-25 05:00 +0200
                          Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Peter Zijlstra <peterz@infradead.org> - 2017-09-25 11:20 +0200
                            Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Paolo Bonzini <pbonzini@redhat.com> - 2017-09-25 17:20 +0200
                          Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2017-09-25 18:30 +0200
            Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Marcelo Tosatti <mtosatti@redhat.com> - 2017-09-22 14:20 +0200
              Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Peter Zijlstra <peterz@infradead.org> - 2017-09-22 14:40 +0200
                Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Marcelo Tosatti <mtosatti@redhat.com> - 2017-09-22 14:40 +0200
                  Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Peter Zijlstra <peterz@infradead.org> - 2017-09-22 15:10 +0200
                    Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Marcelo Tosatti <mtosatti@redhat.com> - 2017-09-25 04:30 +0200
                      Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall Peter Zijlstra <peterz@infradead.org> - 2017-09-25 10:40 +0200
                Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall\ Marcelo Tosatti <mtosatti@redhat.com> - 2017-09-22 14:50 +0200
                  Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall\ Peter Zijlstra <peterz@infradead.org> - 2017-09-22 15:10 +0200
                    Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall\ Marcelo Tosatti <mtosatti@redhat.com> - 2017-09-25 04:30 +0200
                      Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall\ Peter Zijlstra <peterz@infradead.org> - 2017-09-25 11:00 +0200
                      Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall\ Thomas Gleixner <tglx@linutronix.de> - 2017-09-25 12:50 +0200
                        Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO  hypercall\ Jan Kiszka <jan.kiszka@siemens.com> - 2017-09-25 20:40 +0200

Page 2 of 2 — ← Prev page 1 [2]


#1738827 — Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall

FromPeter Zijlstra <peterz@infradead.org>
Date2017-09-25 10:40 +0200
SubjectRe: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall
Message-ID<utx5M-3Sb-7@gated-at.bofh.it>
In reply to#1738712
On Sun, Sep 24, 2017 at 10:52:58PM -0300, Marcelo Tosatti wrote:
> On Fri, Sep 22, 2017 at 02:59:51PM +0200, Peter Zijlstra wrote:

> > Your patch is voodoo programming. You don't solve the actual problem,
> > you try and paper over it.
> 
> Priority boosting on a particular section of code is voodoo programming? 

Yes, because there's nothing that prevents said section of code from
triggering the exact problem you're trying to avoid.

The real fix is making sure the problem cannot happen to begin with,
which is done by fixing the interaction between the VCPU and that
emulation thread thing.

[toc] | [prev] | [next] | [standalone]


#1737476 — Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\

FromMarcelo Tosatti <mtosatti@redhat.com>
Date2017-09-22 14:50 +0200
SubjectRe: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\
Message-ID<usvz5-68T-45@gated-at.bofh.it>
In reply to#1737454
On Fri, Sep 22, 2017 at 02:31:07PM +0200, Peter Zijlstra wrote:
> On Fri, Sep 22, 2017 at 09:16:40AM -0300, Marcelo Tosatti wrote:
> > On Fri, Sep 22, 2017 at 12:00:05PM +0200, Peter Zijlstra wrote:
> > > On Thu, Sep 21, 2017 at 10:10:41PM -0300, Marcelo Tosatti wrote:
> > > > When executing guest vcpu-0 with FIFO:1 priority, which is necessary
> > > > to
> > > > deal with the following situation:
> > > > 
> > > > VCPU-0 (housekeeping VCPU)              VCPU-1 (realtime VCPU)
> > > > 
> > > > raw_spin_lock(A)
> > > > interrupted, schedule task T-1          raw_spin_lock(A) (spin)
> > > > 
> > > > raw_spin_unlock(A)
> > > > 
> > > > Certain operations must interrupt guest vcpu-0 (see trace below).
> > > 
> > > Those traces don't make any sense. All they include is kvm_exit and you
> > > can't tell anything from that.
> > 
> > Hi Peter,
> > 
> > OK lets describe whats happening:
> > 
> > With QEMU emulator thread and vcpu-0 sharing a physical CPU
> > (which is a request from several NFV customers, to improve
> > guest packing), the following occurs when the guest generates 
> > the following pattern:
> > 
> > 		1. submit IO.
> > 		2. busy spin.
> 
> User-space spinning is a bad idea in general and terminally broken in
> a RT setup. Sounds like you need to go fix qemu to not suck.

Are you arguing its invalid for the following application to execute on 
housekeeping vcpu of a realtime system:

void main(void)
{

    submit_IO();
    do {
       computation();
    } while (!interrupted());
}

Really?

Replace "busy spin" by "useful computation until interrupted". 

[toc] | [prev] | [next] | [standalone]


#1737486 — Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\

FromPeter Zijlstra <peterz@infradead.org>
Date2017-09-22 15:10 +0200
SubjectRe: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\
Message-ID<usvSq-6uj-7@gated-at.bofh.it>
In reply to#1737476
On Fri, Sep 22, 2017 at 09:40:05AM -0300, Marcelo Tosatti wrote:

> Are you arguing its invalid for the following application to execute on 
> housekeeping vcpu of a realtime system:
> 
> void main(void)
> {
> 
>     submit_IO();
>     do {
>        computation();
>     } while (!interrupted());
> }
> 
> Really?

No. Nobody cares about random crap tasks.

[toc] | [prev] | [next] | [standalone]


#1738713 — Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\

FromMarcelo Tosatti <mtosatti@redhat.com>
Date2017-09-25 04:30 +0200
SubjectRe: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\
Message-ID<utrjI-8rW-9@gated-at.bofh.it>
In reply to#1737486
On Fri, Sep 22, 2017 at 03:01:41PM +0200, Peter Zijlstra wrote:
> On Fri, Sep 22, 2017 at 09:40:05AM -0300, Marcelo Tosatti wrote:
> 
> > Are you arguing its invalid for the following application to execute on 
> > housekeeping vcpu of a realtime system:
> > 
> > void main(void)
> > {
> > 
> >     submit_IO();
> >     do {
> >        computation();
> >     } while (!interrupted());
> > }
> > 
> > Really?
> 
> No. Nobody cares about random crap tasks.

Nobody has control over all code that runs in userspace Peter. And not
supporting a valid sequence of steps because its "crap" (whatever your 
definition of crap is) makes no sense.

It might be that someone decides to do the above (i really can't see 
any actual reasoning i can follow and agree on your "its crap"
argument), this truly seems valid to me.

So lets follow the reasoning steps:

1) "NACK, because you didnt understand the problem".

	OK thats an invalid NACK, you did understand the problem
	later and now your argument is the following.

2) "NACK, because all VCPUs should be SCHED_FIFO all the time".

But the existence of this code path from userspace:

  submit_IO();
  do {
     computation();
  } while (!interrupted());

Its a supported code sequence, and works fine in a non-RT environment.
Therefore it should work on an -RT environment.
Think of any two applications, such as an IO application
and a CPU bound application. The IO application will be severely
impacted, or never execute, in such scenario.

Is that combination of tasks "random crap tasks" ? (No, its not, which 
makes me think you're just NACKing without giving enough thought to the
problem).

So please give me some logical reasoning for the NACK (people can live with
it, but it has to be good enough to justify the decreasing packing of 
guests in pCPUs):

1) "Voodoo programming" (its hard for me to parse what you mean with
that... do you mean you foresee this style of priority boosting causing
problems in the future? Can you give an example?).

Is there fundamentally wrong about priority boosting in spinlock
sections, or this particular style of priority boosting is wrong?

2) "Pollution of the kernel code path". That makes sense to me, if thats
whats your concerned about.

3) "Reduction of spinlock performance". Its true, but for NFV workloads
people don't care about.

4) "All vcpus should be SCHED_FIFO all the time". OK, why is that?
What dictates that to be true?

What the patch does is the following:
It reduces the window where SCHED_FIFO is applied vcpu0
to those were a spinlock is shared between -RT vcpus and vcpu0
(why: because otherwise, when the emulator thread is sharing a
pCPU with vcpu0, its unable to generate interrupts vcpu0).

And its being rejected because:
Please fill in.

[toc] | [prev] | [next] | [standalone]


#1738841 — Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\

FromPeter Zijlstra <peterz@infradead.org>
Date2017-09-25 11:00 +0200
SubjectRe: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\
Message-ID<utxp8-3ZQ-19@gated-at.bofh.it>
In reply to#1738713
On Sun, Sep 24, 2017 at 11:22:38PM -0300, Marcelo Tosatti wrote:
> On Fri, Sep 22, 2017 at 03:01:41PM +0200, Peter Zijlstra wrote:
> > On Fri, Sep 22, 2017 at 09:40:05AM -0300, Marcelo Tosatti wrote:
> > 
> > > Are you arguing its invalid for the following application to execute on 
> > > housekeeping vcpu of a realtime system:
> > > 
> > > void main(void)
> > > {
> > > 
> > >     submit_IO();
> > >     do {
> > >        computation();
> > >     } while (!interrupted());
> > > }
> > > 
> > > Really?
> > 
> > No. Nobody cares about random crap tasks.
> 
> Nobody has control over all code that runs in userspace Peter. And not
> supporting a valid sequence of steps because its "crap" (whatever your 
> definition of crap is) makes no sense.
> 
> It might be that someone decides to do the above (i really can't see 
> any actual reasoning i can follow and agree on your "its crap"
> argument), this truly seems valid to me.

We don't care what other tasks do. This isn't a hard thing to
understand. You're free to run whatever junk on your CPUs. This doesn't
(much) affect the correct functioning of RT tasks that you also run
there.

> So lets follow the reasoning steps:
> 
> 1) "NACK, because you didnt understand the problem".
> 
> 	OK thats an invalid NACK, you did understand the problem
> 	later and now your argument is the following.

It was a NACK because you wrote a shit changelog that didn't explain the
problem. But yes.

> 2) "NACK, because all VCPUs should be SCHED_FIFO all the time".

Very much, if you want a RT guest, all VCPU's should run at RT prio and
the interaction between the VCPUs and all supporting threads should be
designed for RT.

> But the existence of this code path from userspace:
> 
>   submit_IO();
>   do {
>      computation();
>   } while (!interrupted());
> 
> Its a supported code sequence, and works fine in a non-RT environment.

Who cares about that chunk of code? Have you forgotten to mention that
this is the form of the emulation thread?

> Therefore it should work on an -RT environment.

No, this is where you're wrong. That code works on -RT as long as you
don't expect it to be a valid RT program. -RT kernels will run !RT stuff
just fine.

But the moment you run a program as RT (FIFO/RR/DEADLINE) it had better
damn well be a valid RT program, and that excludes a lot of code.

> So please give me some logical reasoning for the NACK (people can live with
> it, but it has to be good enough to justify the decreasing packing of 
> guests in pCPUs):
> 
> 1) "Voodoo programming" (its hard for me to parse what you mean with
> that... do you mean you foresee this style of priority boosting causing
> problems in the future? Can you give an example?).

Your 'solution' only works if you sacrifice a goat on a full moon,
because only that ensures the guest doesn't VM_EXIT and cause the
self-same problem while you've boosted it.

Because you've _not_ fixed the actual problem!

> Is there fundamentally wrong about priority boosting in spinlock
> sections, or this particular style of priority boosting is wrong?

Yes, its fundamentally crap, because it doesn't guarantee anything.

RT is about making guarantees. An RT program needs a provable forward
progress guarantee at the very least. It including a priority inversion
disqualifies it from being sane.

> 2) "Pollution of the kernel code path". That makes sense to me, if thats
> whats your concerned about.

Also..

> 3) "Reduction of spinlock performance". Its true, but for NFV workloads
> people don't care about.

I've no idea what an NFV is.

> 4) "All vcpus should be SCHED_FIFO all the time". OK, why is that?
> What dictates that to be true?

Solid engineering. Does the guest kernel function as a bunch of
independent CPUs or does it assume all CPUs are equal and have strong
inter-cpu connections? Linux is the latter, therefore if one VCPU is RT
they all should be.

Dammit, you even recognise this in the spin-owner preemption issue
you're hacking around, but then go arse-about-face 'solving' it.

> What the patch does is the following:
> It reduces the window where SCHED_FIFO is applied vcpu0
> to those were a spinlock is shared between -RT vcpus and vcpu0
> (why: because otherwise, when the emulator thread is sharing a
> pCPU with vcpu0, its unable to generate interrupts vcpu0).
> 
> And its being rejected because:

Its not fixing the actual problem. The real problem is the prio
inversion between the VCPU and the emulation thread, _That_ is what
needs fixing.

Rewrite that VCPU/emulator interaction to be a proper RT construct.

Then you can run the VCPU at RT prio as you should, and the guest can
issue all the VM_EXIT things it wants at any time and still function
correctly.

[toc] | [prev] | [next] | [standalone]


#1738918 — Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\

FromThomas Gleixner <tglx@linutronix.de>
Date2017-09-25 12:50 +0200
SubjectRe: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\
Message-ID<utz7B-5aK-19@gated-at.bofh.it>
In reply to#1738713
On Sun, 24 Sep 2017, Marcelo Tosatti wrote:
> On Fri, Sep 22, 2017 at 03:01:41PM +0200, Peter Zijlstra wrote:
> What the patch does is the following:
> It reduces the window where SCHED_FIFO is applied vcpu0
> to those were a spinlock is shared between -RT vcpus and vcpu0
> (why: because otherwise, when the emulator thread is sharing a
> pCPU with vcpu0, its unable to generate interrupts vcpu0).
> 
> And its being rejected because:
> Please fill in.

Your patch is just papering over one particular problem, but it's not
fixing the root cause. That's the worst engineering approach and we all
know how fast this kind of crap falls over.

There are enough other issues which can cause starvation of the RT VCPUs
when the housekeeping VCPU is preempted, not just the particular problem
which you observed.

Back then when I did the first prototype of RT in KVM, I made it entirely
clear, that you have to spend one physical CPU for _each_ VCPU, independent
whether the VCPU is reserved for RT workers or the housekeeping VCPU. The
emulator thread needs to run on a separate physical CPU.

If you want to run the housekeeping VCPU and the emulator thread on the
same physical CPU then you have to make sure that both the emulator and the
housekeeper side of affairs are designed and implemented with RT in
mind. As long as that is not the case, you simply cannot run them on the
same physical CPU. RT is about guarantees and guarantees cannot be achieved
with bandaid engineering.

It's that simple, end of story.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


#1739211 — Re: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\

FromJan Kiszka <jan.kiszka@siemens.com>
Date2017-09-25 20:40 +0200
SubjectRe: [patch 3/3] x86: kvm guest side support for KVM_HC_RT_PRIO hypercall\
Message-ID<utGsp-1GA-7@gated-at.bofh.it>
In reply to#1738918
On 2017-09-25 12:41, Thomas Gleixner wrote:
> On Sun, 24 Sep 2017, Marcelo Tosatti wrote:
>> On Fri, Sep 22, 2017 at 03:01:41PM +0200, Peter Zijlstra wrote:
>> What the patch does is the following:
>> It reduces the window where SCHED_FIFO is applied vcpu0
>> to those were a spinlock is shared between -RT vcpus and vcpu0
>> (why: because otherwise, when the emulator thread is sharing a
>> pCPU with vcpu0, its unable to generate interrupts vcpu0).
>>
>> And its being rejected because:
>> Please fill in.
> 
> Your patch is just papering over one particular problem, but it's not
> fixing the root cause. That's the worst engineering approach and we all
> know how fast this kind of crap falls over.
> 
> There are enough other issues which can cause starvation of the RT VCPUs
> when the housekeeping VCPU is preempted, not just the particular problem
> which you observed.
> 
> Back then when I did the first prototype of RT in KVM, I made it entirely
> clear, that you have to spend one physical CPU for _each_ VCPU, independent
> whether the VCPU is reserved for RT workers or the housekeeping VCPU. The
> emulator thread needs to run on a separate physical CPU.
> 
> If you want to run the housekeeping VCPU and the emulator thread on the
> same physical CPU then you have to make sure that both the emulator and the
> housekeeper side of affairs are designed and implemented with RT in
> mind. As long as that is not the case, you simply cannot run them on the
> same physical CPU. RT is about guarantees and guarantees cannot be achieved
> with bandaid engineering.

It's even more complicated for the guest: It needs to be aware of the
latencies its interaction with a VM - instead of a real machine - may
cause while being in whatever critical sections. That's an additional
design dimension that would be very hard to establish and maintain, even
in Linux.

The only way around that is to truly decouple guest CPUs via full core
isolation inside the Linux guest and have your RT guest application
exploit this partitioning, e.g. by using lock-less inter-core
communication without kernel help.

The reason I was playing with PV-sched back then was to explore how you
could map the guest's task prio dynamically on its host vcpu. That
involved boosting whenever en event (aka irq) came in for the guest
vcpu. It turned out to be a more or less working solution looking for a
real-world problem.

Jan

-- 
Siemens AG, Corporate Technology, CT RDA ITP SES-DE
Corporate Competence Center Embedded Linux

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web