Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1165532 > unrolled thread

Re: [PATCH 4/5] x86/asm/entry/32: Replace RESTORE_RSI_RDI[_RDX] with open-coded 32-bit reads

Started byIngo Molnar <mingo@kernel.org>
First post2015-06-15 22:30 +0200
Last post2015-06-16 02:30 +0200
Articles 2 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH 4/5] x86/asm/entry/32: Replace RESTORE_RSI_RDI[_RDX] with  open-coded 32-bit reads Ingo Molnar <mingo@kernel.org> - 2015-06-15 22:30 +0200
    Re: [PATCH 4/5] x86/asm/entry/32: Replace RESTORE_RSI_RDI[_RDX] with  open-coded 32-bit reads Denys Vlasenko <dvlasenk@redhat.com> - 2015-06-16 02:30 +0200

#1165532 — Re: [PATCH 4/5] x86/asm/entry/32: Replace RESTORE_RSI_RDI[_RDX] with open-coded 32-bit reads

FromIngo Molnar <mingo@kernel.org>
Date2015-06-15 22:30 +0200
SubjectRe: [PATCH 4/5] x86/asm/entry/32: Replace RESTORE_RSI_RDI[_RDX] with open-coded 32-bit reads
Message-ID<pBJb4-5IV-3@gated-at.bofh.it>
* Denys Vlasenko <dvlasenk@redhat.com> wrote:

> On 06/14/2015 10:40 AM, Ingo Molnar wrote:
> > 
> > * Denys Vlasenko <dvlasenk@redhat.com> wrote:
> > 
> >> 	+8b 74 24 68            mov    0x68(%rsp),%esi
> >> 	+8b 7c 24 70            mov    0x70(%rsp),%edi
> >> 	+8b 54 24 60            mov    0x60(%rsp),%edx
> > 
> > Btw., could you (in another patch) order the restoration properly, by pt_regs 
> > memory order, where possible?
> 
> Will do.
> 
> > So this:
> > 
> >> +	movl	RSI(%rsp), %esi
> >> +	movl	RDI(%rsp), %edi
> >> +	movl	RDX(%rsp), %edx
> >>  	movl	RIP(%rsp), %ecx
> >>  	movl	EFLAGS(%rsp), %r11d
> > 
> > would become:
> > 
> > 	movl	RDX(%rsp), %edx
> > 	movl	RSI(%rsp), %esi
> > 	movl	RDI(%rsp), %edi
> > 	movl	RIP(%rsp), %ecx
> >  	movl	EFLAGS(%rsp), %r11d
> > 
> > ... or so.
> 
> Actually, ecx and r11 need to be loaded first. They are not so much "restored" 
> as "prepared for SYSRET insn". Every cycle lost in loading these delays SYSRET. 
> [...]

So in the typical case they will still be cached, and so their max latency should 
be around 3 cycles.

In fact because they are memory loads, they don't really have dependencies, so 
they should be available to SYSRET almost immediately, i.e. within a cycle - and 
there's no reason to believe why these loads wouldn't pipeline properly and 
parallelize with the many other things SYSRET has to do to organize a return to 
user-space, before it can actually use the target RIP and RFLAGS.

So I strongly doubt that the placement of the RCX and R11 load before the SYSRET 
matters to performance.

In any case this should be testable by looking at syscall performance and 
reordering the instructions.

Thanks,

	Ingo
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1165621

FromDenys Vlasenko <dvlasenk@redhat.com>
Date2015-06-16 02:30 +0200
Message-ID<pBMVj-2Nv-7@gated-at.bofh.it>
In reply to#1165532
On 06/15/2015 10:20 PM, Ingo Molnar wrote:
>> Actually, ecx and r11 need to be loaded first. They are not so much "restored" 
>> as "prepared for SYSRET insn". Every cycle lost in loading these delays SYSRET. 
>> [...]
> 
> So in the typical case they will still be cached, and so their max latency should 
> be around 3 cycles.

If syscall flushes caches (say, a large read), or sleeps
and CPU schedules away, then pt_regs->ip,flags are evicted
and need to be reloaded.

> In fact because they are memory loads, they don't really have dependencies,
> they should be available to SYSRET almost immediately,

They depend on the memory data.

> i.e. within a cycle - and 
> there's no reason to believe why these loads wouldn't pipeline properly and 
> parallelize with the many other things SYSRET has to do to organize a return to 
> user-space, before it can actually use the target RIP and RFLAGS.

This does not sound right.

If it takes, say, 20 cycles to pull data from e.g. L3 cache to ECX,
then SYSRET can't possibly complete sooner than in 20 cycles.

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web