Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1450509 > unrolled thread

Re: [PATCH v15 04/10] arm64: Kprobes with single stepping support

Started byDaniel Thompson <daniel.thompson@linaro.org>
First post2016-07-26 12:00 +0200
Last post2016-07-27 15:40 +0200
Articles 8 — 4 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH v15 04/10] arm64: Kprobes with single stepping support Daniel Thompson <daniel.thompson@linaro.org> - 2016-07-26 12:00 +0200
    Re: [PATCH v15 04/10] arm64: Kprobes with single stepping support Catalin Marinas <catalin.marinas@arm.com> - 2016-07-26 19:00 +0200
      Re: [PATCH v15 04/10] arm64: Kprobes with single stepping support Dave Martin <Dave.Martin@arm.com> - 2016-07-27 12:10 +0200
    Re: [PATCH v15 04/10] arm64: Kprobes with single stepping support Mark Rutland <mark.rutland@arm.com> - 2016-07-26 20:00 +0200
      Re: [PATCH v15 04/10] arm64: Kprobes with single stepping support Daniel Thompson <daniel.thompson@linaro.org> - 2016-07-27 13:30 +0200
        Re: [PATCH v15 04/10] arm64: Kprobes with single stepping support Dave Martin <Dave.Martin@arm.com> - 2016-07-27 13:40 +0200
          Re: [PATCH v15 04/10] arm64: Kprobes with single stepping support Daniel Thompson <daniel.thompson@linaro.org> - 2016-07-27 13:50 +0200
        Re: [PATCH v15 04/10] arm64: Kprobes with single stepping support Mark Rutland <mark.rutland@arm.com> - 2016-07-27 15:40 +0200

#1450509 — Re: [PATCH v15 04/10] arm64: Kprobes with single stepping support

FromDaniel Thompson <daniel.thompson@linaro.org>
Date2016-07-26 12:00 +0200
SubjectRe: [PATCH v15 04/10] arm64: Kprobes with single stepping support
Message-ID<rZ7jA-lC-21@gated-at.bofh.it>
On 25/07/16 18:13, Catalin Marinas wrote:
> On Fri, Jul 22, 2016 at 11:51:32AM -0400, David Long wrote:
>> On 07/22/2016 06:16 AM, Catalin Marinas wrote:
>>> On Thu, Jul 21, 2016 at 02:33:52PM -0400, David Long wrote:
>>> [...]
>>> The document states: "Up to MAX_STACK_SIZE bytes are copied". That means
>>> the arch code could always copy less but never more than MAX_STACK_SIZE.
>>> What we are proposing is that we should try to guess how much to copy
>>> based on the FP value (caller's frame) and, if larger than
>>> MAX_STACK_SIZE, skip the probe hook entirely. I don't think this goes
>>> against the kprobes.txt document but at least it (a) may improve the
>>> performance slightly by avoiding unnecessary copy and (b) it avoids
>>> undefined behaviour if we ever encounter a jprobe with arguments passed
>>> on the stack beyond MAX_STACK_SIZE.
>>
>> OK, it sounds like an improvement. I do worry a little about unexpected side
>> effects.
>
> You get more unexpected side effects by not saving/restoring the whole
> stack. We looked into this on Friday and came to the conclusion that
> there is no safe way for kprobes to know which arguments passed on the
> stack should be preserved, at least not with the current API.
>
> Basically the AArch64 PCS states that for arguments passed on the stack
> (e.g. they can't fit in registers), the caller allocates memory for them
> (on its own stack) and passes the pointer to the callee. Unfortunately,
> the frame pointer seems to be decremented correspondingly to cover the
> arguments, so we don't really have a way to tell how much to copy.
> Copying just the caller's stack frame isn't safe either since a
> callee/caller receiving such argument on the stack may passed it down to
> a callee without copying (I couldn't find anything in the PCS stating
> that this isn't allowed).

The PCS[1] seems (at least to me) to be pretty clear that "the address 
of the first stacked argument is defined to be the initial value of SP".

I think it is only the return value (when stacked via the x8 pointer) 
that can be passed through an intermediate function in the way described 
above. Isn't it OK for a jprobe to clobber this memory? The underlying 
function will overwrite whatever the jprobe put there anyway.

Am I overlooking some additional detail in the PCS?


Daniel.


[1] Google presented me revision IHI 0055B (via infocenter.arm.com)

[toc] | [next] | [standalone]


#1450684

FromCatalin Marinas <catalin.marinas@arm.com>
Date2016-07-26 19:00 +0200
Message-ID<rZdS1-4t5-1@gated-at.bofh.it>
In reply to#1450509
On Tue, Jul 26, 2016 at 10:50:08AM +0100, Daniel Thompson wrote:
> On 25/07/16 18:13, Catalin Marinas wrote:
> >On Fri, Jul 22, 2016 at 11:51:32AM -0400, David Long wrote:
> >>On 07/22/2016 06:16 AM, Catalin Marinas wrote:
> >>>On Thu, Jul 21, 2016 at 02:33:52PM -0400, David Long wrote:
> >>>[...]
> >>>The document states: "Up to MAX_STACK_SIZE bytes are copied". That means
> >>>the arch code could always copy less but never more than MAX_STACK_SIZE.
> >>>What we are proposing is that we should try to guess how much to copy
> >>>based on the FP value (caller's frame) and, if larger than
> >>>MAX_STACK_SIZE, skip the probe hook entirely. I don't think this goes
> >>>against the kprobes.txt document but at least it (a) may improve the
> >>>performance slightly by avoiding unnecessary copy and (b) it avoids
> >>>undefined behaviour if we ever encounter a jprobe with arguments passed
> >>>on the stack beyond MAX_STACK_SIZE.
> >>
> >>OK, it sounds like an improvement. I do worry a little about unexpected side
> >>effects.
> >
> >You get more unexpected side effects by not saving/restoring the whole
> >stack. We looked into this on Friday and came to the conclusion that
> >there is no safe way for kprobes to know which arguments passed on the
> >stack should be preserved, at least not with the current API.
> >
> >Basically the AArch64 PCS states that for arguments passed on the stack
> >(e.g. they can't fit in registers), the caller allocates memory for them
> >(on its own stack) and passes the pointer to the callee. Unfortunately,
> >the frame pointer seems to be decremented correspondingly to cover the
> >arguments, so we don't really have a way to tell how much to copy.
> >Copying just the caller's stack frame isn't safe either since a
> >callee/caller receiving such argument on the stack may passed it down to
> >a callee without copying (I couldn't find anything in the PCS stating
> >that this isn't allowed).
> 
> The PCS[1] seems (at least to me) to be pretty clear that "the address of
> the first stacked argument is defined to be the initial value of SP".
> 
> I think it is only the return value (when stacked via the x8 pointer) that
> can be passed through an intermediate function in the way described above.
> Isn't it OK for a jprobe to clobber this memory? The underlying function
> will overwrite whatever the jprobe put there anyway.
> 
> Am I overlooking some additional detail in the PCS?

I'm not sure I fully understand the PCS. I played with some random hacks
to test_kprobes.c (see below) and the address passed for a big struct
didn't look like the bottom of the stack.

diff --git a/kernel/test_kprobes.c b/kernel/test_kprobes.c
index 0dbab6d1acb4..6ed7be02a560 100644
--- a/kernel/test_kprobes.c
+++ b/kernel/test_kprobes.c
@@ -22,14 +22,18 @@
 
 #define div_factor 3
 
+struct dummy {
+	char dummy_array[MAX_STACK_SIZE * 2];
+};
+
 static u32 rand1, preh_val, posth_val, jph_val;
 static int errors, handler_errors, num_tests;
-static u32 (*target)(u32 value);
+static u32 (*target)(u32 value, struct dummy d);
 static u32 (*target2)(u32 value);
 
-static noinline u32 kprobe_target(u32 value)
+static noinline u32 kprobe_target(u32 value, struct dummy d)
 {
-	return (value / div_factor);
+	return (value / div_factor - d.dummy_array[0] + d.dummy_array[1]);
 }
 
 static int kp_pre_handler(struct kprobe *p, struct pt_regs *regs)
@@ -54,9 +58,11 @@ static struct kprobe kp = {
 	.post_handler = kp_post_handler
 };
 
-static int test_kprobe(void)
+static int noinline test_kprobe(void)
 {
 	int ret;
+	static struct dummy dummy;
+	memset(&dummy, 10, sizeof(dummy));
 
 	ret = register_kprobe(&kp);
 	if (ret < 0) {
@@ -64,7 +70,8 @@ static int test_kprobe(void)
 		return ret;
 	}
 
-	ret = target(rand1);
+	ret = target(rand1, dummy);
+	memset(&dummy, 10, sizeof(dummy));
 	unregister_kprobe(&kp);
 
 	if (preh_val == 0) {
@@ -111,6 +118,8 @@ static int test_kprobes(void)
 {
 	int ret;
 	struct kprobe *kps[2] = {&kp, &kp2};
+	struct dummy dummy;
+	memset(&dummy, 10, sizeof(dummy));
 
 	/* addr and flags should be cleard for reusing kprobe. */
 	kp.addr = NULL;
@@ -123,7 +132,7 @@ static int test_kprobes(void)
 
 	preh_val = 0;
 	posth_val = 0;
-	ret = target(rand1);
+	ret = target(rand1, dummy);
 
 	if (preh_val == 0) {
 		pr_err("kprobe pre_handler not called\n");
@@ -154,7 +163,7 @@ static int test_kprobes(void)
 
 }
 
-static u32 j_kprobe_target(u32 value)
+static u32 j_kprobe_target(u32 value, struct dummy d)
 {
 	if (value != rand1) {
 		handler_errors++;
@@ -174,6 +183,8 @@ static struct jprobe jp = {
 static int test_jprobe(void)
 {
 	int ret;
+	struct dummy dummy;
+	memset(&dummy, 10, sizeof(dummy));
 
 	ret = register_jprobe(&jp);
 	if (ret < 0) {
@@ -181,7 +192,7 @@ static int test_jprobe(void)
 		return ret;
 	}
 
-	ret = target(rand1);
+	ret = target(rand1, dummy);
 	unregister_jprobe(&jp);
 	if (jph_val == 0) {
 		pr_err("jprobe handler not called\n");
@@ -200,6 +211,8 @@ static int test_jprobes(void)
 {
 	int ret;
 	struct jprobe *jps[2] = {&jp, &jp2};
+	struct dummy dummy;
+	memset(&dummy, 10, sizeof(dummy));
 
 	/* addr and flags should be cleard for reusing kprobe. */
 	jp.kp.addr = NULL;
@@ -211,7 +224,7 @@ static int test_jprobes(void)
 	}
 
 	jph_val = 0;
-	ret = target(rand1);
+	ret = target(rand1, dummy);
 	if (jph_val == 0) {
 		pr_err("jprobe handler not called\n");
 		handler_errors++;
@@ -262,6 +275,8 @@ static struct kretprobe rp = {
 static int test_kretprobe(void)
 {
 	int ret;
+	struct dummy dummy;
+	memset(&dummy, 10, sizeof(dummy));
 
 	ret = register_kretprobe(&rp);
 	if (ret < 0) {
@@ -269,7 +284,7 @@ static int test_kretprobe(void)
 		return ret;
 	}
 
-	ret = target(rand1);
+	ret = target(rand1, dummy);
 	unregister_kretprobe(&rp);
 	if (krph_val != rand1) {
 		pr_err("kretprobe handler not called\n");
@@ -306,6 +321,8 @@ static int test_kretprobes(void)
 {
 	int ret;
 	struct kretprobe *rps[2] = {&rp, &rp2};
+	struct dummy dummy;
+	memset(&dummy, 10, sizeof(dummy));
 
 	/* addr and flags should be cleard for reusing kprobe. */
 	rp.kp.addr = NULL;
@@ -317,7 +334,7 @@ static int test_kretprobes(void)
 	}
 
 	krph_val = 0;
-	ret = target(rand1);
+	ret = target(rand1, dummy);
 	if (krph_val != rand1) {
 		pr_err("kretprobe handler not called\n");
 		handler_errors++;

-- 
Catalin

[toc] | [prev] | [next] | [standalone]


#1451207

FromDave Martin <Dave.Martin@arm.com>
Date2016-07-27 12:10 +0200
Message-ID<rZtWO-6qK-67@gated-at.bofh.it>
In reply to#1450684
On Tue, Jul 26, 2016 at 05:55:43PM +0100, Catalin Marinas wrote:
> On Tue, Jul 26, 2016 at 10:50:08AM +0100, Daniel Thompson wrote:
> > On 25/07/16 18:13, Catalin Marinas wrote:
> > >On Fri, Jul 22, 2016 at 11:51:32AM -0400, David Long wrote:
> > >>OK, it sounds like an improvement. I do worry a little about unexpected side
> > >>effects.
> > >
> > >You get more unexpected side effects by not saving/restoring the whole
> > >stack. We looked into this on Friday and came to the conclusion that
> > >there is no safe way for kprobes to know which arguments passed on the
> > >stack should be preserved, at least not with the current API.

[...]

Jumping cheekily onto this thread, what if some function does this:

void go_on_jprobe_me()
{
}

void foo()
{
	struct bar baz;

	start_io(&baz);

	/* ... */

	go_on_jprobe_me();

	end_io(&baz);
}

If some I/O is being done on baz asynchronously, via DMA or via another
thread, a jprobe implementation that attempts to save/restore the stack
beyond the arguments of the probed function is going to race with such
I/O and can corrupt data.

This is a risk whenever any thread triggers some other master to operate
on objects on the first thread's stack -- I/O is a contrived example, but
there are likely other ways similar asynchronous access can happen to
a thread's stack.

Worse, annotating go_on_jprobe_me() as un-jprobeable doesn't help --
the un-jprobeableness is a property not of the function itself, but
rather a property of the set of callers of that function.  That set can
change at runtime (consider out-of-tree modules).

Cheers
---Dave

[toc] | [prev] | [next] | [standalone]


#1450726

FromMark Rutland <mark.rutland@arm.com>
Date2016-07-26 20:00 +0200
Message-ID<rZeO6-55G-11@gated-at.bofh.it>
In reply to#1450509
On Tue, Jul 26, 2016 at 10:50:08AM +0100, Daniel Thompson wrote:
> On 25/07/16 18:13, Catalin Marinas wrote:
> >You get more unexpected side effects by not saving/restoring the whole
> >stack. We looked into this on Friday and came to the conclusion that
> >there is no safe way for kprobes to know which arguments passed on the
> >stack should be preserved, at least not with the current API.
> >
> >Basically the AArch64 PCS states that for arguments passed on the stack
> >(e.g. they can't fit in registers), the caller allocates memory for them
> >(on its own stack) and passes the pointer to the callee. Unfortunately,
> >the frame pointer seems to be decremented correspondingly to cover the
> >arguments, so we don't really have a way to tell how much to copy.
> >Copying just the caller's stack frame isn't safe either since a
> >callee/caller receiving such argument on the stack may passed it down to
> >a callee without copying (I couldn't find anything in the PCS stating
> >that this isn't allowed).
> 
> The PCS[1] seems (at least to me) to be pretty clear that "the
> address of the first stacked argument is defined to be the initial
> value of SP".
> 
> I think it is only the return value (when stacked via the x8
> pointer) that can be passed through an intermediate function in the
> way described above. Isn't it OK for a jprobe to clobber this
> memory? The underlying function will overwrite whatever the jprobe
> put there anyway.
> 
> Am I overlooking some additional detail in the PCS?

I suspect that the "initial value of SP" is simply meant to be relative to the
base of the region of stack reserved for callee parameters. While it also uses
the phrase "current stack-pointer value", I suspect that this is overly
prescriptive.

In practice, GCC allocates callee parameters *above* the frame record
for the caller, which is above the SP and FP. e.g. with:

----
#define NLARGE 128

struct large {
	unsigned long v[NLARGE];
};

unsigned long __attribute__ ((noinline)) large_func(const struct large l)
{
	return l.v[0];
}

int main(int argc, char *argv[])
{
	struct large l = {
		.v = { 1, },
	};
	return large_func(l);
}
----

Which yields the following assembly:

----
00000000004005d0 <large_func>:
  4005d0:       f81f0ff3        str     x19, [sp,#-16]!
  4005d4:       aa0003f3        mov     x19, x0
  4005d8:       f9400260        ldr     x0, [x19]
  4005dc:       f84107f3        ldr     x19, [sp],#16
  4005e0:       d65f03c0        ret

00000000004005e4 <main>:
  4005e4:       d12043ff        sub     sp, sp, #0x810
  4005e8:       a9bf7bfd        stp     x29, x30, [sp,#-16]!
  4005ec:       910003fd        mov     x29, sp
  4005f0:       b9041fa0        str     w0, [x29,#1052]
  4005f4:       f9020ba1        str     x1, [x29,#1040]
  4005f8:       911083a0        add     x0, x29, #0x420
  4005fc:       d2808001        mov     x1, #0x400                      // #1024
  400600:       aa0103e2        mov     x2, x1
  400604:       52800001        mov     w1, #0x0                        // #0
  400608:       97ffff92        bl      400450 <memset@plt>
  40060c:       d2800020        mov     x0, #0x1                        // #1
  400610:       f90213a0        str     x0, [x29,#1056]
  400614:       910043a0        add     x0, x29, #0x10
  400618:       911083a1        add     x1, x29, #0x420
  40061c:       d2808002        mov     x2, #0x400                      // #1024
  400620:       97ffff84        bl      400430 <memcpy@plt>
  400624:       910043a0        add     x0, x29, #0x10
  400628:       97ffffea        bl      4005d0 <large_func>
  40062c:       a8c17bfd        ldp     x29, x30, [sp],#16
  400630:       912043ff        add     sp, sp, #0x810
  400634:       d65f03c0        ret
----

Please ignore the redundant copy GCC generates and copies; I can't seem
to convince it to not do that. The important part is that at 400614 the
argument to the function is the address immediately above the frame
record for main.

In local testing, it seems that additional locals can appear between the
frame record and argument.

Given this, callees can't rely on any relationship between their initial sp and
stacked arguments. Given that, I see no reason why an intermediary could not
simply pass the pointer on while creating further intermediary stack frames.

Thanks,
Mark.

[toc] | [prev] | [next] | [standalone]


#1451241

FromDaniel Thompson <daniel.thompson@linaro.org>
Date2016-07-27 13:30 +0200
Message-ID<rZvce-7cz-33@gated-at.bofh.it>
In reply to#1450726
On 26/07/16 18:54, Mark Rutland wrote:
> On Tue, Jul 26, 2016 at 10:50:08AM +0100, Daniel Thompson wrote:
>> On 25/07/16 18:13, Catalin Marinas wrote:
>>> You get more unexpected side effects by not saving/restoring the whole
>>> stack. We looked into this on Friday and came to the conclusion that
>>> there is no safe way for kprobes to know which arguments passed on the
>>> stack should be preserved, at least not with the current API.
>>>
>>> Basically the AArch64 PCS states that for arguments passed on the stack
>>> (e.g. they can't fit in registers), the caller allocates memory for them
>>> (on its own stack) and passes the pointer to the callee. Unfortunately,
>>> the frame pointer seems to be decremented correspondingly to cover the
>>> arguments, so we don't really have a way to tell how much to copy.
>>> Copying just the caller's stack frame isn't safe either since a
>>> callee/caller receiving such argument on the stack may passed it down to
>>> a callee without copying (I couldn't find anything in the PCS stating
>>> that this isn't allowed).
>>
>> The PCS[1] seems (at least to me) to be pretty clear that "the
>> address of the first stacked argument is defined to be the initial
>> value of SP".
>>
>> I think it is only the return value (when stacked via the x8
>> pointer) that can be passed through an intermediate function in the
>> way described above. Isn't it OK for a jprobe to clobber this
>> memory? The underlying function will overwrite whatever the jprobe
>> put there anyway.
>>
>> Am I overlooking some additional detail in the PCS?
>
> I suspect that the "initial value of SP" is simply meant to be relative to the
> base of the region of stack reserved for callee parameters. While it also uses
> the phrase "current stack-pointer value", I suspect that this is overly
> prescriptive.

I don't think so. Whilst writing my reply of yesterday I forced stacked 
arguments by creating a function with nine arguments (rather than large 
values). The ninth argument is, as expected, passed to the callee based 
on the value of the SP.


> In practice, GCC allocates callee parameters *above* the frame record
> for the caller, which is above the SP and FP. e.g. with:
>
> ----
 > <snip>
 > ----
> ----
> 00000000004005d0 <large_func>:
>   4005d0:       f81f0ff3        str     x19, [sp,#-16]!
>   4005d4:       aa0003f3        mov     x19, x0
>   4005d8:       f9400260        ldr     x0, [x19]
>   4005dc:       f84107f3        ldr     x19, [sp],#16
>   4005e0:       d65f03c0        ret
 >   ...
> ----

Thanks for the example.

The large structure is not a stacked argument from the point of view of 
the PCS parameter passing algorithm (which explicitly says how large 
composite types will be allocated). Instead it looks like it has been 
implicitly passed-by-reference and the caller makes this appear as 
call-by-value by allocating from its own stack frame rather than from 
the stacked argument space. The callee joins in by implicitly 
dereferencing the pointer.

It is interesting to note that you force large_func() to stack its 
arguments (by providing 8 dummy int arguments first) then the implicit 
pass-by-reference behavior is still preserved even for a stacked 
argument; large_func() ends up as:

~~~
large_func:
	ldr	x0, [sp]
	ldr	x0, [x0]
	ret
~~~

Only thing is... I *still* haven't found anything in the AArch64 PCS 
which describes this behavior.

I'm coming to believe that this is a mistake and this information (and 
the threshold at which implicit pass-by-reference kicks in) should be 
documented in section 7.

Or if you prefer the short version: I agree 100% with your analysis but 
cannot find the document that supports it.


Daniel.

[toc] | [prev] | [next] | [standalone]


#1451243

FromDave Martin <Dave.Martin@arm.com>
Date2016-07-27 13:40 +0200
Message-ID<rZvlU-7fU-135@gated-at.bofh.it>
In reply to#1451241
On Wed, Jul 27, 2016 at 12:19:59PM +0100, Daniel Thompson wrote:

[...]

> It is interesting to note that you force large_func() to stack its arguments
> (by providing 8 dummy int arguments first) then the implicit
> pass-by-reference behavior is still preserved even for a stacked argument;
> large_func() ends up as:
> 
> ~~~
> large_func:
> 	ldr	x0, [sp]
> 	ldr	x0, [x0]
> 	ret
> ~~~
> 
> Only thing is... I *still* haven't found anything in the AArch64 PCS which
> describes this behavior.
> 
> I'm coming to believe that this is a mistake and this information (and the
> threshold at which implicit pass-by-reference kicks in) should be documented
> in section 7.

Is that answered by this?

    B.3. If the argument type is a Composite Type that is larger than
    16 bytes, then the argument is copied to memory allocated by the
    caller and the argument is replaced by a pointer to the copy.

Experimenting with gcc's behaviour seems to back this up.

Cheers
---Dave

[toc] | [prev] | [next] | [standalone]


#1451247

FromDaniel Thompson <daniel.thompson@linaro.org>
Date2016-07-27 13:50 +0200
Message-ID<rZvvz-7ju-7@gated-at.bofh.it>
In reply to#1451243
On 27/07/16 12:38, Dave Martin wrote:
> On Wed, Jul 27, 2016 at 12:19:59PM +0100, Daniel Thompson wrote:
>
> [...]
>
>> It is interesting to note that you force large_func() to stack its arguments
>> (by providing 8 dummy int arguments first) then the implicit
>> pass-by-reference behavior is still preserved even for a stacked argument;
>> large_func() ends up as:
>>
>> ~~~
>> large_func:
>> 	ldr	x0, [sp]
>> 	ldr	x0, [x0]
>> 	ret
>> ~~~
>>
>> Only thing is... I *still* haven't found anything in the AArch64 PCS which
>> describes this behavior.
>>
>> I'm coming to believe that this is a mistake and this information (and the
>> threshold at which implicit pass-by-reference kicks in) should be documented
>> in section 7.
>
> Is that answered by this?
>
>     B.3. If the argument type is a Composite Type that is larger than
>     16 bytes, then the argument is copied to memory allocated by the
>     caller and the argument is replaced by a pointer to the copy.
>
> Experimenting with gcc's behaviour seems to back this up.

Absolutely answered by that. Thanks (and sorry for the noise)!


Daniel.

[toc] | [prev] | [next] | [standalone]


#1451287

FromMark Rutland <mark.rutland@arm.com>
Date2016-07-27 15:40 +0200
Message-ID<rZxe2-8qU-21@gated-at.bofh.it>
In reply to#1451241
On Wed, Jul 27, 2016 at 12:19:59PM +0100, Daniel Thompson wrote:
> On 26/07/16 18:54, Mark Rutland wrote:
> >On Tue, Jul 26, 2016 at 10:50:08AM +0100, Daniel Thompson wrote:
> >>On 25/07/16 18:13, Catalin Marinas wrote:
> >>>You get more unexpected side effects by not saving/restoring the whole
> >>>stack. We looked into this on Friday and came to the conclusion that
> >>>there is no safe way for kprobes to know which arguments passed on the
> >>>stack should be preserved, at least not with the current API.
> >>>
> >>>Basically the AArch64 PCS states that for arguments passed on the stack
> >>>(e.g. they can't fit in registers), the caller allocates memory for them
> >>>(on its own stack) and passes the pointer to the callee. Unfortunately,
> >>>the frame pointer seems to be decremented correspondingly to cover the
> >>>arguments, so we don't really have a way to tell how much to copy.
> >>>Copying just the caller's stack frame isn't safe either since a
> >>>callee/caller receiving such argument on the stack may passed it down to
> >>>a callee without copying (I couldn't find anything in the PCS stating
> >>>that this isn't allowed).
> >>
> >>The PCS[1] seems (at least to me) to be pretty clear that "the
> >>address of the first stacked argument is defined to be the initial
> >>value of SP".
> >>
> >>I think it is only the return value (when stacked via the x8
> >>pointer) that can be passed through an intermediate function in the
> >>way described above. Isn't it OK for a jprobe to clobber this
> >>memory? The underlying function will overwrite whatever the jprobe
> >>put there anyway.
> >>
> >>Am I overlooking some additional detail in the PCS?
> >
> >I suspect that the "initial value of SP" is simply meant to be relative to the
> >base of the region of stack reserved for callee parameters. While it also uses
> >the phrase "current stack-pointer value", I suspect that this is overly
> >prescriptive.
> 
> I don't think so. Whilst writing my reply of yesterday I forced
> stacked arguments by creating a function with nine arguments (rather
> than large values). The ninth argument is, as expected, passed to
> the callee based on the value of the SP.

Ah. I'd failed to fully appreciate the distinction between large
structures (which get converted to pointers), and basic argument types
(including those converted pointers).

For basic argument types, I think you're right, and my wording above is
wrong.

However, for (large enough) structures I don't think we have any
guarantee as to their location.

Sorry for the confusion there!

Mark.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web