Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1456076 > unrolled thread

Re: "run seccomp after ptrace" changes expose "missing PTRACE_EVENT_EXIT" bug

Started by"Robert O'Callahan" <robert@ocallahan.org>
First post2016-08-04 02:00 +0200
Last post2016-08-10 22:00 +0200
Articles 2 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: "run seccomp after ptrace" changes expose "missing  PTRACE_EVENT_EXIT" bug "Robert O'Callahan" <robert@ocallahan.org> - 2016-08-04 02:00 +0200
    [PATCH] seccomp: suppress fatal signals that will never be delivered before seccomp forces an exit because of said signals Kyle Huey <me@kylehuey.com> - 2016-08-10 22:00 +0200

#1456076 — Re: "run seccomp after ptrace" changes expose "missing PTRACE_EVENT_EXIT" bug

From"Robert O'Callahan" <robert@ocallahan.org>
Date2016-08-04 02:00 +0200
SubjectRe: "run seccomp after ptrace" changes expose "missing PTRACE_EVENT_EXIT" bug
Message-ID<s2eeS-XC-9@gated-at.bofh.it>
I work on rr (http://rr-project.org/), a record-and-replay
reverse-execution debugger which is a heavy user of ptrace and
seccomp. The recent change to perform syscall-entry PTRACE_SYSCALL
stops before PTRACE_EVENT_SECCOMP stops broke rr, which is fine
because I'm fixing rr and this change actually makes rr faster
(thanks!). However, it exposed an existing kernel bug which creates a
problem for us, and which I'm not sure how to fix.

The problem is that if a tracee task is in a PTRACE_EVENT_SECCOMP
trap, or has been resumed after such a trap but not yet been
scheduled, and another task in the thread-group calls exit_group(),
then the tracee task exits without the ptracer receiving a
PTRACE_EVENT_EXIT notification. Small-ish testcase here:
https://gist.github.com/rocallahan/1344f7d01183c233d08a2c6b93413068.

The bug happens because when __seccomp_filter() detects
fatal_signal_pending(), it calls do_exit() without dequeuing the fatal
signal. When do_exit() sends the PTRACE_EVENT_EXIT notification and
that task is descheduled, __schedule() notices that there is a fatal
signal pending and changes its state from TASK_TRACED to TASK_RUNNING.
That prevents the ptracer's waitpid() from returning the ptrace event.
A more detailed analysis is here:
https://github.com/mozilla/rr/issues/1762#issuecomment-237396255.

This bug has been in the kernel for a while. rr never hit it before
because we trace all threads and mostly run only one tracee thread at
a time. Immediately after each PTRACE_EVENT_SECCOMP notification we'd
issue a PTRACE_SYSCALL to get that task to the syscall-entry
PTRACE_SYSCALL stop, so there was never an opportunity for one tracee
thread to call exit_group while another tracee was in the problematic
part of __seccomp_filter(). Unfortunately now there is no way for us
to avoid that possibility.

My guess is that __seccomp_filter() should dequeue the fatal signal it
detects before calling do_exit(), to behave more like get_signal(). Is
that correct, and if so, what would be the right way to do that?

Thanks,
Robert O'Callahan
-- 
lbir ye,ea yer.tnietoehr  rdn rdsme,anea lurpr  edna e hnysnenh hhe uresyf toD
selthor  stor  edna  siewaoeodm  or v sstvr  esBa  kbvted,t rdsme,aoreseoouoto
o l euetiuruewFa  kbn e hnystoivateweh uresyf tulsa rehr  rdm  or rnea lurpr
.a war hsrer holsa rodvted,t  nenh hneireseoouot.tniesiewaoeivatewt sstvr  esn

[toc] | [next] | [standalone]


#1459758 — [PATCH] seccomp: suppress fatal signals that will never be delivered before seccomp forces an exit because of said signals

FromKyle Huey <me@kylehuey.com>
Date2016-08-10 22:00 +0200
Subject[PATCH] seccomp: suppress fatal signals that will never be delivered before seccomp forces an exit because of said signals
Message-ID<s4HPt-16w-47@gated-at.bofh.it>
In reply to#1456076
This fixes rr. It doesn't quite fix the provided testcase, because the testcase fails to wait on the tracee after awakening from the nanosleep. Instead the testcase immediately does a PTHREAD_CONT, discarding the PTHREAD_EVENT_EXIT. The slightly modified testcase at https://gist.github.com/khuey/3c43ac247c72cef8c956c does pass.

I don't see any obvious way to dequeue only the fatal signal, so instead I dequeue them all. Since none of these signals will ever be delivered it shouldn't affect the executing task.

Suggested-by: Robert O'Callahan <robert@ocallahan.org>
Signed-off-by: Kyle Huey <khuey@kylehuey.com>
---
 kernel/seccomp.c | 14 +++++++++++++-
 1 file changed, 13 insertions(+), 1 deletion(-)

diff --git a/kernel/seccomp.c b/kernel/seccomp.c
index ef6c6c3..728074d 100644
--- a/kernel/seccomp.c
+++ b/kernel/seccomp.c
@@ -609,8 +609,20 @@ static int __seccomp_filter(int this_syscall, const struct seccomp_data *sd,
 		 * Terminating the task now avoids executing a system
 		 * call that may not be intended.
 		 */
-		if (fatal_signal_pending(current))
+		if (fatal_signal_pending(current)) {
+			/*
+			 * Swallow the signals we will never deliver.
+			 * If we do not do this, the PTRACE_EVENT_EXIT will
+			 * be suppressed by those signals.
+			 */
+			siginfo_t info;
+
+			spin_lock_irq(&current->sighand->siglock);
+			while (dequeue_signal(current, &current->blocked, &info));
+			spin_unlock_irq(&current->sighand->siglock);
+
 			do_exit(SIGSYS);
+		}
 		/* Check if the tracer forced the syscall to be skipped. */
 		this_syscall = syscall_get_nr(current, task_pt_regs(current));
 		if (this_syscall < 0)
-- 
2.7.4

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web