Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1456076 > unrolled thread
| Started by | "Robert O'Callahan" <robert@ocallahan.org> |
|---|---|
| First post | 2016-08-04 02:00 +0200 |
| Last post | 2016-08-10 22:00 +0200 |
| Articles | 2 — 2 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: "run seccomp after ptrace" changes expose "missing PTRACE_EVENT_EXIT" bug "Robert O'Callahan" <robert@ocallahan.org> - 2016-08-04 02:00 +0200
[PATCH] seccomp: suppress fatal signals that will never be delivered before seccomp forces an exit because of said signals Kyle Huey <me@kylehuey.com> - 2016-08-10 22:00 +0200
| From | "Robert O'Callahan" <robert@ocallahan.org> |
|---|---|
| Date | 2016-08-04 02:00 +0200 |
| Subject | Re: "run seccomp after ptrace" changes expose "missing PTRACE_EVENT_EXIT" bug |
| Message-ID | <s2eeS-XC-9@gated-at.bofh.it> |
I work on rr (http://rr-project.org/), a record-and-replay reverse-execution debugger which is a heavy user of ptrace and seccomp. The recent change to perform syscall-entry PTRACE_SYSCALL stops before PTRACE_EVENT_SECCOMP stops broke rr, which is fine because I'm fixing rr and this change actually makes rr faster (thanks!). However, it exposed an existing kernel bug which creates a problem for us, and which I'm not sure how to fix. The problem is that if a tracee task is in a PTRACE_EVENT_SECCOMP trap, or has been resumed after such a trap but not yet been scheduled, and another task in the thread-group calls exit_group(), then the tracee task exits without the ptracer receiving a PTRACE_EVENT_EXIT notification. Small-ish testcase here: https://gist.github.com/rocallahan/1344f7d01183c233d08a2c6b93413068. The bug happens because when __seccomp_filter() detects fatal_signal_pending(), it calls do_exit() without dequeuing the fatal signal. When do_exit() sends the PTRACE_EVENT_EXIT notification and that task is descheduled, __schedule() notices that there is a fatal signal pending and changes its state from TASK_TRACED to TASK_RUNNING. That prevents the ptracer's waitpid() from returning the ptrace event. A more detailed analysis is here: https://github.com/mozilla/rr/issues/1762#issuecomment-237396255. This bug has been in the kernel for a while. rr never hit it before because we trace all threads and mostly run only one tracee thread at a time. Immediately after each PTRACE_EVENT_SECCOMP notification we'd issue a PTRACE_SYSCALL to get that task to the syscall-entry PTRACE_SYSCALL stop, so there was never an opportunity for one tracee thread to call exit_group while another tracee was in the problematic part of __seccomp_filter(). Unfortunately now there is no way for us to avoid that possibility. My guess is that __seccomp_filter() should dequeue the fatal signal it detects before calling do_exit(), to behave more like get_signal(). Is that correct, and if so, what would be the right way to do that? Thanks, Robert O'Callahan -- lbir ye,ea yer.tnietoehr rdn rdsme,anea lurpr edna e hnysnenh hhe uresyf toD selthor stor edna siewaoeodm or v sstvr esBa kbvted,t rdsme,aoreseoouoto o l euetiuruewFa kbn e hnystoivateweh uresyf tulsa rehr rdm or rnea lurpr .a war hsrer holsa rodvted,t nenh hneireseoouot.tniesiewaoeivatewt sstvr esn
[toc] | [next] | [standalone]
| From | Kyle Huey <me@kylehuey.com> |
|---|---|
| Date | 2016-08-10 22:00 +0200 |
| Subject | [PATCH] seccomp: suppress fatal signals that will never be delivered before seccomp forces an exit because of said signals |
| Message-ID | <s4HPt-16w-47@gated-at.bofh.it> |
| In reply to | #1456076 |
This fixes rr. It doesn't quite fix the provided testcase, because the testcase fails to wait on the tracee after awakening from the nanosleep. Instead the testcase immediately does a PTHREAD_CONT, discarding the PTHREAD_EVENT_EXIT. The slightly modified testcase at https://gist.github.com/khuey/3c43ac247c72cef8c956c does pass.
I don't see any obvious way to dequeue only the fatal signal, so instead I dequeue them all. Since none of these signals will ever be delivered it shouldn't affect the executing task.
Suggested-by: Robert O'Callahan <robert@ocallahan.org>
Signed-off-by: Kyle Huey <khuey@kylehuey.com>
---
kernel/seccomp.c | 14 +++++++++++++-
1 file changed, 13 insertions(+), 1 deletion(-)
diff --git a/kernel/seccomp.c b/kernel/seccomp.c
index ef6c6c3..728074d 100644
--- a/kernel/seccomp.c
+++ b/kernel/seccomp.c
@@ -609,8 +609,20 @@ static int __seccomp_filter(int this_syscall, const struct seccomp_data *sd,
* Terminating the task now avoids executing a system
* call that may not be intended.
*/
- if (fatal_signal_pending(current))
+ if (fatal_signal_pending(current)) {
+ /*
+ * Swallow the signals we will never deliver.
+ * If we do not do this, the PTRACE_EVENT_EXIT will
+ * be suppressed by those signals.
+ */
+ siginfo_t info;
+
+ spin_lock_irq(¤t->sighand->siglock);
+ while (dequeue_signal(current, ¤t->blocked, &info));
+ spin_unlock_irq(¤t->sighand->siglock);
+
do_exit(SIGSYS);
+ }
/* Check if the tracer forced the syscall to be skipped. */
this_syscall = syscall_get_nr(current, task_pt_regs(current));
if (this_syscall < 0)
--
2.7.4
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web