Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1644389 > unrolled thread

[PATCH 2/3] livepatch: send a fake signal to all blocking tasks

Started byMiroslav Benes <mbenes@suse.cz>
First post2017-05-18 14:10 +0200
Last post2017-05-24 10:40 +0200
Articles 9 — 4 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  [PATCH 2/3] livepatch: send a fake signal to all blocking tasks Miroslav Benes <mbenes@suse.cz> - 2017-05-18 14:10 +0200
    Re: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks Libor Pechacek <lpechacek@suse.com> - 2017-05-18 15:20 +0200
      Re: [PATCH 2/3] livepatch: send a fake signal to all blocking  tasks Miroslav Benes <mbenes@suse.cz> - 2017-05-18 15:30 +0200
    Re: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks Oleg Nesterov <oleg@redhat.com> - 2017-05-18 18:50 +0200
      Re: [PATCH 2/3] livepatch: send a fake signal to all blocking  tasks Miroslav Benes <mbenes@suse.cz> - 2017-05-18 20:20 +0200
        Re: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks Oleg Nesterov <oleg@redhat.com> - 2017-05-18 22:00 +0200
          Re: [PATCH 2/3] livepatch: send a fake signal to all blocking  tasks Miroslav Benes <mbenes@suse.cz> - 2017-05-19 10:00 +0200
    Re: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks Josh Poimboeuf <jpoimboe@redhat.com> - 2017-05-23 19:40 +0200
      Re: [PATCH 2/3] livepatch: send a fake signal to all blocking  tasks Miroslav Benes <mbenes@suse.cz> - 2017-05-24 10:40 +0200

#1644389 — [PATCH 2/3] livepatch: send a fake signal to all blocking tasks

FromMiroslav Benes <mbenes@suse.cz>
Date2017-05-18 14:10 +0200
Subject[PATCH 2/3] livepatch: send a fake signal to all blocking tasks
Message-ID<tIspI-2kb-15@gated-at.bofh.it>
Live patching consistency model is of LEAVE_PATCHED_SET and
SWITCH_THREAD. This means that all tasks in the system have to be marked
one by one as safe to call a new patched function. Safe means when a
task is not (sleeping) in a set of patched functions. That is, no
patched function is on the task's stack. Another clearly safe place is
the boundary between kernel and userspace. The patching waits for all
tasks to get outside of the patched set or to cross the boundary. The
transition is completed afterwards.

The problem is that a task can block the transition for quite a long
time, if not forever. It could sleep in a set of patched functions, for
example.  Luckily we can force the task to leave the set by sending it a
fake signal, that is a signal with no data in signal pending structures
(no handler, no sign of proper signal delivered). Suspend/freezer use
this to freeze the tasks as well. The task gets TIF_SIGPENDING set and
is woken up (if it has been sleeping in the kernel before) or kicked by
rescheduling IPI (if it was running on other CPU). This causes the task
to go to kernel/userspace boundary where the signal would be handled and
the task would be marked as safe in terms of live patching.

There are tasks which are not affected by this technique though. The
fake signal is not sent to kthreads. They should be handled in a
different way. They can be woken up so they leave the patched set and
their TIF_PATCH_PENDING can be cleared thanks to stack checking.

For the sake of completeness, if the task is in TASK_RUNNING state but
not currently running on some CPU it doesn't get the IPI, but it would
eventually handle the signal anyway. Second, if the task runs in the
kernel (in TASK_RUNNING state) it gets the IPI, but the signal is not
handled on return from the interrupt. It would be handled on return to
the userspace in the future when the fake signal is sent again. Stack
checking deals with these cases in a better way.

If the task was sleeping in a syscall it would be woken by our fake
signal, it would check if TIF_SIGPENDING is set (by calling
signal_pending() predicate) and return ERESTART* or EINTR. Syscalls with
ERESTART* return values are restarted in case of the fake signal (see
do_signal()). EINTR is propagated back to the userspace program. This
could disturb the program, but...

* each process dealing with signals should react accordingly to EINTR
  return values.
* syscalls returning EINTR happen to be quite common situation in the
  system even if no fake signal is sent.
* freezer sends the fake signal and does not deal with EINTR anyhow.
  Thus EINTR values are returned when the system is resumed.

The very safe marking is done in entry.S on syscall and
interrupt/exception exit paths, and in a stack checking functions of
livepatch.  TIF_PATCH_PENDING is cleared and the next
recalc_sigpending() drops TIF_SIGPENDING.

Note that the fake signal is not sent to stopped/traced tasks. Such task
prevents the patching to finish till it continues again (is not traced
anymore).

Last, sending the fake signal is not automatic. It is done only when
admin requests it by writing 1 to force sysfs attribute in livepatch
sysfs directory.

Cc: Oleg Nesterov <oleg@redhat.com>
Signed-off-by: Miroslav Benes <mbenes@suse.cz>
---
 include/linux/livepatch.h     |  3 +++
 kernel/livepatch/core.c       |  3 +++
 kernel/livepatch/transition.c | 40 ++++++++++++++++++++++++++++++++++++++++
 kernel/livepatch/transition.h |  1 +
 kernel/signal.c               |  4 +++-
 5 files changed, 50 insertions(+), 1 deletion(-)

diff --git a/include/linux/livepatch.h b/include/linux/livepatch.h
index 194991ef9347..43cfeebeb42b 100644
--- a/include/linux/livepatch.h
+++ b/include/linux/livepatch.h
@@ -29,6 +29,9 @@
 
 #include <asm/livepatch.h>
 
+/* values for sysfs force attribute */
+#define KLP_FORCE_FAKE		1
+
 /* task patch states */
 #define KLP_UNDEFINED	-1
 #define KLP_UNPATCHED	 0
diff --git a/kernel/livepatch/core.c b/kernel/livepatch/core.c
index 84f8944704ad..bb3b78fa7d2b 100644
--- a/kernel/livepatch/core.c
+++ b/kernel/livepatch/core.c
@@ -466,6 +466,9 @@ static ssize_t force_store(struct kobject *kobj, struct kobj_attribute *attr,
 	}
 
 	switch (val) {
+	case KLP_FORCE_FAKE:
+		klp_send_fake_signal();
+		break;
 	default:
 		return -EINVAL;
 	}
diff --git a/kernel/livepatch/transition.c b/kernel/livepatch/transition.c
index adc0cc64aa4b..bb61aaa196d3 100644
--- a/kernel/livepatch/transition.c
+++ b/kernel/livepatch/transition.c
@@ -551,3 +551,43 @@ void klp_copy_process(struct task_struct *child)
 
 	/* TIF_PATCH_PENDING gets copied in setup_thread_stack() */
 }
+
+/*
+ * Sends a fake signal to all non-kthread tasks with TIF_PATCH_PENDING set.
+ * Kthreads with TIF_PATCH_PENDING set are woken up. Only admin can request this
+ * action currently.
+ */
+void klp_send_fake_signal(void)
+{
+	struct task_struct *g, *task;
+
+	pr_info("sending a fake signal and waking sleeping kthreads up\n");
+
+	read_lock(&tasklist_lock);
+	for_each_process_thread(g, task) {
+		if (!klp_patch_pending(task))
+			continue;
+
+		/*
+		 * There is a small race here. We could see TIF_PATCH_PENDING
+		 * set and decide to wake up a kthread or send a fake signal.
+		 * Meanwhile the task could migrate itself and the action
+		 * would be meaningless. It is not serious though.
+		 */
+		if (task->flags & PF_KTHREAD) {
+			/*
+			 * Wake up a kthread which still has not been migrated.
+			 */
+			wake_up_process(task);
+		} else {
+			/*
+			 * Send fake signal to all non-kthread tasks which are
+			 * still not migrated.
+			 */
+			spin_lock_irq(&task->sighand->siglock);
+			signal_wake_up(task, 0);
+			spin_unlock_irq(&task->sighand->siglock);
+		}
+	}
+	read_unlock(&tasklist_lock);
+}
diff --git a/kernel/livepatch/transition.h b/kernel/livepatch/transition.h
index ce09b326546c..1c7ede6eaa77 100644
--- a/kernel/livepatch/transition.h
+++ b/kernel/livepatch/transition.h
@@ -10,5 +10,6 @@ void klp_cancel_transition(void);
 void klp_start_transition(void);
 void klp_try_complete_transition(void);
 void klp_reverse_transition(void);
+void klp_send_fake_signal(void);
 
 #endif /* _LIVEPATCH_TRANSITION_H */
diff --git a/kernel/signal.c b/kernel/signal.c
index 7e59ebc2c25e..3a25cc06231d 100644
--- a/kernel/signal.c
+++ b/kernel/signal.c
@@ -39,6 +39,7 @@
 #include <linux/compat.h>
 #include <linux/cn_proc.h>
 #include <linux/compiler.h>
+#include <linux/livepatch.h>
 
 #define CREATE_TRACE_POINTS
 #include <trace/events/signal.h>
@@ -162,7 +163,8 @@ void recalc_sigpending_and_wake(struct task_struct *t)
 
 void recalc_sigpending(void)
 {
-	if (!recalc_sigpending_tsk(current) && !freezing(current))
+	if (!recalc_sigpending_tsk(current) && !freezing(current) &&
+	    !klp_patch_pending(current))
 		clear_thread_flag(TIF_SIGPENDING);
 
 }
-- 
2.12.2

[toc] | [next] | [standalone]


#1644465

FromLibor Pechacek <lpechacek@suse.com>
Date2017-05-18 15:20 +0200
Message-ID<tItvr-33n-1@gated-at.bofh.it>
In reply to#1644389
On Thu 18-05-17 14:00:42, Miroslav Benes wrote:
[...]
> --- a/include/linux/livepatch.h
> +++ b/include/linux/livepatch.h
> @@ -29,6 +29,9 @@
>  
>  #include <asm/livepatch.h>
>  
> +/* values for sysfs force attribute */
> +#define KLP_FORCE_FAKE		1
> +

Should be documented in Documentation/ABI/testing/sysfs-kernel-livepatch for
end users.

Libor
-- 
Libor Pechacek
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1644526 — Re: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks

FromMiroslav Benes <mbenes@suse.cz>
Date2017-05-18 15:30 +0200
SubjectRe: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks
Message-ID<tItF9-37o-63@gated-at.bofh.it>
In reply to#1644465
On Thu, 18 May 2017, Libor Pechacek wrote:

> On Thu 18-05-17 14:00:42, Miroslav Benes wrote:
> [...]
> > --- a/include/linux/livepatch.h
> > +++ b/include/linux/livepatch.h
> > @@ -29,6 +29,9 @@
> >  
> >  #include <asm/livepatch.h>
> >  
> > +/* values for sysfs force attribute */
> > +#define KLP_FORCE_FAKE		1
> > +
> 
> Should be documented in Documentation/ABI/testing/sysfs-kernel-livepatch for
> end users.

Ugh, I forgot. Thanks for reminding me.

Miroslav

[toc] | [prev] | [next] | [standalone]


#1644745

FromOleg Nesterov <oleg@redhat.com>
Date2017-05-18 18:50 +0200
Message-ID<tIwMG-5iy-15@gated-at.bofh.it>
In reply to#1644389
I didn't see other patches in series, not sure I understand...

On 05/18, Miroslav Benes wrote:
>
> The very safe marking is done in entry.S on syscall and
> interrupt/exception exit paths, and in a stack checking functions of
> livepatch.  TIF_PATCH_PENDING is cleared and the next
> recalc_sigpending() drops TIF_SIGPENDING.

Confused. The task can't return from do_signal() is signal_pending() is
true, thus it will spin forever if klp_patch_pending(current)) is true.
"forever" means until something else clears TIF_PATCH_PENDING, of course.

exit_to_usermode_loop() calls do_signal(), then klp_update_patch_state().
So it won't be cleared here.

Even if you change the order, this won't help unless I missed something,
TIF_PATCH_PENDING can be set when this task has already entered do_signal().

> Last, sending the fake signal is not automatic. It is done only when
> admin requests it by writing 1 to force sysfs attribute in livepatch
> sysfs directory.

OK, but see above, even if klp_send_fake_signal() is never called, the
a task will get this fake signal when it calls recalc_sigpending().

Oleg.

[toc] | [prev] | [next] | [standalone]


#1644818 — Re: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks

FromMiroslav Benes <mbenes@suse.cz>
Date2017-05-18 20:20 +0200
SubjectRe: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks
Message-ID<tIybM-6nh-1@gated-at.bofh.it>
In reply to#1644745
On Thu, 18 May 2017, Oleg Nesterov wrote:

> I didn't see other patches in series, not sure I understand...

There is nothing relevant to this patch, I think. I did not want to bother
you with it.

> On 05/18, Miroslav Benes wrote:
> >
> > The very safe marking is done in entry.S on syscall and
> > interrupt/exception exit paths, and in a stack checking functions of
> > livepatch.  TIF_PATCH_PENDING is cleared and the next
> > recalc_sigpending() drops TIF_SIGPENDING.
> 
> Confused. The task can't return from do_signal() is signal_pending() is
> true, thus it will spin forever if klp_patch_pending(current)) is true.
> "forever" means until something else clears TIF_PATCH_PENDING, of course.
>
> exit_to_usermode_loop() calls do_signal(), then klp_update_patch_state().
> So it won't be cleared here.

Ok, so maybe I misunderstand the code. I see the loop in
exit_to_usermode_loop() for processing ALLWORK_MASK. There we call
do_signal(). We go to get_signal(). The infinite loop there is relevant
for us. We call dequeue_signal(). There, if I am not mistaken
__dequeue_signal() would return 0 in our case, because there is no real
signal pending and thus nothing in the signal data structures.
recalc_sigpending() is called and TIF_SIGPENDING is preserved there (I
presume TIF_PATCH_PENDING is set). signr is zero, dequeue_signal() returns
0. Back in get_signal() the loop is broken and zero is return. Then
do_signal() may or may not restart the syscall.
  
If not, we get back to exit_to_usermode_loop() and TIF_PATCH_PENDING is 
cleared. Yes, it is true that TIF_SIGPENDING is still set and we get to
do_signal() once more. But for the last time.

If the syscall is restarted, it may be different. I have to think about
this one. But...

> Even if you change the order, this won't help unless I missed something,
> TIF_PATCH_PENDING can be set when this task has already entered do_signal().

...I think it could be solved with this anyway. And of course it should 
solve the double call to do_signal() I described above.

Damn, I fixed exactly this in SLES a year or so ago and there is a note I 
did the same in proposed version for upstream. It must have fallen through
the cracks.


So, am I wrong somewhere? It could be anywhere, because it is quite 
confusing.

Regards,
Miroslav

[toc] | [prev] | [next] | [standalone]


#1644883

FromOleg Nesterov <oleg@redhat.com>
Date2017-05-18 22:00 +0200
Message-ID<tIzKy-7VB-17@gated-at.bofh.it>
In reply to#1644818
On 05/18, Miroslav Benes wrote:
>
> On Thu, 18 May 2017, Oleg Nesterov wrote:
>
> >
> > exit_to_usermode_loop() calls do_signal(), then klp_update_patch_state().
> > So it won't be cleared here.
>
> Ok, so maybe I misunderstand the code. I see the loop in
> exit_to_usermode_loop() for processing ALLWORK_MASK. There we call
> do_signal(). We go to get_signal(). The infinite loop there is relevant
> for us. We call dequeue_signal(). There, if I am not mistaken
> __dequeue_signal() would return 0

Yes, sorry, I didn't bother to read the code when I looked at your patch
and my memory fooled me.

> If not, we get back to exit_to_usermode_loop() and TIF_PATCH_PENDING is 
> cleared. Yes, it is true that TIF_SIGPENDING is still set and we get to
> do_signal() once more. But for the last time.

Yes, slightly sub-optimal but not really wrong and you can swap
do_signal() and klp_update_patch_state().

> If the syscall is restarted, it may be different. I have to think about
> this one. But...

Afaics, there are no problems.


In short. Thanks for correcting me and sorry for noise!

Oleg.

[toc] | [prev] | [next] | [standalone]


#1645305 — Re: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks

FromMiroslav Benes <mbenes@suse.cz>
Date2017-05-19 10:00 +0200
SubjectRe: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks
Message-ID<tIKZl-7rF-21@gated-at.bofh.it>
In reply to#1644883
> > If not, we get back to exit_to_usermode_loop() and TIF_PATCH_PENDING is 
> > cleared. Yes, it is true that TIF_SIGPENDING is still set and we get to
> > do_signal() once more. But for the last time.
> 
> Yes, slightly sub-optimal but not really wrong and you can swap
> do_signal() and klp_update_patch_state().

Ok. I'll add it to v2.
 
> > If the syscall is restarted, it may be different. I have to think about
> > this one. But...
> 
> Afaics, there are no problems.
> 
> 
> In short. Thanks for correcting me and sorry for noise!

Thanks for the review, Oleg. Much appreciated.

Miroslav

[toc] | [prev] | [next] | [standalone]


#1648296

FromJosh Poimboeuf <jpoimboe@redhat.com>
Date2017-05-23 19:40 +0200
Message-ID<tKlWO-6kH-17@gated-at.bofh.it>
In reply to#1644389
On Thu, May 18, 2017 at 02:00:42PM +0200, Miroslav Benes wrote:
> @@ -551,3 +551,43 @@ void klp_copy_process(struct task_struct *child)
>  
>  	/* TIF_PATCH_PENDING gets copied in setup_thread_stack() */
>  }
> +
> +/*
> + * Sends a fake signal to all non-kthread tasks with TIF_PATCH_PENDING set.
> + * Kthreads with TIF_PATCH_PENDING set are woken up. Only admin can request this
> + * action currently.
> + */
> +void klp_send_fake_signal(void)
> +{
> +	struct task_struct *g, *task;
> +
> +	pr_info("sending a fake signal and waking sleeping kthreads up\n");

Maybe this should be pr_notice(), for consistency with our other
printks.

Also I wonder if the message can be made more meaningful to the user.
The "fake" part of the signal and the "waking sleeping kthreads" bit
could be too much information for the user, IMO.  How about "signaling
remaining tasks"?  Just an idea.

-- 
Josh

[toc] | [prev] | [next] | [standalone]


#1649282 — Re: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks

FromMiroslav Benes <mbenes@suse.cz>
Date2017-05-24 10:40 +0200
SubjectRe: [PATCH 2/3] livepatch: send a fake signal to all blocking tasks
Message-ID<tKzZN-8mT-31@gated-at.bofh.it>
In reply to#1648296
On Tue, 23 May 2017, Josh Poimboeuf wrote:

> On Thu, May 18, 2017 at 02:00:42PM +0200, Miroslav Benes wrote:
> > @@ -551,3 +551,43 @@ void klp_copy_process(struct task_struct *child)
> >  
> >  	/* TIF_PATCH_PENDING gets copied in setup_thread_stack() */
> >  }
> > +
> > +/*
> > + * Sends a fake signal to all non-kthread tasks with TIF_PATCH_PENDING set.
> > + * Kthreads with TIF_PATCH_PENDING set are woken up. Only admin can request this
> > + * action currently.
> > + */
> > +void klp_send_fake_signal(void)
> > +{
> > +	struct task_struct *g, *task;
> > +
> > +	pr_info("sending a fake signal and waking sleeping kthreads up\n");
> 
> Maybe this should be pr_notice(), for consistency with our other
> printks.
> 
> Also I wonder if the message can be made more meaningful to the user.
> The "fake" part of the signal and the "waking sleeping kthreads" bit
> could be too much information for the user, IMO.  How about "signaling
> remaining tasks"?  Just an idea.

Good one. I'll change it in v2.

Thanks,
Miroslav

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web