Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1550104 > unrolled thread

Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE

Started byDaniel Bristot de Oliveira <bristot@redhat.com>
First post2017-01-03 20:10 +0100
Last post2017-01-11 22:20 +0100
Articles 11 — 6 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE Daniel Bristot de Oliveira <bristot@redhat.com> - 2017-01-03 20:10 +0100
    Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE luca abeni <luca.abeni@unitn.it> - 2017-01-03 22:40 +0100
    Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE luca abeni <luca.abeni@unitn.it> - 2017-01-04 13:20 +0100
      Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE Daniel Bristot de Oliveira <bristot@redhat.com> - 2017-01-04 16:20 +0100
        Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE Luca Abeni <luca.abeni@unitn.it> - 2017-01-04 17:50 +0100
          Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE Daniel Bristot de Oliveira <bristot@redhat.com> - 2017-01-04 19:10 +0100
            Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE Luca Abeni <luca.abeni@unitn.it> - 2017-01-04 19:50 +0100
              Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE Juri Lelli <juri.lelli@arm.com> - 2017-01-11 13:20 +0100
                Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE Luca Abeni <luca.abeni@santannapisa.it> - 2017-01-11 14:00 +0100
                  Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE Juri Lelli <juri.lelli@arm.com> - 2017-01-11 16:10 +0100
                    Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE luca abeni <luca.abeni@santannapisa.it> - 2017-01-11 22:20 +0100

#1550104 — Re: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE

FromDaniel Bristot de Oliveira <bristot@redhat.com>
Date2017-01-03 20:10 +0100
SubjectRe: [RFC v4 0/6] CPU reclaiming for SCHED_DEADLINE
Message-ID<sVCD8-2iN-11@gated-at.bofh.it>
On 12/30/2016 12:33 PM, Luca Abeni wrote:
> From: Luca Abeni <luca.abeni@unitn.it>
> 
> Hi all,
> 
> here is a new version of the patchset implementing CPU reclaiming
> (using the GRUB algorithm[1]) for SCHED_DEADLINE.
> Basically, this feature allows SCHED_DEADLINE tasks to consume more
> than their reserved runtime, up to a maximum fraction of the CPU time
> (so that other tasks are left some spare CPU time to execute), if this
> does not break the guarantees of other SCHED_DEADLINE tasks.
> The patchset applies on top of tip/master.
> 
> 
> The implemented CPU reclaiming algorithm is based on tracking the
> utilization U_act of active tasks (first 2 patches), and modifying the
> runtime accounting rule (see patch 0004). The original GRUB algorithm is
> modified as described in [2] to support multiple CPUs (the original
> algorithm only considered one single CPU, this one tracks U_act per
> runqueue) and to leave an "unreclaimable" fraction of CPU time to non
> SCHED_DEADLINE tasks (see patch 0005: the original algorithm can consume
> 100% of the CPU time, starving all the other tasks).
> Patch 0003 uses the newly introduced "inactive timer" (introduced in
> patch 0002) to fix dl_overflow() and __setparam_dl().
> Patch 0006 allows to enable CPU reclaiming only on selected tasks.

Hi,

Today I did some tests in this patch set. Unfortunately, it seems that
there is a problem :-(.

In a four core box, if I dispatch 11 tasks [1] with setup:

  period = 30 ms
  runtime = 10 ms
  flags = 0 (GRUB disabled)

I see this:
------------------------------- HTOP ------------------------------------
  1  [|||||||||||||||||||||92.5%]   Tasks: 128, 259 thr; 14 running
  2  [|||||||||||||||||||||91.0%]   Load average: 4.65 4.66 4.81 
  3  [|||||||||||||||||||||92.5%]   Uptime: 05:12:43
  4  [|||||||||||||||||||||92.5%]
  Mem[|||||||||||||||1.13G/3.78G]
  Swp[                  0K/3.90G]

  PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
16247 root      -101   0  4204   632   564 R 32.4  0.0  2:10.35 d
16249 root	-101   0  4204   624   556 R 32.4  0.0  2:09.80 d
16250 root	-101   0  4204   728   660 R 32.4  0.0  2:09.58 d
16252 root	-101   0  4204   676   608 R 32.4  0.0  2:09.08 d
16253 root	-101   0  4204   636   568 R 32.4  0.0  2:08.85 d
16254 root      -101   0  4204   732   664 R 32.4  0.0  2:08.62 d
16255 root	-101   0  4204   620   556 R 32.4  0.0  2:08.40 d
16257 root	-101   0  4204   708   640 R 32.4  0.0  2:07.98 d
16256 root	-101   0  4204   624   560 R 32.4  0.0  2:08.18 d
16248 root	-101   0  4204   680   612 R 33.0  0.0  2:10.15 d
16251 root	-101   0  4204   676   608 R 33.0  0.0  2:09.34 d
16259 root       20   0  124M  4692  3120 R  1.1  0.1  0:02.82 htop
 2191 bristot    20   0  649M 41312 32048 S  0.0  1.0  0:28.77 gnome-ter
------------------------------- HTOP ------------------------------------

All tasks are using +- the same amount of CPU time, a little bit more
than 30%, as expected. However, if I enable GRUB in the same task set
I get this:

------------------------------- HTOP ------------------------------------
  1  [|||||||||||||||||||||93.8%]   Tasks: 128, 260 thr; 15 running
  2  [|||||||||||||||||||||95.2%]   Load average: 5.13 5.01 4.98 
  3  [|||||||||||||||||||||93.3%]   Uptime: 05:01:02
  4  [|||||||||||||||||||||96.4%]
  Mem[|||||||||||||||1.13G/3.78G]
  Swp[                  0K/3.90G]

  PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
14967 root      -101   0  4204   628   564 R 45.8  0.0  1h07:49 g
14962 root	-101   0  4204   728   660 R 45.8  0.0  1h05:06 g
14959 root	-101   0  4204   680   612 R 45.2  0.0  1h07:29 g
14927 root	-101   0  4204   624   556 R 44.6  0.0  1h04:30 g
14928 root	-101   0  4204   656   588 R 31.1  0.0 47:37.21 g
14961 root	-101   0  4204   684   616 R 31.1  0.0 47:19.75 g
14968 root	-101   0  4204   636   568 R 31.1  0.0 46:27.36 g
14960 root	-101   0  4204   684   616 R 23.8  0.0 37:31.06 g
14969 root	-101   0  4204   684   616 R 23.8  0.0 38:11.50 g
14925 root	-101   0  4204   636   568 R 23.8  0.0 37:34.88 g
14926 root	-101   0  4204   684   616 R 23.8  0.0 38:27.37 g
16182 root	 20   0  124M  3972  3212 R  0.6  0.1  0:00.23 htop
  862 root       20   0  264M  5668  4832 S  0.6  0.1  0:03.30 iio-sensor
 2191 bristot    20   0  649M 41312 32048 S  0.0  1.0  0:27.62 gnome-term
  588 root       20   0  257M  121M  120M S  0.0  3.1  0:13.53 systemd-jo
------------------------------- HTOP ------------------------------------

Some tasks start to use more CPU time, while others seems to use less
CPU than it was reserved for them. See the task 14926, it is using
only 23.8 % of the CPU, which is less than its 10/30 reservation.

I traced this task activation and noticed this:

         swapper     0 [003] 14968.332244: sched:sched_switch: swapper/3:0 [120] R ==> g:14926 [-1]
               g 14926 [003] 14968.339294: sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1]
runtime: 7050 us (14968.339294 - 14968.332244)

period:  29997 us (14968.362241 - 14968.332244)
         swapper     0 [003] 14968.362241: sched:sched_switch: swapper/3:0 [120] R ==> g:14926 [-1]
               g 14926 [003] 14968.369294: sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1]
runtime: 0.007053 us (14968.369294 = 14968.362241)

period: 29994 us (14968.392235 - 14968.362241)
         swapper     0 [003] 14968.392235: sched:sched_switch: swapper/3:0 [120] R ==> g:14926 [-1]
               g 14926 [003] 14968.399301: sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1]
runtime: 7066 us (14968.399301 - 14968.392235)

period:  30008 us (14968.422243 - 14968.392235)
         swapper     0 [003] 14968.422243: sched:sched_switch: swapper/3:0 [120] R ==> g:14926 [-1]
               g 14926 [003] 14968.429294: sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1]
runtime: 7051 us (14968.429294 - 14968.422243)

period:  29995 us (14968.452238 - 14968.422243)
         swapper     0 [003] 14968.452238: sched:sched_switch: swapper/3:0 [120] R ==> g:14926 [-1]
               g 14926 [003] 14968.459293: sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1]
runtime: 7055 us (14968.459293 - 14968.452238)

period:  30055 us (14968.482293 - 14968.452238)
               g 14925 [003] 14968.482293: sched:sched_switch: g:14925 [-1] R ==> g:14926 [-1]
               g 14926 [003] 14968.490293: sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1]
runtime: 8000 us (14968.490293 - 14968.482293)

The task is using less CPU than it was reserved/guaranteed.

After some debugging, it seems that in this case GRUB is also _reducing_
the runtime of the task by making the notion of consumed runtime
be greater than the actual consumed runtime.

You can see this with this code snip:

------------------- %<-------------------
diff --git a/kernel/sched/deadline.c b/kernel/sched/deadline.c
index 93ff400..1abb594 100644
--- a/kernel/sched/deadline.c
+++ b/kernel/sched/deadline.c
@@ -823,9 +823,21 @@ static void update_curr_dl(struct rq *rq)
 
 	sched_rt_avg_update(rq, delta_exec);
 
-	if (unlikely(dl_se->flags & SCHED_FLAG_RECLAIM))
-		delta_exec = grub_reclaim(delta_exec, rq);
-	dl_se->runtime -= delta_exec;
+	if (unlikely(dl_se->flags & SCHED_FLAG_RECLAIM)) {
+		u64 new_delta_exec;
+		new_delta_exec = grub_reclaim(delta_exec, rq);
+		if (new_delta_exec > delta_exec)
+			trace_printk("new delta exec (%llu) is greater than delta exec (%llu) by %llu\n",
+					new_delta_exec,
+					delta_exec,
+					(new_delta_exec - delta_exec));
+		dl_se->runtime -= new_delta_exec;
+	}
+	else {
+		dl_se->runtime -= delta_exec;
+	}
+
+
 
 throttle:
 	if (dl_runtime_exceeded(dl_se) || dl_se->dl_yielded) {
--------------------------- >% -------------

It seems to be related to the "sched/deadline: do not reclaim the whole
CPU bandwidth", because the trace_printk message I put starts to appear
when we start to touch this limit, and the (new_delta_exec - delta_exec)
seems to be somehow limited to the non_deadline_bw.

Output with sysctl -w kernel.sched_rt_runtime_us=950000
               g-1984  [001] d.h1  1108.783349: update_curr_dl: new delta exec (1050043) is greater than delta exec (1000042) by 50001
               g-1983  [002] d.h1  1108.783349: update_curr_dl: new delta exec (1049974) is greater than delta exec (999976) by 49998
               g-1981  [003] d.h1  1108.783350: update_curr_dl: new delta exec (1050054) is greater than delta exec (1000053) by 50001

Output with sysctl -w kernel.sched_rt_runtime_us=900000
               g-1748  [001] d.h1   418.879815: update_curr_dl: new delta exec (1099995) is greater than delta exec (999996) by 99999
               g-1749  [002] d.h1   418.880815: update_curr_dl: new delta exec (1099986) is greater than delta exec (999988) by 99998
               g-1748  [001] d.h1   418.880815: update_curr_dl: new delta exec (1099962) is greater than delta exec (999966) by 99996

In the case of fewer tasks, this error appears just in the
dispatch of a new task, stabilizing after some ms. But it
does not stabilize when we are closer to the limit of the rt
runtime.

That is all I could find today. Am I missing something?

[1] http://bristot.me/lkml/d.c

-- Daniel

[toc] | [next] | [standalone]


#1550224

Fromluca abeni <luca.abeni@unitn.it>
Date2017-01-03 22:40 +0100
Message-ID<sVEYi-3J6-43@gated-at.bofh.it>
In reply to#1550104
Hi Daniel,
(sorry for the previous html email; I replied from my phone and I did
not realise how the email client was configured)

On Tue, 3 Jan 2017 19:58:38 +0100
Daniel Bristot de Oliveira <bristot@redhat.com> wrote:

[...]
> > The implemented CPU reclaiming algorithm is based on tracking the
> > utilization U_act of active tasks (first 2 patches), and modifying
> > the runtime accounting rule (see patch 0004). The original GRUB
> > algorithm is modified as described in [2] to support multiple CPUs
> > (the original algorithm only considered one single CPU, this one
> > tracks U_act per runqueue) and to leave an "unreclaimable" fraction
> > of CPU time to non SCHED_DEADLINE tasks (see patch 0005: the
> > original algorithm can consume 100% of the CPU time, starving all
> > the other tasks). Patch 0003 uses the newly introduced "inactive
> > timer" (introduced in patch 0002) to fix dl_overflow() and
> > __setparam_dl(). Patch 0006 allows to enable CPU reclaiming only on
> > selected tasks.  
> 
> Hi,
> 
> Today I did some tests in this patch set. Unfortunately, it seems that
> there is a problem :-(.
[...]
I reproduced this issue; thanks for the report. It seems to be due to
the fact that the reclaiming tasks are more than the CPU cores and the
load is very high (near to the utilisation limit).

I am investigating it, and will hopefully post an update in the next
days.



			Thanks,
				Luca


> 
> In a four core box, if I dispatch 11 tasks [1] with setup:
> 
>   period = 30 ms
>   runtime = 10 ms
>   flags = 0 (GRUB disabled)
> 
> I see this:
> ------------------------------- HTOP
> ------------------------------------ 1
> [|||||||||||||||||||||92.5%]   Tasks: 128, 259 thr; 14 running 2
> [|||||||||||||||||||||91.0%]   Load average: 4.65 4.66 4.81 3
> [|||||||||||||||||||||92.5%]   Uptime: 05:12:43 4
> [|||||||||||||||||||||92.5%] Mem[|||||||||||||||1.13G/3.78G]
>   Swp[                  0K/3.90G]
> 
>   PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
> 16247 root      -101   0  4204   632   564 R 32.4  0.0  2:10.35 d
> 16249 root	-101   0  4204   624   556 R 32.4  0.0  2:09.80 d
> 16250 root	-101   0  4204   728   660 R 32.4  0.0  2:09.58 d
> 16252 root	-101   0  4204   676   608 R 32.4  0.0  2:09.08 d
> 16253 root	-101   0  4204   636   568 R 32.4  0.0  2:08.85 d
> 16254 root      -101   0  4204   732   664 R 32.4  0.0  2:08.62 d
> 16255 root	-101   0  4204   620   556 R 32.4  0.0  2:08.40 d
> 16257 root	-101   0  4204   708   640 R 32.4  0.0  2:07.98 d
> 16256 root	-101   0  4204   624   560 R 32.4  0.0  2:08.18 d
> 16248 root	-101   0  4204   680   612 R 33.0  0.0  2:10.15 d
> 16251 root	-101   0  4204   676   608 R 33.0  0.0  2:09.34 d
> 16259 root       20   0  124M  4692  3120 R  1.1  0.1  0:02.82 htop
>  2191 bristot    20   0  649M 41312 32048 S  0.0  1.0  0:28.77
> gnome-ter ------------------------------- HTOP
> ------------------------------------
> 
> All tasks are using +- the same amount of CPU time, a little bit more
> than 30%, as expected. However, if I enable GRUB in the same task set
> I get this:
> 
> ------------------------------- HTOP
> ------------------------------------ 1
> [|||||||||||||||||||||93.8%]   Tasks: 128, 260 thr; 15 running 2
> [|||||||||||||||||||||95.2%]   Load average: 5.13 5.01 4.98 3
> [|||||||||||||||||||||93.3%]   Uptime: 05:01:02 4
> [|||||||||||||||||||||96.4%] Mem[|||||||||||||||1.13G/3.78G]
>   Swp[                  0K/3.90G]
> 
>   PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
> 14967 root      -101   0  4204   628   564 R 45.8  0.0  1h07:49 g
> 14962 root	-101   0  4204   728   660 R 45.8  0.0  1h05:06 g
> 14959 root	-101   0  4204   680   612 R 45.2  0.0  1h07:29 g
> 14927 root	-101   0  4204   624   556 R 44.6  0.0  1h04:30 g
> 14928 root	-101   0  4204   656   588 R 31.1  0.0 47:37.21 g
> 14961 root	-101   0  4204   684   616 R 31.1  0.0 47:19.75 g
> 14968 root	-101   0  4204   636   568 R 31.1  0.0 46:27.36 g
> 14960 root	-101   0  4204   684   616 R 23.8  0.0 37:31.06 g
> 14969 root	-101   0  4204   684   616 R 23.8  0.0 38:11.50 g
> 14925 root	-101   0  4204   636   568 R 23.8  0.0 37:34.88 g
> 14926 root	-101   0  4204   684   616 R 23.8  0.0 38:27.37 g
> 16182 root	 20   0  124M  3972  3212 R  0.6  0.1  0:00.23 htop
>   862 root       20   0  264M  5668  4832 S  0.6  0.1  0:03.30
> iio-sensor 2191 bristot    20   0  649M 41312 32048 S  0.0  1.0
> 0:27.62 gnome-term 588 root       20   0  257M  121M  120M S  0.0
> 3.1  0:13.53 systemd-jo ------------------------------- HTOP
> ------------------------------------
> 
> Some tasks start to use more CPU time, while others seems to use less
> CPU than it was reserved for them. See the task 14926, it is using
> only 23.8 % of the CPU, which is less than its 10/30 reservation.
> 
> I traced this task activation and noticed this:
> 
>          swapper     0 [003] 14968.332244: sched:sched_switch:
> swapper/3:0 [120] R ==> g:14926 [-1] g 14926 [003] 14968.339294:
> sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1] runtime: 7050 us
> (14968.339294 - 14968.332244)
> 
> period:  29997 us (14968.362241 - 14968.332244)
>          swapper     0 [003] 14968.362241: sched:sched_switch:
> swapper/3:0 [120] R ==> g:14926 [-1] g 14926 [003] 14968.369294:
> sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1] runtime: 0.007053
> us (14968.369294 = 14968.362241)
> 
> period: 29994 us (14968.392235 - 14968.362241)
>          swapper     0 [003] 14968.392235: sched:sched_switch:
> swapper/3:0 [120] R ==> g:14926 [-1] g 14926 [003] 14968.399301:
> sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1] runtime: 7066 us
> (14968.399301 - 14968.392235)
> 
> period:  30008 us (14968.422243 - 14968.392235)
>          swapper     0 [003] 14968.422243: sched:sched_switch:
> swapper/3:0 [120] R ==> g:14926 [-1] g 14926 [003] 14968.429294:
> sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1] runtime: 7051 us
> (14968.429294 - 14968.422243)
> 
> period:  29995 us (14968.452238 - 14968.422243)
>          swapper     0 [003] 14968.452238: sched:sched_switch:
> swapper/3:0 [120] R ==> g:14926 [-1] g 14926 [003] 14968.459293:
> sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1] runtime: 7055 us
> (14968.459293 - 14968.452238)
> 
> period:  30055 us (14968.482293 - 14968.452238)
>                g 14925 [003] 14968.482293: sched:sched_switch:
> g:14925 [-1] R ==> g:14926 [-1] g 14926 [003] 14968.490293:
> sched:sched_switch: g:14926 [-1] R ==> g:14960 [-1] runtime: 8000 us
> (14968.490293 - 14968.482293)
> 
> The task is using less CPU than it was reserved/guaranteed.
> 
> After some debugging, it seems that in this case GRUB is also
> _reducing_ the runtime of the task by making the notion of consumed
> runtime be greater than the actual consumed runtime.
> 
> You can see this with this code snip:
> 
> ------------------- %<-------------------
> diff --git a/kernel/sched/deadline.c b/kernel/sched/deadline.c
> index 93ff400..1abb594 100644
> --- a/kernel/sched/deadline.c
> +++ b/kernel/sched/deadline.c
> @@ -823,9 +823,21 @@ static void update_curr_dl(struct rq *rq)
>  
>  	sched_rt_avg_update(rq, delta_exec);
>  
> -	if (unlikely(dl_se->flags & SCHED_FLAG_RECLAIM))
> -		delta_exec = grub_reclaim(delta_exec, rq);
> -	dl_se->runtime -= delta_exec;
> +	if (unlikely(dl_se->flags & SCHED_FLAG_RECLAIM)) {
> +		u64 new_delta_exec;
> +		new_delta_exec = grub_reclaim(delta_exec, rq);
> +		if (new_delta_exec > delta_exec)
> +			trace_printk("new delta exec (%llu) is
> greater than delta exec (%llu) by %llu\n",
> +					new_delta_exec,
> +					delta_exec,
> +					(new_delta_exec -
> delta_exec));
> +		dl_se->runtime -= new_delta_exec;
> +	}
> +	else {
> +		dl_se->runtime -= delta_exec;
> +	}
> +
> +
>  
>  throttle:
>  	if (dl_runtime_exceeded(dl_se) || dl_se->dl_yielded) {
> --------------------------- >% -------------  
> 
> It seems to be related to the "sched/deadline: do not reclaim the
> whole CPU bandwidth", because the trace_printk message I put starts
> to appear when we start to touch this limit, and the (new_delta_exec
> - delta_exec) seems to be somehow limited to the non_deadline_bw.
> 
> Output with sysctl -w kernel.sched_rt_runtime_us=950000
>                g-1984  [001] d.h1  1108.783349: update_curr_dl: new
> delta exec (1050043) is greater than delta exec (1000042) by 50001
> g-1983  [002] d.h1  1108.783349: update_curr_dl: new delta exec
> (1049974) is greater than delta exec (999976) by 49998 g-1981  [003]
> d.h1  1108.783350: update_curr_dl: new delta exec (1050054) is
> greater than delta exec (1000053) by 50001
> 
> Output with sysctl -w kernel.sched_rt_runtime_us=900000
>                g-1748  [001] d.h1   418.879815: update_curr_dl: new
> delta exec (1099995) is greater than delta exec (999996) by 99999
> g-1749  [002] d.h1   418.880815: update_curr_dl: new delta exec
> (1099986) is greater than delta exec (999988) by 99998 g-1748  [001]
> d.h1   418.880815: update_curr_dl: new delta exec (1099962) is
> greater than delta exec (999966) by 99996
> 
> In the case of fewer tasks, this error appears just in the
> dispatch of a new task, stabilizing after some ms. But it
> does not stabilize when we are closer to the limit of the rt
> runtime.
> 
> That is all I could find today. Am I missing something?
> 
> [1] http://bristot.me/lkml/d.c
> 
> -- Daniel

[toc] | [prev] | [next] | [standalone]


#1550705

Fromluca abeni <luca.abeni@unitn.it>
Date2017-01-04 13:20 +0100
Message-ID<sVSHT-4z2-7@gated-at.bofh.it>
In reply to#1550104
Hi Daniel,

On Tue, 3 Jan 2017 19:58:38 +0100
Daniel Bristot de Oliveira <bristot@redhat.com> wrote:

[...]
> In a four core box, if I dispatch 11 tasks [1] with setup:
> 
>   period = 30 ms
>   runtime = 10 ms
>   flags = 0 (GRUB disabled)
> 
> I see this:
> ------------------------------- HTOP
> ------------------------------------ 1
> [|||||||||||||||||||||92.5%]   Tasks: 128, 259 thr; 14 running 2
> [|||||||||||||||||||||91.0%]   Load average: 4.65 4.66 4.81 3
> [|||||||||||||||||||||92.5%]   Uptime: 05:12:43 4
> [|||||||||||||||||||||92.5%] Mem[|||||||||||||||1.13G/3.78G]
>   Swp[                  0K/3.90G]
> 
>   PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
> 16247 root      -101   0  4204   632   564 R 32.4  0.0  2:10.35 d
> 16249 root	-101   0  4204   624   556 R 32.4  0.0  2:09.80 d
> 16250 root	-101   0  4204   728   660 R 32.4  0.0  2:09.58 d
> 16252 root	-101   0  4204   676   608 R 32.4  0.0  2:09.08 d
> 16253 root	-101   0  4204   636   568 R 32.4  0.0  2:08.85 d
> 16254 root      -101   0  4204   732   664 R 32.4  0.0  2:08.62 d
> 16255 root	-101   0  4204   620   556 R 32.4  0.0  2:08.40 d
> 16257 root	-101   0  4204   708   640 R 32.4  0.0  2:07.98 d
> 16256 root	-101   0  4204   624   560 R 32.4  0.0  2:08.18 d
> 16248 root	-101   0  4204   680   612 R 33.0  0.0  2:10.15 d
> 16251 root	-101   0  4204   676   608 R 33.0  0.0  2:09.34 d
> 16259 root       20   0  124M  4692  3120 R  1.1  0.1  0:02.82 htop
>  2191 bristot    20   0  649M 41312 32048 S  0.0  1.0  0:28.77
> gnome-ter ------------------------------- HTOP
> ------------------------------------
> 
> All tasks are using +- the same amount of CPU time, a little bit more
> than 30%, as expected.

Notice that, if I understand well, each task should receive 33.33% (1/3)
of CPU time. Anyway, I think this is ok...

> However, if I enable GRUB in the same task set I get this:
> 
> ------------------------------- HTOP
> ------------------------------------ 1
> [|||||||||||||||||||||93.8%]   Tasks: 128, 260 thr; 15 running 2
> [|||||||||||||||||||||95.2%]   Load average: 5.13 5.01 4.98 3
> [|||||||||||||||||||||93.3%]   Uptime: 05:01:02 4
> [|||||||||||||||||||||96.4%] Mem[|||||||||||||||1.13G/3.78G]
>   Swp[                  0K/3.90G]
> 
>   PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
> 14967 root      -101   0  4204   628   564 R 45.8  0.0  1h07:49 g
> 14962 root	-101   0  4204   728   660 R 45.8  0.0  1h05:06 g
> 14959 root	-101   0  4204   680   612 R 45.2  0.0  1h07:29 g
> 14927 root	-101   0  4204   624   556 R 44.6  0.0  1h04:30 g
> 14928 root	-101   0  4204   656   588 R 31.1  0.0 47:37.21 g
> 14961 root	-101   0  4204   684   616 R 31.1  0.0 47:19.75 g
> 14968 root	-101   0  4204   636   568 R 31.1  0.0 46:27.36 g
> 14960 root	-101   0  4204   684   616 R 23.8  0.0 37:31.06 g
> 14969 root	-101   0  4204   684   616 R 23.8  0.0 38:11.50 g
> 14925 root	-101   0  4204   636   568 R 23.8  0.0 37:34.88 g
> 14926 root	-101   0  4204   684   616 R 23.8  0.0 38:27.37 g
> 16182 root	 20   0  124M  3972  3212 R  0.6  0.1  0:00.23 htop
>   862 root       20   0  264M  5668  4832 S  0.6  0.1  0:03.30
> iio-sensor 2191 bristot    20   0  649M 41312 32048 S  0.0  1.0
> 0:27.62 gnome-term 588 root       20   0  257M  121M  120M S  0.0
> 3.1  0:13.53 systemd-jo ------------------------------- HTOP
> ------------------------------------
> 
> Some tasks start to use more CPU time, while others seems to use less
> CPU than it was reserved for them. See the task 14926, it is using
> only 23.8 % of the CPU, which is less than its 10/30 reservation.

What happened here is that some runqueues have an active utilisation
larger than 0.95. So, GRUB is decreasing the amount of time received by
the tasks on those runqueues to consume less than 95%... This is the
reason for the effect you noticed below:


> After some debugging, it seems that in this case GRUB is also
> _reducing_ the runtime of the task by making the notion of consumed
> runtime be greater than the actual consumed runtime.
[...]

Now, this is "kind of expected", because you have 11 tasks each one
having utilisation 1/3, distributed on 4 CPUs... So, some CPU will have
3 tasks on it, resulting in an utilisation = 1 > 0.95. But this should
not result in what you have seen in htop...
The real issue seems to be that at some point some runqueues have an
active utilisation = 1.33 (4 dl tasks in the runqueue), with other
runqueues only having 2 tasks... And this results in the huge imbalance
in utilisations you noticed. I am trying to understand why this
happens... It seems to me that a "pull_dl_task()" might end up pulling
more than 1 task... Is this possible?


			Luca

> 
> You can see this with this code snip:
> 
> ------------------- %<-------------------
> diff --git a/kernel/sched/deadline.c b/kernel/sched/deadline.c
> index 93ff400..1abb594 100644
> --- a/kernel/sched/deadline.c
> +++ b/kernel/sched/deadline.c
> @@ -823,9 +823,21 @@ static void update_curr_dl(struct rq *rq)
>  
>  	sched_rt_avg_update(rq, delta_exec);
>  
> -	if (unlikely(dl_se->flags & SCHED_FLAG_RECLAIM))
> -		delta_exec = grub_reclaim(delta_exec, rq);
> -	dl_se->runtime -= delta_exec;
> +	if (unlikely(dl_se->flags & SCHED_FLAG_RECLAIM)) {
> +		u64 new_delta_exec;
> +		new_delta_exec = grub_reclaim(delta_exec, rq);
> +		if (new_delta_exec > delta_exec)
> +			trace_printk("new delta exec (%llu) is
> greater than delta exec (%llu) by %llu\n",
> +					new_delta_exec,
> +					delta_exec,
> +					(new_delta_exec -
> delta_exec));
> +		dl_se->runtime -= new_delta_exec;
> +	}
> +	else {
> +		dl_se->runtime -= delta_exec;
> +	}
> +
> +
>  
>  throttle:
>  	if (dl_runtime_exceeded(dl_se) || dl_se->dl_yielded) {
> --------------------------- >% -------------  
> 
> It seems to be related to the "sched/deadline: do not reclaim the
> whole CPU bandwidth", because the trace_printk message I put starts
> to appear when we start to touch this limit, and the (new_delta_exec
> - delta_exec) seems to be somehow limited to the non_deadline_bw.
> 
> Output with sysctl -w kernel.sched_rt_runtime_us=950000
>                g-1984  [001] d.h1  1108.783349: update_curr_dl: new
> delta exec (1050043) is greater than delta exec (1000042) by 50001
> g-1983  [002] d.h1  1108.783349: update_curr_dl: new delta exec
> (1049974) is greater than delta exec (999976) by 49998 g-1981  [003]
> d.h1  1108.783350: update_curr_dl: new delta exec (1050054) is
> greater than delta exec (1000053) by 50001
> 
> Output with sysctl -w kernel.sched_rt_runtime_us=900000
>                g-1748  [001] d.h1   418.879815: update_curr_dl: new
> delta exec (1099995) is greater than delta exec (999996) by 99999
> g-1749  [002] d.h1   418.880815: update_curr_dl: new delta exec
> (1099986) is greater than delta exec (999988) by 99998 g-1748  [001]
> d.h1   418.880815: update_curr_dl: new delta exec (1099962) is
> greater than delta exec (999966) by 99996
> 
> In the case of fewer tasks, this error appears just in the
> dispatch of a new task, stabilizing after some ms. But it
> does not stabilize when we are closer to the limit of the rt
> runtime.
> 
> That is all I could find today. Am I missing something?
> 
> [1] http://bristot.me/lkml/d.c
> 
> -- Daniel

[toc] | [prev] | [next] | [standalone]


#1550911

FromDaniel Bristot de Oliveira <bristot@redhat.com>
Date2017-01-04 16:20 +0100
Message-ID<sVVw6-6pX-27@gated-at.bofh.it>
In reply to#1550705
On 01/04/2017 01:17 PM, luca abeni wrote:
> Hi Daniel,
> 
> On Tue, 3 Jan 2017 19:58:38 +0100
> Daniel Bristot de Oliveira <bristot@redhat.com> wrote:
> 
> [...]
>> In a four core box, if I dispatch 11 tasks [1] with setup:
>>
>>   period = 30 ms
>>   runtime = 10 ms
>>   flags = 0 (GRUB disabled)
>>
>> I see this:
>> ------------------------------- HTOP
>> ------------------------------------ 1
>> [|||||||||||||||||||||92.5%]   Tasks: 128, 259 thr; 14 running 2
>> [|||||||||||||||||||||91.0%]   Load average: 4.65 4.66 4.81 3
>> [|||||||||||||||||||||92.5%]   Uptime: 05:12:43 4
>> [|||||||||||||||||||||92.5%] Mem[|||||||||||||||1.13G/3.78G]
>>   Swp[                  0K/3.90G]
>>
>>   PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
>> 16247 root      -101   0  4204   632   564 R 32.4  0.0  2:10.35 d
>> 16249 root	-101   0  4204   624   556 R 32.4  0.0  2:09.80 d
>> 16250 root	-101   0  4204   728   660 R 32.4  0.0  2:09.58 d
>> 16252 root	-101   0  4204   676   608 R 32.4  0.0  2:09.08 d
>> 16253 root	-101   0  4204   636   568 R 32.4  0.0  2:08.85 d
>> 16254 root      -101   0  4204   732   664 R 32.4  0.0  2:08.62 d
>> 16255 root	-101   0  4204   620   556 R 32.4  0.0  2:08.40 d
>> 16257 root	-101   0  4204   708   640 R 32.4  0.0  2:07.98 d
>> 16256 root	-101   0  4204   624   560 R 32.4  0.0  2:08.18 d
>> 16248 root	-101   0  4204   680   612 R 33.0  0.0  2:10.15 d
>> 16251 root	-101   0  4204   676   608 R 33.0  0.0  2:09.34 d
>> 16259 root       20   0  124M  4692  3120 R  1.1  0.1  0:02.82 htop
>>  2191 bristot    20   0  649M 41312 32048 S  0.0  1.0  0:28.77
>> gnome-ter ------------------------------- HTOP
>> ------------------------------------
>>
>> All tasks are using +- the same amount of CPU time, a little bit more
>> than 30%, as expected.
> 
> Notice that, if I understand well, each task should receive 33.33% (1/3)
> of CPU time. Anyway, I think this is ok...

If we think on a partitioned system, yes for the CPUs in which 3 'd'
tasks are able to run. But as sched deadline is global by definition,
the load is:

SUM(U_i)  / M processors.

1/3 * 11  / 4            = 0.916666667

So 10/30 (1/3) of this workload is:
91.6 / 3 = 30.533333333

Well, the rest is probably overheads, like scheduling, migration...

>> However, if I enable GRUB in the same task set I get this:
>>
>> ------------------------------- HTOP
>> ------------------------------------ 1
>> [|||||||||||||||||||||93.8%]   Tasks: 128, 260 thr; 15 running 2
>> [|||||||||||||||||||||95.2%]   Load average: 5.13 5.01 4.98 3
>> [|||||||||||||||||||||93.3%]   Uptime: 05:01:02 4
>> [|||||||||||||||||||||96.4%] Mem[|||||||||||||||1.13G/3.78G]
>>   Swp[                  0K/3.90G]
>>
>>   PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
>> 14967 root      -101   0  4204   628   564 R 45.8  0.0  1h07:49 g
>> 14962 root	-101   0  4204   728   660 R 45.8  0.0  1h05:06 g
>> 14959 root	-101   0  4204   680   612 R 45.2  0.0  1h07:29 g
>> 14927 root	-101   0  4204   624   556 R 44.6  0.0  1h04:30 g
>> 14928 root	-101   0  4204   656   588 R 31.1  0.0 47:37.21 g
>> 14961 root	-101   0  4204   684   616 R 31.1  0.0 47:19.75 g
>> 14968 root	-101   0  4204   636   568 R 31.1  0.0 46:27.36 g
>> 14960 root	-101   0  4204   684   616 R 23.8  0.0 37:31.06 g
>> 14969 root	-101   0  4204   684   616 R 23.8  0.0 38:11.50 g
>> 14925 root	-101   0  4204   636   568 R 23.8  0.0 37:34.88 g
>> 14926 root	-101   0  4204   684   616 R 23.8  0.0 38:27.37 g
>> 16182 root	 20   0  124M  3972  3212 R  0.6  0.1  0:00.23 htop
>>   862 root       20   0  264M  5668  4832 S  0.6  0.1  0:03.30
>> iio-sensor 2191 bristot    20   0  649M 41312 32048 S  0.0  1.0
>> 0:27.62 gnome-term 588 root       20   0  257M  121M  120M S  0.0
>> 3.1  0:13.53 systemd-jo ------------------------------- HTOP
>> ------------------------------------
>>
>> Some tasks start to use more CPU time, while others seems to use less
>> CPU than it was reserved for them. See the task 14926, it is using
>> only 23.8 % of the CPU, which is less than its 10/30 reservation.
> 
> What happened here is that some runqueues have an active utilisation
> larger than 0.95. So, GRUB is decreasing the amount of time received by
> the tasks on those runqueues to consume less than 95%... This is the
> reason for the effect you noticed below:

I see. But, AFAIK, the Linux's sched deadline measures the load
globally, not locally. So, it is not a problem having a load > than 95%
in the local queue if the global queue is < 95%.

Am I missing something?

> 
>> After some debugging, it seems that in this case GRUB is also
>> _reducing_ the runtime of the task by making the notion of consumed
>> runtime be greater than the actual consumed runtime.
> [...]
> 
> Now, this is "kind of expected", because you have 11 tasks each one
> having utilisation 1/3, distributed on 4 CPUs... So, some CPU will have
> 3 tasks on it, resulting in an utilisation = 1 > 0.95. But this should
> not result in what you have seen in htop...

Well, the sched deadline aims to schedule the M highest priority tasks,
and migrates tasks to achieve this goal. However, I am not sure if
having the whole runqueue balance is a goal/restriction/feature of the
deadline scheduler.

Maybe this is the difference between the GRUB and sched deadline
assumptions that is causing the problem. Just thinking aloud.

> The real issue seems to be that at some point some runqueues have an
> active utilisation = 1.33 (4 dl tasks in the runqueue), with other
> runqueues only having 2 tasks... And this results in the huge imbalance
> in utilisations you noticed. I am trying to understand why this
> happens... It seems to me that a "pull_dl_task()" might end up pulling
> more than 1 task... Is this possible?

Yeah, this explain the numbers.

Brainstorm time! (sorry if it sounds obviously unfeasible):
Is it possible to think on GRUB tracking the global utilization?

-- Daniel

[toc] | [prev] | [next] | [standalone]


#1550986

FromLuca Abeni <luca.abeni@unitn.it>
Date2017-01-04 17:50 +0100
Message-ID<sVWVc-7cW-19@gated-at.bofh.it>
In reply to#1550911
Hi Daniel,

2017-01-04 16:14 GMT+01:00, Daniel Bristot de Oliveira <bristot@redhat.com>:
> On 01/04/2017 01:17 PM, luca abeni wrote:
>> Hi Daniel,
>>
>> On Tue, 3 Jan 2017 19:58:38 +0100
>> Daniel Bristot de Oliveira <bristot@redhat.com> wrote:
>>
>> [...]
>>> In a four core box, if I dispatch 11 tasks [1] with setup:
>>>
>>>   period = 30 ms
>>>   runtime = 10 ms
>>>   flags = 0 (GRUB disabled)
>>>
>>> I see this:
>>> ------------------------------- HTOP
>>> ------------------------------------ 1
>>> [|||||||||||||||||||||92.5%]   Tasks: 128, 259 thr; 14 running 2
>>> [|||||||||||||||||||||91.0%]   Load average: 4.65 4.66 4.81 3
>>> [|||||||||||||||||||||92.5%]   Uptime: 05:12:43 4
>>> [|||||||||||||||||||||92.5%] Mem[|||||||||||||||1.13G/3.78G]
>>>   Swp[                  0K/3.90G]
>>>
>>>   PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
>>> 16247 root      -101   0  4204   632   564 R 32.4  0.0  2:10.35 d
>>> 16249 root	-101   0  4204   624   556 R 32.4  0.0  2:09.80 d
>>> 16250 root	-101   0  4204   728   660 R 32.4  0.0  2:09.58 d
>>> 16252 root	-101   0  4204   676   608 R 32.4  0.0  2:09.08 d
>>> 16253 root	-101   0  4204   636   568 R 32.4  0.0  2:08.85 d
>>> 16254 root      -101   0  4204   732   664 R 32.4  0.0  2:08.62 d
>>> 16255 root	-101   0  4204   620   556 R 32.4  0.0  2:08.40 d
>>> 16257 root	-101   0  4204   708   640 R 32.4  0.0  2:07.98 d
>>> 16256 root	-101   0  4204   624   560 R 32.4  0.0  2:08.18 d
>>> 16248 root	-101   0  4204   680   612 R 33.0  0.0  2:10.15 d
>>> 16251 root	-101   0  4204   676   608 R 33.0  0.0  2:09.34 d
>>> 16259 root       20   0  124M  4692  3120 R  1.1  0.1  0:02.82 htop
>>>  2191 bristot    20   0  649M 41312 32048 S  0.0  1.0  0:28.77
>>> gnome-ter ------------------------------- HTOP
>>> ------------------------------------
>>>
>>> All tasks are using +- the same amount of CPU time, a little bit more
>>> than 30%, as expected.
>>
>> Notice that, if I understand well, each task should receive 33.33% (1/3)
>> of CPU time. Anyway, I think this is ok...
>
> If we think on a partitioned system, yes for the CPUs in which 3 'd'
> tasks are able to run. But as sched deadline is global by definition,
> the load is:
>
> SUM(U_i)  / M processors.
>
> 1/3 * 11  / 4            = 0.916666667
>
> So 10/30 (1/3) of this workload is:
> 91.6 / 3 = 30.533333333
>
> Well, the rest is probably overheads, like scheduling, migration...

I do not think this math is correct... Yes, the total utilization of
the taskset is 0.91 (or 3.66, depending on how you define the
utilization...), but I still think that the percentage of CPU time
shown by "top" or "htop" should be 33.33 (or 8.33, depending on how
the tool computes it).
runtime=10 and period=30 means "schedule the task for 10ms every
30ms", so the task will consume 33% of the CPU time of a single core.
In other words, 10/30 is a fraction of the CPU time, not a fraction of
the time consumed by SCHED_DEADLINE tasks.


>>> However, if I enable GRUB in the same task set I get this:
>>>
>>> ------------------------------- HTOP
>>> ------------------------------------ 1
>>> [|||||||||||||||||||||93.8%]   Tasks: 128, 260 thr; 15 running 2
>>> [|||||||||||||||||||||95.2%]   Load average: 5.13 5.01 4.98 3
>>> [|||||||||||||||||||||93.3%]   Uptime: 05:01:02 4
>>> [|||||||||||||||||||||96.4%] Mem[|||||||||||||||1.13G/3.78G]
>>>   Swp[                  0K/3.90G]
>>>
>>>   PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
>>> 14967 root      -101   0  4204   628   564 R 45.8  0.0  1h07:49 g
>>> 14962 root	-101   0  4204   728   660 R 45.8  0.0  1h05:06 g
>>> 14959 root	-101   0  4204   680   612 R 45.2  0.0  1h07:29 g
>>> 14927 root	-101   0  4204   624   556 R 44.6  0.0  1h04:30 g
>>> 14928 root	-101   0  4204   656   588 R 31.1  0.0 47:37.21 g
>>> 14961 root	-101   0  4204   684   616 R 31.1  0.0 47:19.75 g
>>> 14968 root	-101   0  4204   636   568 R 31.1  0.0 46:27.36 g
>>> 14960 root	-101   0  4204   684   616 R 23.8  0.0 37:31.06 g
>>> 14969 root	-101   0  4204   684   616 R 23.8  0.0 38:11.50 g
>>> 14925 root	-101   0  4204   636   568 R 23.8  0.0 37:34.88 g
>>> 14926 root	-101   0  4204   684   616 R 23.8  0.0 38:27.37 g
>>> 16182 root	 20   0  124M  3972  3212 R  0.6  0.1  0:00.23 htop
>>>   862 root       20   0  264M  5668  4832 S  0.6  0.1  0:03.30
>>> iio-sensor 2191 bristot    20   0  649M 41312 32048 S  0.0  1.0
>>> 0:27.62 gnome-term 588 root       20   0  257M  121M  120M S  0.0
>>> 3.1  0:13.53 systemd-jo ------------------------------- HTOP
>>> ------------------------------------
>>>
>>> Some tasks start to use more CPU time, while others seems to use less
>>> CPU than it was reserved for them. See the task 14926, it is using
>>> only 23.8 % of the CPU, which is less than its 10/30 reservation.
>>
>> What happened here is that some runqueues have an active utilisation
>> larger than 0.95. So, GRUB is decreasing the amount of time received by
>> the tasks on those runqueues to consume less than 95%... This is the
>> reason for the effect you noticed below:
>
> I see. But, AFAIK, the Linux's sched deadline measures the load
> globally, not locally. So, it is not a problem having a load > than 95%
> in the local queue if the global queue is < 95%.
>
> Am I missing something?

The version of GRUB reclaiming implemented in my patches tracks a
per-runqueue "active utilization", and uses it for reclaiming.

>>> After some debugging, it seems that in this case GRUB is also
>>> _reducing_ the runtime of the task by making the notion of consumed
>>> runtime be greater than the actual consumed runtime.
>> [...]
>>
>> Now, this is "kind of expected", because you have 11 tasks each one
>> having utilisation 1/3, distributed on 4 CPUs... So, some CPU will have
>> 3 tasks on it, resulting in an utilisation = 1 > 0.95. But this should
>> not result in what you have seen in htop...
>
> Well, the sched deadline aims to schedule the M highest priority tasks,
> and migrates tasks to achieve this goal. However, I am not sure if
> having the whole runqueue balance is a goal/restriction/feature of the
> deadline scheduler.
>
> Maybe this is the difference between the GRUB and sched deadline
> assumptions that is causing the problem. Just thinking aloud.

I think I found some strange behaviour in the push/pull mechanisms (at
least it seems strange to me): a "pull" operation might end up pulling
multiple tasks (I see this can simplify the implementation, but I
think pulling multiple tasks is useless and might introduce some
overhead even independently from my patches), and I suspect (but still
I need to verify this) a "push" operation can push a task to a "wrong"
destination runqueue (I mean, a task is pushed to a runqueue where it
is not the earliest deadline task)...

Without reclaiming, this just results in useless migrations (if I did
not misunderstand something), but with my reclaiming patches this is
probably the source of the strange effect you saw. But I am still
investigating this, so I am not too sure...

>> The real issue seems to be that at some point some runqueues have an
>> active utilisation = 1.33 (4 dl tasks in the runqueue), with other
>> runqueues only having 2 tasks... And this results in the huge imbalance
>> in utilisations you noticed. I am trying to understand why this
>> happens... It seems to me that a "pull_dl_task()" might end up pulling
>> more than 1 task... Is this possible?
>
> Yeah, this explain the numbers.
>
> Brainstorm time! (sorry if it sounds obviously unfeasible):
> Is it possible to think on GRUB tracking the global utilization?

Yes, and I even had a version of my patches using a "per root domain"
global active utilization. If needed I can update my patchset to
implement the global active utilization again.
I switched to per-runqueue active utilization because:
- this can be used for controlling the CPU frequency scaling... And
I've been told that frequency scaling is generally per-core / per-CPU
(but I need to verify this)
- the patches based on global active utilization needed to access this
global utilization in mutual exclusion, so I used a spinlock to
protect it... And I am not sure about scalability issues
- I suspect there were issues when the root domain / exclusive cpuset
is modified.


Thanks,
Luca

[toc] | [prev] | [next] | [standalone]


#1551052

FromDaniel Bristot de Oliveira <bristot@redhat.com>
Date2017-01-04 19:10 +0100
Message-ID<sVYaB-8bd-11@gated-at.bofh.it>
In reply to#1550986
On 01/04/2017 05:42 PM, Luca Abeni wrote:
> Hi Daniel,
> 
> 2017-01-04 16:14 GMT+01:00, Daniel Bristot de Oliveira <bristot@redhat.com>:
>> On 01/04/2017 01:17 PM, luca abeni wrote:
>>> Hi Daniel,
>>>
>>> On Tue, 3 Jan 2017 19:58:38 +0100
>>> Daniel Bristot de Oliveira <bristot@redhat.com> wrote:
>>>
>>> [...]
>>>> In a four core box, if I dispatch 11 tasks [1] with setup:
>>>>
>>>>   period = 30 ms
>>>>   runtime = 10 ms
>>>>   flags = 0 (GRUB disabled)
>>>>
>>>> I see this:
>>>> ------------------------------- HTOP
>>>> ------------------------------------ 1
>>>> [|||||||||||||||||||||92.5%]   Tasks: 128, 259 thr; 14 running 2
>>>> [|||||||||||||||||||||91.0%]   Load average: 4.65 4.66 4.81 3
>>>> [|||||||||||||||||||||92.5%]   Uptime: 05:12:43 4
>>>> [|||||||||||||||||||||92.5%] Mem[|||||||||||||||1.13G/3.78G]
>>>>   Swp[                  0K/3.90G]
>>>>
>>>>   PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
>>>> 16247 root      -101   0  4204   632   564 R 32.4  0.0  2:10.35 d
>>>> 16249 root	-101   0  4204   624   556 R 32.4  0.0  2:09.80 d
>>>> 16250 root	-101   0  4204   728   660 R 32.4  0.0  2:09.58 d
>>>> 16252 root	-101   0  4204   676   608 R 32.4  0.0  2:09.08 d
>>>> 16253 root	-101   0  4204   636   568 R 32.4  0.0  2:08.85 d
>>>> 16254 root      -101   0  4204   732   664 R 32.4  0.0  2:08.62 d
>>>> 16255 root	-101   0  4204   620   556 R 32.4  0.0  2:08.40 d
>>>> 16257 root	-101   0  4204   708   640 R 32.4  0.0  2:07.98 d
>>>> 16256 root	-101   0  4204   624   560 R 32.4  0.0  2:08.18 d
>>>> 16248 root	-101   0  4204   680   612 R 33.0  0.0  2:10.15 d
>>>> 16251 root	-101   0  4204   676   608 R 33.0  0.0  2:09.34 d
>>>> 16259 root       20   0  124M  4692  3120 R  1.1  0.1  0:02.82 htop
>>>>  2191 bristot    20   0  649M 41312 32048 S  0.0  1.0  0:28.77
>>>> gnome-ter ------------------------------- HTOP
>>>> ------------------------------------
>>>>
>>>> All tasks are using +- the same amount of CPU time, a little bit more
>>>> than 30%, as expected.
>>>
>>> Notice that, if I understand well, each task should receive 33.33% (1/3)
>>> of CPU time. Anyway, I think this is ok...
>>
>> If we think on a partitioned system, yes for the CPUs in which 3 'd'
>> tasks are able to run. But as sched deadline is global by definition,
>> the load is:
>>
>> SUM(U_i)  / M processors.
>>
>> 1/3 * 11  / 4            = 0.916666667
>>
>> So 10/30 (1/3) of this workload is:
>> 91.6 / 3 = 30.533333333
>>
>> Well, the rest is probably overheads, like scheduling, migration...
> 
> I do not think this math is correct... Yes, the total utilization of
> the taskset is 0.91 (or 3.66, depending on how you define the
> utilization...), but I still think that the percentage of CPU time
> shown by "top" or "htop" should be 33.33 (or 8.33, depending on how
> the tool computes it).
> runtime=10 and period=30 means "schedule the task for 10ms every
> 30ms", so the task will consume 33% of the CPU time of a single core.
> In other words, 10/30 is a fraction of the CPU time, not a fraction of
> the time consumed by SCHED_DEADLINE tasks.

Ack! you are correct, I was so focused on global utilization that end up
missing this point. For the top/htop it should 33.3%.

> 
>>>> However, if I enable GRUB in the same task set I get this:
>>>>
>>>> ------------------------------- HTOP
>>>> ------------------------------------ 1
>>>> [|||||||||||||||||||||93.8%]   Tasks: 128, 260 thr; 15 running 2
>>>> [|||||||||||||||||||||95.2%]   Load average: 5.13 5.01 4.98 3
>>>> [|||||||||||||||||||||93.3%]   Uptime: 05:01:02 4
>>>> [|||||||||||||||||||||96.4%] Mem[|||||||||||||||1.13G/3.78G]
>>>>   Swp[                  0K/3.90G]
>>>>
>>>>   PID USER      PRI  NI  VIRT   RES   SHR S CPU% MEM%   TIME+  Command
>>>> 14967 root      -101   0  4204   628   564 R 45.8  0.0  1h07:49 g
>>>> 14962 root	-101   0  4204   728   660 R 45.8  0.0  1h05:06 g
>>>> 14959 root	-101   0  4204   680   612 R 45.2  0.0  1h07:29 g
>>>> 14927 root	-101   0  4204   624   556 R 44.6  0.0  1h04:30 g
>>>> 14928 root	-101   0  4204   656   588 R 31.1  0.0 47:37.21 g
>>>> 14961 root	-101   0  4204   684   616 R 31.1  0.0 47:19.75 g
>>>> 14968 root	-101   0  4204   636   568 R 31.1  0.0 46:27.36 g
>>>> 14960 root	-101   0  4204   684   616 R 23.8  0.0 37:31.06 g
>>>> 14969 root	-101   0  4204   684   616 R 23.8  0.0 38:11.50 g
>>>> 14925 root	-101   0  4204   636   568 R 23.8  0.0 37:34.88 g
>>>> 14926 root	-101   0  4204   684   616 R 23.8  0.0 38:27.37 g
>>>> 16182 root	 20   0  124M  3972  3212 R  0.6  0.1  0:00.23 htop
>>>>   862 root       20   0  264M  5668  4832 S  0.6  0.1  0:03.30
>>>> iio-sensor 2191 bristot    20   0  649M 41312 32048 S  0.0  1.0
>>>> 0:27.62 gnome-term 588 root       20   0  257M  121M  120M S  0.0
>>>> 3.1  0:13.53 systemd-jo ------------------------------- HTOP
>>>> ------------------------------------
>>>>
>>>> Some tasks start to use more CPU time, while others seems to use less
>>>> CPU than it was reserved for them. See the task 14926, it is using
>>>> only 23.8 % of the CPU, which is less than its 10/30 reservation.
>>>
>>> What happened here is that some runqueues have an active utilisation
>>> larger than 0.95. So, GRUB is decreasing the amount of time received by
>>> the tasks on those runqueues to consume less than 95%... This is the
>>> reason for the effect you noticed below:
>>
>> I see. But, AFAIK, the Linux's sched deadline measures the load
>> globally, not locally. So, it is not a problem having a load > than 95%
>> in the local queue if the global queue is < 95%.
>>
>> Am I missing something?
> 
> The version of GRUB reclaiming implemented in my patches tracks a
> per-runqueue "active utilization", and uses it for reclaiming.

I _think_ that this might be (one of) the source(s) of the problem...

Just exercising...

For example, with my taskset, with a hypothetical perfect balance of the
whole runqueue, one possible scenario is:

   CPU    0    1     2     3
# TASKS   3    3     3     2

In this case, CPUs 0 1 2 are with 100% of local utilization. Thus, the
current task on these CPUs will have their runtime decreased by GRUB.
Meanwhile, the luck tasks in the CPU 3 would use an additional time that
they "globally" do not have - because the system, globally, has a load
higher than the 66.6...% of the local runqueue. Actually, part of the
time decreased from tasks on [0-2] are being used by the tasks on 3,
until the next migration of any task, which will change the luck
tasks... but without any guaranty that all tasks will be the luck one on
every activation, causing the problem.

Does it make sense?

If it does, this let me think that only with the global track of
utilization we will achieve the correct result... but I may be missing
something... :-).

-- Daniel

[toc] | [prev] | [next] | [standalone]


#1551109

FromLuca Abeni <luca.abeni@unitn.it>
Date2017-01-04 19:50 +0100
Message-ID<sVYNl-8qL-47@gated-at.bofh.it>
In reply to#1551052
2017-01-04 19:00 GMT+01:00, Daniel Bristot de Oliveira <bristot@redhat.com>:
[...]
>>>>> Some tasks start to use more CPU time, while others seems to use less
>>>>> CPU than it was reserved for them. See the task 14926, it is using
>>>>> only 23.8 % of the CPU, which is less than its 10/30 reservation.
>>>>
>>>> What happened here is that some runqueues have an active utilisation
>>>> larger than 0.95. So, GRUB is decreasing the amount of time received by
>>>> the tasks on those runqueues to consume less than 95%... This is the
>>>> reason for the effect you noticed below:
>>>
>>> I see. But, AFAIK, the Linux's sched deadline measures the load
>>> globally, not locally. So, it is not a problem having a load > than 95%
>>> in the local queue if the global queue is < 95%.
>>>
>>> Am I missing something?
>>
>> The version of GRUB reclaiming implemented in my patches tracks a
>> per-runqueue "active utilization", and uses it for reclaiming.
>
> I _think_ that this might be (one of) the source(s) of the problem...
I agree that this can cause some problems, but I am not sure if it
justifies the huge difference in utilisations you observed

> Just exercising...
>
> For example, with my taskset, with a hypothetical perfect balance of the
> whole runqueue, one possible scenario is:
>
>    CPU    0    1     2     3
> # TASKS   3    3     3     2
>
> In this case, CPUs 0 1 2 are with 100% of local utilization. Thus, the
> current task on these CPUs will have their runtime decreased by GRUB.
> Meanwhile, the luck tasks in the CPU 3 would use an additional time that
> they "globally" do not have - because the system, globally, has a load
> higher than the 66.6...% of the local runqueue. Actually, part of the
> time decreased from tasks on [0-2] are being used by the tasks on 3,
> until the next migration of any task, which will change the luck
> tasks... but without any guaranty that all tasks will be the luck one on
> every activation, causing the problem.
>
> Does it make sense?

Yes; but my impression is that gEDF will migrate tasks so that the
distribution of the reclaimed CPU bandwidth is almost uniform...
Instead, you saw huge differences in the utilisations (and I do not
think that "compressing" the utilisations from 100% to 95% can
decrease the utilisation of a task from 33% to 25% / 26%... :)

I suspect there is something more going on here (might be some bug in
one of my patches). I am trying to better understand what happened.

> If it does, this let me think that only with the global track of
> utilization we will achieve the correct result... but I may be missing
> something... :-).

Of course tracking the global active utilisation can be a solution,
but I also want to better understand what is wrong with the current
approach.

Thanks,
Luca

[toc] | [prev] | [next] | [standalone]


#1556437

FromJuri Lelli <juri.lelli@arm.com>
Date2017-01-11 13:20 +0100
Message-ID<sYq2J-6QF-1@gated-at.bofh.it>
In reply to#1551109
Hi,

On 04/01/17 19:30, Luca Abeni wrote:
> 2017-01-04 19:00 GMT+01:00, Daniel Bristot de Oliveira <bristot@redhat.com>:
> [...]
> >>>>> Some tasks start to use more CPU time, while others seems to use less
> >>>>> CPU than it was reserved for them. See the task 14926, it is using
> >>>>> only 23.8 % of the CPU, which is less than its 10/30 reservation.
> >>>>
> >>>> What happened here is that some runqueues have an active utilisation
> >>>> larger than 0.95. So, GRUB is decreasing the amount of time received by
> >>>> the tasks on those runqueues to consume less than 95%... This is the
> >>>> reason for the effect you noticed below:
> >>>
> >>> I see. But, AFAIK, the Linux's sched deadline measures the load
> >>> globally, not locally. So, it is not a problem having a load > than 95%
> >>> in the local queue if the global queue is < 95%.
> >>>
> >>> Am I missing something?
> >>
> >> The version of GRUB reclaiming implemented in my patches tracks a
> >> per-runqueue "active utilization", and uses it for reclaiming.
> >
> > I _think_ that this might be (one of) the source(s) of the problem...
> I agree that this can cause some problems, but I am not sure if it
> justifies the huge difference in utilisations you observed
> 
> > Just exercising...
> >
> > For example, with my taskset, with a hypothetical perfect balance of the
> > whole runqueue, one possible scenario is:
> >
> >    CPU    0    1     2     3
> > # TASKS   3    3     3     2
> >
> > In this case, CPUs 0 1 2 are with 100% of local utilization. Thus, the
> > current task on these CPUs will have their runtime decreased by GRUB.
> > Meanwhile, the luck tasks in the CPU 3 would use an additional time that
> > they "globally" do not have - because the system, globally, has a load
> > higher than the 66.6...% of the local runqueue. Actually, part of the
> > time decreased from tasks on [0-2] are being used by the tasks on 3,
> > until the next migration of any task, which will change the luck
> > tasks... but without any guaranty that all tasks will be the luck one on
> > every activation, causing the problem.
> >
> > Does it make sense?
> 
> Yes; but my impression is that gEDF will migrate tasks so that the
> distribution of the reclaimed CPU bandwidth is almost uniform...
> Instead, you saw huge differences in the utilisations (and I do not
> think that "compressing" the utilisations from 100% to 95% can
> decrease the utilisation of a task from 33% to 25% / 26%... :)
>

I tried to replicate Daniel's experiment, but I don't see such a skewed
allocation. They get a reasonably uniform bandwidth and the trace
looks fairly good as well (all processes get to run on the different
processors at some time).

> I suspect there is something more going on here (might be some bug in
> one of my patches). I am trying to better understand what happened.
> 

However, playing with this a bit further, I found out one thing that
looks counter-intuitive (at least to me :).

Simplifying Daniel's example, let's say that we have one 10/30 task
running on a CPU with a 500/1000 global limit. Applying grub_reclaim()
formula we have:

 delta_exec = delta * (0.5 + 0.333) = delta * 0.833

Which in practice means that 1ms of real delta (at 1000HZ) corresponds
to 0.833ms of virtual delta. Considering this, a 10ms (over 30ms)
reservation gets "extended" to ~12ms (over 30ms), that is to say the
task consumes 0.4 of the CPU's bandwidth. top seems to back what I'm
saying, but am I still talking nonsense? :)

I was expecting that the task could consume 0.5 worth of bandwidth with
the given global limit. Is the current behaviour intended?

If we want to change this behaviour maybe something like the following
might work?

 delta_exec = (delta * to_ratio((1ULL << 20) - rq->dl.non_deadline_bw,
                                rq->dl.running_bw)) >> 20

The idea would be to normalize running_bw over the available dl_bw.

Thoughts?

Best,

- Juri

[toc] | [prev] | [next] | [standalone]


#1556464

FromLuca Abeni <luca.abeni@santannapisa.it>
Date2017-01-11 14:00 +0100
Message-ID<sYqFr-72T-17@gated-at.bofh.it>
In reply to#1556437
Hi Juri,
(I reply from my new email address)

On Wed, 11 Jan 2017 12:19:51 +0000
Juri Lelli <juri.lelli@arm.com> wrote:
[...]
> > > For example, with my taskset, with a hypothetical perfect balance
> > > of the whole runqueue, one possible scenario is:
> > >
> > >    CPU    0    1     2     3
> > > # TASKS   3    3     3     2
> > >
> > > In this case, CPUs 0 1 2 are with 100% of local utilization.
> > > Thus, the current task on these CPUs will have their runtime
> > > decreased by GRUB. Meanwhile, the luck tasks in the CPU 3 would
> > > use an additional time that they "globally" do not have - because
> > > the system, globally, has a load higher than the 66.6...% of the
> > > local runqueue. Actually, part of the time decreased from tasks
> > > on [0-2] are being used by the tasks on 3, until the next
> > > migration of any task, which will change the luck tasks... but
> > > without any guaranty that all tasks will be the luck one on every
> > > activation, causing the problem.
> > >
> > > Does it make sense?  
> > 
> > Yes; but my impression is that gEDF will migrate tasks so that the
> > distribution of the reclaimed CPU bandwidth is almost uniform...
> > Instead, you saw huge differences in the utilisations (and I do not
> > think that "compressing" the utilisations from 100% to 95% can
> > decrease the utilisation of a task from 33% to 25% / 26%... :)
> >  
> 
> I tried to replicate Daniel's experiment, but I don't see such a
> skewed allocation. They get a reasonably uniform bandwidth and the
> trace looks fairly good as well (all processes get to run on the
> different processors at some time).

With some effort, I replicated the issue noticed by Daniel... I think
it also depends on the CPU speed (and on good or bad luck :), but the
"unfair" CPU allocation can actually happen.
I am working on a fix (based on the m-grub modifications proposed at
last April's SAC - in my original patchset, I over-simplified the
algorithm).


> > I suspect there is something more going on here (might be some bug
> > in one of my patches). I am trying to better understand what
> > happened. 
> 
> However, playing with this a bit further, I found out one thing that
> looks counter-intuitive (at least to me :).
> 
> Simplifying Daniel's example, let's say that we have one 10/30 task
> running on a CPU with a 500/1000 global limit. Applying grub_reclaim()
> formula we have:
> 
>  delta_exec = delta * (0.5 + 0.333) = delta * 0.833
> 
> Which in practice means that 1ms of real delta (at 1000HZ) corresponds
> to 0.833ms of virtual delta. Considering this, a 10ms (over 30ms)
> reservation gets "extended" to ~12ms (over 30ms), that is to say the
> task consumes 0.4 of the CPU's bandwidth. top seems to back what I'm
> saying, but am I still talking nonsense? :)

You are right; my "Do not reclaim the whole CPU bandwidth" patch is an
approximation... I hoped that this approximation could be more precise
than what it really is.
I used the "Uact + unreclaimable utilization" equation to avoid
divisions in grub_reclaim(), but the equation should really be "Uact /
reclaimable utilization"... So, in your example it is
	delta * 0.3333 / 0.5 = delta * 0.6666
that results in 15ms over 30ms, as expected.

I'll fix that patch for the next submission.

> I was expecting that the task could consume 0.5 worth of bandwidth
> with the given global limit. Is the current behaviour intended?
> 
> If we want to change this behaviour maybe something like the following
> might work?
> 
>  delta_exec = (delta * to_ratio((1ULL << 20) - rq->dl.non_deadline_bw,
>                                 rq->dl.running_bw)) >> 20
My current patch does
	(delta * rq->dl.running_bw * rq->dl.deadline_bw_inv) >> 20 >> 8;
where rq->dl.deadline_bw_inv has been set to
	to_ratio(global_rt_runtime(), global_rt_period()) >> 12;
	
This seems to work fine, and should introduce less overhead than
to_ratio().


		Thanks,
			Luca

[toc] | [prev] | [next] | [standalone]


#1556581

FromJuri Lelli <juri.lelli@arm.com>
Date2017-01-11 16:10 +0100
Message-ID<sYsHg-8rV-45@gated-at.bofh.it>
In reply to#1556464
On 11/01/17 13:39, Luca Abeni wrote:
> Hi Juri,
> (I reply from my new email address)
> 
> On Wed, 11 Jan 2017 12:19:51 +0000
> Juri Lelli <juri.lelli@arm.com> wrote:
> [...]
> > > > For example, with my taskset, with a hypothetical perfect balance
> > > > of the whole runqueue, one possible scenario is:
> > > >
> > > >    CPU    0    1     2     3
> > > > # TASKS   3    3     3     2
> > > >
> > > > In this case, CPUs 0 1 2 are with 100% of local utilization.
> > > > Thus, the current task on these CPUs will have their runtime
> > > > decreased by GRUB. Meanwhile, the luck tasks in the CPU 3 would
> > > > use an additional time that they "globally" do not have - because
> > > > the system, globally, has a load higher than the 66.6...% of the
> > > > local runqueue. Actually, part of the time decreased from tasks
> > > > on [0-2] are being used by the tasks on 3, until the next
> > > > migration of any task, which will change the luck tasks... but
> > > > without any guaranty that all tasks will be the luck one on every
> > > > activation, causing the problem.
> > > >
> > > > Does it make sense?  
> > > 
> > > Yes; but my impression is that gEDF will migrate tasks so that the
> > > distribution of the reclaimed CPU bandwidth is almost uniform...
> > > Instead, you saw huge differences in the utilisations (and I do not
> > > think that "compressing" the utilisations from 100% to 95% can
> > > decrease the utilisation of a task from 33% to 25% / 26%... :)
> > >  
> > 
> > I tried to replicate Daniel's experiment, but I don't see such a
> > skewed allocation. They get a reasonably uniform bandwidth and the
> > trace looks fairly good as well (all processes get to run on the
> > different processors at some time).
> 
> With some effort, I replicated the issue noticed by Daniel... I think
> it also depends on the CPU speed (and on good or bad luck :), but the
> "unfair" CPU allocation can actually happen.

Yeah, actual allocation in general varies. I guess the question is: do
we care? We currently don't load balance considering utilizations, only
dynamic deadlines matter.

> I am working on a fix (based on the m-grub modifications proposed at
> last April's SAC - in my original patchset, I over-simplified the
> algorithm).
> 

OK, will have a look to next version.

> 
> > > I suspect there is something more going on here (might be some bug
> > > in one of my patches). I am trying to better understand what
> > > happened. 
> > 
> > However, playing with this a bit further, I found out one thing that
> > looks counter-intuitive (at least to me :).
> > 
> > Simplifying Daniel's example, let's say that we have one 10/30 task
> > running on a CPU with a 500/1000 global limit. Applying grub_reclaim()
> > formula we have:
> > 
> >  delta_exec = delta * (0.5 + 0.333) = delta * 0.833
> > 
> > Which in practice means that 1ms of real delta (at 1000HZ) corresponds
> > to 0.833ms of virtual delta. Considering this, a 10ms (over 30ms)
> > reservation gets "extended" to ~12ms (over 30ms), that is to say the
> > task consumes 0.4 of the CPU's bandwidth. top seems to back what I'm
> > saying, but am I still talking nonsense? :)
> 
> You are right; my "Do not reclaim the whole CPU bandwidth" patch is an
> approximation... I hoped that this approximation could be more precise
> than what it really is.
> I used the "Uact + unreclaimable utilization" equation to avoid
> divisions in grub_reclaim(), but the equation should really be "Uact /
> reclaimable utilization"... So, in your example it is
> 	delta * 0.3333 / 0.5 = delta * 0.6666
> that results in 15ms over 30ms, as expected.
> 
> I'll fix that patch for the next submission.
> 

Right, OK.

> > I was expecting that the task could consume 0.5 worth of bandwidth
> > with the given global limit. Is the current behaviour intended?
> > 
> > If we want to change this behaviour maybe something like the following
> > might work?
> > 
> >  delta_exec = (delta * to_ratio((1ULL << 20) - rq->dl.non_deadline_bw,
> >                                 rq->dl.running_bw)) >> 20
> My current patch does
> 	(delta * rq->dl.running_bw * rq->dl.deadline_bw_inv) >> 20 >> 8;
> where rq->dl.deadline_bw_inv has been set to
> 	to_ratio(global_rt_runtime(), global_rt_period()) >> 12;
> 	
> This seems to work fine, and should introduce less overhead than
> to_ratio().
> 

Sure, we don't want to do divisions if we can. Why the intermediate
right shifts, though?

Thanks,

- Juri

[toc] | [prev] | [next] | [standalone]


#1556941

Fromluca abeni <luca.abeni@santannapisa.it>
Date2017-01-11 22:20 +0100
Message-ID<sYytk-3x3-31@gated-at.bofh.it>
In reply to#1556581
On Wed, 11 Jan 2017 15:06:47 +0000
Juri Lelli <juri.lelli@arm.com> wrote:

> On 11/01/17 13:39, Luca Abeni wrote:
> > Hi Juri,
> > (I reply from my new email address)
> > 
> > On Wed, 11 Jan 2017 12:19:51 +0000
> > Juri Lelli <juri.lelli@arm.com> wrote:
> > [...]  
> > > > > For example, with my taskset, with a hypothetical perfect
> > > > > balance of the whole runqueue, one possible scenario is:
> > > > >
> > > > >    CPU    0    1     2     3
> > > > > # TASKS   3    3     3     2
> > > > >
> > > > > In this case, CPUs 0 1 2 are with 100% of local utilization.
> > > > > Thus, the current task on these CPUs will have their runtime
> > > > > decreased by GRUB. Meanwhile, the luck tasks in the CPU 3
> > > > > would use an additional time that they "globally" do not have
> > > > > - because the system, globally, has a load higher than the
> > > > > 66.6...% of the local runqueue. Actually, part of the time
> > > > > decreased from tasks on [0-2] are being used by the tasks on
> > > > > 3, until the next migration of any task, which will change
> > > > > the luck tasks... but without any guaranty that all tasks
> > > > > will be the luck one on every activation, causing the problem.
> > > > >
> > > > > Does it make sense?    
> > > > 
> > > > Yes; but my impression is that gEDF will migrate tasks so that
> > > > the distribution of the reclaimed CPU bandwidth is almost
> > > > uniform... Instead, you saw huge differences in the
> > > > utilisations (and I do not think that "compressing" the
> > > > utilisations from 100% to 95% can decrease the utilisation of a
> > > > task from 33% to 25% / 26%... :) 
> > > 
> > > I tried to replicate Daniel's experiment, but I don't see such a
> > > skewed allocation. They get a reasonably uniform bandwidth and the
> > > trace looks fairly good as well (all processes get to run on the
> > > different processors at some time).  
> > 
> > With some effort, I replicated the issue noticed by Daniel... I
> > think it also depends on the CPU speed (and on good or bad luck :),
> > but the "unfair" CPU allocation can actually happen.  
> 
> Yeah, actual allocation in general varies. I guess the question is: do
> we care? We currently don't load balance considering utilizations,
> only dynamic deadlines matter.

Right... But the problem is that with the version of GRUB I proposed
this unfairness can result in some tasks receiving less CPU time than
the guaranteed amount (because some other tasks receive much more). I
think there are at least two possible ways to fix this (without
changing the migration strategy), and I am working on them...
(hopefully, I'll post something in next week)


> > > I was expecting that the task could consume 0.5 worth of bandwidth
> > > with the given global limit. Is the current behaviour intended?
> > > 
> > > If we want to change this behaviour maybe something like the
> > > following might work?
> > > 
> > >  delta_exec = (delta * to_ratio((1ULL << 20) -
> > > rq->dl.non_deadline_bw, rq->dl.running_bw)) >> 20  
> > My current patch does
> > 	(delta * rq->dl.running_bw * rq->dl.deadline_bw_inv) >> 20
> > >> 8; where rq->dl.deadline_bw_inv has been set to
> > 	to_ratio(global_rt_runtime(), global_rt_period()) >> 12;
> > 	
> > This seems to work fine, and should introduce less overhead than
> > to_ratio().
> >   
> 
> Sure, we don't want to do divisions if we can. Why the intermediate
> right shifts, though?

I wrote it like this to remember that ">> 20" comes from how
"to_ratio()" computes the utilization, and the additional ">> 8"
comes from the fact that deadline_bw_inv is shifted left by 8, to avoid
losing precision (I used 8 insted of 20 so that the computation can be
- hopefully - performed on 32 bits... Of course I can revise this if
needed).

If needed I can change the ">> 20 >> 8" in ">> 28", or remove the
">> 12" from the deadline_bw_inv conmputation (so that we can use
">> 40" or ">> 20 >> 20" in grub_reclaim()).


			Thanks,
				Luca

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web