Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1560846 > unrolled thread

Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes

Started byThomas Gleixner <tglx@linutronix.de>
First post2017-01-17 18:40 +0100
Last post2017-01-20 20:40 +0100
Articles 9 on this page of 29 — 7 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-17 18:40 +0100
    Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-18 03:50 +0100
      Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-18 10:00 +0100
        Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Peter Zijlstra <peterz@infradead.org> - 2017-01-18 11:10 +0100
          Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-19 21:10 +0100
        Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-18 20:50 +0100
        RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Yu, Fenghua" <fenghua.yu@intel.com> - 2017-01-18 22:20 +0100
        Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-18 22:20 +0100
          Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-19 18:50 +0100
            Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-20 08:50 +0100
              Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-20 09:40 +0100
                Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-20 21:30 +0100
        Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-19 03:20 +0100
          Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-19 18:30 +0100
            Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-19 19:50 +0100
        Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Vikas Shivappa <vikas.shivappa@linux.intel.com> - 2017-01-19 03:30 +0100
        Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Stephane Eranian <eranian@google.com> - 2017-01-19 07:50 +0100
          Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-19 19:50 +0100
        Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Vikas Shivappa <vikas.shivappa@linux.intel.com> - 2017-01-20 03:40 +0100
          Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-20 09:00 +0100
            Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-20 15:50 +0100
              Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-20 21:20 +0100
                Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-20 22:10 +0100
                  Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-20 22:50 +0100
                    Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-21 01:00 +0100
                Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-23 11:20 +0100
                  Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Peter Zijlstra <peterz@infradead.org> - 2017-01-23 12:40 +0100
            Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-20 21:50 +0100
          Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Stephane Eranian <eranian@google.com> - 2017-01-20 20:40 +0100

Page 2 of 2 — ← Prev page 1 [2]


#1563681

FromThomas Gleixner <tglx@linutronix.de>
Date2017-01-20 15:50 +0100
Message-ID<t1IFQ-8tF-21@gated-at.bofh.it>
In reply to#1563354
On Thu, 19 Jan 2017, David Carrillo-Cisneros wrote:
> 
> If resctrl groups could lift the restriction of one resctl per CLOSID,
> then the user can create many resctrl in the way perf cgroups are
> created now. The advantage is that there wont be cgroup hierarchy!
> making things much simpler. Also no need to optimize perf event
> context switch to make llc_occupancy work.

So if I understand you correctly, then you want a mechanism to have groups
of entities (tasks, cpus) and associate them to a particular resource
control group.

So they share the CLOSID of the control group and each entity group can
have its own RMID.

Now you want to be able to move the entity groups around between control
groups without losing the RMID associated to the entity group.

So the whole picture would look like this:

rdt ->  CTRLGRP -> CLOSID

mon ->  MONGRP  -> RMID
   
And you want to move MONGRP from one CTRLGRP to another.

Can you please write up in a abstract way what the design requirements are
that you need. So far we are talking about implementation details and
unspecfied wishlists, but what we really need is an abstract requirement.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


#1563892

FromDavid Carrillo-Cisneros <davidcc@google.com>
Date2017-01-20 21:20 +0100
Message-ID<t1NPc-3oi-15@gated-at.bofh.it>
In reply to#1563681
On Fri, Jan 20, 2017 at 5:29 AM Thomas Gleixner <tglx@linutronix.de> wrote:
>
> On Thu, 19 Jan 2017, David Carrillo-Cisneros wrote:
> >
> > If resctrl groups could lift the restriction of one resctl per CLOSID,
> > then the user can create many resctrl in the way perf cgroups are
> > created now. The advantage is that there wont be cgroup hierarchy!
> > making things much simpler. Also no need to optimize perf event
> > context switch to make llc_occupancy work.
>
> So if I understand you correctly, then you want a mechanism to have groups
> of entities (tasks, cpus) and associate them to a particular resource
> control group.
>
> So they share the CLOSID of the control group and each entity group can
> have its own RMID.
>
> Now you want to be able to move the entity groups around between control
> groups without losing the RMID associated to the entity group.
>
> So the whole picture would look like this:
>
> rdt ->  CTRLGRP -> CLOSID
>
> mon ->  MONGRP  -> RMID
>
> And you want to move MONGRP from one CTRLGRP to another.

Almost, but not quite. My idea is no have MONGRP and CTRLGRP to be the
same thing. Details below.

>
> Can you please write up in a abstract way what the design requirements are
> that you need. So far we are talking about implementation details and
> unspecfied wishlists, but what we really need is an abstract requirement.

My pleasure:


Design Proposal for Monitoring of RDT Allocation Groups.
-----------------------------------------------------------------------------

Currently each CTRLGRP has a unique CLOSID and a (most likely) unique
cache bitmask (CBM) per resource. Non-unique CBM are possible although
useless. An unique CLOSID forbids more CTRLGRPs than physical CLOSIDs.
CLOSIDs are much more scarce than RMIDs.

If we lift the condition of unique CLOSID, then the user can create
multiple CTRLGRPs with the same schemata. Internally, those CTRCGRP
would share the CLOSID and RDT_Allocation must maintain the schemata
to CLOSID relationship (similarly to what the previous CAT driver used
to do). Elements in CTRLGRP.tasks and CTRLGRP.cpus behave the same as
now: adding an element removes it from its previous CTRLGRP.


This change would allow further partitioning the allocation groups
into (allocation, monitoring) groups as follows:

With allocation only:
            CTRLGRP0     CTRLGRP_ALLOC_ONLY
schemata:  L3:0=0xff0       L3:0=x00f
tasks:       PID0       P0_0,P0_1,P1_0,P1_1
cpus:        0x3                0xC

If we want to monitor (P0_0,P0_1), (P1_0,P1_1) and CPUs 0xC
independently, with the new model we could create:
            CTRLGRP0     CTRLGRP1     CTRLGRP2        CTRLGRP3
schemata:  L3:0=0xff0   L3:0=x00f    L3:0=0x00f     L3:0=0x00f
tasks:       PID0         <none>      P0_0,P0_1     P1_0, P1_1
cpus:        0x3           0xC          0x0             0x0

Internally, CTRLGRP1, CTRLGRP2, and CTRLGRP2 would share the CLOSID for (L3,0).


Now we can ask perf to monitor any of the CTRLGRPs independently -once
we solve how to pass to perf what (CTRLGRP, resource_id) to monitor-.
The perf_event will reserve and assign the RMID to the monitored
CTRLGRP. The RDT subsystem will context switch the whole PQR_ASSOC MSR
(CLOSID and RMID), so perf won't have to.

If CTRLGRP's schemata changes, the RDT subsystem will find a new
CLOSID for the new schemata (potentially reusing an existing one) or
fail (just like the old CAT used to). The RMID does not change during
schemata updates.

If a CTRLGRP dies, the monitoring perf_event continues to exists as a
useless wraith, just as happens with cgroup events now.

Since CTRLGRPs have no hierarchy. There is no need to handle that in
the new RDT Monitoring PMU, greatly simplifying it over the previously
proposed versions.

A breaking change in user observed behavior with respect to the
existing CQM PMU is that there wouldn't be task events. A task must be
part of a CTRLGRP and events are created per (CTRLGRP, resource_id)
pair. If an user wants to monitor a task across multiple resources
(e.g. l3_occupancy across two packages), she must create one event per
resource_id and add the two counts.

I see this breaking change as an improvement, since hiding the cache
topology to user space introduced lots of ugliness and complexity to
the CQM PMU without improving accuracy over user space adding the
events.

Implementation ideas:

First idea is to expose one monitoring file per resource in a CTRLGRP,
so the list of CTRLGRP's files would be: schemata, tasks, cpus,
monitor_l3_0, monitor_l3_1, ...

the monitor_<resource_id> file descriptor is passed to perf_event_open
in the way cgroup file descriptors are passed now. All events to the
same (CTRLGRP,resource_id) share RMID.

The RMID allocation part can either be handled by RDT Allocation or by
the RDT Monitoring PMU. Either ways, the existence of PMU's
perf_events allocates/releases the RMID.

Also, since this new design removes hierarchy and task events, it
allows for a simple solution of the RMID rotation problem. The removal
of task events eliminates the cgroup vs task event conflict existing
in the upstream version; it also removes the need to ensure that all
active packages have RMIDs at the same time that added complexity to
my version of CQM/CMT. Lastly, the removal of hierarchy removes the
reliance on cgroups, the complex tree based read, and all the hooks
and cgroup files that "raped" the cgroup subsystem.

Thoughts?

Thanks,
David

[toc] | [prev] | [next] | [standalone]


#1563918

FromShivappa Vikas <vikas.shivappa@intel.com>
Date2017-01-20 22:10 +0100
Message-ID<t1OBA-3UM-29@gated-at.bofh.it>
In reply to#1563892

On Fri, 20 Jan 2017, David Carrillo-Cisneros wrote:

> On Fri, Jan 20, 2017 at 5:29 AM Thomas Gleixner <tglx@linutronix.de> wrote:
>>
>> On Thu, 19 Jan 2017, David Carrillo-Cisneros wrote:
>>>
>>> If resctrl groups could lift the restriction of one resctl per CLOSID,
>>> then the user can create many resctrl in the way perf cgroups are
>>> created now. The advantage is that there wont be cgroup hierarchy!
>>> making things much simpler. Also no need to optimize perf event
>>> context switch to make llc_occupancy work.
>>
>> So if I understand you correctly, then you want a mechanism to have groups
>> of entities (tasks, cpus) and associate them to a particular resource
>> control group.
>>
>> So they share the CLOSID of the control group and each entity group can
>> have its own RMID.
>>
>> Now you want to be able to move the entity groups around between control
>> groups without losing the RMID associated to the entity group.
>>
>> So the whole picture would look like this:
>>
>> rdt ->  CTRLGRP -> CLOSID
>>
>> mon ->  MONGRP  -> RMID
>>
>> And you want to move MONGRP from one CTRLGRP to another.
>
> Almost, but not quite. My idea is no have MONGRP and CTRLGRP to be the
> same thing. Details below.
>
>>
>> Can you please write up in a abstract way what the design requirements are
>> that you need. So far we are talking about implementation details and
>> unspecfied wishlists, but what we really need is an abstract requirement.
>
> My pleasure:
>
>
> Design Proposal for Monitoring of RDT Allocation Groups.
> -----------------------------------------------------------------------------
>
> Currently each CTRLGRP has a unique CLOSID and a (most likely) unique
> cache bitmask (CBM) per resource. Non-unique CBM are possible although
> useless. An unique CLOSID forbids more CTRLGRPs than physical CLOSIDs.
> CLOSIDs are much more scarce than RMIDs.
>
> If we lift the condition of unique CLOSID, then the user can create
> multiple CTRLGRPs with the same schemata. Internally, those CTRCGRP
> would share the CLOSID and RDT_Allocation must maintain the schemata
> to CLOSID relationship (similarly to what the previous CAT driver used
> to do). Elements in CTRLGRP.tasks and CTRLGRP.cpus behave the same as
> now: adding an element removes it from its previous CTRLGRP.
>
>
> This change would allow further partitioning the allocation groups
> into (allocation, monitoring) groups as follows:
>
> With allocation only:
>            CTRLGRP0     CTRLGRP_ALLOC_ONLY
> schemata:  L3:0=0xff0       L3:0=x00f
> tasks:       PID0       P0_0,P0_1,P1_0,P1_1
> cpus:        0x3                0xC

Not clear what the PID0 and P0_0 mean ?

If you have to support something like MONGRP and CTRLGRP overall 
you want to allow for a task to be present in multiple groups ?

>
> If we want to monitor (P0_0,P0_1), (P1_0,P1_1) and CPUs 0xC
> independently, with the new model we could create:
>            CTRLGRP0     CTRLGRP1     CTRLGRP2        CTRLGRP3
> schemata:  L3:0=0xff0   L3:0=x00f    L3:0=0x00f     L3:0=0x00f
> tasks:       PID0         <none>      P0_0,P0_1     P1_0, P1_1
> cpus:        0x3           0xC          0x0             0x0
>
> Internally, CTRLGRP1, CTRLGRP2, and CTRLGRP2 would share the CLOSID for (L3,0).
>
>
> Now we can ask perf to monitor any of the CTRLGRPs independently -once
> we solve how to pass to perf what (CTRLGRP, resource_id) to monitor-.
> The perf_event will reserve and assign the RMID to the monitored
> CTRLGRP. The RDT subsystem will context switch the whole PQR_ASSOC MSR
> (CLOSID and RMID), so perf won't have to.

This can be solved by suporting just the -t in perf and a new option in perf to 
suport resctrl group monitoring (something similar to -R). That way we provide 
the flexible granularity to monitor tasks 
independent of whether they are in any resctrl group (and hence also a subset).

CTRLGRP		TASKS		MASK
CTRLGRP1	PID1,PID2	L3:0=0Xf,1=0xf0
CTRLGRP2	PID3,PID4	L3:0=0Xf0,1=0xf00

#perf stat -e llc_occupancy -R CTRLGRP1

#perf stat -e llc_occupancy -t PID3,PID4

The RMID allocation is independent of resctrl CLOSid allocation and hence the 
RMID is not always married to CLOS which seems like the requirement here.

OR

We could have CTRLGRPs with control_only, monitor_only or control_monitor 
options.

now a task could be present in both control_only and monitor_only
group or it could be present only in a control_monitor_group. The transitions 
from one state to another are guarded by this same principle.

CTRLGRP		TASKS		MASK			TYPE
CTRLGRP1	PID1,PID2	L3:0=0Xf,1=0xf0		control_only
CTRLGRP2	PID3,PID4	L3:0=0Xf0,1=0xf00	control_only
CTRLGRP3	PID2,PID3				monitor_only
CTRLGRP4	PID5,PID6	L3:0=0Xf0,1=0xf00	control_monitor

CTRLGRP3 allows you to monitor a set of tasks which is not bound to be in the 
same CTRLGRP and you can add or move tasks into this. The adding and removing 
the tasks is whats easily supported compared to the task granularity although 
such a thing could still be supported with the task granularity.

CTRLGRP4 allows you to tie the monitor and control together so when tasks move 
in and out of this we still have that group to consider. And these groups still 
retain the cpu masks like before so that cpu monitoring is still supported.

In this case we would need a new option to support the ctrlgrp monitoring in 
perf or a new tool to do all this if we dont want to bother perf.

>
> If CTRLGRP's schemata changes, the RDT subsystem will find a new
> CLOSID for the new schemata (potentially reusing an existing one) or
> fail (just like the old CAT used to). The RMID does not change during
> schemata updates.
>
> If a CTRLGRP dies, the monitoring perf_event continues to exists as a
> useless wraith, just as happens with cgroup events now.
>
> Since CTRLGRPs have no hierarchy. There is no need to handle that in
> the new RDT Monitoring PMU, greatly simplifying it over the previously
> proposed versions.
>
> A breaking change in user observed behavior with respect to the
> existing CQM PMU is that there wouldn't be task events. A task must be
> part of a CTRLGRP and events are created per (CTRLGRP, resource_id)
> pair. If an user wants to monitor a task across multiple resources
> (e.g. l3_occupancy across two packages), she must create one event per
> resource_id and add the two counts.
>
> I see this breaking change as an improvement, since hiding the cache
> topology to user space introduced lots of ugliness and complexity to
> the CQM PMU without improving accuracy over user space adding the
> events.
>
> Implementation ideas:
>
> First idea is to expose one monitoring file per resource in a CTRLGRP,
> so the list of CTRLGRP's files would be: schemata, tasks, cpus,
> monitor_l3_0, monitor_l3_1, ...
>
> the monitor_<resource_id> file descriptor is passed to perf_event_open
> in the way cgroup file descriptors are passed now. All events to the
> same (CTRLGRP,resource_id) share RMID.
>
> The RMID allocation part can either be handled by RDT Allocation or by
> the RDT Monitoring PMU. Either ways, the existence of PMU's
> perf_events allocates/releases the RMID.
>
> Also, since this new design removes hierarchy and task events, it
> allows for a simple solution of the RMID rotation problem. The removal
> of task events eliminates the cgroup vs task event conflict existing
> in the upstream version; it also removes the need to ensure that all
> active packages have RMIDs at the same time that added complexity to
> my version of CQM/CMT. Lastly, the removal of hierarchy removes the
> reliance on cgroups, the complex tree based read, and all the hooks
> and cgroup files that "raped" the cgroup subsystem.

Yes, not sure if the view is same after I sent the implementation details in 
documentation :) (most likely it is).
But the option could be to not support perf_cgroup for cqm and 
support a new option in perf to monitor resctrl groups and tasks (or some other 
options like mongrp)

I am so far inclined to creating a new monitoring interface that way we dont try 
to "rape" the existing perf specifics for this RDT or later RDT quirk/features.

>
> Thoughts?
>
> Thanks,
> David
>

[toc] | [prev] | [next] | [standalone]


#1563929

FromDavid Carrillo-Cisneros <davidcc@google.com>
Date2017-01-20 22:50 +0100
Message-ID<t1Pei-48I-13@gated-at.bofh.it>
In reply to#1563918
On Fri, Jan 20, 2017 at 1:08 PM, Shivappa Vikas
<vikas.shivappa@intel.com> wrote:
>
>
> On Fri, 20 Jan 2017, David Carrillo-Cisneros wrote:
>
>> On Fri, Jan 20, 2017 at 5:29 AM Thomas Gleixner <tglx@linutronix.de>
>> wrote:
>>>
>>>
>>> On Thu, 19 Jan 2017, David Carrillo-Cisneros wrote:
>>>>
>>>>
>>>> If resctrl groups could lift the restriction of one resctl per CLOSID,
>>>> then the user can create many resctrl in the way perf cgroups are
>>>> created now. The advantage is that there wont be cgroup hierarchy!
>>>> making things much simpler. Also no need to optimize perf event
>>>> context switch to make llc_occupancy work.
>>>
>>>
>>> So if I understand you correctly, then you want a mechanism to have
>>> groups
>>> of entities (tasks, cpus) and associate them to a particular resource
>>> control group.
>>>
>>> So they share the CLOSID of the control group and each entity group can
>>> have its own RMID.
>>>
>>> Now you want to be able to move the entity groups around between control
>>> groups without losing the RMID associated to the entity group.
>>>
>>> So the whole picture would look like this:
>>>
>>> rdt ->  CTRLGRP -> CLOSID
>>>
>>> mon ->  MONGRP  -> RMID
>>>
>>> And you want to move MONGRP from one CTRLGRP to another.
>>
>>
>> Almost, but not quite. My idea is no have MONGRP and CTRLGRP to be the
>> same thing. Details below.
>>
>>>
>>> Can you please write up in a abstract way what the design requirements
>>> are
>>> that you need. So far we are talking about implementation details and
>>> unspecfied wishlists, but what we really need is an abstract requirement.
>>
>>
>> My pleasure:
>>
>>
>> Design Proposal for Monitoring of RDT Allocation Groups.
>>
>> -----------------------------------------------------------------------------
>>
>> Currently each CTRLGRP has a unique CLOSID and a (most likely) unique
>> cache bitmask (CBM) per resource. Non-unique CBM are possible although
>> useless. An unique CLOSID forbids more CTRLGRPs than physical CLOSIDs.
>> CLOSIDs are much more scarce than RMIDs.
>>
>> If we lift the condition of unique CLOSID, then the user can create
>> multiple CTRLGRPs with the same schemata. Internally, those CTRCGRP
>> would share the CLOSID and RDT_Allocation must maintain the schemata
>> to CLOSID relationship (similarly to what the previous CAT driver used
>> to do). Elements in CTRLGRP.tasks and CTRLGRP.cpus behave the same as
>> now: adding an element removes it from its previous CTRLGRP.
>>
>>
>> This change would allow further partitioning the allocation groups
>> into (allocation, monitoring) groups as follows:
>>
>> With allocation only:
>>            CTRLGRP0     CTRLGRP_ALLOC_ONLY
>> schemata:  L3:0=0xff0       L3:0=x00f
>> tasks:       PID0       P0_0,P0_1,P1_0,P1_1
>> cpus:        0x3                0xC
>
>
> Not clear what the PID0 and P0_0 mean ?

PID0, and P*_* are arbitrary PIDs. The tasks file works the same as it
does now in RDT. I am not changing that.

>
> If you have to support something like MONGRP and CTRLGRP overall you want to
> allow for a task to be present in multiple groups ?

I am not proposing to support MONGRP and CTRLGRP. I am proposing to
allow monitoring of CTRGRPs only.

>>
>> If we want to monitor (P0_0,P0_1), (P1_0,P1_1) and CPUs 0xC
>> independently, with the new model we could create:
>>            CTRLGRP0     CTRLGRP1     CTRLGRP2        CTRLGRP3
>> schemata:  L3:0=0xff0   L3:0=x00f    L3:0=0x00f     L3:0=0x00f
>> tasks:       PID0         <none>      P0_0,P0_1     P1_0, P1_1
>> cpus:        0x3           0xC          0x0             0x0
>>
>> Internally, CTRLGRP1, CTRLGRP2, and CTRLGRP2 would share the CLOSID for
>> (L3,0).
>>
>>
>> Now we can ask perf to monitor any of the CTRLGRPs independently -once
>> we solve how to pass to perf what (CTRLGRP, resource_id) to monitor-.
>> The perf_event will reserve and assign the RMID to the monitored
>> CTRLGRP. The RDT subsystem will context switch the whole PQR_ASSOC MSR
>> (CLOSID and RMID), so perf won't have to.
>
>
> This can be solved by suporting just the -t in perf and a new option in perf
> to suport resctrl group monitoring (something similar to -R). That way we
> provide the flexible granularity to monitor tasks independent of whether
> they are in any resctrl group (and hence also a subset).

One of the key points of my proposal is to remove monitoring PIDs
independently. That simplifies things by letting RDT handle CLOSIDs
and RMIDs together.

>
> CTRLGRP         TASKS           MASK
> CTRLGRP1        PID1,PID2       L3:0=0Xf,1=0xf0
> CTRLGRP2        PID3,PID4       L3:0=0Xf0,1=0xf00
>
> #perf stat -e llc_occupancy -R CTRLGRP1
>
> #perf stat -e llc_occupancy -t PID3,PID4
>
> The RMID allocation is independent of resctrl CLOSid allocation and hence
> the RMID is not always married to CLOS which seems like the requirement
> here.

It is not a requirement. Both the CLOSID and the RMID of a CTRLGRP can
change in my proposal.

>
> OR
>
> We could have CTRLGRPs with control_only, monitor_only or control_monitor
> options.
>
> now a task could be present in both control_only and monitor_only
> group or it could be present only in a control_monitor_group. The
> transitions from one state to another are guarded by this same principle.
>
> CTRLGRP         TASKS           MASK                    TYPE
> CTRLGRP1        PID1,PID2       L3:0=0Xf,1=0xf0         control_only
> CTRLGRP2        PID3,PID4       L3:0=0Xf0,1=0xf00       control_only
> CTRLGRP3        PID2,PID3                               monitor_only
> CTRLGRP4        PID5,PID6       L3:0=0Xf0,1=0xf00       control_monitor
>
> CTRLGRP3 allows you to monitor a set of tasks which is not bound to be in
> the same CTRLGRP and you can add or move tasks into this. The adding and
> removing the tasks is whats easily supported compared to the task
> granularity although such a thing could still be supported with the task
> granularity.
>
> CTRLGRP4 allows you to tie the monitor and control together so when tasks
> move in and out of this we still have that group to consider. And these
> groups still retain the cpu masks like before so that cpu monitoring is
> still supported.

Instead of having 3 types of CTRLGRPl, I am proposing one kind
(equivalent to your control_monitor type) that uses a non-zero RMID
when an appropriate perf_event is attached to it. What advantages do
you see on having 3 distinct types?

>
> In this case we would need a new option to support the ctrlgrp monitoring in
> perf or a new tool to do all this if we dont want to bother perf.
>
>

Agree, I like expanding the cgroup fd option to take CTRLGRP fds, as
described in the Implementation Ideas part of the proposal.

>>
>> If CTRLGRP's schemata changes, the RDT subsystem will find a new
>> CLOSID for the new schemata (potentially reusing an existing one) or
>> fail (just like the old CAT used to). The RMID does not change during
>> schemata updates.
>>
>> If a CTRLGRP dies, the monitoring perf_event continues to exists as a
>> useless wraith, just as happens with cgroup events now.
>>
>> Since CTRLGRPs have no hierarchy. There is no need to handle that in
>> the new RDT Monitoring PMU, greatly simplifying it over the previously
>> proposed versions.
>>
>> A breaking change in user observed behavior with respect to the
>> existing CQM PMU is that there wouldn't be task events. A task must be
>> part of a CTRLGRP and events are created per (CTRLGRP, resource_id)
>> pair. If an user wants to monitor a task across multiple resources
>> (e.g. l3_occupancy across two packages), she must create one event per
>> resource_id and add the two counts.
>>
>> I see this breaking change as an improvement, since hiding the cache
>> topology to user space introduced lots of ugliness and complexity to
>> the CQM PMU without improving accuracy over user space adding the
>> events.
>>
>> Implementation ideas:
>>
>> First idea is to expose one monitoring file per resource in a CTRLGRP,
>> so the list of CTRLGRP's files would be: schemata, tasks, cpus,
>> monitor_l3_0, monitor_l3_1, ...
>>
>> the monitor_<resource_id> file descriptor is passed to perf_event_open
>> in the way cgroup file descriptors are passed now. All events to the
>> same (CTRLGRP,resource_id) share RMID.
>>
>> The RMID allocation part can either be handled by RDT Allocation or by
>> the RDT Monitoring PMU. Either ways, the existence of PMU's
>> perf_events allocates/releases the RMID.
>>
>> Also, since this new design removes hierarchy and task events, it
>> allows for a simple solution of the RMID rotation problem. The removal
>> of task events eliminates the cgroup vs task event conflict existing
>> in the upstream version; it also removes the need to ensure that all
>> active packages have RMIDs at the same time that added complexity to
>> my version of CQM/CMT. Lastly, the removal of hierarchy removes the
>> reliance on cgroups, the complex tree based read, and all the hooks
>> and cgroup files that "raped" the cgroup subsystem.
>
>
> Yes, not sure if the view is same after I sent the implementation details in
> documentation :) (most likely it is).
> But the option could be to not support perf_cgroup for cqm and support a new
> option in perf to monitor resctrl groups and tasks (or some other options
> like mongrp)

Agree with no supporting cgroups. This proposal is about supporting
neither cgroups nor tasks and do all monitoring through CTRLGRPs
through an expansion of an existing perf option.

>
> I am so far inclined to creating a new monitoring interface that way we dont
> try to "rape" the existing perf specifics for this RDT or later RDT
> quirk/features.
>

On first inspection it seems to me like perf would be fine with this
approach. It requires no changes to the system call and just some
changes in the way the cgroup_fd is handled in perf_event_open
(besides making sure that a context-less PMU don't break things). Do
you foresee any conflict with future features?

Thanks,
David

[toc] | [prev] | [next] | [standalone]


#1563999

FromShivappa Vikas <vikas.shivappa@intel.com>
Date2017-01-21 01:00 +0100
Message-ID<t1Rg5-5kW-5@gated-at.bofh.it>
In reply to#1563929

On Fri, 20 Jan 2017, David Carrillo-Cisneros wrote:

> On Fri, Jan 20, 2017 at 1:08 PM, Shivappa Vikas
> <vikas.shivappa@intel.com> wrote:
>>
>>
>> On Fri, 20 Jan 2017, David Carrillo-Cisneros wrote:
>>
>>> On Fri, Jan 20, 2017 at 5:29 AM Thomas Gleixner <tglx@linutronix.de>
>>> wrote:
>>>>
>>>>
>>>> On Thu, 19 Jan 2017, David Carrillo-Cisneros wrote:
>>>>>
>>>>>
>>>>> If resctrl groups could lift the restriction of one resctl per CLOSID,
>>>>> then the user can create many resctrl in the way perf cgroups are
>>>>> created now. The advantage is that there wont be cgroup hierarchy!
>>>>> making things much simpler. Also no need to optimize perf event
>>>>> context switch to make llc_occupancy work.
>>>>
>>>>
>>>> So if I understand you correctly, then you want a mechanism to have
>>>> groups
>>>> of entities (tasks, cpus) and associate them to a particular resource
>>>> control group.
>>>>
>>>> So they share the CLOSID of the control group and each entity group can
>>>> have its own RMID.
>>>>
>>>> Now you want to be able to move the entity groups around between control
>>>> groups without losing the RMID associated to the entity group.
>>>>
>>>> So the whole picture would look like this:
>>>>
>>>> rdt ->  CTRLGRP -> CLOSID
>>>>
>>>> mon ->  MONGRP  -> RMID
>>>>
>>>> And you want to move MONGRP from one CTRLGRP to another.
>>>
>>>
>>> Almost, but not quite. My idea is no have MONGRP and CTRLGRP to be the
>>> same thing. Details below.
>>>
>>>>
>>>> Can you please write up in a abstract way what the design requirements
>>>> are
>>>> that you need. So far we are talking about implementation details and
>>>> unspecfied wishlists, but what we really need is an abstract requirement.
>>>
>>>
>>> My pleasure:
>>>
>>>
>>> Design Proposal for Monitoring of RDT Allocation Groups.
>>>
>>> -----------------------------------------------------------------------------
>>>
>>> Currently each CTRLGRP has a unique CLOSID and a (most likely) unique
>>> cache bitmask (CBM) per resource. Non-unique CBM are possible although
>>> useless. An unique CLOSID forbids more CTRLGRPs than physical CLOSIDs.
>>> CLOSIDs are much more scarce than RMIDs.
>>>
>>> If we lift the condition of unique CLOSID, then the user can create
>>> multiple CTRLGRPs with the same schemata. Internally, those CTRCGRP
>>> would share the CLOSID and RDT_Allocation must maintain the schemata
>>> to CLOSID relationship (similarly to what the previous CAT driver used
>>> to do). Elements in CTRLGRP.tasks and CTRLGRP.cpus behave the same as
>>> now: adding an element removes it from its previous CTRLGRP.
>>>
>>>
>>> This change would allow further partitioning the allocation groups
>>> into (allocation, monitoring) groups as follows:
>>>
>>> With allocation only:
>>>            CTRLGRP0     CTRLGRP_ALLOC_ONLY
>>> schemata:  L3:0=0xff0       L3:0=x00f
>>> tasks:       PID0       P0_0,P0_1,P1_0,P1_1
>>> cpus:        0x3                0xC
>>
>>
>> Not clear what the PID0 and P0_0 mean ?
>
> PID0, and P*_* are arbitrary PIDs. The tasks file works the same as it
> does now in RDT. I am not changing that.
>
>>
>> If you have to support something like MONGRP and CTRLGRP overall you want to
>> allow for a task to be present in multiple groups ?
>
> I am not proposing to support MONGRP and CTRLGRP. I am proposing to
> allow monitoring of CTRGRPs only.
>
>>>
>>> If we want to monitor (P0_0,P0_1), (P1_0,P1_1) and CPUs 0xC
>>> independently, with the new model we could create:
>>>            CTRLGRP0     CTRLGRP1     CTRLGRP2        CTRLGRP3
>>> schemata:  L3:0=0xff0   L3:0=x00f    L3:0=0x00f     L3:0=0x00f
>>> tasks:       PID0         <none>      P0_0,P0_1     P1_0, P1_1
>>> cpus:        0x3           0xC          0x0             0x0
>>>
>>> Internally, CTRLGRP1, CTRLGRP2, and CTRLGRP2 would share the CLOSID for
>>> (L3,0).
>>>
>>>
>>> Now we can ask perf to monitor any of the CTRLGRPs independently -once
>>> we solve how to pass to perf what (CTRLGRP, resource_id) to monitor-.
>>> The perf_event will reserve and assign the RMID to the monitored
>>> CTRLGRP. The RDT subsystem will context switch the whole PQR_ASSOC MSR
>>> (CLOSID and RMID), so perf won't have to.
>>
>>
>> This can be solved by suporting just the -t in perf and a new option in perf
>> to suport resctrl group monitoring (something similar to -R). That way we
>> provide the flexible granularity to monitor tasks independent of whether
>> they are in any resctrl group (and hence also a subset).
>
> One of the key points of my proposal is to remove monitoring PIDs
> independently. That simplifies things by letting RDT handle CLOSIDs
> and RMIDs together.
>
>>
>> CTRLGRP         TASKS           MASK
>> CTRLGRP1        PID1,PID2       L3:0=0Xf,1=0xf0
>> CTRLGRP2        PID3,PID4       L3:0=0Xf0,1=0xf00
>>
>> #perf stat -e llc_occupancy -R CTRLGRP1
>>
>> #perf stat -e llc_occupancy -t PID3,PID4
>>
>> The RMID allocation is independent of resctrl CLOSid allocation and hence
>> the RMID is not always married to CLOS which seems like the requirement
>> here.
>
> It is not a requirement. Both the CLOSID and the RMID of a CTRLGRP can
> change in my proposal.
>
>>
>> OR
>>
>> We could have CTRLGRPs with control_only, monitor_only or control_monitor
>> options.
>>
>> now a task could be present in both control_only and monitor_only
>> group or it could be present only in a control_monitor_group. The
>> transitions from one state to another are guarded by this same principle.
>>
>> CTRLGRP         TASKS           MASK                    TYPE
>> CTRLGRP1        PID1,PID2       L3:0=0Xf,1=0xf0         control_only
>> CTRLGRP2        PID3,PID4       L3:0=0Xf0,1=0xf00       control_only
>> CTRLGRP3        PID2,PID3                               monitor_only
>> CTRLGRP4        PID5,PID6       L3:0=0Xf0,1=0xf00       control_monitor
>>
>> CTRLGRP3 allows you to monitor a set of tasks which is not bound to be in
>> the same CTRLGRP and you can add or move tasks into this. The adding and
>> removing the tasks is whats easily supported compared to the task
>> granularity although such a thing could still be supported with the task
>> granularity.
>>
>> CTRLGRP4 allows you to tie the monitor and control together so when tasks
>> move in and out of this we still have that group to consider. And these
>> groups still retain the cpu masks like before so that cpu monitoring is
>> still supported.
>
> Instead of having 3 types of CTRLGRPl, I am proposing one kind
> (equivalent to your control_monitor type) that uses a non-zero RMID
> when an appropriate perf_event is attached to it. What advantages do
> you see on having 3 distinct types?

Basically I am trying to collage the requirements of what you are Stephan 
mentioned, Thomas and some other OEMs who only care about task monitoring 
(probably for a long time/life time so want them to be efficient etc).

To cover all the scenarios I see these may be the design requirements:

A group here is a group of tasks

1. To setup control groups and be able to monitor the control groups.
2. To be able to monitor the groups from the begining of task creation 
(lifetime or continuous)
3. To be able to monitor groups and not care about the CLOS(or allocation is 
'dont care'). Meaning the monitor groups may be a subset of the control groups 
or a superset or may intersect multiple control groups.
Or IOW , the task set here is arbitrary.
4. To monitor a task without bothering to create a group.

CTRLGRP         TASKS           MASK                    TYPE
CTRLGRP1        PID1,PID2       L3:0=0Xf,1=0xf0         control_only
CTRLGRP2        PID3,PID4       L3:0=0Xf0,1=0xf00       control_only
CTRLGRP3        PID2,PID3                               monitor_only
CTRLGRP4        PID5,PID6       L3:0=0Xf0,1=0xf00       control_monitor

now the implementation of this is say either through adding a new option to perf 
(reusing some of cgroup/or not etc which is implementation specific), or create 
a new tool - by tool i really mean an ioctl mechanism - so we just a syscall 
which can be called. Dont need a user mode tool per say. For ex I use resmon as 
the tool , but that is equivalent to having similar option in perf.

With this now if the user wants to do #1,

# resmon -R CTRLGRP1

#2

The resctrl group adds the children of the tasks in the same group. So we can 
essentially do continuous monitoring.

# echo $$ > ../ctrlgrp1/tasks
# resmon -R ctrlgrp1
# task1 &

#3 above is the situation where you want to first profile a bunch of tasks as to 
what the cache usage is (dont care about the CLOS) and then based on the usage , 
assign them or group them into control groups giving cache. This could be 
typical usage model in large scale cluster workload distribution or even in a 
real time workload.

Create the monitor_only groups for this like the CTRLGRP3 above

# echo PID1, ... PIDn > ../ctrlgrp3/tasks
# resmon -R ctrlgrp3

#4 above is when you say are already monitoring some tasks like PID2 and PID3 
which are part of a CTRLGRP but you want to monitor just PID2 now. Cannot create 
a new CTRLGRP with just PID2 if PID2 is already present in one of the CTRLGRPS 
(if we break this requirement , then it complicates the number of groups that 
can be created as we assume there is no hierarchy).

user can use the -t or task monitoring option to do this. Also the user can use 
this option without bothering to create any resctrl groups at all.

# resmon -t PID4, PID5, PID6

I think the email thread is going very long and we should just meet f2f probably 
next week to iron out the requirements and chalk out a design proposal.

overall it looks like we can do without recycling/rotation and just do reuse 
and throw error when run out of RMIDs, support per-package RMIDs. But need to 
add support for cpu monitoring and resctrl groups monitoring (without loosing 
the option to be able to have the flexibility to support monitoring at other 
granularities than the resctrl groups)

>
>>
>> In this case we would need a new option to support the ctrlgrp monitoring in
>> perf or a new tool to do all this if we dont want to bother perf.
>>
>>
>
> Agree, I like expanding the cgroup fd option to take CTRLGRP fds, as
> described in the Implementation Ideas part of the proposal.
>
>>>
>>> If CTRLGRP's schemata changes, the RDT subsystem will find a new
>>> CLOSID for the new schemata (potentially reusing an existing one) or
>>> fail (just like the old CAT used to). The RMID does not change during
>>> schemata updates.
>>>
>>> If a CTRLGRP dies, the monitoring perf_event continues to exists as a
>>> useless wraith, just as happens with cgroup events now.
>>>
>>> Since CTRLGRPs have no hierarchy. There is no need to handle that in
>>> the new RDT Monitoring PMU, greatly simplifying it over the previously
>>> proposed versions.
>>>
>>> A breaking change in user observed behavior with respect to the
>>> existing CQM PMU is that there wouldn't be task events. A task must be
>>> part of a CTRLGRP and events are created per (CTRLGRP, resource_id)
>>> pair. If an user wants to monitor a task across multiple resources
>>> (e.g. l3_occupancy across two packages), she must create one event per
>>> resource_id and add the two counts.
>>>
>>> I see this breaking change as an improvement, since hiding the cache
>>> topology to user space introduced lots of ugliness and complexity to
>>> the CQM PMU without improving accuracy over user space adding the
>>> events.
>>>
>>> Implementation ideas:
>>>
>>> First idea is to expose one monitoring file per resource in a CTRLGRP,
>>> so the list of CTRLGRP's files would be: schemata, tasks, cpus,
>>> monitor_l3_0, monitor_l3_1, ...
>>>
>>> the monitor_<resource_id> file descriptor is passed to perf_event_open
>>> in the way cgroup file descriptors are passed now. All events to the
>>> same (CTRLGRP,resource_id) share RMID.
>>>
>>> The RMID allocation part can either be handled by RDT Allocation or by
>>> the RDT Monitoring PMU. Either ways, the existence of PMU's
>>> perf_events allocates/releases the RMID.
>>>
>>> Also, since this new design removes hierarchy and task events, it
>>> allows for a simple solution of the RMID rotation problem. The removal
>>> of task events eliminates the cgroup vs task event conflict existing
>>> in the upstream version; it also removes the need to ensure that all
>>> active packages have RMIDs at the same time that added complexity to
>>> my version of CQM/CMT. Lastly, the removal of hierarchy removes the
>>> reliance on cgroups, the complex tree based read, and all the hooks
>>> and cgroup files that "raped" the cgroup subsystem.
>>
>>
>> Yes, not sure if the view is same after I sent the implementation details in
>> documentation :) (most likely it is).
>> But the option could be to not support perf_cgroup for cqm and support a new
>> option in perf to monitor resctrl groups and tasks (or some other options
>> like mongrp)
>
> Agree with no supporting cgroups. This proposal is about supporting
> neither cgroups nor tasks and do all monitoring through CTRLGRPs
> through an expansion of an existing perf option.
>
>>
>> I am so far inclined to creating a new monitoring interface that way we dont
>> try to "rape" the existing perf specifics for this RDT or later RDT
>> quirk/features.
>>
>
> On first inspection it seems to me like perf would be fine with this
> approach. It requires no changes to the system call and just some
> changes in the way the cgroup_fd is handled in perf_event_open
> (besides making sure that a context-less PMU don't break things). Do
> you foresee any conflict with future features?

The new tool would really add a new syscall instead of modifying the existing 
perf open syscall for the cgroup (pls see above)

Thanks,
Vikas

>
> Thanks,
> David
>

[toc] | [prev] | [next] | [standalone]


#1564847

FromThomas Gleixner <tglx@linutronix.de>
Date2017-01-23 11:20 +0100
Message-ID<t2JTc-59S-17@gated-at.bofh.it>
In reply to#1563892
On Fri, 20 Jan 2017, David Carrillo-Cisneros wrote:
> On Fri, Jan 20, 2017 at 5:29 AM Thomas Gleixner <tglx@linutronix.de> wrote:
> > Can you please write up in a abstract way what the design requirements are
> > that you need. So far we are talking about implementation details and
> > unspecfied wishlists, but what we really need is an abstract requirement.
> 
> My pleasure:
> 
> 
> Design Proposal for Monitoring of RDT Allocation Groups.

I was asking for requirements, not a design proposal. In order to make a
design you need a requirements specification.

So again: 

  Can please everyone involved write up their specific requirements
  for CQM and stop spamming us with half baken design proposals?

  And I mean abstract requirements and not again something which is
  referring to existing crap or some desired crap.

The complete list of requirements has to be agreed on before we talk about
anything else.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


#1564900

FromPeter Zijlstra <peterz@infradead.org>
Date2017-01-23 12:40 +0100
Message-ID<t2L8C-5R4-25@gated-at.bofh.it>
In reply to#1564847
On Mon, Jan 23, 2017 at 10:47:44AM +0100, Thomas Gleixner wrote:
> So again: 
> 
>   Can please everyone involved write up their specific requirements
>   for CQM and stop spamming us with half baken design proposals?
> 
>   And I mean abstract requirements and not again something which is
>   referring to existing crap or some desired crap.
> 
> The complete list of requirements has to be agreed on before we talk about
> anything else.

So something along the lines of:

 A) need to create a (named) group of tasks
   1) group composition needs to be dynamic; ie. we can add/remove member
      tasks at any time.
   2) a task can only belong to _one_ group at any one time.
   3) grouping need not be hierarchical?

 B) for each group, we need to set a CAT mask
   1) this CAT mask must be dynamic; ie. we can, during the existence of
      the group, change the mask at any time.

 C) for each group, we need to monitor CQM bits
   1) this monitor need not change


Supporting Use-Cases:

A.1: The Job (or VM) can have a dynamic task set
B.1: Dynamic QoS for each Job (or VM) as demand / load changes



Feel free to expand etc..

[toc] | [prev] | [next] | [standalone]


#1563910

FromShivappa Vikas <vikas.shivappa@intel.com>
Date2017-01-20 21:50 +0100
Message-ID<t1Oie-3yD-13@gated-at.bofh.it>
In reply to#1563354

On Thu, 19 Jan 2017, David Carrillo-Cisneros wrote:

> On Thu, Jan 19, 2017 at 6:32 PM, Vikas Shivappa
> <vikas.shivappa@linux.intel.com> wrote:
>> Resending including Thomas , also with some changes. Sorry for the spam
>>
>> Based on Thomas and Peterz feedback Can think of two design
>> variants which target:
>>
>> -Support monitoring and allocating using the same resctrl group.
>> user can use a resctrl group to allocate resources and also monitor
>> them (with respect to tasks or cpu)
>>
>> -Also allows monitoring outside of resctrl so that user can
>> monitor subgroups who use the same closid. This mode can be used
>> when user wants to monitor more than just the resctrl groups.
>>
>> The first design version uses and modifies perf_cgroup, second version
>> builds a new interface resmon.
>
> The second version would require to build a whole new set of tools,
> deploy them and maintain them. Users will have to run perf for certain
> events and resmon (or whatever is named the new tool) for rdt. I see
> it as too complex and much prefer to keep using perf.

This was so that we have the flexibility to align the tools as per the 
requirement of the feature rather than twisting the perf behaviour and also have 
that flexibility for future when new RDT features are added (something 
similar to what we did by introducing resctrl groups instead of using cgroups 
for CAT)

Sometimes thats a lot simpler as we dont need a lot code given the 
limited/specific syscalls we need to support. Just like the resctrl fs which is 
specific to RDT.

It looks like your requirement is to be able to monitor a group of tasks 
independently apart from the resctrl groups?

The task option should provide that flexibility to monitor a bunch of tasks 
independently apart from whether they are part of resctrl group or not. The 
assignment of RMID is contolled underneat by the kernel so we can optimize the 
usage of RMIDs and also RMIDs are tied to this group of tasks whether its a 
subset of resctrl group or not.

>
>> The first version is close to the patches
>> sent with some additions/changes. This includes details of the design as
>> per Thomas/Peterz feedback.
>>
>> 1> First Design option: without modifying the resctrl and using perf
>> --------------------------------------------------------------------
>> --------------------------------------------------------------------
>>
>> In this design everything in resctrl interface works like
>> before (the info, resource group files like task schemata all remain the
>> same)
>>
>>
>> Monitor cqm using perf
>> ----------------------
>>
>> perf can monitor individual tasks using the -t
>> option just like before.
>>
>> # perf stat -e llc_occupancy -t PID1,PID2
>>
>> user can monitor the cpu occupancy using the -C option in perf:
>>
>> # perf stat -e llc_occupancy -C 5
>>
>> Below shows how user can monitor cgroup occupancy:
>>
>> # mount -t cgroup -o perf_event perf_event /sys/fs/cgroup/perf_event/
>> # mkdir /sys/fs/cgroup/perf_event/g1
>> # mkdir /sys/fs/cgroup/perf_event/g2
>> # echo PID1 > /sys/fs/cgroup/perf_event/g2/tasks
>>
>> # perf stat -e intel_cqm/llc_occupancy/ -a -G g2
>>
>> To monitor a resctrl group, user can group the same tasks in resctrl
>> group into the cgroup.
>>
>> To monitor the tasks in p1 in example 2 below, add the tasks in resctrl
>> group p1 to cgroup g1
>>
>> # echo 5678 > /sys/fs/cgroup/perf_event/g1/tasks
>>
>> Introducing a new option for resctrl may complicate monitoring because
>> supporting cgroup 'task groups' and resctrl 'task groups' leads to
>> situations where:
>> if the groups intersect, then there is no way to know what
>> l3_allocations contribute to which group.
>>
>> ex:
>> p1 has tasks t1, t2, t3
>> g1 has tasks t2, t3, t4
>>
>> The only way to get occupancy for g1 and p1 would be to allocate an RMID
>> for each task which can as well be done with the -t option.
>
> That's simply recreating the resctrl group as a cgroup.
>
> I think that the main advantage of doing allocation first is that we
> could use the context switch in rdt allocation and greatly simplify
> the pmu side of it.
>
> If resctrl groups could lift the restriction of one resctl per CLOSID,
> then the user can create many resctrl in the way perf cgroups are
> created now. The advantage is that there wont be cgroup hierarchy!
> making things much simpler. Also no need to optimize perf event
> context switch to make llc_occupancy work.
>
> Then we only need a way to express that monitoring must happen in a
> resctl to the perf_event_open syscall.
>
> My first thought is to have a "rdt_monitor" file per resctl group. A
> user passes it to perf_event_open in the way cgroups are passed now.
> We could extend the meaning of the flag PERF_FLAG_PID_CGROUP to also
> cover rdt_monitor files. The syscall can figure if it's a cgroup or a
> rdt_group. The rdt_monitoring PMU would only work with rdt_monitor
> groups
>
> Then the rdm_monitoring PMU will be pretty dumb, having neither task
> nor CPU contexts. Just providing the pmu->read and pmu->event_init
> functions.
>
> Task monitoring can be done with resctrl as well by adding the PID to
> a new resctl and opening the event on it. And, since we'd allow CLOSID
> to be shared between resctrl groups, allocation wouldn't break.

It looks like we are trying to create a MONGRP and CTRLGRP like Thomas mentions.

Although resctrl group now does not have a hierarchy a task can be part of 
only one group - breaking this is just equivalent to having a seperate resmon 
group which may group the tasks independent of how they are grouped in the 
resctrl group?

That can be achieved as well with the option to monitor at task granularity ? 
that means if we support task option and the option to monitor resctrl groups we 
obtain the same functionality.

[toc] | [prev] | [next] | [standalone]


#1563884

FromStephane Eranian <eranian@google.com>
Date2017-01-20 20:40 +0100
Message-ID<t1Ncu-2V1-19@gated-at.bofh.it>
In reply to#1563241
On Thu, Jan 19, 2017 at 6:32 PM, Vikas Shivappa
<vikas.shivappa@linux.intel.com> wrote:
>
> Resending including Thomas , also with some changes. Sorry for the spam
>
> Based on Thomas and Peterz feedback Can think of two design
> variants which target:
>
> -Support monitoring and allocating using the same resctrl group.
> user can use a resctrl group to allocate resources and also monitor
> them (with respect to tasks or cpu)
>
> -Also allows monitoring outside of resctrl so that user can
> monitor subgroups who use the same closid. This mode can be used
> when user wants to monitor more than just the resctrl groups.
>
> The first design version uses and modifies perf_cgroup, second version
> builds a new interface resmon. The first version is close to the patches
> sent with some additions/changes. This includes details of the design as
> per Thomas/Peterz feedback.
>
> 1> First Design option: without modifying the resctrl and using perf
> --------------------------------------------------------------------
> --------------------------------------------------------------------
>
> In this design everything in resctrl interface works like
> before (the info, resource group files like task schemata all remain the
> same)
>
>
> Monitor cqm using perf
> ----------------------
>
> perf can monitor individual tasks using the -t
> option just like before.
>
> # perf stat -e llc_occupancy -t PID1,PID2
>
> user can monitor the cpu occupancy using the -C option in perf:
>
> # perf stat -e llc_occupancy -C 5
>
> Below shows how user can monitor cgroup occupancy:
>
> # mount -t cgroup -o perf_event perf_event /sys/fs/cgroup/perf_event/
> # mkdir /sys/fs/cgroup/perf_event/g1
> # mkdir /sys/fs/cgroup/perf_event/g2
> # echo PID1 > /sys/fs/cgroup/perf_event/g2/tasks
>
> # perf stat -e intel_cqm/llc_occupancy/ -a -G g2
>
Presented this way, this  does not quite address the use case I
described earlier here.
We want to be able to monitor the cgroup allocations from first thread
creation. What you have above has a large gap. Many apps do allocations
as their very first steps, so if you do:
$ my_test_prg &
[1456]
$ echo 1456 >/sys/fs/cgroup/perf_event/g2/tasks
$ perf stat -e intel_cqm/llc_occupancy/ -a -G g2

You have a race. But if you allow:

$ perf stat -e intel_cqm/llc_occupancy/ -a -G g2 (i.e, on an empty cgroup)
$ echo $$ >/sys/fs/cgroup/perf_event/g2/tasks (put shell in cgroup, so
my_test_prg runs immediately in the cgroup)
$ my_test_prg &

Then there is a way to avoid the gap.

>
> To monitor a resctrl group, user can group the same tasks in resctrl
> group into the cgroup.
>
> To monitor the tasks in p1 in example 2 below, add the tasks in resctrl
> group p1 to cgroup g1
>
> # echo 5678 > /sys/fs/cgroup/perf_event/g1/tasks
>
> Introducing a new option for resctrl may complicate monitoring because
> supporting cgroup 'task groups' and resctrl 'task groups' leads to
> situations where:
> if the groups intersect, then there is no way to know what
> l3_allocations contribute to which group.
>
> ex:
> p1 has tasks t1, t2, t3
> g1 has tasks t2, t3, t4
>
> The only way to get occupancy for g1 and p1 would be to allocate an RMID
> for each task which can as well be done with the -t option.
>
> Monitoring cqm cgroups Implementation
> -------------------------------------
>
> When monitoring two different cgroups in the same hierarchy (ex say g11
> has an ancestor g1 which are both being monitored as shown below) we
> need the g11 counts to be considered for g1 as well.
>
> # mount -t cgroup -o perf_event perf_event /sys/fs/cgroup/perf_event/
> # mkdir /sys/fs/cgroup/perf_event/g1
> # mkdir /sys/fs/cgroup/perf_event/g1/g11
>
> When measuring for g1 llc_occupancy we cannot write two different RMIDs
> (because we need to count for g11 as well)
> during context switch to measure the occupancy for both g1 and g11.
> Hence the driver maintains this information and writes the RMID of the
> lowest member in the ancestory which is being monitored during ctx
> switch.
>
> The cqm_info is added to the perf_cgroup structure to maintain this
> information. The structure is allocated and destroyed at css_alloc and
> css_free. All the events tied to a cgroup can use the same
> information while reading the counts.
>
> struct perf_cgroup {
> #ifdef CONFIG_INTEL_RDT_M
>         void *cqm_info;
> #endif
> ...
>
>  }
>
> struct cqm_info {
>   bool mon_enabled;
>   int level;
>   u32 *rmid;
>   struct cgrp_cqm_info *mfa;
>   struct list_head tskmon_rlist;
>  };
>
> Due to the hierarchical nature of cgroups, every cgroup just
> monitors for the 'nearest monitored ancestor' at all times.
> Since root cgroup is always monitored, all descendents
> at boot time monitor for root and hence all mfa points to root
> except for root->mfa which is NULL.
>
> 1. RMID setup: When cgroup x start monitoring:
>    for each descendent y, if y's mfa->level < x->level, then
>    y->mfa = x. (Where level of root node = 0...)
> 2. sched_in: During sched_in for x
>    if (x->mon_enabled) choose x->rmid
>      else choose x->mfa->rmid.
> 3. read: for each descendent of cgroup x
>    if (x->monitored) count += rmid_read(x->rmid).
> 4. evt_destroy: for each descendent y of x, if (y->mfa == x)
>    then y->mfa = x->mfa. Meaning if any descendent was monitoring for
>    x, set that descendent to monitor for the cgroup which x was
>    monitoring for.
>
> To monitor a task in a cgroup x along with monitoring cgroup x itself
> cqm_info maintains a list of tasks that are being monitored in the
> cgroup.
>
> When a task which belongs to a cgroup x is being monitored, it
> always uses its own task->rmid even if cgroup x is monitored during sched_in.
> To account for the counts of such tasks, cgroup keeps this list
> and parses it during read.
> taskmon_rlist is used to maintain the list. The list is modified when a
> task is attached to the cgroup or removed from the group.
>
> Example 1 (Some examples modeled from resctrl ui documentation)
> ---------
>
> A single socket system which has real-time tasks running on core 4-7 and
> non real-time workload assigned to core 0-3. The real-time tasks share
> text and data, so a per task association is not required and due to
> interaction with the kernel it's desired that the kernel on these cores shares L3
> with the tasks.
>
> # cd /sys/fs/resctrl
>
> # echo "L3:0=3ff" > schemata
>
> core 0-1 are assigned to the new group and make sure that the
> kernel and the tasks running there get 50% of the cache.
>
> # echo 03 > p0/cpus
>
> monitor the cpus 0-1
>
> # perf stat -e llc_occupancy -C 0-1
>
> Example 2
> ---------
>
> A real time task running on cpu 2-3(socket 0) is allocated a dedicated 25% of the
> cache.
>
> # cd /sys/fs/resctrl
>
> # mkdir p1
> # echo "L3:0=0f00;1=ffff" > p1/schemata
> # echo 5678 > p1/tasks
> # taskset -cp 2-3 5678
>
> To monitor the same group of tasks create a cgroup g1
>
> # mount -t cgroup -o perf_event perf_event /sys/fs/cgroup/perf_event/
> # mkdir /sys/fs/cgroup/perf_event/g1
> # perf stat -e llc_occupancy -a -G g1
>
> Example 3
> ---------
>
> sometimes user may just want to profile the cache occupancy first before
> assigning any CLOSids. Also this provides an override option where user
> can monitor some tasks which have say CLOS 0 that he is about to place
> in a CLOSId based on the amount of cache occupancy. This could apply to
> the same real time tasks above where user is caliberating the % of cache
> thats needed.
>
> # perf stat -e llc_occupancy -t PIDx,PIDy
>
> RMID allocation
> ---------------
>
> RMIDs are allocated per package to achieve better scaling of RMIDs.
> RMIDs are plenty (2-4 per logical processor) and also are per package
> meaning a two socket system would have twice the number of RMIDs.
> If we still run out of RMIDs an error is thrown that monitoring wasnt
> possible as the RMID wasnt available.
>
> Kernel Scheduling
> -----------------
>
> During ctx switch cqm choses the RMID in the following priority
>
> 1. if cpu has a RMID , choose that
> 2. if the task has a RMID directly tied to it choose that (task is
>    monitored)
> 3. choose the RMID of the task's cgroup (by default tasks belong to root
>    cgroup with RMID 0)
>
> Read
> ----
>
> When user calls cqm to retrieve the monitored count, we read the
> counter_msr and return the count. For cgroup hierarcy , the count is
> measured as explained in the cgroup implementation section by traversing
> the cgroup hierarchy.
>
>
> 2> Second Design option: Build a new usermode tool resmon
> ---------------------------------------------------------
> ---------------------------------------------------------
>
> In this design everything in resctrl interface works like
> before (the info, resource group files like task schemata all remain the
> same).
>
> This version supports monitoring resctrl groups directly.
> But we need a user interface for the user to read the counters.  We can
> create one file to set monitoring and one
> file in resctrl directory which will reflect the counts but may not be
> efficient as a lot of times user reads the counts frequently.
>
> Build a new user mode interface resmon
> --------------------------------------
>
> Since modifying the existing perf to
> suit the different h/w architecture seems to not follow the CAT
> interface model, it may well be better to have a different and dedicated
> interface for the RDT monitoring (just like we had a new fs for CAT)
>
> resmon supports monitoring a resctrl group or a task. The two modes may
> provide enough granularity needed for monitoring
> -can monitor cpu data.
> -can monitor per resctrl group data.
> -can choose custom or subset of tasks with in a resctrl group and monitor.
>
> # resmon [<options>]
> -r <resctrl group>
> -t <PID>
> -s <mon_mask>
> -I <time in ms>
>
> "resctrl group": is the resctrl directory.
>
> "mon_mask:  is a bit mask of logical packages which indicates which packages user is
> interested in monitoring.
>
> "time in ms": The time for which the monitoring takes place
> (this can potentially be changed to start and stop/read options)
>
> Example 1 (Some examples modeled from resctrl ui documentation)
> ---------
>
> A single socket system which has real-time tasks running on core 4-7 and
> non real-time workload assigned to core 0-3. The real-time tasks share
> text and data, so a per task association is not required and due to
> interaction with the kernel it's desired that the kernel on these cores shares L3
> with the tasks.
>
> # cd /sys/fs/resctrl
> # mkdir p0
> # echo "L3:0=3ff" > p0/schemata
>
> core 0-1 are assigned to the new group and make sure that the
> kernel and the tasks running there get 50% of the cache.
>
> # echo 03 > p0/cpus
>
> monitor the cpus 0-1 for 10s.
>
> # resmon -r p0 -s 1 -I 10000
>
> Example 2
> ---------
>
> A real time task running on cpu 2-3(socket 0) is allocated a dedicated 25% of the
> cache.
>
> # cd /sys/fs/resctrl
>
> # mkdir p1
> # echo "L3:0=0f00;1=ffff" > p1/schemata
> # echo 5678 > p1/tasks
> # taskset -cp 2-3 5678
>
> Monitor the task for 5s on socket zero
>
> # resmon -r p1 -s 1 -I 5000
>
> Example 3
> ---------
>
> sometimes user may just want to profile the cache occupancy first before
> assigning any CLOSids. Also this provides an override option where user
> can monitor some tasks which have say CLOS 0 that he is about to place
> in a CLOSId based on the amount of cache occupancy. This could apply to
> the same real time tasks above where user is caliberating the % of cache
> thats needed.
>
> # resmon -t PIDx,PIDy -s 1 -I 10000
>
> returns the sum of count of PIDx and PIDy
>
> RMID Allocation
> ---------------
>
> This would remain the same like design version 1, where we support per
> package RMIDs and throw error when out of RMIDs due to h/w limited
> RMIDs.
>
>

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web