Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1571948 > unrolled thread
| Started by | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| First post | 2017-02-01 21:10 +0100 |
| Last post | 2017-02-02 19:00 +0100 |
| Articles | 20 on this page of 28 — 6 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Luck, Tony" <tony.luck@intel.com> - 2017-02-01 21:10 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-02-02 00:20 +0100
RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Luck, Tony" <tony.luck@intel.com> - 2017-02-02 18:40 +0100
RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Luck, Tony" <tony.luck@intel.com> - 2017-02-02 20:40 +0100
RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-02-02 21:20 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-02-02 21:30 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Luck, Tony" <tony.luck@intel.com> - 2017-02-03 00:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-02-03 02:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-02-03 03:20 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Luck, Tony" <tony.luck@intel.com> - 2017-02-03 19:00 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-02-03 22:10 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Luck, Tony" <tony.luck@intel.com> - 2017-02-03 23:30 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Stephane Eranian <eranian@google.com> - 2017-02-07 09:10 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Luck, Tony" <tony.luck@intel.com> - 2017-02-07 20:10 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Stephane Eranian <eranian@google.com> - 2017-02-08 22:40 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-02-07 21:30 +0100
RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Luck, Tony" <tony.luck@intel.com> - 2017-02-06 20:00 +0100
RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Luck, Tony" <tony.luck@intel.com> - 2017-02-06 22:30 +0100
RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-02-06 22:40 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-02-06 22:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-02-06 23:20 +0100
RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Luck, Tony" <tony.luck@intel.com> - 2017-02-07 00:30 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-02-07 01:40 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Andi Kleen <andi@firstfloor.org> - 2017-02-02 01:40 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Andi Kleen <andi@firstfloor.org> - 2017-02-02 02:20 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-02-02 02:20 +0100
RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Yu, Fenghua" <fenghua.yu@intel.com> - 2017-02-02 02:30 +0100
RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-02-02 19:00 +0100
Page 1 of 2 [1] 2 Next page →
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2017-02-01 21:10 +0100 |
| Subject | RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes |
| Message-ID | <t69o5-26I-15@gated-at.bofh.it> |
> I was asking for requirements, not a design proposal. In order to make a > design you need a requirements specification. Here's what I came up with ... not a fully baked list, but should allow for some useful discussion on whether any of these are not really needed, or if there is a glaring hole that misses some use case: 1) Able to measure using all supported events (currently L3 occupancy, Total B/W, Local B/W) 2) Measure per thread 3) Including kernel threads 4) Put multiple threads into a single measurement group (forced by h/w shortage of RMIDs, but probably good to have anyway) 5) New threads created inherit measurement group from parent 6) Report separate results per domain (L3) 7) Must be able to measure based on existing resctrl CAT group 8) Can get measurements for subsets of tasks in a CAT group (to find the guys hogging the resources) 9) Measure per logical CPU (pick active RMID in same precedence for task/cpu as CAT picks CLOSID) 10) Put multiple CPUs into a group Nice to have: 1) Readout using "perf(1)" [subset of modes that make sense ... tying monitoring to resctrl file system will make most command line usage of perf(1) close to impossible. -Tony
[toc] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-02-02 00:20 +0100 |
| Message-ID | <t6clX-4e8-3@gated-at.bofh.it> |
| In reply to | #1571948 |
On Wed, Feb 1, 2017 at 12:08 PM Luck, Tony <tony.luck@intel.com> wrote: > > > I was asking for requirements, not a design proposal. In order to make a > > design you need a requirements specification. > > Here's what I came up with ... not a fully baked list, but should allow for some useful > discussion on whether any of these are not really needed, or if there is a glaring hole > that misses some use case: > > 1) Able to measure using all supported events (currently L3 occupancy, Total B/W, Local B/W) > 2) Measure per thread > 3) Including kernel threads > 4) Put multiple threads into a single measurement group (forced by h/w shortage of RMIDs, but probably good to have anyway) Even with infinite hw RMIDs you want to be able to have one RMID per thread groups to avoid reading a potentially large list of RMIDs every time you measure one group's event (with the delay and error associated to measure many RMIDs whose values fluctuate rapidly). > 5) New threads created inherit measurement group from parent > 6) Report separate results per domain (L3) > 7) Must be able to measure based on existing resctrl CAT group > 8) Can get measurements for subsets of tasks in a CAT group (to find the guys hogging the resources) > 9) Measure per logical CPU (pick active RMID in same precedence for task/cpu as CAT picks CLOSID) I agree that "Measure per logical CPU" is a requirement, but why is "pick active RMID in same precedence for task/cpu as CAT picks CLOSID" one as well? Are we set on handling RMIDs the way CLOSIDs are handled? there are drawbacks to do so, one is that it would make impossible to do CPU monitoring and CPU filtering the way is done for all other PMUs. i.e. the following commands (or their equivalent in whatever other API you create) won't work: a) perf stat -e intel_cqm/total_bytes/ -C 2 or b.1) perf stat -e intel_cqm/total_bytes/ -C 2 <a_measurement_group> or b.2) perf stat -e intel_cqm/llc_occupancy/ -a <a_measurement_group> in (a) because many RMIDs may run in the CPU and, in (b's) because the same measurement group's RMID will be used across all CPUs. I know this is similar to how it is in CAT, but CAT was never intended to do monitoring. We can do the CAT way and the perf way, or not, but if we will drop support for perf's like CPU support, it must be explicitly stated and not an implicit consequence of a design choice leaked into requirements. > 10) Put multiple CPUs into a group 11) Able to measure across CAT groups. So that a user can: A) measure a task that runs on CPUs that are in different CAT groups (one of Thomas' use case FWICT), and B) measure tasks even if they change their CAT group (my use case). > > Nice to have: > 1) Readout using "perf(1)" [subset of modes that make sense ... tying monitoring to resctrl file system will make most command line usage of perf(1) close to impossible. We discussed this offline and I still disagree that it is close to impossible to use perf and perf_event_open. In fact, I think it's very simple : a) We stretch the usage of the pid parameter in perf_event_open to also allow a PMU specific task group fd (as of now it's either a PID or a cgroup fd). b) PMUs that can handle non-cgroup task groups have a special PMU_CAP flag to signal the generic code to not resolve the fd to a cgroup pointer and, instead, save it as is in struct perf_event (a few lines of code). c) The PMU takes care of resolving the task group's fd. The above is ONE way to do it, there may be others. But there is a big advantage on leveraging perf_event_open and ease integration with the perf tool and the myriads of tools that use the perf API. 12) Whatever fs or syscall is provided instead of perf syscalls, it should provide total_time_enabled in the way perf does, otherwise is hard to interpret MBM values. > > -Tony > >
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2017-02-02 18:40 +0100 |
| Message-ID | <t6twt-6ZQ-9@gated-at.bofh.it> |
| In reply to | #1572043 |
>> 7) Must be able to measure based on existing resctrl CAT group
>> 8) Can get measurements for subsets of tasks in a CAT group (to find the guys hogging the resources)
>> 9) Measure per logical CPU (pick active RMID in same precedence for task/cpu as CAT picks CLOSID)
>
> I agree that "Measure per logical CPU" is a requirement, but why is
> "pick active RMID in same precedence for task/cpu as CAT picks CLOSID"
> one as well? Are we set on handling RMIDs the way CLOSIDs are
> handled? there are drawbacks to do so, one is that it would make
> impossible to do CPU monitoring and CPU filtering the way is done for
> all other PMUs.
I'm too focused on monitoring existing CAT groups. If we move the parenthetical remark
from item 9, to item 7, then I think it is better. When monitoring a CAT group we need to
monitor exactly the processes that are controlled by the CAT group. So RMID must match
CLOSID, and the precedence rules make that work.
For other monitoring cases we can do things differently - so long as we have a way
to express what we want, and we don't pile a ton of code into context switch to figure
out which RMID is to be loaded into PQR_ASSOC.
I thought of another requirement this morning:
N+1) When we set up monitoring we must allocate all the resources we need (or fail the setup
if we can't get them). Not allowed to error in the middle of monitoring because we can't
find a free RMID)
-Tony
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2017-02-02 20:40 +0100 |
| Message-ID | <t6voC-8cY-23@gated-at.bofh.it> |
| In reply to | #1572043 |
>> Nice to have:
>> 1) Readout using "perf(1)" [subset of modes that make sense ... tying monitoring
>> to resctrl file system will make most command line usage of perf(1) close to impossible.
>
>
> We discussed this offline and I still disagree that it is close to
> impossible to use perf and perf_event_open. In fact, I think it's very
> simple :
Maybe s/most/many/ ?
The issue here is that we are going to define which tasks and cpus are being
monitored *outside* of the perf command. So usage like:
# perf stat -I 1000 -e intel_cqm/llc_occupancy {command}
are completely out of scope ... we aren't planning to change the perf(1)
command to know about creating a CQM monitor group, assigning the
PID of {command} to it, and then report on llc_occupancy.
So perf(1) usage is only going to support modes where it attaches to some
monitor group that was previously established. The "-C 2" option to monitor
CPU 2 is certainly plausible ... assuming you set up a monitor group to track
what is happening on CPU 2 ... I just don't know how perf(1) would know the
name of that group.
Vikas is pushing for "-R rdtgroup" ... though our offline discussions included
overloading "-g" and have perf(1) pick appropriately from cgroups or rdtgroups
depending on event type.
-Tony
[toc] | [prev] | [next] | [standalone]
| From | Shivappa Vikas <vikas.shivappa@intel.com> |
|---|---|
| Date | 2017-02-02 21:20 +0100 |
| Message-ID | <t6w1k-eC-23@gated-at.bofh.it> |
| In reply to | #1572704 |
Hello Peterz/Andi, On Thu, 2 Feb 2017, Luck, Tony wrote: >>> Nice to have: >>> 1) Readout using "perf(1)" [subset of modes that make sense ... tying monitoring >>> to resctrl file system will make most command line usage of perf(1) close to impossible. >> > Vikas is pushing for "-R rdtgroup" ... though our offline discussions included > overloading "-g" and have perf(1) pick appropriately from cgroups or rdtgroups > depending on event type. Assume we build support to monitor the existing resctrl CAT groups like Thomas suggested. For the perf interface would something like below seems reasonable or a disaster(given that we have a new -R option specific to the PMU/which works only on this PMU) ? # mount -t resctrl resctrl /sys/fs/resctrl # cd /sys/fs/resctrl # mkdir p0 p1 # echo "L3:0=3;1=c" > /sys/fs/resctrl/p0/schemata # echo "L3:0=3;1=3" > /sys/fs/resctrl/p1/schemata Now monitor the group p1 using perf. perf would have a new option -R to monitor the resctrl groups. perf would still have a cqm event like today intel_cqm/llc_occupancy which supports however only one mode -R and not any of -C,-t,-G etc. So pretty much the -R works like a -G .. except that it works on the resctrl fs and not perf_cgroup. PMU would have a flag to indicate the perf user mode to check only the llc_occupancy event is supported for the -R. # perf stat -e intel_cqm/llc_occupancy -R p1 -Vikas > > -Tony >
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-02-02 21:30 +0100 |
| Message-ID | <t6wb0-hZ-17@gated-at.bofh.it> |
| In reply to | #1572704 |
On Thu, Feb 2, 2017 at 11:33 AM, Luck, Tony <tony.luck@intel.com> wrote:
>>> Nice to have:
>>> 1) Readout using "perf(1)" [subset of modes that make sense ... tying monitoring
>>> to resctrl file system will make most command line usage of perf(1) close to impossible.
>>
>>
>> We discussed this offline and I still disagree that it is close to
>> impossible to use perf and perf_event_open. In fact, I think it's very
>> simple :
>
> Maybe s/most/many/ ?
>
> The issue here is that we are going to define which tasks and cpus are being
> monitored *outside* of the perf command. So usage like:
>
> # perf stat -I 1000 -e intel_cqm/llc_occupancy {command}
>
> are completely out of scope ... we aren't planning to change the perf(1)
> command to know about creating a CQM monitor group, assigning the
> PID of {command} to it, and then report on llc_occupancy.
>
> So perf(1) usage is only going to support modes where it attaches to some
> monitor group that was previously established. The "-C 2" option to monitor
> CPU 2 is certainly plausible ... assuming you set up a monitor group to track
> what is happening on CPU 2 ... I just don't know how perf(1) would know the
> name of that group.
There is no need to change perf(1) to support
# perf stat -I 1000 -e intel_cqm/llc_occupancy {command}
the PMU can work with resctrl to provide the support through
perf_event_open, with the advantage that tools other than perf could
also use it.
I'd argue is more stable and has less corner cases if the
task_mongroups get extra RMIDs for the task events attached to them
than having userspace tools create and destroy groups and move tasks
behind the scenes.
I provided implementation details on the write-up I shared offline on
Monday. If "easy monitoring" of stand-alone task becomes a
requirement, we can dig on the pros and cons of implementing in kernel
vs user space.
>
> Vikas is pushing for "-R rdtgroup" ... though our offline discussions included
> overloading "-g" and have perf(1) pick appropriately from cgroups or rdtgroups
> depending on event type.
I see it more like generalizing the -G option to represent a task
group that can be a cgroup or a PMU specific one.
Currently the perf(1) simply translates the argument of the -G option
into a file descriptor. My idea doesn't change that, just makes perf
tool to look for a "task_group_root" file in the PMU folder and use it
to find as base path for the file descriptor. If a PMU doesnt have
such file, then perf(1) uses the perf cgroup mounting point, as it
does now. That makes for a very simple implementation on the perf tool
side.
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2017-02-03 00:50 +0100 |
| Message-ID | <t6ziy-2dQ-17@gated-at.bofh.it> |
| In reply to | #1572733 |
On Thu, Feb 02, 2017 at 12:22:42PM -0800, David Carrillo-Cisneros wrote:
> There is no need to change perf(1) to support
> # perf stat -I 1000 -e intel_cqm/llc_occupancy {command}
>
> the PMU can work with resctrl to provide the support through
> perf_event_open, with the advantage that tools other than perf could
> also use it.
I agree it would be better to expose the counters through
a standard perf_event_open() interface ... but we don't seem
to have had much luck doing that so far.
That would need the requirements to be re-written with the
focus of what does resctrl need to do to support each of the
perf(1) command line modes of operation. The fact that these
counters work rather differently from normal h/w counters
has resulted in massively complex volumes of code trying
to map them into what perf_event_open() expects.
The key points of weirdness seem to be:
1) We need to allocate an RMID for the duration of monitoring. While
there are quite a lot of RMIDs, it is easy to envision scenarios
where there are not enough.
2) We need to load that RMID into PQR_ASSOC on a logical CPU whenever a process
of interest is running.
3) An RMID is shared by llc_occupancy, local_bytes and total_bytes events
4) For llc_occupancy the count can change even when none of the processes
are running becauase cache lines are evicted
5) llc_occupancy measures the delta, not the absolute occupancy. To
get a good result requires monitoring from process creation (or
lots of patience, or the nuclear option "wbinvd").
6) RMID counters are package scoped
These result in all sorts of hard to resolve situations. E.g. you are
monitoring local bandwidth coming from logical CPU2 using RMID=22. I'm
looking at the cache occupancy of PID=234 using RMID=45. The scheduler
decides to run my proocess on your CPU. We can only load one RMID, so
one of us will be disappointed (unless we have some crazy complex code
where your instance of perf borrows RMID=45 and reads out the local
byte count on sched_in() and sched_out() to add to the runing count
you were keeping against RMID=22).
How can we document such restrictions for people who haven't been
digging in this code for over a year?
I think a perf_event_open() interface would make some simple cases
work, but result in some swearing once people start running multiple
complex monitors at the same time.
-Tony
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-02-03 02:50 +0100 |
| Message-ID | <t6BaF-3o4-1@gated-at.bofh.it> |
| In reply to | #1572814 |
On Thu, Feb 2, 2017 at 3:41 PM, Luck, Tony <tony.luck@intel.com> wrote:
> On Thu, Feb 02, 2017 at 12:22:42PM -0800, David Carrillo-Cisneros wrote:
>> There is no need to change perf(1) to support
>> # perf stat -I 1000 -e intel_cqm/llc_occupancy {command}
>>
>> the PMU can work with resctrl to provide the support through
>> perf_event_open, with the advantage that tools other than perf could
>> also use it.
>
> I agree it would be better to expose the counters through
> a standard perf_event_open() interface ... but we don't seem
> to have had much luck doing that so far.
>
> That would need the requirements to be re-written with the
> focus of what does resctrl need to do to support each of the
> perf(1) command line modes of operation. The fact that these
> counters work rather differently from normal h/w counters
> has resulted in massively complex volumes of code trying
> to map them into what perf_event_open() expects.
>
> The key points of weirdness seem to be:
>
> 1) We need to allocate an RMID for the duration of monitoring. While
> there are quite a lot of RMIDs, it is easy to envision scenarios
> where there are not enough.
>
> 2) We need to load that RMID into PQR_ASSOC on a logical CPU whenever a process
> of interest is running.
>
> 3) An RMID is shared by llc_occupancy, local_bytes and total_bytes events
>
> 4) For llc_occupancy the count can change even when none of the processes
> are running becauase cache lines are evicted
>
> 5) llc_occupancy measures the delta, not the absolute occupancy. To
> get a good result requires monitoring from process creation (or
> lots of patience, or the nuclear option "wbinvd").
>
> 6) RMID counters are package scoped
>
>
> These result in all sorts of hard to resolve situations. E.g. you are
> monitoring local bandwidth coming from logical CPU2 using RMID=22. I'm
> looking at the cache occupancy of PID=234 using RMID=45. The scheduler
> decides to run my proocess on your CPU. We can only load one RMID, so
> one of us will be disappointed (unless we have some crazy complex code
> where your instance of perf borrows RMID=45 and reads out the local
> byte count on sched_in() and sched_out() to add to the runing count
> you were keeping against RMID=22).
>
> How can we document such restrictions for people who haven't been
> digging in this code for over a year?
>
> I think a perf_event_open() interface would make some simple cases
> work, but result in some swearing once people start running multiple
> complex monitors at the same time.
More problems:
7) Time multiplexing of RMIDs is hard because llc_occupancy cannot be reset.
8) Only one RMID per CPU can be loaded at a time into PQR_ASSOC.
Most of the complexity in past attempts were mainly caused by:
A. Task events being defined as system-wide and not package-wide.
What you describe in points (4) and (6) made this complicated.
B. The cgroup hierarchy, due to (7) and (8).
A and B caused the bulk of the code by complicating RMID assignment,
reading and rotation.
Now that we've learned from the past experience, we have defined
per-domain monitoring and use flat groups. FWICT, that enough to allow
a simple implementation that can be expressed through perf_event_open.
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-02-03 03:20 +0100 |
| Message-ID | <t6BDH-3Nz-1@gated-at.bofh.it> |
| In reply to | #1572843 |
Something to be aware is that CAT cpus don't work the way CPU filtering works in perf: If I have the following CAT groups: - default group with task TD - group GC1 with CPU0 and CLOSID 1 - group GT1 with no CPUs and task T1 and CLOSID2 - TD and T1 run on CPU0. Then T1 will use CLOSID2 and TD CLOSID1. Some allocations done in CPU0 did not use CLOSID1. Now, if I have the same setup in monitoring groups and I were to read llc_occupancy in the RMID of GC1, I'd read llc_occupancy for TD only, and have a blind spot on T1. That's not how CPU events work on perf. So CPUs have a different meaning on CAT than on perf. The above is another reason to separate the allocation and the monitoring groups. Having - Independent allocation and monitoring groups. - Independent CPU and task grouping. would allow us semantics that monitor CAT groups and eventually can be extended to also monitor the perf way, this is support: - filter by task - filter by task group (cgroup or monitoring group or whatever). - filter by CPU (the perf way) - combinations of task/task_group and CPU (the perf way) If we tie allocation groups and monitoring groups, we are tying the meaning of CPUs and we'll have to choose between the CAT meaning or the perf meaning. Let's allow semantics that will allow perf like monitoring to eventually work, even if its not immediately supported. Thanks, David
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2017-02-03 19:00 +0100 |
| Message-ID | <t6Qjp-4OT-35@gated-at.bofh.it> |
| In reply to | #1572854 |
On Thu, Feb 02, 2017 at 06:14:05PM -0800, David Carrillo-Cisneros wrote: > If we tie allocation groups and monitoring groups, we are tying the > meaning of CPUs and we'll have to choose between the CAT meaning or > the perf meaning. > > Let's allow semantics that will allow perf like monitoring to > eventually work, even if its not immediately supported. Would it work to make monitor groups be "task list only" or "cpu mask only" (unlike control groups that allow mixing). Then the intel_rdt_sched_in() code could pick the RMID in ways that give you the perf(1) meaning. I.e. if you create a monitor group and assign some CPUs to it, then we will always load the RMID for that monitor group when running on those cpus, regardless of what group(s) the current process belongs to. But if you didn't create any cpu-only monitor groups, then we'd assign RMID using same rules as CLOSID (so measurements from a control group would track allocation policies). We are already planning that creating monitor only groups will change what is reported in the control group (e.g. you pull some tasks out of the control group to monitor them separately, so the control group only reports the tasks that you didn't move out for monitoring). -Tony
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-02-03 22:10 +0100 |
| Message-ID | <t6Thf-7ai-15@gated-at.bofh.it> |
| In reply to | #1573316 |
On Fri, Feb 3, 2017 at 9:52 AM, Luck, Tony <tony.luck@intel.com> wrote: > On Thu, Feb 02, 2017 at 06:14:05PM -0800, David Carrillo-Cisneros wrote: >> If we tie allocation groups and monitoring groups, we are tying the >> meaning of CPUs and we'll have to choose between the CAT meaning or >> the perf meaning. >> >> Let's allow semantics that will allow perf like monitoring to >> eventually work, even if its not immediately supported. > > Would it work to make monitor groups be "task list only" or "cpu mask only" > (unlike control groups that allow mixing). That works, but please don't use chmod. Make it explicit by the group position (i.e. mon/cpus/grpCPU1, mon/tasks/grpTasks1). > > Then the intel_rdt_sched_in() code could pick the RMID in ways that > give you the perf(1) meaning. I.e. if you create a monitor group and assign > some CPUs to it, then we will always load the RMID for that monitor group > when running on those cpus, regardless of what group(s) the current process > belongs to. But if you didn't create any cpu-only monitor groups, then we'd > assign RMID using same rules as CLOSID (so measurements from a control group > would track allocation policies). I think that's very confusing for the user. A group's observed behavior should be determined by its attributes and not change depending on how other groups are configured. Think on multiple users monitoring simultaneously. > > We are already planning that creating monitor only groups will change > what is reported in the control group (e.g. you pull some tasks out of > the control group to monitor them separately, so the control group only > reports the tasks that you didn't move out for monitoring). That's also confusing, and the work-around that Vikas proposed of two separate files to enumerate tasks (one for control and one for monitoring) breaks the concept of a task group. From our discussions, we can support the use cases we care about without weird-corner cases, by having: - A set of allocation group as stand now. Either use the current resctrl, or rename it to something like resdir/ctrl (before v4.10 sails). - A set of monitoring task groups. Either in a "tasks" folder in a resmon fs or in resdir/mon/tasks. - A set of monitoring CPU groups. Either in a "cpus" folder in a resmon fs or in resdir/mon/cpus. So when a user measures a group (shown using the -G option, it could as well be the -R Vikas wants): 1) perf stat -e llc_occupancy -G resdir/ctrl/g1 measures the CAT allocation group as if RMIDs were managed like CLOSIDs. 2) perf stat -e llc_occupancy -G resdir/mon/tasks/g1 measures the combined occupancy of all tasks in g1 (like a cgroups in present perf). 3) perf stat -e llc_occupancy -C <some id of resdir/mon/cpus/g1> *XOR* perf stat -e llc_occupancy -G resdir/mon/cpus/g1 measures the combined occupancy of all tasks while ran in any CPU in g1 (perf-like filtering, not the CAT way). I know the present implementation scope is limited, so you could: - support 1) and/or 2) only - do a simple RMID management (e.g. same RMID all packages, allocate RMID on creation or fail) - do the custom fs based tool that Vikas mentioned instead of using perf_event_open (if it's somehow easier to build and maintain a new tool rather than reuse perf(1) ). any or all of the above are fine. But please don't choose group semantics that will prevent us from eventually supporting full perf-like behavior or that we already know explode in user's face. Thanks, David
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2017-02-03 23:30 +0100 |
| Message-ID | <t6UwG-7Wd-13@gated-at.bofh.it> |
| In reply to | #1573439 |
On Fri, Feb 03, 2017 at 01:08:05PM -0800, David Carrillo-Cisneros wrote:
> On Fri, Feb 3, 2017 at 9:52 AM, Luck, Tony <tony.luck@intel.com> wrote:
> > On Thu, Feb 02, 2017 at 06:14:05PM -0800, David Carrillo-Cisneros wrote:
> >> If we tie allocation groups and monitoring groups, we are tying the
> >> meaning of CPUs and we'll have to choose between the CAT meaning or
> >> the perf meaning.
> >>
> >> Let's allow semantics that will allow perf like monitoring to
> >> eventually work, even if its not immediately supported.
> >
> > Would it work to make monitor groups be "task list only" or "cpu mask only"
> > (unlike control groups that allow mixing).
>
> That works, but please don't use chmod. Make it explicit by the group
> position (i.e. mon/cpus/grpCPU1, mon/tasks/grpTasks1).
I had been thinking that after writing a PID to "tasks" we'd disallow
writes to "cpus". But is sounds nicer for the user to declare their
intention upfront. Counter propsosal in the naming war:
.../monitor/bytask/{groupname}
.../monitor/bycpu/{groupname}
> > Then the intel_rdt_sched_in() code could pick the RMID in ways that
> > give you the perf(1) meaning. I.e. if you create a monitor group and assign
> > some CPUs to it, then we will always load the RMID for that monitor group
> > when running on those cpus, regardless of what group(s) the current process
> > belongs to. But if you didn't create any cpu-only monitor groups, then we'd
> > assign RMID using same rules as CLOSID (so measurements from a control group
> > would track allocation policies).
>
> I think that's very confusing for the user. A group's observed
> behavior should be determined by its attributes and not change
> depending on how other groups are configured. Think on multiple users
> monitoring simultaneously.
>
> >
> > We are already planning that creating monitor only groups will change
> > what is reported in the control group (e.g. you pull some tasks out of
> > the control group to monitor them separately, so the control group only
> > reports the tasks that you didn't move out for monitoring).
>
> That's also confusing, and the work-around that Vikas proposed of two
> separate files to enumerate tasks (one for control and one for
> monitoring) breaks the concept of a task group.
There are some simple cases where we can make the data shown in the
original control group look the same. E.g. we move a few tasks over to a
/bytask/ group (or several groups if we want a very fine breakdown) and
then have the report from the control group sum the RMIDs from the monitor
groups and add to the total from the native RMID of the control group.
But this falls apart if the user asks a single monitor group to monitor
tasks from multiple control groups. Perhaps we could disallow this
(when we assign the first task to a monitor group, capture the CLOSID
and then only allow other tasks with the same CLOSID to be added ... unless
the group becomes empty, and which point we can latch onto a new CLOSID).
/bycpu/ monitoring is very resource intensive if we have to preserve
the control group reports. We'd need to allocate MAXCLOSID[1] RMIDs for
each group so that we can keep separate counts for tasks from each
control group that run on our CPUs and then sum them to report the
/bycpu/ data (instead of just one RMID, and no math). This also
puts more memory references into the sched_in path while we
figure out which RMID to load into PQR_ASSOC.
I'd want to warn the user in the Documentation that splitting off
too many monitor groups from a control group will result in less
than stellar accuracy in reporting as the kernel cannot read
multiple RMIDs atomically and data is changing between reads.
> I know the present implementation scope is limited, so you could:
> - support 1) and/or 2) only
> - do a simple RMID management (e.g. same RMID all packages, allocate
> RMID on creation or fail)
> - do the custom fs based tool that Vikas mentioned instead of using
> perf_event_open (if it's somehow easier to build and maintain a new
> tool rather than reuse perf(1) ).
>
> any or all of the above are fine. But please don't choose group
> semantics that will prevent us from eventually supporting full
> perf-like behavior or that we already know explode in user's face.
I'm trying hard to find a way to do this. I.e. start with a patch
that has limited capabilities and needs a custom tool, but can later
grow into something that meets your needs.
-Tony
[1] Lazy allocation means finding we can't find a free RMID in the
middle of context switch ... not willing to go there.
[toc] | [prev] | [next] | [standalone]
| From | Stephane Eranian <eranian@google.com> |
|---|---|
| Date | 2017-02-07 09:10 +0100 |
| Message-ID | <t890C-gF-25@gated-at.bofh.it> |
| In reply to | #1573316 |
Hi,
I wanted to take a few steps back and look at the overall goals for
cache monitoring.
From the various threads and discussion, my understanding is as follows.
I think the design must ensure that the following usage models can be monitored:
- the allocations in your CAT partitions
- the allocations from a task (inclusive of children tasks)
- the allocations from a group of tasks (inclusive of children tasks)
- the allocations from a CPU
- the allocations from a group of CPUs
All cases but first one (CAT) are natural usage. So I want to describe
the CAT in more details.
The goal, as I understand it, it to monitor what is going on inside
the CAT partition to detect
whether it saturates or if it has room to "breathe". Let's take a
simple example.
Suppose, we have a CAT group, cat1:
cat1: 20MB partition (CLOSID1)
CPUs=CPU0,CPU1
TASKs=PID20
There can only be one CLOSID active on a CPU at a time. The kernel
chooses to prioritize tasks over CPU when enforcing cases with multiple
CLOSIDs.
Let's review how this works for cat1 and for each scenario look at how
the kernel enforces or not the cache partition:
1. ENFORCED: PIDx with no CLOSID runs on CPU0 or CPU1
2. NOT ENFORCED: PIDx with CLOSIDx (x!=1) runs on CPU0, CPU1
3. ENFORCED: PID20 runs with CLOSID1 on CPU0, CPU1
4. ENFORCED: PID20 runs with CLOSID1 on CPUx (x!=0,1) with CPU CLOSIDx (x!=1)
5. ENFORCED: PID20 runs with CLOSID1 on CPUx (x!=0,1) with no CLOSID
Now, let's review how we could track the allocations done in cat1 using a single
RMID. There can only be one RMID active at a time per CPU. The kernel
chooses to prioritize tasks over CPU:
cat1: 20MB partition (CLOSID1, RMID1)
CPUs=CPU0,CPU1
TASKs=PID20
1. MONITORED: PIDx with no RMID runs on CPU0 or CPU1
2. NOT MONITORED: PIDx with RMIDx (x!=1) runs on CPU0, CPU1
3. MONITORED: PID20 with RMID1 runs on CPU0, CPU1
4. MONITORED: PID20 with RMD1 runs on CPUx (x!=0,1) with CPU RMIDx (x!=1)
5. MONITORED: PID20 runs with RMID1 on CPUx (x!=0,1) with no RMID
To make sense to a user, the cases where the hardware monitors MUST be
the same as the cases where the hardware enforces the cache
partitioning.
Here we see that it works using a single RMID.
However doing so limits certain monitoring modes where a user might want to
get a breakdown per CPU of the allocations, such as with:
$ perf stat -a -A -e llc_occupancy -R cat1
(where -R points to the monitoring group in rsrcfs). Here this mode would not be
possible because the two CPUs in the group share the same RMID.
Now let's take another scenario, and suppose you have two monitoring groups
as follows:
mon1: RMID1
CPUs=CPU0,CPU1
mon2: RMID2
TASKS=PID20
If PID20 runs on CP0, then RMID2 is activated, and thus allocations
done by PID20 are not counted towards RMID1. There is a blind spot.
Whether or not this is a problem depends on the semantic exported by
the interface for CPU mode:
1-Count all allocations from any tasks running on CPU
2-Count all allocations from tasks which are NOT monitoring themselves
If the kernel choses 1, then there is a blind spot and the measurement
is not as accurate as it could be because of the decision to use only one RDMID.
But if the kernel choses 2, then everything works fine with a single RMID.
If the kernel treats occupancy monitoring as measuring cycles on a CPU, i.e.,
measure any activity from any thread (choice 1), then the single RMID per group
does not work.
If the kernel treats occupancy monitoring as measuring cycles in a cgroup on a
CPU, i.e., measures only when threads of the cgroup run on that CPU, then using
a single RMID per group works.
Hope this helps clarifies the usage model and design choices.
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2017-02-07 20:10 +0100 |
| Message-ID | <t8jjk-6Qz-9@gated-at.bofh.it> |
| In reply to | #1575450 |
On Tue, Feb 07, 2017 at 12:08:09AM -0800, Stephane Eranian wrote: > Hi, > > I wanted to take a few steps back and look at the overall goals for > cache monitoring. > From the various threads and discussion, my understanding is as follows. > > I think the design must ensure that the following usage models can be monitored: > - the allocations in your CAT partitions > - the allocations from a task (inclusive of children tasks) > - the allocations from a group of tasks (inclusive of children tasks) > - the allocations from a CPU > - the allocations from a group of CPUs > > All cases but first one (CAT) are natural usage. So I want to describe > the CAT in more details. > The goal, as I understand it, it to monitor what is going on inside > the CAT partition to detect > whether it saturates or if it has room to "breathe". Let's take a > simple example. By "natural usage" you mean "like perf(1) provides for other events"? But we are trying to figure out requirements here ... what data do people need to manage caches and memory bandwidth. So from this perspective monitoring a CAT group is a natural first choice ... did we provision this group with too much, or too little cache. From that starting point I can see that a possible next step when finding that a CAT group has too small a cache is to drill down to find out how the tasks in the group are using cache. Armed with that information you could move tasks that hog too much cache (and are believed to be streaming through memory) into a different CAT group. What I'm not seeing is how drilling to CPUs helps you. Say you have CPUs=CPU0,CPU1 in the CAT group and you collect data that shows that 75% of the cache occupancy is attributed to CPU0, and only 25% to CPU1. What can you do with this information to improve things? If it is deemed too complex (from a kernel code perspective) to implement per-CPU reporting how bad a loss would that be? -Tony
[toc] | [prev] | [next] | [standalone]
| From | Stephane Eranian <eranian@google.com> |
|---|---|
| Date | 2017-02-08 22:40 +0100 |
| Message-ID | <t8I82-5Aj-15@gated-at.bofh.it> |
| In reply to | #1575963 |
Tony, On Tue, Feb 7, 2017 at 10:52 AM, Luck, Tony <tony.luck@intel.com> wrote: > On Tue, Feb 07, 2017 at 12:08:09AM -0800, Stephane Eranian wrote: >> Hi, >> >> I wanted to take a few steps back and look at the overall goals for >> cache monitoring. >> From the various threads and discussion, my understanding is as follows. >> >> I think the design must ensure that the following usage models can be monitored: >> - the allocations in your CAT partitions >> - the allocations from a task (inclusive of children tasks) >> - the allocations from a group of tasks (inclusive of children tasks) >> - the allocations from a CPU >> - the allocations from a group of CPUs >> >> All cases but first one (CAT) are natural usage. So I want to describe >> the CAT in more details. >> The goal, as I understand it, it to monitor what is going on inside >> the CAT partition to detect >> whether it saturates or if it has room to "breathe". Let's take a >> simple example. > > By "natural usage" you mean "like perf(1) provides for other events"? > Yes, people are used to monitoring events per task or per CPU. In that sense, it is the common usage model. Cgroup monitoring is a derivative of per-cpu mode. > But we are trying to figure out requirements here ... what data do people > need to manage caches and memory bandwidth. So from this perspective > monitoring a CAT group is a natural first choice ... did we provision > this group with too much, or too little cache. > I am not saying CAT is not natural. I am saying it is a justified requirement but a new one and thus need to make sure it is understood and that the kernel must track CAT partition and CAT partition cache occupancy monitoring similarly. > From that starting point I can see that a possible next step when > finding that a CAT group has too small a cache is to drill down to > find out how the tasks in the group are using cache. Armed with that > information you could move tasks that hog too much cache (and are believed > to be streaming through memory) into a different CAT group. > This is a valid usage model. But you have people who care about monitoring occupancy but do not necessarily use CAT partitions. Yet in this case, the occupancy data is still very useful to gauge cache footprint of a workload. Therefore this usage model should not be discounted. > What I'm not seeing is how drilling to CPUs helps you. > Looking for imbalance, for instance. Are all the allocations done from only a subset of the CPUs? > Say you have CPUs=CPU0,CPU1 in the CAT group and you collect data that > shows that 75% of the cache occupancy is attributed to CPU0, and only > 25% to CPU1. What can you do with this information to improve things? > If it is deemed too complex (from a kernel code perspective) to > implement per-CPU reporting how bad a loss would that be? > It is okay to first focus on per-task and per-CAT partition. What I'd like to see is an API that could possibly be extended later on to do per-CPU only mode. I am okay with having only per-CAT and per-task groups initially to keep things simpler. But the rsrcfs interface should allow extension to per-CPU only mode. Then the kernel implementation would take care of allocating the RMID accordingly. The key is always to ensure allocations can be tracked since inception of the group be it CAT, tasks, CPU. > -Tony
[toc] | [prev] | [next] | [standalone]
| From | Shivappa Vikas <vikas.shivappa@intel.com> |
|---|---|
| Date | 2017-02-07 21:30 +0100 |
| Message-ID | <t8kyJ-7xK-7@gated-at.bofh.it> |
| In reply to | #1575450 |
On Tue, 7 Feb 2017, Stephane Eranian wrote:
> Hi,
>
> I wanted to take a few steps back and look at the overall goals for
> cache monitoring.
> From the various threads and discussion, my understanding is as follows.
>
> I think the design must ensure that the following usage models can be monitored:
> - the allocations in your CAT partitions
> - the allocations from a task (inclusive of children tasks)
> - the allocations from a group of tasks (inclusive of children tasks)
> - the allocations from a CPU
> - the allocations from a group of CPUs
>
> All cases but first one (CAT) are natural usage. So I want to describe
> the CAT in more details.
> The goal, as I understand it, it to monitor what is going on inside
> the CAT partition to detect
> whether it saturates or if it has room to "breathe". Let's take a
> simple example.
>
> Suppose, we have a CAT group, cat1:
>
> cat1: 20MB partition (CLOSID1)
> CPUs=CPU0,CPU1
> TASKs=PID20
>
> There can only be one CLOSID active on a CPU at a time. The kernel
> chooses to prioritize tasks over CPU when enforcing cases with multiple
> CLOSIDs.
>
> Let's review how this works for cat1 and for each scenario look at how
> the kernel enforces or not the cache partition:
>
> 1. ENFORCED: PIDx with no CLOSID runs on CPU0 or CPU1
> 2. NOT ENFORCED: PIDx with CLOSIDx (x!=1) runs on CPU0, CPU1
> 3. ENFORCED: PID20 runs with CLOSID1 on CPU0, CPU1
> 4. ENFORCED: PID20 runs with CLOSID1 on CPUx (x!=0,1) with CPU CLOSIDx (x!=1)
> 5. ENFORCED: PID20 runs with CLOSID1 on CPUx (x!=0,1) with no CLOSID
>
> Now, let's review how we could track the allocations done in cat1 using a single
> RMID. There can only be one RMID active at a time per CPU. The kernel
> chooses to prioritize tasks over CPU:
>
> cat1: 20MB partition (CLOSID1, RMID1)
> CPUs=CPU0,CPU1
> TASKs=PID20
>
> 1. MONITORED: PIDx with no RMID runs on CPU0 or CPU1
> 2. NOT MONITORED: PIDx with RMIDx (x!=1) runs on CPU0, CPU1
> 3. MONITORED: PID20 with RMID1 runs on CPU0, CPU1
> 4. MONITORED: PID20 with RMD1 runs on CPUx (x!=0,1) with CPU RMIDx (x!=1)
> 5. MONITORED: PID20 runs with RMID1 on CPUx (x!=0,1) with no RMID
>
> To make sense to a user, the cases where the hardware monitors MUST be
> the same as the cases where the hardware enforces the cache
> partitioning.
>
> Here we see that it works using a single RMID.
>
> However doing so limits certain monitoring modes where a user might want to
> get a breakdown per CPU of the allocations, such as with:
> $ perf stat -a -A -e llc_occupancy -R cat1
> (where -R points to the monitoring group in rsrcfs). Here this mode would not be
> possible because the two CPUs in the group share the same RMID.
In the requirements here https://marc.info/?l=linux-kernel&m=148597969808732
8) Can get measurements for subsets of tasks in a CAT group (to find the
guys hogging the resources).
This should also applies to the subsets of cpus.
That would let you monitor on CPUs that is a subset or different from a CAT
group. That should let you create mon groups like in the second example you
mention along with the control groups above.
mon0: RMID0
CPUs=CPU0
mon1: RMID1
CPUs=CPU1
mon2: RMID2
CPUs=CPU2
...
>
> Now let's take another scenario, and suppose you have two monitoring groups
> as follows:
>
> mon1: RMID1
> CPUs=CPU0,CPU1
> mon2: RMID2
> TASKS=PID20
>
> If PID20 runs on CP0, then RMID2 is activated, and thus allocations
> done by PID20 are not counted towards RMID1. There is a blind spot.
>
> Whether or not this is a problem depends on the semantic exported by
> the interface for CPU mode:
> 1-Count all allocations from any tasks running on CPU
> 2-Count all allocations from tasks which are NOT monitoring themselves
>
> If the kernel choses 1, then there is a blind spot and the measurement
> is not as accurate as it could be because of the decision to use only one RDMID.
> But if the kernel choses 2, then everything works fine with a single RMID.
>
> If the kernel treats occupancy monitoring as measuring cycles on a CPU, i.e.,
> measure any activity from any thread (choice 1), then the single RMID per group
> does not work.
>
> If the kernel treats occupancy monitoring as measuring cycles in a cgroup on a
> CPU, i.e., measures only when threads of the cgroup run on that CPU, then using
> a single RMID per group works.
>
Agree there are blind spots in both. But the requirements is trying to be based
on the resctrl allocation as Thomas suggested.
Which is aligned to monitoring real time tasks as i understand.
for the above example, some tasks which donot have an RMID(say in the root
group) are the real time tasks that are specially configured to running on a cpux which need to be
allocated or monitored.
> Hope this helps clarifies the usage model and design choices.
>
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2017-02-06 20:00 +0100 |
| Message-ID | <t7WG5-ld-5@gated-at.bofh.it> |
| In reply to | #1572043 |
Digging through the e-mails from last week to generate a new version
of the requirements I looked harder at this:
> 12) Whatever fs or syscall is provided instead of perf syscalls, it
> should provide total_time_enabled in the way perf does, otherwise is
> hard to interpret MBM values.
This looks tricky if we are piggy-backing on the CAT code to switch
RMID along with CLOSID at context switch time. We could get an
approximation by adding:
if (newRMID != oldRMID) {
now = grab current time in some format
atomic_add(rmid_enabled_time[oldRMID], now - this_cpu_read(rmid_time));
this_cpu_write(rmid_time, now);
}
but:
1) that would only work on a single socket machine (we'd really want rmid_enabled_time
separately for each socket)
2) when we want to read that enabled time, we'd really need to add time for all the
threads currently running on CPUs across the system since we last switched RMID
3) reading the time and doing atomic ops in context switch code won't be popular
:-(
-Tony
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2017-02-06 22:30 +0100 |
| Message-ID | <t7Z1g-1XE-35@gated-at.bofh.it> |
| In reply to | #1572043 |
> 12) Whatever fs or syscall is provided instead of perf syscalls, it > should provide total_time_enabled in the way perf does, otherwise is > hard to interpret MBM values. It seems that it is hard to define what we even mean by memory bandwidth. If you are measuring just one task and you find that the total number of bytes read is 1GB at some point, and one second later the total bytes is 2GB, then it is clear that the average bandwidth for this process is 1GB/s. If you know that the task was only running for 50% of the cycles during that 1s interval, you could say that it is doing 2GB/s ... which is I believe what you were thinking when you wrote #12 above. But whether that is right depends a bit on *why* it only ran 50% of the time. If it was time-sliced out by the scheduler ... then it may have been trying to be a 2GB/s app. But if it was waiting for packets from the network, then it really is using 1 GB/s. All bets are off if you are measuring a service that consists of several tasks running concurrently. All you can really talk about is the aggregate average bandwidth (total bytes / wall-clock time). It makes no sense to try and factor in how much cpu time each of the individual tasks got. -Tony
[toc] | [prev] | [next] | [standalone]
| From | Shivappa Vikas <vikas.shivappa@intel.com> |
|---|---|
| Date | 2017-02-06 22:40 +0100 |
| Message-ID | <t7ZaW-214-17@gated-at.bofh.it> |
| In reply to | #1575158 |
On Mon, 6 Feb 2017, Luck, Tony wrote: >> 12) Whatever fs or syscall is provided instead of perf syscalls, it >> should provide total_time_enabled in the way perf does, otherwise is >> hard to interpret MBM values. > > It seems that it is hard to define what we even mean by memory bandwidth. > > If you are measuring just one task and you find that the total number of bytes > read is 1GB at some point, and one second later the total bytes is 2GB, then > it is clear that the average bandwidth for this process is 1GB/s. If you know > that the task was only running for 50% of the cycles during that 1s interval, > you could say that it is doing 2GB/s ... which is I believe what you were > thinking when you wrote #12 above. But whether that is right depends a > bit on *why* it only ran 50% of the time. If it was time-sliced out by the > scheduler ... then it may have been trying to be a 2GB/s app. But if it > was waiting for packets from the network, then it really is using 1 GB/s. Is the requirement is to have both enabled and run time or just enabled time (enabled time must be easy to report - just the wall time from start trace to end trace)? This is not reported correctly in the upstream perf cqm and for cgroup -C we dont report it either (since we report the package). Thanks, Vikas > > All bets are off if you are measuring a service that consists of several > tasks running concurrently. All you can really talk about is the aggregate > average bandwidth (total bytes / wall-clock time). It makes no sense to > try and factor in how much cpu time each of the individual tasks got. > > -Tony >
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-02-06 22:50 +0100 |
| Message-ID | <t7ZkB-24P-7@gated-at.bofh.it> |
| In reply to | #1575169 |
On Mon, Feb 6, 2017 at 1:36 PM, Shivappa Vikas <vikas.shivappa@intel.com> wrote: > > > On Mon, 6 Feb 2017, Luck, Tony wrote: > >>> 12) Whatever fs or syscall is provided instead of perf syscalls, it >>> should provide total_time_enabled in the way perf does, otherwise is >>> hard to interpret MBM values. >> >> >> It seems that it is hard to define what we even mean by memory bandwidth. >> >> If you are measuring just one task and you find that the total number of >> bytes >> read is 1GB at some point, and one second later the total bytes is 2GB, >> then >> it is clear that the average bandwidth for this process is 1GB/s. If you >> know >> that the task was only running for 50% of the cycles during that 1s >> interval, >> you could say that it is doing 2GB/s ... which is I believe what you were >> thinking when you wrote #12 above. But whether that is right depends a >> bit on *why* it only ran 50% of the time. If it was time-sliced out by the >> scheduler ... then it may have been trying to be a 2GB/s app. But if it >> was waiting for packets from the network, then it really is using 1 GB/s. > > > Is the requirement is to have both enabled and run time or just enabled time > (enabled time must be easy to report - just the wall time from start trace > to end trace)? Both, but since the original requirements dropped rotation, then total_running == total_enabled. > > This is not reported correctly in the upstream perf cqm and for > cgroup -C we dont report it either (since we report the package). using the -x option shows the run time and the % enabled. Many tools uses that csv output. > > Thanks, > Vikas > > >> >> All bets are off if you are measuring a service that consists of several >> tasks running concurrently. All you can really talk about is the aggregate >> average bandwidth (total bytes / wall-clock time). It makes no sense to >> try and factor in how much cpu time each of the individual tasks got. >> >> -Tony >> >
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web