Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1560846 > unrolled thread
| Started by | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| First post | 2017-01-17 18:40 +0100 |
| Last post | 2017-01-20 20:40 +0100 |
| Articles | 20 on this page of 29 — 7 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-17 18:40 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-18 03:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-18 10:00 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Peter Zijlstra <peterz@infradead.org> - 2017-01-18 11:10 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-19 21:10 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-18 20:50 +0100
RE: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes "Yu, Fenghua" <fenghua.yu@intel.com> - 2017-01-18 22:20 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-18 22:20 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-19 18:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-20 08:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-20 09:40 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-20 21:30 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-19 03:20 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-19 18:30 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-19 19:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Vikas Shivappa <vikas.shivappa@linux.intel.com> - 2017-01-19 03:30 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Stephane Eranian <eranian@google.com> - 2017-01-19 07:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-19 19:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Vikas Shivappa <vikas.shivappa@linux.intel.com> - 2017-01-20 03:40 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-20 09:00 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-20 15:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-20 21:20 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-20 22:10 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes David Carrillo-Cisneros <davidcc@google.com> - 2017-01-20 22:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-21 01:00 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Thomas Gleixner <tglx@linutronix.de> - 2017-01-23 11:20 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Peter Zijlstra <peterz@infradead.org> - 2017-01-23 12:40 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Shivappa Vikas <vikas.shivappa@intel.com> - 2017-01-20 21:50 +0100
Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes Stephane Eranian <eranian@google.com> - 2017-01-20 20:40 +0100
Page 1 of 2 [1] 2 Next page →
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-01-17 18:40 +0100 |
| Subject | Re: [PATCH 00/12] Cqm2: Intel Cache quality monitoring fixes |
| Message-ID | <t0FTJ-1vq-39@gated-at.bofh.it> |
On Fri, 6 Jan 2017, Vikas Shivappa wrote: > Cqm(cache quality monitoring) is part of Intel RDT(resource director > technology) which enables monitoring and controlling of processor shared > resources via MSR interface. We know that already. No need for advertising this over and over. > Below are the issues and the fixes we attempt- Emphasis on attempt. > - Issue(1): Inaccurate data for per package data, systemwide. Just prints > zeros or arbitrary numbers. > > Fix: Patches fix this by just throwing an error if the mode is not supported. > The modes supported is task monitoring and cgroup monitoring. > Also the per package > data for say socket x is returned with the -C <cpu on socketx> -G cgrpy option. > The systemwide data can be looked up by monitoring root cgroup. Fine. That just lacks any comment in the implementation. Otherwise I would not have asked the question about cpu monitoring. Though I fundamentaly hate the idea of requiring cgroups for this to work. If I just want to look at CPU X why on earth do I have to set up all that cgroup muck? Just because your main focus is cgroups? > - Issue(2): RMIDs are global and dont scale with more packages and hence > also run out of RMIDs very soon. > > Fix: Support per pkg RMIDs hence scale better with more > packages, and get more RMIDs to use and use when needed (ie when tasks > are actually scheduled on the package). That's fine, just the implementation is completely broken > - Issue(5): CAT and cqm/mbm write the same PQR_ASSOC_MSR seperately > Fix: Integrate the sched in code and write the PQR_MSR only once every switch_to Brilliant stuff that. I bet that the seperate MSR writes which we have now are actually faster even in the worst case of two writes. > [PATCH 02/12] x86/cqm: Remove cqm recycling/conflict handling That's the only undisputed and probably correct patch in that whole series. Executive summary: The whole series except 2/12 is a complete trainwreck. Major issues: - Patch split is random and leads to non compilable and non functional steps aside of creating a unreviewable mess. - The core code is updated with random hooks which claim to be a generic framework but are completely hardcoded for that particular CQM case. The introduced hooks are neither fully documented (semantical and functional) nor justified why they are required. - The code quality is horrible in coding style, technical correctness and design. So how can we proceed from here? I want to see small patch series which change a certain aspect of the implementation and only that. They need to be split in preparatory changes, which refactor code or add new interfaces to the core code, and the actual implementation in CQM/MBM. Each of the patches must have a proper changelog explaining the WHY and if required the semantical properties of a new interface. Each of the patches must compile without warnigns/errors and be fully functional vs. the changes they do. Any attempt to resend this disaster as a whole will be NACKed w/o even looking at it. I've wasted enough time with this wholesale approach in the past and it did not work out. I'm not going to play that wasteful game over and over again. Thanks, tglx
[toc] | [next] | [standalone]
| From | Shivappa Vikas <vikas.shivappa@intel.com> |
|---|---|
| Date | 2017-01-18 03:50 +0100 |
| Message-ID | <t0OtY-6Fn-1@gated-at.bofh.it> |
| In reply to | #1560846 |
On Tue, 17 Jan 2017, Thomas Gleixner wrote: > On Fri, 6 Jan 2017, Vikas Shivappa wrote: >> Cqm(cache quality monitoring) is part of Intel RDT(resource director >> technology) which enables monitoring and controlling of processor shared >> resources via MSR interface. > > We know that already. No need for advertising this over and over. > >> Below are the issues and the fixes we attempt- > > Emphasis on attempt. > >> - Issue(1): Inaccurate data for per package data, systemwide. Just prints >> zeros or arbitrary numbers. >> >> Fix: Patches fix this by just throwing an error if the mode is not supported. >> The modes supported is task monitoring and cgroup monitoring. >> Also the per package >> data for say socket x is returned with the -C <cpu on socketx> -G cgrpy option. >> The systemwide data can be looked up by monitoring root cgroup. > > Fine. That just lacks any comment in the implementation. Otherwise I would > not have asked the question about cpu monitoring. Though I fundamentaly > hate the idea of requiring cgroups for this to work. > > If I just want to look at CPU X why on earth do I have to set up all that > cgroup muck? Just because your main focus is cgroups? The upstream per cpu data is broken because its not overriding the other task event RMIDs on that cpu with the cpu event RMID. Can be fixed by adding a percpu struct to hold the RMID thats affinitized to the cpu, however then we miss all the task llc_occupancy in that - still evaluating it. > >> - Issue(2): RMIDs are global and dont scale with more packages and hence >> also run out of RMIDs very soon. >> >> Fix: Support per pkg RMIDs hence scale better with more >> packages, and get more RMIDs to use and use when needed (ie when tasks >> are actually scheduled on the package). > > That's fine, just the implementation is completely broken > >> - Issue(5): CAT and cqm/mbm write the same PQR_ASSOC_MSR seperately >> Fix: Integrate the sched in code and write the PQR_MSR only once every switch_to > > Brilliant stuff that. I bet that the seperate MSR writes which we have now > are actually faster even in the worst case of two writes. > >> [PATCH 02/12] x86/cqm: Remove cqm recycling/conflict handling > > That's the only undisputed and probably correct patch in that whole series. > > Executive summary: The whole series except 2/12 is a complete trainwreck. > > Major issues: > > - Patch split is random and leads to non compilable and non functional > steps aside of creating a unreviewable mess. Will split as per the comments into divisions that are compilable without warnings. The series was tested for compile and build but wasnt for warnings. (except for the int if_cgroup_event which you pointed..) > > - The core code is updated with random hooks which claim to be a generic > framework but are completely hardcoded for that particular CQM case. > > The introduced hooks are neither fully documented (semantical and > functional) nor justified why they are required. > > - The code quality is horrible in coding style, technical correctness and > design. > > So how can we proceed from here? > > I want to see small patch series which change a certain aspect of the > implementation and only that. They need to be split in preparatory changes, > which refactor code or add new interfaces to the core code, and the actual > implementation in CQM/MBM. Appreciate your time for review and all the feedback. Will plan to send a smaller patch series which includes the following contents at first and then follow up other series: - the current 2/12 - add the per package RMID with the fixes mentioned. - The reuse patch which reuses the RMIDs that are freed from events. - Patch to indicate error if RMID was still not available. - Any relavant sched and hot cpu updates. That way its a functional set where tasks would be able to use different RMIDs on different packages and provide correct data for 'task events'. Either the data would be correct or would indicate an error that we were limited by h/w RMIDs. Is that a reasonable set to target ? Then followup with different series to get correct data for 'cgroup events' and support cgroup and task together and other issues, system/ per cpu events etc. Thanks, Vikas > > Each of the patches must have a proper changelog explaining the WHY and if > required the semantical properties of a new interface. > > Each of the patches must compile without warnigns/errors and be fully > functional vs. the changes they do. > > Any attempt to resend this disaster as a whole will be NACKed w/o even > looking at it. I've wasted enough time with this wholesale approach in the > past and it did not work out. I'm not going to play that wasteful game > over and over again. > > Thanks, > > tglx > > >
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-01-18 10:00 +0100 |
| Message-ID | <t0Ug1-1O0-11@gated-at.bofh.it> |
| In reply to | #1561230 |
On Tue, 17 Jan 2017, Shivappa Vikas wrote: > On Tue, 17 Jan 2017, Thomas Gleixner wrote: > > On Fri, 6 Jan 2017, Vikas Shivappa wrote: > > > - Issue(1): Inaccurate data for per package data, systemwide. Just prints > > > zeros or arbitrary numbers. > > > > > > Fix: Patches fix this by just throwing an error if the mode is not > > > supported. > > > The modes supported is task monitoring and cgroup monitoring. > > > Also the per package > > > data for say socket x is returned with the -C <cpu on socketx> -G cgrpy > > > option. > > > The systemwide data can be looked up by monitoring root cgroup. > > > > Fine. That just lacks any comment in the implementation. Otherwise I would > > not have asked the question about cpu monitoring. Though I fundamentaly > > hate the idea of requiring cgroups for this to work. > > > > If I just want to look at CPU X why on earth do I have to set up all that > > cgroup muck? Just because your main focus is cgroups? > > The upstream per cpu data is broken because its not overriding the other task > event RMIDs on that cpu with the cpu event RMID. > > Can be fixed by adding a percpu struct to hold the RMID thats affinitized > to the cpu, however then we miss all the task llc_occupancy in that - still > evaluating it. The point here is that CQM is closely connected to the cache allocation technology. After a lengthy discussion we ended up having - per cpu CLOSID - per task CLOSID where all tasks which do not have a CLOSID assigned use the CLOSID which is assigned to the CPU they are running on. So if I configure a system by simply partitioning the cache per cpu, which is the proper way to do it for HPC and RT usecases where workloads are partitioned on CPUs as well, then I really want to have an equaly simple way to monitor the occupancy for that reservation. And looking at that from the CAT point of view, which is the proper way to do it, makes it obvious that CQM should be modeled to match CAT. So lets assume the following: CPU 0-3 default CLOSID 0 CPU 4 CLOSID 1 CPU 5 CLOSID 2 CPU 6 CLOSID 3 CPU 7 CLOSID 3 T1 CLOSID 4 T2 CLOSID 5 T3 CLOSID 6 T4 CLOSID 6 All other tasks use the per cpu defaults, i.e. the CLOSID of the CPU they run on. then the obvious basic monitoring requirement is to have a RMID for each CLOSID. So when I monitor CPU4, i.e. CLOSID 1 and T1 runs on CPU4, then I do not care at all about the occupancy of T1 simply because that is running on a seperate reservation. Trying to make that an aggregated value in the first place is completely wrong. If you want an aggregate, which is pretty much useless, then user space tools can generate it easily. The whole approach you and David have taken is to whack some desired cgroup functionality and whatever into CQM without rethinking the overall design. And that's fundamentaly broken because it does not take cache (and memory bandwidth) allocation into account. I seriously doubt, that the existing CQM/MBM code can be refactored in any useful way. As Peter Zijlstra said before: Remove the existing cruft completely and start with completely new design from scratch. And this new design should start from the allocation angle and then add the whole other muck on top so far its possible. Allocation related monitoring must be the primary focus, everything else is just tinkering. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-01-18 11:10 +0100 |
| Message-ID | <t0VlM-2Ex-9@gated-at.bofh.it> |
| In reply to | #1561363 |
On Wed, Jan 18, 2017 at 09:53:02AM +0100, Thomas Gleixner wrote: > The whole approach you and David have taken is to whack some desired cgroup > functionality and whatever into CQM without rethinking the overall > design. And that's fundamentaly broken because it does not take cache (and > memory bandwidth) allocation into account. > > I seriously doubt, that the existing CQM/MBM code can be refactored in any > useful way. As Peter Zijlstra said before: Remove the existing cruft > completely and start with completely new design from scratch. > > And this new design should start from the allocation angle and then add the > whole other muck on top so far its possible. Allocation related monitoring > must be the primary focus, everything else is just tinkering. Agreed, the little I have seen of these patches is quite horrible. And there seems to be a definite lack of design; or at the very least an utter lack of communication of it. The approach, in so far that I could make sense of it, seems to utterly rape perf-cgroup. I think Thomas makes a sensible point in trying to match it to the CAT stuffs.
[toc] | [prev] | [next] | [standalone]
| From | Shivappa Vikas <vikas.shivappa@intel.com> |
|---|---|
| Date | 2017-01-19 21:10 +0100 |
| Message-ID | <t1rbY-64j-23@gated-at.bofh.it> |
| In reply to | #1561431 |
Hello Peterz, On Wed, 18 Jan 2017, Peter Zijlstra wrote: > On Wed, Jan 18, 2017 at 09:53:02AM +0100, Thomas Gleixner wrote: >> The whole approach you and David have taken is to whack some desired cgroup >> functionality and whatever into CQM without rethinking the overall >> design. And that's fundamentaly broken because it does not take cache (and >> memory bandwidth) allocation into account. >> >> I seriously doubt, that the existing CQM/MBM code can be refactored in any >> useful way. As Peter Zijlstra said before: Remove the existing cruft >> completely and start with completely new design from scratch. >> >> And this new design should start from the allocation angle and then add the >> whole other muck on top so far its possible. Allocation related monitoring >> must be the primary focus, everything else is just tinkering. > > Agreed, the little I have seen of these patches is quite horrible. And > there seems to be a definite lack of design; or at the very least an > utter lack of communication of it. the 1/12 Documentation patch describes the interface. Basically we are just trying to support the task and cgroup monitoring. By the design document, do you want a document describing how we enable the cgroup for cqm since its a special case? (which would include all the arch_info in the perf_cgroup we add to keep track of hierarchy in the driver , etc ..) Thanks, Vikas > > The approach, in so far that I could make sense of it, seems to utterly > rape perf-cgroup. I think Thomas makes a sensible point in trying to > match it to the CAT stuffs. >
[toc] | [prev] | [next] | [standalone]
| From | Shivappa Vikas <vikas.shivappa@intel.com> |
|---|---|
| Date | 2017-01-18 20:50 +0100 |
| Message-ID | <t14p4-8e2-25@gated-at.bofh.it> |
| In reply to | #1561363 |
On Wed, 18 Jan 2017, Thomas Gleixner wrote: > On Tue, 17 Jan 2017, Shivappa Vikas wrote: >> On Tue, 17 Jan 2017, Thomas Gleixner wrote: >>> On Fri, 6 Jan 2017, Vikas Shivappa wrote: >>>> - Issue(1): Inaccurate data for per package data, systemwide. Just prints >>>> zeros or arbitrary numbers. >>>> >>>> Fix: Patches fix this by just throwing an error if the mode is not >>>> supported. >>>> The modes supported is task monitoring and cgroup monitoring. >>>> Also the per package >>>> data for say socket x is returned with the -C <cpu on socketx> -G cgrpy >>>> option. >>>> The systemwide data can be looked up by monitoring root cgroup. >>> >>> Fine. That just lacks any comment in the implementation. Otherwise I would >>> not have asked the question about cpu monitoring. Though I fundamentaly >>> hate the idea of requiring cgroups for this to work. >>> >>> If I just want to look at CPU X why on earth do I have to set up all that >>> cgroup muck? Just because your main focus is cgroups? >> >> The upstream per cpu data is broken because its not overriding the other task >> event RMIDs on that cpu with the cpu event RMID. >> >> Can be fixed by adding a percpu struct to hold the RMID thats affinitized >> to the cpu, however then we miss all the task llc_occupancy in that - still >> evaluating it. > > The point here is that CQM is closely connected to the cache allocation > technology. After a lengthy discussion we ended up having > > - per cpu CLOSID > - per task CLOSID > > where all tasks which do not have a CLOSID assigned use the CLOSID which is > assigned to the CPU they are running on. > > So if I configure a system by simply partitioning the cache per cpu, which > is the proper way to do it for HPC and RT usecases where workloads are > partitioned on CPUs as well, then I really want to have an equaly simple > way to monitor the occupancy for that reservation. > > And looking at that from the CAT point of view, which is the proper way to > do it, makes it obvious that CQM should be modeled to match CAT. Ok , makes sense. Tony and Fenghua had suggested some ideas to model the two more close together. Let me do some more brainstorming and try to come up with a draft that can be discussed. > > So lets assume the following: > > CPU 0-3 default CLOSID 0 > CPU 4 CLOSID 1 > CPU 5 CLOSID 2 > CPU 6 CLOSID 3 > CPU 7 CLOSID 3 > > T1 CLOSID 4 > T2 CLOSID 5 > T3 CLOSID 6 > T4 CLOSID 6 > > All other tasks use the per cpu defaults, i.e. the CLOSID of the CPU > they run on. > > then the obvious basic monitoring requirement is to have a RMID for each > CLOSID. > > So when I monitor CPU4, i.e. CLOSID 1 and T1 runs on CPU4, then I do not > care at all about the occupancy of T1 simply because that is running on a > seperate reservation. Ok, then we can give the cpu monitoring a priority just like CAT. Trying to make that an aggregated value in the first > place is completely wrong. If you want an aggregate, which is pretty much > useless, then user space tools can generate it easily. > > The whole approach you and David have taken is to whack some desired cgroup > functionality and whatever into CQM without rethinking the overall > design. And that's fundamentaly broken because it does not take cache (and > memory bandwidth) allocation into account. > > I seriously doubt, that the existing CQM/MBM code can be refactored in any > useful way. As Peter Zijlstra said before: Remove the existing cruft > completely and start with completely new design from scratch. I missed Peterz indicated new design from scratch. Was only bothered with the implementations given that CAt was still going on. Since CAT is up now we may be able to do better. Thanks, Vikas > > And this new design should start from the allocation angle and then add the > whole other muck on top so far its possible. Allocation related monitoring > must be the primary focus, everything else is just tinkering. > > Thanks, > > tglx > > > > > > > > >
[toc] | [prev] | [next] | [standalone]
| From | "Yu, Fenghua" <fenghua.yu@intel.com> |
|---|---|
| Date | 2017-01-18 22:20 +0100 |
| Message-ID | <t15O9-RQ-13@gated-at.bofh.it> |
| In reply to | #1561363 |
> From: Thomas Gleixner [mailto:tglx@linutronix.de] > On Tue, 17 Jan 2017, Shivappa Vikas wrote: > > On Tue, 17 Jan 2017, Thomas Gleixner wrote: > > > On Fri, 6 Jan 2017, Vikas Shivappa wrote: > > > > - Issue(1): Inaccurate data for per package data, systemwide. Just > > > > prints zeros or arbitrary numbers. > > > > > > > > Fix: Patches fix this by just throwing an error if the mode is not > > > > supported. > > > > The modes supported is task monitoring and cgroup monitoring. > > > > Also the per package > > > > data for say socket x is returned with the -C <cpu on socketx> -G > > > > cgrpy option. > > > > The systemwide data can be looked up by monitoring root cgroup. > > > > > > Fine. That just lacks any comment in the implementation. Otherwise I > > > would not have asked the question about cpu monitoring. Though I > > > fundamentaly hate the idea of requiring cgroups for this to work. > > > > > > If I just want to look at CPU X why on earth do I have to set up all > > > that cgroup muck? Just because your main focus is cgroups? > > > > The upstream per cpu data is broken because its not overriding the > > other task event RMIDs on that cpu with the cpu event RMID. > > > > Can be fixed by adding a percpu struct to hold the RMID thats > > affinitized to the cpu, however then we miss all the task > > llc_occupancy in that - still evaluating it. > > The point here is that CQM is closely connected to the cache allocation > technology. After a lengthy discussion we ended up having > > - per cpu CLOSID > - per task CLOSID > > where all tasks which do not have a CLOSID assigned use the CLOSID which is > assigned to the CPU they are running on. > > So if I configure a system by simply partitioning the cache per cpu, which is > the proper way to do it for HPC and RT usecases where workloads are > partitioned on CPUs as well, then I really want to have an equaly simple way > to monitor the occupancy for that reservation. > > And looking at that from the CAT point of view, which is the proper way to do > it, makes it obvious that CQM should be modeled to match CAT. > > So lets assume the following: > > CPU 0-3 default CLOSID 0 > CPU 4 CLOSID 1 > CPU 5 CLOSID 2 > CPU 6 CLOSID 3 > CPU 7 CLOSID 3 > > T1 CLOSID 4 > T2 CLOSID 5 > T3 CLOSID 6 > T4 CLOSID 6 > > All other tasks use the per cpu defaults, i.e. the CLOSID of the CPU > they run on. > > then the obvious basic monitoring requirement is to have a RMID for each > CLOSID. So the mapping between RMID and CLOSID is 1:1 mapping, right? Then changing the current resctrl interface in kernel as follows: 1. In rdtgroup_mkdir() (i.e. creating a partition in resctrl), allocate one RMID for the partition. Then the mapping between RMID and CLOSID is set up in mkdir. 2. In rdtgroup_rmdir() (i.e. removing a partition in resctrl), free the RMID. Then the mapping between RMID and CLOSID is dismissed. In user space: 1. Create a partition in resctrl and allocate L3 CBM in schemata and assign a PID in "tasks". 2. Start a user monitoring tool (e.g. perf) to monitor the PID. The monitoring tool needs to be updated to know resctrl interface. We may update perf to work with resctrl interface. Since the PID is assigned to the partition which has a CLOSID and RMID mapping, the PID is monitored while it's running in the allocated portion of L3. Is above proposal the right way to go? Thanks. -Fenghua
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-01-18 22:20 +0100 |
| Message-ID | <t15Oa-RQ-19@gated-at.bofh.it> |
| In reply to | #1561363 |
On Wed, Jan 18, 2017 at 12:53 AM, Thomas Gleixner <tglx@linutronix.de> wrote: > On Tue, 17 Jan 2017, Shivappa Vikas wrote: >> On Tue, 17 Jan 2017, Thomas Gleixner wrote: >> > On Fri, 6 Jan 2017, Vikas Shivappa wrote: >> > > - Issue(1): Inaccurate data for per package data, systemwide. Just prints >> > > zeros or arbitrary numbers. >> > > >> > > Fix: Patches fix this by just throwing an error if the mode is not >> > > supported. >> > > The modes supported is task monitoring and cgroup monitoring. >> > > Also the per package >> > > data for say socket x is returned with the -C <cpu on socketx> -G cgrpy >> > > option. >> > > The systemwide data can be looked up by monitoring root cgroup. >> > >> > Fine. That just lacks any comment in the implementation. Otherwise I would >> > not have asked the question about cpu monitoring. Though I fundamentaly >> > hate the idea of requiring cgroups for this to work. >> > >> > If I just want to look at CPU X why on earth do I have to set up all that >> > cgroup muck? Just because your main focus is cgroups? >> >> The upstream per cpu data is broken because its not overriding the other task >> event RMIDs on that cpu with the cpu event RMID. >> >> Can be fixed by adding a percpu struct to hold the RMID thats affinitized >> to the cpu, however then we miss all the task llc_occupancy in that - still >> evaluating it. > > The point here is that CQM is closely connected to the cache allocation > technology. After a lengthy discussion we ended up having > > - per cpu CLOSID > - per task CLOSID > > where all tasks which do not have a CLOSID assigned use the CLOSID which is > assigned to the CPU they are running on. > > So if I configure a system by simply partitioning the cache per cpu, which > is the proper way to do it for HPC and RT usecases where workloads are > partitioned on CPUs as well, then I really want to have an equaly simple > way to monitor the occupancy for that reservation. > > And looking at that from the CAT point of view, which is the proper way to > do it, makes it obvious that CQM should be modeled to match CAT. > > So lets assume the following: > > CPU 0-3 default CLOSID 0 > CPU 4 CLOSID 1 > CPU 5 CLOSID 2 > CPU 6 CLOSID 3 > CPU 7 CLOSID 3 > > T1 CLOSID 4 > T2 CLOSID 5 > T3 CLOSID 6 > T4 CLOSID 6 > > All other tasks use the per cpu defaults, i.e. the CLOSID of the CPU > they run on. > > then the obvious basic monitoring requirement is to have a RMID for each > CLOSID. There are use cases where the RMID to CLOSID mapping is not that simple. Some of them are: 1. Fine-tuning of cache allocation. We may want to have a CLOSID for a thread during phases that initialize relevant data, while changing it to another during phases that pollute cache. Yet, we want the RMID to remain the same. A different variation is to change CLOSID to increase/decrease the size of the allocated cache when high/low contention is detected. 2. Contention detection. I start with: - T1 has RMID 1. - T1 changes RMID to 2. will expect llc_occupancy(1) to decrease while llc_occupancy(2) increases. The rate of change will be relative to the level of cache contention present at the time. This all happens without changing the CLOSID. > > So when I monitor CPU4, i.e. CLOSID 1 and T1 runs on CPU4, then I do not > care at all about the occupancy of T1 simply because that is running on a > seperate reservation. It is not useless for scenarios where CLOSID and RMIDs change dynamically See above. > Trying to make that an aggregated value in the first > place is completely wrong. If you want an aggregate, which is pretty much > useless, then user space tools can generate it easily. Not useless, see above. Having user space tools to aggregate implies wasting some of the already scarce RMIDs. > > The whole approach you and David have taken is to whack some desired cgroup > functionality and whatever into CQM without rethinking the overall > design. And that's fundamentaly broken because it does not take cache (and > memory bandwidth) allocation into account. Monitoring and allocation are closely related yet independent. I see the advantages of allowing a per-cpu RMID as you describe in the example. Yet, RMIDs and CLOSIDs should remain independent to allow use cases beyond one simply monitoring occupancy per allocation. > > I seriously doubt, that the existing CQM/MBM code can be refactored in any > useful way. As Peter Zijlstra said before: Remove the existing cruft > completely and start with completely new design from scratch. > > And this new design should start from the allocation angle and then add the > whole other muck on top so far its possible. Allocation related monitoring > must be the primary focus, everything else is just tinkering. Assuming that my stated need for more than one RMID per CLOSID or more than one CLOSID per RMID is recognized, what would be the advantage of starting the design of monitoring from the allocation perspective? It's quite doable to create a new version of CQM/CMT without all the cgroup murk. We can also create an easy way to open events to monitor CLOSIDs. Yet, I don't see the advantage of dissociating monitoring from perf and directly building in on top of allocation without the assumption of 1 CLOSID : 1 RMID. Thanks, David > > Thanks, > > tglx > > > > > > > >
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-01-19 18:50 +0100 |
| Message-ID | <t1p0t-4y6-1@gated-at.bofh.it> |
| In reply to | #1562184 |
On Wed, 18 Jan 2017, David Carrillo-Cisneros wrote: > On Wed, Jan 18, 2017 at 12:53 AM, Thomas Gleixner <tglx@linutronix.de> wrote: > There are use cases where the RMID to CLOSID mapping is not that simple. > Some of them are: > > 1. Fine-tuning of cache allocation. We may want to have a CLOSID for a thread > during phases that initialize relevant data, while changing it to another during > phases that pollute cache. Yet, we want the RMID to remain the same. That's fine. I did not say that you need fixed RMD <-> CLOSID mappings. The point is that monitoring across different CLOSID domains is pointless. I have no idea how you want to do that with the proposed implementation to switch the RMID of the thread on the fly, but that's a different story. > A different variation is to change CLOSID to increase/decrease the size of the > allocated cache when high/low contention is detected. > > 2. Contention detection. I start with: > - T1 has RMID 1. > - T1 changes RMID to 2. > will expect llc_occupancy(1) to decrease while llc_occupancy(2) increases. Of course does RMID1 decrease because it's not longer in use. Oh well. > The rate of change will be relative to the level of cache contention present > at the time. This all happens without changing the CLOSID. See above. > > > > So when I monitor CPU4, i.e. CLOSID 1 and T1 runs on CPU4, then I do not > > care at all about the occupancy of T1 simply because that is running on a > > seperate reservation. > > It is not useless for scenarios where CLOSID and RMIDs change dynamically > See above. Above you are talking about the same CLOSID and different RMIDS and not about changing both. > > Trying to make that an aggregated value in the first > > place is completely wrong. If you want an aggregate, which is pretty much > > useless, then user space tools can generate it easily. > > Not useless, see above. It is prettey useless, because CPU4 has CLOSID1 while T1 has CLOSID4 and making an aggregate over those two has absolutely nothing to do with your scenario above. If you want the aggregate value, then create it in user space and oracle (or should I say google) out of it whatever you want, but do not impose that to the kernel. > Having user space tools to aggregate implies wasting some of the already > scarce RMIDs. Oh well. Can you please explain how you want to monitor the scenario I explained above: CPU4 CLOSID 1 T1 CLOSID 4 So if T1 runs on CPU4 then it uses CLOSID 4 which does not at all affect the cache occupancy of CLOSID 1. So if you use the same RMID then you pollute either the information of CPU4 (CLOSID1) or the information of T1 (CLOSID4) To gather any useful information for both CPU1 and T1 you need TWO RMIDs. Everything else is voodoo and crystal ball analysis and we are not going to support that. > > The whole approach you and David have taken is to whack some desired cgroup > > functionality and whatever into CQM without rethinking the overall > > design. And that's fundamentaly broken because it does not take cache (and > > memory bandwidth) allocation into account. > > Monitoring and allocation are closely related yet independent. Independent to some degree. Sure you can claim they are completely independent, but lots of the resulting combinations make absolutely no sense at all. And we really don't want to support non-sensical measurements just because we can. The outcome of this is complexity, inaccuracy and code which is too horrible to look at. > I see the advantages of allowing a per-cpu RMID as you describe in the example. > > Yet, RMIDs and CLOSIDs should remain independent to allow use cases beyond > one simply monitoring occupancy per allocation. I agree there are use cases where you want to monitor across allocations, like monitoring a task which has no CLOSID assigned and runs on different CPUs and therefor potentially on different CLOSIDs which are assigned to the different CPUs. That's fine and you want a seperate RMID for this. But once you have a fixed CLOSID association then reusing and aggregating across CLOSID domains is more than useless. > > I seriously doubt, that the existing CQM/MBM code can be refactored in any > > useful way. As Peter Zijlstra said before: Remove the existing cruft > > completely and start with completely new design from scratch. > > > > And this new design should start from the allocation angle and then add the > > whole other muck on top so far its possible. Allocation related monitoring > > must be the primary focus, everything else is just tinkering. > > Assuming that my stated need for more than one RMID per CLOSID or more > than one CLOSID per RMID is recognized, what would be the advantage of > starting the design of monitoring from the allocation perspective? > > It's quite doable to create a new version of CQM/CMT without all the > cgroup murk. > > We can also create an easy way to open events to monitor CLOSIDs. Yet, I > don't see the advantage of dissociating monitoring from perf and directly > building in on top of allocation without the assumption of 1 CLOSID : 1 > RMID. I did not say that you need to remove it from perf. perf is still going to be the interface to interact with monitoring, but it needs to be done in a way which makes sense. The current cgroup focussed proposal which is completely oblivious of the allocation mechanism does not make any sense to me at all. Starting the design from the allocation POV makes a lot of sense because that's the point where you start to make the decisions about useful and useless monitoring choices. And limiting the choices is the best way to limit the RMID exhaustion in the first place. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-01-20 08:50 +0100 |
| Message-ID | <t1C7n-4pX-9@gated-at.bofh.it> |
| In reply to | #1562975 |
On Thu, Jan 19, 2017 at 9:41 AM, Thomas Gleixner <tglx@linutronix.de> wrote: > On Wed, 18 Jan 2017, David Carrillo-Cisneros wrote: >> On Wed, Jan 18, 2017 at 12:53 AM, Thomas Gleixner <tglx@linutronix.de> wrote: >> There are use cases where the RMID to CLOSID mapping is not that simple. >> Some of them are: >> >> 1. Fine-tuning of cache allocation. We may want to have a CLOSID for a thread >> during phases that initialize relevant data, while changing it to another during >> phases that pollute cache. Yet, we want the RMID to remain the same. > > That's fine. I did not say that you need fixed RMD <-> CLOSID mappings. The > point is that monitoring across different CLOSID domains is pointless. > > I have no idea how you want to do that with the proposed implementation to > switch the RMID of the thread on the fly, but that's a different story. > >> A different variation is to change CLOSID to increase/decrease the size of the >> allocated cache when high/low contention is detected. >> >> 2. Contention detection. I start with: >> - T1 has RMID 1. >> - T1 changes RMID to 2. >> will expect llc_occupancy(1) to decrease while llc_occupancy(2) increases. > > Of course does RMID1 decrease because it's not longer in use. Oh well. > >> The rate of change will be relative to the level of cache contention present >> at the time. This all happens without changing the CLOSID. > > See above. > >> > >> > So when I monitor CPU4, i.e. CLOSID 1 and T1 runs on CPU4, then I do not >> > care at all about the occupancy of T1 simply because that is running on a >> > seperate reservation. >> >> It is not useless for scenarios where CLOSID and RMIDs change dynamically >> See above. > > Above you are talking about the same CLOSID and different RMIDS and not > about changing both. The scenario I talked about implies changing CLOSID without affecting monitoring. It happens when the allocation needs for a thread/cgroup/CPU change dynamically. Forcing to change the RMID together with the CLOSID would give wrong monitoring values unless the old RMID is kept around until becomes free, which is ugly and would waste a RMID. > >> > Trying to make that an aggregated value in the first >> > place is completely wrong. If you want an aggregate, which is pretty much >> > useless, then user space tools can generate it easily. >> >> Not useless, see above. > > It is prettey useless, because CPU4 has CLOSID1 while T1 has CLOSID4 and > making an aggregate over those two has absolutely nothing to do with your > scenario above. That's true. It is useless in the case you mentioned. I erroneously interpreted the "useless" in your comment as a general statement about aggregating RMID occupancies. > > If you want the aggregate value, then create it in user space and oracle > (or should I say google) out of it whatever you want, but do not impose > that to the kernel. > >> Having user space tools to aggregate implies wasting some of the already >> scarce RMIDs. > > Oh well. Can you please explain how you want to monitor the scenario I > explained above: > > CPU4 CLOSID 1 > T1 CLOSID 4 > > So if T1 runs on CPU4 then it uses CLOSID 4 which does not at all affect > the cache occupancy of CLOSID 1. So if you use the same RMID then you > pollute either the information of CPU4 (CLOSID1) or the information of T1 > (CLOSID4) > > To gather any useful information for both CPU1 and T1 you need TWO > RMIDs. Everything else is voodoo and crystal ball analysis and we are not > going to support that. > Correct. Yet, having two RMIDs to monitor the same task/cgroup/CPU just because the CLOSID changed is wasteful. >> > The whole approach you and David have taken is to whack some desired cgroup >> > functionality and whatever into CQM without rethinking the overall >> > design. And that's fundamentaly broken because it does not take cache (and >> > memory bandwidth) allocation into account. >> >> Monitoring and allocation are closely related yet independent. > > Independent to some degree. Sure you can claim they are completely > independent, but lots of the resulting combinations make absolutely no > sense at all. And we really don't want to support non-sensical measurements > just because we can. The outcome of this is complexity, inaccuracy and code > which is too horrible to look at. > >> I see the advantages of allowing a per-cpu RMID as you describe in the example. >> >> Yet, RMIDs and CLOSIDs should remain independent to allow use cases beyond >> one simply monitoring occupancy per allocation. > > I agree there are use cases where you want to monitor across allocations, > like monitoring a task which has no CLOSID assigned and runs on different > CPUs and therefor potentially on different CLOSIDs which are assigned to > the different CPUs. > > That's fine and you want a seperate RMID for this. > > But once you have a fixed CLOSID association then reusing and aggregating > across CLOSID domains is more than useless. > Correct. But there may not be a fixed CLOSID association if loads exhibit dynamic behavior and/or system load changes dynamically. >> > I seriously doubt, that the existing CQM/MBM code can be refactored in any >> > useful way. As Peter Zijlstra said before: Remove the existing cruft >> > completely and start with completely new design from scratch. >> > >> > And this new design should start from the allocation angle and then add the >> > whole other muck on top so far its possible. Allocation related monitoring >> > must be the primary focus, everything else is just tinkering. >> >> Assuming that my stated need for more than one RMID per CLOSID or more >> than one CLOSID per RMID is recognized, what would be the advantage of >> starting the design of monitoring from the allocation perspective? >> >> It's quite doable to create a new version of CQM/CMT without all the >> cgroup murk. >> >> We can also create an easy way to open events to monitor CLOSIDs. Yet, I >> don't see the advantage of dissociating monitoring from perf and directly >> building in on top of allocation without the assumption of 1 CLOSID : 1 >> RMID. > > I did not say that you need to remove it from perf. perf is still going to > be the interface to interact with monitoring, but it needs to be done in a > way which makes sense. The current cgroup focussed proposal which is > completely oblivious of the allocation mechanism does not make any sense to > me at all. > > Starting the design from the allocation POV makes a lot of sense because > that's the point where you start to make the decisions about useful and > useless monitoring choices. And limiting the choices is the best way to > limit the RMID exhaustion in the first place. Thanks for the extra explanation. David > > Thanks, > > tglx > >
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-01-20 09:40 +0100 |
| Message-ID | <t1CTM-4W3-21@gated-at.bofh.it> |
| In reply to | #1563343 |
On Thu, 19 Jan 2017, David Carrillo-Cisneros wrote: > On Thu, Jan 19, 2017 at 9:41 AM, Thomas Gleixner <tglx@linutronix.de> wrote: > > Above you are talking about the same CLOSID and different RMIDS and not > > about changing both. > > The scenario I talked about implies changing CLOSID without affecting > monitoring. > It happens when the allocation needs for a thread/cgroup/CPU change > dynamically. Forcing to change the RMID together with the CLOSID would > give wrong monitoring values unless the old RMID is kept around until > becomes free, which is ugly and would waste a RMID. When the allocation needs for a resource control group change, then we simply update the allocation constraints of that group without chaning the CLOSID. So everything just stays the same. If you move entities to a different group then of course the CLOSID changes and then it's a different story how to deal with monitoring. > > To gather any useful information for both CPU1 and T1 you need TWO > > RMIDs. Everything else is voodoo and crystal ball analysis and we are not > > going to support that. > > > > Correct. Yet, having two RMIDs to monitor the same task/cgroup/CPU > just because the CLOSID changed is wasteful. Again, the CLOSID only changes if you move entities to a different resource control group and in that case the RMID change is the least of your worries. > Correct. But there may not be a fixed CLOSID association if loads > exhibit dynamic behavior and/or system load changes dynamically. So, you really want to move entities around between resource control groups dynamically? I'm not seing why you would want to do that, but I'm all ear to get educated on that. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-01-20 21:30 +0100 |
| Message-ID | <t1NYS-3rw-19@gated-at.bofh.it> |
| In reply to | #1563381 |
On Fri, Jan 20, 2017 at 12:30 AM, Thomas Gleixner <tglx@linutronix.de> wrote: > On Thu, 19 Jan 2017, David Carrillo-Cisneros wrote: >> On Thu, Jan 19, 2017 at 9:41 AM, Thomas Gleixner <tglx@linutronix.de> wrote: >> > Above you are talking about the same CLOSID and different RMIDS and not >> > about changing both. >> >> The scenario I talked about implies changing CLOSID without affecting >> monitoring. >> It happens when the allocation needs for a thread/cgroup/CPU change >> dynamically. Forcing to change the RMID together with the CLOSID would >> give wrong monitoring values unless the old RMID is kept around until >> becomes free, which is ugly and would waste a RMID. > > When the allocation needs for a resource control group change, then we > simply update the allocation constraints of that group without chaning the > CLOSID. So everything just stays the same. > > If you move entities to a different group then of course the CLOSID > changes and then it's a different story how to deal with monitoring. > >> > To gather any useful information for both CPU1 and T1 you need TWO >> > RMIDs. Everything else is voodoo and crystal ball analysis and we are not >> > going to support that. >> > >> >> Correct. Yet, having two RMIDs to monitor the same task/cgroup/CPU >> just because the CLOSID changed is wasteful. > > Again, the CLOSID only changes if you move entities to a different resource > control group and in that case the RMID change is the least of your worries. > >> Correct. But there may not be a fixed CLOSID association if loads >> exhibit dynamic behavior and/or system load changes dynamically. > > So, you really want to move entities around between resource control groups > dynamically? I'm not seing why you would want to do that, but I'm all ear > to get educated on that. No, I don't want to move entities across resource control groups. I was confused by the idea of CLOSIDs being married to control groups, but now is clear even to me that that was never the intention. Thanks, David > > Thanks, > > tglx
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-01-19 03:20 +0100 |
| Message-ID | <t1auu-3KN-13@gated-at.bofh.it> |
| In reply to | #1561363 |
On Wed, Jan 18, 2017 at 12:53 AM, Thomas Gleixner <tglx@linutronix.de> wrote: > On Tue, 17 Jan 2017, Shivappa Vikas wrote: >> On Tue, 17 Jan 2017, Thomas Gleixner wrote: >> > On Fri, 6 Jan 2017, Vikas Shivappa wrote: >> > > - Issue(1): Inaccurate data for per package data, systemwide. Just prints >> > > zeros or arbitrary numbers. >> > > >> > > Fix: Patches fix this by just throwing an error if the mode is not >> > > supported. >> > > The modes supported is task monitoring and cgroup monitoring. >> > > Also the per package >> > > data for say socket x is returned with the -C <cpu on socketx> -G cgrpy >> > > option. >> > > The systemwide data can be looked up by monitoring root cgroup. >> > >> > Fine. That just lacks any comment in the implementation. Otherwise I would >> > not have asked the question about cpu monitoring. Though I fundamentaly >> > hate the idea of requiring cgroups for this to work. >> > >> > If I just want to look at CPU X why on earth do I have to set up all that >> > cgroup muck? Just because your main focus is cgroups? >> >> The upstream per cpu data is broken because its not overriding the other task >> event RMIDs on that cpu with the cpu event RMID. >> >> Can be fixed by adding a percpu struct to hold the RMID thats affinitized >> to the cpu, however then we miss all the task llc_occupancy in that - still >> evaluating it. > > The point here is that CQM is closely connected to the cache allocation > technology. After a lengthy discussion we ended up having > > - per cpu CLOSID > - per task CLOSID > > where all tasks which do not have a CLOSID assigned use the CLOSID which is > assigned to the CPU they are running on. > > So if I configure a system by simply partitioning the cache per cpu, which > is the proper way to do it for HPC and RT usecases where workloads are > partitioned on CPUs as well, then I really want to have an equaly simple > way to monitor the occupancy for that reservation. > > And looking at that from the CAT point of view, which is the proper way to > do it, makes it obvious that CQM should be modeled to match CAT. > > So lets assume the following: > > CPU 0-3 default CLOSID 0 > CPU 4 CLOSID 1 > CPU 5 CLOSID 2 > CPU 6 CLOSID 3 > CPU 7 CLOSID 3 > > T1 CLOSID 4 > T2 CLOSID 5 > T3 CLOSID 6 > T4 CLOSID 6 > > All other tasks use the per cpu defaults, i.e. the CLOSID of the CPU > they run on. > > then the obvious basic monitoring requirement is to have a RMID for each > CLOSID. > > So when I monitor CPU4, i.e. CLOSID 1 and T1 runs on CPU4, then I do not > care at all about the occupancy of T1 simply because that is running on a > seperate reservation. Trying to make that an aggregated value in the first > place is completely wrong. If you want an aggregate, which is pretty much > useless, then user space tools can generate it easily. > > The whole approach you and David have taken is to whack some desired cgroup > functionality and whatever into CQM without rethinking the overall > design. And that's fundamentaly broken because it does not take cache (and > memory bandwidth) allocation into account. > > I seriously doubt, that the existing CQM/MBM code can be refactored in any > useful way. As Peter Zijlstra said before: Remove the existing cruft > completely and start with completely new design from scratch. > > And this new design should start from the allocation angle and then add the > whole other muck on top so far its possible. Allocation related monitoring > must be the primary focus, everything else is just tinkering. > If in this email you meant "Resource group" where you wrote "CLOSID", then please disregard my previous email. It seems like a good idea to me to have a 1:1 mapping between RMIDs and "Resource groups". The distinction matter because changing the schemata in the resource group would likely trigger a change of CLOSID, which is useful. Thanks, David
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-01-19 18:30 +0100 |
| Message-ID | <t1oH7-4qY-11@gated-at.bofh.it> |
| In reply to | #1562381 |
On Wed, Jan 18, 2017 at 6:09 PM, David Carrillo-Cisneros <davidcc@google.com> wrote: > On Wed, Jan 18, 2017 at 12:53 AM, Thomas Gleixner <tglx@linutronix.de> wrote: >> On Tue, 17 Jan 2017, Shivappa Vikas wrote: >>> On Tue, 17 Jan 2017, Thomas Gleixner wrote: >>> > On Fri, 6 Jan 2017, Vikas Shivappa wrote: >>> > > - Issue(1): Inaccurate data for per package data, systemwide. Just prints >>> > > zeros or arbitrary numbers. >>> > > >>> > > Fix: Patches fix this by just throwing an error if the mode is not >>> > > supported. >>> > > The modes supported is task monitoring and cgroup monitoring. >>> > > Also the per package >>> > > data for say socket x is returned with the -C <cpu on socketx> -G cgrpy >>> > > option. >>> > > The systemwide data can be looked up by monitoring root cgroup. >>> > >>> > Fine. That just lacks any comment in the implementation. Otherwise I would >>> > not have asked the question about cpu monitoring. Though I fundamentaly >>> > hate the idea of requiring cgroups for this to work. >>> > >>> > If I just want to look at CPU X why on earth do I have to set up all that >>> > cgroup muck? Just because your main focus is cgroups? >>> >>> The upstream per cpu data is broken because its not overriding the other task >>> event RMIDs on that cpu with the cpu event RMID. >>> >>> Can be fixed by adding a percpu struct to hold the RMID thats affinitized >>> to the cpu, however then we miss all the task llc_occupancy in that - still >>> evaluating it. >> >> The point here is that CQM is closely connected to the cache allocation >> technology. After a lengthy discussion we ended up having >> >> - per cpu CLOSID >> - per task CLOSID >> >> where all tasks which do not have a CLOSID assigned use the CLOSID which is >> assigned to the CPU they are running on. >> >> So if I configure a system by simply partitioning the cache per cpu, which >> is the proper way to do it for HPC and RT usecases where workloads are >> partitioned on CPUs as well, then I really want to have an equaly simple >> way to monitor the occupancy for that reservation. >> >> And looking at that from the CAT point of view, which is the proper way to >> do it, makes it obvious that CQM should be modeled to match CAT. >> >> So lets assume the following: >> >> CPU 0-3 default CLOSID 0 >> CPU 4 CLOSID 1 >> CPU 5 CLOSID 2 >> CPU 6 CLOSID 3 >> CPU 7 CLOSID 3 >> >> T1 CLOSID 4 >> T2 CLOSID 5 >> T3 CLOSID 6 >> T4 CLOSID 6 >> >> All other tasks use the per cpu defaults, i.e. the CLOSID of the CPU >> they run on. >> >> then the obvious basic monitoring requirement is to have a RMID for each >> CLOSID. >> >> So when I monitor CPU4, i.e. CLOSID 1 and T1 runs on CPU4, then I do not >> care at all about the occupancy of T1 simply because that is running on a >> seperate reservation. Trying to make that an aggregated value in the first >> place is completely wrong. If you want an aggregate, which is pretty much >> useless, then user space tools can generate it easily. >> >> The whole approach you and David have taken is to whack some desired cgroup >> functionality and whatever into CQM without rethinking the overall >> design. And that's fundamentaly broken because it does not take cache (and >> memory bandwidth) allocation into account. >> >> I seriously doubt, that the existing CQM/MBM code can be refactored in any >> useful way. As Peter Zijlstra said before: Remove the existing cruft >> completely and start with completely new design from scratch. >> >> And this new design should start from the allocation angle and then add the >> whole other muck on top so far its possible. Allocation related monitoring >> must be the primary focus, everything else is just tinkering. >> > > If in this email you meant "Resource group" where you wrote "CLOSID", then > please disregard my previous email. It seems like a good idea to me to have > a 1:1 mapping between RMIDs and "Resource groups". > > The distinction matter because changing the schemata in the resource group > would likely trigger a change of CLOSID, which is useful. > Just realized that the sharing of CLOSIDs is not part of the accepted version of RDT. My mental model was still on the old CAT driver that did allow sharing of CLOSIDs between cgroups. Now I understand why CLOSID was assumed to be equal with "Resource groups". Sorry for the noise. Then the comments in my previous email hold. In summary and addition to latest emails: A 1:1 mapping between CLOSID/"Resource group" to RMID, as Fenghua suggested is very problematic because the number of CLOSIDs is much much smaller than the number of RMIDs, and, as Stephane mentioned it's a common use case to want to independently monitor many task/cgroups inside an allocation partition. A 1:many mapping of CLOSID to RMIDs may work as a cheap replacement of cgroup monitoring but the case where CLOSID changes would be messy. In llc_occupancy, if RMIDs are changed, old RMIDs still hold valid occupancy for indefinite time, so either RMIDs would be preserved (breaking the 1:many mapping) or old RMIDs should be tracked while they are dirty. Thanks, David
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-01-19 19:50 +0100 |
| Message-ID | <t1pWy-58h-21@gated-at.bofh.it> |
| In reply to | #1562959 |
On Thu, 19 Jan 2017, David Carrillo-Cisneros wrote: > A 1:1 mapping between CLOSID/"Resource group" to RMID, as Fenghua suggested > is very problematic because the number of CLOSIDs is much much smaller than the > number of RMIDs, and, as Stephane mentioned it's a common use case to want to > independently monitor many task/cgroups inside an allocation partition. Again, that was not my intention. I just want to limit the combinations. > A 1:many mapping of CLOSID to RMIDs may work as a cheap replacement of > cgroup monitoring but the case where CLOSID changes would be messy. In CLOSIDs of RDT groups do not change. They are allocated when the group is created. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Vikas Shivappa <vikas.shivappa@linux.intel.com> |
|---|---|
| Date | 2017-01-19 03:30 +0100 |
| Message-ID | <t1aE9-3NW-7@gated-at.bofh.it> |
| In reply to | #1561363 |
Based on Thomas and Peterz feedback Can think of two variants which target:
-Support monitoring and allocating using the same resctrl group.
user can use a resctrl group to allocate resources and also monitor
them (with respect to tasks or cpu)
-allows 'task only' monitoring outside of resctrl. This mode can be used
when user wants to override the RMIDs in the resctrl or when he wants
to monitor more than just the resctrl groups.
option 1> without modifying the resctrl
In this design everything in resctrl interface works like
before (the info, resource group files like task schemata all remain the
same) but the resctrl groups are mapped to one RMID as well as a CLOSID.
But we need a user interface for the user to read the counters. We can
create one file to set monitoring and one
file in resctrl directory which will reflect the counts but may not be
efficient as a lot of times user reads the counts frequently.
For the user interface there may be two options to do this:
1.a> Build a new user mode interface resmon
Since modifying the existing perf to
suit the different h/w architecture seems to not follow the CAT
interface model, it may well be better to have a different and dedicated
interface for the RDT monitoring (just like we had a new fs for CAT)
$resmon -r <resctrl group> -s <mon_mask> -I <time in ms>
"resctrl group": is the resctrl directory.
"mon_mask: is a bit mask of logical packages which indicates which packages user is
interested in monitoring.
"time in ms": The time for which the monitoring takes place
(this can potentially be changed to start and stop/read options)
Example 1 (Some examples modeled from resctrl ui documentation)
---------
A single socket system which has real-time tasks running on core 4-7 and
non real-time workload assigned to core 0-3. The real-time tasks share
text and data, so a per task association is not required and due to
interaction with the kernel it's desired that the kernel on these cores shares L3
with the tasks.
# cd /sys/fs/resctrl
# echo "L3:0=3ff" > schemata
core 0-1 are assigned to the new group and make sure that the
kernel and the tasks running there get 50% of the cache.
# echo 03 > p0/cpus
monitor the cpus 0-1 for 5s
# resmon -r p0 -s 1 -I 5000
Example 2
---------
A real time task running on cpu 2-3(socket 0) is allocated a dedicated 25% of the
cache.
# cd /sys/fs/resctrl
# mkdir p1
# echo "L3:0=0f00;1=ffff" > p1/schemata
# echo 5678 > p1/tasks
# taskset -cp 2-3 5678
Monitor the task for 5s on socket zero
# resmon -r p0 -s 1 -I 5000
Example 3
---------
sometimes user may just want to profile the cache occupancy first before
assigning any CLOSids. Also this provides an override option where user
can monitor some tasks which have say CLOS 0 that he is about to place
in a CLOSId based on the amount of cache occupancy. This could apply to
the same real time tasks above where user is caliberating the % of cache
thats needed.
monitor a task PIDx on socket 0 for 10s
# resmon -t PIDx -s 1 -I 10000
1.b> Add a new option to perf apart from supporting the task monitoring
in perf.
- Monitor a resctrl group.
Introduce a new option for perf "-R" which indicates to monitor a
resctrl group.
$mkdir /sys/fs/resctrl/p1
$echo PID1 > /sys/fs/resctrl/p1/tasks
$echo PID2 > /sys/fs/resctrl/p1/tasks
$perf stat -e llc_occupancy -R p1
would return the count for the resctrl group p1.
- Monitor a task outside of resctrl group ('task only')
In this case , the perf can also monitor individual tasks using the -t
option just like before.
$perf stat -e llc_occupancy -t p1
- Monitor CPUs.
For the example 1 above , perf can be used to monitor the resctrl group
p0
$perf stat -e llc_occupancy -t p0
The issue with both options may be what happens when we run out of
RMIDs. For the resctrl groups , since we know the max groups that can be
created and the # of CLOSIds is very less compared to # of RMIDs we
reserve an RMID for each resctrl group so there is never a case that
RMID is not available for resctrl group.
For task monitoring , it can use the rest of the RMIDs.
Why do we need seperate 'task only' monitoring ?
-----------------------------------------
The seperate task monitoring option lets
the user use the RMIDs effectively and not be restricted to # of
CLOSids. Also deal with the scenarios of example 3.
RMID allocation/init
--------------------
resctrl monitoring:
RMIDs are allocated when CLOSIds are allocated during mkdir. One RMId is
allocated per socket just like CLOSid.
task monitoring:
When task events are created, RMIDs are allocated. Can also do a lazy
allocation of RMIDs when the tasks are actually scheduled in on a
socket.
Kernel Scheduling
-----------------
During ctx switch cqm choses the RMID in the following priority (1-
highest priority)
1. if cpu has a RMID , choose that
2. if the task has a RMID directly tied to it choose that
3. choose the RMID of the task's resctrl
Read
----
When user calls cqm to retrieve the monitored count, we read the
counter_msr and return the count.
option 2> Modifying the resctrl
This changes the resctrl interface schemata where user inputs the
CLOSids and RMIDs instead of CBMs.
# cd /sys/fs/resctrl
# mkdir p0 p1
# echo "L3:0=<closidx>;1=<closidy>" > /sys/fs/resctrl/p0/schemata
There is a mapping between closid and cbm which the user can change.
# echo 0xff > .../config/L3/0/cbm
Display the CLOSids
# ls .../config/L3/
0
1
2
.
.
.
15
As an extension to cqm , this schemata can be modified to also have the
RMIDs be chosen by the user. That way user can configure different RMIDs
for the same CLOSid if needed like in example 3. and also since we have
so many more RMIDs than CLOSids , user is not restricted by the number
of resctrl groups he can create (With the current model, user cannot
create more directories than the number of CLOSIds)
# echo "L3:0=<closidx>,<RMID1>;1=<closidy>,<RMID2>" > /sys/fs/resctrl/p0/schemata
user interface to monitor can be same as shown in the design variant #1
with the difference that this may have a lesser need for the 'task only'
monitoring.
[toc] | [prev] | [next] | [standalone]
| From | Stephane Eranian <eranian@google.com> |
|---|---|
| Date | 2017-01-19 07:50 +0100 |
| Message-ID | <t1eHM-6qK-9@gated-at.bofh.it> |
| In reply to | #1561363 |
On Wed, Jan 18, 2017 at 12:53 AM, Thomas Gleixner <tglx@linutronix.de> wrote: > On Tue, 17 Jan 2017, Shivappa Vikas wrote: >> On Tue, 17 Jan 2017, Thomas Gleixner wrote: >> > On Fri, 6 Jan 2017, Vikas Shivappa wrote: >> > > - Issue(1): Inaccurate data for per package data, systemwide. Just prints >> > > zeros or arbitrary numbers. >> > > >> > > Fix: Patches fix this by just throwing an error if the mode is not >> > > supported. >> > > The modes supported is task monitoring and cgroup monitoring. >> > > Also the per package >> > > data for say socket x is returned with the -C <cpu on socketx> -G cgrpy >> > > option. >> > > The systemwide data can be looked up by monitoring root cgroup. >> > >> > Fine. That just lacks any comment in the implementation. Otherwise I would >> > not have asked the question about cpu monitoring. Though I fundamentaly >> > hate the idea of requiring cgroups for this to work. >> > >> > If I just want to look at CPU X why on earth do I have to set up all that >> > cgroup muck? Just because your main focus is cgroups? >> >> The upstream per cpu data is broken because its not overriding the other task >> event RMIDs on that cpu with the cpu event RMID. >> >> Can be fixed by adding a percpu struct to hold the RMID thats affinitized >> to the cpu, however then we miss all the task llc_occupancy in that - still >> evaluating it. > > The point here is that CQM is closely connected to the cache allocation > technology. After a lengthy discussion we ended up having > > - per cpu CLOSID > - per task CLOSID > > where all tasks which do not have a CLOSID assigned use the CLOSID which is > assigned to the CPU they are running on. > > So if I configure a system by simply partitioning the cache per cpu, which > is the proper way to do it for HPC and RT usecases where workloads are > partitioned on CPUs as well, then I really want to have an equaly simple > way to monitor the occupancy for that reservation. > Your use case is specific to HPC and not Web workloads we run. Jobs run in cgroups which may span all the CPUs of the machine. CAT may be used to partition the cache. Cgroups would run inside a partition. There may be multiple cgroups running in the same partition. I can understand the value of tracking occupancy per CLOSID, however that granularity is not enough for our use case. Inside a partition, we want to know the occupancy of each cgroup to be able to assign blame to the top consumer. Thus, there needs to be a way to monitor occupancy per cgroup. I'd like to understand how your proposal would cover this use case. Another important aspect is that CQM measures new allocations, thus to get total occupancy you need to be able to monitor the thread, CPU, CLOSid or cgroup from the beginning of execution. In the case of a cgroup from the moment where the first thread is scheduled into the cgroup. To do this a RMID needs to be assigned from the beginning to the entity to be monitored. It could be by creating a CQM event just to cause an RMID to be assigned as discussed earlier on this thread. And then if a perf stat is launched later it will get the same RMID and report full occupancy. But that requires the first event to remain alive, i.e., some process must keep the file descriptor open, i.e., need some daemon or a perf stat running in the background. There are also use cases where you want CQM without necessarily enabling CAT, for instance, if you want to know the cache footprint of a workload to estimate how if it could be co-located with others. I think any viable proposal needs to be able to support your use case as well as the ones I described above. > And looking at that from the CAT point of view, which is the proper way to > do it, makes it obvious that CQM should be modeled to match CAT. > > So lets assume the following: > > CPU 0-3 default CLOSID 0 > CPU 4 CLOSID 1 > CPU 5 CLOSID 2 > CPU 6 CLOSID 3 > CPU 7 CLOSID 3 > > T1 CLOSID 4 > T2 CLOSID 5 > T3 CLOSID 6 > T4 CLOSID 6 > > All other tasks use the per cpu defaults, i.e. the CLOSID of the CPU > they run on. > > then the obvious basic monitoring requirement is to have a RMID for each > CLOSID. > > So when I monitor CPU4, i.e. CLOSID 1 and T1 runs on CPU4, then I do not > care at all about the occupancy of T1 simply because that is running on a > seperate reservation. Trying to make that an aggregated value in the first > place is completely wrong. If you want an aggregate, which is pretty much > useless, then user space tools can generate it easily. > > The whole approach you and David have taken is to whack some desired cgroup > functionality and whatever into CQM without rethinking the overall > design. And that's fundamentaly broken because it does not take cache (and > memory bandwidth) allocation into account. > > I seriously doubt, that the existing CQM/MBM code can be refactored in any > useful way. As Peter Zijlstra said before: Remove the existing cruft > completely and start with completely new design from scratch. > > And this new design should start from the allocation angle and then add the > whole other muck on top so far its possible. Allocation related monitoring > must be the primary focus, everything else is just tinkering. > > Thanks, > > tglx > > > > > > > >
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-01-19 19:50 +0100 |
| Message-ID | <t1pWy-58h-23@gated-at.bofh.it> |
| In reply to | #1562436 |
On Wed, 18 Jan 2017, Stephane Eranian wrote: > On Wed, Jan 18, 2017 at 12:53 AM, Thomas Gleixner <tglx@linutronix.de> wrote: > > > Your use case is specific to HPC and not Web workloads we run. Jobs run > in cgroups which may span all the CPUs of the machine. CAT may be used > to partition the cache. Cgroups would run inside a partition. There may > be multiple cgroups running in the same partition. I can understand the > value of tracking occupancy per CLOSID, however that granularity is not > enough for our use case. Inside a partition, we want to know the > occupancy of each cgroup to be able to assign blame to the top > consumer. Thus, there needs to be a way to monitor occupancy per > cgroup. I'd like to understand how your proposal would cover this use > case. The point I'm making as I explained to David is that we need to start from the allocation angle. Of course can you monitor different tasks or task groups inside an allocation. > Another important aspect is that CQM measures new allocations, thus to > get total occupancy you need to be able to monitor the thread, CPU, > CLOSid or cgroup from the beginning of execution. In the case of a cgroup > from the moment where the first thread is scheduled into the cgroup. To > do this a RMID needs to be assigned from the beginning to the entity to > be monitored. It could be by creating a CQM event just to cause an RMID > to be assigned as discussed earlier on this thread. And then if a perf > stat is launched later it will get the same RMID and report full > occupancy. But that requires the first event to remain alive, i.e., some > process must keep the file descriptor open, i.e., need some daemon or a > perf stat running in the background. That's fine, but there must be a less convoluted way to do that. The currently proposed stuff is simply horrible because it lacks any form of design and is just hacked into submission. > There are also use cases where you want CQM without necessarily enabling > CAT, for instance, if you want to know the cache footprint of a workload > to estimate how if it could be co-located with others. That's a subset of the other stuff because it's all bound to CLOSID 0. So you can again monitor tasks or tasks groups seperately. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Vikas Shivappa <vikas.shivappa@linux.intel.com> |
|---|---|
| Date | 2017-01-20 03:40 +0100 |
| Message-ID | <t1xho-1jm-5@gated-at.bofh.it> |
| In reply to | #1561363 |
Resending including Thomas , also with some changes. Sorry for the spam
Based on Thomas and Peterz feedback Can think of two design
variants which target:
-Support monitoring and allocating using the same resctrl group.
user can use a resctrl group to allocate resources and also monitor
them (with respect to tasks or cpu)
-Also allows monitoring outside of resctrl so that user can
monitor subgroups who use the same closid. This mode can be used
when user wants to monitor more than just the resctrl groups.
The first design version uses and modifies perf_cgroup, second version
builds a new interface resmon. The first version is close to the patches
sent with some additions/changes. This includes details of the design as
per Thomas/Peterz feedback.
1> First Design option: without modifying the resctrl and using perf
--------------------------------------------------------------------
--------------------------------------------------------------------
In this design everything in resctrl interface works like
before (the info, resource group files like task schemata all remain the
same)
Monitor cqm using perf
----------------------
perf can monitor individual tasks using the -t
option just like before.
# perf stat -e llc_occupancy -t PID1,PID2
user can monitor the cpu occupancy using the -C option in perf:
# perf stat -e llc_occupancy -C 5
Below shows how user can monitor cgroup occupancy:
# mount -t cgroup -o perf_event perf_event /sys/fs/cgroup/perf_event/
# mkdir /sys/fs/cgroup/perf_event/g1
# mkdir /sys/fs/cgroup/perf_event/g2
# echo PID1 > /sys/fs/cgroup/perf_event/g2/tasks
# perf stat -e intel_cqm/llc_occupancy/ -a -G g2
To monitor a resctrl group, user can group the same tasks in resctrl
group into the cgroup.
To monitor the tasks in p1 in example 2 below, add the tasks in resctrl
group p1 to cgroup g1
# echo 5678 > /sys/fs/cgroup/perf_event/g1/tasks
Introducing a new option for resctrl may complicate monitoring because
supporting cgroup 'task groups' and resctrl 'task groups' leads to
situations where:
if the groups intersect, then there is no way to know what
l3_allocations contribute to which group.
ex:
p1 has tasks t1, t2, t3
g1 has tasks t2, t3, t4
The only way to get occupancy for g1 and p1 would be to allocate an RMID
for each task which can as well be done with the -t option.
Monitoring cqm cgroups Implementation
-------------------------------------
When monitoring two different cgroups in the same hierarchy (ex say g11
has an ancestor g1 which are both being monitored as shown below) we
need the g11 counts to be considered for g1 as well.
# mount -t cgroup -o perf_event perf_event /sys/fs/cgroup/perf_event/
# mkdir /sys/fs/cgroup/perf_event/g1
# mkdir /sys/fs/cgroup/perf_event/g1/g11
When measuring for g1 llc_occupancy we cannot write two different RMIDs
(because we need to count for g11 as well)
during context switch to measure the occupancy for both g1 and g11.
Hence the driver maintains this information and writes the RMID of the
lowest member in the ancestory which is being monitored during ctx
switch.
The cqm_info is added to the perf_cgroup structure to maintain this
information. The structure is allocated and destroyed at css_alloc and
css_free. All the events tied to a cgroup can use the same
information while reading the counts.
struct perf_cgroup {
#ifdef CONFIG_INTEL_RDT_M
void *cqm_info;
#endif
...
}
struct cqm_info {
bool mon_enabled;
int level;
u32 *rmid;
struct cgrp_cqm_info *mfa;
struct list_head tskmon_rlist;
};
Due to the hierarchical nature of cgroups, every cgroup just
monitors for the 'nearest monitored ancestor' at all times.
Since root cgroup is always monitored, all descendents
at boot time monitor for root and hence all mfa points to root
except for root->mfa which is NULL.
1. RMID setup: When cgroup x start monitoring:
for each descendent y, if y's mfa->level < x->level, then
y->mfa = x. (Where level of root node = 0...)
2. sched_in: During sched_in for x
if (x->mon_enabled) choose x->rmid
else choose x->mfa->rmid.
3. read: for each descendent of cgroup x
if (x->monitored) count += rmid_read(x->rmid).
4. evt_destroy: for each descendent y of x, if (y->mfa == x)
then y->mfa = x->mfa. Meaning if any descendent was monitoring for
x, set that descendent to monitor for the cgroup which x was
monitoring for.
To monitor a task in a cgroup x along with monitoring cgroup x itself
cqm_info maintains a list of tasks that are being monitored in the
cgroup.
When a task which belongs to a cgroup x is being monitored, it
always uses its own task->rmid even if cgroup x is monitored during sched_in.
To account for the counts of such tasks, cgroup keeps this list
and parses it during read.
taskmon_rlist is used to maintain the list. The list is modified when a
task is attached to the cgroup or removed from the group.
Example 1 (Some examples modeled from resctrl ui documentation)
---------
A single socket system which has real-time tasks running on core 4-7 and
non real-time workload assigned to core 0-3. The real-time tasks share
text and data, so a per task association is not required and due to
interaction with the kernel it's desired that the kernel on these cores shares L3
with the tasks.
# cd /sys/fs/resctrl
# echo "L3:0=3ff" > schemata
core 0-1 are assigned to the new group and make sure that the
kernel and the tasks running there get 50% of the cache.
# echo 03 > p0/cpus
monitor the cpus 0-1
# perf stat -e llc_occupancy -C 0-1
Example 2
---------
A real time task running on cpu 2-3(socket 0) is allocated a dedicated 25% of the
cache.
# cd /sys/fs/resctrl
# mkdir p1
# echo "L3:0=0f00;1=ffff" > p1/schemata
# echo 5678 > p1/tasks
# taskset -cp 2-3 5678
To monitor the same group of tasks create a cgroup g1
# mount -t cgroup -o perf_event perf_event /sys/fs/cgroup/perf_event/
# mkdir /sys/fs/cgroup/perf_event/g1
# perf stat -e llc_occupancy -a -G g1
Example 3
---------
sometimes user may just want to profile the cache occupancy first before
assigning any CLOSids. Also this provides an override option where user
can monitor some tasks which have say CLOS 0 that he is about to place
in a CLOSId based on the amount of cache occupancy. This could apply to
the same real time tasks above where user is caliberating the % of cache
thats needed.
# perf stat -e llc_occupancy -t PIDx,PIDy
RMID allocation
---------------
RMIDs are allocated per package to achieve better scaling of RMIDs.
RMIDs are plenty (2-4 per logical processor) and also are per package
meaning a two socket system would have twice the number of RMIDs.
If we still run out of RMIDs an error is thrown that monitoring wasnt
possible as the RMID wasnt available.
Kernel Scheduling
-----------------
During ctx switch cqm choses the RMID in the following priority
1. if cpu has a RMID , choose that
2. if the task has a RMID directly tied to it choose that (task is
monitored)
3. choose the RMID of the task's cgroup (by default tasks belong to root
cgroup with RMID 0)
Read
----
When user calls cqm to retrieve the monitored count, we read the
counter_msr and return the count. For cgroup hierarcy , the count is
measured as explained in the cgroup implementation section by traversing
the cgroup hierarchy.
2> Second Design option: Build a new usermode tool resmon
---------------------------------------------------------
---------------------------------------------------------
In this design everything in resctrl interface works like
before (the info, resource group files like task schemata all remain the
same).
This version supports monitoring resctrl groups directly.
But we need a user interface for the user to read the counters. We can
create one file to set monitoring and one
file in resctrl directory which will reflect the counts but may not be
efficient as a lot of times user reads the counts frequently.
Build a new user mode interface resmon
--------------------------------------
Since modifying the existing perf to
suit the different h/w architecture seems to not follow the CAT
interface model, it may well be better to have a different and dedicated
interface for the RDT monitoring (just like we had a new fs for CAT)
resmon supports monitoring a resctrl group or a task. The two modes may
provide enough granularity needed for monitoring
-can monitor cpu data.
-can monitor per resctrl group data.
-can choose custom or subset of tasks with in a resctrl group and monitor.
# resmon [<options>]
-r <resctrl group>
-t <PID>
-s <mon_mask>
-I <time in ms>
"resctrl group": is the resctrl directory.
"mon_mask: is a bit mask of logical packages which indicates which packages user is
interested in monitoring.
"time in ms": The time for which the monitoring takes place
(this can potentially be changed to start and stop/read options)
Example 1 (Some examples modeled from resctrl ui documentation)
---------
A single socket system which has real-time tasks running on core 4-7 and
non real-time workload assigned to core 0-3. The real-time tasks share
text and data, so a per task association is not required and due to
interaction with the kernel it's desired that the kernel on these cores shares L3
with the tasks.
# cd /sys/fs/resctrl
# mkdir p0
# echo "L3:0=3ff" > p0/schemata
core 0-1 are assigned to the new group and make sure that the
kernel and the tasks running there get 50% of the cache.
# echo 03 > p0/cpus
monitor the cpus 0-1 for 10s.
# resmon -r p0 -s 1 -I 10000
Example 2
---------
A real time task running on cpu 2-3(socket 0) is allocated a dedicated 25% of the
cache.
# cd /sys/fs/resctrl
# mkdir p1
# echo "L3:0=0f00;1=ffff" > p1/schemata
# echo 5678 > p1/tasks
# taskset -cp 2-3 5678
Monitor the task for 5s on socket zero
# resmon -r p1 -s 1 -I 5000
Example 3
---------
sometimes user may just want to profile the cache occupancy first before
assigning any CLOSids. Also this provides an override option where user
can monitor some tasks which have say CLOS 0 that he is about to place
in a CLOSId based on the amount of cache occupancy. This could apply to
the same real time tasks above where user is caliberating the % of cache
thats needed.
# resmon -t PIDx,PIDy -s 1 -I 10000
returns the sum of count of PIDx and PIDy
RMID Allocation
---------------
This would remain the same like design version 1, where we support per
package RMIDs and throw error when out of RMIDs due to h/w limited
RMIDs.
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2017-01-20 09:00 +0100 |
| Message-ID | <t1Ch4-4tl-23@gated-at.bofh.it> |
| In reply to | #1563241 |
On Thu, Jan 19, 2017 at 6:32 PM, Vikas Shivappa
<vikas.shivappa@linux.intel.com> wrote:
> Resending including Thomas , also with some changes. Sorry for the spam
>
> Based on Thomas and Peterz feedback Can think of two design
> variants which target:
>
> -Support monitoring and allocating using the same resctrl group.
> user can use a resctrl group to allocate resources and also monitor
> them (with respect to tasks or cpu)
>
> -Also allows monitoring outside of resctrl so that user can
> monitor subgroups who use the same closid. This mode can be used
> when user wants to monitor more than just the resctrl groups.
>
> The first design version uses and modifies perf_cgroup, second version
> builds a new interface resmon.
The second version would require to build a whole new set of tools,
deploy them and maintain them. Users will have to run perf for certain
events and resmon (or whatever is named the new tool) for rdt. I see
it as too complex and much prefer to keep using perf.
> The first version is close to the patches
> sent with some additions/changes. This includes details of the design as
> per Thomas/Peterz feedback.
>
> 1> First Design option: without modifying the resctrl and using perf
> --------------------------------------------------------------------
> --------------------------------------------------------------------
>
> In this design everything in resctrl interface works like
> before (the info, resource group files like task schemata all remain the
> same)
>
>
> Monitor cqm using perf
> ----------------------
>
> perf can monitor individual tasks using the -t
> option just like before.
>
> # perf stat -e llc_occupancy -t PID1,PID2
>
> user can monitor the cpu occupancy using the -C option in perf:
>
> # perf stat -e llc_occupancy -C 5
>
> Below shows how user can monitor cgroup occupancy:
>
> # mount -t cgroup -o perf_event perf_event /sys/fs/cgroup/perf_event/
> # mkdir /sys/fs/cgroup/perf_event/g1
> # mkdir /sys/fs/cgroup/perf_event/g2
> # echo PID1 > /sys/fs/cgroup/perf_event/g2/tasks
>
> # perf stat -e intel_cqm/llc_occupancy/ -a -G g2
>
> To monitor a resctrl group, user can group the same tasks in resctrl
> group into the cgroup.
>
> To monitor the tasks in p1 in example 2 below, add the tasks in resctrl
> group p1 to cgroup g1
>
> # echo 5678 > /sys/fs/cgroup/perf_event/g1/tasks
>
> Introducing a new option for resctrl may complicate monitoring because
> supporting cgroup 'task groups' and resctrl 'task groups' leads to
> situations where:
> if the groups intersect, then there is no way to know what
> l3_allocations contribute to which group.
>
> ex:
> p1 has tasks t1, t2, t3
> g1 has tasks t2, t3, t4
>
> The only way to get occupancy for g1 and p1 would be to allocate an RMID
> for each task which can as well be done with the -t option.
That's simply recreating the resctrl group as a cgroup.
I think that the main advantage of doing allocation first is that we
could use the context switch in rdt allocation and greatly simplify
the pmu side of it.
If resctrl groups could lift the restriction of one resctl per CLOSID,
then the user can create many resctrl in the way perf cgroups are
created now. The advantage is that there wont be cgroup hierarchy!
making things much simpler. Also no need to optimize perf event
context switch to make llc_occupancy work.
Then we only need a way to express that monitoring must happen in a
resctl to the perf_event_open syscall.
My first thought is to have a "rdt_monitor" file per resctl group. A
user passes it to perf_event_open in the way cgroups are passed now.
We could extend the meaning of the flag PERF_FLAG_PID_CGROUP to also
cover rdt_monitor files. The syscall can figure if it's a cgroup or a
rdt_group. The rdt_monitoring PMU would only work with rdt_monitor
groups
Then the rdm_monitoring PMU will be pretty dumb, having neither task
nor CPU contexts. Just providing the pmu->read and pmu->event_init
functions.
Task monitoring can be done with resctrl as well by adding the PID to
a new resctl and opening the event on it. And, since we'd allow CLOSID
to be shared between resctrl groups, allocation wouldn't break.
It's a first idea, so please dont hate too hard ;) .
David
>
> Monitoring cqm cgroups Implementation
> -------------------------------------
>
> When monitoring two different cgroups in the same hierarchy (ex say g11
> has an ancestor g1 which are both being monitored as shown below) we
> need the g11 counts to be considered for g1 as well.
>
> # mount -t cgroup -o perf_event perf_event /sys/fs/cgroup/perf_event/
> # mkdir /sys/fs/cgroup/perf_event/g1
> # mkdir /sys/fs/cgroup/perf_event/g1/g11
>
> When measuring for g1 llc_occupancy we cannot write two different RMIDs
> (because we need to count for g11 as well)
> during context switch to measure the occupancy for both g1 and g11.
> Hence the driver maintains this information and writes the RMID of the
> lowest member in the ancestory which is being monitored during ctx
> switch.
>
> The cqm_info is added to the perf_cgroup structure to maintain this
> information. The structure is allocated and destroyed at css_alloc and
> css_free. All the events tied to a cgroup can use the same
> information while reading the counts.
>
> struct perf_cgroup {
> #ifdef CONFIG_INTEL_RDT_M
> void *cqm_info;
> #endif
> ...
>
> }
>
> struct cqm_info {
> bool mon_enabled;
> int level;
> u32 *rmid;
> struct cgrp_cqm_info *mfa;
> struct list_head tskmon_rlist;
> };
>
> Due to the hierarchical nature of cgroups, every cgroup just
> monitors for the 'nearest monitored ancestor' at all times.
> Since root cgroup is always monitored, all descendents
> at boot time monitor for root and hence all mfa points to root
> except for root->mfa which is NULL.
>
> 1. RMID setup: When cgroup x start monitoring:
> for each descendent y, if y's mfa->level < x->level, then
> y->mfa = x. (Where level of root node = 0...)
> 2. sched_in: During sched_in for x
> if (x->mon_enabled) choose x->rmid
> else choose x->mfa->rmid.
> 3. read: for each descendent of cgroup x
> if (x->monitored) count += rmid_read(x->rmid).
> 4. evt_destroy: for each descendent y of x, if (y->mfa == x)
> then y->mfa = x->mfa. Meaning if any descendent was monitoring for
> x, set that descendent to monitor for the cgroup which x was
> monitoring for.
>
> To monitor a task in a cgroup x along with monitoring cgroup x itself
> cqm_info maintains a list of tasks that are being monitored in the
> cgroup.
>
> When a task which belongs to a cgroup x is being monitored, it
> always uses its own task->rmid even if cgroup x is monitored during sched_in.
> To account for the counts of such tasks, cgroup keeps this list
> and parses it during read.
> taskmon_rlist is used to maintain the list. The list is modified when a
> task is attached to the cgroup or removed from the group.
>
> Example 1 (Some examples modeled from resctrl ui documentation)
> ---------
>
> A single socket system which has real-time tasks running on core 4-7 and
> non real-time workload assigned to core 0-3. The real-time tasks share
> text and data, so a per task association is not required and due to
> interaction with the kernel it's desired that the kernel on these cores shares L3
> with the tasks.
>
> # cd /sys/fs/resctrl
>
> # echo "L3:0=3ff" > schemata
>
> core 0-1 are assigned to the new group and make sure that the
> kernel and the tasks running there get 50% of the cache.
>
> # echo 03 > p0/cpus
>
> monitor the cpus 0-1
>
> # perf stat -e llc_occupancy -C 0-1
>
> Example 2
> ---------
>
> A real time task running on cpu 2-3(socket 0) is allocated a dedicated 25% of the
> cache.
>
> # cd /sys/fs/resctrl
>
> # mkdir p1
> # echo "L3:0=0f00;1=ffff" > p1/schemata
> # echo 5678 > p1/tasks
> # taskset -cp 2-3 5678
>
> To monitor the same group of tasks create a cgroup g1
>
> # mount -t cgroup -o perf_event perf_event /sys/fs/cgroup/perf_event/
> # mkdir /sys/fs/cgroup/perf_event/g1
> # perf stat -e llc_occupancy -a -G g1
>
> Example 3
> ---------
>
> sometimes user may just want to profile the cache occupancy first before
> assigning any CLOSids. Also this provides an override option where user
> can monitor some tasks which have say CLOS 0 that he is about to place
> in a CLOSId based on the amount of cache occupancy. This could apply to
> the same real time tasks above where user is caliberating the % of cache
> thats needed.
>
> # perf stat -e llc_occupancy -t PIDx,PIDy
>
> RMID allocation
> ---------------
>
> RMIDs are allocated per package to achieve better scaling of RMIDs.
> RMIDs are plenty (2-4 per logical processor) and also are per package
> meaning a two socket system would have twice the number of RMIDs.
> If we still run out of RMIDs an error is thrown that monitoring wasnt
> possible as the RMID wasnt available.
>
> Kernel Scheduling
> -----------------
>
> During ctx switch cqm choses the RMID in the following priority
>
> 1. if cpu has a RMID , choose that
> 2. if the task has a RMID directly tied to it choose that (task is
> monitored)
> 3. choose the RMID of the task's cgroup (by default tasks belong to root
> cgroup with RMID 0)
>
> Read
> ----
>
> When user calls cqm to retrieve the monitored count, we read the
> counter_msr and return the count. For cgroup hierarcy , the count is
> measured as explained in the cgroup implementation section by traversing
> the cgroup hierarchy.
>
>
> 2> Second Design option: Build a new usermode tool resmon
> ---------------------------------------------------------
> ---------------------------------------------------------
>
> In this design everything in resctrl interface works like
> before (the info, resource group files like task schemata all remain the
> same).
>
> This version supports monitoring resctrl groups directly.
> But we need a user interface for the user to read the counters. We can
> create one file to set monitoring and one
> file in resctrl directory which will reflect the counts but may not be
> efficient as a lot of times user reads the counts frequently.
>
> Build a new user mode interface resmon
> --------------------------------------
>
> Since modifying the existing perf to
> suit the different h/w architecture seems to not follow the CAT
> interface model, it may well be better to have a different and dedicated
> interface for the RDT monitoring (just like we had a new fs for CAT)
>
> resmon supports monitoring a resctrl group or a task. The two modes may
> provide enough granularity needed for monitoring
> -can monitor cpu data.
> -can monitor per resctrl group data.
> -can choose custom or subset of tasks with in a resctrl group and monitor.
>
> # resmon [<options>]
> -r <resctrl group>
> -t <PID>
> -s <mon_mask>
> -I <time in ms>
>
> "resctrl group": is the resctrl directory.
>
> "mon_mask: is a bit mask of logical packages which indicates which packages user is
> interested in monitoring.
>
> "time in ms": The time for which the monitoring takes place
> (this can potentially be changed to start and stop/read options)
>
> Example 1 (Some examples modeled from resctrl ui documentation)
> ---------
>
> A single socket system which has real-time tasks running on core 4-7 and
> non real-time workload assigned to core 0-3. The real-time tasks share
> text and data, so a per task association is not required and due to
> interaction with the kernel it's desired that the kernel on these cores shares L3
> with the tasks.
>
> # cd /sys/fs/resctrl
> # mkdir p0
> # echo "L3:0=3ff" > p0/schemata
>
> core 0-1 are assigned to the new group and make sure that the
> kernel and the tasks running there get 50% of the cache.
>
> # echo 03 > p0/cpus
>
> monitor the cpus 0-1 for 10s.
>
> # resmon -r p0 -s 1 -I 10000
>
> Example 2
> ---------
>
> A real time task running on cpu 2-3(socket 0) is allocated a dedicated 25% of the
> cache.
>
> # cd /sys/fs/resctrl
>
> # mkdir p1
> # echo "L3:0=0f00;1=ffff" > p1/schemata
> # echo 5678 > p1/tasks
> # taskset -cp 2-3 5678
>
> Monitor the task for 5s on socket zero
>
> # resmon -r p1 -s 1 -I 5000
>
> Example 3
> ---------
>
> sometimes user may just want to profile the cache occupancy first before
> assigning any CLOSids. Also this provides an override option where user
> can monitor some tasks which have say CLOS 0 that he is about to place
> in a CLOSId based on the amount of cache occupancy. This could apply to
> the same real time tasks above where user is caliberating the % of cache
> thats needed.
>
> # resmon -t PIDx,PIDy -s 1 -I 10000
>
> returns the sum of count of PIDx and PIDy
>
> RMID Allocation
> ---------------
>
> This would remain the same like design version 1, where we support per
> package RMIDs and throw error when out of RMIDs due to h/w limited
> RMIDs.
>
>
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web