Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1194760 > unrolled thread
| Started by | "Auld, Will" <will.auld@intel.com> |
|---|---|
| First post | 2015-07-29 03:30 +0200 |
| Last post | 2015-07-29 22:10 +0200 |
| Articles | 10 — 4 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
RE: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide "Auld, Will" <will.auld@intel.com> - 2015-07-29 03:30 +0200
Re: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide Marcelo Tosatti <mtosatti@redhat.com> - 2015-07-29 21:40 +0200
Re: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide Marcelo Tosatti <mtosatti@redhat.com> - 2015-07-30 22:10 +0200
Re: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide Marcelo Tosatti <mtosatti@redhat.com> - 2015-07-31 17:40 +0200
Re: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide Martin Kletzander <mkletzan@redhat.com> - 2015-08-02 17:50 +0200
Re: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide Marcelo Tosatti <mtosatti@redhat.com> - 2015-08-03 17:20 +0200
Re: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide Vikas Shivappa <vikas.shivappa@intel.com> - 2015-08-03 20:30 +0200
Re: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide Marcelo Tosatti <mtosatti@redhat.com> - 2015-07-31 17:20 +0200
[summary] Re: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide Vikas Shivappa <vikas.shivappa@intel.com> - 2015-07-31 18:50 +0200
RE: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide Vikas Shivappa <vikas.shivappa@intel.com> - 2015-07-29 22:10 +0200
| From | "Auld, Will" <will.auld@intel.com> |
|---|---|
| Date | 2015-07-29 03:30 +0200 |
| Subject | RE: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide |
| Message-ID | <pRolX-1At-3@gated-at.bofh.it> |
> -----Original Message----- > From: Shivappa, Vikas > Sent: Tuesday, July 28, 2015 5:07 PM > To: Marcelo Tosatti > Cc: Vikas Shivappa; linux-kernel@vger.kernel.org; Shivappa, Vikas; > x86@kernel.org; hpa@zytor.com; tglx@linutronix.de; mingo@kernel.org; > tj@kernel.org; peterz@infradead.org; Fleming, Matt; Auld, Will; Williamson, > Glenn P; Juvva, Kanaka D > Subject: Re: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and > cgroup usage guide > > > > On Tue, 28 Jul 2015, Marcelo Tosatti wrote: > > > On Wed, Jul 01, 2015 at 03:21:04PM -0700, Vikas Shivappa wrote: > >> Adds a description of Cache allocation technology, overview of kernel > >> implementation and usage of Cache Allocation cgroup interface. > >> > >> Cache allocation is a sub-feature of Resource Director > >> Technology(RDT) Allocation or Platform Shared resource control which > >> provides support to control Platform shared resources like L3 cache. > >> Currently L3 Cache is the only resource that is supported in RDT. > >> More information can be found in the Intel SDM, Volume 3, section 17.15. > >> > >> Cache Allocation Technology provides a way for the Software (OS/VMM) > >> to restrict cache allocation to a defined 'subset' of cache which may > >> be overlapping with other 'subsets'. This feature is used when > >> allocating a line in cache ie when pulling new data into the cache. > >> > >> Signed-off-by: Vikas Shivappa <vikas.shivappa@linux.intel.com> > >> --- > >> Documentation/cgroups/rdt.txt | 215 > >> ++++++++++++++++++++++++++++++++++++++++++ > >> 1 file changed, 215 insertions(+) > >> create mode 100644 Documentation/cgroups/rdt.txt > >> > >> diff --git a/Documentation/cgroups/rdt.txt > >> b/Documentation/cgroups/rdt.txt new file mode 100644 index > >> 0000000..dfff477 > >> --- /dev/null > >> +++ b/Documentation/cgroups/rdt.txt > >> @@ -0,0 +1,215 @@ > >> + RDT > >> + --- > >> + > >> +Copyright (C) 2014 Intel Corporation Written by > >> +vikas.shivappa@linux.intel.com (based on contents and format from > >> +cpusets.txt) > >> + > >> +CONTENTS: > >> +========= > >> + > >> +1. Cache Allocation Technology > >> + 1.1 What is RDT and Cache allocation ? > >> + 1.2 Why is Cache allocation needed ? > >> + 1.3 Cache allocation implementation overview > >> + 1.4 Assignment of CBM and CLOS > >> + 1.5 Scheduling and Context Switch > >> +2. Usage Examples and Syntax > >> + > >> +1. Cache Allocation Technology(Cache allocation) > >> +=================================== > >> + > >> +1.1 What is RDT and Cache allocation > >> +------------------------------------ > >> + > >> +Cache allocation is a sub-feature of Resource Director > >> +Technology(RDT) Allocation or Platform Shared resource control which > >> +provides support to control Platform shared resources like L3 cache. > >> +Currently L3 Cache is the only resource that is supported in RDT. > >> +More information can be found in the Intel SDM, Volume 3, section 17.15. > >> + > >> +Cache Allocation Technology provides a way for the Software (OS/VMM) > >> +to restrict cache allocation to a defined 'subset' of cache which > >> +may be overlapping with other 'subsets'. This feature is used when > >> +allocating a line in cache ie when pulling new data into the cache. > >> +The programming of the h/w is done via programming MSRs. > >> + > >> +The different cache subsets are identified by CLOS identifier (class > >> +of service) and each CLOS has a CBM (cache bit mask). The CBM is a > >> +contiguous set of bits which defines the amount of cache resource > >> +that is available for each 'subset'. > >> + > >> +1.2 Why is Cache allocation needed > >> +---------------------------------- > >> + > >> +In todays new processors the number of cores is continuously > >> +increasing, especially in large scale usage models where VMs are > >> +used like webservers and datacenters. The number of cores increase > >> +the number of threads or workloads that can simultaneously be run. > >> +When multi-threaded-applications, VMs, workloads run concurrently > >> +they compete for shared resources including L3 cache. > >> + > >> +The Cache allocation enables more cache resources to be made > >> +available for higher priority applications based on guidance from > >> +the execution environment. > >> + > >> +The architecture also allows dynamically changing these subsets > >> +during runtime to further optimize the performance of the higher > >> +priority application with minimal degradation to the low priority app. > >> +Additionally, resources can be rebalanced for system throughput benefit. > >> + > >> +This technique may be useful in managing large computer systems > >> +which large L3 cache. Examples may be large servers running > >> +instances of webservers or database servers. In such complex > >> +systems, these subsets can be used for more careful placing of the > >> +available cache resources. > >> + > >> +1.3 Cache allocation implementation Overview > >> +-------------------------------------------- > >> + > >> +Kernel implements a cgroup subsystem to support cache allocation. > >> + > >> +Each cgroup has a CLOSid <-> CBM(cache bit mask) mapping. > >> +A CLOS(Class of service) is represented by a CLOSid.CLOSid is > >> +internal to the kernel and not exposed to user. Each cgroup would > >> +have one CBM and would just represent one cache 'subset'. > >> + > >> +The cgroup follows cgroup hierarchy ,mkdir and adding tasks to the > >> +cgroup never fails. When a child cgroup is created it inherits the > >> +CLOSid and the CBM from its parent. When a user changes the default > >> +CBM for a cgroup, a new CLOSid may be allocated if the CBM was not > >> +used before. The changing of 'l3_cache_mask' may fail with -ENOSPC > >> +once the kernel runs out of maximum CLOSids it can support. > >> +User can create as many cgroups as he wants but having different > >> +CBMs at the same time is restricted by the maximum number of CLOSids > >> +(multiple cgroups can have the same CBM). > >> +Kernel maintains a CLOSid<->cbm mapping which keeps reference > >> +counter for each cgroup using a CLOSid. > >> + > >> +The tasks in the cgroup would get to fill the L3 cache represented > >> +by the cgroup's 'l3_cache_mask' file. > >> + > >> +Root directory would have all available bits set in 'l3_cache_mask' > >> +file by default. > >> + > >> +Each RDT cgroup directory has the following files. Some of them may > >> +be a part of common RDT framework or be specific to RDT sub-features > >> +like cache allocation. > >> + > >> + - intel_rdt.l3_cache_mask: The cache bitmask(CBM) is represented by > >> + this file. The bitmask must be contiguous and would have a 1 or 2 > >> + bit minimum length. > >> + > >> +1.4 Assignment of CBM,CLOS > >> +-------------------------- > >> + > >> +The 'l3_cache_mask' needs to be a subset of the parent node's > >> +'l3_cache_mask'. Any contiguous subset of these bits(with a minimum > >> +of 2 bits on hsw SKUs) maybe set to indicate the cache mapping > >> +desired. The 'l3_cache_mask' between 2 directories can overlap. The > >> +'l3_cache_mask' would represent the cache 'subset' of the Cache > >> +allocation cgroup. For ex: on a system with 16 bits of max cbm bits, > >> +if the directory has the least significant 4 bits set in its 'l3_cache_mask' > file(meaning the 'l3_cache_mask' > >> +is just 0xf), it would be allocated the right quarter of the Last > >> +level cache which means the tasks belonging to this Cache allocation > >> +cgroup can use the right quarter of the cache to fill. If it has the > >> +most significant 8 bits set ,it would be allocated the left half of > >> +the cache(8 bits out of 16 represents 50%). > >> + > >> +The cache portion defined in the CBM file is available to all tasks > >> +within the cgroup to fill and these task are not allowed to allocate > >> +space in other parts of the cache. > >> + > >> +1.5 Scheduling and Context Switch > >> +--------------------------------- > >> + > >> +During context switch kernel implements this by writing the CLOSid > >> +(internally maintained by kernel) of the cgroup to which the task > >> +belongs to the CPU's IA32_PQR_ASSOC MSR. The MSR is only written > >> +when there is a change in the CLOSid for the CPU in order to > >> +minimize the latency incurred during context switch. > >> + > >> +The following considerations are done for the PQR MSR write so that > >> +it has minimal impact on scheduling hot path: > >> +- This path doesnt exist on any non-intel platforms. > >> +- On Intel platforms, this would not exist by default unless > >> +CGROUP_RDT is enabled. > >> +- remains a no-op when CGROUP_RDT is enabled and intel hardware does > >> +not support the feature. > >> +- When feature is available, still remains a no-op till the user > >> +manually creates a cgroup *and* assigns a new cache mask. Since the > >> +child node inherits the parents cache mask , by cgroup creation > >> +there is no scheduling hot path impact from the new cgroup. > >> +- per cpu PQR values are cached and the MSR write is only done when > >> +there is a task with different PQR is scheduled on the CPU. > >> +Typically if the task groups are bound to be scheduled on a set of > >> +CPUs , the number of MSR writes is greatly reduced. > >> + > >> +2. Usage examples and syntax > >> +============================ > >> + > >> +To check if Cache allocation was enabled on your system > >> + > >> +dmesg | grep -i intel_rdt > >> +should output : intel_rdt: Max bitmask length: xx,Max ClosIds: xx > >> +the length of l3_cache_mask and CLOS should depend on the system you > use. > >> + > >> +Also /proc/cpuinfo would have rdt(if rdt is enabled) and cat_l3( if L3 > >> + cache allocation is enabled). > >> + > >> +Following would mount the cache allocation cgroup subsystem and > >> +create > >> +2 directories. Please refer to Documentation/cgroups/cgroups.txt on > >> +details about how to use cgroups. > >> + > >> + cd /sys/fs/cgroup > >> + mkdir rdt > >> + mount -t cgroup -ointel_rdt intel_rdt /sys/fs/cgroup/rdt cd rdt > >> + > >> +Create 2 rdt cgroups > >> + > >> + mkdir group1 > >> + mkdir group2 > >> + > >> +Following are some of the Files in the directory > >> + > >> + ls > >> + rdt.l3_cache_mask > >> + tasks > >> + > >> +Say if the cache is 2MB and cbm supports 16 bits, then setting the > >> +below allocates the 'right 1/4th(512KB)' of the cache to group2 > >> + > >> +Edit the CBM for group2 to set the least significant 4 bits. This > >> +allocates 'right quarter' of the cache. > >> + > >> + cd group2 > >> + /bin/echo 0xf > rdt.l3_cache_mask > >> + > >> + > >> +Edit the CBM for group2 to set the least significant 8 bits.This > >> +allocates the right half of the cache to 'group2'. > >> + > >> + cd group2 > >> + /bin/echo 0xff > rdt.l3_cache_mask > >> + > >> +Assign tasks to the group2 > >> + > >> + /bin/echo PID1 > tasks > >> + /bin/echo PID2 > tasks > >> + > >> + Meaning now threads > >> + PID1 and PID2 get to fill the 'right half' of the cache as the > >> + belong to cgroup group2. > >> + > >> +Create a group under group2 > >> + > >> + cd group2 > >> + mkdir group21 > >> + cat rdt.l3_cache_mask > >> + 0xff - inherits parents mask. > >> + > >> + /bin/echo 0xfff > rdt.l3_cache_mask - throws error as mask has to > >> + parent's mask's subset > >> + > >> +In order to restrict RDT cgroups to specific set of CPUs rdt can be > >> +comounted with cpusets. > >> -- > >> 1.9.1 > > > > Vikas, > > > > Can you give an example of comounting with cpusets? What do you mean > > by restrict RDT cgroups to specific set of CPUs? > > I was going to edit the documentation soon as i see a lot of feedback on the > same. It may have caused confusion. > > I mean just pinning down tasks to a set of cpus. This does not mean we make the > cache exclusive to the tasks.. > > > > > Another limitation of this interface is that it assumes the task <-> > > control group assignment is pertinent, that is: > > > > | taskgroup, L3 policy|: > > > > | taskgroupA, 50% L3 exclusive |, > > | taskgroupB, 50% L3 |, > > | taskgroupC, 50% L3 |. > > > > Whenever taskgroup A is empty (that is no runnable task in it), you > > waste 50% of > > L3 cache. > > Cgroup masks can always overlap , and hence wont have exclusive cache > allocation. > > > > > I think this problem and the similar problem of L3 reservation with > > CPU isolation can be solved in this way: whenever a task from cgroupE > > with exclusive way access is migrated to a new die, impose the > > exclusivity (by removing access to that way by other cgroups). > > > > Whenever cgroupE has zero tasks, remove exclusivity (by allowing other > > cgroups to use the exclusive ways of it). > > Same comment as above - Cgroup masks can always overlap and other cgroups > can allocate the same cache , and hence wont have exclusive cache allocation. [Auld, Will] You can define all the cbm to provide one clos with an exclusive area > > So natuarally the cgroup with tasks would get to use the cache if it has the same > mask (say representing 50% of cache in your example) as others . [Auld, Will] automatic adjustment of the cbm make me nervous. There are times when we want to limit the cache for a process independent of whether there is lots of unused cache. > (assume there are 8 bits max cbm) > cgroupa - mask - 0xf > cgroupb - mask - 0xf . Now if cgroupa has no tasks , cgroupb naturally gets all > the cache. > > Thanks, > Vikas > > > > > I'll cook a patch. > > > > > > > > > > -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Marcelo Tosatti <mtosatti@redhat.com> |
|---|---|
| Date | 2015-07-29 21:40 +0200 |
| Message-ID | <pRFmO-LX-27@gated-at.bofh.it> |
| In reply to | #1194760 |
On Wed, Jul 29, 2015 at 01:28:38AM +0000, Auld, Will wrote: > > > Whenever cgroupE has zero tasks, remove exclusivity (by allowing other > > > cgroups to use the exclusive ways of it). > > > > Same comment as above - Cgroup masks can always overlap and other cgroups > > can allocate the same cache , and hence wont have exclusive cache allocation. > > [Auld, Will] You can define all the cbm to provide one clos with an exclusive area > > > > > So natuarally the cgroup with tasks would get to use the cache if it has the same > > mask (say representing 50% of cache in your example) as others . > > [Auld, Will] automatic adjustment of the cbm make me nervous. There are times > when we want to limit the cache for a process independent of whether there is > lots of unused cache. How about this: desiredclos (closid p1 p2 p3 p4) 1 1 0 0 0 2 0 0 0 1 3 0 1 1 0 p means part. closid 1 is a exclusive cgroup. closid 2 is a "cache hog" class. closid 3 is "default closid". Desiredclos is what user has specified. Transition 1: desiredclos --> effectiveclos Clean all bits of unused closid's (that must be updated whenever a closid1 cgroup goes from empty->nonempty and vice-versa). effectiveclos (closid p1 p2 p3 p4) 1 0 0 0 0 2 0 0 0 1 3 0 1 1 0 Transition 2: effectiveclos --> expandedclos expandedclos (closid p1 p2 p3 p4) 1 0 0 0 0 2 0 0 0 1 3 1 1 1 0 Then you have different inplacecos for each CPU (see pseudo-code below): On the following events. - task migration to new pCPU: - task creation: id = smp_processor_id(); for (part = desiredclos.p1; ...; part++) /* if my cosid is set and any other cosid is clear, for the part, synchronize desiredclos --> inplacecos */ if (part[mycosid] == 1 && part[any_othercosid] == 0) wrmsr(part, desiredclos); -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Marcelo Tosatti <mtosatti@redhat.com> |
|---|---|
| Date | 2015-07-30 22:10 +0200 |
| Message-ID | <pS2jn-dO-11@gated-at.bofh.it> |
| In reply to | #1195462 |
On Thu, Jul 30, 2015 at 10:47:23AM -0700, Vikas Shivappa wrote:
>
>
> Marcello,
>
>
> On Wed, 29 Jul 2015, Marcelo Tosatti wrote:
> >
> >How about this:
> >
> >desiredclos (closid p1 p2 p3 p4)
> > 1 1 0 0 0
> > 2 0 0 0 1
> > 3 0 1 1 0
>
> #1 Currently in the rdt cgroup , the root cgroup always has all the
> bits set and cant be changed (because the cgroup hierarchy would by
> default make this to have all bits as all the children need to have
> a subset of the root's bitmask). So if the user creates a cgroup and
> not put any task in it , the tasks in the root cgroup could be still
> using that part of the cache. Thats the reason i say we can have
> really 'exclusive' masks.
>
> Or in other words - there is always a desired clos (0) which has all
> parts set which acts like a default pool.
>
> Also the parts can overlap. Please apply this for all the below
> comments which will change the way they work.
>
> >
> >p means part.
>
> I am assuming p = (a contiguous cache capacity bit mask)
Yes.
> >closid 1 is a exclusive cgroup.
> >closid 2 is a "cache hog" class.
> >closid 3 is "default closid".
> >
> >Desiredclos is what user has specified.
> >
> >Transition 1: desiredclos --> effectiveclos
> >Clean all bits of unused closid's
> >(that must be updated whenever a
> >closid1 cgroup goes from empty->nonempty
> >and vice-versa).
> >
> >effectiveclos (closid p1 p2 p3 p4)
> > 1 0 0 0 0
> > 2 0 0 0 1
> > 3 0 1 1 0
>
> >
> >Transition 2: effectiveclos --> expandedclos
> >expandedclos (closid p1 p2 p3 p4)
> > 1 0 0 0 0
> > 2 0 0 0 1
> > 3 1 1 1 0
> >Then you have different inplacecos for each
> >CPU (see pseudo-code below):
> >
> >On the following events.
> >
> >- task migration to new pCPU:
> >- task creation:
> >
> > id = smp_processor_id();
> > for (part = desiredclos.p1; ...; part++)
> > /* if my cosid is set and any other
> > cosid is clear, for the part,
> > synchronize desiredclos --> inplacecos */
> > if (part[mycosid] == 1 &&
> > part[any_othercosid] == 0)
> > wrmsr(part, desiredclos);
> >
>
> Currently the root cgroup would have all the bits set which will act
> like a default cgroup where all the otherwise unused parts (assuming
> they are a set of contiguous cache capacity bits) will be used.
>
> Otherwise the question is in the expandedclos - who decides to
> expand the closx parts to include some of the unused parts.. - that
> could just be a default root always ?
Right, so the problem is for certain closid's you might never want
to expand (because doing so would cause data to be cached in a
cache way which might have high eviction rate in the future).
See the example from Will.
But for the default cache (that is "unclassified applications"
i suppose it is beneficial to expand in most cases, that is,
use maximum amount of cache irrespective of eviction rate, which
is the behaviour that exists now without CAT).
So perhaps a new flag "expand=y/n" can be added to the cgroup
directories... What do you say?
Userspace representation of CAT
-------------------------------
Usage model:
1) measure application performance without L3 cache reservation.
2) measure application perf with L3 cache reservation and
X number of cache ways until desired performance is attained.
Requirements:
1) Persistency of CLOS configuration across hardware. On migration
of operating system or application between different hardware
systems we'd like the following to be maintained:
- exclusive number of bytes (*) reserved to a certain CLOSid.
- shared number of bytes (*) reserved between a certain group
of CLOSid's.
For both code and data, rounded down or up in cache way size.
2) Reasoning:
Different CBM masks in different hardware platforms might be necessary
to specify the same CLOS configuration, in terms of exclusive number of
bytes and shared number of bytes. (cache-way rounded number of bytes).
For example, due to L3 allocation by other hardware entities in certain parts
of the cache it might be necessary to relocate CBM mask to achieve
the same CLOS configuration.
3) Proposed format:
sharedregionK.exclusive - Number of exclusive cache bytes reserved for
shared region.
sharedregionK.excl_data - Number of exclusive cache data bytes reserved for
shared region.
sharedregionK.excl_bytes - Number of exclusive cache code bytes reserved for
shared region.
sharedregionK.round_down - Round down to cache way bytes from respective number
specification (default is round up).
sharedregionK.expand - y/n - Expand shared region to more cache ways
when available (default N).
cgroupN.exclusive - Number of exclusive L3 cache bytes reserved
for cgroup.
cgroupN.excl_data - Number of exclusive L3 data cache bytes reserved
for cgroup.
cgroupN.excl_code - Number of exclusive L3 code cache bytes reserved
for cgroup.
cgroupN.round_down - Round down to cache way bytes from respective number
specification (default is round up).
cgroupN.expand - y/n - Expand shared region to more cache ways when
available (default N).
cgroupN.shared = { sharedregion1, sharedregion2, ... } (list of shared
regions)
Example 1:
One application with 2M exclusive cache, two applications
with 1M exclusive each, sharing an expansive shared region of 1M.
cgroup1.exclusive = 2M
sharedregion1.exclusive = 1M
sharedregion1.expand = Y
cgroup2.exclusive = 1M
cgroup2.shared = sharedregion1
cgroup3.exclusive = 1M
cgroup3.shared = sharedregion1
Example 2:
3 high performance applications running, one of which is a cache hog
with no cache locality.
cgroup1.exclusive = 8M
cgroup2.exclusive = 8M
cgroup3.exclusive = 512K
cgroup3.round_down = Y
In all cases the default cgroup (which requires no explicit
specification) is expansive and uses the remaining cache
ways, including the ways shared by other hardware entities.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Marcelo Tosatti <mtosatti@redhat.com> |
|---|---|
| Date | 2015-07-31 17:40 +0200 |
| Message-ID | <pSkzF-1wQ-15@gated-at.bofh.it> |
| In reply to | #1196407 |
On Thu, Jul 30, 2015 at 05:08:13PM -0300, Marcelo Tosatti wrote:
> On Thu, Jul 30, 2015 at 10:47:23AM -0700, Vikas Shivappa wrote:
> >
> >
> > Marcello,
> >
> >
> > On Wed, 29 Jul 2015, Marcelo Tosatti wrote:
> > >
> > >How about this:
> > >
> > >desiredclos (closid p1 p2 p3 p4)
> > > 1 1 0 0 0
> > > 2 0 0 0 1
> > > 3 0 1 1 0
> >
> > #1 Currently in the rdt cgroup , the root cgroup always has all the
> > bits set and cant be changed (because the cgroup hierarchy would by
> > default make this to have all bits as all the children need to have
> > a subset of the root's bitmask). So if the user creates a cgroup and
> > not put any task in it , the tasks in the root cgroup could be still
> > using that part of the cache. Thats the reason i say we can have
> > really 'exclusive' masks.
> >
> > Or in other words - there is always a desired clos (0) which has all
> > parts set which acts like a default pool.
> >
> > Also the parts can overlap. Please apply this for all the below
> > comments which will change the way they work.
> >
> > >
> > >p means part.
> >
> > I am assuming p = (a contiguous cache capacity bit mask)
>
> Yes.
>
> > >closid 1 is a exclusive cgroup.
> > >closid 2 is a "cache hog" class.
> > >closid 3 is "default closid".
> > >
> > >Desiredclos is what user has specified.
> > >
> > >Transition 1: desiredclos --> effectiveclos
> > >Clean all bits of unused closid's
> > >(that must be updated whenever a
> > >closid1 cgroup goes from empty->nonempty
> > >and vice-versa).
> > >
> > >effectiveclos (closid p1 p2 p3 p4)
> > > 1 0 0 0 0
> > > 2 0 0 0 1
> > > 3 0 1 1 0
> >
> > >
> > >Transition 2: effectiveclos --> expandedclos
> > >expandedclos (closid p1 p2 p3 p4)
> > > 1 0 0 0 0
> > > 2 0 0 0 1
> > > 3 1 1 1 0
> > >Then you have different inplacecos for each
> > >CPU (see pseudo-code below):
> > >
> > >On the following events.
> > >
> > >- task migration to new pCPU:
> > >- task creation:
> > >
> > > id = smp_processor_id();
> > > for (part = desiredclos.p1; ...; part++)
> > > /* if my cosid is set and any other
> > > cosid is clear, for the part,
> > > synchronize desiredclos --> inplacecos */
> > > if (part[mycosid] == 1 &&
> > > part[any_othercosid] == 0)
> > > wrmsr(part, desiredclos);
> > >
> >
> > Currently the root cgroup would have all the bits set which will act
> > like a default cgroup where all the otherwise unused parts (assuming
> > they are a set of contiguous cache capacity bits) will be used.
> >
> > Otherwise the question is in the expandedclos - who decides to
> > expand the closx parts to include some of the unused parts.. - that
> > could just be a default root always ?
>
> Right, so the problem is for certain closid's you might never want
> to expand (because doing so would cause data to be cached in a
> cache way which might have high eviction rate in the future).
> See the example from Will.
>
> But for the default cache (that is "unclassified applications"
> i suppose it is beneficial to expand in most cases, that is,
> use maximum amount of cache irrespective of eviction rate, which
> is the behaviour that exists now without CAT).
>
> So perhaps a new flag "expand=y/n" can be added to the cgroup
> directories... What do you say?
>
> Userspace representation of CAT
> -------------------------------
>
> Usage model:
> 1) measure application performance without L3 cache reservation.
> 2) measure application perf with L3 cache reservation and
> X number of cache ways until desired performance is attained.
>
> Requirements:
> 1) Persistency of CLOS configuration across hardware. On migration
> of operating system or application between different hardware
> systems we'd like the following to be maintained:
> - exclusive number of bytes (*) reserved to a certain CLOSid.
> - shared number of bytes (*) reserved between a certain group
> of CLOSid's.
>
> For both code and data, rounded down or up in cache way size.
>
> 2) Reasoning:
> Different CBM masks in different hardware platforms might be necessary
> to specify the same CLOS configuration, in terms of exclusive number of
> bytes and shared number of bytes. (cache-way rounded number of bytes).
> For example, due to L3 allocation by other hardware entities in certain parts
> of the cache it might be necessary to relocate CBM mask to achieve
> the same CLOS configuration.
>
> 3) Proposed format:
>
> sharedregionK.exclusive - Number of exclusive cache bytes reserved for
> shared region.
> sharedregionK.excl_data - Number of exclusive cache data bytes reserved for
> shared region.
> sharedregionK.excl_bytes - Number of exclusive cache code bytes reserved for
> shared region.
> sharedregionK.round_down - Round down to cache way bytes from respective number
> specification (default is round up).
> sharedregionK.expand - y/n - Expand shared region to more cache ways
> when available (default N).
>
> cgroupN.exclusive - Number of exclusive L3 cache bytes reserved
> for cgroup.
> cgroupN.excl_data - Number of exclusive L3 data cache bytes reserved
> for cgroup.
> cgroupN.excl_code - Number of exclusive L3 code cache bytes reserved
> for cgroup.
> cgroupN.round_down - Round down to cache way bytes from respective number
> specification (default is round up).
> cgroupN.expand - y/n - Expand shared region to more cache ways when
> available (default N).
> cgroupN.shared = { sharedregion1, sharedregion2, ... } (list of shared
> regions)
>
> Example 1:
> One application with 2M exclusive cache, two applications
> with 1M exclusive each, sharing an expansive shared region of 1M.
>
> cgroup1.exclusive = 2M
>
> sharedregion1.exclusive = 1M
> sharedregion1.expand = Y
>
> cgroup2.exclusive = 1M
> cgroup2.shared = sharedregion1
>
> cgroup3.exclusive = 1M
> cgroup3.shared = sharedregion1
>
> Example 2:
> 3 high performance applications running, one of which is a cache hog
> with no cache locality.
>
> cgroup1.exclusive = 8M
> cgroup2.exclusive = 8M
>
> cgroup3.exclusive = 512K
> cgroup3.round_down = Y
>
> In all cases the default cgroup (which requires no explicit
> specification) is expansive and uses the remaining cache
> ways, including the ways shared by other hardware entities.
>
Moving this discussion from another to this thread, sorry.
> >Second question:
> >Do you envision any use case which the placement of cache
> >and not the quantity of cache is a criteria for decision?
> >That is, two cases with the same amount of cache for each CLOSid,
> >but with different locations inside the cache?
> >(except sharing of ways by two CLOSid's, of course).
> >
>
> cbm max - 16 bits. 000f - allocate right quarter. f000 - allocate
> left quarter.. ? extend the case to any number of valid contiguous
> bits.
Yes, the hardware allows you to specify the same number of cache ways
to a given COSid in different cache locations. The question was whether
do you envision any use case where different locations make a
difference?
I can't see any (except for hardware users of cache ways, which
the OS could control automatically, all it needs to know from
the user configuration is whether a given cgroup is a "exclusive user"
of a number of cache ways, in which case it should not use a
This information is crucial because if there are no forseeable use
cases then that can simplify the interface enormously (could have the
kernel handle the issues that my "userspace interface" proposal is
handling).
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Martin Kletzander <mkletzan@redhat.com> |
|---|---|
| Date | 2015-08-02 17:50 +0200 |
| Message-ID | <pT3Gp-8ld-15@gated-at.bofh.it> |
| In reply to | #1196407 |
[Multipart message — attachments visible in raw view] — view raw
On Thu, Jul 30, 2015 at 05:08:13PM -0300, Marcelo Tosatti wrote:
>On Thu, Jul 30, 2015 at 10:47:23AM -0700, Vikas Shivappa wrote:
>>
>>
>> Marcello,
>>
>>
>> On Wed, 29 Jul 2015, Marcelo Tosatti wrote:
>> >
>> >How about this:
>> >
>> >desiredclos (closid p1 p2 p3 p4)
>> > 1 1 0 0 0
>> > 2 0 0 0 1
>> > 3 0 1 1 0
>>
>> #1 Currently in the rdt cgroup , the root cgroup always has all the
>> bits set and cant be changed (because the cgroup hierarchy would by
>> default make this to have all bits as all the children need to have
>> a subset of the root's bitmask). So if the user creates a cgroup and
>> not put any task in it , the tasks in the root cgroup could be still
>> using that part of the cache. Thats the reason i say we can have
>> really 'exclusive' masks.
>>
>> Or in other words - there is always a desired clos (0) which has all
>> parts set which acts like a default pool.
>>
>> Also the parts can overlap. Please apply this for all the below
>> comments which will change the way they work.
>>
>> >
>> >p means part.
>>
>> I am assuming p = (a contiguous cache capacity bit mask)
>
>Yes.
>
>> >closid 1 is a exclusive cgroup.
>> >closid 2 is a "cache hog" class.
>> >closid 3 is "default closid".
>> >
>> >Desiredclos is what user has specified.
>> >
>> >Transition 1: desiredclos --> effectiveclos
>> >Clean all bits of unused closid's
>> >(that must be updated whenever a
>> >closid1 cgroup goes from empty->nonempty
>> >and vice-versa).
>> >
>> >effectiveclos (closid p1 p2 p3 p4)
>> > 1 0 0 0 0
>> > 2 0 0 0 1
>> > 3 0 1 1 0
>>
>> >
>> >Transition 2: effectiveclos --> expandedclos
>> >expandedclos (closid p1 p2 p3 p4)
>> > 1 0 0 0 0
>> > 2 0 0 0 1
>> > 3 1 1 1 0
>> >Then you have different inplacecos for each
>> >CPU (see pseudo-code below):
>> >
>> >On the following events.
>> >
>> >- task migration to new pCPU:
>> >- task creation:
>> >
>> > id = smp_processor_id();
>> > for (part = desiredclos.p1; ...; part++)
>> > /* if my cosid is set and any other
>> > cosid is clear, for the part,
>> > synchronize desiredclos --> inplacecos */
>> > if (part[mycosid] == 1 &&
>> > part[any_othercosid] == 0)
>> > wrmsr(part, desiredclos);
>> >
>>
>> Currently the root cgroup would have all the bits set which will act
>> like a default cgroup where all the otherwise unused parts (assuming
>> they are a set of contiguous cache capacity bits) will be used.
>>
>> Otherwise the question is in the expandedclos - who decides to
>> expand the closx parts to include some of the unused parts.. - that
>> could just be a default root always ?
>
>Right, so the problem is for certain closid's you might never want
>to expand (because doing so would cause data to be cached in a
>cache way which might have high eviction rate in the future).
>See the example from Will.
>
>But for the default cache (that is "unclassified applications"
>i suppose it is beneficial to expand in most cases, that is,
>use maximum amount of cache irrespective of eviction rate, which
>is the behaviour that exists now without CAT).
>
>So perhaps a new flag "expand=y/n" can be added to the cgroup
>directories... What do you say?
>
>Userspace representation of CAT
>-------------------------------
>
>Usage model:
>1) measure application performance without L3 cache reservation.
>2) measure application perf with L3 cache reservation and
>X number of cache ways until desired performance is attained.
>
>Requirements:
>1) Persistency of CLOS configuration across hardware. On migration
>of operating system or application between different hardware
>systems we'd like the following to be maintained:
> - exclusive number of bytes (*) reserved to a certain CLOSid.
> - shared number of bytes (*) reserved between a certain group
> of CLOSid's.
>
>For both code and data, rounded down or up in cache way size.
>
>2) Reasoning:
>Different CBM masks in different hardware platforms might be necessary
>to specify the same CLOS configuration, in terms of exclusive number of
>bytes and shared number of bytes. (cache-way rounded number of bytes).
>For example, due to L3 allocation by other hardware entities in certain parts
>of the cache it might be necessary to relocate CBM mask to achieve
>the same CLOS configuration.
>
>3) Proposed format:
>
Few questions from a random listener, I apologise if some of them are
in a wrong place due to me missing some information from past threads.
I'm not sure whether the following proposal to the format is the
internal structure or what's going to be in cgroups. If this is
user-visible interface, I think it could be a little less detailed.
>sharedregionK.exclusive - Number of exclusive cache bytes reserved for
> shared region.
>sharedregionK.excl_data - Number of exclusive cache data bytes reserved for
> shared region.
>sharedregionK.excl_bytes - Number of exclusive cache code bytes reserved for
> shared region.
>sharedregionK.round_down - Round down to cache way bytes from respective number
> specification (default is round up).
>sharedregionK.expand - y/n - Expand shared region to more cache ways
> when available (default N).
>
>cgroupN.exclusive - Number of exclusive L3 cache bytes reserved
> for cgroup.
>cgroupN.excl_data - Number of exclusive L3 data cache bytes reserved
> for cgroup.
>cgroupN.excl_code - Number of exclusive L3 code cache bytes reserved
> for cgroup.
By exclusive, you mean that it's exclusive to the tasks in this
cgroup?
The thing is that we must differentiate between limiting some
process's from hogging the memory (like example 2 below) and making
some part of the cache exclusive for particular application (example 1
below).
I just hope we won't need to add something similar to 'isolcpus=' just
so we can make sure none of the tasks in the root cgroup can spoil the
part of the cache we need to have exclusive.
I'm not sure creating a new subgroup and moving all the tasks there
would work, It certainly is not possible with other cgroups, like the
cpuset cgroup mentioned beforehand.
I also don't quite fully understand how the co-mounting with the
cpuset cgroup should work, but that's not design-related.
One more question, how does this work on systems with multiple L3
caches (e.g. large NUMA node systems)? I'm guessing if the process is
running only on some CPUs, the wrmsr() will be called on that
particular CPU(s), right?
>cgroupN.round_down - Round down to cache way bytes from respective number
> specification (default is round up).
>cgroupN.expand - y/n - Expand shared region to more cache ways when
> available (default N).
>cgroupN.shared = { sharedregion1, sharedregion2, ... } (list of shared
>regions)
>
>Example 1:
>One application with 2M exclusive cache, two applications
>with 1M exclusive each, sharing an expansive shared region of 1M.
>
>cgroup1.exclusive = 2M
>
>sharedregion1.exclusive = 1M
>sharedregion1.expand = Y
>
>cgroup2.exclusive = 1M
>cgroup2.shared = sharedregion1
>
>cgroup3.exclusive = 1M
>cgroup3.shared = sharedregion1
>
>Example 2:
>3 high performance applications running, one of which is a cache hog
>with no cache locality.
>
>cgroup1.exclusive = 8M
>cgroup2.exclusive = 8M
>
>cgroup3.exclusive = 512K
>cgroup3.round_down = Y
>
>In all cases the default cgroup (which requires no explicit
>specification) is expansive and uses the remaining cache
>ways, including the ways shared by other hardware entities.
>
>--
>To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
>the body of a message to majordomo@vger.kernel.org
>More majordomo info at http://vger.kernel.org/majordomo-info.html
>Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Marcelo Tosatti <mtosatti@redhat.com> |
|---|---|
| Date | 2015-08-03 17:20 +0200 |
| Message-ID | <pTpGW-6PV-21@gated-at.bofh.it> |
| In reply to | #1198397 |
On Sun, Aug 02, 2015 at 05:48:07PM +0200, Martin Kletzander wrote:
> On Thu, Jul 30, 2015 at 05:08:13PM -0300, Marcelo Tosatti wrote:
> >On Thu, Jul 30, 2015 at 10:47:23AM -0700, Vikas Shivappa wrote:
> >>
> >>
> >>Marcello,
> >>
> >>
> >>On Wed, 29 Jul 2015, Marcelo Tosatti wrote:
> >>>
> >>>How about this:
> >>>
> >>>desiredclos (closid p1 p2 p3 p4)
> >>> 1 1 0 0 0
> >>> 2 0 0 0 1
> >>> 3 0 1 1 0
> >>
> >>#1 Currently in the rdt cgroup , the root cgroup always has all the
> >>bits set and cant be changed (because the cgroup hierarchy would by
> >>default make this to have all bits as all the children need to have
> >>a subset of the root's bitmask). So if the user creates a cgroup and
> >>not put any task in it , the tasks in the root cgroup could be still
> >>using that part of the cache. Thats the reason i say we can have
> >>really 'exclusive' masks.
> >>
> >>Or in other words - there is always a desired clos (0) which has all
> >>parts set which acts like a default pool.
> >>
> >>Also the parts can overlap. Please apply this for all the below
> >>comments which will change the way they work.
> >>
> >>>
> >>>p means part.
> >>
> >>I am assuming p = (a contiguous cache capacity bit mask)
> >
> >Yes.
> >
> >>>closid 1 is a exclusive cgroup.
> >>>closid 2 is a "cache hog" class.
> >>>closid 3 is "default closid".
> >>>
> >>>Desiredclos is what user has specified.
> >>>
> >>>Transition 1: desiredclos --> effectiveclos
> >>>Clean all bits of unused closid's
> >>>(that must be updated whenever a
> >>>closid1 cgroup goes from empty->nonempty
> >>>and vice-versa).
> >>>
> >>>effectiveclos (closid p1 p2 p3 p4)
> >>> 1 0 0 0 0
> >>> 2 0 0 0 1
> >>> 3 0 1 1 0
> >>
> >>>
> >>>Transition 2: effectiveclos --> expandedclos
> >>>expandedclos (closid p1 p2 p3 p4)
> >>> 1 0 0 0 0
> >>> 2 0 0 0 1
> >>> 3 1 1 1 0
> >>>Then you have different inplacecos for each
> >>>CPU (see pseudo-code below):
> >>>
> >>>On the following events.
> >>>
> >>>- task migration to new pCPU:
> >>>- task creation:
> >>>
> >>> id = smp_processor_id();
> >>> for (part = desiredclos.p1; ...; part++)
> >>> /* if my cosid is set and any other
> >>> cosid is clear, for the part,
> >>> synchronize desiredclos --> inplacecos */
> >>> if (part[mycosid] == 1 &&
> >>> part[any_othercosid] == 0)
> >>> wrmsr(part, desiredclos);
> >>>
> >>
> >>Currently the root cgroup would have all the bits set which will act
> >>like a default cgroup where all the otherwise unused parts (assuming
> >>they are a set of contiguous cache capacity bits) will be used.
> >>
> >>Otherwise the question is in the expandedclos - who decides to
> >>expand the closx parts to include some of the unused parts.. - that
> >>could just be a default root always ?
> >
> >Right, so the problem is for certain closid's you might never want
> >to expand (because doing so would cause data to be cached in a
> >cache way which might have high eviction rate in the future).
> >See the example from Will.
> >
> >But for the default cache (that is "unclassified applications"
> >i suppose it is beneficial to expand in most cases, that is,
> >use maximum amount of cache irrespective of eviction rate, which
> >is the behaviour that exists now without CAT).
> >
> >So perhaps a new flag "expand=y/n" can be added to the cgroup
> >directories... What do you say?
> >
> >Userspace representation of CAT
> >-------------------------------
> >
> >Usage model:
> >1) measure application performance without L3 cache reservation.
> >2) measure application perf with L3 cache reservation and
> >X number of cache ways until desired performance is attained.
> >
> >Requirements:
> >1) Persistency of CLOS configuration across hardware. On migration
> >of operating system or application between different hardware
> >systems we'd like the following to be maintained:
> > - exclusive number of bytes (*) reserved to a certain CLOSid.
> > - shared number of bytes (*) reserved between a certain group
> > of CLOSid's.
> >
> >For both code and data, rounded down or up in cache way size.
> >
> >2) Reasoning:
> >Different CBM masks in different hardware platforms might be necessary
> >to specify the same CLOS configuration, in terms of exclusive number of
> >bytes and shared number of bytes. (cache-way rounded number of bytes).
> >For example, due to L3 allocation by other hardware entities in certain parts
> >of the cache it might be necessary to relocate CBM mask to achieve
> >the same CLOS configuration.
> >
> >3) Proposed format:
> >
>
> Few questions from a random listener, I apologise if some of them are
> in a wrong place due to me missing some information from past threads.
>
> I'm not sure whether the following proposal to the format is the
> internal structure or what's going to be in cgroups. If this is
> user-visible interface, I think it could be a little less detailed.
User visible interface. The idea is to have userspace code that performs
[ user visible specification ] ----> [ cbm bitmasks on present hardware
platform ]
In systemd, probably (or whatever is between the user and the cgroup
interface).
> >sharedregionK.exclusive - Number of exclusive cache bytes reserved for
> > shared region.
> >sharedregionK.excl_data - Number of exclusive cache data bytes reserved for
> > shared region.
> >sharedregionK.excl_bytes - Number of exclusive cache code bytes reserved for
> > shared region.
> >sharedregionK.round_down - Round down to cache way bytes from respective number
> > specification (default is round up).
> >sharedregionK.expand - y/n - Expand shared region to more cache ways
> > when available (default N).
> >
> >cgroupN.exclusive - Number of exclusive L3 cache bytes reserved
> > for cgroup.
> >cgroupN.excl_data - Number of exclusive L3 data cache bytes reserved
> > for cgroup.
> >cgroupN.excl_code - Number of exclusive L3 code cache bytes reserved
> > for cgroup.
>
> By exclusive, you mean that it's exclusive to the tasks in this
> cgroup?
Correct.
> The thing is that we must differentiate between limiting some
> process's from hogging the memory (like example 2 below) and making
> some part of the cache exclusive for particular application (example 1
> below).
AFAICS there is no difference because: both require exclusive cache
access: the hog wants exclusive access between any other user of its
cachelines will be penalized. the high performance application wants
exclusive cache access because any other user of its cachelines will
penalize it.
Where do you see the need to differentiate?
> I just hope we won't need to add something similar to 'isolcpus=' just
> so we can make sure none of the tasks in the root cgroup can spoil the
> part of the cache we need to have exclusive.
>
> I'm not sure creating a new subgroup and moving all the tasks there
> would work, It certainly is not possible with other cgroups, like the
> cpuset cgroup mentioned beforehand.
Why not? Should be able to place all tasks in a given cgroup? (trying
to setup systemd to do that now...).
> I also don't quite fully understand how the co-mounting with the
> cpuset cgroup should work, but that's not design-related.
Neither do I.
> One more question, how does this work on systems with multiple L3
> caches (e.g. large NUMA node systems)? I'm guessing if the process is
> running only on some CPUs, the wrmsr() will be called on that
> particular CPU(s), right?
Not in the current patchset, that has to be fixed...
> >cgroupN.round_down - Round down to cache way bytes from respective number
> > specification (default is round up).
> >cgroupN.expand - y/n - Expand shared region to more cache ways when
> > available (default N).
> >cgroupN.shared = { sharedregion1, sharedregion2, ... } (list of shared
> >regions)
> >
> >Example 1:
> >One application with 2M exclusive cache, two applications
> >with 1M exclusive each, sharing an expansive shared region of 1M.
> >
> >cgroup1.exclusive = 2M
> >
> >sharedregion1.exclusive = 1M
> >sharedregion1.expand = Y
> >
> >cgroup2.exclusive = 1M
> >cgroup2.shared = sharedregion1
> >
> >cgroup3.exclusive = 1M
> >cgroup3.shared = sharedregion1
> >
> >Example 2:
> >3 high performance applications running, one of which is a cache hog
> >with no cache locality.
> >
> >cgroup1.exclusive = 8M
> >cgroup2.exclusive = 8M
> >
> >cgroup3.exclusive = 512K
> >cgroup3.round_down = Y
> >
> >In all cases the default cgroup (which requires no explicit
> >specification) is expansive and uses the remaining cache
> >ways, including the ways shared by other hardware entities.
> >
> >--
> >To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
> >the body of a message to majordomo@vger.kernel.org
> >More majordomo info at http://vger.kernel.org/majordomo-info.html
> >Please read the FAQ at http://www.tux.org/lkml/
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vikas Shivappa <vikas.shivappa@intel.com> |
|---|---|
| Date | 2015-08-03 20:30 +0200 |
| Message-ID | <pTsEO-2EE-41@gated-at.bofh.it> |
| In reply to | #1198938 |
Hello Marcelo/Martin,
Like I mentioned let me modify the documentation to better help understand
the usage. Things like updating each package bitmask is already in the patches.
Lets discuss offline and come up a well defined proposal for change if any and
then update that in next series. We seem to be just looping over same items.
Thanks,
Vikas
On Mon, 3 Aug 2015, Marcelo Tosatti wrote:
> On Sun, Aug 02, 2015 at 05:48:07PM +0200, Martin Kletzander wrote:
>> On Thu, Jul 30, 2015 at 05:08:13PM -0300, Marcelo Tosatti wrote:
>>> On Thu, Jul 30, 2015 at 10:47:23AM -0700, Vikas Shivappa wrote:
>>>>
>>>>
>>>> Marcello,
>>>>
>>>>
>>>> On Wed, 29 Jul 2015, Marcelo Tosatti wrote:
>>>>>
>>>>> How about this:
>>>>>
>>>>> desiredclos (closid p1 p2 p3 p4)
>>>>> 1 1 0 0 0
>>>>> 2 0 0 0 1
>>>>> 3 0 1 1 0
>>>>
>>>> #1 Currently in the rdt cgroup , the root cgroup always has all the
>>>> bits set and cant be changed (because the cgroup hierarchy would by
>>>> default make this to have all bits as all the children need to have
>>>> a subset of the root's bitmask). So if the user creates a cgroup and
>>>> not put any task in it , the tasks in the root cgroup could be still
>>>> using that part of the cache. Thats the reason i say we can have
>>>> really 'exclusive' masks.
>>>>
>>>> Or in other words - there is always a desired clos (0) which has all
>>>> parts set which acts like a default pool.
>>>>
>>>> Also the parts can overlap. Please apply this for all the below
>>>> comments which will change the way they work.
>>>>
>>>>>
>>>>> p means part.
>>>>
>>>> I am assuming p = (a contiguous cache capacity bit mask)
>>>
>>> Yes.
>>>
>>>>> closid 1 is a exclusive cgroup.
>>>>> closid 2 is a "cache hog" class.
>>>>> closid 3 is "default closid".
>>>>>
>>>>> Desiredclos is what user has specified.
>>>>>
>>>>> Transition 1: desiredclos --> effectiveclos
>>>>> Clean all bits of unused closid's
>>>>> (that must be updated whenever a
>>>>> closid1 cgroup goes from empty->nonempty
>>>>> and vice-versa).
>>>>>
>>>>> effectiveclos (closid p1 p2 p3 p4)
>>>>> 1 0 0 0 0
>>>>> 2 0 0 0 1
>>>>> 3 0 1 1 0
>>>>
>>>>>
>>>>> Transition 2: effectiveclos --> expandedclos
>>>>> expandedclos (closid p1 p2 p3 p4)
>>>>> 1 0 0 0 0
>>>>> 2 0 0 0 1
>>>>> 3 1 1 1 0
>>>>> Then you have different inplacecos for each
>>>>> CPU (see pseudo-code below):
>>>>>
>>>>> On the following events.
>>>>>
>>>>> - task migration to new pCPU:
>>>>> - task creation:
>>>>>
>>>>> id = smp_processor_id();
>>>>> for (part = desiredclos.p1; ...; part++)
>>>>> /* if my cosid is set and any other
>>>>> cosid is clear, for the part,
>>>>> synchronize desiredclos --> inplacecos */
>>>>> if (part[mycosid] == 1 &&
>>>>> part[any_othercosid] == 0)
>>>>> wrmsr(part, desiredclos);
>>>>>
>>>>
>>>> Currently the root cgroup would have all the bits set which will act
>>>> like a default cgroup where all the otherwise unused parts (assuming
>>>> they are a set of contiguous cache capacity bits) will be used.
>>>>
>>>> Otherwise the question is in the expandedclos - who decides to
>>>> expand the closx parts to include some of the unused parts.. - that
>>>> could just be a default root always ?
>>>
>>> Right, so the problem is for certain closid's you might never want
>>> to expand (because doing so would cause data to be cached in a
>>> cache way which might have high eviction rate in the future).
>>> See the example from Will.
>>>
>>> But for the default cache (that is "unclassified applications"
>>> i suppose it is beneficial to expand in most cases, that is,
>>> use maximum amount of cache irrespective of eviction rate, which
>>> is the behaviour that exists now without CAT).
>>>
>>> So perhaps a new flag "expand=y/n" can be added to the cgroup
>>> directories... What do you say?
>>>
>>> Userspace representation of CAT
>>> -------------------------------
>>>
>>> Usage model:
>>> 1) measure application performance without L3 cache reservation.
>>> 2) measure application perf with L3 cache reservation and
>>> X number of cache ways until desired performance is attained.
>>>
>>> Requirements:
>>> 1) Persistency of CLOS configuration across hardware. On migration
>>> of operating system or application between different hardware
>>> systems we'd like the following to be maintained:
>>> - exclusive number of bytes (*) reserved to a certain CLOSid.
>>> - shared number of bytes (*) reserved between a certain group
>>> of CLOSid's.
>>>
>>> For both code and data, rounded down or up in cache way size.
>>>
>>> 2) Reasoning:
>>> Different CBM masks in different hardware platforms might be necessary
>>> to specify the same CLOS configuration, in terms of exclusive number of
>>> bytes and shared number of bytes. (cache-way rounded number of bytes).
>>> For example, due to L3 allocation by other hardware entities in certain parts
>>> of the cache it might be necessary to relocate CBM mask to achieve
>>> the same CLOS configuration.
>>>
>>> 3) Proposed format:
>>>
>>
>> Few questions from a random listener, I apologise if some of them are
>> in a wrong place due to me missing some information from past threads.
>>
>> I'm not sure whether the following proposal to the format is the
>> internal structure or what's going to be in cgroups. If this is
>> user-visible interface, I think it could be a little less detailed.
>
> User visible interface. The idea is to have userspace code that performs
>
> [ user visible specification ] ----> [ cbm bitmasks on present hardware
> platform ]
>
> In systemd, probably (or whatever is between the user and the cgroup
> interface).
>
>>> sharedregionK.exclusive - Number of exclusive cache bytes reserved for
>>> shared region.
>>> sharedregionK.excl_data - Number of exclusive cache data bytes reserved for
>>> shared region.
>>> sharedregionK.excl_bytes - Number of exclusive cache code bytes reserved for
>>> shared region.
>>> sharedregionK.round_down - Round down to cache way bytes from respective number
>>> specification (default is round up).
>>> sharedregionK.expand - y/n - Expand shared region to more cache ways
>>> when available (default N).
>>>
>>> cgroupN.exclusive - Number of exclusive L3 cache bytes reserved
>>> for cgroup.
>>> cgroupN.excl_data - Number of exclusive L3 data cache bytes reserved
>>> for cgroup.
>>> cgroupN.excl_code - Number of exclusive L3 code cache bytes reserved
>>> for cgroup.
>>
>> By exclusive, you mean that it's exclusive to the tasks in this
>> cgroup?
>
> Correct.
>
>> The thing is that we must differentiate between limiting some
>> process's from hogging the memory (like example 2 below) and making
>> some part of the cache exclusive for particular application (example 1
>> below).
>
> AFAICS there is no difference because: both require exclusive cache
> access: the hog wants exclusive access between any other user of its
> cachelines will be penalized. the high performance application wants
> exclusive cache access because any other user of its cachelines will
> penalize it.
>
> Where do you see the need to differentiate?
>
>> I just hope we won't need to add something similar to 'isolcpus=' just
>> so we can make sure none of the tasks in the root cgroup can spoil the
>> part of the cache we need to have exclusive.
>>
>> I'm not sure creating a new subgroup and moving all the tasks there
>> would work, It certainly is not possible with other cgroups, like the
>> cpuset cgroup mentioned beforehand.
>
> Why not? Should be able to place all tasks in a given cgroup? (trying
> to setup systemd to do that now...).
>
>> I also don't quite fully understand how the co-mounting with the
>> cpuset cgroup should work, but that's not design-related.
>
> Neither do I.
>
>> One more question, how does this work on systems with multiple L3
>> caches (e.g. large NUMA node systems)? I'm guessing if the process is
>> running only on some CPUs, the wrmsr() will be called on that
>> particular CPU(s), right?
>
> Not in the current patchset, that has to be fixed...
>
>>> cgroupN.round_down - Round down to cache way bytes from respective number
>>> specification (default is round up).
>>> cgroupN.expand - y/n - Expand shared region to more cache ways when
>>> available (default N).
>>> cgroupN.shared = { sharedregion1, sharedregion2, ... } (list of shared
>>> regions)
>>>
>>> Example 1:
>>> One application with 2M exclusive cache, two applications
>>> with 1M exclusive each, sharing an expansive shared region of 1M.
>>>
>>> cgroup1.exclusive = 2M
>>>
>>> sharedregion1.exclusive = 1M
>>> sharedregion1.expand = Y
>>>
>>> cgroup2.exclusive = 1M
>>> cgroup2.shared = sharedregion1
>>>
>>> cgroup3.exclusive = 1M
>>> cgroup3.shared = sharedregion1
>>>
>>> Example 2:
>>> 3 high performance applications running, one of which is a cache hog
>>> with no cache locality.
>>>
>>> cgroup1.exclusive = 8M
>>> cgroup2.exclusive = 8M
>>>
>>> cgroup3.exclusive = 512K
>>> cgroup3.round_down = Y
>>>
>>> In all cases the default cgroup (which requires no explicit
>>> specification) is expansive and uses the remaining cache
>>> ways, including the ways shared by other hardware entities.
>>>
>>> --
>>> To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
>>> the body of a message to majordomo@vger.kernel.org
>>> More majordomo info at http://vger.kernel.org/majordomo-info.html
>>> Please read the FAQ at http://www.tux.org/lkml/
>
>
>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Marcelo Tosatti <mtosatti@redhat.com> |
|---|---|
| Date | 2015-07-31 17:20 +0200 |
| Message-ID | <pSkgj-19Y-39@gated-at.bofh.it> |
| In reply to | #1195462 |
On Thu, Jul 30, 2015 at 04:03:07PM -0700, Vikas Shivappa wrote: > > > On Thu, 30 Jul 2015, Marcelo Tosatti wrote: > > >On Thu, Jul 30, 2015 at 10:47:23AM -0700, Vikas Shivappa wrote: > >> > >> > >>Marcello, > >> > >> > >>On Wed, 29 Jul 2015, Marcelo Tosatti wrote: > >>> > >>>How about this: > >>> > >>>desiredclos (closid p1 p2 p3 p4) > >>> 1 1 0 0 0 > >>> 2 0 0 0 1 > >>> 3 0 1 1 0 > >> > >>#1 Currently in the rdt cgroup , the root cgroup always has all the > >>bits set and cant be changed (because the cgroup hierarchy would by > >>default make this to have all bits as all the children need to have > >>a subset of the root's bitmask). So if the user creates a cgroup and > >>not put any task in it , the tasks in the root cgroup could be still > >>using that part of the cache. Thats the reason i say we can have > >>really 'exclusive' masks. > >> > >>Or in other words - there is always a desired clos (0) which has all > >>parts set which acts like a default pool. > >> > >>Also the parts can overlap. Please apply this for all the below > >>comments which will change the way they work. > > > > > >> > >>> > >>>p means part. > >> > >>I am assuming p = (a contiguous cache capacity bit mask) > >> > >>>closid 1 is a exclusive cgroup. > >>>closid 2 is a "cache hog" class. > >>>closid 3 is "default closid". > >>> > >>>Desiredclos is what user has specified. > >>> > >>>Transition 1: desiredclos --> effectiveclos > >>>Clean all bits of unused closid's > >>>(that must be updated whenever a > >>>closid1 cgroup goes from empty->nonempty > >>>and vice-versa). > >>> > >>>effectiveclos (closid p1 p2 p3 p4) > >>> 1 0 0 0 0 > >>> 2 0 0 0 1 > >>> 3 0 1 1 0 > >> > >>> > >>>Transition 2: effectiveclos --> expandedclos > >>>expandedclos (closid p1 p2 p3 p4) > >>> 1 0 0 0 0 > >>> 2 0 0 0 1 > >>> 3 1 1 1 0 > >>>Then you have different inplacecos for each > >>>CPU (see pseudo-code below): > >>> > >>>On the following events. > >>> > >>>- task migration to new pCPU: > >>>- task creation: > >>> > >>> id = smp_processor_id(); > >>> for (part = desiredclos.p1; ...; part++) > >>> /* if my cosid is set and any other > >>> cosid is clear, for the part, > >>> synchronize desiredclos --> inplacecos */ > >>> if (part[mycosid] == 1 && > >>> part[any_othercosid] == 0) > >>> wrmsr(part, desiredclos); > >>> > >> > >>Currently the root cgroup would have all the bits set which will act > >>like a default cgroup where all the otherwise unused parts (assuming > >>they are a set of contiguous cache capacity bits) will be used. > > > >Right, but we don't want to place tasks in there in case one cgroup > >wants exclusive cache access. > > > >So whenever you want an exclusive cgroup you'd do: > > > >create cgroup-exclusive; reserve desired part of the cache > >for it. > >create cgroup-default; reserved all cache minus that of cgroup-exclusive > >for it. > > > >place tasks that belong to cgroup-exclusive into it. > >place all other tasks (including init) into cgroup-default. > > > >Is that right? > > Yes you could do that. > > You can create cgroups to have masks which are exclusive in todays > implementation, just that you could also created more cgroups to > overlap the masks again.. iow we dont have an exclusive flag for the > cgroup mask. > Is that a common use case in the server environment that you need to > prevent other cgroups from using a certain mask ? (since the root > user should control these allocations .. he should know?) Yes, there are two known use-cases that have this characteristic: 1) High performance numeric application which has been optimized to a certain fraction of the cache. 2) Low latency application in multi-application OS. For both cases exclusive cache access is wanted. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vikas Shivappa <vikas.shivappa@intel.com> |
|---|---|
| Date | 2015-07-31 18:50 +0200 |
| Subject | [summary] Re: [PATCH 3/9] x86/intel_rdt: Cache Allocation documentation and cgroup usage guide |
| Message-ID | <pSlFp-33y-21@gated-at.bofh.it> |
| In reply to | #1197107 |
To summarize the ever growing thread : 1. the rdt_cgroup can be used to configure exclusive cache bitmaps for the child nodes which can be used for the scenarios which Marcello mentions. simle examples which were mentioned : max bitmask length : 16 . hence full mask is 0xffff groupx_realtime - 0xff . group2_systemtraffic - 0xf. : put a lot of tasks from root node to here or which ever is offending and thrashing. groupy_<mytraffic> - 0x0f Now the groupx has its own area of cache that can used by the realtime/(specific scenario) apps. Similarly configure any groupy. 2. Can the maps can let you specify which cache ways ways the cache is allocated ? - No , this is implementation specific as mentioned in the SDM. So when we configure a mask , you really dont know which ways or which exact lines are used on which SKUs .. We may not see any use case as well which is needed for apps to allocate cache in specific areas and the h/w does not support this as well. 3. Letting the user specify size in bytes instead of bitmap : we have already gone through this discussion in older versions. The user can simply check the size of the total cache and understand what map could be what size. I dont see a special need to specify an interface to enter the cache in bytes and then round off - user could instead use the roundoff values before hand or iow it automatically does when he specifies the bitmask. ex: find cache size from /proc/cpuinfo. - say 20MB bitmask max - 0xfffff. This means the roundoff(chunk) size supported is only 1MB , so when you specify the mask say 0x3(2MB) thats already taken care of. Same applies to percentage - the masks automatically round off the percentage. Please note that this is quite different from the way we can allocate memory in bytes and needs to be treated differently given that the hardware provides interface in a particular way. 4. Letting the kernel automatically extend the bitmap may affect a lot of other things and will need a lot of heuristics - note that we have overlapping masks . This interface lets the super-user control the cache allocation and it may be very confusing for the user if he has allocated a cache mask and suddenly from under the floor the kernel changes it. Thanks, Vikas On Fri, 31 Jul 2015, Marcelo Tosatti wrote: > On Thu, Jul 30, 2015 at 04:03:07PM -0700, Vikas Shivappa wrote: >> >> >> On Thu, 30 Jul 2015, Marcelo Tosatti wrote: >> >>> On Thu, Jul 30, 2015 at 10:47:23AM -0700, Vikas Shivappa wrote: >>>> >>>> >>>> Marcello, >>>> >>>> >>>> On Wed, 29 Jul 2015, Marcelo Tosatti wrote: >>>>> >>>>> How about this: >>>>> >>>>> desiredclos (closid p1 p2 p3 p4) >>>>> 1 1 0 0 0 >>>>> 2 0 0 0 1 >>>>> 3 0 1 1 0 >>>> >>>> #1 Currently in the rdt cgroup , the root cgroup always has all the >>>> bits set and cant be changed (because the cgroup hierarchy would by >>>> default make this to have all bits as all the children need to have >>>> a subset of the root's bitmask). So if the user creates a cgroup and >>>> not put any task in it , the tasks in the root cgroup could be still >>>> using that part of the cache. Thats the reason i say we can have >>>> really 'exclusive' masks. >>>> >>>> Or in other words - there is always a desired clos (0) which has all >>>> parts set which acts like a default pool. >>>> >>>> Also the parts can overlap. Please apply this for all the below >>>> comments which will change the way they work. >>> >>> >>>> >>>>> >>>>> p means part. >>>> >>>> I am assuming p = (a contiguous cache capacity bit mask) >>>> >>>>> closid 1 is a exclusive cgroup. >>>>> closid 2 is a "cache hog" class. >>>>> closid 3 is "default closid". >>>>> >>>>> Desiredclos is what user has specified. >>>>> >>>>> Transition 1: desiredclos --> effectiveclos >>>>> Clean all bits of unused closid's >>>>> (that must be updated whenever a >>>>> closid1 cgroup goes from empty->nonempty >>>>> and vice-versa). >>>>> >>>>> effectiveclos (closid p1 p2 p3 p4) >>>>> 1 0 0 0 0 >>>>> 2 0 0 0 1 >>>>> 3 0 1 1 0 >>>> >>>>> >>>>> Transition 2: effectiveclos --> expandedclos >>>>> expandedclos (closid p1 p2 p3 p4) >>>>> 1 0 0 0 0 >>>>> 2 0 0 0 1 >>>>> 3 1 1 1 0 >>>>> Then you have different inplacecos for each >>>>> CPU (see pseudo-code below): >>>>> >>>>> On the following events. >>>>> >>>>> - task migration to new pCPU: >>>>> - task creation: >>>>> >>>>> id = smp_processor_id(); >>>>> for (part = desiredclos.p1; ...; part++) >>>>> /* if my cosid is set and any other >>>>> cosid is clear, for the part, >>>>> synchronize desiredclos --> inplacecos */ >>>>> if (part[mycosid] == 1 && >>>>> part[any_othercosid] == 0) >>>>> wrmsr(part, desiredclos); >>>>> >>>> >>>> Currently the root cgroup would have all the bits set which will act >>>> like a default cgroup where all the otherwise unused parts (assuming >>>> they are a set of contiguous cache capacity bits) will be used. >>> >>> Right, but we don't want to place tasks in there in case one cgroup >>> wants exclusive cache access. >>> >>> So whenever you want an exclusive cgroup you'd do: >>> >>> create cgroup-exclusive; reserve desired part of the cache >>> for it. >>> create cgroup-default; reserved all cache minus that of cgroup-exclusive >>> for it. >>> >>> place tasks that belong to cgroup-exclusive into it. >>> place all other tasks (including init) into cgroup-default. >>> >>> Is that right? >> >> Yes you could do that. >> >> You can create cgroups to have masks which are exclusive in todays >> implementation, just that you could also created more cgroups to >> overlap the masks again.. iow we dont have an exclusive flag for the >> cgroup mask. >> Is that a common use case in the server environment that you need to >> prevent other cgroups from using a certain mask ? (since the root >> user should control these allocations .. he should know?) > > Yes, there are two known use-cases that have this characteristic: > > 1) High performance numeric application which has been optimized > to a certain fraction of the cache. > > 2) Low latency application in multi-application OS. > > For both cases exclusive cache access is wanted. > > -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vikas Shivappa <vikas.shivappa@intel.com> |
|---|---|
| Date | 2015-07-29 22:10 +0200 |
| Message-ID | <pRFPQ-1zd-13@gated-at.bofh.it> |
| In reply to | #1194760 |
On Tue, 28 Jul 2015, Auld, Will wrote: > > >> -----Original Message----- >> >> Same comment as above - Cgroup masks can always overlap and other cgroups >> can allocate the same cache , and hence wont have exclusive cache allocation. > > [Auld, Will] You can define all the cbm to provide one clos with an exclusive area Do you mean a CLOS that has all the bits set. We donot support exclusive area today. The bits in the mask can overlap .. hence can always share the same cache allocation . > >> >> So natuarally the cgroup with tasks would get to use the cache if it has the same >> mask (say representing 50% of cache in your example) as others . > > [Auld, Will] automatic adjustment of the cbm make me nervous. There are times > when we want to limit the cache for a process independent of whether there is > lots of unused cache. > Please see example below - In general , I just mean the cache mask can have bits that can overlap - does not matter whether there is tasks in it or not. > >> (assume there are 8 bits max cbm) >> cgroupa - mask - 0xf >> cgroupb - mask - 0xf . Now if cgroupa has no tasks , cgroupb naturally gets all >> the cache. >> >> Thanks, >> Vikas -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web