Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1441777 > unrolled thread

[PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface

Started by"Fenghua Yu" <fenghua.yu@intel.com>
First post2016-07-13 00:10 +0200
Last post2016-07-19 14:40 +0200
Articles 6 — 3 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:10 +0200
    Re: [PATCH 13/32] Documentation, x86: Documentation for Intel resource  allocation user interface Thomas Gleixner <tglx@linutronix.de> - 2016-07-13 15:00 +0200
      Re: [PATCH 13/32] Documentation, x86: Documentation for Intel  resource allocation user interface "Luck, Tony" <tony.luck@intel.com> - 2016-07-13 19:20 +0200
        Re: [PATCH 13/32] Documentation, x86: Documentation for Intel resource  allocation user interface Thomas Gleixner <tglx@linutronix.de> - 2016-07-14 09:00 +0200
          Re: [PATCH 13/32] Documentation, x86: Documentation for Intel  resource allocation user interface "Luck, Tony" <tony.luck@intel.com> - 2016-07-14 19:20 +0200
            Re: [PATCH 13/32] Documentation, x86: Documentation for Intel resource  allocation user interface Thomas Gleixner <tglx@linutronix.de> - 2016-07-19 14:40 +0200

#1441777 — [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface

From"Fenghua Yu" <fenghua.yu@intel.com>
Date2016-07-13 00:10 +0200
Subject[PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface
Message-ID<rUe2o-2xx-83@gated-at.bofh.it>
From: Fenghua Yu <fenghua.yu@intel.com>

The documentation describes user interface of how to allocate resource
in Intel RDT.

Please note that the documentation covers generic user interface. Current
patch set code only implemente CAT L3. CAT L2 code will be sent later.

Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Reviewed-by: Tony Luck <tony.luck@intel.com>
---
 Documentation/x86/intel_rdt_ui.txt | 268 +++++++++++++++++++++++++++++++++++++
 1 file changed, 268 insertions(+)
 create mode 100644 Documentation/x86/intel_rdt_ui.txt

diff --git a/Documentation/x86/intel_rdt_ui.txt b/Documentation/x86/intel_rdt_ui.txt
new file mode 100644
index 0000000..c52baf5
--- /dev/null
+++ b/Documentation/x86/intel_rdt_ui.txt
@@ -0,0 +1,268 @@
+User Interface for Resource Allocation in Intel Resource Director Technology
+
+Copyright (C) 2016 Intel Corporation
+
+Fenghua Yu <fenghua.yu@intel.com>
+
+We create a new file system rscctrl in /sys/fs as user interface for Cache
+Allocation Technology (CAT) and future resource allocations in Intel
+Resource Director Technology (RDT). User can allocate cache or other
+resources to tasks or cpus through this interface.
+
+CONTENTS
+========
+
+	1. Terms
+	2. Mount rscctrl file system
+	3. Hierarchy in rscctrl
+	4. Create and remove sub-directory
+	5. Add/remove a task in a partition
+	6. Add/remove a CPU in a partition
+	7. Some usage examples
+
+
+1. Terms
+========
+
+We use the following terms and concepts in this documentation.
+
+RDT: Intel Resoure Director Technology
+
+CAT: Cache Allocation Technology
+
+CDP: Code and Data Prioritization
+
+CBM: Cache Bit Mask
+
+Cache ID: A cache identification. It is unique in one cache index on the
+platform. User can find cache ID in cache sysfs interface:
+/sys/devices/system/cpu/cpu*/cache/index*/id
+
+Share resource domain: A few different resources can share same QoS mask
+MSRs array. For example, one L2 cache can share QoS MSRs with its next level
+L3 cache. A domain number represents the L2 cache, the L3 cache, the L2
+cache's shared cpumask, and the L3 cache's shared cpumask.
+
+2. Mount rscctrl file system
+============================
+
+Like other file systems, the rscctrl file system needs to be mounted before
+it can be used.
+
+mount -t rscctrl rscctrl <-o cdp,verbose> /sys/fs/rscctrl
+
+This command mounts the rscctrl file system under /sys/fs/rscctrl.
+
+Options are optional:
+
+cdp: Enable Code and Data Prioritization (CDP). Without the option, CDP
+is disabled.
+
+verbose: Output more info in the "info" file under info directory and in
+dmesg. This is mainly for debug.
+
+
+3. Hierarchy in rscctrl
+=======================
+
+The initial hierarchy of the rscctrl file system is as follows after mount:
+
+/sys/fs/rscctrl/info/info
+		    /<resource0>/<resource0 specific info files>
+		    /<resource1>/<resource1 specific info files>
+			....
+	       /tasks
+	       /cpus
+	       /schemas
+
+There are a few files and sub-directories in the hierarchy.
+
+3.1. info
+---------
+
+The read-only sub-directory "info" in root directory has RDT related
+system info.
+
+The "info" file under the info sub-directory shows general info of the system.
+It shows shared domain and the resources within this domain.
+
+Each resource has its own info sub-directory. User can read the information
+for allocation. For example, l3 directory has max_closid, max_cbm_len,
+domain_to_cache_id.
+
+3.2. tasks
+----------
+
+The file "tasks" has all task ids in the root directory initially. The
+thread ids in the file will be added or removed among sub-directories or
+partitions. A task id only stays in one directory at the same time.
+
+3.3. cpus
+
+The file "cpus" has a cpu mask that specifies the CPUs that are bound to the
+schemas. Any tasks scheduled on the cpus will use the schemas. User can set
+both "cpus" and "tasks" to share the same schema in one directory. But when
+a CPU is bound to a schema, a task running on the CPU uses this schema and
+kernel will ignore scheam set up for the task in "tasks".
+
+Initial value is all zeros which means there is no CPU bound to the schemas
+in the root directory and tasks use the schemas.
+
+3.4. schemas
+------------
+
+The file "schemas" has default allocation masks/values for all resources on
+each socket/cpu. Format of the file "schemas" is in multiple lines and each
+line represents masks or values for one resource.
+
+Format of one resource schema line is as follows:
+
+<resource name>:<resource id0>=<schema>;<resource id1>=<schema>;...
+
+As one example, CAT L3's schema format is:
+
+L3:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...
+
+On a two socket machine, L3's schema line could be:
+
+L3:0=ff;1=c0
+
+which means this line in "schemas" file is for CAT L3, L3 cache id 0's CBM
+is 0xff, and L3 cache id 1's CBM is 0xc0.
+
+If one resource is disabled, its line is not shown in schemas file.
+
+The schema line can be expended for situations. L3 cbms format can be
+expended to CDP enabled L3 cbms format:
+
+L3:<cache_id0>=<d_cbm>,<i_cbm>;<cache_id1>=<d_cbm>,<i_cbm>;...
+
+Initial value is all ones which means all tasks use all resources initially.
+
+4. Create and remove sub-directory
+===================================
+
+User can create a sub-directory under the root directory by "mkdir" command.
+User can remove the sub-directory by "rmdir" command.
+
+Each sub-directory represents a resource allocation policy that user can
+allocate resources for tasks or cpus.
+
+Each directory has three files "tasks", "cpus", and "schemas". The meaning
+of each file is same as the files in the root directory.
+
+When a directory is created, initial contents of the files are:
+
+tasks: Empty. This means no task currently uses this allocation schemas.
+cpus: All zeros. This means no CPU uses this allocation schemas.
+schemas: All ones. This means all resources can be used in this allocation.
+
+5. Add/remove a task in a partition
+===================================
+
+User can add/remove a task by writing its PID in "tasks" in a partition.
+User can read PIDs stored in one "tasks" file.
+
+One task PID only exists in one partition/directory at the same time. If PID
+is written in a new directory, it's removed automatically from its last
+directory.
+
+6. Add/remove a CPU in a partition
+==================================
+
+User can add/remove a CPU by writing its bit in "cpus" in a partition.
+User can read CPUs stored in one "cpus" file.
+
+One CPU only exists in one partition/directory if user wants it to be bound
+to any "schemas". Kernel guarantees uniqueness of the CPU in the whole
+directory to make sure it only uses one schemas. If a CPU is written in one
+new directory, it's automatically removed from its original directory if it
+exists in the original directory.
+
+Or it doesn't exist in the whole directory if user doesn't bind it to any
+"schemas".
+
+7. Some usage examples
+======================
+
+7.1 Example 1 for sharing CLOSID on socket 0 between two partitions
+
+Only L3 cbm is enabled. Assume the machine is 2-socket and dual-core without
+hyperthreading.
+
+#mount -t rscctrl rscctrl /sys/fs/rscctrl
+#cd /sys/fs/rscctrl
+#mkdir p0 p1
+#echo "L3:0=3;1=c" > /sys/fs/rscctrl/p0/schemas
+#echo "L3:0=3;1=3" > /sys/fs/rscctrl/p1/schemas
+
+In partition p0, kernel allocates CLOSID 0 for L3 cbm=0x3 on socket 0 and
+CLOSID 0 for cbm=0xc on socket 1.
+
+In partition p1, kernel allocates CLOSID 0 for L3 cbm=0x3 on socket 0 and
+CLOSID 1 for cbm=0x3 on socket 1.
+
+When p1/schemas is updated for socket 0, kernel searches existing
+IA32_L3_QOS_MASK_n MSR registers and finds that 0x3 is in IA32_L3_QOS_MASK_0
+register already. Therefore CLOSID 0 is shared between partition 0 and
+partition 1 on socket 0.
+
+When p1/schemas is udpated for socket 1, kernel searches existing
+IA32_L3_QOS_MASK_n registers and doesn't find a matching cbm. Therefore
+CLOSID 1 is created and IA32_L3_QOS_MASK_1=0xc.
+
+7.2 Example 2 for allocating L3 cache for real-time apps
+
+Two real time tasks pid=1234 running on processor 0 and pid=5678 running on
+processor 1 on socket 0 on a 2-socket and dual core machine. To avoid noisy
+neighbors, each of the two real-time tasks exclusively occupies one quarter
+of L3 cache on socket 0. Assume L3 cbm max width is 20 bits.
+
+#mount -t rscctrl rscctrl /sys/fs/rscctrl
+#cd /sys/fs/rscctrl
+#mkdir p0 p1
+#taskset 0x1 1234
+#taskset 0x2 5678
+#cd /sys/fs/rscctrl/
+#edit schemas to have following allocation:
+L3:0=3ff;1=fffff
+
+which means that all tasks use whole L3 cache 1 and half of L3 cache 0.
+
+#cd ..
+#mkdir p1 p2
+#cd p1
+#echo 1234 >tasks
+#edit schemas to have following two lines:
+L3:0=f8000;1=fffff
+
+which means task 1234 uses L3 cbm=0xf8000, i.e. one quarter of L3 cache 0
+and whole L3 cache 1.
+
+Since 1234 is tied to processor 0, it actually uses the quarter of L3
+on socket 0 only.
+
+#cd ../p2
+#echo 5678 >tasks
+#edit schemas to have following two lines:
+L3:0=7c00;1=fffff
+
+Which means that task 5678 uses L3 cbm=0x7c00, another quarter of L3 cache 0
+and whole L3 cache 1.
+
+Since 5678 is tied to processor 1, it actually only uses the quarter of L3
+on socket 0.
+
+Internally three CLOSIDs are allocated on L3 cache 0:
+IA32_L3_QOS_MASK_0 = 0x3ff
+IA32_L3_QOS_MASK_1 = 0xf8000
+IA32_L3_QOS_MASK_2 = 0x7c00.
+
+Each CLOSID's reference count=1 on L3 cache 0. There is no shared cbms on
+cache 0.
+
+Only one CLOSID is allocated on L3 cache 1:
+
+IA32_L3_QOS_MASK_0=0xfffff. It's shared by root, p1 and p2.
+
+Therefore CLOSID 0's reference count=3 on L3 cache 1.
-- 
2.5.0

[toc] | [next] | [standalone]


#1442416 — Re: [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface

FromThomas Gleixner <tglx@linutronix.de>
Date2016-07-13 15:00 +0200
SubjectRe: [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface
Message-ID<rUrVD-3bS-3@gated-at.bofh.it>
In reply to#1441777
On Tue, 12 Jul 2016, Fenghua Yu wrote:
> +3. Hierarchy in rscctrl
> +=======================

What means rscctrl?

You were not able to find a more cryptic acronym?

> +
> +The initial hierarchy of the rscctrl file system is as follows after mount:
> +
> +/sys/fs/rscctrl/info/info
> +		    /<resource0>/<resource0 specific info files>
> +		    /<resource1>/<resource1 specific info files>
> +			....
> +	       /tasks
> +	       /cpus
> +	       /schemas
> +
> +There are a few files and sub-directories in the hierarchy.

Shouldn't that read:

The following files and sub-directories are available:

> +3.1. info
> +---------

Those sub points want to be indented so it's clear where they belong to.

> +
> +The read-only sub-directory "info" in root directory has RDT related
> +system info.
> +
> +The "info" file under the info sub-directory shows general info of the system.
> +It shows shared domain and the resources within this domain.
> +
> +Each resource has its own info sub-directory. User can read the information
> +for allocation. For example, l3 directory has max_closid, max_cbm_len,
> +domain_to_cache_id.

Can you please restructure this so it's more obvious what you want to explain.

    The "info" directory contains read-only system information:
  
    3.1.1 info

    The read-only file 'info' contains general information of the resource
    control facility:

    - Shared domains and the resources associated to those domains

    3.1.2 resources

    Each resource has its seperate sub directory, which contains resource
    specific information.

    3.1.2.1 L3 specific files

    - max_closid:		The maximum number of available closids
      				(explain closid ....)
    - max_cbm_len:     	       	....
    - domain_to_cache_id:      	....

So when you add L2 then you can add a proper description of the L2 related
files.

> +3.2. tasks
> +----------
> +
> +The file "tasks" has all task ids in the root directory initially.

This does not make sense.

   The tasks file contains all thread ids which are associated to the root
   resource partition. Initially all threads are associated to this.

   Threads can be moved to other tasks files in resource partitions. A thread
   can only be associated with a single resource partition.

> +thread ids in the file will be added or removed among sub-directories or
> +partitions. A task id only stays in one directory at the same time.

Is a task required to be associated to at least one 'tasks' file?

> +3.3. cpus
> +
> +The file "cpus" has a cpu mask that specifies the CPUs that are bound to the
> +schemas.

Please explain the concept of schemata (I prefer schemata as plural of schema,
but that's just my preference) before explaining what the cpumask in this file
means.

> +Any tasks scheduled on the cpus will use the schemas. User can set
> +both "cpus" and "tasks" to share the same schema in one directory. But when
> +a CPU is bound to a schema, a task running on the CPU uses this schema and
> +kernel will ignore scheam set up for the task in "tasks".

This does not make any sense. 

When a task is bound to a schema then this should have preference over the
schema which is associated to the CPU. The CPU association is meant for tasks
which are not bound to a particular partition/schema.

So the initial setup should be:

   - All CPUs are associated to the root resource partition

   - No thread is associated to a particular resource partition

When a thread is added to a 'tasks' file of a partition then this partition
takes preference. If it's removed, i.e. the association to a partition is
undone, then the CPU association is used.

I have no idea why you think that all threads should be in a tasks file by
default. Associating CPUs in the first place makes a lot more sense as it
represents the topology of the system nicely.

> +Initial value is all zeros which means there is no CPU bound to the schemas
> +in the root directory and tasks use the schemas.

As I said above this is backwards.

> +3.4. schemas
> +------------
> +
> +The file "schemas" has default allocation masks/values for all resources on
> +each socket/cpu. Format of the file "schemas" is in multiple lines and each
> +line represents masks or values for one resource.

You really want to explain that the 'tasks', 'cpus' and 'schemata' files are
available on all levels of the resource hierarchy. The special case of the
files in the root partition is, that there default values are set when the
facility is initialized.

> +Format of one resource schema line is as follows:
> +
> +<resource name>:<resource id0>=<schema>;<resource id1>=<schema>;...

> +As one example, CAT L3's schema format is:

That's crap. You want a proper sub point explaining the L3 schema format and
not 'one example'.

  3.4.1 L3 schema

  L3 resource ids are the L3 domains, which are currently per socket.

  The format for CBM only partitioning is:

      L3:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...

      <cbm> is the cache allocation bitmask in hex    

      Example:

	L3:0=ff;1=c0;

      	Explanation of example ....
 
  For CBM and CDP paritioning the format is:

      L3:<cache_id0>=<d_cbm>,<i_cbm>;<cache_id1>=<d_cbm>,<i_cbm>;...

      Example:
         ....

> +If one resource is disabled, its line is not shown in schemas file.

That means:	  

     Resources which are not described in a schemata file are disabled for
     that particular partition.

Right?

Now that raises the question how this is supposed to work. Let's assume that
we have a partition 'foo' and thread X is in the tasks file of that
partition. The schema of that partition contains only an L2 entry. What's the
L3 association for thread X? Nothing at all?

> +The schema line can be expended for situations. L3 cbms format can be

You probably wanted to say extended, right?

> +4. Create and remove sub-directory
> +===================================

What is the meaning of a 'sub-directory'. I assume it's a resource
partition. So this chapter should be named so. The fact that the partition is
based on a directory is just an implementation detail.

> +User can create a sub-directory under the root directory by "mkdir" command.
> +User can remove the sub-directory by "rmdir" command.

User? Any user?

> +
> +Each sub-directory represents a resource allocation policy that user can
> +allocate resources for tasks or cpus.
> +
> +Each directory has three files "tasks", "cpus", and "schemas". The meaning
> +of each file is same as the files in the root directory.
> +
> +When a directory is created, initial contents of the files are:
> +
> +tasks: Empty. This means no task currently uses this allocation schemas.
> +cpus: All zeros. This means no CPU uses this allocation schemas.
> +schemas: All ones. This means all resources can be used in this allocation.

> +5. Add/remove a task in a partition
> +===================================
> +
> +User can add/remove a task by writing its PID in "tasks" in a partition.
> +User can read PIDs stored in one "tasks" file.
> +
> +One task PID only exists in one partition/directory at the same time. If PID
> +is written in a new directory, it's removed automatically from its last
> +directory.

Please use partition consistently. Aside of that this belongs to the
description of the 'tasks' file.

> +
> +6. Add/remove a CPU in a partition
> +==================================
> +
> +User can add/remove a CPU by writing its bit in "cpus" in a partition.
> +User can read CPUs stored in one "cpus" file.

Any (l)user ?

> +One CPU only exists in one partition/directory if user wants it to be bound
> +to any "schemas". Kernel guarantees uniqueness of the CPU in the whole
> +directory to make sure it only uses one schemas. If a CPU is written in one

   ^^^^^^^^^
You mean hierarchy here, right?

> +new directory, it's automatically removed from its original directory if it
> +exists in the original directory.

Please use partition not directory.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


#1442650 — Re: [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface

From"Luck, Tony" <tony.luck@intel.com>
Date2016-07-13 19:20 +0200
SubjectRe: [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface
Message-ID<rUvZg-63d-29@gated-at.bofh.it>
In reply to#1442416
On Wed, Jul 13, 2016 at 02:47:30PM +0200, Thomas Gleixner wrote:
> On Tue, 12 Jul 2016, Fenghua Yu wrote:
> > +3. Hierarchy in rscctrl
> > +=======================
> 
> What means rscctrl?
> 
> You were not able to find a more cryptic acronym?

rscctrl == resource control

Intel marketing would (probably) like us to use:

   /sys/fs/Intel(R) Resource Director Technology(TM)/

Happy to take suggestions for something in between those
extremes :-)

> > +Any tasks scheduled on the cpus will use the schemas. User can set
> > +both "cpus" and "tasks" to share the same schema in one directory. But when
> > +a CPU is bound to a schema, a task running on the CPU uses this schema and
> > +kernel will ignore scheam set up for the task in "tasks".
> 
> This does not make any sense. 
> 
> When a task is bound to a schema then this should have preference over the
> schema which is associated to the CPU. The CPU association is meant for tasks
> which are not bound to a particular partition/schema.
> 
> So the initial setup should be:
> 
>    - All CPUs are associated to the root resource partition
> 
>    - No thread is associated to a particular resource partition
> 
> When a thread is added to a 'tasks' file of a partition then this partition
> takes preference. If it's removed, i.e. the association to a partition is
> undone, then the CPU association is used.
> 
> I have no idea why you think that all threads should be in a tasks file by
> default. Associating CPUs in the first place makes a lot more sense as it
> represents the topology of the system nicely.

If we did it that way, it would be harder to change the default
resources.  E.g. now we start with all processes in the root
rdtgroup.  We can change the schema for the root group and restrict
them to, say, 60% of L3 cache on one (or all) sockets - giving us
40% of cache to give out to one or more groups.

So what we've implemented (and perhaps need to explain better here)
is that every thread always belongs to one (and only one) rdtgroup.
It will use the resources described in that group whereever it runs,
except in the case where we have designated some cpus as special snowflakes.
When a cpu is assigned to an rdtgroup the schema for the cpu has
precedence (i.e. we write the MSR with a CLOSID once, and then it
never changes).

Some of this is confusing because people will very likely also use
cpu affinity to control where their processes run. But affinity is
orthogonal to rdtgroup membership.

I think what we have allows you to so all the things we talked about.
But if we are missing a case, or if things can be simplified while
still retaining the same functionality then lets discuss that.
Otherwise we can revise the documentation to explain all this better.

> 
> > +Initial value is all zeros which means there is no CPU bound to the schemas
> > +in the root directory and tasks use the schemas.
> 
> As I said above this is backwards.

> > +If one resource is disabled, its line is not shown in schemas file.
> 
> That means:	  
> 
>      Resources which are not described in a schemata file are disabled for
>      that particular partition.
> 
> Right?
> 
> Now that raises the question how this is supposed to work. Let's assume that
> we have a partition 'foo' and thread X is in the tasks file of that
> partition. The schema of that partition contains only an L2 entry. What's the
> L3 association for thread X? Nothing at all?

Resources are either enabled or disabled globally. Each schema file
must provide details for every enabled resource. So if we are on a
processor that supports both L2 and L3, we will normally have schema
files that specify both.  We could boot with the "disable_cat_l2"
kernel command line option and then every schema file would just
specify L3 (and the MSRs for L2 would all be set to all-ones so that
everyone had full access to the L2 on each core).

> > +User can create a sub-directory under the root directory by "mkdir" command.
> > +User can remove the sub-directory by "rmdir" command.
> 
> User? Any user?

Well if someone did:
 # chmod 777 /sys/fs/rscctrl
then any user could make directories.  That would be inadvisable.
You could use 775 and let a trusted group have control so that you
didn't require root access to modify things.

Should we say "system administrator" rather than "user"?

-Tony

[toc] | [prev] | [next] | [standalone]


#1443106 — Re: [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface

FromThomas Gleixner <tglx@linutronix.de>
Date2016-07-14 09:00 +0200
SubjectRe: [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface
Message-ID<rUIMN-64Y-1@gated-at.bofh.it>
In reply to#1442650
On Wed, 13 Jul 2016, Luck, Tony wrote:
> On Wed, Jul 13, 2016 at 02:47:30PM +0200, Thomas Gleixner wrote:
> > On Tue, 12 Jul 2016, Fenghua Yu wrote:
> > > +3. Hierarchy in rscctrl
> > > +=======================
> > 
> > What means rscctrl?
> > 
> > You were not able to find a more cryptic acronym?
> 
> rscctrl == resource control
> 
> Intel marketing would (probably) like us to use:
> 
>    /sys/fs/Intel(R) Resource Director Technology(TM)/
> 
> Happy to take suggestions for something in between those
> extremes :-)

I'd suggest "resctrl" and the abbreviation dictionaries tell me that the most
common ones for resource are: R, RESORC, RES

> > > +Any tasks scheduled on the cpus will use the schemas. User can set
> > > +both "cpus" and "tasks" to share the same schema in one directory. But when
> > > +a CPU is bound to a schema, a task running on the CPU uses this schema and
> > > +kernel will ignore scheam set up for the task in "tasks".
> > 
> > This does not make any sense. 
> > 
> > When a task is bound to a schema then this should have preference over the
> > schema which is associated to the CPU. The CPU association is meant for tasks
> > which are not bound to a particular partition/schema.
> > 
> > So the initial setup should be:
> > 
> >    - All CPUs are associated to the root resource partition
> > 
> >    - No thread is associated to a particular resource partition
> > 
> > When a thread is added to a 'tasks' file of a partition then this partition
> > takes preference. If it's removed, i.e. the association to a partition is
> > undone, then the CPU association is used.
> > 
> > I have no idea why you think that all threads should be in a tasks file by
> > default. Associating CPUs in the first place makes a lot more sense as it
> > represents the topology of the system nicely.
> 
> If we did it that way, it would be harder to change the default
> resources.  E.g. now we start with all processes in the root
> rdtgroup.  We can change the schema for the root group and restrict
> them to, say, 60% of L3 cache on one (or all) sockets - giving us
> 40% of cache to give out to one or more groups.

I tend to disagree.

If you start up with all resources assigned to all CPUs and all tasks are set
to use the CPU default, then you still can restrict the root CPU defaults to
60% L3 which gives you 40% of cache to hand out.

What's hard about this?

Now you can start to create new partitions and either assign CPU or tasks to
them.

As a side effect that avoids the whole 'find all tasks' on mount machinery
simply because the CPU defaults do not change at all.
 
> So what we've implemented (and perhaps need to explain better here)
> is that every thread always belongs to one (and only one) rdtgroup.
> It will use the resources described in that group whereever it runs,
> except in the case where we have designated some cpus as special snowflakes.

I don't think that case as special snowflakes. Due to the very limited number
of cosids the CPU association is going to be a very useful tool.

> When a cpu is assigned to an rdtgroup the schema for the cpu has
> precedence (i.e. we write the MSR with a CLOSID once, and then it
> never changes).
> 
> Some of this is confusing because people will very likely also use
> cpu affinity to control where their processes run. But affinity is
> orthogonal to rdtgroup membership.

Right. It's confusing and what's even more confusing is that you have no way
to figure out what a particular task is actually using. With the 'use CPU
defaults, if not assigned to a partition' scheme you can very easy figure out
what a task is using because its either in a partition task list or not.

> I think what we have allows you to so all the things we talked about.
> But if we are missing a case, or if things can be simplified while
> still retaining the same functionality then lets discuss that.

It covers almost everything except the case I outlined before:

   Isolated CPU	 	    Important Task runs on isolated CPU
   5% exclusive cache	    10% exclusive cache

That's impossible with your scheme, but it's something which matters. You want
to make sure that the system services on that isolated CPU stay cache hot
without hurting the cache locality of your isolated task.

> Otherwise we can revise the documentation to explain all this better.

That needs to be done in any case. The existing one does not really qualify as
proper documentation. It's closer to a fairy tale :)

I really have to ask why you did not take the time and include all the
information you gave now into that documentation file in the first place.

> > > +Initial value is all zeros which means there is no CPU bound to the schemas
> > > +in the root directory and tasks use the schemas.
> > 
> > As I said above this is backwards.
> 
> > > +If one resource is disabled, its line is not shown in schemas file.
> > 
> > That means:	  
> > 
> >      Resources which are not described in a schemata file are disabled for
> >      that particular partition.
> > 
> > Right?
> > 
> > Now that raises the question how this is supposed to work. Let's assume that
> > we have a partition 'foo' and thread X is in the tasks file of that
> > partition. The schema of that partition contains only an L2 entry. What's the
> > L3 association for thread X? Nothing at all?
> 
> Resources are either enabled or disabled globally. Each schema file
> must provide details for every enabled resource. So if we are on a
> processor that supports both L2 and L3, we will normally have schema
> files that specify both.
> We could boot with the "disable_cat_l2"
> kernel command line option and then every schema file would just
> specify L3 (and the MSRs for L2 would all be set to all-ones so that
> everyone had full access to the L2 on each core).

So the above should read:

   Each schema file must provide configuration for all resource controls which
   are enabled in the system.

Right?
 
> > > +User can create a sub-directory under the root directory by "mkdir" command.
> > > +User can remove the sub-directory by "rmdir" command.
> > 
> > User? Any user?
> 
> Well if someone did:
>  # chmod 777 /sys/fs/rscctrl
> then any user could make directories.  That would be inadvisable.
> You could use 775 and let a trusted group have control so that you
> didn't require root access to modify things.
> 
> Should we say "system administrator" rather than "user"?

Yes. Because the default should be 755 which is the obvious choice for all
root/admin controlled things. If root decides to change it to 777 then it's
not the kernels problem. But documentation should clearly say: It's a root
controlled resource.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


#1443615 — Re: [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface

From"Luck, Tony" <tony.luck@intel.com>
Date2016-07-14 19:20 +0200
SubjectRe: [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface
Message-ID<rUSsO-45I-21@gated-at.bofh.it>
In reply to#1443106
On Thu, Jul 14, 2016 at 08:53:17AM +0200, Thomas Gleixner wrote:
> > Happy to take suggestions for something in between those
> > extremes :-)
> 
> I'd suggest "resctrl" and the abbreviation dictionaries tell me that the most
> common ones for resource are: R, RESORC, RES

OK. "resctrl" it is.

> As a side effect that avoids the whole 'find all tasks' on mount machinery
> simply because the CPU defaults do not change at all.

That's a very good side effect.

It just means that the "tasks" file in the root of the hierachy will
need different read/write functions from those in sub-directories.

read: scan all tasks, print pid for ones with task->rdtgroup == NULL

write: remove task from the rdtgroup list that it was on; set task->rdtgroup = NULL;

> It covers almost everything except the case I outlined before:
> 
>    Isolated CPU	 	    Important Task runs on isolated CPU
>    5% exclusive cache	    10% exclusive cache
> 
> That's impossible with your scheme, but it's something which matters. You want
> to make sure that the system services on that isolated CPU stay cache hot
> without hurting the cache locality of your isolated task.

So the core part of __intel_rdt_sched_in() will look like:

	/*
	 * Precedence rules:
	 * Processes assigned to an rdtgroup use that group
	 * wherever they run. If they don't have an rdtgroup
	 * we see if the current cpu has one and use it.
	 * If no specific rdtgroup was provided, we use the
	 * root_rdtgroup
	 */
	rdtgrp = current->rdtgroup;
	if (!rdtgrp) {
		rdtgrp = per_cpu(cpu_rdtgroup, cpu);
		if (!rdtgrp)
			rdtgrp = root_rdtgroup;
	}

> > Otherwise we can revise the documentation to explain all this better.
> 
> That needs to be done in any case. The existing one does not really qualify as
> proper documentation. It's closer to a fairy tale :)

Yes. We will re-write.

-Tony

[toc] | [prev] | [next] | [standalone]


#1446393 — Re: [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface

FromThomas Gleixner <tglx@linutronix.de>
Date2016-07-19 14:40 +0200
SubjectRe: [PATCH 13/32] Documentation, x86: Documentation for Intel resource allocation user interface
Message-ID<rWCtA-40N-17@gated-at.bofh.it>
In reply to#1443615
On Thu, 14 Jul 2016, Luck, Tony wrote:
> So the core part of __intel_rdt_sched_in() will look like:
> 
> 	/*
> 	 * Precedence rules:
> 	 * Processes assigned to an rdtgroup use that group
> 	 * wherever they run. If they don't have an rdtgroup
> 	 * we see if the current cpu has one and use it.
> 	 * If no specific rdtgroup was provided, we use the
> 	 * root_rdtgroup
> 	 */
> 	rdtgrp = current->rdtgroup;
> 	if (!rdtgrp) {
> 		rdtgrp = per_cpu(cpu_rdtgroup, cpu);
> 		if (!rdtgrp)
> 			rdtgrp = root_rdtgroup;
> 	}

That can be done simpler. The default cpu_rdtgroup should be root_rdtgroup. So
you spare one conditional.

Thanks

	tglx

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web