Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1501214 > unrolled thread
| Started by | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| First post | 2016-10-15 01:20 +0200 |
| Last post | 2016-10-17 13:10 +0200 |
| Articles | 20 on this page of 49 — 5 participants |
Back to article view | Back to linux.kernel
[PATCH v4 00/18] Intel Cache Allocation Technology "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
[PATCH v4 02/18] cacheinfo: Introduce cache id "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 02/18] cacheinfo: Introduce cache id Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 12:40 +0200
[PATCH v4 16/18] x86/intel_rdt: Add schemata file "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 16/18] x86/intel_rdt: Add schemata file Thomas Gleixner <tglx@linutronix.de> - 2016-10-18 00:40 +0200
[PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 23:20 +0200
Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system "Luck, Tony" <tony.luck@intel.com> - 2016-10-18 00:00 +0200
Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system Thomas Gleixner <tglx@linutronix.de> - 2016-10-18 01:00 +0200
RE: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system Thomas Gleixner <tglx@linutronix.de> - 2016-10-18 01:10 +0200
RE: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system "Luck, Tony" <tony.luck@intel.com> - 2016-10-18 01:20 +0200
RE: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system Thomas Gleixner <tglx@linutronix.de> - 2016-10-18 01:30 +0200
RE: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system "Luck, Tony" <tony.luck@intel.com> - 2016-10-18 01:10 +0200
Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system Fenghua Yu <fenghua.yu@intel.com> - 2016-10-18 00:20 +0200
Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system Thomas Gleixner <tglx@linutronix.de> - 2016-10-18 01:30 +0200
Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system "Luck, Tony" <tony.luck@intel.com> - 2016-10-18 01:40 +0200
Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system Fenghua Yu <fenghua.yu@intel.com> - 2016-10-18 02:00 +0200
Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system Thomas Gleixner <tglx@linutronix.de> - 2016-10-18 12:50 +0200
[PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 15:50 +0200
Re: [PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID Fenghua Yu <fenghua.yu@intel.com> - 2016-10-17 17:10 +0200
Re: [PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID "Luck, Tony" <tony.luck@intel.com> - 2016-10-17 18:40 +0200
RE: [PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID "Yu, Fenghua" <fenghua.yu@intel.com> - 2016-10-17 18:50 +0200
Re: [PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID "Luck, Tony" <tony.luck@intel.com> - 2016-10-17 22:30 +0200
Re: [PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 19:10 +0200
Re: [PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 19:10 +0200
Re: [PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 19:10 +0200
Re: [PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID Fenghua Yu <fenghua.yu@intel.com> - 2016-10-17 20:20 +0200
[PATCH v4 01/18] Documentation, ABI: Add a document entry for cache id "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 01/18] Documentation, ABI: Add a document entry for cache id Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 12:40 +0200
[PATCH v4 18/18] MAINTAINERS: Add maintainer for Intel RDT resource allocation "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
[PATCH v4 04/18] x86/intel_rdt: Feature discovery "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
[PATCH v4 11/18] x86/intel_rdt: Add basic resctrl filesystem support "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 11/18] x86/intel_rdt: Add basic resctrl filesystem support Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 21:40 +0200
[PATCH v4 05/18] Documentation, x86: Documentation for Intel resource allocation user interface "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
[PATCH v4 17/18] x86/intel_rdt: Add scheduler hook "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
[PATCH v4 14/18] x86/intel_rdt: Add cpus file "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 14/18] x86/intel_rdt: Add cpus file Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 23:40 +0200
[PATCH v4 12/18] x86/intel_rdt: Add "info" files to resctrl file system "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 12/18] x86/intel_rdt: Add "info" files to resctrl file system Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 21:50 +0200
[PATCH v4 15/18] x86/intel_rdt: Add tasks files "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 15/18] x86/intel_rdt: Add tasks files Thomas Gleixner <tglx@linutronix.de> - 2016-10-18 00:10 +0200
Re: [PATCH v4 15/18] x86/intel_rdt: Add tasks files "Luck, Tony" <tony.luck@intel.com> - 2016-10-18 00:20 +0200
[PATCH v4 10/18] x86/intel_rdt: Build structures for each resource based on cache topology "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 10/18] x86/intel_rdt: Build structures for each resource based on cache topology Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 16:50 +0200
[PATCH v4 03/18] x86, intel_cacheinfo: Enable cache id in x86 "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 03/18] x86, intel_cacheinfo: Enable cache id in x86 Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 13:00 +0200
[PATCH v4 07/18] x86/intel_rdt: Add Haswell feature discovery "Fenghua Yu" <fenghua.yu@intel.com> - 2016-10-15 01:20 +0200
Re: [PATCH v4 07/18] x86/intel_rdt: Add Haswell feature discovery Thomas Gleixner <tglx@linutronix.de> - 2016-10-17 13:10 +0200
Page 1 of 3 [1] 2 3 Next page →
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-10-15 01:20 +0200 |
| Subject | [PATCH v4 00/18] Intel Cache Allocation Technology |
| Message-ID | <ssjVD-7Za-5@gated-at.bofh.it> |
From: Fenghua Yu <fenghua.yu@intel.com>
Change log in v4:
* Changed CONFIG_INTEL_RDT to CONFIG_RDT_A. Plain "RDT" refers to all
resource features, both monitoring (CQM, MBM, ...) and control (CAT L3,
L3/CDP, L2, ...). Adding the "_A" matches with the feature bit name
for all the control features "X86_FEATURE_RDT_A".
* Nilay: Cleaned up duplicate declarations (and ones that appear in the
wrong patch)
* Boris: Add comment on which specific Haswell models support CAT L3
* Boris: Use boot_cpu_data and boot_cpu_has()
* Don't call smp_call_function_single() from hot add notifier. We are already
on the right cpu.
* Thomas (from earlier postings) check return value from rdtgroup_init() and
cleanup if it failed
* Boris: Be more descriptive than "cache allocation" (we now say "RDT cache
allocation")
* Nilay: Move some code out of mount() [mostly into closid_alloc()]
* Nilay: Drop wrapper functions rdtgroup_{alloc,free}, just call kzalloc/kfree
* smp_call_function_many() skips current cpu (to Tony's surprise). Make sure
we make the call locally if current cpu is included in the mask.
* Also be preempt aware around calls to smp_call_function_many()
* Nilay: Add extra parens in:
if ((attr == &dev_attr_id.attr) && (this_leaf->attributes & CACHE_ID))
* Added a comment for closid_alloc that current allocation is global and will
do per resource domain allocation later.
* Added CAT L2 support and changed commit message in patch #8.
The patches are in the same order as V3 and have small commit messages changes
in patch #7 and #8.
0001-Documentation-ABI-Add-a-document-entry-for-cache-id.patch
0002-cacheinfo-Introduce-cache-id.patch
0003-x86-intel_cacheinfo-Enable-cache-id-in-x86.patch
These three define an "id" for each cache ... we need a "name"
for a cache so we can say what restrictions to apply to each
cache in the system. All you will see at this point is an
extra "id" file in each /sys/devices/system/cpu/cpu*/cache/index*/
directory.
0004-x86-intel_rdt-Feature-discovery.patch
Look at CPUID for the features related to cache allocation.
At this point /proc/cpuinfo shows extra flags for the features
found on your system.
0005-Documentation-x86-Documentation-for-Intel-resource-a.patch
Documentation patch could be anywhere in this sequence. We
put in early so you can read it to see how to use the
interface.
0006-x86-intel_rdt-Add-CONFIG-Makefile-and-basic-initiali.patch
Add CONFIG_INTEL_RDT (default "n" ... you'll have to set
it to have this, and all the following patches do anything).
Template driver here just checks for features and spams
the console with one line for each.
0007-x86-intel_rdt-Add-Haswell-feature-discovery.patch
There are some Haswell systems that support cache allocation,
but they were made before the CPUID bits were fully defined.
So we check by probing the CBM base MSR to see if CLOSID
bits stick. Unless you have one of these Haswells, you won't
see any difference here.
0008-x86-intel_rdt-Pick-up-L3-L2-RDT-parameters-from-CPUID.patch
This is all new code, not seen in the previous versions of this
patch series. L3 and L2 cache allocations are just the first of
several resource control features. Define rdt_resource structure
that contains all the useful things we need to know about a
resource. Pick up the parameters for the resource from CPUID.
The console spam strings change format here.
0009-x86-cqm-Move-PQR_ASSOC-management-code-into-generic-.patch
The PQR_ASSOC MSR has a field for the CLOSID (which we need
define which allocation rules are in effect). But it also
contains the RMID (used by CQM and MBM perf monitoring).
The perf code got here first, but defined structures that
make it easy for the two systems to co-exist without stomping
on each other. This patch moves the relevant parts into a
common header file and changes the scope from "static" to
global so we can access them. No visible change.
0010-x86-intel_rdt-Build-structures-for-each-resource-bas.patch
For each enabled resource, we build a list of "rdt_domains" based
on hotplug cpu notifications. Since we only have L3 at this point,
this is just a list of L3 caches (named by the "id" established
in the first three patches). As each cache is found we initialize
the array of CBMs (cache bit masks). No visible change here.
0011-x86-intel_rdt-Add-basic-resctrl-filesystem-support.patch
Our interface is a kernfs backed file system. Establish the
mount point, and provide mount/unmount functionality.
At this point "/sys/fs/resctrl" appears. You can mount and
unmount the resctrl file system (if your system supports
code/data prioritization, you can use the "cdp" mount option).
The file system is empty and doesn't allow creation of any
files or subdirectories.
0012-x86-intel_rdt-Add-info-files-to-resctrl-file-system.patch
Parameters for each resource are buried in CPUID leaf 0x10.
This isn't very user friendly for scripts and applications
that want to configure resource allocation. Create an
"info" directory, with a subdirectory for each resource
containing a couple of useful parameters. Visible change:
$ ls -l /sys/fs/resctrl/info/L3
total 0
-r--r--r-- 1 root root 0 Oct 7 11:20 cbm_val
-r--r--r-- 1 root root 0 Oct 7 11:20 num_closid
0013-x86-intel_rdt-Add-mkdir-to-resctrl-file-system.patch
Each resource group is represented by a directory in the
resctrl file system. The root directory is the default group.
Use "mkdir" to create new groups and "rmdir" to remove them.
The maximum number of groups is defined by the effective
number of CLOSIDs.
Visible change: If you have CDP (and enable with the "cdp"
mount option) you will find that you can only create half
as many groups as without (e.g. 8 vs. 16 on Broadwell, but
the default group uses one ... so actually 7, 15).
0014-x86-intel_rdt-Add-cpus-file.patch
One of the control mechanisms for a resource group is the
logical CPU. Initially all CPUs are assigned to the default
group. They can be reassigned to other groups by writing
a cpumask to the "cpus" file. See the documentation for what
this means.
Visible change: "cpus" file in the root, and automatically
in each created subdirectory. You can "echo" masks to these
files and watch as CPUs added to one group are removed from
whatever group they previously belonged to. Removing a directory
will give all CPUs owned by it back to the default (root)
group.
0015-x86-intel_rdt-Add-tasks-files.patch
Tasks can be assigned to resource groups by writing their PID
to a "tasks" file (which removes the task from its previous
group). Forked/cloned tasks inherit the group from their
parent. You cannot remove a group (directory) that has any
tasks assigned.
Visible change: "tasks" files appear. E.g. (we see two tasks
in the group, our shell, and the "cat" that it spawned).
# echo $$ > p0/tasks; cat p0/tasks
268890
268914
0016-x86-intel_rdt-Add-schemata-file.patch
The "schemata" file in each group/directory defines what
access tasks controlled by this resource are permitted.
One line per resource type. Fields for each instance of
the resource. You redefine the access by wrting to the
file in the same format.
Visible change: "schemata" file which starts out with maximum
allowed resources. E.g.
$ cat schemata
L3:0=fffff;1=fffff
Now restrict this group to just 20% of L3 on first cache, but
allow 50% on the second
# echo L3:0=f;1=3ff > schemata
0017-x86-intel_rdt-Add-scheduler-hook.patch
When context switching we check if we are changing resource
groups for the new process, and update the PQR_ASSOC MSR with
the new CLOSID if needed.
Visble change: Everything should be working now. Tasks run with
the permitted access to L3 cache.
0018-MAINTAINERS-Add-maintainer-for-Intel-RDT-resource-al.patch
New files ... need a maintainer. Fenghua has the job.
Fenghua Yu (15):
Documentation, ABI: Add a document entry for cache id
cacheinfo: Introduce cache id
x86, intel_cacheinfo: Enable cache id in x86
x86/intel_rdt: Feature discovery
Documentation, x86: Documentation for Intel resource allocation user
interface
x86/intel_rdt: Add CONFIG, Makefile, and basic initialization
x86/intel_rdt: Add Haswell feature discovery
x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID
x86/cqm: Move PQR_ASSOC management code into generic code used by both
CQM and CAT
x86/intel_rdt: Add basic resctrl filesystem support
x86/intel_rdt: Add "info" files to resctrl file system
x86/intel_rdt: Add mkdir to resctrl file system
x86/intel_rdt: Add tasks files
x86/intel_rdt: Add scheduler hook
MAINTAINERS: Add maintainer for Intel RDT resource allocation
Tony Luck (3):
x86/intel_rdt: Build structures for each resource based on cache
topology
x86/intel_rdt: Add cpus file
x86/intel_rdt: Add schemata file
Documentation/ABI/testing/sysfs-devices-system-cpu | 16 +
Documentation/x86/intel_rdt_ui.txt | 162 ++++
MAINTAINERS | 8 +
arch/x86/Kconfig | 12 +
arch/x86/events/intel/cqm.c | 23 +-
arch/x86/include/asm/cpufeatures.h | 5 +
arch/x86/include/asm/intel_rdt.h | 206 +++++
arch/x86/include/asm/intel_rdt_common.h | 27 +
arch/x86/kernel/cpu/Makefile | 2 +
arch/x86/kernel/cpu/intel_cacheinfo.c | 20 +
arch/x86/kernel/cpu/intel_rdt.c | 304 +++++++
arch/x86/kernel/cpu/intel_rdt_rdtgroup.c | 973 +++++++++++++++++++++
arch/x86/kernel/cpu/intel_rdt_schemata.c | 266 ++++++
arch/x86/kernel/cpu/scattered.c | 3 +
arch/x86/kernel/process_32.c | 4 +
arch/x86/kernel/process_64.c | 4 +
drivers/base/cacheinfo.c | 5 +
include/linux/cacheinfo.h | 3 +
include/linux/sched.h | 3 +
include/uapi/linux/magic.h | 1 +
20 files changed, 2026 insertions(+), 21 deletions(-)
create mode 100644 Documentation/x86/intel_rdt_ui.txt
create mode 100644 arch/x86/include/asm/intel_rdt.h
create mode 100644 arch/x86/include/asm/intel_rdt_common.h
create mode 100644 arch/x86/kernel/cpu/intel_rdt.c
create mode 100644 arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
create mode 100644 arch/x86/kernel/cpu/intel_rdt_schemata.c
--
2.5.0
[toc] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-10-15 01:20 +0200 |
| Subject | [PATCH v4 02/18] cacheinfo: Introduce cache id |
| Message-ID | <ssjVD-7Za-9@gated-at.bofh.it> |
| In reply to | #1501214 |
From: Fenghua Yu <fenghua.yu@intel.com>
Cache management software needs a name for each instance of a cache of
a particular type.
The current cacheinfo structure does not provide any information about
the underlying hardware so there is no way to expose it.
Hardware with cache management features provides means (cpuid, enumeration
etc.) to retrieve the hardware id of a particular cache instance. Cache
instances which share hardware have the same hardware id.
Add an 'id' field to struct cacheinfo to store this information. Expose
this information under the /sys/devices/system/cpu/cpu*/cache/index*/
directory as well.
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Signed-off-by: Tony Luck <tony.luck@intel.com>
---
drivers/base/cacheinfo.c | 5 +++++
include/linux/cacheinfo.h | 3 +++
2 files changed, 8 insertions(+)
diff --git a/drivers/base/cacheinfo.c b/drivers/base/cacheinfo.c
index e9fd32e..00a9688 100644
--- a/drivers/base/cacheinfo.c
+++ b/drivers/base/cacheinfo.c
@@ -233,6 +233,7 @@ static ssize_t file_name##_show(struct device *dev, \
return sprintf(buf, "%u\n", this_leaf->object); \
}
+show_one(id, id);
show_one(level, level);
show_one(coherency_line_size, coherency_line_size);
show_one(number_of_sets, number_of_sets);
@@ -314,6 +315,7 @@ static ssize_t write_policy_show(struct device *dev,
return n;
}
+static DEVICE_ATTR_RO(id);
static DEVICE_ATTR_RO(level);
static DEVICE_ATTR_RO(type);
static DEVICE_ATTR_RO(coherency_line_size);
@@ -327,6 +329,7 @@ static DEVICE_ATTR_RO(shared_cpu_list);
static DEVICE_ATTR_RO(physical_line_partition);
static struct attribute *cache_default_attrs[] = {
+ &dev_attr_id.attr,
&dev_attr_type.attr,
&dev_attr_level.attr,
&dev_attr_shared_cpu_map.attr,
@@ -350,6 +353,8 @@ cache_default_attrs_is_visible(struct kobject *kobj,
const struct cpumask *mask = &this_leaf->shared_cpu_map;
umode_t mode = attr->mode;
+ if ((attr == &dev_attr_id.attr) && (this_leaf->attributes & CACHE_ID))
+ return mode;
if ((attr == &dev_attr_type.attr) && this_leaf->type)
return mode;
if ((attr == &dev_attr_level.attr) && this_leaf->level)
diff --git a/include/linux/cacheinfo.h b/include/linux/cacheinfo.h
index 2189935..0bcbb67 100644
--- a/include/linux/cacheinfo.h
+++ b/include/linux/cacheinfo.h
@@ -18,6 +18,7 @@ enum cache_type {
/**
* struct cacheinfo - represent a cache leaf node
+ * @id: This cache's id. It is unique among caches with the same (type, level).
* @type: type of the cache - data, inst or unified
* @level: represents the hierarchy in the multi-level cache
* @coherency_line_size: size of each cache line usually representing
@@ -44,6 +45,7 @@ enum cache_type {
* keeping, the remaining members form the core properties of the cache
*/
struct cacheinfo {
+ unsigned int id;
enum cache_type type;
unsigned int level;
unsigned int coherency_line_size;
@@ -61,6 +63,7 @@ struct cacheinfo {
#define CACHE_WRITE_ALLOCATE BIT(3)
#define CACHE_ALLOCATE_POLICY_MASK \
(CACHE_READ_ALLOCATE | CACHE_WRITE_ALLOCATE)
+#define CACHE_ID BIT(4)
struct device_node *of_node;
bool disable_sysfs;
--
2.5.0
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-10-17 12:40 +0200 |
| Subject | Re: [PATCH v4 02/18] cacheinfo: Introduce cache id |
| Message-ID | <stduN-2hG-1@gated-at.bofh.it> |
| In reply to | #1501215 |
On Fri, 14 Oct 2016, Fenghua Yu wrote: > From: Fenghua Yu <fenghua.yu@intel.com> > > Cache management software needs a name for each instance of a cache of > a particular type. s/name/id/ Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-10-15 01:20 +0200 |
| Subject | [PATCH v4 16/18] x86/intel_rdt: Add schemata file |
| Message-ID | <ssjVD-7Za-13@gated-at.bofh.it> |
| In reply to | #1501214 |
From: Tony Luck <tony.luck@intel.com>
Last of the per resource group files. Also mode 0644. This one shows
the resources available to the group. Syntax depends on whether the
"cdp" mount option was given. With code/data prioritization disabled
it is simply a list of masks for each cache domain. Initial value
allows access to all of the L3 cache on all domains. E.g. on a 2 socket
Broadwell:
L3:0=fffff;1=fffff
With CDP enabled, separate masks for data and instructions are provided:
L3:0=fffff,fffff;1=fffff,fffff
Signed-off-by: Tony Luck <tony.luck@intel.com>
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
---
arch/x86/include/asm/intel_rdt.h | 6 +
arch/x86/kernel/cpu/Makefile | 2 +-
arch/x86/kernel/cpu/intel_rdt_rdtgroup.c | 7 +
arch/x86/kernel/cpu/intel_rdt_schemata.c | 266 +++++++++++++++++++++++++++++++
4 files changed, 280 insertions(+), 1 deletion(-)
create mode 100644 arch/x86/kernel/cpu/intel_rdt_schemata.c
diff --git a/arch/x86/include/asm/intel_rdt.h b/arch/x86/include/asm/intel_rdt.h
index 6ab31ba..8435ec6 100644
--- a/arch/x86/include/asm/intel_rdt.h
+++ b/arch/x86/include/asm/intel_rdt.h
@@ -70,6 +70,7 @@ struct rftype {
* @cdp_capable: Code/Data Prioritization available
* @cdp_enabled: Code/Data Prioritization enabled
* @tmp_cbms: Scratch space when updating schemata
+ * @num_cbms: Number of CBMs in tmp_cbms
* @cache_level: Which cache level defines scope of this domain
*/
struct rdt_resource {
@@ -86,6 +87,7 @@ struct rdt_resource {
bool cdp_capable;
bool cdp_enabled;
u32 *tmp_cbms;
+ int num_cbms;
int cache_level;
};
@@ -156,4 +158,8 @@ DECLARE_PER_CPU_READ_MOSTLY(int, cpu_closid);
void rdt_cbm_update(void *arg);
struct rdtgroup *rdtgroup_kn_lock_live(struct kernfs_node *kn);
void rdtgroup_kn_unlock(struct kernfs_node *kn);
+ssize_t rdtgroup_schemata_write(struct kernfs_open_file *of,
+ char *buf, size_t nbytes, loff_t off);
+int rdtgroup_schemata_show(struct kernfs_open_file *of,
+ struct seq_file *s, void *v);
#endif /* _ASM_X86_INTEL_RDT_H */
diff --git a/arch/x86/kernel/cpu/Makefile b/arch/x86/kernel/cpu/Makefile
index b4334e8..c9f8c81 100644
--- a/arch/x86/kernel/cpu/Makefile
+++ b/arch/x86/kernel/cpu/Makefile
@@ -34,7 +34,7 @@ obj-$(CONFIG_CPU_SUP_CENTAUR) += centaur.o
obj-$(CONFIG_CPU_SUP_TRANSMETA_32) += transmeta.o
obj-$(CONFIG_CPU_SUP_UMC_32) += umc.o
-obj-$(CONFIG_INTEL_RDT_A) += intel_rdt.o intel_rdt_rdtgroup.o
+obj-$(CONFIG_INTEL_RDT_A) += intel_rdt.o intel_rdt_rdtgroup.o intel_rdt_schemata.o
obj-$(CONFIG_X86_MCE) += mcheck/
obj-$(CONFIG_MTRR) += mtrr/
diff --git a/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c b/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
index bdbe2d1..6787fd8 100644
--- a/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
+++ b/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
@@ -442,6 +442,13 @@ static struct rftype rdtgroup_base_files[] = {
.seq_show = rdtgroup_tasks_show,
},
{
+ .name = "schemata",
+ .mode = 0644,
+ .kf_ops = &rdtgroup_kf_single_ops,
+ .write = rdtgroup_schemata_write,
+ .seq_show = rdtgroup_schemata_show,
+ },
+ {
/* NULL terminated */
}
};
diff --git a/arch/x86/kernel/cpu/intel_rdt_schemata.c b/arch/x86/kernel/cpu/intel_rdt_schemata.c
new file mode 100644
index 0000000..ead4f43
--- /dev/null
+++ b/arch/x86/kernel/cpu/intel_rdt_schemata.c
@@ -0,0 +1,266 @@
+/*
+ * Resource Director Technology(RDT)
+ * - Cache Allocation code.
+ *
+ * Copyright (C) 2016 Intel Corporation
+ *
+ * Authors:
+ * Fenghua Yu <fenghua.yu@intel.com>
+ * Tony Luck <tony.luck@intel.com>
+ *
+ * This program is free software; you can redistribute it and/or modify it
+ * under the terms and conditions of the GNU General Public License,
+ * version 2, as published by the Free Software Foundation.
+ *
+ * This program is distributed in the hope it will be useful, but WITHOUT
+ * ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or
+ * FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for
+ * more details.
+ *
+ * More information about RDT be found in the Intel (R) x86 Architecture
+ * Software Developer Manual June 2016, volume 3, section 17.17.
+ */
+
+#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt
+
+#include <linux/kernfs.h>
+#include <linux/seq_file.h>
+#include <linux/slab.h>
+#include <asm/intel_rdt.h>
+
+/*
+ * Check whether a cache bit mask is valid. The SDM says:
+ * Please note that all (and only) contiguous '1' combinations
+ * are allowed (e.g. FFFFH, 0FF0H, 003CH, etc.).
+ * Additionally Haswell requires at least two bits set.
+ */
+static bool cbm_validate(unsigned long var, struct rdt_resource *r)
+{
+ unsigned long first_bit, zero_bit;
+
+ if (var == 0 || var > r->max_cbm)
+ return false;
+
+ first_bit = find_first_bit(&var, r->cbm_len);
+ zero_bit = find_next_zero_bit(&var, r->cbm_len, first_bit);
+
+ if (find_next_bit(&var, r->cbm_len, zero_bit) < r->cbm_len)
+ return false;
+
+ if ((zero_bit - first_bit) < r->min_cbm_bits)
+ return false;
+ return true;
+}
+
+/*
+ * Read one cache bit mask (hex). Check that it is valid for the current
+ * resource type.
+ */
+static int parse_cbm_token(char *tok, struct rdt_resource *r)
+{
+ unsigned long data;
+ int ret;
+
+ ret = kstrtoul(tok, 16, &data);
+ if (ret)
+ return ret;
+ if (!cbm_validate(data, r))
+ return -EINVAL;
+ r->tmp_cbms[r->num_cbms++] = data;
+ return 0;
+}
+
+/*
+ * If code/data prioritization is enabled for this resource we need
+ * two bit masks separated by a ",". Otherwise a single bit mask.
+ */
+static int parse_cbm(char *buf, struct rdt_resource *r)
+{
+ char *cbm1 = buf;
+ int ret;
+
+ if (r->cdp_enabled)
+ cbm1 = strsep(&buf, ",");
+ if (!cbm1 || !buf)
+ return 1;
+ ret = parse_cbm_token(cbm1, r);
+ if (ret)
+ return ret;
+ if (r->cdp_enabled)
+ return parse_cbm_token(buf, r);
+ return 0;
+}
+
+/*
+ * For each domain in this resource we expect to find a series of:
+ * id=mask[,mask]
+ * separated by ";". The "id" is in decimal, and must appear in the
+ * right order.
+ */
+static int parse_line(char *line, struct rdt_resource *r)
+{
+ struct list_head *l;
+ struct rdt_domain *d;
+ char *dom = NULL, *id;
+ unsigned long dom_id;
+
+ list_for_each(l, &r->domains) {
+ d = list_entry(l, struct rdt_domain, list);
+ dom = strsep(&line, ";");
+ if (!dom)
+ return -EINVAL;
+ id = strsep(&dom, "=");
+ if (kstrtoul(id, 10, &dom_id) || dom_id != d->id)
+ return -EINVAL;
+ if (parse_cbm(dom, r))
+ return -EINVAL;
+ }
+
+ /* Any garbage at the end of the line? */
+ if (line && line[0])
+ return -EINVAL;
+ return 0;
+}
+
+static void update_domains(struct rdt_resource *r, int closid)
+{
+ int cpu, idx = 0;
+ struct list_head *l;
+ struct rdt_domain *d;
+ struct msr_param msr_param;
+ struct cpumask cpu_mask;
+
+ cpumask_clear(&cpu_mask);
+ msr_param.low = closid << r->cdp_enabled;
+ msr_param.high = msr_param.low + 1 + r->cdp_enabled;
+ msr_param.res = r;
+
+ list_for_each(l, &r->domains) {
+ d = list_entry(l, struct rdt_domain, list);
+ cpumask_set_cpu(cpumask_any(&d->cpu_mask), &cpu_mask);
+ d->cbm[msr_param.low] = r->tmp_cbms[idx++];
+ if (r->cdp_enabled)
+ d->cbm[msr_param.low + 1] = r->tmp_cbms[idx++];
+ }
+ cpu = get_cpu();
+ /* Update CBM on this cpu if it's in cpu_mask. */
+ if (cpumask_test_cpu(cpu, &cpu_mask))
+ rdt_cbm_update(&msr_param);
+ /* Update CBM on other cpus. */
+ smp_call_function_many(&cpu_mask, rdt_cbm_update, &msr_param, 1);
+ put_cpu();
+}
+
+ssize_t rdtgroup_schemata_write(struct kernfs_open_file *of,
+ char *buf, size_t nbytes, loff_t off)
+{
+ char *tok, *resname;
+ struct rdtgroup *rdtgrp;
+ struct rdt_resource *r;
+ int closid, ret = 0;
+ u32 *l3_cbms = NULL;
+
+ /* Legal input requires a trailing newline */
+ if (nbytes == 0 || buf[nbytes - 1] != '\n')
+ return -EINVAL;
+ buf[nbytes - 1] = '\0';
+
+ rdtgrp = rdtgroup_kn_lock_live(of->kn);
+ if (!rdtgrp) {
+ rdtgroup_kn_unlock(of->kn);
+ return -ENOENT;
+ }
+
+ closid = rdtgrp->closid;
+
+ /* get scratch space to save all the masks while we validate input */
+ for_each_rdt_resource(r) {
+ r->tmp_cbms = kcalloc(r->num_domains << r->cdp_enabled,
+ sizeof(*l3_cbms), GFP_KERNEL);
+ if (!r->tmp_cbms) {
+ ret = -ENOMEM;
+ goto fail;
+ }
+ r->num_cbms = 0;
+ }
+
+ while ((tok = strsep(&buf, "\n")) != NULL) {
+ resname = strsep(&tok, ":");
+ if (!tok) {
+ ret = -EINVAL;
+ goto fail;
+ }
+ for_each_rdt_resource(r) {
+ if (!strcmp(resname, r->name) &&
+ closid < r->num_closid) {
+ ret = parse_line(tok, r);
+ if (ret)
+ goto fail;
+ break;
+ }
+ }
+ if (!r->name) {
+ ret = -EINVAL;
+ goto fail;
+ }
+ }
+
+ /* Did the parser find all the masks we need? */
+ for_each_rdt_resource(r) {
+ if (r->num_cbms != r->num_domains << r->cdp_enabled) {
+ ret = -EINVAL;
+ goto fail;
+ }
+ }
+
+ for_each_rdt_resource(r)
+ update_domains(r, closid);
+
+fail:
+ rdtgroup_kn_unlock(of->kn);
+ for_each_rdt_resource(r) {
+ kfree(r->tmp_cbms);
+ r->tmp_cbms = NULL;
+ }
+ return ret ?: nbytes;
+}
+
+static void show_doms(struct seq_file *s, struct rdt_resource *r, int closid)
+{
+ struct list_head *l;
+ struct rdt_domain *dom;
+ int idx = closid << r->cdp_enabled;
+ bool sep = false;
+
+ seq_printf(s, "%s:", r->name);
+ list_for_each(l, &r->domains) {
+ dom = list_entry(l, struct rdt_domain, list);
+ if (sep)
+ seq_puts(s, ";");
+ seq_printf(s, "%d=%x", dom->id, dom->cbm[idx]);
+ if (r->cdp_enabled)
+ seq_printf(s, ",%x", dom->cbm[idx + 1]);
+ sep = true;
+ }
+ seq_puts(s, "\n");
+}
+
+int rdtgroup_schemata_show(struct kernfs_open_file *of,
+ struct seq_file *s, void *v)
+{
+ struct rdtgroup *rdtgrp;
+ struct rdt_resource *r;
+ int closid, ret = 0;
+
+ rdtgrp = rdtgroup_kn_lock_live(of->kn);
+ if (rdtgrp) {
+ closid = rdtgrp->closid;
+ for_each_rdt_resource(r)
+ if (closid < r->num_closid)
+ show_doms(s, r, closid);
+ } else {
+ ret = -ENOENT;
+ }
+ rdtgroup_kn_unlock(of->kn);
+ return ret;
+}
--
2.5.0
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-10-18 00:40 +0200 |
| Subject | Re: [PATCH v4 16/18] x86/intel_rdt: Add schemata file |
| Message-ID | <stoJA-1Co-19@gated-at.bofh.it> |
| In reply to | #1501216 |
On Fri, 14 Oct 2016, Fenghua Yu wrote:
> +static void update_domains(struct rdt_resource *r, int closid)
> +{
> + int cpu, idx = 0;
> + struct list_head *l;
> + struct rdt_domain *d;
> + struct msr_param msr_param;
> + struct cpumask cpu_mask;
Again. No cpumasks on stack.
> +ssize_t rdtgroup_schemata_write(struct kernfs_open_file *of,
> + char *buf, size_t nbytes, loff_t off)
> +{
> + char *tok, *resname;
> + struct rdtgroup *rdtgrp;
> + struct rdt_resource *r;
> + int closid, ret = 0;
> + u32 *l3_cbms = NULL;
> +
> + /* Legal input requires a trailing newline */
s/Legal/Valid/ please. There is no law which enforces this.
> + if (nbytes == 0 || buf[nbytes - 1] != '\n')
> + return -EINVAL;
> + buf[nbytes - 1] = '\0';
> +
> + rdtgrp = rdtgroup_kn_lock_live(of->kn);
> + if (!rdtgrp) {
> + rdtgroup_kn_unlock(of->kn);
> + return -ENOENT;
> + }
> +
> + closid = rdtgrp->closid;
> +
> + /* get scratch space to save all the masks while we validate input */
> + for_each_rdt_resource(r) {
> + r->tmp_cbms = kcalloc(r->num_domains << r->cdp_enabled,
> + sizeof(*l3_cbms), GFP_KERNEL);
> + if (!r->tmp_cbms) {
> + ret = -ENOMEM;
> + goto fail;
> + }
> + r->num_cbms = 0;
This wants to be r->num_tmp_cbms for clarity.
Thanks,
tglx
[toc] | [prev] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-10-15 01:20 +0200 |
| Subject | [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <ssjVD-7Za-15@gated-at.bofh.it> |
| In reply to | #1501214 |
From: Fenghua Yu <fenghua.yu@intel.com>
Resource control groups are represented as directories in the resctrl
file system. The root directory describes the default resources available
to tasks that have not been assigned specific resources. Other directories
can be created at the root level to make new resource groups. It is not
permitted to make directories within other directories.
Hardware uses a CLOSID (Class of service ID) to determine which resource
limits are currently in effect. The exact number available is enumerated
by CPUID leaf 0x10, but on current implementations it is a small number.
We implement a simple bitmask allocator for CLOSIDs.
Each resource control group uses one CLOSID, which limits the total number
of directories that can be created.
Resource groups can be removed using rmdir.
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Signed-off-by: Tony Luck <tony.luck@intel.com>
---
arch/x86/include/asm/intel_rdt.h | 9 ++
arch/x86/kernel/cpu/intel_rdt_rdtgroup.c | 255 +++++++++++++++++++++++++++++++
2 files changed, 264 insertions(+)
diff --git a/arch/x86/include/asm/intel_rdt.h b/arch/x86/include/asm/intel_rdt.h
index ea8c09b3..7eb8078 100644
--- a/arch/x86/include/asm/intel_rdt.h
+++ b/arch/x86/include/asm/intel_rdt.h
@@ -9,13 +9,20 @@
* @kn: kernfs node
* @rdtgroup_list: linked list for all rdtgroups
* @closid: closid for this rdtgroup
+ * @flags: status bits
+ * @waitcount: how many cpus expect to find this
*/
struct rdtgroup {
struct kernfs_node *kn;
struct list_head rdtgroup_list;
int closid;
+ int flags;
+ atomic_t waitcount;
};
+/* rdtgroup.flags */
+#define RDT_DELETED 1
+
/* List of all resource groups */
extern struct list_head rdt_all_groups;
@@ -142,4 +149,6 @@ union cpuid_0x10_1_edx {
};
void rdt_cbm_update(void *arg);
+struct rdtgroup *rdtgroup_kn_lock_live(struct kernfs_node *kn);
+void rdtgroup_kn_unlock(struct kernfs_node *kn);
#endif /* _ASM_X86_INTEL_RDT_H */
diff --git a/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c b/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
index 316fa0c..9e0044d 100644
--- a/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
+++ b/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
@@ -26,16 +26,79 @@
#include <linux/seq_file.h>
#include <linux/sched.h>
#include <linux/slab.h>
+#include <linux/cpu.h>
#include <uapi/linux/magic.h>
#include <asm/intel_rdt.h>
+#include <asm/intel_rdt_common.h>
DEFINE_STATIC_KEY_FALSE(rdt_enable_key);
struct kernfs_root *rdt_root;
struct rdtgroup rdtgroup_default;
LIST_HEAD(rdt_all_groups);
+/*
+ * Trivial allocator for CLOSIDs. Since h/w only supports a small number,
+ * we can keep a bitmap of free CLOSIDs in a single integer.
+ *
+ * Please note: This only supports global CLOSID across multiple
+ * resources and multiple sockets. User can create rdtgroups including root
+ * rdtgroup up to the number of CLOSIDs, which is 16 on Broadwell. When
+ * number of caches is big or number of supported resources sharing CLOSID
+ * is growing, it's getting harder to find usable rdtgroups which is limited
+ * by the small number of CLOSIDs.
+ *
+ * In the future, if it's necessary, we can implement more complex CLOSID
+ * allocation per socket/per resource domain and utilize CLOSIDs as many
+ * as possible. E.g. on 2-socket Broadwell, user can create upto 16x16=256
+ * rdtgroups and each rdtgroup has different combination of two L3 CBMs.
+ */
+static int closid_free_map;
+
+static void closid_init(void)
+{
+ struct rdt_resource *r = &rdt_resources_all[RDT_RESOURCE_L3];
+ int rdt_max_closid;
+
+ /* Enabling L3 CDP halves the number of CLOSIDs */
+ if (r->cdp_enabled)
+ r->num_closid = r->max_closid / 2;
+ else
+ r->num_closid = r->max_closid;
+
+ /* Compute rdt_max_closid across all resources */
+ rdt_max_closid = 0;
+ for_each_rdt_resource(r)
+ rdt_max_closid = max(rdt_max_closid, r->num_closid);
+ if (rdt_max_closid > 32) {
+ pr_warn_once("Only using 32/%d CLOSIDs\n", rdt_max_closid);
+ rdt_max_closid = 32;
+ }
+
+ closid_free_map = BIT_MASK(rdt_max_closid) - 1;
+
+ /* CLOSID 0 is always reserved for the default group */
+ closid_free_map &= ~1;
+}
+
+int closid_alloc(void)
+{
+ int closid = ffs(closid_free_map);
+
+ if (closid == 0)
+ return -ENOSPC;
+ closid--;
+ closid_free_map &= ~(1 << closid);
+
+ return closid;
+}
+
+static void closid_free(int closid)
+{
+ closid_free_map |= 1 << closid;
+}
+
/* set uid and gid of rdtgroup dirs and files to that of the creator */
static int rdtgroup_kn_set_ugid(struct kernfs_node *kn)
{
@@ -249,6 +312,54 @@ static int parse_rdtgroupfs_options(char *data)
return 0;
}
+/*
+ * We don't allow rdtgroup directories to be created anywhere
+ * except the root directory. Thus when looking for the rdtgroup
+ * structure for a kernfs node we are either looking at a directory,
+ * in which case the rdtgroup structure is pointed at by the "priv"
+ * field, otherwise we have a file, and need only look to the parent
+ * to find the rdtgroup.
+ */
+static struct rdtgroup *kernfs_to_rdtgroup(struct kernfs_node *kn)
+{
+ if (kernfs_type(kn) == KERNFS_DIR)
+ return kn->priv;
+ else
+ return kn->parent->priv;
+}
+
+struct rdtgroup *rdtgroup_kn_lock_live(struct kernfs_node *kn)
+{
+ struct rdtgroup *rdtgrp = kernfs_to_rdtgroup(kn);
+
+ atomic_inc(&rdtgrp->waitcount);
+ kernfs_break_active_protection(kn);
+
+ mutex_lock(&rdtgroup_mutex);
+
+ /* Was this group deleted while we waited? */
+ if (rdtgrp->flags & RDT_DELETED)
+ return NULL;
+
+ return rdtgrp;
+}
+
+void rdtgroup_kn_unlock(struct kernfs_node *kn)
+{
+ struct rdtgroup *rdtgrp = kernfs_to_rdtgroup(kn);
+
+ mutex_unlock(&rdtgroup_mutex);
+
+ if (atomic_dec_and_test(&rdtgrp->waitcount) &&
+ (rdtgrp->flags & RDT_DELETED)) {
+ kernfs_unbreak_active_protection(kn);
+ kernfs_put(kn);
+ kfree(rdtgrp);
+ } else {
+ kernfs_unbreak_active_protection(kn);
+ }
+}
+
static struct dentry *rdt_mount(struct file_system_type *fs_type,
int flags, const char *unused_dev_name,
void *data)
@@ -272,6 +383,8 @@ static struct dentry *rdt_mount(struct file_system_type *fs_type,
goto out;
}
+ closid_init();
+
dentry = kernfs_mount(fs_type, flags, rdt_root,
RDTGROUP_SUPER_MAGIC, &new_sb);
if (IS_ERR(dentry))
@@ -319,6 +432,39 @@ static void reset_all_cbms(struct rdt_resource *r)
put_cpu();
}
+static void rdt_reset_pqr_assoc_closid(void *v)
+{
+ struct intel_pqr_state *state = this_cpu_ptr(&pqr_state);
+
+ state->closid = 0;
+ wrmsr(MSR_IA32_PQR_ASSOC, state->rmid, 0);
+}
+
+/*
+ * Forcibly remove all of subdirectories under root.
+ */
+static void rmdir_all_sub(void)
+{
+ struct rdtgroup *rdtgrp;
+ struct list_head *l, *next;
+
+ get_cpu();
+ /* Reset PQR_ASSOC MSR on this cpu. */
+ rdt_reset_pqr_assoc_closid(NULL);
+ /* Reset PQR_ASSOC MSR on the rest of cpus. */
+ smp_call_function_many(cpu_online_mask, rdt_reset_pqr_assoc_closid,
+ NULL, 1);
+ put_cpu();
+ list_for_each_safe(l, next, &rdt_all_groups) {
+ rdtgrp = list_entry(l, struct rdtgroup, rdtgroup_list);
+ if (rdtgrp == &rdtgroup_default)
+ continue;
+ kernfs_remove(rdtgrp->kn);
+ list_del(&rdtgrp->rdtgroup_list);
+ kfree(rdtgrp);
+ }
+}
+
static void rdt_kill_sb(struct super_block *sb)
{
struct rdt_resource *r;
@@ -334,6 +480,7 @@ static void rdt_kill_sb(struct super_block *sb)
set_l3_qos_cfg(r);
}
+ rmdir_all_sub();
static_branch_disable(&rdt_enable_key);
kernfs_kill_sb(sb);
mutex_unlock(&rdtgroup_mutex);
@@ -345,7 +492,115 @@ static struct file_system_type rdt_fs_type = {
.kill_sb = rdt_kill_sb,
};
+static int rdtgroup_mkdir(struct kernfs_node *parent_kn, const char *name,
+ umode_t mode)
+{
+ struct rdtgroup *parent, *rdtgrp;
+ struct kernfs_node *kn;
+ int ret, closid;
+
+ /* Only allow mkdir in the root directory */
+ if (parent_kn != rdtgroup_default.kn)
+ return -EPERM;
+
+ /* Do not accept '\n' to avoid unparsable situation. */
+ if (strchr(name, '\n'))
+ return -EINVAL;
+
+ parent = rdtgroup_kn_lock_live(parent_kn);
+ if (!parent) {
+ ret = -ENODEV;
+ goto out_unlock;
+ }
+
+ ret = closid_alloc();
+ if (ret < 0)
+ goto out_unlock;
+ closid = ret;
+
+ /* allocate the rdtgroup. */
+ rdtgrp = kzalloc(sizeof(*rdtgrp), GFP_KERNEL);
+ if (!rdtgrp) {
+ ret = -ENOSPC;
+ goto out_closid_free;
+ }
+ rdtgrp->closid = closid;
+ list_add(&rdtgrp->rdtgroup_list, &rdt_all_groups);
+
+ /* kernfs creates the directory for rdtgrp */
+ kn = kernfs_create_dir(parent->kn, name, mode, rdtgrp);
+ if (IS_ERR(kn)) {
+ ret = PTR_ERR(kn);
+ goto out_cancel_ref;
+ }
+ rdtgrp->kn = kn;
+
+ /*
+ * This extra ref will be put in kernfs_remove() and guarantees
+ * that @rdtgrp->kn is always accessible.
+ */
+ kernfs_get(kn);
+
+ ret = rdtgroup_kn_set_ugid(kn);
+ if (ret)
+ goto out_destroy;
+
+ kernfs_activate(kn);
+
+ ret = 0;
+ goto out_unlock;
+
+out_destroy:
+ kernfs_remove(rdtgrp->kn);
+out_cancel_ref:
+ kfree(rdtgrp);
+out_closid_free:
+ closid_free(closid);
+out_unlock:
+ rdtgroup_kn_unlock(parent_kn);
+ return ret;
+}
+
+static int rdtgroup_rmdir(struct kernfs_node *kn)
+{
+ struct rdtgroup *rdtgrp;
+ int ret = 0;
+
+ rdtgrp = rdtgroup_kn_lock_live(kn);
+ if (!rdtgrp) {
+ rdtgroup_kn_unlock(kn);
+ return -ENOENT;
+ }
+
+ /*
+ * rmdir is for deleting resource groups. Don't
+ * allow deletion of "info" or any of its subdirectories
+ */
+ if (!rdtgrp) {
+ mutex_unlock(&rdtgroup_mutex);
+ kernfs_unbreak_active_protection(kn);
+ return -EPERM;
+ }
+
+ rdtgrp->flags = RDT_DELETED;
+ closid_free(rdtgrp->closid);
+ list_del(&rdtgrp->rdtgroup_list);
+
+ /*
+ * one extra hold on this, will drop when we kfree(rdtgrp)
+ * in rdtgroup_kn_unlock()
+ */
+ kernfs_get(kn);
+ kernfs_remove(rdtgrp->kn);
+
+ rdtgroup_kn_unlock(kn);
+
+ return ret;
+}
+
static struct kernfs_syscall_ops rdtgroup_kf_syscall_ops = {
+ .mkdir = rdtgroup_mkdir,
+ .rmdir = rdtgroup_rmdir,
};
static int __init rdtgroup_setup_root(void)
--
2.5.0
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-10-17 23:20 +0200 |
| Subject | Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <stnua-De-17@gated-at.bofh.it> |
| In reply to | #1501217 |
On Fri, 14 Oct 2016, Fenghua Yu wrote:
> +/*
> + * Trivial allocator for CLOSIDs. Since h/w only supports a small number,
> + * we can keep a bitmap of free CLOSIDs in a single integer.
> + *
> + * Please note: This only supports global CLOSID across multiple
> + * resources and multiple sockets. User can create rdtgroups including root
> + * rdtgroup up to the number of CLOSIDs, which is 16 on Broadwell. When
> + * number of caches is big or number of supported resources sharing CLOSID
> + * is growing, it's getting harder to find usable rdtgroups which is limited
> + * by the small number of CLOSIDs.
> + *
> + * In the future, if it's necessary, we can implement more complex CLOSID
> + * allocation per socket/per resource domain and utilize CLOSIDs as many
> + * as possible. E.g. on 2-socket Broadwell, user can create upto 16x16=256
> + * rdtgroups and each rdtgroup has different combination of two L3 CBMs.
I'm confused as usual, but a two socket broadwell has exactly two L3 cache
domains and exactly 16 CLOSIDs per cache domain.
If you take CDP into account then the number of CLOSIDs is reduced to 8 per
cache domains.
So we never can have more than nr(CLOSIDs) * nr(L3 cache domains) unique
settings. So for a two socket broadwell its 32 for !CDP and 16 for CDP.
With the proposed user interface the number of unique rdtgroups is simply
the number of CLOSIDs because we handle the cache domains already per
resource, i.e. the meaning of CLOSID can be set independently per cache
domain.
Can you please explain why you think that we can have 16x16 unique
rdtgroups if we just have 16 resp. 8 CLOSIDs available?
> +static int closid_free_map;
> +
> +static void closid_init(void)
> +{
> + struct rdt_resource *r = &rdt_resources_all[RDT_RESOURCE_L3];
> + int rdt_max_closid;
> +
> + /* Enabling L3 CDP halves the number of CLOSIDs */
> + if (r->cdp_enabled)
> + r->num_closid = r->max_closid / 2;
> + else
> + r->num_closid = r->max_closid;
If you do the L3Data/L3Core thingy then this can go away.
> + /* Compute rdt_max_closid across all resources */
> + rdt_max_closid = 0;
> + for_each_rdt_resource(r)
> + rdt_max_closid = max(rdt_max_closid, r->num_closid);
Oh no! This needs to be min().
Assume you have a system with L3 and L2 CAT. L2 reports COS_MAX=16, L3
reports COS_MAX=16 as well. Then you enabled CDP which cuts L3 COS_MAX in
half. So the real usable number of CLOSIDs is going to be 8 for both L2 and
L3 simply because you do not have a seperation of L2 and L3 in
MSR_PQR_ASSOC. And if you allow 16 then any CLOSID > 8 will result in
undefined behaviour. See SDM:
"When CDP is enabled, specifying a COS value in IA32_PQR_ASSOC.COS outside
of the lower half of the COS space will cause undefined performance impact
to code and data fetches due to MSR space re-indexing into code/data masks
when CDP is enabled."
> + if (rdt_max_closid > 32) {
> + pr_warn_once("Only using 32/%d CLOSIDs\n", rdt_max_closid);
You have a fancy for pr_*_once(). How often is this going to be executed?
Once per mount and we can really print it each time. mount is hardly a fast
path operation.
Also please spell out: "Using only 32 of %d CLOSIDs"
because "Using only 32/64 COSIDs" does not make any sense.
> + rdt_max_closid = 32;
> + }
> +static void rdt_reset_pqr_assoc_closid(void *v)
> +{
> + struct intel_pqr_state *state = this_cpu_ptr(&pqr_state);
> +
> + state->closid = 0;
> + wrmsr(MSR_IA32_PQR_ASSOC, state->rmid, 0);
So this is called with preemption disabled, but is thread context the only
context in which pqr_state can be modified? If yes, please add a comment
explaining this clearly, otherwise the protection is not sufficient. I'm
too lazy to stare into the monitoring code ...
> +/*
> + * Forcibly remove all of subdirectories under root.
> + */
> +static void rmdir_all_sub(void)
> +{
> + struct rdtgroup *rdtgrp;
> + struct list_head *l, *next;
> +
> + get_cpu();
> + /* Reset PQR_ASSOC MSR on this cpu. */
> + rdt_reset_pqr_assoc_closid(NULL);
> + /* Reset PQR_ASSOC MSR on the rest of cpus. */
> + smp_call_function_many(cpu_online_mask, rdt_reset_pqr_assoc_closid,
> + NULL, 1);
So you have reset the CLOSIDs to 0, but what resets L3/2_QOS_MASK_0 to the
default value (all valid bits set)? Further what clears CDP?
> +static int rdtgroup_mkdir(struct kernfs_node *parent_kn, const char *name,
> + umode_t mode)
> +{
....
> + /* allocate the rdtgroup. */
> + rdtgrp = kzalloc(sizeof(*rdtgrp), GFP_KERNEL);
> + if (!rdtgrp) {
> + ret = -ENOSPC;
> + goto out_closid_free;
> + }
> + rdtgrp->closid = closid;
> + list_add(&rdtgrp->rdtgroup_list, &rdt_all_groups);
> +
> + /* kernfs creates the directory for rdtgrp */
> + kn = kernfs_create_dir(parent->kn, name, mode, rdtgrp);
> + if (IS_ERR(kn)) {
> + ret = PTR_ERR(kn);
> + goto out_cancel_ref;
Yuck! You are going to free rdtgrp, which is still enqueued in
rdt_all_groups...
> +static int rdtgroup_rmdir(struct kernfs_node *kn)
> +{
> + struct rdtgroup *rdtgrp;
> + int ret = 0;
> +
> + rdtgrp = rdtgroup_kn_lock_live(kn);
> + if (!rdtgrp) {
> + rdtgroup_kn_unlock(kn);
> + return -ENOENT;
> + }
> +
> + /*
> + * rmdir is for deleting resource groups. Don't
> + * allow deletion of "info" or any of its subdirectories
> + */
> + if (!rdtgrp) {
And how are we going to reach this? You checked !rdtgrp already above ....
> + mutex_unlock(&rdtgroup_mutex);
> + kernfs_unbreak_active_protection(kn);
> + return -EPERM;
> + }
> +
> + rdtgrp->flags = RDT_DELETED;
> + closid_free(rdtgrp->closid);
> + list_del(&rdtgrp->rdtgroup_list);
> +
> + /*
> + * one extra hold on this, will drop when we kfree(rdtgrp)
> + * in rdtgroup_kn_unlock()
> + */
> + kernfs_get(kn);
So in rdtgrp_mkdir you have this:
> + /*
> + * This extra ref will be put in kernfs_remove() and guarantees
> + * that @rdtgrp->kn is always accessible.
> + */
> + kernfs_get(kn);
> +
That smells fishy ,..
Thanks,
tglx
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2016-10-18 00:00 +0200 |
| Subject | Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <sto6R-YG-3@gated-at.bofh.it> |
| In reply to | #1502468 |
On Mon, Oct 17, 2016 at 11:14:55PM +0200, Thomas Gleixner wrote: > > + /* Compute rdt_max_closid across all resources */ > > + rdt_max_closid = 0; > > + for_each_rdt_resource(r) > > + rdt_max_closid = max(rdt_max_closid, r->num_closid); > > Oh no! This needs to be min(). > > Assume you have a system with L3 and L2 CAT. L2 reports COS_MAX=16, L3 > reports COS_MAX=16 as well. Then you enabled CDP which cuts L3 COS_MAX in > half. So the real usable number of CLOSIDs is going to be 8 for both L2 and > L3 simply because you do not have a seperation of L2 and L3 in > MSR_PQR_ASSOC. And if you allow 16 then any CLOSID > 8 will result in > undefined behaviour. See SDM: > > "When CDP is enabled, specifying a COS value in IA32_PQR_ASSOC.COS outside > of the lower half of the COS space will cause undefined performance impact > to code and data fetches due to MSR space re-indexing into code/data masks > when CDP is enabled." Bother. The SDM also has this gem: 17.17.5.1 Cache Allocation Technology Dynamic Configuration Both the CAT masks and CQM registers are accessible and modifiable at any time during execution using RDMSR/WRMSR unless otherwise noted. When writing to these MSRs a #GP(0) will be generated if any of the following conditions occur: * Writing a COS greater than the supported maximum (specified as the maximum value of CPUID.(EAX=10H, ECX=ResID):EDX[15:0] for all valid ResID values) is written to the IA32_PQR_ASSOC.CLOS field. With the intent here being that if you have more of one resource than another, you can use all of the resources in the larger (with the resource with fewer mask registers defaulting to the maximum value when PQR_ASSOC.COS is too large [and I can't find the text that talks about that default behaviour :-( ] I think this all means that L3/CDP is "special". If CDP is on, we can't use the top half of the CLOSID space that CPUID.(EAX=10H, ECX=1):EDX[15:0] told us is present. So we can't exceed the half-way point. But in other cases like L3 with max=16 and L2 with max=8 we should allow 16 groups, but 0-7 allow control of L3 and L2, while 8-15 only allow L3 control (the schemata code enforces this when you get to that part). Perhaps we can encode this in another field in the rdt_resource structure that says that some maximums globally override all others, while some can be legitimately exceeded. I'll have some words with the h/w architect, and get the SDM fixed in the next edition. -Tony
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-10-18 01:00 +0200 |
| Subject | Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <stp2V-1JC-1@gated-at.bofh.it> |
| In reply to | #1502502 |
On Mon, 17 Oct 2016, Luck, Tony wrote: > On Mon, Oct 17, 2016 at 11:14:55PM +0200, Thomas Gleixner wrote: > > > + /* Compute rdt_max_closid across all resources */ > > > + rdt_max_closid = 0; > > > + for_each_rdt_resource(r) > > > + rdt_max_closid = max(rdt_max_closid, r->num_closid); > > > > Oh no! This needs to be min(). > > > > Assume you have a system with L3 and L2 CAT. L2 reports COS_MAX=16, L3 > > reports COS_MAX=16 as well. Then you enabled CDP which cuts L3 COS_MAX in > > half. So the real usable number of CLOSIDs is going to be 8 for both L2 and > > L3 simply because you do not have a seperation of L2 and L3 in > > MSR_PQR_ASSOC. And if you allow 16 then any CLOSID > 8 will result in > > undefined behaviour. See SDM: > > > > "When CDP is enabled, specifying a COS value in IA32_PQR_ASSOC.COS outside > > of the lower half of the COS space will cause undefined performance impact > > to code and data fetches due to MSR space re-indexing into code/data masks > > when CDP is enabled." > > Bother. The SDM also has this gem: > > 17.17.5.1 Cache Allocation Technology Dynamic Configuration > Both the CAT masks and CQM registers are accessible and modifiable at > any time during execution using RDMSR/WRMSR unless otherwise noted. When > writing to these MSRs a #GP(0) will be generated if any of the following > conditions occur: > > * Writing a COS greater than the supported maximum (specified as the > maximum value of CPUID.(EAX=10H, ECX=ResID):EDX[15:0] for all valid > ResID values) is written to the IA32_PQR_ASSOC.CLOS field. > > With the intent here being that if you have more of one resource than > another, you can use all of the resources in the larger (with the > resource with fewer mask registers defaulting to the maximum value > when PQR_ASSOC.COS is too large [and I can't find the text that talks > about that default behaviour :-( ] > > I think this all means that L3/CDP is "special". If CDP is on, we can't > use the top half of the CLOSID space that CPUID.(EAX=10H, ECX=1):EDX[15:0] > told us is present. So we can't exceed the half-way point. That's my understanding. > But in other cases like L3 with max=16 and L2 with max=8 we should allow > 16 groups, but 0-7 allow control of L3 and L2, while 8-15 only allow L3 > control (the schemata code enforces this when you get to that part). Cute. Welcome to configuration hell! So how are we going to deal with that in the schematas? Assume the L3=16 and L2=8 case(no CDP). So effectively any write of L2 to CLOSID=0 will affect the setting of L2 in CLOSID=8. Will the code tell the user that L2 cannot be set for CLOSID >= 8? Will it print the setting of CLOSID - 8 for CLOSID >= 8 when you read the schemata file of CLOSID >= 8? And of course this becomes very interesting when CLOSID 1 is deleted, then what happens to CLOSID 9? Not to talk about the case when CLOSID 1 is reused again. Seems to be a well thought out hardware feature once again. > Perhaps we can encode this in another field in the rdt_resource structure > that says that some maximums globally override all others, while some > can be legitimately exceeded. That should work. > I'll have some words with the h/w architect, and get the SDM fixed in > the next edition. Whatever the outcome will be :) Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-10-18 01:10 +0200 |
| Subject | RE: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <stpcB-226-5@gated-at.bofh.it> |
| In reply to | #1502568 |
On Mon, 17 Oct 2016, Luck, Tony wrote: > > So how are we going to deal with that in the schematas? Assume the L3=16 > > and L2=8 case(no CDP). So effectively any write of L2 to CLOSID=0 will > > affect the setting of L2 in CLOSID=8. > > > > Will the code tell the user that L2 cannot be set for CLOSID >= 8? > > > > Will it print the setting of CLOSID - 8 for CLOSID >= 8 when you read the > > schemata file of CLOSID >= 8? > > The default and first 7 directories you make will have lines in schemata for > both L3 and L2. The next 8 will just have L3 (and won't let you add L2). > > > And of course this becomes very interesting when CLOSID 1 is deleted, then > > what happens to CLOSID 9? Not to talk about the case when CLOSID 1 is > > reused again. > > Deleting CLOSID 1 won't affect CLOSID 9. How so? CLOSID 9 is using CLOSID 1 L2 settings. Are we just keeping the L2 setting of CLOSID 1 around and do not reset it to default? > When you make a new directory that gets allocated CLOSID 1, it will have > both L3 and L2 lines. Now CLOSID 1 is reused and reset to all 1's per default and even if not then the operator will set a new L2 value and therefor wreckaging CLOSID 9. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2016-10-18 01:20 +0200 |
| Subject | RE: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <stpmh-25m-5@gated-at.bofh.it> |
| In reply to | #1502573 |
> How so? CLOSID 9 is using CLOSID 1 L2 settings. Are we just keeping the L2 > setting of CLOSID 1 around and do not reset it to default? No. When CLOSID 9 arrives at the L2 h/w, it doesn't just take the bits it likes an discard the high bits to map to L2_CBM[1]. It just turns into into the maximum allowed value for an L2 CBM. -Tony
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-10-18 01:30 +0200 |
| Subject | RE: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <stpvX-28I-9@gated-at.bofh.it> |
| In reply to | #1502579 |
On Mon, 17 Oct 2016, Luck, Tony wrote: > > How so? CLOSID 9 is using CLOSID 1 L2 settings. Are we just keeping the L2 > > setting of CLOSID 1 around and do not reset it to default? > > No. When CLOSID 9 arrives at the L2 h/w, it doesn't just take the bits it > likes an discard the high bits to map to L2_CBM[1]. It just turns into > into the maximum allowed value for an L2 CBM. So all CLOSIDs >= 8 will use all valid CBM bits for L2? That's even more insane as this breaks any L2 partitioning which is set up by CLOSIDs < 8. So effectively CLOSIDs >= 8 are useless in the L3=16 and L2=8 case. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2016-10-18 01:10 +0200 |
| Subject | RE: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <stpcB-226-7@gated-at.bofh.it> |
| In reply to | #1502568 |
> So how are we going to deal with that in the schematas? Assume the L3=16 > and L2=8 case(no CDP). So effectively any write of L2 to CLOSID=0 will > affect the setting of L2 in CLOSID=8. > > Will the code tell the user that L2 cannot be set for CLOSID >= 8? > > Will it print the setting of CLOSID - 8 for CLOSID >= 8 when you read the > schemata file of CLOSID >= 8? The default and first 7 directories you make will have lines in schemata for both L3 and L2. The next 8 will just have L3 (and won't let you add L2). > And of course this becomes very interesting when CLOSID 1 is deleted, then > what happens to CLOSID 9? Not to talk about the case when CLOSID 1 is > reused again. Deleting CLOSID 1 won't affect CLOSID 9. When you make a new directory that gets allocated CLOSID 1, it will have both L3 and L2 lines. -Tony
[toc] | [prev] | [next] | [standalone]
| From | Fenghua Yu <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-10-18 00:20 +0200 |
| Subject | Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <stoqd-1pY-13@gated-at.bofh.it> |
| In reply to | #1502468 |
On Mon, Oct 17, 2016 at 11:14:55PM +0200, Thomas Gleixner wrote: > On Fri, 14 Oct 2016, Fenghua Yu wrote: > > +/* > > + * Trivial allocator for CLOSIDs. Since h/w only supports a small number, > > + * we can keep a bitmap of free CLOSIDs in a single integer. > > + * > > + * Please note: This only supports global CLOSID across multiple > > + * resources and multiple sockets. User can create rdtgroups including root > > + * rdtgroup up to the number of CLOSIDs, which is 16 on Broadwell. When > > + * number of caches is big or number of supported resources sharing CLOSID > > + * is growing, it's getting harder to find usable rdtgroups which is limited > > + * by the small number of CLOSIDs. > > + * > > + * In the future, if it's necessary, we can implement more complex CLOSID > > + * allocation per socket/per resource domain and utilize CLOSIDs as many > > + * as possible. E.g. on 2-socket Broadwell, user can create upto 16x16=256 > > + * rdtgroups and each rdtgroup has different combination of two L3 CBMs. > > I'm confused as usual, but a two socket broadwell has exactly two L3 cache > domains and exactly 16 CLOSIDs per cache domain. > > If you take CDP into account then the number of CLOSIDs is reduced to 8 per > cache domains. > > So we never can have more than nr(CLOSIDs) * nr(L3 cache domains) unique > settings. So for a two socket broadwell its 32 for !CDP and 16 for CDP. > > With the proposed user interface the number of unique rdtgroups is simply > the number of CLOSIDs because we handle the cache domains already per > resource, i.e. the meaning of CLOSID can be set independently per cache > domain. > > Can you please explain why you think that we can have 16x16 unique > rdtgroups if we just have 16 resp. 8 CLOSIDs available? To simplify, we only consider CAT case. CDP has similar similar situation. For a two socket broadwell, we have schemata format "L3:0=x;1=y" suppose the two cache ids are 0 and 1, and x is cache 0's cbm and y is cache 1's cbm. Then kernel allocates one closid and its cbm=x for cache 0 and one closid and its cbm=y for cache 1. So we can have the following 16x16 different partitions/rdtgroups. Each partition/rdgroup has its name and has its own unique closid combinations on two caches. If a task is assigned to any of partition, the task has its unique combination of closids when running on cache 0 and when running on cache 1. Belowing cbm values are example values. name schemata closids on cache 0 and 1 allocated by kernel ---- -------- -------------------------------------------- (closid 0 on cache0 combined with 16 different closid on cache1) part0: L3:0=1;1=1 closid0/cbm=1 on cache0 and closid0/cbm=1 on cache1 part1: L3:0=1;1=3 closid0/cbm=1 on cache0 and closid1/cbm=3 on cache1 part2: L3:0=1;1=7 closid0/cbm=1 on cache0 and closid2/cbm=7 on cache1 part3: L3:0=1;1=f closid0/cbm=1 on cache0 and closid3/cbm=f on cache1 part4: L3:0=1;1=1f closid0/cbm=1 on cache0 and closid4/cbm=1f on cache1 part5: L3:0=1;1=3f closid0/cbm=1 on cache0 and closid5/cbm=3f on cache1 part6: L3:0=1;1=7f closid0/cbm=1 on cache0 and closid6/cbm=7f on cache1 part7: L3:0=1;1=ff closid0/cbm=1 on cache0 and closid7/cbm=ff on cache1 part8: L3:0=1;1=1ff closid0/cbm=1 on cache0 and closid8/cbm=1ff on cache1 part9: L3:0=1;1=3ff closid0/cbm=1 on cache0 and closid9/cbm=3ff on cache1 part10: L3:0=1;1=7ff closid0/cbm=1 on cache0 and closid10/cbm=7ff on cache1 part11: L3:0=1;1=fff closid0/cbm=1 on cache0 and closid11/cbm=fff on cache1 part12: L3:0=1;1=1fff closid0/cbm=1 on cache0 and closid12/cbm=1fff on cache1 part13: L3:0=1;1=3fff closid0/cbm=1 on cache0 and closid13/cbm=3fff on cache1 part14: L3:0=1;1=7fff closid0/cbm=1 on cache0 and closid14/cbm=7fff on cache1 part15: L3:0=1;1=ffff closid0/cbm=1 on cache0 and closid15/cbm=ffff on cache1 (closid 1 on cache0 combined with 16 different closid on cache1) part16: L3:0=3;1=1 closid1/cbm=3 on cache0 and closid0/cbm=1 on cache1 part17: L3:0=3;1=3 closid1/cbm=3 on cache0 and closid1/cbm=3 on cache1 ... part31: L3:0=3;1=ffff closid1/cbm=3 on cache0 and closid15/cbm=ffff on cache1 (closid 2 on cache0 combined with 16 different closid on cache1) part16: L3:0=7;1=1 closid2/cbm=3 on cache0 and closid0/cbm=1 on cache1 part17: L3:0=7;1=3 closid2/cbm=3 on cache0 and closid1/cbm=3 on cache1 ... part31: L3:0=7;1=ffff closid2/cbm=3 on cache0 and closid15/cbm=ffff on cache1 (closid 3 on cache0 combined with 16 different closids on cache1) ... (closid 4 on cache0 combined with 16 different closids on cache1) ... (closid 5 on cache0 combined with 16 different closids on cache1) ... (closid 6 on cache0 combined with 16 different closids on cache1) ... (closid 7 on cache0 combined with 16 different closids on cache1) ... (closid 8 on cache0 combined with 16 different closids on cache1) ... (closid 9 on cache0 combined with 16 different closids on cache1) ... (closid 10 on cache0 combined with 16 different closids on cache1) ... (closid 11 on cache0 combined with 16 different closids on cache1) ... (closid 12 on cache0 combined with 16 different closids on cache1) ... (closid 13 on cache0 combined with 16 different closids on cache1) ... (closid 14 on cache0 combined with 16 different closids on cache1) ... (closid 15 on cache0 combined with 16 different closids on cache1) ... part254: L3:0=ffff;1=7fff closid15/cbm=ffff on cache0 and closid14/cbm=7fff on cache1 part255: L3:0=ffff;1=ffff closid15/cbm=ffff on cache0 and closid15/cbm=ffff on cache1 To utilize as much combinations as possbile, we may implement a more complex allocation than current one. Does this make sense? Thanks. -Fenghua
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-10-18 01:30 +0200 |
| Subject | Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <stpvX-28I-5@gated-at.bofh.it> |
| In reply to | #1502525 |
On Mon, 17 Oct 2016, Fenghua Yu wrote: > part0: L3:0=1;1=1 closid0/cbm=1 on cache0 and closid0/cbm=1 on cache1 > (closid 15 on cache0 combined with 16 different closids on cache1) > ... > part254: L3:0=ffff;1=7fff closid15/cbm=ffff on cache0 and closid14/cbm=7fff on cache1 > part255: L3:0=ffff;1=ffff closid15/cbm=ffff on cache0 and closid15/cbm=ffff on cache1 > > To utilize as much combinations as possbile, we may implement a > more complex allocation than current one. > > Does this make sense? Thanks for the explanation. I knew that I'm missing something. But how is that supposed to work? The schemata files have no idea of closids simply because the closids are assigned automatically. And that makes the whole thing exponentially complex. You must allow to create ALL rdt groups (initialy as a copy of the root group) and then when the schemata file is written you have to look whether the particular CBM value for a particular domain is already used and assign the same cosid for this domain. That of course makes the whole L2 business completely diffuse because you might end up with: Dom0 = COSID1 and DOM1 = COSID9 So you can set the L2 for Dom0, but not for DOM1 and then if you set L2 for Dom0 you must find a new COSID for Dom0. If there is none, then you must reject the write and leave the admin puzzled. There is a reason why I suggested: https://lkml.kernel.org/r/alpine.DEB.2.11.1511181534450.3761@nanos It's certainly not perfect (missing L2 etc.), but clearly avoids exactly the above issues. And it would allow you to utilize the 256 groups in an understandable way. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | "Luck, Tony" <tony.luck@intel.com> |
|---|---|
| Date | 2016-10-18 01:40 +0200 |
| Subject | Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <stpFD-2bU-1@gated-at.bofh.it> |
| In reply to | #1502582 |
On Tue, Oct 18, 2016 at 01:20:36AM +0200, Thomas Gleixner wrote: > On Mon, 17 Oct 2016, Fenghua Yu wrote: > > part0: L3:0=1;1=1 closid0/cbm=1 on cache0 and closid0/cbm=1 on cache1 > > (closid 15 on cache0 combined with 16 different closids on cache1) > > ... > > part254: L3:0=ffff;1=7fff closid15/cbm=ffff on cache0 and closid14/cbm=7fff on cache1 > > part255: L3:0=ffff;1=ffff closid15/cbm=ffff on cache0 and closid15/cbm=ffff on cache1 > > > > To utilize as much combinations as possbile, we may implement a > > more complex allocation than current one. > > > > Does this make sense? > > Thanks for the explanation. I knew that I'm missing something. > > But how is that supposed to work? The schemata files have no idea of > closids simply because the closids are assigned automatically. And that > makes the whole thing exponentially complex. You must allow to create ALL > rdt groups (initialy as a copy of the root group) and then when the > schemata file is written you have to look whether the particular CBM value > for a particular domain is already used and assign the same cosid for this > domain. That of course makes the whole L2 business completely diffuse > because you might end up with: > > Dom0 = COSID1 and DOM1 = COSID9 > > So you can set the L2 for Dom0, but not for DOM1 and then if you set L2 for > Dom0 you must find a new COSID for Dom0. If there is none, then you must > reject the write and leave the admin puzzled. > > There is a reason why I suggested: > > https://lkml.kernel.org/r/alpine.DEB.2.11.1511181534450.3761@nanos > > It's certainly not perfect (missing L2 etc.), but clearly avoids exactly > the above issues. And it would allow you to utilize the 256 groups in an > understandable way. If you head down that path someone with a 4-socket system will try to make 16x16x16x16 = 65536 groups and "understandable" takes a bit of a beating. The eight socket system with 16^8 = 4G groups defies any rationale hope. Best not to think about 16 sockets. The L2 + L3 configuration space gets unbelievably messy too. There's a reason why I ripped out the allocation code and went with a simple global allocator in this version. If we decide we need something fancier we can adapt later. Some solutions might be transparent to applications, others might add a "closid" file into each directory to give 2nd generation applications hooks to view (and maybe control) which closid is used by each group. -Tony
[toc] | [prev] | [next] | [standalone]
| From | Fenghua Yu <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-10-18 02:00 +0200 |
| Subject | Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <stpZ0-2jK-47@gated-at.bofh.it> |
| In reply to | #1502584 |
On Mon, Oct 17, 2016 at 04:37:30PM -0700, Luck, Tony wrote: > On Tue, Oct 18, 2016 at 01:20:36AM +0200, Thomas Gleixner wrote: > > On Mon, 17 Oct 2016, Fenghua Yu wrote: > > > part0: L3:0=1;1=1 closid0/cbm=1 on cache0 and closid0/cbm=1 on cache1 > > > (closid 15 on cache0 combined with 16 different closids on cache1) > > > ... > > > part254: L3:0=ffff;1=7fff closid15/cbm=ffff on cache0 and closid14/cbm=7fff on cache1 > > > part255: L3:0=ffff;1=ffff closid15/cbm=ffff on cache0 and closid15/cbm=ffff on cache1 > > > > > > To utilize as much combinations as possbile, we may implement a > > > more complex allocation than current one. > > > > > > Does this make sense? > > > > Thanks for the explanation. I knew that I'm missing something. > > > > But how is that supposed to work? The schemata files have no idea of > > closids simply because the closids are assigned automatically. And that > > makes the whole thing exponentially complex. You must allow to create ALL > > rdt groups (initialy as a copy of the root group) and then when the > > schemata file is written you have to look whether the particular CBM value > > for a particular domain is already used and assign the same cosid for this > > domain. That of course makes the whole L2 business completely diffuse > > because you might end up with: > > > > Dom0 = COSID1 and DOM1 = COSID9 > > > > So you can set the L2 for Dom0, but not for DOM1 and then if you set L2 for > > Dom0 you must find a new COSID for Dom0. If there is none, then you must > > reject the write and leave the admin puzzled. > > > > There is a reason why I suggested: > > > > https://lkml.kernel.org/r/alpine.DEB.2.11.1511181534450.3761@nanos > > > > It's certainly not perfect (missing L2 etc.), but clearly avoids exactly > > the above issues. And it would allow you to utilize the 256 groups in an > > understandable way. > > If you head down that path someone with a 4-socket system will try to > make 16x16x16x16 = 65536 groups and "understandable" takes a bit of > a beating. The eight socket system with 16^8 = 4G groups defies any > rationale hope. Best not to think about 16 sockets. The number of 16^L3 cache numbers is max partition number limitation that a sysadmin can create in theory. Beyond the number, allocation returns no space. It's kind of like other cases eg many many mkdir in one directory can fail at one point because mkdir run out of disk space etc. > > The L2 + L3 configuration space gets unbelievably messy too. > > There's a reason why I ripped out the allocation code and went with > a simple global allocator in this version. If we decide we need something > fancier we can adapt later. Some solutions might be transparent to > applications, others might add a "closid" file into each directory to > give 2nd generation applications hooks to view (and maybe control) > which closid is used by each group. Fully agree with Tony. We understand the complexity of the situation and just have a simple and working solution for the first version. Thanks. -Fenghua
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-10-18 12:50 +0200 |
| Subject | Re: [PATCH v4 13/18] x86/intel_rdt: Add mkdir to resctrl file system |
| Message-ID | <stA81-sA-1@gated-at.bofh.it> |
| In reply to | #1502584 |
On Mon, 17 Oct 2016, Luck, Tony wrote: > On Tue, Oct 18, 2016 at 01:20:36AM +0200, Thomas Gleixner wrote: > > It's certainly not perfect (missing L2 etc.), but clearly avoids exactly > > the above issues. And it would allow you to utilize the 256 groups in an > > understandable way. > > If you head down that path someone with a 4-socket system will try to > make 16x16x16x16 = 65536 groups and "understandable" takes a bit of > a beating. The eight socket system with 16^8 = 4G groups defies any > rationale hope. Best not to think about 16 sockets. > > The L2 + L3 configuration space gets unbelievably messy too. > > There's a reason why I ripped out the allocation code and went with > a simple global allocator in this version. If we decide we need something > fancier we can adapt later. Some solutions might be transparent to > applications, others might add a "closid" file into each directory to > give 2nd generation applications hooks to view (and maybe control) > which closid is used by each group. I'm not saying that we want something fancier. I fully agree with your decision to make a simple global allocator. I was just puzzled by the 16*16 comment and wondered what this is about. Looking at Fenghuas explanation and the examples there is nothing which really looks like we ever want it. In fact the fancy CLOSID matrix does not make much sense at all. So I rather would like to see a comment clearly explaining why the chosen allocator (grouping) gives us the most straight forward way to utilize the hardware. It surely restricts the theoretical choices, but it limits them to the subset which makes technically sense. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-10-15 01:20 +0200 |
| Subject | [PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID |
| Message-ID | <ssjVD-7Za-17@gated-at.bofh.it> |
| In reply to | #1501214 |
From: Fenghua Yu <fenghua.yu@intel.com>
Define struct rdt_resource to hold all the parameterized
values for an RDT resource. Fill in some of those values
from CPUID leaf 0x10 (on Haswell we hard code them).
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Signed-off-by: Tony Luck <tony.luck@intel.com>
---
arch/x86/include/asm/intel_rdt.h | 62 ++++++++++++++++++++++++++++++++++
arch/x86/kernel/cpu/intel_rdt.c | 72 +++++++++++++++++++++++++++++++++++++---
2 files changed, 130 insertions(+), 4 deletions(-)
diff --git a/arch/x86/include/asm/intel_rdt.h b/arch/x86/include/asm/intel_rdt.h
index 3aca86d..8c61d83 100644
--- a/arch/x86/include/asm/intel_rdt.h
+++ b/arch/x86/include/asm/intel_rdt.h
@@ -2,5 +2,67 @@
#define _ASM_X86_INTEL_RDT_H
#define IA32_L3_CBM_BASE 0xc90
+#define IA32_L2_CBM_BASE 0xd10
+/**
+ * struct rdt_resource - attributes of an RDT resource
+ * @enabled: Is this feature enabled on this machine
+ * @name: Name to use in "schemata" file
+ * @max_closid: Maximum number of CLOSIDs supported
+ * @num_closid: Current number of CLOSIDs available
+ * @max_cbm: Largest Cache Bit Mask allowed
+ * @min_cbm_bits: Minimum number of bits to be set in a cache
+ * bit mask
+ * @domains: All domains for this resource
+ * @num_domains: Number of domains active
+ * @msr_base: Base MSR address for CBMs
+ * @cdp_capable: Code/Data Prioritization available
+ * @cdp_enabled: Code/Data Prioritization enabled
+ * @tmp_cbms: Scratch space when updating schemata
+ * @cache_level: Which cache level defines scope of this domain
+ */
+struct rdt_resource {
+ bool enabled;
+ char *name;
+ int max_closid;
+ int num_closid;
+ int cbm_len;
+ int min_cbm_bits;
+ u32 max_cbm;
+ struct list_head domains;
+ int num_domains;
+ int msr_base;
+ bool cdp_capable;
+ bool cdp_enabled;
+ u32 *tmp_cbms;
+ int cache_level;
+};
+
+#define for_each_rdt_resource(r) \
+ for (r = rdt_resources_all; r->name; r++) \
+ if (r->enabled)
+
+#define IA32_L3_CBM_BASE 0xc90
+extern struct rdt_resource rdt_resources_all[];
+
+enum {
+ RDT_RESOURCE_L3,
+ RDT_RESOURCE_L2,
+};
+
+/* CPUID.(EAX=10H, ECX=ResID=1).EAX */
+union cpuid_0x10_1_eax {
+ struct {
+ unsigned int cbm_len:5;
+ } split;
+ unsigned int full;
+};
+
+/* CPUID.(EAX=10H, ECX=ResID=1).EDX */
+union cpuid_0x10_1_edx {
+ struct {
+ unsigned int cos_max:16;
+ } split;
+ unsigned int full;
+};
#endif /* _ASM_X86_INTEL_RDT_H */
diff --git a/arch/x86/kernel/cpu/intel_rdt.c b/arch/x86/kernel/cpu/intel_rdt.c
index 9d55942..87f9650 100644
--- a/arch/x86/kernel/cpu/intel_rdt.c
+++ b/arch/x86/kernel/cpu/intel_rdt.c
@@ -31,6 +31,28 @@
#include <asm/intel-family.h>
#include <asm/intel_rdt.h>
+#define domain_init(name) LIST_HEAD_INIT(rdt_resources_all[name].domains)
+
+struct rdt_resource rdt_resources_all[] = {
+ {
+ .name = "L3",
+ .domains = domain_init(RDT_RESOURCE_L3),
+ .msr_base = IA32_L3_CBM_BASE,
+ .min_cbm_bits = 1,
+ .cache_level = 3
+ },
+ {
+ .name = "L2",
+ .domains = domain_init(RDT_RESOURCE_L2),
+ .msr_base = IA32_L2_CBM_BASE,
+ .min_cbm_bits = 1,
+ .cache_level = 2
+ },
+ {
+ /* NULL terminated */
+ }
+};
+
/*
* cache_alloc_hsw_probe() - Have to probe for Intel haswell server CPUs
* as they do not have CPUID enumeration support for Cache allocation.
@@ -53,6 +75,7 @@ static inline bool cache_alloc_hsw_probe(void)
{
u32 l, h;
u32 max_cbm = BIT_MASK(20) - 1;
+ struct rdt_resource *r = &rdt_resources_all[RDT_RESOURCE_L3];
if (wrmsr_safe(IA32_L3_CBM_BASE, max_cbm, 0))
return false;
@@ -60,11 +83,19 @@ static inline bool cache_alloc_hsw_probe(void)
if (l != max_cbm)
return false;
+ r->max_closid = 4;
+ r->num_closid = r->max_closid;
+ r->cbm_len = 20;
+ r->max_cbm = max_cbm;
+ r->min_cbm_bits = 2;
+ r->enabled = true;
+
return true;
}
static inline bool get_rdt_resources(void)
{
+ struct rdt_resource *r;
bool ret = false;
if (boot_cpu_data.x86_vendor == X86_VENDOR_INTEL &&
@@ -74,20 +105,53 @@ static inline bool get_rdt_resources(void)
if (!boot_cpu_has(X86_FEATURE_RDT_A))
return false;
- if (boot_cpu_has(X86_FEATURE_CAT_L3))
+ if (boot_cpu_has(X86_FEATURE_CAT_L3)) {
+ union cpuid_0x10_1_eax eax;
+ union cpuid_0x10_1_edx edx;
+ u32 ebx, ecx;
+
+ r = &rdt_resources_all[RDT_RESOURCE_L3];
+ cpuid_count(0x00000010, 1, &eax.full, &ebx, &ecx, &edx.full);
+ r->max_closid = edx.split.cos_max + 1;
+ r->num_closid = r->max_closid;
+ r->cbm_len = eax.split.cbm_len + 1;
+ r->max_cbm = BIT_MASK(eax.split.cbm_len + 1) - 1;
+ if (boot_cpu_has(X86_FEATURE_CDP_L3))
+ r->cdp_capable = true;
+ r->enabled = true;
+
ret = true;
+ }
+ if (boot_cpu_has(X86_FEATURE_CAT_L2)) {
+ union cpuid_0x10_1_eax eax;
+ union cpuid_0x10_1_edx edx;
+ u32 ebx, ecx;
+
+ /* CPUID 0x10.2 fields are same format at 0x10.1 */
+ r = &rdt_resources_all[RDT_RESOURCE_L2];
+ cpuid_count(0x00000010, 2, &eax.full, &ebx, &ecx, &edx.full);
+ r->max_closid = edx.split.cos_max + 1;
+ r->num_closid = r->max_closid;
+ r->cbm_len = eax.split.cbm_len + 1;
+ r->max_cbm = BIT_MASK(eax.split.cbm_len + 1) - 1;
+ r->enabled = true;
+
+ ret = true;
+ }
return ret;
}
static int __init intel_rdt_late_init(void)
{
+ struct rdt_resource *r;
+
if (!get_rdt_resources())
return -ENODEV;
- pr_info("Intel RDT cache allocation detected\n");
- if (boot_cpu_has(X86_FEATURE_CDP_L3))
- pr_info("Intel RDT code data prioritization detected\n");
+ for_each_rdt_resource(r)
+ pr_info("Intel RDT %s allocation %s detected\n", r->name,
+ r->cdp_capable ? " (with CDP)" : "");
return 0;
}
--
2.5.0
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-10-17 15:50 +0200 |
| Subject | Re: [PATCH v4 08/18] x86/intel_rdt: Pick up L3/L2 RDT parameters from CPUID |
| Message-ID | <stgsG-47u-15@gated-at.bofh.it> |
| In reply to | #1501218 |
On Fri, 14 Oct 2016, Fenghua Yu wrote:
> +/**
> + * struct rdt_resource - attributes of an RDT resource
> + * @enabled: Is this feature enabled on this machine
> + * @name: Name to use in "schemata" file
> + * @max_closid: Maximum number of CLOSIDs supported
> + * @num_closid: Current number of CLOSIDs available
> + * @max_cbm: Largest Cache Bit Mask allowed
> + * @min_cbm_bits: Minimum number of bits to be set in a cache
That should be 'number of consecutive bits', right?
> + * bit mask
> + * @domains: All domains for this resource
> + * @num_domains: Number of domains active
> + * @msr_base: Base MSR address for CBMs
> + * @cdp_capable: Code/Data Prioritization available
> + * @cdp_enabled: Code/Data Prioritization enabled
I wonder whether this is the proper abstraction level. We might as well do
the following:
rdtresources[] = {
{
.name = "L3",
},
{
.name = "L3Data",
},
{
.name = "L3Code",
},
and enable either L3 or L3Data+L3Code. Not sure if that makes things
simpler, but it's definitely worth a thought or two.
> +#define for_each_rdt_resource(r) \
> + for (r = rdt_resources_all; r->name; r++) \
> + if (r->enabled)
So the resource array must be NULL terminated, right? You might as well use
r < rdt_resources_all + ARRAY_SIZE(rdt_resources_all)
as the loop condition. So you avoid the NULL termination.
> +
> +#define IA32_L3_CBM_BASE 0xc90
> +extern struct rdt_resource rdt_resources_all[];
Please visually split this. CBM_BASE has nothing to do with the resource
array.
> +#define domain_init(name) LIST_HEAD_INIT(rdt_resources_all[name].domains)
name is really misleading here. Please use id and make this an inline
function.
> +struct rdt_resource rdt_resources_all[] = {
> static inline bool get_rdt_resources(void)
> {
> + struct rdt_resource *r;
> bool ret = false;
>
> if (boot_cpu_data.x86_vendor == X86_VENDOR_INTEL &&
> @@ -74,20 +105,53 @@ static inline bool get_rdt_resources(void)
>
> if (!boot_cpu_has(X86_FEATURE_RDT_A))
> return false;
> - if (boot_cpu_has(X86_FEATURE_CAT_L3))
> + if (boot_cpu_has(X86_FEATURE_CAT_L3)) {
> + union cpuid_0x10_1_eax eax;
> + union cpuid_0x10_1_edx edx;
> + u32 ebx, ecx;
> +
> + r = &rdt_resources_all[RDT_RESOURCE_L3];
> + cpuid_count(0x00000010, 1, &eax.full, &ebx, &ecx, &edx.full);
> + r->max_closid = edx.split.cos_max + 1;
> + r->num_closid = r->max_closid;
> + r->cbm_len = eax.split.cbm_len + 1;
> + r->max_cbm = BIT_MASK(eax.split.cbm_len + 1) - 1;
> + if (boot_cpu_has(X86_FEATURE_CDP_L3))
> + r->cdp_capable = true;
> + r->enabled = true;
> +
> ret = true;
> + }
> + if (boot_cpu_has(X86_FEATURE_CAT_L2)) {
> + union cpuid_0x10_1_eax eax;
> + union cpuid_0x10_1_edx edx;
> + u32 ebx, ecx;
> +
> + /* CPUID 0x10.2 fields are same format at 0x10.1 */
> + r = &rdt_resources_all[RDT_RESOURCE_L2];
> + cpuid_count(0x00000010, 2, &eax.full, &ebx, &ecx, &edx.full);
> + r->max_closid = edx.split.cos_max + 1;
> + r->num_closid = r->max_closid;
> + r->cbm_len = eax.split.cbm_len + 1;
> + r->max_cbm = BIT_MASK(eax.split.cbm_len + 1) - 1;
> + r->enabled = true;
Copy and paste is a wonderful thing, right?
static void rdt_get_config(int idx, struct rdt_resource *r)
{
union cpuid_0x10_1_eax eax;
union cpuid_0x10_1_edx edx;
u32 ebx, ecx;
cpuid_count(0x00000010, idx, &eax.full, &ebx, &ecx, &edx.full);
r->max_closid = edx.split.cos_max + 1;
r->num_closid = r->max_closid;
r->cbm_len = eax.split.cbm_len + 1;
r->max_cbm = BIT_MASK(eax.split.cbm_len + 1) - 1;
r->enabled = true;
}
and and the call site:
if (boot_cpu_has(X86_FEATURE_CAT_L3)) {
rdt_get_config(1, &rdt_resources_all[RDT_RESOURCE_L3]);
if (boot_cpu_has(X86_FEATURE_CDP_L3))
r->cdp_capable = true;
ret = true;
}
if (boot_cpu_has(X86_FEATURE_CAT_L2)) {
rdt_get_config(2, &rdt_resources_all[RDT_RESOURCE_L2]);
ret = true;
}
Hmm?
> + for_each_rdt_resource(r)
> + pr_info("Intel RDT %s allocation %s detected\n", r->name,
> + r->cdp_capable ? " (with CDP)" : "");
This want's curly braces around the loop.
Thanks,
tglx
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | linux.kernel
csiph-web