Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1441783 > unrolled thread
| Started by | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| First post | 2016-07-13 00:10 +0200 |
| Last post | 2016-07-25 20:10 +0200 |
| Articles | 20 on this page of 28 — 8 participants |
Back to article view | Back to linux.kernel
[PATCH 00/32] Enable Intel Resource Allocation in Resource Director Technology "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:10 +0200
[PATCH 06/32] x86/intel_rdt: Hot cpu support for Cache Allocation "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:10 +0200
Re: [PATCH 06/32] x86/intel_rdt: Hot cpu support for Cache Allocation Thomas Gleixner <tglx@linutronix.de> - 2016-07-13 11:30 +0200
Re: [PATCH 06/32] x86/intel_rdt: Hot cpu support for Cache Allocation Shivappa Vikas <vikas.shivappa@linux.intel.com> - 2016-07-21 21:50 +0200
Re: [PATCH 06/32] x86/intel_rdt: Hot cpu support for Cache Allocation David Carrillo-Cisneros <davidcc@google.com> - 2016-07-14 02:50 +0200
RE: [PATCH 06/32] x86/intel_rdt: Hot cpu support for Cache Allocation "Yu, Fenghua" <fenghua.yu@intel.com> - 2016-07-15 01:00 +0200
[PATCH 19/32] sched.h: Add rg_list and rdtgroup in task_struct "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:10 +0200
Re: [PATCH 19/32] sched.h: Add rg_list and rdtgroup in task_struct Thomas Gleixner <tglx@linutronix.de> - 2016-07-13 15:10 +0200
RE: [PATCH 19/32] sched.h: Add rg_list and rdtgroup in task_struct "Yu, Fenghua" <fenghua.yu@intel.com> - 2016-07-13 20:00 +0200
[PATCH 21/32] x86/intel_rdt.h: Header for inter_rdt.c "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:10 +0200
[PATCH 12/32] x86/intel_rdt: Hot cpu update for code data prioritization "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:10 +0200
[PATCH 27/32] x86/intel_rdt_rdtgroup.c: Implement rscctrl file system commands "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:10 +0200
[PATCH 30/32] x86/intel_rdt_rdtgroup.c: Process schemas input from rscctrl interface "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:10 +0200
Re: [PATCH 30/32] x86/intel_rdt_rdtgroup.c: Process schemas input from rscctrl interface David Carrillo-Cisneros <davidcc@google.com> - 2016-07-14 02:50 +0200
RE: [PATCH 30/32] x86/intel_rdt_rdtgroup.c: Process schemas input from rscctrl interface "Yu, Fenghua" <fenghua.yu@intel.com> - 2016-07-14 08:20 +0200
Re: [PATCH 30/32] x86/intel_rdt_rdtgroup.c: Process schemas input from rscctrl interface Thomas Gleixner <tglx@linutronix.de> - 2016-07-14 08:20 +0200
RE: [PATCH 30/32] x86/intel_rdt_rdtgroup.c: Process schemas input from rscctrl interface "Yu, Fenghua" <fenghua.yu@intel.com> - 2016-07-14 08:40 +0200
[PATCH 02/32] x86/intel_rdt: Add support for Cache Allocation detection "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:10 +0200
[PATCH 01/32] x86/intel_rdt: Cache Allocation documentation "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:10 +0200
[PATCH 28/32] x86/intel_rdt_rdtgroup.c: Read and write cpus "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:10 +0200
[PATCH 03/32] x86/intel_rdt: Add Class of service management "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:20 +0200
[PATCH 04/32] x86/intel_rdt: Add L3 cache capacity bitmask management "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:20 +0200
Re: [PATCH 04/32] x86/intel_rdt: Add L3 cache capacity bitmask management Marcelo Tosatti <mtosatti@redhat.com> - 2016-07-22 23:10 +0200
Re: [PATCH 04/32] x86/intel_rdt: Add L3 cache capacity bitmask management "Luck, Tony" <tony.luck@intel.com> - 2016-07-22 23:50 +0200
[PATCH 05/32] x86/intel_rdt: Implement scheduling support for Intel RDT "Fenghua Yu" <fenghua.yu@intel.com> - 2016-07-13 00:20 +0200
Re: [PATCH 05/32] x86/intel_rdt: Implement scheduling support for Intel RDT Nilay Vaish <nilayvaish@gmail.com> - 2016-07-25 18:30 +0200
Re: [PATCH 05/32] x86/intel_rdt: Implement scheduling support for Intel RDT Nilay Vaish <nilayvaish@gmail.com> - 2016-07-25 18:40 +0200
Re: [PATCH 05/32] x86/intel_rdt: Implement scheduling support for Intel RDT "Luck, Tony" <tony.luck@intel.com> - 2016-07-25 20:10 +0200
Page 1 of 2 [1] 2 Next page →
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-13 00:10 +0200 |
| Subject | [PATCH 00/32] Enable Intel Resource Allocation in Resource Director Technology |
| Message-ID | <rUe2l-2xx-3@gated-at.bofh.it> |
From: Fenghua Yu <fenghua.yu@intel.com>
L3 cache allocation allows per task control over which areas of the last
level cache are available for allocation. It is the first resource that
can be controlled as part of Intel Resource Director Technology (RDT).
This patch series creates a framework that will make it easy to add
additional resources (like L2 cache).
See Intel Software Developer manual volume 3, chapter 17 for architectural
details. Also Documentation/x86/intel_rdt.txt and
Documentation/x86/intel_rdt_ui.txt (in parts 0001 & 0013 of this patch
series).
A previous implementation used "cgroups" as the user interface. This was
rejected.
The new interface:
1) Aligns better with the h/w capabilities provided
2) Gives finer control (per thread instead of process)
3) Gives control over kernel threads as well as user threads
4) Allows resource allocation policies to be tied to certain cpus across
all contexts (tglx request)
Note that parts 1-12 are largely unchanged from what was posted last year
except for the removal of cgroup pieces and dynamic CAT/CDP switch.
Fenghua Yu (20):
Documentation, x86: Documentation for Intel resource allocation user
interface
x86/cpufeatures: Get max closid and max cbm len and clean feature
comments and code
cacheinfo: Introduce cache id
Documentation, ABI: Add a document entry for cache id
x86, intel_cacheinfo: Enable cache id in x86
drivers/base/cacheinfo.c: Export some cacheinfo functions for others
to use
sched.h: Add rg_list and rdtgroup in task_struct
magic number for rscctrl file system
x86/intel_rdt.h: Header for inter_rdt.c
x86/intel_rdt_rdtgroup.h: Header for user interface
x86/intel_rdt.c: Extend RDT to per cache and per resources
Task fork and exit for rdtgroup
x86/intel_rdt_rdtgroup.c: User interface for RDT
x86/intel_rdt_rdtgroup.c: Create info directory
x86/intel_rdt_rdtgroup.c: Implement rscctrl file system commands
x86/intel_rdt_rdtgroup.c: Read and write cpus
x86/intel_rdt_rdtgroup.c: Tasks iterator and write
x86/intel_rdt_rdtgroup.c: Process schemas input from rscctrl interface
MAINTAINERS: Add maintainer for Intel RDT resource allocation
x86/Makefile: Build intel_rdt_rdtgroup.c
Vikas Shivappa (12):
x86/intel_rdt: Cache Allocation documentation
x86/intel_rdt: Add support for Cache Allocation detection
x86/intel_rdt: Add Class of service management
x86/intel_rdt: Add L3 cache capacity bitmask management
x86/intel_rdt: Implement scheduling support for Intel RDT
x86/intel_rdt: Hot cpu support for Cache Allocation
x86/intel_rdt: Intel haswell Cache Allocation enumeration
Define CONFIG_INTEL_RDT
x86/intel_rdt: Intel Code Data Prioritization detection
x86/intel_rdt: Adds support to enable Code Data Prioritization
x86/intel_rdt: Class of service and capacity bitmask management for
CDP
x86/intel_rdt: Hot cpu update for code data prioritization
Documentation/ABI/testing/sysfs-devices-system-cpu | 17 +
Documentation/x86/intel_rdt.txt | 109 +
Documentation/x86/intel_rdt_ui.txt | 268 +++
MAINTAINERS | 8 +
arch/x86/Kconfig | 12 +
arch/x86/events/intel/cqm.c | 26 +-
arch/x86/include/asm/cpufeature.h | 2 +
arch/x86/include/asm/cpufeatures.h | 13 +-
arch/x86/include/asm/intel_rdt.h | 139 ++
arch/x86/include/asm/intel_rdt_rdtgroup.h | 229 ++
arch/x86/include/asm/pqr_common.h | 27 +
arch/x86/include/asm/processor.h | 3 +
arch/x86/kernel/cpu/Makefile | 2 +
arch/x86/kernel/cpu/common.c | 19 +
arch/x86/kernel/cpu/intel_cacheinfo.c | 20 +
arch/x86/kernel/cpu/intel_rdt.c | 838 ++++++++
arch/x86/kernel/cpu/intel_rdt_rdtgroup.c | 2230 ++++++++++++++++++++
arch/x86/kernel/process_64.c | 6 +
drivers/base/cacheinfo.c | 7 +-
include/linux/cacheinfo.h | 5 +
include/linux/sched.h | 4 +
include/uapi/linux/magic.h | 2 +
kernel/exit.c | 2 +
kernel/fork.c | 4 +
24 files changed, 3965 insertions(+), 27 deletions(-)
create mode 100644 Documentation/x86/intel_rdt.txt
create mode 100644 Documentation/x86/intel_rdt_ui.txt
create mode 100644 arch/x86/include/asm/intel_rdt.h
create mode 100644 arch/x86/include/asm/intel_rdt_rdtgroup.h
create mode 100644 arch/x86/include/asm/pqr_common.h
create mode 100644 arch/x86/kernel/cpu/intel_rdt.c
create mode 100644 arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
--
2.5.0
[toc] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-13 00:10 +0200 |
| Subject | [PATCH 06/32] x86/intel_rdt: Hot cpu support for Cache Allocation |
| Message-ID | <rUe2p-2xx-109@gated-at.bofh.it> |
| In reply to | #1441783 |
From: Vikas Shivappa <vikas.shivappa@linux.intel.com>
This patch adds hot plug cpu support for Intel Cache allocation. Support
includes updating the cache bitmask MSRs IA32_L3_QOS_n when a new CPU
package comes online or goes offline. The IA32_L3_QOS_n MSRs are one per
Class of service on each CPU package. The new package's MSRs are
synchronized with the values of existing MSRs. Also the software cache
for IA32_PQR_ASSOC MSRs are reset during hot cpu notifications.
Signed-off-by: Vikas Shivappa <vikas.shivappa@linux.intel.com>
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Reviewed-by: Tony Luck <tony.luck@intel.com>
---
arch/x86/kernel/cpu/intel_rdt.c | 76 +++++++++++++++++++++++++++++++++++++++++
1 file changed, 76 insertions(+)
diff --git a/arch/x86/kernel/cpu/intel_rdt.c b/arch/x86/kernel/cpu/intel_rdt.c
index 8379df8..31f8588 100644
--- a/arch/x86/kernel/cpu/intel_rdt.c
+++ b/arch/x86/kernel/cpu/intel_rdt.c
@@ -24,6 +24,7 @@
#include <linux/slab.h>
#include <linux/err.h>
+#include <linux/cpu.h>
#include <linux/sched.h>
#include <asm/pqr_common.h>
#include <asm/intel_rdt.h>
@@ -234,6 +235,75 @@ static inline bool rdt_cpumask_update(int cpu)
return false;
}
+/*
+ * cbm_update_msrs() - Updates all the existing IA32_L3_MASK_n MSRs
+ * which are one per CLOSid on the current package.
+ */
+static void cbm_update_msrs(void *dummy)
+{
+ int maxid = boot_cpu_data.x86_cache_max_closid;
+ struct rdt_remote_data info;
+ unsigned int i;
+
+ for (i = 0; i < maxid; i++) {
+ if (cctable[i].clos_refcnt) {
+ info.msr = CBM_FROM_INDEX(i);
+ info.val = cctable[i].l3_cbm;
+ msr_cpu_update(&info);
+ }
+ }
+}
+
+static inline void intel_rdt_cpu_start(int cpu)
+{
+ struct intel_pqr_state *state = &per_cpu(pqr_state, cpu);
+
+ state->closid = 0;
+ mutex_lock(&rdt_group_mutex);
+ if (rdt_cpumask_update(cpu))
+ smp_call_function_single(cpu, cbm_update_msrs, NULL, 1);
+ mutex_unlock(&rdt_group_mutex);
+}
+
+static void intel_rdt_cpu_exit(unsigned int cpu)
+{
+ int i;
+
+ mutex_lock(&rdt_group_mutex);
+ if (!cpumask_test_and_clear_cpu(cpu, &rdt_cpumask)) {
+ mutex_unlock(&rdt_group_mutex);
+ return;
+ }
+
+ cpumask_and(&tmp_cpumask, topology_core_cpumask(cpu), cpu_online_mask);
+ cpumask_clear_cpu(cpu, &tmp_cpumask);
+ i = cpumask_any(&tmp_cpumask);
+
+ if (i < nr_cpu_ids)
+ cpumask_set_cpu(i, &rdt_cpumask);
+ mutex_unlock(&rdt_group_mutex);
+}
+
+static int intel_rdt_cpu_notifier(struct notifier_block *nb,
+ unsigned long action, void *hcpu)
+{
+ unsigned int cpu = (unsigned long)hcpu;
+
+ switch (action) {
+ case CPU_DOWN_FAILED:
+ case CPU_ONLINE:
+ intel_rdt_cpu_start(cpu);
+ break;
+ case CPU_DOWN_PREPARE:
+ intel_rdt_cpu_exit(cpu);
+ break;
+ default:
+ break;
+ }
+
+ return NOTIFY_OK;
+}
+
static int __init intel_rdt_late_init(void)
{
struct cpuinfo_x86 *c = &boot_cpu_data;
@@ -261,9 +331,15 @@ static int __init intel_rdt_late_init(void)
goto out_err;
}
+ cpu_notifier_register_begin();
+
for_each_online_cpu(i)
rdt_cpumask_update(i);
+ __hotcpu_notifier(intel_rdt_cpu_notifier, 0);
+
+ cpu_notifier_register_done();
+
static_key_slow_inc(&rdt_enable_key);
pr_info("Intel cache allocation enabled\n");
out_err:
--
2.5.0
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-07-13 11:30 +0200 |
| Subject | Re: [PATCH 06/32] x86/intel_rdt: Hot cpu support for Cache Allocation |
| Message-ID | <rUoEp-13x-3@gated-at.bofh.it> |
| In reply to | #1441784 |
On Tue, 12 Jul 2016, Fenghua Yu wrote:
> static int __init intel_rdt_late_init(void)
> {
> struct cpuinfo_x86 *c = &boot_cpu_data;
> @@ -261,9 +331,15 @@ static int __init intel_rdt_late_init(void)
> goto out_err;
> }
>
> + cpu_notifier_register_begin();
> +
> for_each_online_cpu(i)
> rdt_cpumask_update(i);
>
> + __hotcpu_notifier(intel_rdt_cpu_notifier, 0);
CPU hotplug notifiers are phased out. Please use the new state machine
interfaces.
Thanks,
tglx
[toc] | [prev] | [next] | [standalone]
| From | Shivappa Vikas <vikas.shivappa@linux.intel.com> |
|---|---|
| Date | 2016-07-21 21:50 +0200 |
| Subject | Re: [PATCH 06/32] x86/intel_rdt: Hot cpu support for Cache Allocation |
| Message-ID | <rXs8N-3wK-1@gated-at.bofh.it> |
| In reply to | #1442251 |
On Wed, 13 Jul 2016, Thomas Gleixner wrote:
> On Tue, 12 Jul 2016, Fenghua Yu wrote:
>> static int __init intel_rdt_late_init(void)
>> {
>> struct cpuinfo_x86 *c = &boot_cpu_data;
>> @@ -261,9 +331,15 @@ static int __init intel_rdt_late_init(void)
>> goto out_err;
>> }
>>
>> + cpu_notifier_register_begin();
>> +
>> for_each_online_cpu(i)
>> rdt_cpumask_update(i);
>>
>> + __hotcpu_notifier(intel_rdt_cpu_notifier, 0);
>
> CPU hotplug notifiers are phased out. Please use the new state machine
> interfaces.
Ok, I just see the patch for cqm with the new state machine. Also would need to
remove the usage of the static tmp cpumask like we did in cqm.
Thanks,
Vikas
>
> Thanks,
>
> tglx
>
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2016-07-14 02:50 +0200 |
| Subject | Re: [PATCH 06/32] x86/intel_rdt: Hot cpu support for Cache Allocation |
| Message-ID | <rUD0J-2cF-5@gated-at.bofh.it> |
| In reply to | #1441784 |
> +static inline void intel_rdt_cpu_start(int cpu)
> +{
> + struct intel_pqr_state *state = &per_cpu(pqr_state, cpu);
> +
> + state->closid = 0;
> + mutex_lock(&rdt_group_mutex);
> + if (rdt_cpumask_update(cpu))
> + smp_call_function_single(cpu, cbm_update_msrs, NULL, 1);
> + mutex_unlock(&rdt_group_mutex);
what happens if cpu's with a cache_id not available at boot comes online?
[toc] | [prev] | [next] | [standalone]
| From | "Yu, Fenghua" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-15 01:00 +0200 |
| Subject | RE: [PATCH 06/32] x86/intel_rdt: Hot cpu support for Cache Allocation |
| Message-ID | <rUXLP-7gd-3@gated-at.bofh.it> |
| In reply to | #1442968 |
>
> > +static inline void intel_rdt_cpu_start(int cpu) {
> > + struct intel_pqr_state *state = &per_cpu(pqr_state, cpu);
> > +
> > + state->closid = 0;
> > + mutex_lock(&rdt_group_mutex);
> > + if (rdt_cpumask_update(cpu))
> > + smp_call_function_single(cpu, cbm_update_msrs, NULL, 1);
> > + mutex_unlock(&rdt_group_mutex);
>
> what happens if cpu's with a cache_id not available at boot comes online?
For L3, that case happens when a new socket is hot plugged into the platform.
We don't handle that right now because that needs platform support and I don't
have that kind of platform to test.
But maybe I can add that support in code and do a test in a simulated mode.
Basically that will create a new domain for the new cache_id.
Thanks.
-Fenghua
[toc] | [prev] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-13 00:10 +0200 |
| Subject | [PATCH 19/32] sched.h: Add rg_list and rdtgroup in task_struct |
| Message-ID | <rUe2p-2xx-103@gated-at.bofh.it> |
| In reply to | #1441783 |
From: Fenghua Yu <fenghua.yu@intel.com>
rg_list is linked list to connect to other tasks in a rdtgroup.
The point of rdtgroup allows the task to access its own rdtgroup directly.
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Reviewed-by: Tony Luck <tony.luck@intel.com>
---
include/linux/sched.h | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/include/linux/sched.h b/include/linux/sched.h
index 253538f..55adf17 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -1761,6 +1761,10 @@ struct task_struct {
/* cg_list protected by css_set_lock and tsk->alloc_lock */
struct list_head cg_list;
#endif
+#ifdef CONFIG_INTEL_RDT
+ struct list_head rg_list;
+ struct rdtgroup *rdtgroup;
+#endif
#ifdef CONFIG_FUTEX
struct robust_list_head __user *robust_list;
#ifdef CONFIG_COMPAT
--
2.5.0
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-07-13 15:10 +0200 |
| Subject | Re: [PATCH 19/32] sched.h: Add rg_list and rdtgroup in task_struct |
| Message-ID | <rUs5k-3uO-7@gated-at.bofh.it> |
| In reply to | #1441785 |
On Tue, 12 Jul 2016, Fenghua Yu wrote: > From: Fenghua Yu <fenghua.yu@intel.com> > > rg_list is linked list to connect to other tasks in a rdtgroup. Can you please prefix that member proper, e.g. rdt_group_list > The point of rdtgroup allows the task to access its own rdtgroup directly. 'pointer' perhaps? Also please use a proper prefix: rdt_group A proper description might be: rdt_group points to the rdt group to which the task belongs. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | "Yu, Fenghua" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-13 20:00 +0200 |
| Subject | RE: [PATCH 19/32] sched.h: Add rg_list and rdtgroup in task_struct |
| Message-ID | <rUwBY-6kl-23@gated-at.bofh.it> |
| In reply to | #1442428 |
> From: Thomas Gleixner [mailto:tglx@linutronix.de] > Sent: Wednesday, July 13, 2016 5:56 AM > On Tue, 12 Jul 2016, Fenghua Yu wrote: > > From: Fenghua Yu <fenghua.yu@intel.com> > > > > rg_list is linked list to connect to other tasks in a rdtgroup. > > Can you please prefix that member proper, e.g. rdt_group_list There is another similar name in task_struct, which is cg_list for cgroup list. I just follow that name to have a rg_list for rdtgroup list. If you think rdt_group_list is a better name, I sure will change rg_list to rdt_group_list. > > > The point of rdtgroup allows the task to access its own rdtgroup directly. > > 'pointer' perhaps? Also please use a proper prefix: rdt_group > > A proper description might be: > > rdt_group points to the rdt group to which the task belongs. In the patch set, I use the name "rdtgroup" to represent a rdt group. Should I change the name "rdtgroup" to "rdt_group" in all patches Including descriptions and code? Thanks. -Fenghua
[toc] | [prev] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-13 00:10 +0200 |
| Subject | [PATCH 21/32] x86/intel_rdt.h: Header for inter_rdt.c |
| Message-ID | <rUe2p-2xx-105@gated-at.bofh.it> |
| In reply to | #1441783 |
From: Fenghua Yu <fenghua.yu@intel.com>
The header mainly provides functions to call from the user interface
file intel_rdt_rdtgroup.c.
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Reviewed-by: Tony Luck <tony.luck@intel.com>
---
arch/x86/include/asm/intel_rdt.h | 87 +++++++++++++++++++++++++++++++++++++---
1 file changed, 81 insertions(+), 6 deletions(-)
diff --git a/arch/x86/include/asm/intel_rdt.h b/arch/x86/include/asm/intel_rdt.h
index f2cb91d..4c5e0ac 100644
--- a/arch/x86/include/asm/intel_rdt.h
+++ b/arch/x86/include/asm/intel_rdt.h
@@ -3,27 +3,99 @@
#ifdef CONFIG_INTEL_RDT
+#include <linux/seq_file.h>
#include <linux/jump_label.h>
-#define MAX_CBM_LENGTH 32
#define IA32_L3_CBM_BASE 0xc90
-#define CBM_FROM_INDEX(x) (IA32_L3_CBM_BASE + x)
-#define MSR_IA32_PQOS_CFG 0xc81
+#define L3_CBM_FROM_INDEX(x) (IA32_L3_CBM_BASE + x)
+
+#define MSR_IA32_L3_QOS_CFG 0xc81
+
+enum resource_type {
+ RESOURCE_L3 = 0,
+ RESOURCE_NUM = 1,
+};
+
+#define MAX_CACHE_LEAVES 4
+#define MAX_CACHE_DOMAINS 64
+
+DECLARE_PER_CPU_READ_MOSTLY(int, cpu_l3_domain);
+DECLARE_PER_CPU_READ_MOSTLY(struct rdtgroup *, cpu_rdtgroup);
extern struct static_key rdt_enable_key;
void __intel_rdt_sched_in(void *dummy);
+extern bool use_rdtgroup_tasks;
+
+extern bool cdp_enabled;
+
+struct rdt_opts {
+ bool cdp_enabled;
+ bool verbose;
+ bool simulate_cat_l3;
+};
+
+struct cache_domain {
+ cpumask_t shared_cpu_map[MAX_CACHE_DOMAINS];
+ unsigned int max_cache_domains_num;
+ unsigned int level;
+ unsigned int shared_cache_id[MAX_CACHE_DOMAINS];
+};
+
+extern struct rdt_opts rdt_opts;
struct clos_cbm_table {
- unsigned long l3_cbm;
+ unsigned long cbm;
unsigned int clos_refcnt;
};
struct clos_config {
- unsigned long *closmap;
+ unsigned long **closmap;
u32 max_closid;
- u32 closids_used;
};
+struct shared_domain {
+ struct cpumask cpumask;
+ int l3_domain;
+};
+
+#define for_each_cache_domain(domain, start_domain, max_domain) \
+ for (domain = start_domain; domain < max_domain; domain++)
+
+extern struct clos_config cconfig;
+extern struct shared_domain *shared_domain;
+extern int shared_domain_num;
+
+extern struct rdtgroup *root_rdtgrp;
+extern void rdtgroup_fork(struct task_struct *child);
+extern void rdtgroup_post_fork(struct task_struct *child);
+
+extern struct clos_cbm_table **l3_cctable;
+
+extern unsigned int min_bitmask_len;
+extern void msr_cpu_update(void *arg);
+extern inline void closid_get(u32 closid, int domain);
+extern void closid_put(u32 closid, int domain);
+extern void closid_free(u32 closid, int domain, int level);
+extern int closid_alloc(u32 *closid, int domain);
+extern bool cat_l3_enabled;
+extern unsigned int get_domain_num(int level);
+extern struct shared_domain *shared_domain;
+extern int shared_domain_num;
+extern inline int get_dcbm_table_index(int x);
+extern inline int get_icbm_table_index(int x);
+
+extern int get_cache_leaf(int level, int cpu);
+
+extern void cbm_update_l3_msr(void *pindex);
+extern int level_to_leaf(int level);
+
+extern void init_msrs(bool cdpenabled);
+extern bool cat_enabled(int level);
+extern u64 max_cbm(int level);
+extern u32 max_cbm_len(int level);
+
+extern void rdtgroup_exit(struct task_struct *tsk);
+
/*
* intel_rdt_sched_in() - Writes the task's CLOSid to IA32_PQR_MSR
*
@@ -54,6 +126,9 @@ static inline void intel_rdt_sched_in(void)
#else
static inline void intel_rdt_sched_in(void) {}
+static inline void rdtgroup_fork(struct task_struct *child) {}
+static inline void rdtgroup_post_fork(struct task_struct *child) {}
+static inline void rdtgroup_exit(struct task_struct *tsk) {}
#endif
#endif
--
2.5.0
[toc] | [prev] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-13 00:10 +0200 |
| Subject | [PATCH 12/32] x86/intel_rdt: Hot cpu update for code data prioritization |
| Message-ID | <rUe2p-2xx-107@gated-at.bofh.it> |
| In reply to | #1441783 |
From: Vikas Shivappa <vikas.shivappa@linux.intel.com>
Updates hot cpu notification handling for code data prioritization(cdp).
The capacity bitmask(cbm) is global for both data and instruction and we
need to update the new online package with all the cbms by writing to
the IA32_L3_QOS_n MSRs.
Signed-off-by: Vikas Shivappa <vikas.shivappa@linux.intel.com>
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Reviewed-by: Tony Luck <tony.luck@intel.com>
---
arch/x86/kernel/cpu/intel_rdt.c | 27 +++++++++++++++++++++------
1 file changed, 21 insertions(+), 6 deletions(-)
diff --git a/arch/x86/kernel/cpu/intel_rdt.c b/arch/x86/kernel/cpu/intel_rdt.c
index 7a03671..057aef1 100644
--- a/arch/x86/kernel/cpu/intel_rdt.c
+++ b/arch/x86/kernel/cpu/intel_rdt.c
@@ -330,6 +330,26 @@ static inline bool rdt_cpumask_update(int cpu)
return false;
}
+static void cbm_update_msr(u32 index)
+{
+ struct rdt_remote_data info;
+ int dindex;
+
+ dindex = DCBM_TABLE_INDEX(index);
+ if (cctable[dindex].clos_refcnt) {
+
+ info.msr = CBM_FROM_INDEX(dindex);
+ info.val = cctable[dindex].l3_cbm;
+ msr_cpu_update((void *) &info);
+
+ if (cdp_enabled) {
+ info.msr = __ICBM_MSR_INDEX(index);
+ info.val = cctable[dindex + 1].l3_cbm;
+ msr_cpu_update((void *) &info);
+ }
+ }
+}
+
/*
* cbm_update_msrs() - Updates all the existing IA32_L3_MASK_n MSRs
* which are one per CLOSid on the current package.
@@ -337,15 +357,10 @@ static inline bool rdt_cpumask_update(int cpu)
static void cbm_update_msrs(void *dummy)
{
int maxid = cconfig.max_closid;
- struct rdt_remote_data info;
unsigned int i;
for (i = 0; i < maxid; i++) {
- if (cctable[i].clos_refcnt) {
- info.msr = CBM_FROM_INDEX(i);
- info.val = cctable[i].l3_cbm;
- msr_cpu_update((void *) &info);
- }
+ cbm_update_msr(i);
}
}
--
2.5.0
[toc] | [prev] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-13 00:10 +0200 |
| Subject | [PATCH 27/32] x86/intel_rdt_rdtgroup.c: Implement rscctrl file system commands |
| Message-ID | <rUe2p-2xx-119@gated-at.bofh.it> |
| In reply to | #1441783 |
From: Fenghua Yu <fenghua.yu@intel.com>
Four basic file system commands are implement for rscctrl.
mount, umount, mkdir, and rmdir.
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Reviewed-by: Tony Luck <tony.luck@intel.com>
---
arch/x86/kernel/cpu/intel_rdt_rdtgroup.c | 237 +++++++++++++++++++++++++++++++
1 file changed, 237 insertions(+)
diff --git a/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c b/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
index b2140a8..91ea3509 100644
--- a/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
+++ b/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
@@ -856,6 +856,139 @@ static void rdtgroup_idr_remove(struct idr *idr, int id)
spin_unlock_bh(&rdtgroup_idr_lock);
}
+
+static int rdtgroup_mkdir(struct kernfs_node *parent_kn, const char *name,
+ umode_t mode)
+{
+ struct rdtgroup *parent, *rdtgrp;
+ struct rdtgroup_root *root;
+ struct kernfs_node *kn;
+ int level, ret;
+
+ if (parent_kn != root_rdtgrp->kn)
+ return -EPERM;
+
+ /* Do not accept '\n' to avoid unparsable situation.
+ */
+ if (strchr(name, '\n'))
+ return -EINVAL;
+
+ parent = rdtgroup_kn_lock_live(parent_kn);
+ if (!parent)
+ return -ENODEV;
+ root = parent->root;
+ level = parent->level + 1;
+
+ /* allocate the rdtgroup and its ID, 0 is reserved for the root */
+ rdtgrp = kzalloc(sizeof(*rdtgrp) +
+ sizeof(rdtgrp->ancestor_ids[0]) * (level + 1),
+ GFP_KERNEL);
+ if (!rdtgrp) {
+ ret = -ENOMEM;
+ goto out_unlock;
+ }
+
+ /*
+ * Temporarily set the pointer to NULL, so idr_find() won't return
+ * a half-baked rdtgroup.
+ */
+ rdtgrp->id = rdtgroup_idr_alloc(&root->rdtgroup_idr, NULL, 2, 0,
+ GFP_KERNEL);
+ if (rdtgrp->id < 0) {
+ ret = -ENOMEM;
+ goto out_cancel_ref;
+ }
+
+ INIT_LIST_HEAD(&rdtgrp->pset.tasks);
+
+ init_rdtgroup_housekeeping(rdtgrp);
+ cpumask_clear(&rdtgrp->cpu_mask);
+
+ rdtgrp->root = root;
+ rdtgrp->level = level;
+
+ if (test_bit(RDTGRP_CPUSET_CLONE_CHILDREN, &parent->flags))
+ set_bit(RDTGRP_CPUSET_CLONE_CHILDREN, &rdtgrp->flags);
+
+ /* create the directory */
+ kn = kernfs_create_dir(parent->kn, name, mode, rdtgrp);
+ if (IS_ERR(kn)) {
+ ret = PTR_ERR(kn);
+ goto out_free_id;
+ }
+ rdtgrp->kn = kn;
+
+ /*
+ * This extra ref will be put in kernfs_remove() and guarantees
+ * that @rdtgrp->kn is always accessible.
+ */
+ kernfs_get(kn);
+
+ atomic_inc(&root->nr_rdtgrps);
+
+ /*
+ * @rdtgrp is now fully operational. If something fails after this
+ * point, it'll be released via the normal destruction path.
+ */
+ rdtgroup_idr_replace(&root->rdtgroup_idr, rdtgrp, rdtgrp->id);
+
+ ret = rdtgroup_kn_set_ugid(kn);
+ if (ret)
+ goto out_destroy;
+
+ ret = rdtgroup_partition_populate_dir(kn);
+ if (ret)
+ goto out_destroy;
+
+ kernfs_activate(kn);
+
+ list_add_tail(&rdtgrp->rdtgroup_list, &rdtgroup_lists);
+ /* Generate default schema for rdtgrp. */
+ ret = get_default_resources(rdtgrp);
+ if (ret)
+ goto out_destroy;
+
+ ret = 0;
+ goto out_unlock;
+
+out_free_id:
+ rdtgroup_idr_remove(&root->rdtgroup_idr, rdtgrp->id);
+out_cancel_ref:
+ kfree(rdtgrp);
+out_unlock:
+ rdtgroup_kn_unlock(parent_kn);
+ return ret;
+
+out_destroy:
+ rdtgroup_destroy_locked(rdtgrp);
+ goto out_unlock;
+}
+
+static int rdtgroup_rmdir(struct kernfs_node *kn)
+{
+ struct rdtgroup *rdtgrp;
+ int cpu;
+ int ret = 0;
+
+ rdtgrp = rdtgroup_kn_lock_live(kn);
+ if (!rdtgrp)
+ return -ENODEV;
+
+ if (!list_empty(&rdtgrp->pset.tasks)) {
+ ret = -EBUSY;
+ goto out;
+ }
+
+ for_each_cpu(cpu, &rdtgrp->cpu_mask)
+ per_cpu(cpu_rdtgroup, cpu) = 0;
+
+ ret = rdtgroup_destroy_locked(rdtgrp);
+
+out:
+ rdtgroup_kn_unlock(kn);
+ return ret;
+}
+
static int
rdtgroup_move_task_all(struct rdtgroup *src_rdtgrp, struct rdtgroup *dst_rdtgrp)
{
@@ -957,6 +1090,11 @@ static void release_root_closid(void)
}
}
+static struct kernfs_syscall_ops rdtgroup_kf_syscall_ops = {
+ .mkdir = rdtgroup_mkdir,
+ .rmdir = rdtgroup_rmdir,
+};
+
static void setup_task_rg_lists(struct rdtgroup *rdtgrp, bool enable)
{
struct task_struct *p, *g;
@@ -1009,6 +1147,105 @@ static void setup_task_rg_lists(struct rdtgroup *rdtgrp, bool enable)
*/
static bool rdtgrp_dfl_root_visible;
+bool rdtgroup_mounted;
+
+static struct dentry *rdt_mount(struct file_system_type *fs_type,
+ int flags, const char *unused_dev_name,
+ void *data)
+{
+ struct super_block *pinned_sb = NULL;
+ struct rdtgroup_root *root;
+ struct dentry *dentry;
+ int ret;
+ bool new_sb;
+
+ /*
+ * The first time anyone tries to mount a rdtgroup, enable the list
+ * linking tasks and fix up all existing tasks.
+ */
+ if (rdtgroup_mounted)
+ return ERR_PTR(-EBUSY);
+
+ rdt_opts.cdp_enabled = false;
+ rdt_opts.verbose = false;
+ cdp_enabled = false;
+
+ ret = parse_rdtgroupfs_options(data);
+ if (ret)
+ goto out_mount;
+
+ if (rdt_opts.cdp_enabled) {
+ cdp_enabled = true;
+ cconfig.max_closid >>= cdp_enabled;
+ pr_info("CDP is enabled\n");
+ }
+
+ init_msrs(cdp_enabled);
+
+ rdtgrp_dfl_root_visible = true;
+ root = &rdtgrp_dfl_root;
+
+ ret = get_default_resources(&root->rdtgrp);
+ if (ret)
+ return ERR_PTR(-ENOSPC);
+
+out_mount:
+ dentry = kernfs_mount(fs_type, flags, root->kf_root,
+ RDTGROUP_SUPER_MAGIC,
+ &new_sb);
+ if (IS_ERR(dentry) || !new_sb)
+ goto out_unlock;
+
+ /*
+ * If @pinned_sb, we're reusing an existing root and holding an
+ * extra ref on its sb. Mount is complete. Put the extra ref.
+ */
+ if (pinned_sb) {
+ WARN_ON(new_sb);
+ deactivate_super(pinned_sb);
+ }
+
+ setup_task_rg_lists(&root->rdtgrp, true);
+
+ cpumask_clear(&root->rdtgrp.cpu_mask);
+ rdtgroup_mounted = true;
+
+ return dentry;
+
+out_unlock:
+ return ERR_PTR(ret);
+}
+
+static void rdt_kill_sb(struct super_block *sb)
+{
+ int ret;
+
+ mutex_lock(&rdtgroup_mutex);
+
+ ret = rmdir_all_sub();
+ if (ret)
+ goto out_unlock;
+
+ setup_task_rg_lists(root_rdtgrp, false);
+ release_root_closid();
+ root_rdtgrp->resource.valid = false;
+
+ /* Restore max_closid to original value. */
+ cconfig.max_closid <<= cdp_enabled;
+
+ kernfs_kill_sb(sb);
+ rdtgroup_mounted = false;
+out_unlock:
+
+ mutex_unlock(&rdtgroup_mutex);
+}
+
+static struct file_system_type rdt_fs_type = {
+ .name = "rscctrl",
+ .mount = rdt_mount,
+ .kill_sb = rdt_kill_sb,
+};
+
static ssize_t rdtgroup_file_write(struct kernfs_open_file *of, char *buf,
size_t nbytes, loff_t off)
{
--
2.5.0
[toc] | [prev] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-13 00:10 +0200 |
| Subject | [PATCH 30/32] x86/intel_rdt_rdtgroup.c: Process schemas input from rscctrl interface |
| Message-ID | <rUe2p-2xx-111@gated-at.bofh.it> |
| In reply to | #1441783 |
From: Fenghua Yu <fenghua.yu@intel.com>
There is one "schemas" file in each rdtgroup directory. User can input
schemas in the file to control how to allocate resources.
The input schemas first needs to pass validation. If there is no syntax
issue, kernel digests the input schemas and find CLOSID for each
domain for each resource.
A shared domain covers a few different resource domains which share
the same CLOSID. Kernel will find a CLOSID in each shared domain. If
an existing CLOSID and its CBMs match input schemas, the CLOSID is
shared by this rdtgroup. Otherwise, kernel tries to alloc a new
CLOSID for this rdtgroup. If a new CLOSID is available, update QoS MASK
MSRs. If no more CLOSID is available, kernel report ENODEV to user.
A shared domain is in preparation for multiple resources (like L2)
that will be added very soon.
User can read the schemas saved in the file.
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Reviewed-by: Tony Luck <tony.luck@intel.com>
---
arch/x86/kernel/cpu/intel_rdt_rdtgroup.c | 673 +++++++++++++++++++++++++++++++
1 file changed, 673 insertions(+)
diff --git a/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c b/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
index e6e8757..bb85995 100644
--- a/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
+++ b/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
@@ -817,6 +817,679 @@ end:
return nbytes;
}
+static int get_res_type(char **res, enum resource_type *res_type)
+{
+ char *tok;
+
+ tok = strsep(res, ":");
+ if (tok == NULL)
+ return -EINVAL;
+
+ if (!strcmp(tok, "L3")) {
+ *res_type = RESOURCE_L3;
+ return 0;
+ }
+
+ return -EINVAL;
+}
+
+static int divide_resources(char *buf, char *resources[RESOURCE_NUM])
+{
+ char *tok;
+ unsigned int resource_num = 0;
+ int ret = 0;
+ char *res;
+ char *res_block;
+ size_t size;
+ enum resource_type res_type;
+
+ size = strlen(buf) + 1;
+ res = kzalloc(size, GFP_KERNEL);
+ if (!res) {
+ ret = -ENOSPC;
+ goto out;
+ }
+
+ while ((tok = strsep(&buf, "\n")) != NULL) {
+ if (strlen(tok) == 0)
+ break;
+ if (resource_num++ >= 1) {
+ pr_info("More than one line of resource input!\n");
+ ret = -EINVAL;
+ goto out;
+ }
+ strcpy(res, tok);
+ }
+
+ res_block = res;
+ ret = get_res_type(&res_block, &res_type);
+ if (ret) {
+ pr_info("Unknown resource type!");
+ goto out;
+ }
+
+ if (res_type == RESOURCE_L3 && cat_enabled(CACHE_LEVEL3)) {
+ strcpy(resources[RESOURCE_L3], res_block);
+ } else {
+ pr_info("Invalid resource type!");
+ goto out;
+ }
+
+ ret = 0;
+
+out:
+ kfree(res);
+ return ret;
+}
+
+static bool cbm_validate(unsigned long var, int level)
+{
+ u32 maxcbmlen = max_cbm_len(level);
+ unsigned long first_bit, zero_bit;
+
+ if (bitmap_weight(&var, maxcbmlen) < min_bitmask_len)
+ return false;
+
+ if (var & ~max_cbm(level))
+ return false;
+
+ first_bit = find_first_bit(&var, maxcbmlen);
+ zero_bit = find_next_zero_bit(&var, maxcbmlen, first_bit);
+
+ if (find_next_bit(&var, maxcbmlen, zero_bit) < maxcbmlen)
+ return false;
+
+ return true;
+}
+
+static int get_input_cbm(char *tok, struct cache_resource *l,
+ int input_domain_num, int level)
+{
+ int ret;
+
+ if (!cdp_enabled) {
+ if (tok == NULL)
+ return -EINVAL;
+
+ ret = kstrtoul(tok, 16,
+ (unsigned long *)&l->cbm[input_domain_num]);
+ if (ret)
+ return ret;
+
+ if (!cbm_validate(l->cbm[input_domain_num], level))
+ return -EINVAL;
+ } else {
+ char *input_cbm1_str;
+
+ input_cbm1_str = strsep(&tok, ",");
+ if (input_cbm1_str == NULL || tok == NULL)
+ return -EINVAL;
+
+ ret = kstrtoul(input_cbm1_str, 16,
+ (unsigned long *)&l->cbm[input_domain_num]);
+ if (ret)
+ return ret;
+
+ if (!cbm_validate(l->cbm[input_domain_num], level))
+ return -EINVAL;
+
+ ret = kstrtoul(tok, 16,
+ (unsigned long *)&l->cbm2[input_domain_num]);
+ if (ret)
+ return ret;
+
+ if (!cbm_validate(l->cbm2[input_domain_num], level))
+ return -EINVAL;
+ }
+
+ return 0;
+}
+
+
+static int get_cache_schema(char *buf, struct cache_resource *l, int level,
+ struct rdtgroup *rdtgrp)
+{
+ char *tok, *tok_cache_id;
+ int ret;
+ int domain_num;
+ int input_domain_num;
+ int len;
+ unsigned int input_cache_id;
+ unsigned int cid;
+ unsigned int leaf;
+
+ if (!cat_enabled(level) && strcmp(buf, ";")) {
+ pr_info("Disabled resource should have empty schema\n");
+ return -EINVAL;
+ }
+
+ len = strlen(buf);
+ /*
+ * Translate cache id based cbm from one line string with format
+ * "<cache prefix>:<cache id0>=xxxx;<cache id1>=xxxx;..." for
+ * disabled cdp.
+ * Or
+ * "<cache prefix>:<cache id0>=xxxxx,xxxxx;<cache id1>=xxxxx,xxxxx;..."
+ * for enabled cdp.
+ */
+ input_domain_num = 0;
+ while ((tok = strsep(&buf, ";")) != NULL) {
+ tok_cache_id = strsep(&tok, "=");
+ if (tok_cache_id == NULL)
+ goto cache_id_err;
+
+ ret = kstrtouint(tok_cache_id, 16, &input_cache_id);
+ if (ret)
+ goto cache_id_err;
+
+ leaf = level_to_leaf(level);
+ cid = cache_domains[leaf].shared_cache_id[input_domain_num];
+ if (input_cache_id != cid)
+ goto cache_id_err;
+
+ ret = get_input_cbm(tok, l, input_domain_num, level);
+ if (ret)
+ goto cbm_err;
+
+ input_domain_num++;
+ if (input_domain_num > get_domain_num(level)) {
+ pr_info("domain number is more than max %d\n",
+ MAX_CACHE_DOMAINS);
+ return -EINVAL;
+ }
+ }
+
+ domain_num = get_domain_num(level);
+ if (domain_num != input_domain_num) {
+ pr_info("%s input domain number %d doesn't match domain number %d\n",
+ "l3",
+ input_domain_num, domain_num);
+
+ return -EINVAL;
+ }
+
+ return 0;
+
+cache_id_err:
+ pr_info("Invalid cache id in field %d for L%1d\n", input_domain_num,
+ level);
+ return -EINVAL;
+
+cbm_err:
+ pr_info("Invalid cbm in field %d for cache L%d\n",
+ input_domain_num, level);
+ return -EINVAL;
+}
+
+struct resources {
+ struct cache_resource *l3;
+};
+
+static bool cbm_found(struct cache_resource *l, struct rdtgroup *r,
+ int domain, int level)
+{
+ int closid;
+ int l3_domain;
+ u64 cctable_cbm;
+ u64 cbm;
+ int dindex;
+
+ closid = r->resource.closid[domain];
+
+ if (level == CACHE_LEVEL3) {
+ l3_domain = shared_domain[domain].l3_domain;
+ cbm = l->cbm[l3_domain];
+ dindex = get_dcbm_table_index(closid);
+ cctable_cbm = l3_cctable[l3_domain][dindex].cbm;
+ if (cdp_enabled) {
+ u64 icbm;
+ u64 cctable_icbm;
+ int iindex;
+
+ icbm = l->cbm2[l3_domain];
+ iindex = get_icbm_table_index(closid);
+ cctable_icbm = l3_cctable[l3_domain][iindex].cbm;
+
+ return cbm == cctable_cbm && icbm == cctable_icbm;
+ }
+
+ return cbm == cctable_cbm;
+ }
+
+ return false;
+}
+
+enum {
+ CURRENT_CLOSID,
+ REUSED_OWN_CLOSID,
+ REUSED_OTHER_CLOSID,
+ NEW_CLOSID,
+};
+
+/*
+ * Check if the reference counts are all ones in rdtgrp's domain.
+ */
+static bool one_refcnt(struct rdtgroup *rdtgrp, int domain)
+{
+ int refcnt;
+ int closid;
+
+ closid = rdtgrp->resource.closid[domain];
+ if (cat_l3_enabled) {
+ int l3_domain;
+ int dindex;
+
+ l3_domain = shared_domain[domain].l3_domain;
+ dindex = get_dcbm_table_index(closid);
+ refcnt = l3_cctable[l3_domain][dindex].clos_refcnt;
+ if (refcnt != 1)
+ return false;
+
+ if (cdp_enabled) {
+ int iindex;
+
+ iindex = get_icbm_table_index(closid);
+ refcnt = l3_cctable[l3_domain][iindex].clos_refcnt;
+
+ if (refcnt != 1)
+ return false;
+ }
+ }
+
+ return true;
+}
+
+/*
+ * Go through all shared domains. Check if there is an existing closid
+ * in all rdtgroups that matches l3 cbms in the shared
+ * domain. If find one, reuse the closid. Otherwise, allocate a new one.
+ */
+static int get_rdtgroup_resources(struct resources *resources_set,
+ struct rdtgroup *rdtgrp)
+{
+ struct cache_resource *l3;
+ bool l3_cbm_found;
+ struct list_head *l;
+ struct rdtgroup *r;
+ u64 cbm;
+ int rdt_closid[MAX_CACHE_DOMAINS];
+ int rdt_closid_type[MAX_CACHE_DOMAINS];
+ int domain;
+ int closid;
+ int ret;
+
+ l3 = resources_set->l3;
+ memcpy(rdt_closid, rdtgrp->resource.closid,
+ shared_domain_num * sizeof(int));
+ for (domain = 0; domain < shared_domain_num; domain++) {
+ if (rdtgrp->resource.valid) {
+ /*
+ * If current rdtgrp is the only user of cbms in
+ * this domain, will replace the cbms with the input
+ * cbms and reuse its own closid.
+ */
+ if (one_refcnt(rdtgrp, domain)) {
+ closid = rdtgrp->resource.closid[domain];
+ rdt_closid[domain] = closid;
+ rdt_closid_type[domain] = REUSED_OWN_CLOSID;
+ continue;
+ }
+
+ l3_cbm_found = true;
+
+ if (cat_l3_enabled)
+ l3_cbm_found = cbm_found(l3, rdtgrp, domain,
+ CACHE_LEVEL3);
+
+ /*
+ * If the cbms in this shared domain are already
+ * existing in current rdtgrp, record the closid
+ * and its type.
+ */
+ if (l3_cbm_found) {
+ closid = rdtgrp->resource.closid[domain];
+ rdt_closid[domain] = closid;
+ rdt_closid_type[domain] = CURRENT_CLOSID;
+ continue;
+ }
+ }
+
+ /*
+ * If the cbms are not found in this rdtgrp, search other
+ * rdtgroups and see if there are matched cbms.
+ */
+ l3_cbm_found = cat_l3_enabled ? false : true;
+ list_for_each(l, &rdtgroup_lists) {
+ r = list_entry(l, struct rdtgroup, rdtgroup_list);
+ if (r == rdtgrp || !r->resource.valid)
+ continue;
+
+ if (cat_l3_enabled)
+ l3_cbm_found = cbm_found(l3, r, domain,
+ CACHE_LEVEL3);
+
+ if (l3_cbm_found) {
+ /* Get the closid that matches l3 cbms.*/
+ closid = r->resource.closid[domain];
+ rdt_closid[domain] = closid;
+ rdt_closid_type[domain] = REUSED_OTHER_CLOSID;
+ break;
+ }
+ }
+
+ if (!l3_cbm_found) {
+ /*
+ * If no existing closid is found, allocate
+ * a new one.
+ */
+ ret = closid_alloc(&closid, domain);
+ if (ret)
+ goto err;
+ rdt_closid[domain] = closid;
+ rdt_closid_type[domain] = NEW_CLOSID;
+ }
+ }
+
+ /*
+ * Now all closid are ready in rdt_closid. Update rdtgrp's closid.
+ */
+ for_each_cache_domain(domain, 0, shared_domain_num) {
+ /*
+ * Nothing is changed if the same closid and same cbms were
+ * found in this rdtgrp's domain.
+ */
+ if (rdt_closid_type[domain] == CURRENT_CLOSID)
+ continue;
+
+ /*
+ * Put rdtgroup closid. No need to put the closid if we
+ * just change cbms and keep the closid (REUSED_OWN_CLOSID).
+ */
+ if (rdtgrp->resource.valid &&
+ rdt_closid_type[domain] != REUSED_OWN_CLOSID) {
+ /* Put old closid in this rdtgrp's domain if valid. */
+ closid = rdtgrp->resource.closid[domain];
+ closid_put(closid, domain);
+ }
+
+ /*
+ * Replace the closid in this rdtgrp's domain with saved
+ * closid that was newly allocted (NEW_CLOSID), or found in
+ * another rdtgroup's domains (REUSED_CLOSID), or found in
+ * this rdtgrp (REUSED_OWN_CLOSID).
+ */
+ closid = rdt_closid[domain];
+ rdtgrp->resource.closid[domain] = closid;
+
+ /*
+ * Get the reused other rdtgroup's closid. No need to get the
+ * closid newly allocated (NEW_CLOSID) because it's been
+ * already got in closid_alloc(). And no need to get the closid
+ * for resued own closid (REUSED_OWN_CLOSID).
+ */
+ if (rdt_closid_type[domain] == REUSED_OTHER_CLOSID)
+ closid_get(closid, domain);
+
+ /*
+ * If the closid comes from a newly allocated closid
+ * (NEW_CLOSID), or found in this rdtgrp (REUSED_OWN_CLOSID),
+ * cbms for this closid will be updated in MSRs.
+ */
+ if (rdt_closid_type[domain] == NEW_CLOSID ||
+ rdt_closid_type[domain] == REUSED_OWN_CLOSID) {
+ /*
+ * Update cbm in cctable with the newly allocated
+ * closid.
+ */
+ if (cat_l3_enabled) {
+ int cpu;
+ struct cpumask *mask;
+ int dindex;
+ int l3_domain = shared_domain[domain].l3_domain;
+ int leaf = level_to_leaf(CACHE_LEVEL3);
+
+ cbm = l3->cbm[l3_domain];
+ dindex = get_dcbm_table_index(closid);
+ l3_cctable[l3_domain][dindex].cbm = cbm;
+ if (cdp_enabled) {
+ int iindex;
+
+ cbm = l3->cbm2[l3_domain];
+ iindex = get_icbm_table_index(closid);
+ l3_cctable[l3_domain][iindex].cbm = cbm;
+ }
+
+ mask =
+ &cache_domains[leaf].shared_cpu_map[l3_domain];
+
+ cpu = cpumask_first(mask);
+ smp_call_function_single(cpu, cbm_update_l3_msr,
+ &closid, 1);
+ }
+ }
+ }
+
+ rdtgrp->resource.valid = true;
+
+ return 0;
+err:
+ /* Free previously allocated closid. */
+ for_each_cache_domain(domain, 0, shared_domain_num) {
+ if (rdt_closid_type[domain] != NEW_CLOSID)
+ continue;
+
+ closid_put(rdt_closid[domain], domain);
+
+ }
+
+ return ret;
+}
+
+static void init_cache_resource(struct cache_resource *l)
+{
+ l->cbm = NULL;
+ l->cbm2 = NULL;
+ l->closid = NULL;
+ l->refcnt = NULL;
+}
+
+static void free_cache_resource(struct cache_resource *l)
+{
+ kfree(l->cbm);
+ kfree(l->cbm2);
+ kfree(l->closid);
+ kfree(l->refcnt);
+}
+
+static int alloc_cache_resource(struct cache_resource *l, int level)
+{
+ int domain_num = get_domain_num(level);
+
+ l->cbm = kcalloc(domain_num, sizeof(*l->cbm), GFP_KERNEL);
+ l->cbm2 = kcalloc(domain_num, sizeof(*l->cbm2), GFP_KERNEL);
+ l->closid = kcalloc(domain_num, sizeof(*l->closid), GFP_KERNEL);
+ l->refcnt = kcalloc(domain_num, sizeof(*l->refcnt), GFP_KERNEL);
+ if (l->cbm && l->cbm2 && l->closid && l->refcnt)
+ return 0;
+
+ return -ENOMEM;
+}
+
+/*
+ * This function digests schemas given in text buf. If the schemas are in
+ * right format and there is enough closid, input the schemas in rdtgrp
+ * and update resource cctables.
+ *
+ * Inputs:
+ * buf: string buffer containing schemas
+ * rdtgrp: current rdtgroup holding schemas.
+ *
+ * Return:
+ * 0 on success or error code.
+ */
+static int get_resources(char *buf, struct rdtgroup *rdtgrp)
+{
+ char *resources[RESOURCE_NUM];
+ struct cache_resource l3;
+ struct resources resources_set;
+ int ret;
+ char *resources_block;
+ int i;
+ int size = strlen(buf) + 1;
+
+ resources_block = kcalloc(RESOURCE_NUM, size, GFP_KERNEL);
+ if (!resources_block)
+ return -ENOMEM;
+
+ for (i = 0; i < RESOURCE_NUM; i++)
+ resources[i] = (char *)(resources_block + i * size);
+
+ ret = divide_resources(buf, resources);
+ if (ret) {
+ kfree(resources_block);
+ return -EINVAL;
+ }
+
+ init_cache_resource(&l3);
+
+ if (cat_l3_enabled) {
+ ret = alloc_cache_resource(&l3, CACHE_LEVEL3);
+ if (ret)
+ goto out;
+
+ ret = get_cache_schema(resources[RESOURCE_L3], &l3,
+ CACHE_LEVEL3, rdtgrp);
+ if (ret)
+ goto out;
+
+ resources_set.l3 = &l3;
+ } else
+ resources_set.l3 = NULL;
+
+ ret = get_rdtgroup_resources(&resources_set, rdtgrp);
+
+out:
+ kfree(resources_block);
+ free_cache_resource(&l3);
+
+ return ret;
+}
+
+static void gen_cache_prefix(char *buf, int level)
+{
+ sprintf(buf, "L%1d:", level == CACHE_LEVEL3 ? 3 : 2);
+}
+
+static int get_cache_id(int domain, int level)
+{
+ return cache_domains[level_to_leaf(level)].shared_cache_id[domain];
+}
+
+static void gen_cache_buf(char *buf, int level)
+{
+ int domain;
+ char buf1[1024];
+ int domain_num;
+ u64 val;
+
+ gen_cache_prefix(buf, level);
+
+ domain_num = get_domain_num(level);
+
+ val = max_cbm(level);
+
+ for (domain = 0; domain < domain_num; domain++) {
+ sprintf(buf1, "%d=%lx", get_cache_id(domain, level),
+ (unsigned long)val);
+ strcat(buf, buf1);
+ if (cdp_enabled) {
+ sprintf(buf1, ",%lx", (unsigned long)val);
+ strcat(buf, buf1);
+ }
+ if (domain < domain_num - 1)
+ sprintf(buf1, ";");
+ else
+ sprintf(buf1, "\n");
+ strcat(buf, buf1);
+ }
+}
+
+/*
+ * Set up schemas in root rdtgroup. All schemas in all resources are default
+ * values (all 1's) for all domains.
+ *
+ * Input: root rdtgroup.
+ * Return: 0: successful
+ * non-0: error code
+ */
+static int get_default_resources(struct rdtgroup *rdtgrp)
+{
+ char schema[1024];
+ int ret = 0;
+
+ strcpy(rdtgrp->schema, "");
+
+ if (cat_enabled(CACHE_LEVEL3)) {
+ gen_cache_buf(schema, CACHE_LEVEL3);
+
+ if (strlen(schema)) {
+ char buf[1024];
+
+ strcpy(buf, schema);
+ ret = get_resources(buf, rdtgrp);
+ if (ret)
+ return ret;
+ }
+ strcat(rdtgrp->schema, schema);
+ }
+
+ return ret;
+}
+
+static ssize_t rdtgroup_schemas_write(struct kernfs_open_file *of,
+ char *buf, size_t nbytes, loff_t off)
+{
+ int ret = 0;
+ struct rdtgroup *rdtgrp;
+ char *schema;
+
+ rdtgrp = rdtgroup_kn_lock_live(of->kn);
+ if (!rdtgrp)
+ return -ENODEV;
+
+ schema = kzalloc(sizeof(char) * strlen(buf) + 1, GFP_KERNEL);
+ if (!schema) {
+ ret = -ENOMEM;
+ goto out_unlock;
+ }
+
+ memcpy(schema, buf, strlen(buf) + 1);
+
+ ret = get_resources(buf, rdtgrp);
+ if (ret)
+ goto out;
+
+ memcpy(rdtgrp->schema, schema, strlen(schema) + 1);
+
+out:
+ kfree(schema);
+
+out_unlock:
+ rdtgroup_kn_unlock(of->kn);
+ return nbytes;
+}
+
+static int rdtgroup_schemas_show(struct seq_file *s, void *v)
+{
+ struct kernfs_open_file *of = s->private;
+ struct rdtgroup *rdtgrp;
+
+ rdtgrp = rdtgroup_kn_lock_live(of->kn);
+ seq_printf(s, "%s", rdtgrp->schema);
+ rdtgroup_kn_unlock(of->kn);
+ return 0;
+}
+
static void show_rdt_tasks(struct list_head *tasks, struct seq_file *s)
{
struct list_head *pos;
--
2.5.0
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2016-07-14 02:50 +0200 |
| Subject | Re: [PATCH 30/32] x86/intel_rdt_rdtgroup.c: Process schemas input from rscctrl interface |
| Message-ID | <rUD0J-2cF-7@gated-at.bofh.it> |
| In reply to | #1441789 |
> +static int get_res_type(char **res, enum resource_type *res_type)
> +{
> + char *tok;
> +
> + tok = strsep(res, ":");
> + if (tok == NULL)
> + return -EINVAL;
> +
> + if (!strcmp(tok, "L3")) {
Maybe use strstrip to allow a more readable input ? i.e. "L3 : <schema> "
> + *res_type = RESOURCE_L3;
> + return 0;
> + }
> +
> + return -EINVAL;
> +}
> +
> +static int divide_resources(char *buf, char *resources[RESOURCE_NUM])
> +{
> + char *tok;
> + unsigned int resource_num = 0;
> + int ret = 0;
> + char *res;
> + char *res_block;
> + size_t size;
> + enum resource_type res_type;
> +
> + size = strlen(buf) + 1;
> + res = kzalloc(size, GFP_KERNEL);
> + if (!res) {
> + ret = -ENOSPC;
-ENOMEM?
> +
> + res_block = res;
> + ret = get_res_type(&res_block, &res_type);
> + if (ret) {
> + pr_info("Unknown resource type!");
> + goto out;
> + }
does this work if res_block doesn't have ":"? don't you need to check res_block?
> +static int get_cache_schema(char *buf, struct cache_resource *l, int level,
> + struct rdtgroup *rdtgrp)
> +{
> + char *tok, *tok_cache_id;
> + int ret;
> + int domain_num;
> + int input_domain_num;
> + int len;
> + unsigned int input_cache_id;
> + unsigned int cid;
> + unsigned int leaf;
> +
> + if (!cat_enabled(level) && strcmp(buf, ";")) {
> + pr_info("Disabled resource should have empty schema\n");
> + return -EINVAL;
> + }
> +
> + len = strlen(buf);
> + /*
> + * Translate cache id based cbm from one line string with format
> + * "<cache prefix>:<cache id0>=xxxx;<cache id1>=xxxx;..." for
> + * disabled cdp.
> + * Or
> + * "<cache prefix>:<cache id0>=xxxxx,xxxxx;<cache id1>=xxxxx,xxxxx;..."
> + * for enabled cdp.
> + */
> + input_domain_num = 0;
> + while ((tok = strsep(&buf, ";")) != NULL) {
> + tok_cache_id = strsep(&tok, "=");
> + if (tok_cache_id == NULL)
> + goto cache_id_err;
what if no "=" ? , also would be nice to allow spaces around "=" .
> +
> + ret = kstrtouint(tok_cache_id, 16, &input_cache_id);
> + if (ret)
> + goto cache_id_err;
> +
> + leaf = level_to_leaf(level);
why is this in the loop?
> + cid = cache_domains[leaf].shared_cache_id[input_domain_num];
> + if (input_cache_id != cid)
> + goto cache_id_err;
so schemata must be present for all cache_id's and sorted in
increasing order of cache_id? what's the point of having the cache_id#
then?
> +
> +/*
> + * Check if the reference counts are all ones in rdtgrp's domain.
> + */
> +static bool one_refcnt(struct rdtgroup *rdtgrp, int domain)
> +{
> + int refcnt;
> + int closid;
> +
> + closid = rdtgrp->resource.closid[domain];
> + if (cat_l3_enabled) {
if cat_l3_enabled == false, then reference counts are always one?
> + * Go through all shared domains. Check if there is an existing closid
> + * in all rdtgroups that matches l3 cbms in the shared
> + * domain. If find one, reuse the closid. Otherwise, allocate a new one.
> + */
> +static int get_rdtgroup_resources(struct resources *resources_set,
> + struct rdtgroup *rdtgrp)
> +{
> + struct cache_resource *l3;
> + bool l3_cbm_found;
> + struct list_head *l;
> + struct rdtgroup *r;
> + u64 cbm;
> + int rdt_closid[MAX_CACHE_DOMAINS];
> + int rdt_closid_type[MAX_CACHE_DOMAINS];
> + int domain;
> + int closid;
> + int ret;
> +
> + l3 = resources_set->l3;
l3 is NULL if cat_l3_enabled == false but it seems like it may be used
later even though.
> + memcpy(rdt_closid, rdtgrp->resource.closid,
> + shared_domain_num * sizeof(int));
> + for (domain = 0; domain < shared_domain_num; domain++) {
> + if (rdtgrp->resource.valid) {
> + /*
> + * If current rdtgrp is the only user of cbms in
> + * this domain, will replace the cbms with the input
> + * cbms and reuse its own closid.
> + */
> + if (one_refcnt(rdtgrp, domain)) {
> + closid = rdtgrp->resource.closid[domain];
> + rdt_closid[domain] = closid;
> + rdt_closid_type[domain] = REUSED_OWN_CLOSID;
> + continue;
> + }
> +
> + l3_cbm_found = true;
> +
> + if (cat_l3_enabled)
> + l3_cbm_found = cbm_found(l3, rdtgrp, domain,
> + CACHE_LEVEL3);
> +
> + /*
> + * If the cbms in this shared domain are already
> + * existing in current rdtgrp, record the closid
> + * and its type.
> + */
> + if (l3_cbm_found) {
> + closid = rdtgrp->resource.closid[domain];
> + rdt_closid[domain] = closid;
> + rdt_closid_type[domain] = CURRENT_CLOSID;
a new l3 resource will be created if cat_l3_enabled is false.
> +static void init_cache_resource(struct cache_resource *l)
> +{
> + l->cbm = NULL;
> + l->cbm2 = NULL;
is cbm2 the data bitmask for when CDP is enabled? if so, a more
descriptive name may help.
> + l->closid = NULL;
> + l->refcnt = NULL;
> +}
> +
> +static void free_cache_resource(struct cache_resource *l)
> +{
> + kfree(l->cbm);
> + kfree(l->cbm2);
> + kfree(l->closid);
> + kfree(l->refcnt);
this function is used to clean up alloc_cache_resource in the error
path of get_resources where it's not necessarily true that all of l's
members were allocated.
[toc] | [prev] | [next] | [standalone]
| From | "Yu, Fenghua" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-14 08:20 +0200 |
| Subject | RE: [PATCH 30/32] x86/intel_rdt_rdtgroup.c: Process schemas input from rscctrl interface |
| Message-ID | <rUIa6-5Pz-13@gated-at.bofh.it> |
| In reply to | #1442969 |
> From: Thomas Gleixner [mailto:tglx@linutronix.de]
> Sent: Wednesday, July 13, 2016 11:11 PM
> On Wed, 13 Jul 2016, David Carrillo-Cisneros wrote:
> > > +static void free_cache_resource(struct cache_resource *l) {
> > > + kfree(l->cbm);
> > > + kfree(l->cbm2);
> > > + kfree(l->closid);
> > > + kfree(l->refcnt);
> >
> > this function is used to clean up alloc_cache_resource in the error
> > path of get_resources where it's not necessarily true that all of l's
> > members were allocated.
>
> kfree handles kfree(NULL) nicely.....
Yes, that's right. If I check the pointer before kfree(), checkpatch.pl will
report warning for that and suggest kfree(NULL) is safe and code is
short.
Thanks.
-Fenghua
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-07-14 08:20 +0200 |
| Subject | Re: [PATCH 30/32] x86/intel_rdt_rdtgroup.c: Process schemas input from rscctrl interface |
| Message-ID | <rUIa6-5Pz-11@gated-at.bofh.it> |
| In reply to | #1442969 |
On Wed, 13 Jul 2016, David Carrillo-Cisneros wrote:
> > +static void free_cache_resource(struct cache_resource *l)
> > +{
> > + kfree(l->cbm);
> > + kfree(l->cbm2);
> > + kfree(l->closid);
> > + kfree(l->refcnt);
>
> this function is used to clean up alloc_cache_resource in the error
> path of get_resources where it's not necessarily true that all of l's
> members were allocated.
kfree handles kfree(NULL) nicely.....
Thanks,
tglx
[toc] | [prev] | [next] | [standalone]
| From | "Yu, Fenghua" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-14 08:40 +0200 |
| Subject | RE: [PATCH 30/32] x86/intel_rdt_rdtgroup.c: Process schemas input from rscctrl interface |
| Message-ID | <rUItr-5Xh-5@gated-at.bofh.it> |
| In reply to | #1442969 |
> From: David Carrillo-Cisneros [mailto:davidcc@google.com]
> > +static int get_res_type(char **res, enum resource_type *res_type) {
> > + char *tok;
> > +
> > + tok = strsep(res, ":");
> > + if (tok == NULL)
> > + return -EINVAL;
> > +
> > + if (!strcmp(tok, "L3")) {
>
> Maybe use strstrip to allow a more readable input ? i.e. "L3 : <schema> "
>
> > + *res_type = RESOURCE_L3;
> > + return 0;
> > + }
> > +
> > + return -EINVAL;
> > +}
> > +
> > +static int divide_resources(char *buf, char *resources[RESOURCE_NUM])
> > +{
> > + char *tok;
> > + unsigned int resource_num = 0;
> > + int ret = 0;
> > + char *res;
> > + char *res_block;
> > + size_t size;
> > + enum resource_type res_type;
> > +
> > + size = strlen(buf) + 1;
> > + res = kzalloc(size, GFP_KERNEL);
> > + if (!res) {
> > + ret = -ENOSPC;
>
> -ENOMEM?
Will change to -ENOMEM.
>
> > +
> > + res_block = res;
> > + ret = get_res_type(&res_block, &res_type);
> > + if (ret) {
> > + pr_info("Unknown resource type!");
> > + goto out;
> > + }
>
> does this work if res_block doesn't have ":"? don't you need to check
> res_block?
get_res_type() checks ":" and return -EINVAL if no ":".
>
> > +static int get_cache_schema(char *buf, struct cache_resource *l, int level,
> > + struct rdtgroup *rdtgrp) {
> > + char *tok, *tok_cache_id;
> > + int ret;
> > + int domain_num;
> > + int input_domain_num;
> > + int len;
> > + unsigned int input_cache_id;
> > + unsigned int cid;
> > + unsigned int leaf;
> > +
> > + if (!cat_enabled(level) && strcmp(buf, ";")) {
> > + pr_info("Disabled resource should have empty schema\n");
> > + return -EINVAL;
> > + }
> > +
> > + len = strlen(buf);
> > + /*
> > + * Translate cache id based cbm from one line string with format
> > + * "<cache prefix>:<cache id0>=xxxx;<cache id1>=xxxx;..." for
> > + * disabled cdp.
> > + * Or
> > + * "<cache prefix>:<cache id0>=xxxxx,xxxxx;<cache
> id1>=xxxxx,xxxxx;..."
> > + * for enabled cdp.
> > + */
> > + input_domain_num = 0;
> > + while ((tok = strsep(&buf, ";")) != NULL) {
> > + tok_cache_id = strsep(&tok, "=");
> > + if (tok_cache_id == NULL)
> > + goto cache_id_err;
>
> what if no "=" ? , also would be nice to allow spaces around "=" .
Without "=-", reports id error. Sure I can strip the spaces around "=".
>
> > +
> > + ret = kstrtouint(tok_cache_id, 16, &input_cache_id);
> > + if (ret)
> > + goto cache_id_err;
> > +
> > + leaf = level_to_leaf(level);
>
> why is this in the loop?
Leaf is the cache index number which is contiguous starting from 0. We need to save and get
cache id info from cache index.
<leaf, cache id> uniquely identifies a cache.
Architecturally level can not be used to identify a cache. There could be 2 caches in one level, ie. icache and dcache.
So we need to translate from level
>
> > + cid = cache_domains[leaf].shared_cache_id[input_domain_num];
> > + if (input_cache_id != cid)
> > + goto cache_id_err;
>
> so schemata must be present for all cache_id's and sorted in increasing order
> of cache_id? what's the point of having the cache_id# then?
Cache_id may be not be contiguous. It can not be used directly as array index.
For user interface, user inputs cache_id to identify a cache. Internally kernel uses
domain, which is contiguous and used as index for internally saved cbm. Kernel
interface code does the mapping between cache_id and domain number.
>
> > +
> > +/*
> > + * Check if the reference counts are all ones in rdtgrp's domain.
> > + */
> > +static bool one_refcnt(struct rdtgroup *rdtgrp, int domain) {
> > + int refcnt;
> > + int closid;
> > +
> > + closid = rdtgrp->resource.closid[domain];
> > + if (cat_l3_enabled) {
>
> if cat_l3_enabled == false, then reference counts are always one?
I can change to return false if cat_l3_enabled==false.
>
> > + * Go through all shared domains. Check if there is an existing
> > +closid
> > + * in all rdtgroups that matches l3 cbms in the shared
> > + * domain. If find one, reuse the closid. Otherwise, allocate a new one.
> > + */
> > +static int get_rdtgroup_resources(struct resources *resources_set,
> > + struct rdtgroup *rdtgrp) {
> > + struct cache_resource *l3;
> > + bool l3_cbm_found;
> > + struct list_head *l;
> > + struct rdtgroup *r;
> > + u64 cbm;
> > + int rdt_closid[MAX_CACHE_DOMAINS];
> > + int rdt_closid_type[MAX_CACHE_DOMAINS];
> > + int domain;
> > + int closid;
> > + int ret;
> > +
> > + l3 = resources_set->l3;
>
> l3 is NULL if cat_l3_enabled == false but it seems like it may be used later
> even though.
>
> > + memcpy(rdt_closid, rdtgrp->resource.closid,
> > + shared_domain_num * sizeof(int));
> > + for (domain = 0; domain < shared_domain_num; domain++) {
> > + if (rdtgrp->resource.valid) {
> > + /*
> > + * If current rdtgrp is the only user of cbms in
> > + * this domain, will replace the cbms with the input
> > + * cbms and reuse its own closid.
> > + */
> > + if (one_refcnt(rdtgrp, domain)) {
> > + closid = rdtgrp->resource.closid[domain];
> > + rdt_closid[domain] = closid;
> > + rdt_closid_type[domain] = REUSED_OWN_CLOSID;
> > + continue;
> > + }
> > +
> > + l3_cbm_found = true;
> > +
> > + if (cat_l3_enabled)
> > + l3_cbm_found = cbm_found(l3, rdtgrp, domain,
> > +
> > + CACHE_LEVEL3);
> > +
> > + /*
> > + * If the cbms in this shared domain are already
> > + * existing in current rdtgrp, record the closid
> > + * and its type.
> > + */
> > + if (l3_cbm_found) {
> > + closid = rdtgrp->resource.closid[domain];
> > + rdt_closid[domain] = closid;
> > + rdt_closid_type[domain] =
> > + CURRENT_CLOSID;
>
> a new l3 resource will be created if cat_l3_enabled is false.
This code actually handles both l3 and l2.
L2 patches will be sent out later on top of this patch set.
With this patch set alone, when l3 is disabled, we will not come
to this function because initialization will fail and there is no
rscctrl file system.
>
>
> > +static void init_cache_resource(struct cache_resource *l) {
> > + l->cbm = NULL;
> > + l->cbm2 = NULL;
>
> is cbm2 the data bitmask for when CDP is enabled? if so, a more descriptive
> name may help.
Yes, that's right. It's the second CBM when CDP is enabled. Maybe I change
to a nicer name.
>
> > + l->closid = NULL;
> > + l->refcnt = NULL;
> > +}
> > +
> > +static void free_cache_resource(struct cache_resource *l) {
> > + kfree(l->cbm);
> > + kfree(l->cbm2);
> > + kfree(l->closid);
> > + kfree(l->refcnt);
>
> this function is used to clean up alloc_cache_resource in the error path of
> get_resources where it's not necessarily true that all of l's members were
> allocated.
As Thomas already said, kfree(NULL) is safe. Don't need to check pointer
before kfree() and code is short.
Thanks.
-Fenghua
[toc] | [prev] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-13 00:10 +0200 |
| Subject | [PATCH 02/32] x86/intel_rdt: Add support for Cache Allocation detection |
| Message-ID | <rUe2p-2xx-123@gated-at.bofh.it> |
| In reply to | #1441783 |
From: Vikas Shivappa <vikas.shivappa@linux.intel.com>
This patch includes CPUID enumeration routines for Cache allocation and
new values to track resources to the cpuinfo_x86 structure.
Cache allocation provides a way for the Software (OS/VMM) to restrict
cache allocation to a defined 'subset' of cache which may be overlapping
with other 'subsets'. This feature is used when allocating a line in
cache ie when pulling new data into the cache. The programming of the
hardware is done via programming MSRs (model specific registers).
Signed-off-by: Vikas Shivappa <vikas.shivappa@linux.intel.com>
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Reviewed-by: Tony Luck <tony.luck@intel.com>
---
arch/x86/include/asm/cpufeatures.h | 10 +++++++---
arch/x86/include/asm/processor.h | 3 +++
arch/x86/kernel/cpu/Makefile | 2 ++
arch/x86/kernel/cpu/common.c | 15 ++++++++++++++
arch/x86/kernel/cpu/intel_rdt.c | 40 ++++++++++++++++++++++++++++++++++++++
5 files changed, 67 insertions(+), 3 deletions(-)
create mode 100644 arch/x86/kernel/cpu/intel_rdt.c
diff --git a/arch/x86/include/asm/cpufeatures.h b/arch/x86/include/asm/cpufeatures.h
index 4a41348..667acf3 100644
--- a/arch/x86/include/asm/cpufeatures.h
+++ b/arch/x86/include/asm/cpufeatures.h
@@ -12,7 +12,7 @@
/*
* Defines x86 CPU feature bits
*/
-#define NCAPINTS 18 /* N 32-bit words worth of info */
+#define NCAPINTS 19 /* N 32-bit words worth of info */
#define NBUGINTS 1 /* N 32-bit bug flags */
/*
@@ -220,6 +220,7 @@
#define X86_FEATURE_RTM ( 9*32+11) /* Restricted Transactional Memory */
#define X86_FEATURE_CQM ( 9*32+12) /* Cache QoS Monitoring */
#define X86_FEATURE_MPX ( 9*32+14) /* Memory Protection Extension */
+#define X86_FEATURE_RDT ( 9*32+15) /* Resource Allocation */
#define X86_FEATURE_AVX512F ( 9*32+16) /* AVX-512 Foundation */
#define X86_FEATURE_AVX512DQ ( 9*32+17) /* AVX-512 DQ (Double/Quad granular) Instructions */
#define X86_FEATURE_RDSEED ( 9*32+18) /* The RDSEED instruction */
@@ -284,8 +285,11 @@
/* AMD-defined CPU features, CPUID level 0x80000007 (ebx), word 17 */
#define X86_FEATURE_OVERFLOW_RECOV (17*32+0) /* MCA overflow recovery support */
-#define X86_FEATURE_SUCCOR (17*32+1) /* Uncorrectable error containment and recovery */
-#define X86_FEATURE_SMCA (17*32+3) /* Scalable MCA */
+#define X86_FEATURE_SUCCOR (17*32+ 1) /* Uncorrectable error containment and recovery */
+#define X86_FEATURE_SMCA (17*32+ 3) /* Scalable MCA */
+
+/* Intel-defined CPU features, CPUID level 0x00000010:0 (ebx), word 18 */
+#define X86_FEATURE_CAT_L3 (18*32+ 1) /* Cache Allocation L3 */
/*
* BUG word(s)
diff --git a/arch/x86/include/asm/processor.h b/arch/x86/include/asm/processor.h
index 62c6cc3..598c9bc 100644
--- a/arch/x86/include/asm/processor.h
+++ b/arch/x86/include/asm/processor.h
@@ -119,6 +119,9 @@ struct cpuinfo_x86 {
int x86_cache_occ_scale; /* scale to bytes */
int x86_power;
unsigned long loops_per_jiffy;
+ /* Cache Allocation values: */
+ u16 x86_cache_max_cbm_len;
+ u16 x86_cache_max_closid;
/* cpuid returned max cores value: */
u16 x86_max_cores;
u16 apicid;
diff --git a/arch/x86/kernel/cpu/Makefile b/arch/x86/kernel/cpu/Makefile
index 4a8697f..39b8e6f 100644
--- a/arch/x86/kernel/cpu/Makefile
+++ b/arch/x86/kernel/cpu/Makefile
@@ -42,6 +42,8 @@ obj-$(CONFIG_X86_LOCAL_APIC) += perfctr-watchdog.o
obj-$(CONFIG_HYPERVISOR_GUEST) += vmware.o hypervisor.o mshyperv.o
+obj-$(CONFIG_INTEL_RDT) += intel_rdt.o
+
ifdef CONFIG_X86_FEATURE_NAMES
quiet_cmd_mkcapflags = MKCAP $@
cmd_mkcapflags = $(CONFIG_SHELL) $(srctree)/$(src)/mkcapflags.sh $< $@
diff --git a/arch/x86/kernel/cpu/common.c b/arch/x86/kernel/cpu/common.c
index 0fe6953..42c90cb 100644
--- a/arch/x86/kernel/cpu/common.c
+++ b/arch/x86/kernel/cpu/common.c
@@ -711,6 +711,21 @@ void get_cpu_cap(struct cpuinfo_x86 *c)
}
}
+ /* Additional Intel-defined flags: level 0x00000010 */
+ if (c->cpuid_level >= 0x00000010) {
+ u32 eax, ebx, ecx, edx;
+
+ cpuid_count(0x00000010, 0, &eax, &ebx, &ecx, &edx);
+ c->x86_capability[14] = ebx;
+
+ if (cpu_has(c, X86_FEATURE_CAT_L3)) {
+
+ cpuid_count(0x00000010, 1, &eax, &ebx, &ecx, &edx);
+ c->x86_cache_max_closid = edx + 1;
+ c->x86_cache_max_cbm_len = eax + 1;
+ }
+ }
+
/* AMD-defined flags: level 0x80000001 */
eax = cpuid_eax(0x80000000);
c->extended_cpuid_level = eax;
diff --git a/arch/x86/kernel/cpu/intel_rdt.c b/arch/x86/kernel/cpu/intel_rdt.c
new file mode 100644
index 0000000..f49e970
--- /dev/null
+++ b/arch/x86/kernel/cpu/intel_rdt.c
@@ -0,0 +1,40 @@
+/*
+ * Resource Director Technology(RDT)
+ * - Cache Allocation code.
+ *
+ * Copyright (C) 2014 Intel Corporation
+ *
+ * 2015-05-25 Written by
+ * Vikas Shivappa <vikas.shivappa@intel.com>
+ *
+ * This program is free software; you can redistribute it and/or modify it
+ * under the terms and conditions of the GNU General Public License,
+ * version 2, as published by the Free Software Foundation.
+ *
+ * This program is distributed in the hope it will be useful, but WITHOUT
+ * ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or
+ * FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for
+ * more details.
+ *
+ * More information about RDT be found in the Intel (R) x86 Architecture
+ * Software Developer Manual June 2015, volume 3, section 17.15.
+ */
+
+#define pr_fmt(fmt) KBUILD_MODNAME ": " fmt
+
+#include <linux/slab.h>
+#include <linux/err.h>
+
+static int __init intel_rdt_late_init(void)
+{
+ struct cpuinfo_x86 *c = &boot_cpu_data;
+
+ if (!cpu_has(c, X86_FEATURE_CAT_L3))
+ return -ENODEV;
+
+ pr_info("Intel cache allocation detected\n");
+
+ return 0;
+}
+
+late_initcall(intel_rdt_late_init);
--
2.5.0
[toc] | [prev] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-13 00:10 +0200 |
| Subject | [PATCH 01/32] x86/intel_rdt: Cache Allocation documentation |
| Message-ID | <rUe2p-2xx-121@gated-at.bofh.it> |
| In reply to | #1441783 |
From: Vikas Shivappa <vikas.shivappa@linux.intel.com> Adds a description of Cache allocation technology, overview of kernel framework implementation. The framework has APIs to manage class of service, capacity bitmask(CBM), scheduling support and other architecture specific implementation. The APIs are used to build the rscctrl interface in later patches. Cache allocation is a sub-feature of Resource Director Technology (RDT) or Platform Shared resource control which provides support to control Platform shared resources like L3 cache. Cache Allocation Technology provides a way for the Software (OS/VMM) to restrict cache allocation to a defined 'subset' of cache which may be overlapping with other 'subsets'. This feature is used when allocating a line in cache ie when pulling new data into the cache. The tasks are grouped into CLOS (class of service). OS uses MSR writes to indicate the CLOSid of the thread when scheduling in and to indicate the cache capacity associated with the CLOSid. Currently cache allocation is supported for L3 cache. More information can be found in the Intel SDM June 2015, Volume 3, section 17.16. Signed-off-by: Vikas Shivappa <vikas.shivappa@linux.intel.com> Signed-off-by: Fenghua Yu <fenghua.yu@intel.com> Reviewed-by: Tony Luck <tony.luck@intel.com> --- Documentation/x86/intel_rdt.txt | 109 ++++++++++++++++++++++++++++++++++++++++ 1 file changed, 109 insertions(+) create mode 100644 Documentation/x86/intel_rdt.txt diff --git a/Documentation/x86/intel_rdt.txt b/Documentation/x86/intel_rdt.txt new file mode 100644 index 0000000..05ec819 --- /dev/null +++ b/Documentation/x86/intel_rdt.txt @@ -0,0 +1,109 @@ + Intel RDT + --------- + +Copyright (C) 2014 Intel Corporation +Written by vikas.shivappa@linux.intel.com + +CONTENTS: +========= + +1. Cache Allocation Technology + 1.1 What is RDT and Cache allocation ? + 1.2 Why is Cache allocation needed ? + 1.3 Cache allocation implementation overview + 1.4 Assignment of CBM and CLOS + 1.5 Scheduling and Context Switch + +1. Cache Allocation Technology +=================================== + +1.1 What is RDT and Cache allocation +------------------------------------ + +Cache allocation is a sub-feature of Resource Director Technology (RDT) +Allocation or Platform Shared resource control which provides support to +control Platform shared resources like L3 cache. Currently L3 Cache is +the only resource that is supported in RDT. More information can be +found in the Intel SDM June 2015, Volume 3, section 17.16. + +Cache Allocation Technology provides a way for the Software (OS/VMM) to +restrict cache allocation to a defined 'subset' of cache which may be +overlapping with other 'subsets'. This feature is used when allocating a +line in cache ie when pulling new data into the cache. The programming +of the h/w is done via programming MSRs. + +The different cache subsets are identified by CLOS identifier (class of +service) and each CLOS has a CBM (cache bit mask). The CBM is a +contiguous set of bits which defines the amount of cache resource that +is available for each 'subset'. + +1.2 Why is Cache allocation needed +---------------------------------- + +In todays new processors the number of cores is continuously increasing +especially in large scale usage models where VMs are used like +webservers and datacenters. The number of cores increase the number of +threads or workloads that can simultaneously be run. When +multi-threaded-applications, VMs, workloads run concurrently they +compete for shared resources including L3 cache. + +The architecture also allows dynamically changing these subsets during +runtime to further optimize the performance of the higher priority +application with minimal degradation to the low priority app. +Additionally, resources can be rebalanced for system throughput benefit. + +This technique may be useful in managing large computer server systems +with large L3 cache, in the cloud and container context. Examples may be +large servers running instances of webservers or database servers. In +such complex systems, these subsets can be used for more careful placing +of the available cache resources by a centralized root accessible +interface. + +A specific use case may be to solve the noisy neighbour issue when a app +which is constantly copying data like streaming app is using large +amount of cache which could have otherwise been used by a high priority +computing application. Using the cache allocation feature, the streaming +application can be confined to use a smaller cache and the high priority +application be awarded a larger amount of cache space. + +1.3 Cache allocation implementation Overview +-------------------------------------------- + +Kernel has a new field in the task_struct called 'closid' which +represents the Class of service ID of the task. + +There is a 1:1 CLOSid <-> CBM (capacity bit mask) mapping. A CLOS (Class +of service) is represented by a CLOSid. Each closid would have one CBM +and would just represent one cache 'subset'. The tasks would get to +fill the L3 cache represented by the capacity bit mask or CBM. + +The APIs to manage the closid and CBM can be used to develop user +interfaces. + +1.4 Assignment of CBM, CLOS +--------------------------- + +The framework provides APIs to manage the closid and CBM which can be +used to develop user/kernel mode interfaces. + +1.5 Scheduling and Context Switch +--------------------------------- + +During context switch kernel implements this by writing the CLOSid of +the task to the CPU's IA32_PQR_ASSOC MSR. The MSR is only written when +there is a change in the CLOSid for the CPU in order to minimize the +latency incurred during context switch. + +The following considerations are done for the PQR MSR write so that it +has minimal impact on scheduling hot path: + - This path doesn't exist on any non-intel platforms. + - On Intel platforms, this would not exist by default unless INTEL_RDT + is enabled. + - remains a no-op when INTEL_RDT is enabled and intel hardware does + not support the feature. + - When feature is available, does not do MSR write till the user + starts using the feature *and* assigns a new cache capacity mask. + - per cpu PQR values are cached and the MSR write is only done when + there is a task with different PQR is scheduled on the CPU. Typically + if the task groups are bound to be scheduled on a set of CPUs, the + number of MSR writes is greatly reduced. -- 2.5.0
[toc] | [prev] | [next] | [standalone]
| From | "Fenghua Yu" <fenghua.yu@intel.com> |
|---|---|
| Date | 2016-07-13 00:10 +0200 |
| Subject | [PATCH 28/32] x86/intel_rdt_rdtgroup.c: Read and write cpus |
| Message-ID | <rUe2p-2xx-125@gated-at.bofh.it> |
| In reply to | #1441783 |
From: Fenghua Yu <fenghua.yu@intel.com>
Normally each task is associated with one rdtgroup and we use the schema
for that rdtgroup whenever the task is running. The user can designate
some cpus to always use the same schema, regardless of which task is
running. To do that the user write a cpumask bit string to the "cpus"
file.
A cpu can only be listed in one rdtgroup. If the user specifies a cpu
that is currently assigned to a different rdtgroup, it is removed
from that rdtgroup.
See Documentation/x86/intel_rdt_ui.txt
Signed-off-by: Fenghua Yu <fenghua.yu@intel.com>
Reviewed-by: Tony Luck <tony.luck@intel.com>
---
arch/x86/kernel/cpu/intel_rdt_rdtgroup.c | 54 ++++++++++++++++++++++++++++++++
1 file changed, 54 insertions(+)
diff --git a/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c b/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
index 91ea3509..b5f42f5 100644
--- a/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
+++ b/arch/x86/kernel/cpu/intel_rdt_rdtgroup.c
@@ -767,6 +767,60 @@ void rdtgroup_exit(struct task_struct *tsk)
static struct rdtgroup *rdtgroup_kn_lock_live(struct kernfs_node *kn);
static void rdtgroup_kn_unlock(struct kernfs_node *kn);
+static int rdtgroup_cpus_show(struct seq_file *s, void *v)
+{
+ struct kernfs_open_file *of = s->private;
+ struct rdtgroup *rdtgrp;
+
+ rdtgrp = rdtgroup_kn_lock_live(of->kn);
+ seq_printf(s, "%*pb\n", cpumask_pr_args(&rdtgrp->cpu_mask));
+ rdtgroup_kn_unlock(of->kn);
+
+ return 0;
+}
+
+static ssize_t rdtgroup_cpus_write(struct kernfs_open_file *of,
+ char *buf, size_t nbytes, loff_t off)
+{
+ struct rdtgroup *rdtgrp;
+ unsigned long bitmap[BITS_TO_LONGS(NR_CPUS)];
+ struct cpumask *cpumask;
+ int cpu;
+ struct list_head *l;
+ struct rdtgroup *r;
+
+ if (!buf)
+ return -EINVAL;
+
+ rdtgrp = rdtgroup_kn_lock_live(of->kn);
+ if (!rdtgrp)
+ return -ENODEV;
+
+ if (list_empty(&rdtgroup_lists))
+ goto end;
+
+ __bitmap_parse(buf, strlen(buf), 0, bitmap, nr_cpu_ids);
+
+ cpumask = to_cpumask(bitmap);
+
+ list_for_each(l, &rdtgroup_lists) {
+ r = list_entry(l, struct rdtgroup, rdtgroup_list);
+ if (r == rdtgrp)
+ continue;
+
+ for_each_cpu_and(cpu, &r->cpu_mask, cpumask)
+ cpumask_clear_cpu(cpu, &r->cpu_mask);
+ }
+
+ cpumask_copy(&rdtgrp->cpu_mask, cpumask);
+ for_each_cpu(cpu, cpumask)
+ per_cpu(cpu_rdtgroup, cpu) = rdtgrp;
+
+end:
+ rdtgroup_kn_unlock(of->kn);
+
+ return nbytes;
+}
static struct rftype rdtgroup_partition_base_files[] = {
{
--
2.5.0
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web