Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1220388 > unrolled thread
| Started by | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| First post | 2015-09-07 22:50 +0200 |
| Last post | 2015-09-10 19:50 +0200 |
| Articles | 20 on this page of 40 — 5 participants |
Back to article view | Back to linux.kernel
[PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-07 22:50 +0200
[PATCH 3/7] devcg: Added infrastructure for rdma device cgroup. Parav Pandit <pandit.parav@gmail.com> - 2015-09-07 22:50 +0200
Re: [PATCH 3/7] devcg: Added infrastructure for rdma device cgroup. Parav Pandit <pandit.parav@gmail.com> - 2015-09-08 09:10 +0200
[PATCH 4/7] devcg: Added rdma resource tracker object per task Parav Pandit <pandit.parav@gmail.com> - 2015-09-07 22:50 +0200
Re: [PATCH 4/7] devcg: Added rdma resource tracker object per task Parav Pandit <pandit.parav@gmail.com> - 2015-09-08 09:10 +0200
Re: [PATCH 4/7] devcg: Added rdma resource tracker object per task Parav Pandit <pandit.parav@gmail.com> - 2015-09-08 10:30 +0200
[PATCH 7/7] devcg: Added Documentation of RDMA device cgroup. Parav Pandit <pandit.parav@gmail.com> - 2015-09-07 22:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-07 23:00 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-08 17:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-09 06:00 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-10 18:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-10 19:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-10 22:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-11 05:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 06:10 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Doug Ledford <dledford@redhat.com> - 2015-09-11 06:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 17:00 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-11 18:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-11 18:40 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 21:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-14 12:20 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 18:40 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-11 18:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 21:10 +0200
RE: [PATCH 0/7] devcg: device cgroup extension for rdma resource "Hefty, Sean" <sean.hefty@intel.com> - 2015-09-11 21:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2015-09-11 21:50 +0200
RE: [PATCH 0/7] devcg: device cgroup extension for rdma resource "Hefty, Sean" <sean.hefty@intel.com> - 2015-09-11 22:10 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-14 13:10 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-14 16:10 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-14 17:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2015-09-14 19:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-14 21:00 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2015-09-14 22:20 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-15 05:10 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2015-09-15 05:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-16 06:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-14 12:20 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-11 06:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 17:10 +0200
RE: [PATCH 0/7] devcg: device cgroup extension for rdma resource "Hefty, Sean" <sean.hefty@intel.com> - 2015-09-10 19:50 +0200
Page 1 of 2 [1] 2 Next page →
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-07 22:50 +0200 |
| Subject | [PATCH 0/7] devcg: device cgroup extension for rdma resource |
| Message-ID | <q6bwt-863-3@gated-at.bofh.it> |
Currently user space applications can easily take away all the rdma device specific resources such as AH, CQ, QP, MR etc. Due to which other applications in other cgroup or kernel space ULPs may not even get chance to allocate any rdma resources. This patch-set allows limiting rdma resources to set of processes. It extend device cgroup controller for limiting rdma device limits. With this patch, user verbs module queries rdma device cgroup controller to query process's limit to consume such resource. It uncharge resource counter after resource is being freed. It extends the task structure to hold the statistic information about process's rdma resource usage so that when process migrates from one to other controller, right amount of resources can be migrated from one to other cgroup. Future patches will support RDMA flows resource and will be enhanced further to enforce limit of other resources and capabilities. Parav Pandit (7): devcg: Added user option to rdma resource tracking. devcg: Added rdma resource tracking module. devcg: Added infrastructure for rdma device cgroup. devcg: Added rdma resource tracker object per task devcg: device cgroup's extension for RDMA resource. devcg: Added support to use RDMA device cgroup. devcg: Added Documentation of RDMA device cgroup. Documentation/cgroups/devices.txt | 32 ++- drivers/infiniband/core/uverbs_cmd.c | 139 +++++++++-- drivers/infiniband/core/uverbs_main.c | 39 +++- include/linux/device_cgroup.h | 53 +++++ include/linux/device_rdma_cgroup.h | 83 +++++++ include/linux/sched.h | 12 +- init/Kconfig | 12 + security/Makefile | 1 + security/device_cgroup.c | 119 +++++++--- security/device_rdma_cgroup.c | 422 ++++++++++++++++++++++++++++++++++ 10 files changed, 850 insertions(+), 62 deletions(-) create mode 100644 include/linux/device_rdma_cgroup.h create mode 100644 security/device_rdma_cgroup.c -- 1.8.3.1 -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-07 22:50 +0200 |
| Subject | [PATCH 3/7] devcg: Added infrastructure for rdma device cgroup. |
| Message-ID | <q6bwu-863-27@gated-at.bofh.it> |
| In reply to | #1220388 |
1. Moved necessary functions and data structures to header file to
reuse them at device cgroup white list functionality and for rdma
functionality.
2. Added infrastructure to invoke RDMA specific routines for resource
configuration, query and during fork handling.
3. Added sysfs interface files for configuring max limit of each rdma
resource and one file for querying controllers current resource usage.
Signed-off-by: Parav Pandit <pandit.parav@gmail.com>
---
include/linux/device_cgroup.h | 53 +++++++++++++++++++
security/device_cgroup.c | 119 +++++++++++++++++++++++++++++-------------
2 files changed, 136 insertions(+), 36 deletions(-)
diff --git a/include/linux/device_cgroup.h b/include/linux/device_cgroup.h
index 8b64221..cdbdd60 100644
--- a/include/linux/device_cgroup.h
+++ b/include/linux/device_cgroup.h
@@ -1,6 +1,57 @@
+#ifndef _DEVICE_CGROUP
+#define _DEVICE_CGROUP
+
#include <linux/fs.h>
+#include <linux/cgroup.h>
+#include <linux/device_rdma_cgroup.h>
#ifdef CONFIG_CGROUP_DEVICE
+
+enum devcg_behavior {
+ DEVCG_DEFAULT_NONE,
+ DEVCG_DEFAULT_ALLOW,
+ DEVCG_DEFAULT_DENY,
+};
+
+/*
+ * exception list locking rules:
+ * hold devcgroup_mutex for update/read.
+ * hold rcu_read_lock() for read.
+ */
+
+struct dev_exception_item {
+ u32 major, minor;
+ short type;
+ short access;
+ struct list_head list;
+ struct rcu_head rcu;
+};
+
+struct dev_cgroup {
+ struct cgroup_subsys_state css;
+ struct list_head exceptions;
+ enum devcg_behavior behavior;
+
+#ifdef CONFIG_CGROUP_RDMA_RESOURCE
+ struct devcgroup_rdma rdma;
+#endif
+};
+
+static inline struct dev_cgroup *css_to_devcgroup(struct cgroup_subsys_state *s)
+{
+ return s ? container_of(s, struct dev_cgroup, css) : NULL;
+}
+
+static inline struct dev_cgroup *parent_devcgroup(struct dev_cgroup *dev_cg)
+{
+ return css_to_devcgroup(dev_cg->css.parent);
+}
+
+static inline struct dev_cgroup *task_devcgroup(struct task_struct *task)
+{
+ return css_to_devcgroup(task_css(task, devices_cgrp_id));
+}
+
extern int __devcgroup_inode_permission(struct inode *inode, int mask);
extern int devcgroup_inode_mknod(int mode, dev_t dev);
static inline int devcgroup_inode_permission(struct inode *inode, int mask)
@@ -17,3 +68,5 @@ static inline int devcgroup_inode_permission(struct inode *inode, int mask)
static inline int devcgroup_inode_mknod(int mode, dev_t dev)
{ return 0; }
#endif
+
+#endif
diff --git a/security/device_cgroup.c b/security/device_cgroup.c
index 188c1d2..a0b3239 100644
--- a/security/device_cgroup.c
+++ b/security/device_cgroup.c
@@ -25,42 +25,6 @@
static DEFINE_MUTEX(devcgroup_mutex);
-enum devcg_behavior {
- DEVCG_DEFAULT_NONE,
- DEVCG_DEFAULT_ALLOW,
- DEVCG_DEFAULT_DENY,
-};
-
-/*
- * exception list locking rules:
- * hold devcgroup_mutex for update/read.
- * hold rcu_read_lock() for read.
- */
-
-struct dev_exception_item {
- u32 major, minor;
- short type;
- short access;
- struct list_head list;
- struct rcu_head rcu;
-};
-
-struct dev_cgroup {
- struct cgroup_subsys_state css;
- struct list_head exceptions;
- enum devcg_behavior behavior;
-};
-
-static inline struct dev_cgroup *css_to_devcgroup(struct cgroup_subsys_state *s)
-{
- return s ? container_of(s, struct dev_cgroup, css) : NULL;
-}
-
-static inline struct dev_cgroup *task_devcgroup(struct task_struct *task)
-{
- return css_to_devcgroup(task_css(task, devices_cgrp_id));
-}
-
/*
* called under devcgroup_mutex
*/
@@ -223,6 +187,9 @@ devcgroup_css_alloc(struct cgroup_subsys_state *parent_css)
INIT_LIST_HEAD(&dev_cgroup->exceptions);
dev_cgroup->behavior = DEVCG_DEFAULT_NONE;
+#ifdef CONFIG_CGROUP_RDMA_RESOURCE
+ init_devcgroup_rdma_tracker(dev_cgroup);
+#endif
return &dev_cgroup->css;
}
@@ -234,6 +201,25 @@ static void devcgroup_css_free(struct cgroup_subsys_state *css)
kfree(dev_cgroup);
}
+#ifdef CONFIG_CGROUP_RDMA_RESOURCE
+static int devcgroup_can_attach(struct cgroup_subsys_state *dst_css,
+ struct cgroup_taskset *tset)
+{
+ return devcgroup_rdma_can_attach(dst_css, tset);
+}
+
+static void devcgroup_cancel_attach(struct cgroup_subsys_state *dst_css,
+ struct cgroup_taskset *tset)
+{
+ devcgroup_cancel_attach(dst_css, tset);
+}
+
+static void devcgroup_fork(struct task_struct *task, void *priv)
+{
+ devcgroup_rdma_fork(task, priv);
+}
+#endif
+
#define DEVCG_ALLOW 1
#define DEVCG_DENY 2
#define DEVCG_LIST 3
@@ -788,6 +774,62 @@ static struct cftype dev_cgroup_files[] = {
.seq_show = devcgroup_seq_show,
.private = DEVCG_LIST,
},
+
+#ifdef CONFIG_CGROUP_RDMA_RESOURCE
+ {
+ .name = "rdma.resource.uctx.max",
+ .write = devcgroup_rdma_set_max_resource,
+ .seq_show = devcgroup_rdma_get_max_resource,
+ .private = DEVCG_RDMA_RES_TYPE_UCTX,
+ },
+ {
+ .name = "rdma.resource.cq.max",
+ .write = devcgroup_rdma_set_max_resource,
+ .seq_show = devcgroup_rdma_get_max_resource,
+ .private = DEVCG_RDMA_RES_TYPE_CQ,
+ },
+ {
+ .name = "rdma.resource.ah.max",
+ .write = devcgroup_rdma_set_max_resource,
+ .seq_show = devcgroup_rdma_get_max_resource,
+ .private = DEVCG_RDMA_RES_TYPE_AH,
+ },
+ {
+ .name = "rdma.resource.pd.max",
+ .write = devcgroup_rdma_set_max_resource,
+ .seq_show = devcgroup_rdma_get_max_resource,
+ .private = DEVCG_RDMA_RES_TYPE_PD,
+ },
+ {
+ .name = "rdma.resource.flow.max",
+ .write = devcgroup_rdma_set_max_resource,
+ .seq_show = devcgroup_rdma_get_max_resource,
+ .private = DEVCG_RDMA_RES_TYPE_FLOW,
+ },
+ {
+ .name = "rdma.resource.srq.max",
+ .write = devcgroup_rdma_set_max_resource,
+ .seq_show = devcgroup_rdma_get_max_resource,
+ .private = DEVCG_RDMA_RES_TYPE_SRQ,
+ },
+ {
+ .name = "rdma.resource.qp.max",
+ .write = devcgroup_rdma_set_max_resource,
+ .seq_show = devcgroup_rdma_get_max_resource,
+ .private = DEVCG_RDMA_RES_TYPE_QP,
+ },
+ {
+ .name = "rdma.resource.mr.max",
+ .write = devcgroup_rdma_set_max_resource,
+ .seq_show = devcgroup_rdma_get_max_resource,
+ .private = DEVCG_RDMA_RES_TYPE_MR,
+ },
+ {
+ .name = "rdma.resource.usage",
+ .seq_show = devcgroup_rdma_show_usage,
+ .private = DEVCG_RDMA_LIST_USAGE,
+ },
+#endif
{ } /* terminate */
};
@@ -796,6 +838,11 @@ struct cgroup_subsys devices_cgrp_subsys = {
.css_free = devcgroup_css_free,
.css_online = devcgroup_online,
.css_offline = devcgroup_offline,
+#ifdef CONFIG_CGROUP_RDMA_RESOURCE
+ .fork = devcgroup_fork,
+ .can_attach = devcgroup_can_attach,
+ .cancel_attach = devcgroup_cancel_attach,
+#endif
.legacy_cftypes = dev_cgroup_files,
};
--
1.8.3.1
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-08 09:10 +0200 |
| Subject | Re: [PATCH 3/7] devcg: Added infrastructure for rdma device cgroup. |
| Message-ID | <q6lct-5wa-11@gated-at.bofh.it> |
| In reply to | #1220390 |
On Tue, Sep 8, 2015 at 11:01 AM, Haggai Eran <haggaie@mellanox.com> wrote: > On 07/09/2015 23:38, Parav Pandit wrote: >> diff --git a/include/linux/device_cgroup.h b/include/linux/device_cgroup.h >> index 8b64221..cdbdd60 100644 >> --- a/include/linux/device_cgroup.h >> +++ b/include/linux/device_cgroup.h >> @@ -1,6 +1,57 @@ >> +#ifndef _DEVICE_CGROUP >> +#define _DEVICE_CGROUP >> + >> #include <linux/fs.h> >> +#include <linux/cgroup.h> >> +#include <linux/device_rdma_cgroup.h> > > You cannot add this include line before adding the device_rdma_cgroup.h > (added in patch 5). You should reorder the patches so that after each > patch the kernel builds correctly. > o.k. got it. I will send V1 with this suggested changes. > I also noticed in patch 2 you add device_rdma_cgroup.o to the Makefile > before it was added to the kernel. > o.k. > Regards, > Haggai -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-07 22:50 +0200 |
| Subject | [PATCH 4/7] devcg: Added rdma resource tracker object per task |
| Message-ID | <q6bwv-863-33@gated-at.bofh.it> |
| In reply to | #1220388 |
Added RDMA device resource tracking object per task.
Added comments to capture usage of task lock by device cgroup
for rdma.
Signed-off-by: Parav Pandit <pandit.parav@gmail.com>
---
include/linux/sched.h | 12 +++++++++++-
1 file changed, 11 insertions(+), 1 deletion(-)
diff --git a/include/linux/sched.h b/include/linux/sched.h
index ae21f15..a5f79b6 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -1334,6 +1334,8 @@ union rcu_special {
};
struct rcu_node;
+struct task_rdma_res_counter;
+
enum perf_event_task_context {
perf_invalid_context = -1,
perf_hw_context = 0,
@@ -1637,6 +1639,14 @@ struct task_struct {
struct css_set __rcu *cgroups;
/* cg_list protected by css_set_lock and tsk->alloc_lock */
struct list_head cg_list;
+
+#ifdef CONFIG_CGROUP_RDMA_RESOURCE
+ /* RDMA resource accounting counters, allocated only
+ * when RDMA resources are created by a task.
+ */
+ struct task_rdma_res_counter *rdma_res_counter;
+#endif
+
#endif
#ifdef CONFIG_FUTEX
struct robust_list_head __user *robust_list;
@@ -2676,7 +2686,7 @@ static inline int thread_group_empty(struct task_struct *p)
* Protects ->fs, ->files, ->mm, ->group_info, ->comm, keyring
* subscriptions and synchronises with wait4(). Also used in procfs. Also
* pins the final release of task.io_context. Also protects ->cpuset and
- * ->cgroup.subsys[]. And ->vfork_done.
+ * ->cgroup.subsys[]. Also projtects ->vfork_done and ->rdma_res_counter.
*
* Nests both inside and outside of read_lock(&tasklist_lock).
* It must not be nested with write_lock_irq(&tasklist_lock),
--
1.8.3.1
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-08 09:10 +0200 |
| Subject | Re: [PATCH 4/7] devcg: Added rdma resource tracker object per task |
| Message-ID | <q6lct-5wa-3@gated-at.bofh.it> |
| In reply to | #1220391 |
On Tue, Sep 8, 2015 at 11:18 AM, Haggai Eran <haggaie@mellanox.com> wrote: > On 07/09/2015 23:38, Parav Pandit wrote: >> @@ -2676,7 +2686,7 @@ static inline int thread_group_empty(struct task_struct *p) >> * Protects ->fs, ->files, ->mm, ->group_info, ->comm, keyring >> * subscriptions and synchronises with wait4(). Also used in procfs. Also >> * pins the final release of task.io_context. Also protects ->cpuset and >> - * ->cgroup.subsys[]. And ->vfork_done. >> + * ->cgroup.subsys[]. Also projtects ->vfork_done and ->rdma_res_counter. > s/projtects/protects/ >> * >> * Nests both inside and outside of read_lock(&tasklist_lock). >> * It must not be nested with write_lock_irq(&tasklist_lock), > Hi Haggai Eran, Did you miss to put comments or I missed something? Parav -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-08 10:30 +0200 |
| Subject | Re: [PATCH 4/7] devcg: Added rdma resource tracker object per task |
| Message-ID | <q6mrT-7dD-1@gated-at.bofh.it> |
| In reply to | #1220530 |
On Tue, Sep 8, 2015 at 1:54 PM, Haggai Eran <haggaie@mellanox.com> wrote: > On 08/09/2015 10:04, Parav Pandit wrote: >> On Tue, Sep 8, 2015 at 11:18 AM, Haggai Eran <haggaie@mellanox.com> wrote: >>> On 07/09/2015 23:38, Parav Pandit wrote: >>>> @@ -2676,7 +2686,7 @@ static inline int thread_group_empty(struct task_struct *p) >>>> * Protects ->fs, ->files, ->mm, ->group_info, ->comm, keyring >>>> * subscriptions and synchronises with wait4(). Also used in procfs. Also >>>> * pins the final release of task.io_context. Also protects ->cpuset and >>>> - * ->cgroup.subsys[]. And ->vfork_done. >>>> + * ->cgroup.subsys[]. Also projtects ->vfork_done and ->rdma_res_counter. >>> s/projtects/protects/ >>>> * >>>> * Nests both inside and outside of read_lock(&tasklist_lock). >>>> * It must not be nested with write_lock_irq(&tasklist_lock), >>> >> >> Hi Haggai Eran, >> Did you miss to put comments or I missed something? > > Yes, I wrote "s/projtects/protects/" to tell you that you have a typo in > your comment. You should change the word "projtects" to "protects". > > Haggai > ah. ok. Right. Will correct it. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-07 22:50 +0200 |
| Subject | [PATCH 7/7] devcg: Added Documentation of RDMA device cgroup. |
| Message-ID | <q6bwv-863-35@gated-at.bofh.it> |
| In reply to | #1220388 |
Modified device cgroup documentation to reflect its dual purpose without creating new cgroup subsystem for rdma. Added documentation to describe functionality and usage of device cgroup extension for RDMA. Signed-off-by: Parav Pandit <pandit.parav@gmail.com> --- Documentation/cgroups/devices.txt | 32 +++++++++++++++++++++++++++++--- 1 file changed, 29 insertions(+), 3 deletions(-) diff --git a/Documentation/cgroups/devices.txt b/Documentation/cgroups/devices.txt index 3c1095c..eca5b70 100644 --- a/Documentation/cgroups/devices.txt +++ b/Documentation/cgroups/devices.txt @@ -1,9 +1,12 @@ -Device Whitelist Controller +Device Controller 1. Description: -Implement a cgroup to track and enforce open and mknod restrictions -on device files. A device cgroup associates a device access +Device controller implements a cgroup for two purposes. + +1.1 Device white list controller +It implement a cgroup to track and enforce open and mknod +restrictions on device files. A device cgroup associates a device access whitelist with each cgroup. A whitelist entry has 4 fields. 'type' is a (all), c (char), or b (block). 'all' means it applies to all types and all major and minor numbers. Major and minor are @@ -15,8 +18,15 @@ cgroup gets a copy of the parent. Administrators can then remove devices from the whitelist or add new entries. A child cgroup can never receive a device access which is denied by its parent. +1.2 RDMA device resource controller +It implements a cgroup to limit various RDMA device resources for +a controller. Such resource includes RDMA PD, CQ, AH, MR, SRQ, QP, FLOW. +It limits RDMA resources access to tasks of the cgroup across multiple +RDMA devices. + 2. User Interface +2.1 Device white list controller An entry is added using devices.allow, and removed using devices.deny. For instance @@ -33,6 +43,22 @@ will remove the default 'a *:* rwm' entry. Doing will add the 'a *:* rwm' entry to the whitelist. +2.2 RDMA device controller + +RDMA resources are limited using devices.rdma.resource.max.<resource_name>. +Doing + echo 200 > /sys/fs/cgroup/1/rdma.resource.max_qp +will limit maximum number of QP across all the process of cgroup to 200. + +More examples: + echo 200 > /sys/fs/cgroup/1/rdma.resource.max_flow + echo 10 > /sys/fs/cgroup/1/rdma.resource.max_pd + echo 15 > /sys/fs/cgroup/1/rdma.resource.max_srq + echo 1 > /sys/fs/cgroup/1/rdma.resource.max_uctx + +RDMA resource current usage can be tracked using devices.rdma.resource.usage + cat /sys/fs/cgroup/1/devices.rdma.resource.usage + 3. Security Any task can move itself between cgroups. This clearly won't -- 1.8.3.1 -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-07 23:00 +0200 |
| Message-ID | <q6bG9-8hQ-5@gated-at.bofh.it> |
| In reply to | #1220388 |
Hi Doug, Tejun, This is from cgroups for-4.3 branch. linux-rdma trunk will face compilation error as its behind Tejun's for-4.3 branch. Patch has dependency on the some of the cgroup subsystem functionality for fork(). Therefore its required to merge those changes first to linux-rdma trunk. Parav On Tue, Sep 8, 2015 at 2:08 AM, Parav Pandit <pandit.parav@gmail.com> wrote: > Currently user space applications can easily take away all the rdma > device specific resources such as AH, CQ, QP, MR etc. Due to which other > applications in other cgroup or kernel space ULPs may not even get chance > to allocate any rdma resources. > > This patch-set allows limiting rdma resources to set of processes. > It extend device cgroup controller for limiting rdma device limits. > > With this patch, user verbs module queries rdma device cgroup controller > to query process's limit to consume such resource. It uncharge resource > counter after resource is being freed. > > It extends the task structure to hold the statistic information about process's > rdma resource usage so that when process migrates from one to other controller, > right amount of resources can be migrated from one to other cgroup. > > Future patches will support RDMA flows resource and will be enhanced further > to enforce limit of other resources and capabilities. > > Parav Pandit (7): > devcg: Added user option to rdma resource tracking. > devcg: Added rdma resource tracking module. > devcg: Added infrastructure for rdma device cgroup. > devcg: Added rdma resource tracker object per task > devcg: device cgroup's extension for RDMA resource. > devcg: Added support to use RDMA device cgroup. > devcg: Added Documentation of RDMA device cgroup. > > Documentation/cgroups/devices.txt | 32 ++- > drivers/infiniband/core/uverbs_cmd.c | 139 +++++++++-- > drivers/infiniband/core/uverbs_main.c | 39 +++- > include/linux/device_cgroup.h | 53 +++++ > include/linux/device_rdma_cgroup.h | 83 +++++++ > include/linux/sched.h | 12 +- > init/Kconfig | 12 + > security/Makefile | 1 + > security/device_cgroup.c | 119 +++++++--- > security/device_rdma_cgroup.c | 422 ++++++++++++++++++++++++++++++++++ > 10 files changed, 850 insertions(+), 62 deletions(-) > create mode 100644 include/linux/device_rdma_cgroup.h > create mode 100644 security/device_rdma_cgroup.c > > -- > 1.8.3.1 > -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-08 17:30 +0200 |
| Message-ID | <q6t0m-8ch-5@gated-at.bofh.it> |
| In reply to | #1220388 |
Hello, Parav. On Tue, Sep 08, 2015 at 02:08:16AM +0530, Parav Pandit wrote: > Currently user space applications can easily take away all the rdma > device specific resources such as AH, CQ, QP, MR etc. Due to which other > applications in other cgroup or kernel space ULPs may not even get chance > to allocate any rdma resources. Is there something simple I can read up on what each resource is? What's the usual access control mechanism? > This patch-set allows limiting rdma resources to set of processes. > It extend device cgroup controller for limiting rdma device limits. I don't think this belongs to devcg. If these make sense as a set of resources to be controlled via cgroup, the right way prolly would be a separate controller. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-09 06:00 +0200 |
| Message-ID | <q6EI9-89L-1@gated-at.bofh.it> |
| In reply to | #1220925 |
On Tue, Sep 8, 2015 at 8:53 PM, Tejun Heo <tj@kernel.org> wrote:
> Hello, Parav.
>
> On Tue, Sep 08, 2015 at 02:08:16AM +0530, Parav Pandit wrote:
>> Currently user space applications can easily take away all the rdma
>> device specific resources such as AH, CQ, QP, MR etc. Due to which other
>> applications in other cgroup or kernel space ULPs may not even get chance
>> to allocate any rdma resources.
>
> Is there something simple I can read up on what each resource is?
> What's the usual access control mechanism?
>
Hi Tejun,
This is one old white paper, but most of the reasoning still holds true on RDMA.
http://h10032.www1.hp.com/ctg/Manual/c00257031.pdf
More notes on RDMA resources and summary:
RDMA allows data transport from one system to other system where RDMA
device implements OSI layers 4 to 1 typically in hardware, drivers.
RDMA device provides data path semantics to perform data transfer in
zero copy manner from one to other host, very similar to local dma
controller.
It also allows data transfer operation from user space application of
one to other system.
In order to do so, all the resources are created using trusted kernel
space which also provides isolation among applications.
These resources include are- QP (queue pair) to transfer data, CQ
(Completion queue) to indicate completion of data transfer operation,
MR (memory region) to represent user application memory as source or
destination for data transfer.
Common resources are QP, SRQ (shared received queue), CQ, MR, AH
(Address handle), FLOW, PD (protection domain), user context etc.
>> This patch-set allows limiting rdma resources to set of processes.
>> It extend device cgroup controller for limiting rdma device limits.
>
> I don't think this belongs to devcg. If these make sense as a set of
> resources to be controlled via cgroup, the right way prolly would be a
> separate controller.
>
In past there has been similar comment to have dedicated cgroup
controller for RDMA instead of merging with device cgroup.
I am ok with both the approach, however I prefer to utilize device
controller instead of spinning of new controller for new devices
category.
I anticipate more such need would arise and for new device category,
it might not be worth to have new cgroup controller.
RapidIO though very less popular and upcoming PCIe are on horizon to
offer similar benefits as that of RDMA and in future having one
controller for each of them again would not be right approach.
I certainly seek your and others inputs in this email thread here whether
(a) to continue to extend device cgroup (which support character,
block devices white list) and now RDMA devices
or
(b) to spin of new controller, if so what are the compelling reasons
that it can provide compare to extension.
Current scope of the patch is limited to RDMA resources as first
patch, but for fact I am sure that there are more functionality in
pipe to support via this cgroup by me and others.
So keeping atleast these two aspects in mind, I need input on
direction of dedicated controller or new one.
In future, I anticipate that we might have sub directory to device
cgroup for individual device class to control.
such as,
<sys/fs/cgroup/devices/
/char
/block
/rdma
/pcie
/child_cgroup..1..N
Each controllers cgroup access files would remain within their own
scope. We are not there yet from base infrastructure but something to
be done as it matures and users start using it.
> Thanks.
>
> --
> tejun
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-10 18:50 +0200 |
| Message-ID | <q7dcR-79C-5@gated-at.bofh.it> |
| In reply to | #1221215 |
Hello, Parav. On Wed, Sep 09, 2015 at 09:27:40AM +0530, Parav Pandit wrote: > This is one old white paper, but most of the reasoning still holds true on RDMA. > http://h10032.www1.hp.com/ctg/Manual/c00257031.pdf Just read it. Much appreciated. ... > These resources include are- QP (queue pair) to transfer data, CQ > (Completion queue) to indicate completion of data transfer operation, > MR (memory region) to represent user application memory as source or > destination for data transfer. > Common resources are QP, SRQ (shared received queue), CQ, MR, AH > (Address handle), FLOW, PD (protection domain), user context etc. It's kinda bothering that all these are disparate resources. I suppose that each restriction comes from the underlying hardware and there's no accepted higher level abstraction for these things? > >> This patch-set allows limiting rdma resources to set of processes. > >> It extend device cgroup controller for limiting rdma device limits. > > > > I don't think this belongs to devcg. If these make sense as a set of > > resources to be controlled via cgroup, the right way prolly would be a > > separate controller. > > > > In past there has been similar comment to have dedicated cgroup > controller for RDMA instead of merging with device cgroup. > I am ok with both the approach, however I prefer to utilize device > controller instead of spinning of new controller for new devices > category. > I anticipate more such need would arise and for new device category, > it might not be worth to have new cgroup controller. > RapidIO though very less popular and upcoming PCIe are on horizon to > offer similar benefits as that of RDMA and in future having one > controller for each of them again would not be right approach. > > I certainly seek your and others inputs in this email thread here whether > (a) to continue to extend device cgroup (which support character, > block devices white list) and now RDMA devices > or > (b) to spin of new controller, if so what are the compelling reasons > that it can provide compare to extension. I'm doubtful that these things are gonna be mainstream w/o building up higher level abstractions on top and if we ever get there we won't be talking about MR or CQ or whatever. Also, whatever next-gen is unlikely to have enough commonalities when the proposed resource knobs are this low level, so let's please keep it separate, so that if/when this goes out of fashion for one reason or another, the controller can silently wither away too. > Current scope of the patch is limited to RDMA resources as first > patch, but for fact I am sure that there are more functionality in > pipe to support via this cgroup by me and others. > So keeping atleast these two aspects in mind, I need input on > direction of dedicated controller or new one. > > In future, I anticipate that we might have sub directory to device > cgroup for individual device class to control. > such as, > <sys/fs/cgroup/devices/ > /char > /block > /rdma > /pcie > /child_cgroup..1..N > Each controllers cgroup access files would remain within their own > scope. We are not there yet from base infrastructure but something to > be done as it matures and users start using it. I don't think that jives with the rest of cgroup and what generic block or pcie attributes are directly exposed to applications and need to be hierarchically controlled via cgroup? Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-10 19:50 +0200 |
| Message-ID | <q7e8W-8uF-3@gated-at.bofh.it> |
| In reply to | #1222307 |
On Thu, Sep 10, 2015 at 10:19 PM, Tejun Heo <tj@kernel.org> wrote: > Hello, Parav. > > On Wed, Sep 09, 2015 at 09:27:40AM +0530, Parav Pandit wrote: >> This is one old white paper, but most of the reasoning still holds true on RDMA. >> http://h10032.www1.hp.com/ctg/Manual/c00257031.pdf > > Just read it. Much appreciated. > > ... >> These resources include are- QP (queue pair) to transfer data, CQ >> (Completion queue) to indicate completion of data transfer operation, >> MR (memory region) to represent user application memory as source or >> destination for data transfer. >> Common resources are QP, SRQ (shared received queue), CQ, MR, AH >> (Address handle), FLOW, PD (protection domain), user context etc. > > It's kinda bothering that all these are disparate resources. Actually not. They are linked resources. Every QP needs associated one or two CQ, one PD. Every QP will use few MRs for data transfer. Here is the good programming guide of the RDMA APIs exposed to the user space application. http://www.mellanox.com/related-docs/prod_software/RDMA_Aware_Programming_user_manual.pdf So first version of the cgroups patch will address the control operation for section 3.4. > I suppose that each restriction comes from the underlying hardware and > there's no accepted higher level abstraction for these things? > There is higher level abstraction which is through the verbs layer currently which does actually expose the hardware resource but in vendor agnostic way. There are many vendors who support these verbs layer, some of them which I know are Mellanox, Intel, Chelsio, Avago/Emulex whose drivers which support these verbs are in <drivers/infiniband/hw/> kernel tree. There is higher level APIs above the verb layer, such as MPI, libfabric, rsocket, rds, pgas, dapl which uses underlying verbs layer. They all rely on the hardware resource. All of these higher level abstraction is accepted and well used by certain application class. It would be long discussion to go over them here. >> >> This patch-set allows limiting rdma resources to set of processes. >> >> It extend device cgroup controller for limiting rdma device limits. >> > >> > I don't think this belongs to devcg. If these make sense as a set of >> > resources to be controlled via cgroup, the right way prolly would be a >> > separate controller. >> > >> >> In past there has been similar comment to have dedicated cgroup >> controller for RDMA instead of merging with device cgroup. >> I am ok with both the approach, however I prefer to utilize device >> controller instead of spinning of new controller for new devices >> category. >> I anticipate more such need would arise and for new device category, >> it might not be worth to have new cgroup controller. >> RapidIO though very less popular and upcoming PCIe are on horizon to >> offer similar benefits as that of RDMA and in future having one >> controller for each of them again would not be right approach. >> >> I certainly seek your and others inputs in this email thread here whether >> (a) to continue to extend device cgroup (which support character, >> block devices white list) and now RDMA devices >> or >> (b) to spin of new controller, if so what are the compelling reasons >> that it can provide compare to extension. > > I'm doubtful that these things are gonna be mainstream w/o building up > higher level abstractions on top and if we ever get there we won't be > talking about MR or CQ or whatever. Some of the higher level examples I gave above will adapt to resource allocation failure. Some are actually adaptive to few resource allocation failure, they do query resources. But its not completely there yet. Once we have this notion of limited resource in place, abstraction layer would adapt to relatively smaller value of such resource. These higher level abstraction is mainstream. Its shipped at least in Redhat Enterprise Linux. > Also, whatever next-gen is > unlikely to have enough commonalities when the proposed resource knobs > are this low level, I agree that resource won't be common in next-gen other transport whenever they arrive. But with my existing background working on some of those transport, they appear similar in nature and it might seek similar knobs. > so let's please keep it separate, so that if/when > this goes out of fashion for one reason or another, the controller can > silently wither away too. > >> Current scope of the patch is limited to RDMA resources as first >> patch, but for fact I am sure that there are more functionality in >> pipe to support via this cgroup by me and others. >> So keeping atleast these two aspects in mind, I need input on >> direction of dedicated controller or new one. >> >> In future, I anticipate that we might have sub directory to device >> cgroup for individual device class to control. >> such as, >> <sys/fs/cgroup/devices/ >> /char >> /block >> /rdma >> /pcie >> /child_cgroup..1..N >> Each controllers cgroup access files would remain within their own >> scope. We are not there yet from base infrastructure but something to >> be done as it matures and users start using it. > > I don't think that jives with the rest of cgroup and what generic > block or pcie attributes are directly exposed to applications and need > to be hierarchically controlled via cgroup? > I do agree that currently cgroup doesn't have notion of sub cgroup or above hierarchy today. so until than I was considering to implement it under devices cgroup as generic place without the hierarchy shown above. Therefore current interface is at device cgroup level. If you are suggesting to have rdma cgroup as separate entity for near future, its fine with me. Later on when next-gen arrives we might have scope to make rdma cgroup as more generic one. But than it might look like what I described above. In past I have discussions with Liran Liss from Mellanox as well on this topic and we also agreed to have such cgroup controller. He has recent presentation at Linux foundation event indicating to have cgroup for RDMA. Below is the link to it. http://events.linuxfoundation.org/sites/events/files/slides/containing_rdma_final.pdf Slides 1 to 7 and slide 13 will give you more insight to it. Liran and I had similar presentation to RDMA audience with less slides in RDMA openfabrics summit in March 2015. I am ok to create separate cgroup for rdma, if community thinks that way. My preference would be still use device cgroup for above extensions unless there are fundamental issues that I am missing. I would let you make the call. Rdma and other is just another type of device with different characteristics than character or block, so one device cgroup with sub functionalities can allow setting knobs. Every device category will have their own set of knobs for resources, ACL, limits, policy. And I think cgroup is certainly better control point than sysfs or spinning of new control infrastructure for this. That said, I would like to hear your and communities view on how they would like to see this shaping up. > Thanks. > > -- > tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-10 22:30 +0200 |
| Message-ID | <q7gDM-3PH-13@gated-at.bofh.it> |
| In reply to | #1222332 |
Hello, Parav. On Thu, Sep 10, 2015 at 11:16:49PM +0530, Parav Pandit wrote: > >> These resources include are- QP (queue pair) to transfer data, CQ > >> (Completion queue) to indicate completion of data transfer operation, > >> MR (memory region) to represent user application memory as source or > >> destination for data transfer. > >> Common resources are QP, SRQ (shared received queue), CQ, MR, AH > >> (Address handle), FLOW, PD (protection domain), user context etc. > > > > It's kinda bothering that all these are disparate resources. > > Actually not. They are linked resources. Every QP needs associated one > or two CQ, one PD. > Every QP will use few MRs for data transfer. So, if that's the case, let's please implement something higher level. The goal is providing reasonable isolation or protection. If that can be achieved at a higher level of abstraction, please do that. > Here is the good programming guide of the RDMA APIs exposed to the > user space application. > > http://www.mellanox.com/related-docs/prod_software/RDMA_Aware_Programming_user_manual.pdf > So first version of the cgroups patch will address the control > operation for section 3.4. > > > I suppose that each restriction comes from the underlying hardware and > > there's no accepted higher level abstraction for these things? > > There is higher level abstraction which is through the verbs layer > currently which does actually expose the hardware resource but in > vendor agnostic way. > There are many vendors who support these verbs layer, some of them > which I know are Mellanox, Intel, Chelsio, Avago/Emulex whose drivers > which support these verbs are in <drivers/infiniband/hw/> kernel tree. > > There is higher level APIs above the verb layer, such as MPI, > libfabric, rsocket, rds, pgas, dapl which uses underlying verbs layer. > They all rely on the hardware resource. All of these higher level > abstraction is accepted and well used by certain application class. It > would be long discussion to go over them here. Well, the programming interface that userland builds on top doesn't matter too much here but if there is a common resource abstraction which can be made in terms of constructs that consumers of the facility would care about, that likely is a better choice than exposing whatever hardware exposes. > > I'm doubtful that these things are gonna be mainstream w/o building up > > higher level abstractions on top and if we ever get there we won't be > > talking about MR or CQ or whatever. > > Some of the higher level examples I gave above will adapt to resource > allocation failure. Some are actually adaptive to few resource > allocation failure, they do query resources. But its not completely > there yet. Once we have this notion of limited resource in place, > abstraction layer would adapt to relatively smaller value of such > resource. > > These higher level abstraction is mainstream. Its shipped at least in > Redhat Enterprise Linux. Again, I was talking more about resource abstraction - e.g. something along the line of "I want N command buffers". > > Also, whatever next-gen is > > unlikely to have enough commonalities when the proposed resource knobs > > are this low level, > > I agree that resource won't be common in next-gen other transport > whenever they arrive. > But with my existing background working on some of those transport, > they appear similar in nature and it might seek similar knobs. I don't know. What's proposed in this thread seems way too low level to be useful anywhere else. Also, what if there are multiple devices? Is that a problem to worry about? > In past I have discussions with Liran Liss from Mellanox as well on > this topic and we also agreed to have such cgroup controller. > He has recent presentation at Linux foundation event indicating to > have cgroup for RDMA. > Below is the link to it. > http://events.linuxfoundation.org/sites/events/files/slides/containing_rdma_final.pdf > Slides 1 to 7 and slide 13 will give you more insight to it. > Liran and I had similar presentation to RDMA audience with less slides > in RDMA openfabrics summit in March 2015. > > I am ok to create separate cgroup for rdma, if community thinks that way. > My preference would be still use device cgroup for above extensions > unless there are fundamental issues that I am missing. The thing is that they aren't related at all in any way. There's no reason to tie them together. In fact, the way we did devcg is backward. The ideal solution would have been extending the usual ACL to understand cgroups so that it's a natural growth of the permission system. You're talking about actual hardware resources. That has nothing to do with access permissions on device nodes. > I would let you make the call. > Rdma and other is just another type of device with different > characteristics than character or block, so one device cgroup with sub > functionalities can allow setting knobs. > Every device category will have their own set of knobs for resources, > ACL, limits, policy. I'm kinda doubtful we're gonna have too many of these. Hardware details being exposed to userland this directly isn't common. > And I think cgroup is certainly better control point than sysfs or > spinning of new control infrastructure for this. > That said, I would like to hear your and communities view on how they > would like to see this shaping up. I'd say keep it simple and do the minimum. :) Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-11 05:50 +0200 |
| Message-ID | <q7nvA-6BK-3@gated-at.bofh.it> |
| In reply to | #1222383 |
On Fri, Sep 11, 2015 at 1:52 AM, Tejun Heo <tj@kernel.org> wrote: > Hello, Parav. > > On Thu, Sep 10, 2015 at 11:16:49PM +0530, Parav Pandit wrote: >> >> These resources include are- QP (queue pair) to transfer data, CQ >> >> (Completion queue) to indicate completion of data transfer operation, >> >> MR (memory region) to represent user application memory as source or >> >> destination for data transfer. >> >> Common resources are QP, SRQ (shared received queue), CQ, MR, AH >> >> (Address handle), FLOW, PD (protection domain), user context etc. >> > >> > It's kinda bothering that all these are disparate resources. >> >> Actually not. They are linked resources. Every QP needs associated one >> or two CQ, one PD. >> Every QP will use few MRs for data transfer. > > So, if that's the case, let's please implement something higher level. > The goal is providing reasonable isolation or protection. If that can > be achieved at a higher level of abstraction, please do that. > >> Here is the good programming guide of the RDMA APIs exposed to the >> user space application. >> >> http://www.mellanox.com/related-docs/prod_software/RDMA_Aware_Programming_user_manual.pdf >> So first version of the cgroups patch will address the control >> operation for section 3.4. >> >> > I suppose that each restriction comes from the underlying hardware and >> > there's no accepted higher level abstraction for these things? >> >> There is higher level abstraction which is through the verbs layer >> currently which does actually expose the hardware resource but in >> vendor agnostic way. >> There are many vendors who support these verbs layer, some of them >> which I know are Mellanox, Intel, Chelsio, Avago/Emulex whose drivers >> which support these verbs are in <drivers/infiniband/hw/> kernel tree. >> >> There is higher level APIs above the verb layer, such as MPI, >> libfabric, rsocket, rds, pgas, dapl which uses underlying verbs layer. >> They all rely on the hardware resource. All of these higher level >> abstraction is accepted and well used by certain application class. It >> would be long discussion to go over them here. > > Well, the programming interface that userland builds on top doesn't > matter too much here but if there is a common resource abstraction > which can be made in terms of constructs that consumers of the > facility would care about, that likely is a better choice than > exposing whatever hardware exposes. > Tejun, The fact is that user level application uses hardware resources. Verbs layer is software abstraction for it. Drivers are hiding how they implement this QP or CQ or whatever hardware resource they project via API layer. For all of the userland on top of verb layer I mentioned above, the common resource abstraction is these resources AH, QP, CQ, MR etc. Hardware (and driver) might have different view of this resource in their real implementation. For example, verb layer can say that it has 100 QPs, but hardware might actually have 20 QPs that driver decide how to efficiently use it. >> > I'm doubtful that these things are gonna be mainstream w/o building up >> > higher level abstractions on top and if we ever get there we won't be >> > talking about MR or CQ or whatever. >> >> Some of the higher level examples I gave above will adapt to resource >> allocation failure. Some are actually adaptive to few resource >> allocation failure, they do query resources. But its not completely >> there yet. Once we have this notion of limited resource in place, >> abstraction layer would adapt to relatively smaller value of such >> resource. >> >> These higher level abstraction is mainstream. Its shipped at least in >> Redhat Enterprise Linux. > > Again, I was talking more about resource abstraction - e.g. something > along the line of "I want N command buffers". > Yes. We are still talking of resource abstraction here. RDMA and IBTA defines these resources. On top of these resources various frameworks are build. so for example, User land is tuning environment deploying for MPI application, it would configure: 10 processes from the PID controller, 10 CPUs in cpuset controller, 1 PD, 20 CQ, 10 QP, 100 MRs in rdma controller, say user land is tuning environment for deploying rsocket application for 100 connections, it would configure, 100 PD, 100 QP, 200 MR. When verb layer see failure with it, they will adapt to live with what they have at lower performance. Since every higher level which I mentioned in different in the way, it uses RDMA resources, we cannot generalize it as "N command buffers". That generalization in my mind is the - rdma resources - central common entity. >> > Also, whatever next-gen is >> > unlikely to have enough commonalities when the proposed resource knobs >> > are this low level, >> >> I agree that resource won't be common in next-gen other transport >> whenever they arrive. >> But with my existing background working on some of those transport, >> they appear similar in nature and it might seek similar knobs. > > I don't know. What's proposed in this thread seems way too low level > to be useful anywhere else. Also, what if there are multiple devices? > Is that a problem to worry about? > o.k. It doesn't have to be useful anywhere else. If it suffice the need of RDMA applications, its fine for near future. This patch allows limiting resources across multiple devices. As we go along the path, and if requirement come up to have knob on per device basis, thats something we can extend in future. > >> I would let you make the call. >> Rdma and other is just another type of device with different >> characteristics than character or block, so one device cgroup with sub >> functionalities can allow setting knobs. >> Every device category will have their own set of knobs for resources, >> ACL, limits, policy. > > I'm kinda doubtful we're gonna have too many of these. Hardware > details being exposed to userland this directly isn't common. > Its common in RDMA applications. Again they may not be real hardware resource, its just API layer which defines those RDMA constructs. >> And I think cgroup is certainly better control point than sysfs or >> spinning of new control infrastructure for this. >> That said, I would like to hear your and communities view on how they >> would like to see this shaping up. > > I'd say keep it simple and do the minimum. :) > o.k. In that case new rdma cgroup controller which does rdma resource accounting is possibly the most simplest form? Make sense? > Thanks. > > -- > tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-11 06:10 +0200 |
| Message-ID | <q7nOV-7dg-1@gated-at.bofh.it> |
| In reply to | #1222516 |
Hello, Parav. On Fri, Sep 11, 2015 at 09:09:58AM +0530, Parav Pandit wrote: > The fact is that user level application uses hardware resources. > Verbs layer is software abstraction for it. Drivers are hiding how > they implement this QP or CQ or whatever hardware resource they > project via API layer. > For all of the userland on top of verb layer I mentioned above, the > common resource abstraction is these resources AH, QP, CQ, MR etc. > Hardware (and driver) might have different view of this resource in > their real implementation. > For example, verb layer can say that it has 100 QPs, but hardware > might actually have 20 QPs that driver decide how to efficiently use > it. My uneducated suspicion is that the abstraction is just not developed enough. It should be possible to virtualize these resources through, most likely, time-sharing to the level where userland simply says "I want this chunk transferred there" and OS schedules the transfer prioritizing competing requests. It could be that given the use cases rdma might not need such level of abstraction - e.g. most users want to be and are pretty close to bare metal, but, if that's true, it also kinda is weird to build hierarchical resource distribution scheme on top of such bare abstraction. ... > > I don't know. What's proposed in this thread seems way too low level > > to be useful anywhere else. Also, what if there are multiple devices? > > Is that a problem to worry about? > > o.k. It doesn't have to be useful anywhere else. If it suffice the > need of RDMA applications, its fine for near future. > This patch allows limiting resources across multiple devices. > As we go along the path, and if requirement come up to have knob on > per device basis, thats something we can extend in future. You kinda have to decide that upfront cuz it gets baked into the interface. > > I'm kinda doubtful we're gonna have too many of these. Hardware > > details being exposed to userland this directly isn't common. > > Its common in RDMA applications. Again they may not be real hardware > resource, its just API layer which defines those RDMA constructs. It's still a very low level of abstraction which pretty much gets decided by what the hardware and driver decide to do. > > I'd say keep it simple and do the minimum. :) > > o.k. In that case new rdma cgroup controller which does rdma resource > accounting is possibly the most simplest form? > Make sense? So, this fits cgroup's purpose to certain level but it feels like we're trying to build too much on top of something which hasn't developed sufficiently. I suppose it could be that this is the level of development that rdma is gonna reach and dumb cgroup controller can be useful for some use cases. I don't know, so, yeah, let's keep it simple and avoid doing crazy stuff. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Doug Ledford <dledford@redhat.com> |
|---|---|
| Date | 2015-09-11 06:30 +0200 |
| Message-ID | <q7o8h-7BP-1@gated-at.bofh.it> |
| In reply to | #1222518 |
[Multipart message — attachments visible in raw view] — view raw
On 09/11/2015 12:04 AM, Tejun Heo wrote:
> Hello, Parav.
>
> On Fri, Sep 11, 2015 at 09:09:58AM +0530, Parav Pandit wrote:
>> The fact is that user level application uses hardware resources.
>> Verbs layer is software abstraction for it. Drivers are hiding how
>> they implement this QP or CQ or whatever hardware resource they
>> project via API layer.
>> For all of the userland on top of verb layer I mentioned above, the
>> common resource abstraction is these resources AH, QP, CQ, MR etc.
>> Hardware (and driver) might have different view of this resource in
>> their real implementation.
>> For example, verb layer can say that it has 100 QPs, but hardware
>> might actually have 20 QPs that driver decide how to efficiently use
>> it.
>
> My uneducated suspicion is that the abstraction is just not developed
> enough.
The abstraction is 10+ years old. It has had plenty of time to ferment
and something better for the specific use case has not emerged.
> It should be possible to virtualize these resources through,
> most likely, time-sharing to the level where userland simply says "I
> want this chunk transferred there" and OS schedules the transfer
> prioritizing competing requests.
No. And if you think this, then you miss the *entire* point of RDMA
technologies. An analogy that I have used many times in presentations
is that, in the networking world, the kernel is both a postman and a
copy machine. It receives all incoming packets and must sort them to
the right recipient (the postman job) and when the user space
application is ready to use the information it must copy it into the
user's VM space because it couldn't just put the user's data buffer on
the RX buffer list since each buffer might belong to anyone (the copy
machine). In the RDMA world, you create a new queue pair, it is often a
long lived connection (like a socket), but it belongs now to the app and
the app can directly queue both send and receive buffers to the card and
on incoming packets the card will be able to know that the packet
belongs to a specific queue pair and will immediately go to that apps
buffer. You can *not* do this with TCP without moving to complete TCP
offload on the card, registration of specific sockets on the card, and
then allowing the application to pre-register receive buffers for a
specific socket to the card so that incoming data on the wire can go
straight to the right place. If you ever get to the point of "OS
schedules the transfer" then you might as well throw RDMA out the window
because you have totally trashed the benefit it provides.
> It could be that given the use cases rdma might not need such level of
> abstraction - e.g. most users want to be and are pretty close to bare
> metal, but, if that's true, it also kinda is weird to build
> hierarchical resource distribution scheme on top of such bare
> abstraction.
Not really. If you are going to have a bare abstraction, this one isn't
really a bad one. You have devices. On a device, you allocate
protection domains (PDs). If you don't care about cross connection
issues, you ignore this and only use one. If you do care, this acts
like a process's unique VM space only for RDMA buffers, it is a domain
to protect the data of one connection from another. Then you have queue
pairs (QPs) which are roughly the equivalent of a socket. Each QP has
at least one Completion Queue where you get the events that tell you
things have completed (although they often use two, one for send
completions and one for receive completions). And then you use some
number of memory registrations (MRs) and address handles (AHs) depending
on your usage. Since RDMA stands for Remote Direct Memory Access, as
you can imagine, giving a remote machine free reign to access all of the
physical memory in your machine is a security issue. The MRs help to
control what memory the remote host on a specific QP has access to. The
AHs control how we actually route packets from ourselves to the remote host.
Here's the deal. You might be able to create an abstraction above this
that hides *some* of this. But it can't hide even nearly all of it
without loosing significant functionality. The problem here is that you
are thinking about RDMA connections like sockets. They aren't. Not
even close. They are "how do I allow a remote machine to directly read
and write into my machines physical memory in an even remotely close to
secure manner?" These resources aren't hardware resources, they are the
abstraction resources needed to answer that question.
> ...
>>> I don't know. What's proposed in this thread seems way too low level
>>> to be useful anywhere else. Also, what if there are multiple devices?
>>> Is that a problem to worry about?
>>
>> o.k. It doesn't have to be useful anywhere else. If it suffice the
>> need of RDMA applications, its fine for near future.
>> This patch allows limiting resources across multiple devices.
>> As we go along the path, and if requirement come up to have knob on
>> per device basis, thats something we can extend in future.
>
> You kinda have to decide that upfront cuz it gets baked into the
> interface.
>
>>> I'm kinda doubtful we're gonna have too many of these. Hardware
>>> details being exposed to userland this directly isn't common.
>>
>> Its common in RDMA applications. Again they may not be real hardware
>> resource, its just API layer which defines those RDMA constructs.
>
> It's still a very low level of abstraction which pretty much gets
> decided by what the hardware and driver decide to do.
>
>>> I'd say keep it simple and do the minimum. :)
>>
>> o.k. In that case new rdma cgroup controller which does rdma resource
>> accounting is possibly the most simplest form?
>> Make sense?
>
> So, this fits cgroup's purpose to certain level but it feels like
> we're trying to build too much on top of something which hasn't
> developed sufficiently. I suppose it could be that this is the level
> of development that rdma is gonna reach and dumb cgroup controller can
> be useful for some use cases. I don't know, so, yeah, let's keep it
> simple and avoid doing crazy stuff.
>
> Thanks.
>
--
Doug Ledford <dledford@redhat.com>
GPG KeyID: 0E572FDD
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-11 17:00 +0200 |
| Message-ID | <q7xXY-4O5-13@gated-at.bofh.it> |
| In reply to | #1222522 |
Hello, Doug. On Fri, Sep 11, 2015 at 12:24:33AM -0400, Doug Ledford wrote: > > My uneducated suspicion is that the abstraction is just not developed > > enough. > > The abstraction is 10+ years old. It has had plenty of time to ferment > and something better for the specific use case has not emerged. I think that is likely more reflective of the use cases rather than anything inherent in the concept. > > It should be possible to virtualize these resources through, > > most likely, time-sharing to the level where userland simply says "I > > want this chunk transferred there" and OS schedules the transfer > > prioritizing competing requests. > > No. And if you think this, then you miss the *entire* point of RDMA > technologies. An analogy that I have used many times in presentations > is that, in the networking world, the kernel is both a postman and a > copy machine. It receives all incoming packets and must sort them to > the right recipient (the postman job) and when the user space > application is ready to use the information it must copy it into the > user's VM space because it couldn't just put the user's data buffer on > the RX buffer list since each buffer might belong to anyone (the copy > machine). In the RDMA world, you create a new queue pair, it is often a > long lived connection (like a socket), but it belongs now to the app and > the app can directly queue both send and receive buffers to the card and > on incoming packets the card will be able to know that the packet > belongs to a specific queue pair and will immediately go to that apps > buffer. You can *not* do this with TCP without moving to complete TCP > offload on the card, registration of specific sockets on the card, and > then allowing the application to pre-register receive buffers for a > specific socket to the card so that incoming data on the wire can go > straight to the right place. If you ever get to the point of "OS > schedules the transfer" then you might as well throw RDMA out the window > because you have totally trashed the benefit it provides. I don't know. This sounds like classic "this is painful so it must be good" bare metal fantasy. I get that rdma succeeds at bypassing a lot of overhead. That's great but that really isn't exclusive with having more accessible mechanisms built on top. The crux of cost saving is the hardware knowing where the incoming data belongs and putting it there directly. Everything else is there to facilitate that and if you're declaring that it's impossible to build accessible abstractions for that, I can't agree with you. Note that this is not to say that rdma should do that in the operating system. As you said, people have been happy with the bare abstraction for a long time and, given relatively specialized use cases, that can be completely fine but please do note that the lack of proper abstraction isn't an inherent feature. It's just easier that way and putting in more effort hasn't been necessary. > > It could be that given the use cases rdma might not need such level of > > abstraction - e.g. most users want to be and are pretty close to bare > > metal, but, if that's true, it also kinda is weird to build > > hierarchical resource distribution scheme on top of such bare > > abstraction. > > Not really. If you are going to have a bare abstraction, this one isn't > really a bad one. You have devices. On a device, you allocate > protection domains (PDs). If you don't care about cross connection > issues, you ignore this and only use one. If you do care, this acts > like a process's unique VM space only for RDMA buffers, it is a domain > to protect the data of one connection from another. Then you have queue > pairs (QPs) which are roughly the equivalent of a socket. Each QP has > at least one Completion Queue where you get the events that tell you > things have completed (although they often use two, one for send > completions and one for receive completions). And then you use some > number of memory registrations (MRs) and address handles (AHs) depending > on your usage. Since RDMA stands for Remote Direct Memory Access, as > you can imagine, giving a remote machine free reign to access all of the > physical memory in your machine is a security issue. The MRs help to > control what memory the remote host on a specific QP has access to. The > AHs control how we actually route packets from ourselves to the remote host. > > Here's the deal. You might be able to create an abstraction above this > that hides *some* of this. But it can't hide even nearly all of it > without loosing significant functionality. The problem here is that you > are thinking about RDMA connections like sockets. They aren't. Not > even close. They are "how do I allow a remote machine to directly read > and write into my machines physical memory in an even remotely close to > secure manner?" These resources aren't hardware resources, they are the > abstraction resources needed to answer that question. So, the existence of resource limitations is fine. That's what we deal with all the time. The problem usually with this sort of interfaces which expose implementation details to users directly is that it severely limits engineering manuevering space. You usually want your users to express their intentions and a mechanism to arbitrate resources to satisfy those intentions (and in a way more graceful than "we can't, maybe try later?"); otherwise, implementing any sort of high level resource distribution scheme becomes painful and usually the only thing possible is preventing runaway disasters - you don't wanna pin unused resource permanently if there actually is contention around it, so usually all you can do with hard limits is overcommiting limits so that it at least prevents disasters. cpuset is a special case but think of cpu, memory or io controllers. Their resource distribution schemes are a lot more developed than what's proposed in this patchset and that's a necessity because nobody wants to cripple their machines for resource control. This is a lot more like the pids controller and that controller's almost sole purpose is preventing runaway workload wrecking the whole machine. It's getting rambly but the point is that if the resource being controlled by this controller is actually contended for performance reasons, this sort of hard limiting is inherently unlikely to be very useful. If the resource isn't and the main goal is preventing runaway hogs, it'll be able to do that but is that the goal here? For this to be actually useful for performance contended cases, it'd need higher level abstractions. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-11 18:30 +0200 |
| Message-ID | <q7zn3-6Xj-15@gated-at.bofh.it> |
| In reply to | #1222900 |
> If the resource isn't and the main goal is preventing runaway > hogs, it'll be able to do that but is that the goal here? For this to > be actually useful for performance contended cases, it'd need higher > level abstractions. > Resource run away by application can lead to (a) kernel and (b) other applications left out with no resources situation. Both the problems are the target of this patch set by accounting via cgroup. Performance contention can be resolved with higher level user space, which will tune it. Threshold and fail counters are on the way in follow on patch. > Thanks. > > -- > tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-11 18:40 +0200 |
| Message-ID | <q7zwK-78E-19@gated-at.bofh.it> |
| In reply to | #1222950 |
On Fri, Sep 11, 2015 at 10:04 PM, Tejun Heo <tj@kernel.org> wrote: > Hello, Parav. > > On Fri, Sep 11, 2015 at 09:56:31PM +0530, Parav Pandit wrote: >> Resource run away by application can lead to (a) kernel and (b) other >> applications left out with no resources situation. > > Yeap, that this controller would be able to prevent to a reasonable > extent. > >> Both the problems are the target of this patch set by accounting via cgroup. >> >> Performance contention can be resolved with higher level user space, >> which will tune it. > > If individual applications are gonna be allowed to do that, what's to > prevent them from jacking up their limits? I should have been more explicit. I didnt mean the application to control which is allocating it. > So, I assume you're > thinking of a central authority overseeing distribution and enforcing > the policy through cgroups? > Exactly. >> Threshold and fail counters are on the way in follow on patch. > > If you're planning on following what the existing memcg did in this > area, it's unlikely to go well. Would you mind sharing what you have > on mind in the long term? Where do you see this going? > At least current thoughts are: central entity authority monitors fail count and new threashold count. Fail count - as similar to other indicates how many time resource failure occured threshold count - indicates upto what this resource has gone upto in usage. (application might not be able to poll on thousands of such resources entries). So based on fail count and threshold count, it can tune it further. > Thanks. > > -- > tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-11 21:30 +0200 |
| Message-ID | <q7Cbg-2BT-27@gated-at.bofh.it> |
| In reply to | #1222962 |
Hello, Parav. On Fri, Sep 11, 2015 at 10:09:48PM +0530, Parav Pandit wrote: > > If you're planning on following what the existing memcg did in this > > area, it's unlikely to go well. Would you mind sharing what you have > > on mind in the long term? Where do you see this going? > > At least current thoughts are: central entity authority monitors fail > count and new threashold count. > Fail count - as similar to other indicates how many time resource > failure occured > threshold count - indicates upto what this resource has gone upto in > usage. (application might not be able to poll on thousands of such > resources entries). > So based on fail count and threshold count, it can tune it further. So, regardless of the specific resource in question, implementing adaptive resource distribution requires more than simple thresholds and failcnts. The very minimum would be a way to exert reclaim pressure and then a way to measure how much lack of a given resource is affecting the workload. Maybe it can adaptively lower the limits and then watch how often allocation fails but that's highly unlikely to be an effective measure as it can't do anything to hoarders and the frequency of allocation failure doesn't necessarily correlate with the amount of impact the workload is getting (it's not a measure of usage). This is what I'm awry about. The kernel-userland interface here is cut pretty low in the stack leaving most of arbitration and management logic in the userland, which seems to be what people wanted and that's fine, but then you're trying to implement an intelligent resource control layer which straddles across kernel and userland with those low level primitives which inevitably would increase the required interface surface as nobody has enough information. Just to illustrate the point, please think of the alsa interface. We expose hardware capabilities pretty much as-is leaving management and multiplexing to userland and there's nothing wrong with it. It fits better that way; however, we don't then go try to implement cgroup controller for PCM channels. To do any high-level resource management, you gotta do it where the said resource is actually managed and arbitrated. What's the allocation frequency you're expecting? It might be better to just let allocations themselves go through the agent that you're planning. You sure can use cgroup membership to identify who's asking tho. Given how the whole thing is architectured, I'd suggest thinking more about how the whole thing should turn out eventually. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web