Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1220388 > unrolled thread
| Started by | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| First post | 2015-09-07 22:50 +0200 |
| Last post | 2015-09-10 19:50 +0200 |
| Articles | 20 on this page of 40 — 5 participants |
Back to article view | Back to linux.kernel
[PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-07 22:50 +0200
[PATCH 3/7] devcg: Added infrastructure for rdma device cgroup. Parav Pandit <pandit.parav@gmail.com> - 2015-09-07 22:50 +0200
Re: [PATCH 3/7] devcg: Added infrastructure for rdma device cgroup. Parav Pandit <pandit.parav@gmail.com> - 2015-09-08 09:10 +0200
[PATCH 4/7] devcg: Added rdma resource tracker object per task Parav Pandit <pandit.parav@gmail.com> - 2015-09-07 22:50 +0200
Re: [PATCH 4/7] devcg: Added rdma resource tracker object per task Parav Pandit <pandit.parav@gmail.com> - 2015-09-08 09:10 +0200
Re: [PATCH 4/7] devcg: Added rdma resource tracker object per task Parav Pandit <pandit.parav@gmail.com> - 2015-09-08 10:30 +0200
[PATCH 7/7] devcg: Added Documentation of RDMA device cgroup. Parav Pandit <pandit.parav@gmail.com> - 2015-09-07 22:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-07 23:00 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-08 17:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-09 06:00 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-10 18:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-10 19:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-10 22:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-11 05:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 06:10 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Doug Ledford <dledford@redhat.com> - 2015-09-11 06:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 17:00 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-11 18:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-11 18:40 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 21:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-14 12:20 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 18:40 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-11 18:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 21:10 +0200
RE: [PATCH 0/7] devcg: device cgroup extension for rdma resource "Hefty, Sean" <sean.hefty@intel.com> - 2015-09-11 21:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2015-09-11 21:50 +0200
RE: [PATCH 0/7] devcg: device cgroup extension for rdma resource "Hefty, Sean" <sean.hefty@intel.com> - 2015-09-11 22:10 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-14 13:10 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-14 16:10 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-14 17:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2015-09-14 19:30 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-14 21:00 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2015-09-14 22:20 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-15 05:10 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2015-09-15 05:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-16 06:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-14 12:20 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Parav Pandit <pandit.parav@gmail.com> - 2015-09-11 06:50 +0200
Re: [PATCH 0/7] devcg: device cgroup extension for rdma resource Tejun Heo <tj@kernel.org> - 2015-09-11 17:10 +0200
RE: [PATCH 0/7] devcg: device cgroup extension for rdma resource "Hefty, Sean" <sean.hefty@intel.com> - 2015-09-10 19:50 +0200
Page 2 of 2 — ← Prev page 1 [2]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-14 12:20 +0200 |
| Message-ID | <q8z1E-3fP-21@gated-at.bofh.it> |
| In reply to | #1223058 |
On Sat, Sep 12, 2015 at 12:55 AM, Tejun Heo <tj@kernel.org> wrote: > Hello, Parav. > > On Fri, Sep 11, 2015 at 10:09:48PM +0530, Parav Pandit wrote: >> > If you're planning on following what the existing memcg did in this >> > area, it's unlikely to go well. Would you mind sharing what you have >> > on mind in the long term? Where do you see this going? >> >> At least current thoughts are: central entity authority monitors fail >> count and new threashold count. >> Fail count - as similar to other indicates how many time resource >> failure occured >> threshold count - indicates upto what this resource has gone upto in >> usage. (application might not be able to poll on thousands of such >> resources entries). >> So based on fail count and threshold count, it can tune it further. > > So, regardless of the specific resource in question, implementing > adaptive resource distribution requires more than simple thresholds > and failcnts. May be yes. Buts in difficult to go through the whole design to shape up right now. This is the infrastructure getting build with few capabilities. I see this as starting point instead of end point. > The very minimum would be a way to exert reclaim > pressure and then a way to measure how much lack of a given resource > is affecting the workload. Maybe it can adaptively lower the limits > and then watch how often allocation fails but that's highly unlikely > to be an effective measure as it can't do anything to hoarders and the > frequency of allocation failure doesn't necessarily correlate with the > amount of impact the workload is getting (it's not a measure of > usage). It can always kill the hoarding process(es), which is holding up the resources without using it. Such processes will eventually will get restarted but will not be able to hoard so much because its been on the radar for hoarding and its limits have been reduced. > > This is what I'm awry about. The kernel-userland interface here is > cut pretty low in the stack leaving most of arbitration and management > logic in the userland, which seems to be what people wanted and that's > fine, but then you're trying to implement an intelligent resource > control layer which straddles across kernel and userland with those > low level primitives which inevitably would increase the required > interface surface as nobody has enough information. > We might be able to get the information as we go along. Such arbitration and management layer outside (instead of inside) has more visibility into multiple systems which are part of single cluster and processes are spreaded across cgroup in each such system. While a logic inside can manage just a manage a process of single node which are using multiple cgroups. > Just to illustrate the point, please think of the alsa interface. We > expose hardware capabilities pretty much as-is leaving management and > multiplexing to userland and there's nothing wrong with it. It fits > better that way; however, we don't then go try to implement cgroup > controller for PCM channels. To do any high-level resource > management, you gotta do it where the said resource is actually > managed and arbitrated. > > What's the allocation frequency you're expecting? It might be better > to just let allocations themselves go through the agent that you're > planning. In that case we might need to build FUSE style infrastructure. Frequency for RDMA resource allocation is certainly less than read/write calls. > You sure can use cgroup membership to identify who's asking > tho. Given how the whole thing is architectured, I'd suggest thinking > more about how the whole thing should turn out eventually. > Yes, I agree. At this point, its software solution to provide resource isolation in simple manner which has scope to become adaptive in future. > Thanks. > > -- > tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-11 18:40 +0200 |
| Message-ID | <q7zwK-78E-21@gated-at.bofh.it> |
| In reply to | #1222950 |
Hello, Parav. On Fri, Sep 11, 2015 at 09:56:31PM +0530, Parav Pandit wrote: > Resource run away by application can lead to (a) kernel and (b) other > applications left out with no resources situation. Yeap, that this controller would be able to prevent to a reasonable extent. > Both the problems are the target of this patch set by accounting via cgroup. > > Performance contention can be resolved with higher level user space, > which will tune it. If individual applications are gonna be allowed to do that, what's to prevent them from jacking up their limits? So, I assume you're thinking of a central authority overseeing distribution and enforcing the policy through cgroups? > Threshold and fail counters are on the way in follow on patch. If you're planning on following what the existing memcg did in this area, it's unlikely to go well. Would you mind sharing what you have on mind in the long term? Where do you see this going? Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-11 18:50 +0200 |
| Message-ID | <q7zGq-7k9-13@gated-at.bofh.it> |
| In reply to | #1222900 |
> cpuset is a special case but think of cpu, memory or io controllers. > Their resource distribution schemes are a lot more developed than > what's proposed in this patchset and that's a necessity because nobody > wants to cripple their machines for resource control. IO controller and applications are mature in nature. When IO controller throttles the IO, applications are pretty mature where if IO takes longer to complete, there is possibly almost no way to cancel the system call or rather application might not want to cancel the IO at least the non asynchronous one. So application just notice lower performance than throttled way. Its really not possible at RDMA level with RDMA resource to hold up resource creation call for longer time, because reusing existing resource with failed status can likely to give better performance. As Doug explained in his example, many RDMA resources as its been used by applications are relatively long lived. So holding ups resource creation while its taken by other process will certainly will look bad on application performance front compare to returning failure and reusing existing one once its available or once new one is available. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-11 21:10 +0200 |
| Message-ID | <q7BRV-2f9-51@gated-at.bofh.it> |
| In reply to | #1222971 |
Hello, Parav. On Fri, Sep 11, 2015 at 10:17:42PM +0530, Parav Pandit wrote: > IO controller and applications are mature in nature. > When IO controller throttles the IO, applications are pretty mature > where if IO takes longer to complete, there is possibly almost no way > to cancel the system call or rather application might not want to > cancel the IO at least the non asynchronous one. I was more talking about the fact that they allow resources to be consumed when they aren't contended. > So application just notice lower performance than throttled way. > Its really not possible at RDMA level with RDMA resource to hold up > resource creation call for longer time, because reusing existing > resource with failed status can likely to give better performance. > As Doug explained in his example, many RDMA resources as its been used > by applications are relatively long lived. So holding ups resource > creation while its taken by other process will certainly will look bad > on application performance front compare to returning failure and > reusing existing one once its available or once new one is available. I'm not really sold on the idea that this can be used to implement performance based resource distribution. I'll write more about that on the other subthread. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Hefty, Sean" <sean.hefty@intel.com> |
|---|---|
| Date | 2015-09-11 21:30 +0200 |
| Message-ID | <q7Cbg-2BT-25@gated-at.bofh.it> |
| In reply to | #1222900 |
> So, the existence of resource limitations is fine. That's what we > deal with all the time. The problem usually with this sort of > interfaces which expose implementation details to users directly is > that it severely limits engineering manuevering space. You usually > want your users to express their intentions and a mechanism to > arbitrate resources to satisfy those intentions (and in a way more > graceful than "we can't, maybe try later?"); otherwise, implementing > any sort of high level resource distribution scheme becomes painful > and usually the only thing possible is preventing runaway disasters - > you don't wanna pin unused resource permanently if there actually is > contention around it, so usually all you can do with hard limits is > overcommiting limits so that it at least prevents disasters. I agree with Tejun that this proposal is at the wrong level of abstraction. If you look at just trying to limit QPs, it's not clear what that attempts to accomplish. Conceptually, a QP is little more than an addressable endpoint. It may or may not map to HW resources (for Intel NICs it does not). Even when HW resources do back the QP, the hardware is limited by how many QPs can realistically be active at any one time, based on how much caching is available in the NIC. Trying to limit the number of QPs that an app can allocate, therefore, just limits how much of the address space an app can use. There's no clear link between QP limits and HW resource limits, unless you assume a very specific underlying implementation. - Sean -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2015-09-11 21:50 +0200 |
| Message-ID | <q7CuC-2YI-17@gated-at.bofh.it> |
| In reply to | #1223057 |
On Fri, Sep 11, 2015 at 07:22:56PM +0000, Hefty, Sean wrote: > Trying to limit the number of QPs that an app can allocate, > therefore, just limits how much of the address space an app can use. > There's no clear link between QP limits and HW resource limits, > unless you assume a very specific underlying implementation. Isn't that the point though? We have several vendors with hardware that does impose hard limits on specific resources. There is no way to avoid that, and ultimately, those exact HW resources need to be limited. If we want to talk about abstraction, then I'd suggest something very general and simple - two limits: '% of the RDMA hardware resource pool' (per device or per ep?) 'bytes of kernel memory for RDMA structures' (all devices) That comfortably covers all the various kinds of hardware we support in a reasonable fashion. Unless there really is a reason why we need to constrain exactly and precisely PD/QP/MR/AH (I can't think of one off hand) The 'RDMA hardware resource pool' is a vendor-driver-device specific thing, with no generic definition beyond something that doesn't fit in the other limit. Jason -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Hefty, Sean" <sean.hefty@intel.com> |
|---|---|
| Date | 2015-09-11 22:10 +0200 |
| Message-ID | <q7CNX-3AD-5@gated-at.bofh.it> |
| In reply to | #1223069 |
> > Trying to limit the number of QPs that an app can allocate, > > therefore, just limits how much of the address space an app can use. > > There's no clear link between QP limits and HW resource limits, > > unless you assume a very specific underlying implementation. > > Isn't that the point though? We have several vendors with hardware > that does impose hard limits on specific resources. There is no way to > avoid that, and ultimately, those exact HW resources need to be > limited. My point is that limiting the number of QPs that an app can allocate doesn't necessarily mean anything. Is allocating 1000 QPs with 1 entry each better or worse than 1 QP with 10,000 entries? Who knows? > If we want to talk about abstraction, then I'd suggest something very > general and simple - two limits: > '% of the RDMA hardware resource pool' (per device or per ep?) > 'bytes of kernel memory for RDMA structures' (all devices) Yes - this makes more sense to me. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-14 13:10 +0200 |
| Message-ID | <q8zO2-4pn-7@gated-at.bofh.it> |
| In reply to | #1223076 |
On Sat, Sep 12, 2015 at 1:36 AM, Hefty, Sean <sean.hefty@intel.com> wrote: >> > Trying to limit the number of QPs that an app can allocate, >> > therefore, just limits how much of the address space an app can use. >> > There's no clear link between QP limits and HW resource limits, >> > unless you assume a very specific underlying implementation. >> >> Isn't that the point though? We have several vendors with hardware >> that does impose hard limits on specific resources. There is no way to >> avoid that, and ultimately, those exact HW resources need to be >> limited. > > My point is that limiting the number of QPs that an app can allocate doesn't necessarily mean anything. Is allocating 1000 QPs with 1 entry each better or worse than 1 QP with 10,000 entries? Who knows? I think it means if its RDMA RC QP, than whether you can talk to 1000 nodes or 1 node in network. When we deploy MPI application, it know the rank of the application, we know the cluster size of the deployment and based on that resource allocation can be done. If you meant to say from performance point of view, than resource count is possibly not the right measure. Just because we have not defined those interface for performance today in this patch set, doesn't mean that we won't do it. I could easily see a number_of_messages/sec as one interface to be added in future. But that won't stop process hoarders to stop taking away all the QPs, just the way we needed PID controller. Now when it comes to Intel implementation, if it driver layer knows (in future we new APIs) that whether 10 or 100 user QPs should map to few hw-QPs or more hw-QPs (uSNIC). so that hw-QP exposed to one cgroup is isolated from hw-QP exposed to other cgroup. If hw- implementation doesn't require isolation, it could just continue from single pool, its left to the vendor implementation on how to use this information (this API is not present in the patch). So cgroup can also provides a control point for vendor layer to tune internal resource allocation based on provided matrix, which cannot be done by just providing "memory usage by RDMA structures". If I have to compare it with other cgroup knobs, low level individual knobs by itself, doesn't serve any meaningful purpose either. Just by defined how much CPU to use or how much memory to use, it cannot define the application performance either. I am not sure, whether iocontroller can achieve 10 million IOPs by defining single CPU and 64KB of memory. all the knobs needs to be set in right way to reach desired number. In similar line RDMA resource knobs as individual knobs are not definition of performance, its just another knob. > >> If we want to talk about abstraction, then I'd suggest something very >> general and simple - two limits: >> '% of the RDMA hardware resource pool' (per device or per ep?) >> 'bytes of kernel memory for RDMA structures' (all devices) > > Yes - this makes more sense to me. > Sean, Jason, Help me to understand this scheme. 1. How does the % of resource, is different than absolute number? With rest of the cgroups systems we define absolute number at most places to my knowledge. Such as (a) number_of_tcp_bytes, (b) IOPs of block device, (c) cpu cycles etc. 20% of QP = 20 QPs when 100 QPs are with hw. I prefer to keep the resource scheme consistent with other resource control points - i.e. absolute number. 2. bytes of kernel memory for RDMA structures One QP of one vendor might consume X bytes and other Y bytes. How does the application knows how much memory to give. application can allocate 100 QP of each 1 entry deep or 1 QP of 100 entries deep as in Sean's example. Both might consume almost same memory. Application doing 100 QP allocation, still within limit of memory of cgroup leaves other applications without any QP. I don't see a point of memory footprint based scheme, as memory limits are well addressed by more smarter memory controller anyway. I do agree with Tejun, Sean on the point that abstraction level has to be different for using RDMA and thats why libfabrics and other interfaces are emerging which will take its own time to get stabilize, integrated. Until pure IB style RDMA programming model exist - based on RDMA resource based scheme, I think control point also has to be on resources. Once a stable abstraction level is on table (possibly across fabric not just RDMA), than a right resource controller can be implemented. Even when RDMA abstraction layer arrives, as Jason mentioned, at the end it would consume some hw resource anyway, that needs to be controlled too. Jason, If the hardware vendor defines the resource pool without saying its resource QP or MR, how would actually management/control point can decide what should be controlled to what limit? We will need additional user space library component to decode than, after that it needs to be abstracted out as QP or MR so that it can be deal in vendor agnostic way as application layer. and than it would look similar to what is being proposed here? -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-14 16:10 +0200 |
| Message-ID | <q8CCe-8ro-37@gated-at.bofh.it> |
| In reply to | #1224027 |
Hi Tejun, I missed to acknowledge your point that we need both - hard limit and soft limit/weight. Current patchset is only based on hard limit. I see that weight would be another helfpul layer in chain that we can implement after this as incremental that makes review, debugging manageable? Parav On Mon, Sep 14, 2015 at 4:39 PM, Parav Pandit <pandit.parav@gmail.com> wrote: > On Sat, Sep 12, 2015 at 1:36 AM, Hefty, Sean <sean.hefty@intel.com> wrote: >>> > Trying to limit the number of QPs that an app can allocate, >>> > therefore, just limits how much of the address space an app can use. >>> > There's no clear link between QP limits and HW resource limits, >>> > unless you assume a very specific underlying implementation. >>> >>> Isn't that the point though? We have several vendors with hardware >>> that does impose hard limits on specific resources. There is no way to >>> avoid that, and ultimately, those exact HW resources need to be >>> limited. >> >> My point is that limiting the number of QPs that an app can allocate doesn't necessarily mean anything. Is allocating 1000 QPs with 1 entry each better or worse than 1 QP with 10,000 entries? Who knows? > > I think it means if its RDMA RC QP, than whether you can talk to 1000 > nodes or 1 node in network. > When we deploy MPI application, it know the rank of the application, > we know the cluster size of the deployment and based on that resource > allocation can be done. > If you meant to say from performance point of view, than resource > count is possibly not the right measure. > > Just because we have not defined those interface for performance today > in this patch set, doesn't mean that we won't do it. > I could easily see a number_of_messages/sec as one interface to be > added in future. > But that won't stop process hoarders to stop taking away all the QPs, > just the way we needed PID controller. > > Now when it comes to Intel implementation, if it driver layer knows > (in future we new APIs) that whether 10 or 100 user QPs should map to > few hw-QPs or more hw-QPs (uSNIC). > so that hw-QP exposed to one cgroup is isolated from hw-QP exposed to > other cgroup. > If hw- implementation doesn't require isolation, it could just > continue from single pool, its left to the vendor implementation on > how to use this information (this API is not present in the patch). > > So cgroup can also provides a control point for vendor layer to tune > internal resource allocation based on provided matrix, which cannot be > done by just providing "memory usage by RDMA structures". > > If I have to compare it with other cgroup knobs, low level individual > knobs by itself, doesn't serve any meaningful purpose either. > Just by defined how much CPU to use or how much memory to use, it > cannot define the application performance either. > I am not sure, whether iocontroller can achieve 10 million IOPs by > defining single CPU and 64KB of memory. > all the knobs needs to be set in right way to reach desired number. > > In similar line RDMA resource knobs as individual knobs are not > definition of performance, its just another knob. > >> >>> If we want to talk about abstraction, then I'd suggest something very >>> general and simple - two limits: >>> '% of the RDMA hardware resource pool' (per device or per ep?) >>> 'bytes of kernel memory for RDMA structures' (all devices) >> >> Yes - this makes more sense to me. >> > > Sean, Jason, > Help me to understand this scheme. > > 1. How does the % of resource, is different than absolute number? With > rest of the cgroups systems we define absolute number at most places > to my knowledge. > Such as (a) number_of_tcp_bytes, (b) IOPs of block device, (c) cpu cycles etc. > 20% of QP = 20 QPs when 100 QPs are with hw. > I prefer to keep the resource scheme consistent with other resource > control points - i.e. absolute number. > > 2. bytes of kernel memory for RDMA structures > One QP of one vendor might consume X bytes and other Y bytes. How does > the application knows how much memory to give. > application can allocate 100 QP of each 1 entry deep or 1 QP of 100 > entries deep as in Sean's example. > Both might consume almost same memory. > Application doing 100 QP allocation, still within limit of memory of > cgroup leaves other applications without any QP. > I don't see a point of memory footprint based scheme, as memory limits > are well addressed by more smarter memory controller anyway. > > I do agree with Tejun, Sean on the point that abstraction level has to > be different for using RDMA and thats why libfabrics and other > interfaces are emerging which will take its own time to get stabilize, > integrated. > > Until pure IB style RDMA programming model exist - based on RDMA > resource based scheme, I think control point also has to be on > resources. > Once a stable abstraction level is on table (possibly across fabric > not just RDMA), than a right resource controller can be implemented. > Even when RDMA abstraction layer arrives, as Jason mentioned, at the > end it would consume some hw resource anyway, that needs to be > controlled too. > > Jason, > If the hardware vendor defines the resource pool without saying its > resource QP or MR, how would actually management/control point can > decide what should be controlled to what limit? > We will need additional user space library component to decode than, > after that it needs to be abstracted out as QP or MR so that it can be > deal in vendor agnostic way as application layer. > and than it would look similar to what is being proposed here? -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-14 17:30 +0200 |
| Message-ID | <q8DRD-1Is-17@gated-at.bofh.it> |
| In reply to | #1224187 |
Hello, Parav. On Mon, Sep 14, 2015 at 07:34:09PM +0530, Parav Pandit wrote: > I missed to acknowledge your point that we need both - hard limit and > soft limit/weight. Current patchset is only based on hard limit. > I see that weight would be another helfpul layer in chain that we can > implement after this as incremental that makes review, debugging > manageable? At this point, I'm very unsure that doing this as a cgroup controller is a good direction. From userland interface standpoint, publishing a cgroup controller is a big commitment. It is true that we haven't been doing a good job of gatekeeping or polishing controller interfaces but we're trying hard to change that and what's being proposed in this thread doesn't really seem to be mature enough. It's not even clear what's being identified as resources here are things that the users would actually care about or if it's even possible to implement sensible resource control in the kernel via the proposed resource restrictions. So, I'd suggest going back to the board and figuring out what the actual resources are, their distribution strategies should be and at which layer such strategies can be implemented best. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2015-09-14 19:30 +0200 |
| Message-ID | <q8FJM-4pD-5@gated-at.bofh.it> |
| In reply to | #1224027 |
On Mon, Sep 14, 2015 at 04:39:33PM +0530, Parav Pandit wrote: > 1. How does the % of resource, is different than absolute number? With > rest of the cgroups systems we define absolute number at most places > to my knowledge. There isn't really much choice if the abstraction is a bundle of all resources. You can't use an absolute number unless every possible hardware limited resource is defined, which doesn't seem smart to me either. It is not abstract enough, and doesn't match our universe of hardware very well. > 2. bytes of kernel memory for RDMA structures > One QP of one vendor might consume X bytes and other Y bytes. How does > the application knows how much memory to give. I don't see this distinction being useful at such a fine granularity where the control side needs to distinguish between 1 and 2 QPs. The majority use for control groups has been along with containers to prevent a container for exhausting resources in a way that impacts another. In that use model limiting each container to N MB of kernel memory makes it straightforward to reason about resource exhaustion in a multi-tennant environment. We have other controllers that do this, just more indirectly (ie limiting the number of inotifies, or the number of fds indirectly cap kernel memory consumption) ie Presumably some fairly small limitation like 10MB is enough for most non-MPI jobs. > Application doing 100 QP allocation, still within limit of memory of > cgroup leaves other applications without any QP. No, if the HW has a fixed QP pool then it would hit #1 above. Both are active at once. For example you'd say a container cannot use more than 10% of the device's hardware resources, or more than 10MB of kernel memory. If on an mlx card, you probably hit the 10% of QP resources first. If on an qib card there is no HW QP pool (well, almost, QPNs are always limited), so you'd hit the memory limit instead. In either case, we don't want to see a container able to exhaust either all of kernel memory or all of the HW resources to deny other containers. If you have a non-container use case in mind I'd be curious to hear it.. > I don't see a point of memory footprint based scheme, as memory limits > are well addressed by more smarter memory controller anyway. I don't thing #1 is controlled but another controller. This is long lived kernel-side memory allocations to support RDMA resource allocation - we certainly have nothing in the rdma layer that is tracking this stuff. > If the hardware vendor defines the resource pool without saying its > resource QP or MR, how would actually management/control point can > decide what should be controlled to what limit? In the kernel each HW driver has to be involved to declare what it's hardware resource limits are. In user space, it is just a simple limiter knob to prevent resource exhaustion. UAPI wise, nobdy has to care if the limit is actually # of QPs or something else. Jason -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-14 21:00 +0200 |
| Message-ID | <q8H8R-6mh-1@gated-at.bofh.it> |
| In reply to | #1224359 |
On Mon, Sep 14, 2015 at 10:58 PM, Jason Gunthorpe <jgunthorpe@obsidianresearch.com> wrote: > On Mon, Sep 14, 2015 at 04:39:33PM +0530, Parav Pandit wrote: > >> 1. How does the % of resource, is different than absolute number? With >> rest of the cgroups systems we define absolute number at most places >> to my knowledge. > > There isn't really much choice if the abstraction is a bundle of all > resources. You can't use an absolute number unless every possible > hardware limited resource is defined, which doesn't seem smart to me > either. Absolute number of percentage is representation for a given property. That property needs definition. Isn't it? How do we say that "Some undefined" resource you give certain amount, which user doesn't know about what to administer, or configure. It has to be quantifiable entity. It is not abstract enough, and doesn't match our universe of > hardware very well. > Why does the user need to know the actual hardware resource limits or define hardware based resource. RDMA verbs is the abstraction point. We could well define (a) how many number of RDMA connections are allowed instead of QP, or CQ or AH. (b) how many data transfer buffers to use. The fact is we have so many mid layers, which uses these resources differently, above abstraction does not fit the bill. But we know the mid layers how they operate, and how they use the RDMA resource keeping. So if we deploy MPI application for given cluster of container, we can accurately configure the RDMA resource, isn't it? Another example would be, if we don't want only 50% resources to be given to all containers and rest 50% to kernel consumers such as NFS, all containers can reside in single rdma cgroup limited to given limits. >> 2. bytes of kernel memory for RDMA structures >> One QP of one vendor might consume X bytes and other Y bytes. How does >> the application knows how much memory to give. > > I don't see this distinction being useful at such a fine granularity > where the control side needs to distinguish between 1 and 2 QPs. > > The majority use for control groups has been along with containers to > prevent a container for exhausting resources in a way that impacts > another. > Right. Thats the intention. > In that use model limiting each container to N MB of kernel memory > makes it straightforward to reason about resource exhaustion in a > multi-tennant environment. We have other controllers that do this, > just more indirectly (ie limiting the number of inotifies, or the > number of fds indirectly cap kernel memory consumption) > > ie Presumably some fairly small limitation like 10MB is enough for > most non-MPI jobs. Container application always write a simple for loop code to take away majority of QP with 10MB limit. > >> Application doing 100 QP allocation, still within limit of memory of >> cgroup leaves other applications without any QP. > > No, if the HW has a fixed QP pool then it would hit #1 above. Both are > active at once. For example you'd say a container cannot use more than > 10% of the device's hardware resources, or more than 10MB of kernel > memory. > Right. we need to define this resource pool, right? Why it cannot be verbs abstraction? How many resources are really used to implement verb layer in reality is left to hardware vendor Abstract pool just added confusion instead of clarity. Imagine instead of tcp_bytes or kmem bytes, its "some memory resource", how would someone debug/tune a system with abstract knobs? > If on an mlx card, you probably hit the 10% of QP resources first. If > on an qib card there is no HW QP pool (well, almost, QPNs are always > limited), so you'd hit the memory limit instead. > > In either case, we don't want to see a container able to exhaust > either all of kernel memory or all of the HW resources to deny other > containers. > > If you have a non-container use case in mind I'd be curious to hear > it.. Container is the prime case. Additionally equally prime case of non container use case. Today, application can take up all the resource being first class citizan, and NFS mount will fail. So without container also we should be able to restrict resources to user mode app. > >> I don't see a point of memory footprint based scheme, as memory limits >> are well addressed by more smarter memory controller anyway. > > I don't thing #1 is controlled but another controller. This is long > lived kernel-side memory allocations to support RDMA resource > allocation - we certainly have nothing in the rdma layer that is > tracking this stuff. > Some drivers performs mmap() of kernel memory to user space, some drivers does user space page allocation and maps to device. Putting or tracking all those is just so intrusive changes spreading down the vendor drivers or ib layer which may not be right way to track. Memory allocation tracking I believe should be left to memcg. >> If the hardware vendor defines the resource pool without saying its >> resource QP or MR, how would actually management/control point can >> decide what should be controlled to what limit? > > In the kernel each HW driver has to be involved to declare what it's > hardware resource limits are. > > In user space, it is just a simple limiter knob to prevent resource > exhaustion. > > UAPI wise, nobdy has to care if the limit is actually # of QPs or > something else. > If we dont care about resource, we cannot tune or limit it. number of MRs used by MPI vs rsocket vs accelio is way different. > Jason -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2015-09-14 22:20 +0200 |
| Message-ID | <q8Ioh-8kx-7@gated-at.bofh.it> |
| In reply to | #1224406 |
On Tue, Sep 15, 2015 at 12:24:41AM +0530, Parav Pandit wrote: > On Mon, Sep 14, 2015 at 10:58 PM, Jason Gunthorpe > <jgunthorpe@obsidianresearch.com> wrote: > > On Mon, Sep 14, 2015 at 04:39:33PM +0530, Parav Pandit wrote: > > > >> 1. How does the % of resource, is different than absolute number? With > >> rest of the cgroups systems we define absolute number at most places > >> to my knowledge. > > > > There isn't really much choice if the abstraction is a bundle of all > > resources. You can't use an absolute number unless every possible > > hardware limited resource is defined, which doesn't seem smart to me > > either. > > Absolute number of percentage is representation for a given property. > That property needs definition. Isn't it? > How do we say that "Some undefined" resource you give certain amount, > which user doesn't know about what to administer, or configure. > It has to be quantifiable entity. Each vendor can quantify exactly what HW resources their implementation has and how the above limit impacts their card. There will be many variations, and IIRC, some vendors have resource pools not directly related to the standard PD/QP/MR/CQ/AH verbs resources. > > It is not abstract enough, and doesn't match our universe of > > hardware very well. > Why does the user need to know the actual hardware resource limits or > define hardware based resource. Because actual hardware resources *ARE* the limit. We cannot abstract it away. The hardware/driver has real, fixed, immutable limits. No API abstraction can possibly change that. The limits are such there *IS NO* API boundary that can bundle them into something simpler. There will always be apps that require wildly different ratios of the basic verbs resources (PD/QP/CQ/AH/MR) Either we control each and every vendor's limited resource directly (which is where you started), or we just roll them up into a 'all resource' bundle and control them indirectly. There just isn't a mythical third 'better API' choice with the hardware we have today. > (a) how many number of RDMA connections are allowed instead of QP, or CQ or AH. > (b) how many data transfer buffers to use. None of that accurately reflects what the real HW limits actually are. > > ie Presumably some fairly small limitation like 10MB is enough for > > most non-MPI jobs. > > Container application always write a simple for loop code to take away > majority of QP with 10MB limit. No, the HW and kmem limits must work together, the HW limit would prevent exhaustion outside the container. > Imagine instead of tcp_bytes or kmem bytes, its "some memory > resource", how would someone debug/tune a system with abstract knobs? Well, we have the memcg controller that does track kmem. The subsystem specific kmem limit is to force fair sharing of the limited kmem resource within the overall memcg limit. They are complementary. A fictional rdma_kmem and tcp_kmem would serve very similar purposes. > > UAPI wise, nobdy has to care if the limit is actually # of QPs or > > something else. > If we dont care about resource, we cannot tune or limit it. number of > MRs used by MPI vs rsocket vs accelio is way different. So? I don't think it is really important to have an exact, precise, limit. The HW pools are pretty big, unless you plan to run tens of thousands of containers eacg with tiny RDMA limits, it is fine to talk in broader terms (ie 10% of all HW limited resource) which is totally adaquate to hard-prevent run away or exhaustion scenarios. Jason -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-15 05:10 +0200 |
| Message-ID | <q8ON4-GN-7@gated-at.bofh.it> |
| In reply to | #1224437 |
> Because actual hardware resources *ARE* the limit. We cannot abstract > it away. The hardware/driver has real, fixed, immutable limits. No API > abstraction can possibly change that. > > The limits are such there *IS NO* API boundary that can bundle them > into something simpler. There will always be apps that require wildly > different ratios of the basic verbs resources (PD/QP/CQ/AH/MR) > > Either we control each and every vendor's limited resource directly > (which is where you started), or we just roll them up into a 'all > resource' bundle and control them indirectly. There just isn't a > mythical third 'better API' choice with the hardware we have today. > As you precisely described, about wild ratio, we are asking vendor driver (bottom most layer) to statically define what the resource pool is, without telling him which application are we going to run to use those pool. Therefore vendor layer cannot ever define "right" resource pool. If we try to fix defining "right" resource pool, we will have to come up with API to modify/tune individual element of the pool. Once we bring that complexity, it becomes what is proposed in this pachset. Instead of bringing such complex solution, that affecting all the layers which solves the same problem as this patch, its better to keep definition of "bundle" in the user library/application deployment engine. where bundle is set of those resources. May be instead of having invidividual files for each resource, at user interface level, we can have rdma.bundle file. this bundle cgroup file defines these resources such as "ah 100 mr 100 qp 10" > So? I don't think it is really important to have an exact, precise, > limit. The HW pools are pretty big, unless you plan to run tens of > thousands of containers eacg with tiny RDMA limits, it is fine to talk > in broader terms (ie 10% of all HW limited resource) which is totally > adaquate to hard-prevent run away or exhaustion scenarios. > rdma cgroup will allow us to run post 512 or 1024 containers without using PCIe SR-IOV, without creating any vendor specific resource pools. > Jason -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2015-09-15 05:50 +0200 |
| Message-ID | <q8PpM-1pA-1@gated-at.bofh.it> |
| In reply to | #1224614 |
On Tue, Sep 15, 2015 at 08:38:54AM +0530, Parav Pandit wrote: > As you precisely described, about wild ratio, > we are asking vendor driver (bottom most layer) to statically define > what the resource pool is, without telling him which application are > we going to run to use those pool. > Therefore vendor layer cannot ever define "right" resource pool. No, I'm saying the resource pool is *well defined* and *fixed* by each hardware. The only question is how do we expose the N resource limits, the list of which is totally vendor specific. Yes, using a % scheme fixes the ratios, 1% is going to be a certain number of PD's, QP's, MRs, CQ's, etc at a ratio fixed by the driver configuration. That is the trade off for API simplicity. Yes, this results in some resources being over provisioned. I have no idea if that is usable for the workloads people want to run.. But *there is no middle option*. Either each and every single hardware limited resources has a dedicated per-container limit, or they are *somehow* bundled and the ratios become fixed. If Tejun says we can't have something so emphemeral as a vendor specific list of hardware resource pools - then what choice is left? > Instead of bringing such complex solution, that affecting all the > layers which solves the same problem as this patch, > its better to keep definition of "bundle" in the user > library/application deployment engine. > where bundle is set of those resources. The kernel has to do the restriction, so at some point you are telling the kernel to limit each and every unique resource the HW has, which is back to the original patch set, munging how the data is passed makes no difference to the basic objection, IMHO. > rdma cgroup will allow us to run post 512 or 1024 containers without > using PCIe SR-IOV, without creating any vendor specific resource > pools. If you ignore any vendor specific resource limits then you've just left open a hole, a wayward container can exhaust all others - so what was the point of doing all this work? Jason -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-16 06:50 +0200 |
| Message-ID | <q9cPn-2q1-7@gated-at.bofh.it> |
| In reply to | #1224622 |
Hi Jason, Sean, Tejun, I am in process of defining new approach, design based on the feedback given here for new RDMA cgroup from all of you. I have also collected feedback from Liran yesterday and ORNL folks too. Soon I will post the new approach, high level APIs and functionality for review before submitting actual implementation. Regards, Parav Pandit On Tue, Sep 15, 2015 at 9:15 AM, Jason Gunthorpe <jgunthorpe@obsidianresearch.com> wrote: > On Tue, Sep 15, 2015 at 08:38:54AM +0530, Parav Pandit wrote: > >> As you precisely described, about wild ratio, >> we are asking vendor driver (bottom most layer) to statically define >> what the resource pool is, without telling him which application are >> we going to run to use those pool. >> Therefore vendor layer cannot ever define "right" resource pool. > > No, I'm saying the resource pool is *well defined* and *fixed* by each > hardware. > > The only question is how do we expose the N resource limits, the list > of which is totally vendor specific. > >> rdma cgroup will allow us to run post 512 or 1024 containers without >> using PCIe SR-IOV, without creating any vendor specific resource >> pools. > > If you ignore any vendor specific resource limits then you've just > left open a hole, a wayward container can exhaust all others - so what > was the point of doing all this work? > > Jason -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-14 12:20 +0200 |
| Message-ID | <q8z1F-3fP-25@gated-at.bofh.it> |
| In reply to | #1223057 |
On Sat, Sep 12, 2015 at 12:52 AM, Hefty, Sean <sean.hefty@intel.com> wrote: >> So, the existence of resource limitations is fine. That's what we >> deal with all the time. The problem usually with this sort of >> interfaces which expose implementation details to users directly is >> that it severely limits engineering manuevering space. You usually >> want your users to express their intentions and a mechanism to >> arbitrate resources to satisfy those intentions (and in a way more >> graceful than "we can't, maybe try later?"); otherwise, implementing >> any sort of high level resource distribution scheme becomes painful >> and usually the only thing possible is preventing runaway disasters - >> you don't wanna pin unused resource permanently if there actually is >> contention around it, so usually all you can do with hard limits is >> overcommiting limits so that it at least prevents disasters. > > I agree with Tejun that this proposal is at the wrong level of abstraction. > > If you look at just trying to limit QPs, it's not clear what that attempts to accomplish. Conceptually, a QP is little more than an addressable endpoint. It may or may not map to HW resources (for Intel NICs it does not). Even when HW resources do back the QP, the hardware is limited by how many QPs can realistically be active at any one time, based on how much caching is available in the NIC. > cgroups as it stands today provides resource controls in effective manner of existing defined resource, such as cpu cycles, memory in user and kernel space, tcp bytes, IOPS etc. Similarly RDMA programming model defines its own set of resources which is used by applications which accesses those resources directly. What we are debating here is that, RDMA exposing hardware resources is not correct, and therefore whether a cgroup controller is needed or not. There are two points here. 1. Whether RDMA programming model is correct or not which works on defined resources of IB spec. 2. Assuming that programming model is fine, (because we have actively maintained IB stack in kernel and adoption of user space components in OS), whether we need to control those resources or not via cgroup. Tejun trying to say that because point_1 is doesn't seem to be right way to solve problem, point_2 should not be done or done at different level of abstraction. More questions/comments in Jason and Sean thread. Sean, Even though there is no one to one map of verb-QP to hw-QP, in order for driver or lower layer to effectively map the right verb-QP to hw-QP, such vendor specific layer needs to know how is it going to be used. Otherwise two contending applications for a QP may not get the right number of hw-QPs to use. > Trying to limit the number of QPs that an app can allocate, therefore, just limits how much of the address space an app can use. There's no clear link between QP limits and HW resource limits, unless you assume a very specific underlying implementation. > > - Sean -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Parav Pandit <pandit.parav@gmail.com> |
|---|---|
| Date | 2015-09-11 06:50 +0200 |
| Message-ID | <q7orE-7Yk-5@gated-at.bofh.it> |
| In reply to | #1222518 |
On Fri, Sep 11, 2015 at 9:34 AM, Tejun Heo <tj@kernel.org> wrote: > Hello, Parav. > > On Fri, Sep 11, 2015 at 09:09:58AM +0530, Parav Pandit wrote: >> The fact is that user level application uses hardware resources. >> Verbs layer is software abstraction for it. Drivers are hiding how >> they implement this QP or CQ or whatever hardware resource they >> project via API layer. >> For all of the userland on top of verb layer I mentioned above, the >> common resource abstraction is these resources AH, QP, CQ, MR etc. >> Hardware (and driver) might have different view of this resource in >> their real implementation. >> For example, verb layer can say that it has 100 QPs, but hardware >> might actually have 20 QPs that driver decide how to efficiently use >> it. > > My uneducated suspicion is that the abstraction is just not developed > enough. It should be possible to virtualize these resources through, > most likely, time-sharing to the level where userland simply says "I > want this chunk transferred there" and OS schedules the transfer > prioritizing competing requests. Tejun, That is such a perfect abstraction to have at OS level, but not sure how much close it can be to bare metal RDMA it can be. I have started discussion on that front as well as part of other thread, but its certainly long way to go. Most want to enjoy the performance benefit of the bare metal interfaces it provides. Such abstraction that you mentioned, exists, the only difference is instead of its OS as central entity, its the higher level libraries, drivers and hw together does it today for the applications. > > It could be that given the use cases rdma might not need such level of > abstraction - e.g. most users want to be and are pretty close to bare > metal, but, if that's true, it also kinda is weird to build > hierarchical resource distribution scheme on top of such bare > abstraction. > > ... >> > I don't know. What's proposed in this thread seems way too low level >> > to be useful anywhere else. Also, what if there are multiple devices? >> > Is that a problem to worry about? >> >> o.k. It doesn't have to be useful anywhere else. If it suffice the >> need of RDMA applications, its fine for near future. >> This patch allows limiting resources across multiple devices. >> As we go along the path, and if requirement come up to have knob on >> per device basis, thats something we can extend in future. > > You kinda have to decide that upfront cuz it gets baked into the > interface. Well, all the interfaces are not yet defined. Except the test and benchmark utilities, real world applications wouldn't really bother much about which device are they are going through. so I expect that per device level control would nice for very specific applications, but I don't anticipate that in first place. If others have different view, I would be happy to hear that. Even if we extend per device control, I would expect per cgroup control at top level without which its uncontrolled access. > >> > I'm kinda doubtful we're gonna have too many of these. Hardware >> > details being exposed to userland this directly isn't common. >> >> Its common in RDMA applications. Again they may not be real hardware >> resource, its just API layer which defines those RDMA constructs. > > It's still a very low level of abstraction which pretty much gets > decided by what the hardware and driver decide to do. > >> > I'd say keep it simple and do the minimum. :) >> >> o.k. In that case new rdma cgroup controller which does rdma resource >> accounting is possibly the most simplest form? >> Make sense? > > So, this fits cgroup's purpose to certain level but it feels like > we're trying to build too much on top of something which hasn't > developed sufficiently. I suppose it could be that this is the level > of development that rdma is gonna reach and dumb cgroup controller can > be useful for some use cases. I don't know, so, yeah, let's keep it > simple and avoid doing crazy stuff. > o.k. thanks. I would wait for some more time to collect more feedback. In absence of that, I will send updated patch V1 which will include, (a) functionality of this patch in new rdma cgroup as you recommended, (b) fixes for comments from Haggai for this patch (c) more fixes which I have done in mean time > Thanks. > > -- > tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-11 17:10 +0200 |
| Message-ID | <q7y7D-5eQ-21@gated-at.bofh.it> |
| In reply to | #1222526 |
Hello, Parav. On Fri, Sep 11, 2015 at 10:13:59AM +0530, Parav Pandit wrote: > > My uneducated suspicion is that the abstraction is just not developed > > enough. It should be possible to virtualize these resources through, > > most likely, time-sharing to the level where userland simply says "I > > want this chunk transferred there" and OS schedules the transfer > > prioritizing competing requests. > > Tejun, > That is such a perfect abstraction to have at OS level, but not sure > how much close it can be to bare metal RDMA it can be. > I have started discussion on that front as well as part of other > thread, but its certainly long way to go. > Most want to enjoy the performance benefit of the bare metal > interfaces it provides. Yeah, sure, I'm not trying to say that rdma needs or should do that. > Such abstraction that you mentioned, exists, the only difference is > instead of its OS as central entity, its the higher level libraries, > drivers and hw together does it today for the applications. But more that having resource control in the OS and actual arbitration higher up in the stack isn't likely to lead to an effective resource distribution scheme. > > You kinda have to decide that upfront cuz it gets baked into the > > interface. > > Well, all the interfaces are not yet defined. Except the test and I meant the cgroup interface. > benchmark utilities, real world applications wouldn't really bother > much about which device are they are going through. Weights can work fine across multiple devices. Hard limits don't. It just doesn't make any sense. Unless you can exclude multiple device scenarios, you'll have to implement per-device limits. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Hefty, Sean" <sean.hefty@intel.com> |
|---|---|
| Date | 2015-09-10 19:50 +0200 |
| Message-ID | <q7e8W-8uF-7@gated-at.bofh.it> |
| In reply to | #1222307 |
> > In past there has been similar comment to have dedicated cgroup > > controller for RDMA instead of merging with device cgroup. > > I am ok with both the approach, however I prefer to utilize device > > controller instead of spinning of new controller for new devices > > category. > > I anticipate more such need would arise and for new device category, > > it might not be worth to have new cgroup controller. > > RapidIO though very less popular and upcoming PCIe are on horizon to > > offer similar benefits as that of RDMA and in future having one > > controller for each of them again would not be right approach. > > > > I certainly seek your and others inputs in this email thread here > whether > > (a) to continue to extend device cgroup (which support character, > > block devices white list) and now RDMA devices > > or > > (b) to spin of new controller, if so what are the compelling reasons > > that it can provide compare to extension. > > I'm doubtful that these things are gonna be mainstream w/o building up > higher level abstractions on top and if we ever get there we won't be > talking about MR or CQ or whatever. Also, whatever next-gen is > unlikely to have enough commonalities when the proposed resource knobs > are this low level, so let's please keep it separate, so that if/when > this goes out of fashion for one reason or another, the controller can > silently wither away too. As an attempt to abstract the hardware resources only, what these devices are exposing to apps can be viewed as command queues (RDMA QPs and SRQs), notification queues (RDMA CQs and EQs), and space in the device cache and allocated memory (RDMA MRs and AHs, maybe PDs). If one wanted a higher level of abstraction, associations exist between these resources. For example, command queues feed into notification queues. Address handles are required resources to use an unconnected queue pair. - Sean -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web