Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1372625 > unrolled thread
| Started by | Tejun Heo <tj@kernel.org> |
|---|---|
| First post | 2016-04-06 18:00 +0200 |
| Last post | 2016-04-08 05:20 +0200 |
| Articles | 20 on this page of 25 — 4 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Tejun Heo <tj@kernel.org> - 2016-04-06 18:00 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Peter Zijlstra <peterz@infradead.org> - 2016-04-07 08:50 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Johannes Weiner <hannes@cmpxchg.org> - 2016-04-07 09:40 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Peter Zijlstra <peterz@infradead.org> - 2016-04-07 10:10 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Johannes Weiner <hannes@cmpxchg.org> - 2016-04-07 11:30 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Peter Zijlstra <peterz@infradead.org> - 2016-04-07 12:50 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Tejun Heo <tj@kernel.org> - 2016-04-07 21:50 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Peter Zijlstra <peterz@infradead.org> - 2016-04-07 22:30 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Tejun Heo <tj@kernel.org> - 2016-04-08 22:20 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-04-09 08:20 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Peter Zijlstra <peterz@infradead.org> - 2016-04-09 15:40 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Tejun Heo <tj@kernel.org> - 2016-04-13 00:30 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-04-13 09:50 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Tejun Heo <tj@kernel.org> - 2016-04-13 18:10 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-04-13 21:20 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-04-14 08:10 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Tejun Heo <tj@kernel.org> - 2016-04-14 22:00 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-04-15 04:50 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Peter Zijlstra <peterz@infradead.org> - 2016-04-09 18:10 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-04-07 10:10 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Peter Zijlstra <peterz@infradead.org> - 2016-04-07 10:30 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Johannes Weiner <hannes@cmpxchg.org> - 2016-04-07 21:10 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Peter Zijlstra <peterz@infradead.org> - 2016-04-07 21:40 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Johannes Weiner <hannes@cmpxchg.org> - 2016-04-07 22:30 +0200
Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-04-08 05:20 +0200
Page 1 of 2 [1] 2 Next page →
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2016-04-06 18:00 +0200 |
| Subject | Re: [PATCHSET RFC cgroup/for-4.6] cgroup, sched: implement resource group and PRIO_RGRP |
| Message-ID | <rkY26-68Y-5@gated-at.bofh.it> |
Hello, Peter.
Sorry about the delay.
On Mon, Mar 14, 2016 at 12:30:13PM +0100, Peter Zijlstra wrote:
> On Fri, Mar 11, 2016 at 10:41:18AM -0500, Tejun Heo wrote:
> > * A rgroup is a cgroup which is invisible on and transparent to the
> > system-level cgroupfs interface.
> >
> > * A rgroup can be created by specifying CLONE_NEWRGRP flag, along with
> > CLONE_THREAD, during clone(2). A new rgroup is created under the
> > parent thread's cgroup and the new thread is created in it.
>
> This seems overly restrictive. As you well know there's people moving
> threads about after creation.
Will get to this later.
> Also, with this interface the whole thing cannot be used until your
> libc's pthread_create() has been patched to allow use of this new flag.
This isn't difficult to change but is this a problem in the long term?
Once added, this is gonna be a permanent part of API and I think we
better get it right than quick. If this is a concern, we can go for a
setsid(2) style syscall so that users can have an easier access to it.
> > * A rgroup is automatically destroyed when empty.
>
> Except for Zombies it appears..
Zombies do hold onto its rgroup but at that point the rgroup is
draining refs and can't be populated again. The same state as rmdir'd
cgroup with zombies.
> > * A top-level rgroup of a process is a rgroup whose parent cgroup is a
> > sgroup. A process may have multiple top-level rgroups and thus
> > multiple rgroup subtrees under the same parent sgroup.
> >
> > * Unlike sgroups, rgroups are allowed to compete against peer threads.
> > Each rgroup behaves equivalent to a sibling task.
> >
> > * rgroup subtrees are local to the process. When the process forks or
> > execs, its rgroup subtrees are collapsed.
> >
> > * When a process is migrated to a different cgroup, its rgroup
> > subtrees are preserved.
>
> This all makes it impossible to say put a single thread outside of the
> hierarchy forced upon it by the process. Like putting a RT thread in an
> isolated group on the side.
>
> Which is a rather common thing to do.
I don't think the mentioned RT case is problematic. Depending on the
desired outcome,
1. If the admin doesn't want the program to be able to meddle with the
cpu resource control at all, it can just disable the cpu controller
in subtree_control (this is tied to the parent control now but will
be moved to the associated cgroup itself). The application will
create rgroup hierarchy but won't be able to use CPU resource
control and the admin would be able to treat all threads as if they
don't have rgroups at all.
2. If the admin still wants to allow the application to retain CPU
resource control, unless the said program is actively getting in
the way, the admin can set the limits the way it wants along the
hierarchy down to the specific thread.
Note that #1 can be done after-the-fact. The admin can revoke CPU
controller access anytime. For example, assuming the following
hierarchy (cX is a cgroup, rX is a rgroup, NNN are threads).
cA - 234
+ r235 - 235
+ 236
If the process 234 configured CPU resource control in a specific way
and the admin wants to override, the admin can simply do "echo -cpu >
cA.subtree_control". Afterwards, as far as CPU resource control is
concerned, all threads will behave as if there are no rgroups at all
and the admin can tweak the settings of individual threads using the
usual scheduler systemcalls.
> > rgroup lays the foundation for other kernel mechanisms to make use of
> > resource controllers while providing proper isolation between system
> > management and in-process operations removing the awkward and
> > layer-violating requirement for coordination between individual
> > applications and system management. On top of the rgroup mechanism,
> > PRIO_RGRP is implemented for {set|get}priority(2).
> >
> > * PRIO_RGRP can only be used if the target task is already in a
> > rgroup. If setpriority(2) is used and cpu controller is available,
> > cpu controller is enabled until the target rgroup is covered and the
> > specified nice value is set as the weight of the rgroup.
> >
> > * The specified nice value has the same meaning as for tasks. For
> > example, a rgroup and a task competing under the same parent would
> > behave exactly the same as two tasks.
> >
> > * For top-level rgroups, PRIO_RGRP follows the same rlimit
> > restrictions as PRIO_PROCESS; however, as nested rgroups only
> > distribute CPU cycles which are allocated to the process, no
> > restriction is applied.
>
> While this appears neat, I doubt it will remain so in the face of this:
>
> > * A mechanism that applications can use to publish certain rgroups so
> > that external entities can determine which IDs to use to change
> > rgroup settings. I already have interface and implementation design
> > mostly pinned down.
>
> So you need some new fangled way to set/query all the other possible
> cgroup parameters supported, and then suddenly you have one that has two
> possible interface. That's way ugly.
So, the above response is a bit confusing because publishing rgroups
doesn't require setting or querying all other possible cgroup
parameters.
Regarding the need to add separeate interface for each control knob
for rgroups,
1. There aren't many knobs which make sense for in-process control to
begin with.
2. As shown in this patchset's modifiction to setpriority(2), for
stuff which makes sense, we're likely to already have constructs
which already deal with the issue (it is a needed capability with
or without cgroup). The right way forward is seamlessly extending
existing interfaces.
3. If there is no exactly matching interface, we want to add them for
both groups and threads in a way which is consistent with other
syscalls which deal with related issues, especially for the
scheduler.
Just in case, here's a more concerete explanation about publishing
rgroups. The only addition needed for external access is a way to
determine which ID maps to which rgroup - a proc file listing (rgroup
name, ID) pairs.
Let's say the program creates the following internal hierarchy.
cgroup - service0 - highpri_workers
+ lowpri_workers
+ service1 - highpri_workers
+ lowpri_workers
The rgroups are published by a member thread performing, for example,
prctl(PR_SET_RGROUP_NAME, "service0.highpri_workers"). The only thing
it does is pinning the pid so that it stays associated with the rgroup
and publishes it in proc as follows.
# cat /proc/234/rgroups
service0.highpri_workers 240
service0.lowpri_workers 241
service1.highpri_workers 248
service1.lowpri_workers 249
From tooling side, renice(2) can be extended to understand rgroups so
that something like the following works.
# renice -n -10 -r 234:service1.highpri_workers
Thanks.
--
tejun
[toc] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-04-07 08:50 +0200 |
| Message-ID | <rlbVo-8hU-5@gated-at.bofh.it> |
| In reply to | #1372625 |
So I recently got made aware of the fact that cgroupv2 doesn't allow tasks to be associated with !leaf cgroups, this is yet another capability of cpu-cgroup you've destroyed. At this point I really don't see why I should spend another second considering anything v2. So full NAK and stop wasting my time.
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2016-04-07 09:40 +0200 |
| Message-ID | <rlcHL-pI-1@gated-at.bofh.it> |
| In reply to | #1373108 |
On Thu, Apr 07, 2016 at 08:45:49AM +0200, Peter Zijlstra wrote: > So I recently got made aware of the fact that cgroupv2 doesn't allow > tasks to be associated with !leaf cgroups, this is yet another > capability of cpu-cgroup you've destroyed. May I ask how you are using that? The behavior for tasks in !leaf groups was fairly inconsistent across controllers because they all did different things, or didn't handle it at all. For example, the block controller in v1 implements separate weight knobs for the group as a subtree root as well as for the tasks only inside the group itself. But it didn't do so for bandwith limits. The memory controller on the other hand only had a singular set of the controls that applied to both the local tasks and all subgroups. And I know Google had a lot of trouble with that because they ended up with basically uncontrollable leftover cache in the top-level group of some subtree that would put pressure on the real workload leafgroups below. There was a lot of back and forth whether we should add a second set of knobs just to control the local tasks separately from the subtree, but ended up concluding that the situation can be expressed more clearly by creating dedicated leaf subgroups for stuff like management software and launchers instead, so that their memory pools/LRUs are clearly delineated from other groups and seperately controllable. And we couldn't think of any meaningful configuration that could not be expressed in that scheme. I mean, it's the same thing, right? Only that with tasks in !leaf groups the controller would have to emulate a hidden leaf subgroup and provide additional interfacing, and without it the leaf groups are explicit and a single set of knobs suffices. I.e. it seems more of a convenience thing than actual functionality, but one that forces ugly redundancy in the interface. So it was a nice cleanup for the memory controller and I believe the IO controller as well. I'd be curious how it'd be a problem for CPU?
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-04-07 10:10 +0200 |
| Message-ID | <rldaN-RS-1@gated-at.bofh.it> |
| In reply to | #1373122 |
On Thu, Apr 07, 2016 at 03:35:47AM -0400, Johannes Weiner wrote: > On Thu, Apr 07, 2016 at 08:45:49AM +0200, Peter Zijlstra wrote: > > So I recently got made aware of the fact that cgroupv2 doesn't allow > > tasks to be associated with !leaf cgroups, this is yet another > > capability of cpu-cgroup you've destroyed. > > May I ask how you are using that? _I_ use a kernel with CONFIG_CGROUPS=n (yes really). But seriously? You have to ask? The root cgroup is per definition not a leaf, and all tasks start life there, and some cannot be ever moved out. Therefore _everybody_ uses this. > The behavior for tasks in !leaf groups was fairly inconsistent across > controllers because they all did different things, or didn't handle it > at all. Then they're all bloody broken, because fully hierarchical was an early requirement for cgroups; I know, because I had to throw away many days of work and start over with cgroup support when they did that. > So it was a nice cleanup for the memory controller and I believe the > IO controller as well. I'd be curious how it'd be a problem for CPU? The full hierarchy took years to make work and is fully ingrained with how the thing words, changing it isn't going to be nice or easy. So sure, go with a lowest common denominator, instead of fixing shit, yay for progress :/
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2016-04-07 11:30 +0200 |
| Message-ID | <rleqf-1G9-23@gated-at.bofh.it> |
| In reply to | #1373133 |
On Thu, Apr 07, 2016 at 10:08:33AM +0200, Peter Zijlstra wrote: > On Thu, Apr 07, 2016 at 03:35:47AM -0400, Johannes Weiner wrote: > > On Thu, Apr 07, 2016 at 08:45:49AM +0200, Peter Zijlstra wrote: > > > So I recently got made aware of the fact that cgroupv2 doesn't allow > > > tasks to be associated with !leaf cgroups, this is yet another > > > capability of cpu-cgroup you've destroyed. > > > > May I ask how you are using that? > > _I_ use a kernel with CONFIG_CGROUPS=n (yes really). > > But seriously? You have to ask? > > The root cgroup is per definition not a leaf, and all tasks start life > there, and some cannot be ever moved out. > > Therefore _everybody_ uses this. Hm? The root group can always contain tasks. It's not the only thing the root is exempt from, it can't control any resources either: sched_group_set_shares(): /* * We can't change the weight of the root cgroup. */ if (!tg->se[0]) return -EINVAL; tg_set_cfs_bandwidth(): if (tg == &root_task_group) return -EINVAL; etc. and all the problems that led to this rule stem from resource control. > > The behavior for tasks in !leaf groups was fairly inconsistent across > > controllers because they all did different things, or didn't handle it > > at all. > > Then they're all bloody broken, because fully hierarchical was an early > requirement for cgroups; I know, because I had to throw away many days > of work and start over with cgroup support when they did that. I think we're talking past each other. They're all fully hierarchical in the sense of accounting and divvying up resources along a tree structure, and configurable groups competing with other configurable groups or subtrees. That all works perfectly fine. It's the concept of loose unconfigurable tasks competing with configured groups or subtrees that invites problems. It's not a question of implementation, it's that the configurations that people created with e.g. the memory controller repeatedly ended up creating the same problems and the same stupid patches to add the local-only knobs (which the cpu cgroup doesn't have either AFAICS). This is not some gratuitous cutting away of convenience, it's hours and hours of discussions both on the mailinglists and at conferences about such lovely stuff as to whether the memory lowlimit (softlimit) should apply to only the local memory pool or hierarchically because that user happened to have memory pools in !leaf nodes which they had to control somehow. Swear to god. [ And yes, the root group IS "loose unconfigurable tasks" that compete with configured subtrees. But that is very explicit in the interface and you move stuff that consumes significant resources and needs to be controlled out of the root group; it doesn't have the same issue. ] If that happens once or twice I'm willing to write it off as PEBCAK, but if otherwise competent users like Google repeatedly create configurations that lead to these problems, and then end up pushing and lobbying in this case for non-hierarchical knobs to work around problems in the structural organization of the workloads, it's more likely that the interface is shit. So we added a rule that doesn't take away any functionality, but it forces you to organize your workloads more explicitely to take away that ambiguity.
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-04-07 12:50 +0200 |
| Message-ID | <rlfFE-2v3-15@gated-at.bofh.it> |
| In reply to | #1373201 |
On Thu, Apr 07, 2016 at 05:28:24AM -0400, Johannes Weiner wrote: > Hm? The root group can always contain tasks. It's not the only thing > the root is exempt from, it can't control any resources either: it does in fact control resouces; the hierarchy directly affects the proportional distribution of time. > sched_group_set_shares(): > > /* > * We can't change the weight of the root cgroup. > */ > if (!tg->se[0]) > return -EINVAL; The root has, per definition, no siblings, so setting a weight is entirely pointless. > tg_set_cfs_bandwidth(): > > if (tg == &root_task_group) > return -EINVAL; > We have had patches to implement this, but have so far held off because they add a bunch of cycles to some really hot paths and we'd rather not do that. Its not impossible, or unthinkable to do this otherwise.
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2016-04-07 21:50 +0200 |
| Message-ID | <rlo6e-o4-17@gated-at.bofh.it> |
| In reply to | #1373133 |
Hello, Peter. On Thu, Apr 07, 2016 at 10:08:33AM +0200, Peter Zijlstra wrote: > On Thu, Apr 07, 2016 at 03:35:47AM -0400, Johannes Weiner wrote: > > So it was a nice cleanup for the memory controller and I believe the > > IO controller as well. I'd be curious how it'd be a problem for CPU? > > The full hierarchy took years to make work and is fully ingrained with > how the thing words, changing it isn't going to be nice or easy. > > So sure, go with a lowest common denominator, instead of fixing shit, > yay for progress :/ It's easy to get fixated on what each subsystem can do and develop towards different directions siloed in each subsystem. That's what we've had for quite a while in cgroup. Expectedly, this sends off controllers towards different directions. Direct competion between tasks and child cgroups was one of the main sources of balkanization. The balkanization was no coincidence either. Tasks and cgroups are different types of entities and don't have the same control knobs or follow the same lifetime rules. For absolute limits, it isn't clear how much of the parent's resources should be distributed to internal children as opposed to child cgroups. People end up depending on specific implementation details and proposing one-off hacks and interface additions. Proportional weights aren't much better either. CPU has internal mapping between nice values and shares and treat them equally, which can get confusing as the configured weights behave differently depending on how many threads are in the parent cgroup which often is opaque and can't be controlled from outside. Widely diverging from CPU's behavior, IO grouped all internal tasks into an internal leaf node and used to assign a fixed weight to it. Now, you might think that none of it matters and each subsystem treating cgroup hierarchy as arbitrary and orthogonal collections of bean counters is fine; however, that makes it impossible to account for and control operations which span different types of resources. This prevented us from implementing resource control over frigging buffered writes, making the whole IO control thing a joke. While CPU currently doesn't directly tie into it, that is only because CPU cycles spent during writeback isn't yet properly accounted. The structural constraints and resulting consistency don't just subtract from the abilities of each controller. It establishes a common base, the shared resource domains and consistent behaviors on top of them, that further capabilities can be built upon, capabilities as fundamental as comprehensive resource control over buffered writeback. It can be convenient to have subsystem-specific raw bean counters. If that's what the use case calls for, individual controllers can easily be moved to a separate hierarchy although it would naturally lose the capabilities coming from cooperating over shared resource domains. However, please understand that there are a lot of use cases where comprehensive and consistent resource accounting and control over all major resources is useful and necessary. Thanks. -- tejun
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-04-07 22:30 +0200 |
| Message-ID | <rloIV-Ti-3@gated-at.bofh.it> |
| In reply to | #1373712 |
On Thu, Apr 07, 2016 at 03:45:55PM -0400, Tejun Heo wrote: > Hello, Peter. > > On Thu, Apr 07, 2016 at 10:08:33AM +0200, Peter Zijlstra wrote: > > On Thu, Apr 07, 2016 at 03:35:47AM -0400, Johannes Weiner wrote: > > > So it was a nice cleanup for the memory controller and I believe the > > > IO controller as well. I'd be curious how it'd be a problem for CPU? > > > > The full hierarchy took years to make work and is fully ingrained with > > how the thing words, changing it isn't going to be nice or easy. > > > > So sure, go with a lowest common denominator, instead of fixing shit, > > yay for progress :/ > > It's easy to get fixated on what each subsystem can do and develop > towards different directions siloed in each subsystem. That's what > we've had for quite a while in cgroup. Expectedly, this sends off > controllers towards different directions. Direct competion between > tasks and child cgroups was one of the main sources of balkanization. > > The balkanization was no coincidence either. Tasks and cgroups are > different types of entities and don't have the same control knobs or > follow the same lifetime rules. For absolute limits, it isn't clear > how much of the parent's resources should be distributed to internal > children as opposed to child cgroups. People end up depending on > specific implementation details and proposing one-off hacks and > interface additions. Yes, I'm familiar with the problem; but simply mandating leaf only nodes is not a solution, for the very simple fact that there are tasks in the root cgroup that cannot ever be moved out, so we _must_ be able to deal with !leaf nodes containing tasks. A consistent interface for absolute controllers to divvy up the resources between local tasks and child cgroups isn't _that_ hard. And this leaf only business totally screwed over anything proportional. This simply cannot work. > Proportional weights aren't much better either. CPU has internal > mapping between nice values and shares and treat them equally, which > can get confusing as the configured weights behave differently > depending on how many threads are in the parent cgroup which often is > opaque and can't be controlled from outside. Huh what? There's nothing confusing there, the nice to weight mapping is static and can easily be consulted. Alternatively we can make an interface where you can set weight through nice values, for those people that are afraid of numbers. But the configured weights do _not_ behave differently depending on the number of tasks, they behave exactly as specified in the proportional weight based rate distribution. We've done the math.. > Widely diverging from > CPU's behavior, IO grouped all internal tasks into an internal leaf > node and used to assign a fixed weight to it. That's just plain broken... That is not how a proportional weight based hierarchical controller works. > Now, you might think that none of it matters and each subsystem > treating cgroup hierarchy as arbitrary and orthogonal collections of > bean counters is fine; however, that makes it impossible to account > for and control operations which span different types of resources. > This prevented us from implementing resource control over frigging > buffered writes, making the whole IO control thing a joke. While CPU > currently doesn't directly tie into it, that is only because CPU > cycles spent during writeback isn't yet properly accounted. CPU cycles spend in waitqueues aren't properly accounted to whoever queued the job either, and there's a metric ton of async stuff that's not properly accounted, so what? > However, please understand that there are a lot of use cases where > comprehensive and consistent resource accounting and control over all > major resources is useful and necessary. Maybe, but so far I've only heard people complain this v2 thing didn't work for them, and as far as I can see the whole v2 model is internally inconsistent and impossible to implement. The suggestion by Johannes to adjust the leaf node weight depending on the number of tasks in is so ludicrous I don't even know where to start enumerating the fail.
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2016-04-08 22:20 +0200 |
| Message-ID | <rlL2O-Ou-19@gated-at.bofh.it> |
| In reply to | #1373720 |
Hello, Peter. On Thu, Apr 07, 2016 at 10:25:42PM +0200, Peter Zijlstra wrote: > > The balkanization was no coincidence either. Tasks and cgroups are > > different types of entities and don't have the same control knobs or > > follow the same lifetime rules. For absolute limits, it isn't clear > > how much of the parent's resources should be distributed to internal > > children as opposed to child cgroups. People end up depending on > > specific implementation details and proposing one-off hacks and > > interface additions. > > Yes, I'm familiar with the problem; but simply mandating leaf only nodes > is not a solution, for the very simple fact that there are tasks in the > root cgroup that cannot ever be moved out, so we _must_ be able to deal > with !leaf nodes containing tasks. As Johannes already pointed out, the root cgroup has always been special. While pure practicality, performance implications and implementation convenience do play important roles in the special treatment, another constributing aspect is avoiding exposing statistics and control knobs which are duplicates of and/or conflicting with what's already available at the system level. It's never fun to have multiple sources of truth. > A consistent interface for absolute controllers to divvy up the > resources between local tasks and child cgroups isn't _that_ hard. I've spent months thinking about it and didn't get too far. If you have a good solution, I'd be happy to be enlightened. Also, please note that the current solution is based on restricting certain configurations. If we can find a better solution, we can relax the relevant constraints and move onto it without breaking compatibility. > And this leaf only business totally screwed over anything proportional. > > This simply cannot work. Will get to this below. > > Proportional weights aren't much better either. CPU has internal > > mapping between nice values and shares and treat them equally, which > > can get confusing as the configured weights behave differently > > depending on how many threads are in the parent cgroup which often is > > opaque and can't be controlled from outside. > > Huh what? There's nothing confusing there, the nice to weight mapping is > static and can easily be consulted. Alternatively we can make an > interface where you can set weight through nice values, for those people > that are afraid of numbers. > > But the configured weights do _not_ behave differently depending on the > number of tasks, they behave exactly as specified in the proportional > weight based rate distribution. We've done the math.. Yes, once one understands what's going on, it isn't confusing. It's just not something users can intuitively understand from the presented interface. The confusion of course is worsened severely by different controller behaviors. > > Widely diverging from > > CPU's behavior, IO grouped all internal tasks into an internal leaf > > node and used to assign a fixed weight to it. > > That's just plain broken... That is not how a proportional weight based > hierarchical controller works. That's a strong statement. When the hierarchy is composed of equivalent objects as in CPU, not distinguishing internal and leaf nodes would be a more natural way to organize; however, it isn't necessarily true in all cases. For example, while a writeback IO would be issued by some task, the task itself might not have done anything to cause that IO and the IO would essentially be anonymous in the resource domain. Also, different controllers use different units of organization - CPU sees threads, IO sees IO contexts which are usually shared in a process. The difference would lead to differing scaling behaviors in proportional distribution. While the separate buckets and entities model may not be as elegant as tree of uniform objects, it is far from uncommon and more robust when dealing with different types of objects. > > Now, you might think that none of it matters and each subsystem > > treating cgroup hierarchy as arbitrary and orthogonal collections of > > bean counters is fine; however, that makes it impossible to account > > for and control operations which span different types of resources. > > This prevented us from implementing resource control over frigging > > buffered writes, making the whole IO control thing a joke. While CPU > > currently doesn't directly tie into it, that is only because CPU > > cycles spent during writeback isn't yet properly accounted. > > CPU cycles spend in waitqueues aren't properly accounted to whoever > queued the job either, and there's a metric ton of async stuff that's > not properly accounted, so what? The ultimate goal of cgroup resource control is accounting and controlling all significant resource consumptions as configured. Some system operations are inherently global and others are simply too cheap to justify the overhead; however, there still are significant aggregate operations which are being missed out including almost everything taking place in the writeback path. So, yes, we eventually want to be able to account for them, of course in a way which doesn't get in the way of actual operation. > > However, please understand that there are a lot of use cases where > > comprehensive and consistent resource accounting and control over all > > major resources is useful and necessary. > > Maybe, but so far I've only heard people complain this v2 thing didn't > work for them, and as far as I can see the whole v2 model is internally > inconsistent and impossible to implement. I suppose we live in different bubbles. Can you please elaborate which parts of cgroup v2 model are internally inconsistent and impossible to implement? I'd be happy to rectify the situation. > The suggestion by Johannes to adjust the leaf node weight depending on > the number of tasks in is so ludicrous I don't even know where to start > enumerating the fail. That sounds like a pretty uncharitable way to read his message. I think he was trying to find out the underlying requirements so that a way forward can be discussed. I do have the same question. It's difficult to have discussions about trade-offs without knowing where the requirements are coming from. Do you have something on mind for cases where internal tasks have to compete with sibling cgroups? Thanks. -- tejun
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-04-09 08:20 +0200 |
| Message-ID | <rlUpr-FX-3@gated-at.bofh.it> |
| In reply to | #1374451 |
On Fri, 2016-04-08 at 16:11 -0400, Tejun Heo wrote: > > That's just plain broken... That is not how a proportional weight based > > hierarchical controller works. > > That's a strong statement. When the hierarchy is composed of > equivalent objects as in CPU, not distinguishing internal and leaf > nodes would be a more natural way to organize; however... You almost said it yourself, you want to make the natural organization of cpu, cpuacct and cpuset controllers a prohibited organization. There is no "however..." that can possibly justify that. It's akin to mandating: henceforth "horse" shall be spelled "cow", riders thereof shall teach their "cow" the proper enunciation of "moo". It's silly. Like it or not, these controllers have thread encoded in their DNA, it's an integral part of what they are, and how real users in the real world use them. No rationalization will change that cold hard fact. Make an "Aunt Tilly" button for those incapable of comprehending the complexities if you will, but please don't make cgroups so rigid and idiot proof that only idiots (hi system thing) can use it. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-04-09 15:40 +0200 |
| Message-ID | <rm1hg-5Mq-15@gated-at.bofh.it> |
| In reply to | #1374451 |
On Fri, Apr 08, 2016 at 04:11:35PM -0400, Tejun Heo wrote: > > > Widely diverging from > > > CPU's behavior, IO grouped all internal tasks into an internal leaf > > > node and used to assign a fixed weight to it. > > > > That's just plain broken... That is not how a proportional weight based > > hierarchical controller works. > > That's a strong statement. No its plain fact. If you modify a graph, it is not the same graph. Even if you argue by merit of the function on this graph, and state that only the result of this function is important, and any modification to the graph that leaves this result in tact is good; ie. a modification invariant to the function, this fails. Because for proportional controllers all that matters is the number and weight of edges leaving a node. The modification described above does clearly change the outcome and is not invariant under the proportional weight distribution function. > When the hierarchy is composed of > equivalent objects as in CPU, not distinguishing internal and leaf > nodes would be a more natural way to organize; however, it isn't > necessarily true in all cases. For example, while a writeback IO > would be issued by some task, the task itself might not have done > anything to cause that IO and the IO would essentially be anonymous in > the resource domain. Also, different controllers use different units > of organization - CPU sees threads, IO sees IO contexts which are > usually shared in a process. The difference would lead to differing > scaling behaviors in proportional distribution. > > While the separate buckets and entities model may not be as elegant as > tree of uniform objects, it is far from uncommon and more robust when > dealing with different types of objects. The graph does not care about the type of objects the nodes represent, and proportional weight distribution only cares about the edges. With cpu-cgroup the nodes are not of uniform type either, they can be a group or a task. You get runtime type identification and make it work. There just isn't an excuse for crazy crap like this. Its wrong, no two ways about it.
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2016-04-13 00:30 +0200 |
| Message-ID | <rneYO-6JD-1@gated-at.bofh.it> |
| In reply to | #1374630 |
Hello, Peter. On Sat, Apr 09, 2016 at 03:39:17PM +0200, Peter Zijlstra wrote: > > While the separate buckets and entities model may not be as elegant as > > tree of uniform objects, it is far from uncommon and more robust when > > dealing with different types of objects. > > The graph does not care about the type of objects the nodes represent, > and proportional weight distribution only cares about the edges. > > With cpu-cgroup the nodes are not of uniform type either, they can be a > group or a task. You get runtime type identification and make it work. > > There just isn't an excuse for crazy crap like this. Its wrong, no two > ways about it. Abstracing tasks and groups as equivalent objects works well for the scheduler and that's great. This is also because the domain lends itself very well to such simple and elegant approach. The only entities of interest are tasks, as you and Mike pointed out earlier in the thread, and group priority can be easily mapped to task priority. However, this isn't necessarily the case for other controllers. There's also the issue of mapping the model to absolute controllers. For the uniform model to work, there must be a way to treat internal and leaf entities in the same way. For memory, the leaf entities are processes and applying the same model would mean that memory controller would have to implement equivalent per-process control knobs. We don't have that. In fact, we can't have that - a significant part of memory consumption can't be attached to a single process. There is a fundamental distinction between internal and leaf nodes in the memory resource graph. We aren't designing a spherical cow in a vacuum, and, I believe, should aspire to make pragmatic trade-offs of all involved factors. If multiple controllers co-operating on the same resource domains is beneficial and required, we should figure out a way to make different controllers agree and that way most likely will require some trade-offs from various controllers. Given the currently known requirements and constraints, restricting internal competition is a simple and straight-forward way to isolate leaf node handling details of different controllers. The cost is part aesthetical and part practical. While less elegant than tree of uniform objects, it seems a stretch to call internal / leaf node distinction broken especially given that the model is natural to some controllers. The practical cost is loss of the ability to let leaf entities compete against groups. However, we can't evaluate how important such capability is without actual use-cases. If there are important ones, please bring them up, so that we can examine the actual requirements and try to find a good trade-off to support them. I understand that CPU controller getting constrained due to other controllers can feel frustrating; however, the constraint is there to solve practical problems which hopefully are being explained in this conversation. If there is a better trade-off, we can easily get rid of it and move on, but such decision can only be made considering all the relevant factors. If you can think of a better solution, let's please discuss it. Thanks. -- tejun
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-04-13 09:50 +0200 |
| Message-ID | <rnnIK-64i-7@gated-at.bofh.it> |
| In reply to | #1377356 |
On Tue, 2016-04-12 at 18:29 -0400, Tejun Heo wrote: > Hello, Peter. > > On Sat, Apr 09, 2016 at 03:39:17PM +0200, Peter Zijlstra wrote: > > > While the separate buckets and entities model may not be as elegant as > > > tree of uniform objects, it is far from uncommon and more robust when > > > dealing with different types of objects. > > > > The graph does not care about the type of objects the nodes represent, > > and proportional weight distribution only cares about the edges. > > > > With cpu-cgroup the nodes are not of uniform type either, they can be a > > group or a task. You get runtime type identification and make it work. > > > > There just isn't an excuse for crazy crap like this. Its wrong, no two > > ways about it. > > Abstracing tasks and groups as equivalent objects works well for the > scheduler and that's great. This is also because the domain lends > itself very well to such simple and elegant approach. The only > entities of interest are tasks, as you and Mike pointed out earlier in > the thread, and group priority can be easily mapped to task priority. > However, this isn't necessarily the case for other controllers. > > There's also the issue of mapping the model to absolute controllers. > For the uniform model to work, there must be a way to treat internal > and leaf entities in the same way. For memory, the leaf entities are > processes and applying the same model would mean that memory > controller would have to implement equivalent per-process control > knobs. We don't have that. In fact, we can't have that - a > significant part of memory consumption can't be attached to a single > process. There is a fundamental distinction between internal and leaf > nodes in the memory resource graph. > > We aren't designing a spherical cow in a vacuum, and, I believe, > should aspire to make pragmatic trade-offs of all involved factors. > If multiple controllers co-operating on the same resource domains is > beneficial and required, we should figure out a way to make different > controllers agree and that way most likely will require some > trade-offs from various controllers. > > Given the currently known requirements and constraints, restricting > internal competition is a simple and straight-forward way to isolate > leaf node handling details of different controllers. > > The cost is part aesthetical and part practical. While less elegant > than tree of uniform objects, it seems a stretch to call internal / > leaf node distinction broken especially given that the model is > natural to some controllers. That justifies prohibiting proper usages of three controllers, cpu, cpuacct and cpuset? > The practical cost is loss of the ability to let leaf entities compete > against groups. However, we can't evaluate how important such > capability is without actual use-cases. If there are important ones, > please bring them up, so that we can examine the actual requirements > and try to find a good trade-off to support them. Hm, I though Google did that, and I know I mentioned another gigabuck sized outfit. Whatever, ob trade-off.. Another cpuset example is something I was asked to look into recently. There are folks out in the real world who want to run RT guests. Now VIRTUAL REALtime tickles my funny-bone, but I piddled around with it nonetheless to see what such can deliver (not much). System thing and/or libvirt created a cpuset home for qemu, but with VPUs sharing CPU with other qemu threads and the rest of the world, RT performance in little virtual box was as pathetic as one would expect. What did I do about it? Among others, the obvious, I created an exclusive cpuset, and distributed qemu contexts having different requirements among context containment vessels having the required properties. I won't be doing any more of that particular scenario, but certainly will want to distribute various contexts among various context containment vessels in future. I soon enough won't care about cgroups, but others will surely expect cpu, cpuacct and cpuset controllers to continue to function properly. > I understand that CPU controller getting constrained due to other > controllers can feel frustrating; however, the constraint is there to > solve practical problems which hopefully are being explained in this > conversation. If there is a better trade-off, we can easily get rid > of it and move on, but such decision can only be made considering all > the relevant factors. If you can think of a better solution, let's > please discuss it. None here. Any artificial restriction placed on controllers will render same broken in one way or another that will matter to someone somewhere. Making something less than it was will do that. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2016-04-13 18:10 +0200 |
| Message-ID | <rnvwE-3Pk-49@gated-at.bofh.it> |
| In reply to | #1377611 |
Hello, Mike. On Wed, Apr 13, 2016 at 09:43:01AM +0200, Mike Galbraith wrote: > > The cost is part aesthetical and part practical. While less elegant > > than tree of uniform objects, it seems a stretch to call internal / > > leaf node distinction broken especially given that the model is > > natural to some controllers. > > That justifies prohibiting proper usages of three controllers, cpu, > cpuacct and cpuset? Neither cpuacct or cpuset loses any capability from the constraint as there is no difference between tasks being in an internal cgroup or a leaf cgroup nested under it. The only practical impact is that we lose the ability to let internal tasks compete against sibling cgroups for proportional control. > > The practical cost is loss of the ability to let leaf entities compete > > against groups. However, we can't evaluate how important such > > capability is without actual use-cases. If there are important ones, > > please bring them up, so that we can examine the actual requirements > > and try to find a good trade-off to support them. > > Hm, I though Google did that, and I know I mentioned another gigabuck > sized outfit. Whatever, ob trade-off.. Are you saying that you're aware that google or another big outfit is making active use of internal tasks competing against sibling cgroups for proportional CPU distribution? If so, can you please be more specific? > Another cpuset example is something I was asked to look into recently. First of all, as mentioned above, cpuset isn't affected at all in practical terms. Besides, for a very specialized cpuset setup, the cpuset configuration might not have anything to do with the resource domains other controllers use and it might make sense to keep cpuset on a separate hierarchy. > > I understand that CPU controller getting constrained due to other > > controllers can feel frustrating; however, the constraint is there to > > solve practical problems which hopefully are being explained in this > > conversation. If there is a better trade-off, we can easily get rid > > of it and move on, but such decision can only be made considering all > > the relevant factors. If you can think of a better solution, let's > > please discuss it. > > None here. Any artificial restriction placed on controllers will > render same broken in one way or another that will matter to someone > somewhere. Making something less than it was will do that. The specifics of gains and losses are what I've been trying to clarify in this thread. Hopefully, what we can gain from sharing common resource domains is clear by now. The practical cost is loss of the capability to let internal tasks compete against sibling cgroups for proportional control. However, to determine the weight of this cost, we have to know which use-cases call for it. Thanks. -- tejun
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-04-13 21:20 +0200 |
| Message-ID | <rnyuu-6aN-15@gated-at.bofh.it> |
| In reply to | #1378085 |
On Wed, 2016-04-13 at 11:59 -0400, Tejun Heo wrote: > Are you saying that you're aware that google or another big outfit is > making active use of internal tasks competing against sibling cgroups > for proportional CPU distribution? If so, can you please be more > specific? What I'm aware of is a big outfit that moves thread pool workers in/out of a large number of cpu/cpuacct cgroups. What all a worker thread may spawn in a cgroup, or find already there upon arrival and thus compete with I do not know. I'm baffled by why anyone would care which entity competes with which other entity. An entity is an entity is an entity. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-04-14 08:10 +0200 |
| Message-ID | <rnIDw-5In-13@gated-at.bofh.it> |
| In reply to | #1378085 |
On Wed, 2016-04-13 at 11:59 -0400, Tejun Heo wrote: > Hello, Mike. > > On Wed, Apr 13, 2016 at 09:43:01AM +0200, Mike Galbraith wrote: > > > The cost is part aesthetical and part practical. While less > > > elegant > > > than tree of uniform objects, it seems a stretch to call internal > > > / > > > leaf node distinction broken especially given that the model is > > > natural to some controllers. > > > > That justifies prohibiting proper usages of three controllers, cpu, > > cpuacct and cpuset? > > Neither cpuacct or cpuset loses any capability from the constraint as > there is no difference between tasks being in an internal cgroup or a > leaf cgroup nested under it. The only practical impact is that we > lose the ability to let internal tasks compete against sibling cgroups > for proportional control. I'm not getting it. A. entity = task[s] | cgroup[s] B. entity = task[s] ^ cgroup[s] A I get, B I don't, but you seem to be saying B, else we get the task competes with sibling cgroup business. Let /foo be an exclusive cpuset containing exclusive subset bar. How can any task acquire set foo affinity if B really really applies? My box calls me a dummy if I try to create a "proper" home for tasks, one with both no snobby neighbors and proper affinity. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2016-04-14 22:00 +0200 |
| Message-ID | <rnVAJ-7pv-11@gated-at.bofh.it> |
| In reply to | #1378502 |
Hello, Mike.
On Thu, Apr 14, 2016 at 08:07:37AM +0200, Mike Galbraith wrote:
> Let /foo be an exclusive cpuset containing exclusive subset bar.
> How can any task acquire set foo affinity if B really really
> applies? My box calls me a dummy if I try to create a "proper" home
> for tasks, one with both no snobby neighbors and proper affinity.
I'm not sure I quite understand what you're saying. Are you referring
to the cpuset.{cpu|mem}_exclusive knobs?
Thanks.
--
tejun
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-04-15 04:50 +0200 |
| Message-ID | <ro1Zv-41K-1@gated-at.bofh.it> |
| In reply to | #1379278 |
On Thu, 2016-04-14 at 15:57 -0400, Tejun Heo wrote:
> Hello, Mike.
>
> On Thu, Apr 14, 2016 at 08:07:37AM +0200, Mike Galbraith wrote:
> > Let /foo be an exclusive cpuset containing exclusive subset bar.
> > How can any task acquire set foo affinity if B really really
> > applies? My box calls me a dummy if I try to create a "proper"
> > home
> > for tasks, one with both no snobby neighbors and proper affinity.
>
> I'm not sure I quite understand what you're saying. Are you referring
> to the cpuset.{cpu|mem}_exclusive knobs?
Yes.
-Mike
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-04-09 18:10 +0200 |
| Message-ID | <rm3Cq-7IV-7@gated-at.bofh.it> |
| In reply to | #1374451 |
On Fri, Apr 08, 2016 at 04:11:35PM -0400, Tejun Heo wrote: > > Yes, I'm familiar with the problem; but simply mandating leaf only nodes > > is not a solution, for the very simple fact that there are tasks in the > > root cgroup that cannot ever be moved out, so we _must_ be able to deal > > with !leaf nodes containing tasks. > > As Johannes already pointed out, the root cgroup has always been > special. The root of the tree isn't special except for 2 properties. - it _is_ a root; iow, it doesn't have any incoming edges. This also means it doesn't have a parent; nor can have a weight, since that is an edge propery, not a node property. - it always exists; for without a root there is no tree. Making it _more_ special is silly. > > Maybe, but so far I've only heard people complain this v2 thing didn't > > work for them, and as far as I can see the whole v2 model is internally > > inconsistent and impossible to implement. > > I suppose we live in different bubbles. Can you please elaborate > which parts of cgroup v2 model are internally inconsistent and > impossible to implement? I'd be happy to rectify the situation. The fact that we have to deal with tasks in the root cgroup while not allowing tasks in any other node is internally inconsistent. If I can deal with tasks in one node (root) I can equally deal with tasks in any other node in exactly the same manner. Making it different is actually _more_ code.
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-04-07 10:10 +0200 |
| Message-ID | <rldaP-RS-15@gated-at.bofh.it> |
| In reply to | #1373122 |
On Thu, 2016-04-07 at 03:35 -0400, Johannes Weiner wrote: > On Thu, Apr 07, 2016 at 08:45:49AM +0200, Peter Zijlstra wrote: > > So I recently got made aware of the fact that cgroupv2 doesn't > > allow > > tasks to be associated with !leaf cgroups, this is yet another > > capability of cpu-cgroup you've destroyed. > > May I ask how you are using that? One real world usage is the thread pool servicing customers on a cgroup = account basis that I outlined. -Mike
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web