Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1573733 > unrolled thread
| Started by | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| First post | 2017-02-04 20:20 +0100 |
| Last post | 2017-02-07 20:50 +0100 |
| Articles | 20 on this page of 42 — 12 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-04 20:20 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Amir Goldstein <amir73il@gmail.com> - 2017-02-05 09:00 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-06 02:20 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Amir Goldstein <amir73il@gmail.com> - 2017-02-06 08:00 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-06 15:50 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount "J. R. Okajima" <hooanon05g@gmail.com> - 2017-02-06 04:30 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Amir Goldstein <amir73il@gmail.com> - 2017-02-06 07:40 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-06 17:40 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-06 07:50 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Theodore Ts'o <tytso@mit.edu> - 2017-02-06 16:00 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-06 16:20 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount lkml@pengaru.com - 2017-02-06 16:40 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-06 18:40 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount bfields@fieldses.org (J. Bruce Fields) - 2017-02-06 23:00 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-07 01:20 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount "J. Bruce Fields" <bfields@fieldses.org> - 2017-02-07 02:40 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-07 20:10 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Christoph Hellwig <hch@infradead.org> - 2017-02-07 20:50 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount "J. R. Okajima" <hooanon05g@gmail.com> - 2017-02-06 17:30 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Christoph Hellwig <hch@infradead.org> - 2017-02-07 10:20 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Djalal Harouni <tixxdz@gmail.com> - 2017-02-07 10:40 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Christoph Hellwig <hch@infradead.org> - 2017-02-07 11:00 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-07 17:40 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Amir Goldstein <amir73il@gmail.com> - 2017-02-07 19:00 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Christoph Hellwig <hch@infradead.org> - 2017-02-07 19:20 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-07 20:10 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Christoph Hellwig <hch@infradead.org> - 2017-02-07 21:00 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-07 21:10 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Amir Goldstein <amir73il@gmail.com> - 2017-02-07 22:10 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Christoph Hellwig <hch@infradead.org> - 2017-02-07 23:50 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-08 00:50 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Amir Goldstein <amir73il@gmail.com> - 2017-02-08 08:00 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Konstantin Khlebnikov <khlebnikov@yandex-team.ru> - 2017-02-08 13:10 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-08 16:00 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-08 16:30 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Josh Triplett <josh@joshtriplett.org> - 2017-02-08 03:00 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-08 16:30 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Josh Triplett <josh@joshtriplett.org> - 2017-02-09 11:40 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-09 17:40 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount ebiederm@xmission.com (Eric W. Biederman) - 2017-02-13 11:30 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount James Bottomley <James.Bottomley@HansenPartnership.com> - 2017-02-07 19:30 +0100
Re: [RFC 1/1] shiftfs: uid/gid shifting bind mount Djalal Harouni <tixxdz@gmail.com> - 2017-02-07 20:50 +0100
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Djalal Harouni <tixxdz@gmail.com> |
|---|---|
| Date | 2017-02-07 10:40 +0100 |
| Message-ID | <t8apI-12P-19@gated-at.bofh.it> |
| In reply to | #1575497 |
Hi, On Tue, Feb 7, 2017 at 10:19 AM, Christoph Hellwig <hch@infradead.org> wrote: > On Sat, Feb 04, 2017 at 11:19:32AM -0800, James Bottomley wrote: >> This allows any subtree to be uid/gid shifted and bound elsewhere. It >> does this by operating simlarly to overlayfs. Its primary use is for >> shifting the underlying uids of filesystems used to support >> unpriviliged (uid shifted) containers. The usual use case here is >> that the container is operating with an uid shifted unprivileged root >> but sometimes needs to make use of or work with a filesystem image >> that has root at real uid 0. >> >> The mechanism is to allow any subordinate mount namespace to mount a >> shiftfs filesystem (by marking it FS_USERNS_MOUNT) but only allowing >> it to mount marked subtrees (using the -o mark option as root). Once >> mounted, the subtree is mapped via the super block user namespace so >> that the interior ids of the mounting user namespace are the ids >> written to the filesystem. > > Please move this into VFS instead of a stackable fs. We might need > addtional parameters to getattr/setattr to specify the ID translation, > but that's why better than a horrible hack like this. I proposed an RFC months ago which implements all of this at the VFS layer [1], I received some feedback especially from Dave Chinner, however I failed to fix my bugs and improve it not enough resources... The problems discussed here about a new filesystem: inodes numbers, quota and many other things where all noted in that thread and previous threads about shiftfs. We are turning this to a heavy problem compared to all other namespaces... other namespaces integrate perfectly with other subsystems and the rest of layers, there is no special treatment... Christoph, for the getattr/setattr it won't work since internally the resolved path may point to a different mount context where we do not want the ID translation, and we may end up using the wrong vfsmount. A simple getattr/setattr won't work unless there are bigger changes too... [1] https://lkml.org/lkml/2016/5/4/411 -- tixxdz
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2017-02-07 11:00 +0100 |
| Message-ID | <t8aJ4-1am-15@gated-at.bofh.it> |
| In reply to | #1575513 |
On Tue, Feb 07, 2017 at 10:39:48AM +0100, Djalal Harouni wrote: > I proposed an RFC months ago which implements all of this at the VFS > layer [1], I received some feedback especially from Dave Chinner, > however I failed to fix my bugs and improve it not enough resources... And none of the issues goes away by hiding them in a stackable fs, in fact many of them are getting worse.
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2017-02-07 17:40 +0100 |
| Message-ID | <t8gYa-5gJ-29@gated-at.bofh.it> |
| In reply to | #1575497 |
On Tue, 2017-02-07 at 01:19 -0800, Christoph Hellwig wrote: > On Sat, Feb 04, 2017 at 11:19:32AM -0800, James Bottomley wrote: > > This allows any subtree to be uid/gid shifted and bound elsewhere. > > It does this by operating simlarly to overlayfs. Its primary use > > is for shifting the underlying uids of filesystems used to support > > unpriviliged (uid shifted) containers. The usual use case here is > > that the container is operating with an uid shifted unprivileged > > root but sometimes needs to make use of or work with a filesystem > > image that has root at real uid 0. > > > > The mechanism is to allow any subordinate mount namespace to mount > > a shiftfs filesystem (by marking it FS_USERNS_MOUNT) but only > > allowing it to mount marked subtrees (using the -o mark option as > > root). Once mounted, the subtree is mapped via the super block > > user namespace so that the interior ids of the mounting user > > namespace are the ids written to the filesystem. > > Please move this into VFS instead of a stackable fs. We might need > addtional parameters to getattr/setattr to specify the ID > translation, but that's why better than a horrible hack like this. I would need a lot more than that: getattr controls the cosmetic permission display to the user, but enforcement is done in the core permission checks which are inode based. To make this a real bind mount, the core permission checks will have to become subtree aware because knowledge of whether we need a uid shift in the permission check becomes a subtree property. Effectively inode_permission would become dentry_permission and generic_permission would take a dentry instead of an inode. This will be a huge amount of VFS and underlying filesystem churn, since the permissions calls are threaded through a huge chunk of code. Is this the approach that you really want? I suppose I could see the security people linking it because all the security hooks in the permission code become path aware. James
[toc] | [prev] | [next] | [standalone]
| From | Amir Goldstein <amir73il@gmail.com> |
|---|---|
| Date | 2017-02-07 19:00 +0100 |
| Message-ID | <t8idA-5XI-19@gated-at.bofh.it> |
| In reply to | #1575861 |
On Tue, Feb 7, 2017 at 6:37 PM, James Bottomley <James.Bottomley@hansenpartnership.com> wrote: > On Tue, 2017-02-07 at 01:19 -0800, Christoph Hellwig wrote: >> On Sat, Feb 04, 2017 at 11:19:32AM -0800, James Bottomley wrote: >> > This allows any subtree to be uid/gid shifted and bound elsewhere. >> > It does this by operating simlarly to overlayfs. Its primary use >> > is for shifting the underlying uids of filesystems used to support >> > unpriviliged (uid shifted) containers. The usual use case here is >> > that the container is operating with an uid shifted unprivileged >> > root but sometimes needs to make use of or work with a filesystem >> > image that has root at real uid 0. >> > >> > The mechanism is to allow any subordinate mount namespace to mount >> > a shiftfs filesystem (by marking it FS_USERNS_MOUNT) but only >> > allowing it to mount marked subtrees (using the -o mark option as >> > root). Once mounted, the subtree is mapped via the super block >> > user namespace so that the interior ids of the mounting user >> > namespace are the ids written to the filesystem. >> >> Please move this into VFS instead of a stackable fs. We might need >> addtional parameters to getattr/setattr to specify the ID >> translation, but that's why better than a horrible hack like this. > > I would need a lot more than that: getattr controls the cosmetic > permission display to the user, but enforcement is done in the core > permission checks which are inode based. To make this a real bind > mount, the core permission checks will have to become subtree aware > because knowledge of whether we need a uid shift in the permission > check becomes a subtree property. Effectively inode_permission would > become dentry_permission and generic_permission would take a dentry > instead of an inode. This will be a huge amount of VFS and underlying > filesystem churn, since the permissions calls are threaded through a > huge chunk of code. > I am not even sure that would be enough. dentry does not contain information about the mount user came from, and sb contains only information about the user ns of the mounter of the file system, not the mounter of the bind mount, right? I think I am missing some big pieces of the big picture. Would love to hear what Eric has to say.
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2017-02-07 19:20 +0100 |
| Message-ID | <t8iwV-6k3-9@gated-at.bofh.it> |
| In reply to | #1575929 |
On Tue, Feb 07, 2017 at 07:59:00PM +0200, Amir Goldstein wrote: > I am not even sure that would be enough. > dentry does not contain information about the mount user came from, > and sb contains only information about the user ns of the mounter of > the file system, not the mounter of the bind mount, right? > I think I am missing some big pieces of the big picture. > Would love to hear what Eric has to say. IFF we want to do what shiftfs does properly we need vfsmount + inode, no need for the dentry. But maybe we need to go back and decice if we want to allow uid/gid remapping for arbitrary subtrees anyway. Another option would be to require something like a project as used for project quotas as the root. This would also be conveniant as it could storge the used remapping tables.
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2017-02-07 20:10 +0100 |
| Message-ID | <t8jjk-6Qz-15@gated-at.bofh.it> |
| In reply to | #1575938 |
On Tue, 2017-02-07 at 10:10 -0800, Christoph Hellwig wrote: > On Tue, Feb 07, 2017 at 07:59:00PM +0200, Amir Goldstein wrote: > > I am not even sure that would be enough. > > dentry does not contain information about the mount user came from, > > and sb contains only information about the user ns of the mounter > > of > > the file system, not the mounter of the bind mount, right? > > I think I am missing some big pieces of the big picture. > > Would love to hear what Eric has to say. > > IFF we want to do what shiftfs does properly we need vfsmount + > inode, no need for the dentry. Yes, sorry ... I was thinking the dentry contained the mnt, but it doesn't, that's the path. However, threading the mnt through looks substantially harder. > But maybe we need to go back and decice if we want to allow uid/gid > remapping for arbitrary subtrees anyway. So those were the original patches Djalal was referring to. The problem there is that a lot of orchestration systems don't store images they want to bind mount into containers on separately mounted filesystems, which is what's needed to avoid this being per-subtree. However, the clinching argument for me is that the canonical container image *is* a subtree (unlike a vm image which has to be mounted). If we don't make this work on subtrees people go back to daft stacks for containers like copying the image subtree into a loopback mounted filesystem just to make this all work (and then complain about performance and caching and so on). > Another option would be to require something like a project as used > for project quotas as the root. This would also be conveniant as it > could storge the used remapping tables. So this would be like the current project quota except set on a subtree? I could see it being done that way but I don't see what advantage it has over using flags in the subtree itself (the mapping is known based on the mount namespace, so there's really only a single bit of information to store). James
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2017-02-07 21:00 +0100 |
| Message-ID | <t8k5J-78e-43@gated-at.bofh.it> |
| In reply to | #1575967 |
On Tue, Feb 07, 2017 at 11:02:03AM -0800, James Bottomley wrote: > > Another option would be to require something like a project as used > > for project quotas as the root. This would also be conveniant as it > > could storge the used remapping tables. > > So this would be like the current project quota except set on a > subtree? I could see it being done that way but I don't see what > advantage it has over using flags in the subtree itself (the mapping is > known based on the mount namespace, so there's really only a single bit > of information to store). projects (which are the underling concept for project quotas) are per-subtree in practice - the flag is set on an inode and then all directories and files underneath inherit the project ID, hardlinking outside a project is prohinited.
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2017-02-07 21:10 +0100 |
| Message-ID | <t8kfo-7qC-3@gated-at.bofh.it> |
| In reply to | #1576006 |
On Tue, 2017-02-07 at 11:49 -0800, Christoph Hellwig wrote: > On Tue, Feb 07, 2017 at 11:02:03AM -0800, James Bottomley wrote: > > > Another option would be to require something like a project as > > > used > > > for project quotas as the root. This would also be conveniant as > > > it > > > could storge the used remapping tables. > > > > So this would be like the current project quota except set on a > > subtree? I could see it being done that way but I don't see what > > advantage it has over using flags in the subtree itself (the > > mapping is > > known based on the mount namespace, so there's really only a single > > bit > > of information to store). > > projects (which are the underling concept for project quotas) are > per-subtree in practice - the flag is set on an inode and then > all directories and files underneath inherit the project ID, > hardlinking outside a project is prohinited. OK, this is what I don't understand: how is something that's inode based limited to be per-subtree? The way I've seen the VFS operate it seems that any given inode (and indeed dentry) can appear in many subtrees so how do I limit them to just one? James
[toc] | [prev] | [next] | [standalone]
| From | Amir Goldstein <amir73il@gmail.com> |
|---|---|
| Date | 2017-02-07 22:10 +0100 |
| Message-ID | <t8lbs-81u-7@gated-at.bofh.it> |
| In reply to | #1576007 |
On Tue, Feb 7, 2017 at 10:05 PM, James Bottomley <James.Bottomley@hansenpartnership.com> wrote: > On Tue, 2017-02-07 at 11:49 -0800, Christoph Hellwig wrote: >> On Tue, Feb 07, 2017 at 11:02:03AM -0800, James Bottomley wrote: >> > > Another option would be to require something like a project as >> > > used >> > > for project quotas as the root. This would also be conveniant as >> > > it >> > > could storge the used remapping tables. >> > >> > So this would be like the current project quota except set on a >> > subtree? I could see it being done that way but I don't see what >> > advantage it has over using flags in the subtree itself (the >> > mapping is >> > known based on the mount namespace, so there's really only a single >> > bit >> > of information to store). >> >> projects (which are the underling concept for project quotas) are >> per-subtree in practice - the flag is set on an inode and then >> all directories and files underneath inherit the project ID, >> hardlinking outside a project is prohinited. > > OK, this is what I don't understand: how is something that's inode > based limited to be per-subtree? The way I've seen the VFS operate it > seems that any given inode (and indeed dentry) can appear in many > subtrees so how do I limit them to just one? > Project id's are not exactly "subtree" semantic, but inheritance semantics, which is not the same when non empty directories get their project id changed. Here is a recap: https://lwn.net/Articles/623835/ So if you created an empty directory and "marked" it for shiftuid and all descendants inherited this property you would be able to check that property on a per inode basis. Not sure that is what you are looking for? I guess we should define the semantics for the required sub-tree marking, before we can talk about solutions.
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2017-02-07 23:50 +0100 |
| Message-ID | <t8mKe-q3-3@gated-at.bofh.it> |
| In reply to | #1576061 |
On Tue, Feb 07, 2017 at 11:01:29PM +0200, Amir Goldstein wrote: > Project id's are not exactly "subtree" semantic, but inheritance semantics, > which is not the same when non empty directories get their project id changed. > Here is a recap: > https://lwn.net/Articles/623835/ Yes - but if we abuse them for containers we could refine the semantics to simply not allow change of project ids from inside containers based on say capabilities. > I guess we should define the semantics for the required sub-tree marking, > before we can talk about solutions. Good plan.
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2017-02-08 00:50 +0100 |
| Message-ID | <t8nGh-ZA-15@gated-at.bofh.it> |
| In reply to | #1576130 |
On Tue, 2017-02-07 at 14:25 -0800, Christoph Hellwig wrote: > On Tue, Feb 07, 2017 at 11:01:29PM +0200, Amir Goldstein wrote: > > Project id's are not exactly "subtree" semantic, but inheritance > > semantics, > > which is not the same when non empty directories get their project > > id changed. > > Here is a recap: > > https://lwn.net/Articles/623835/ > > Yes - but if we abuse them for containers we could refine the > semantics to simply not allow change of project ids from inside > containers based on say capabilities. We can't really abuse projectid, it's part of the user namespace mapping (for project quota). What we can do is have a new id that behaves like it. But like I said, we don't really need a ful ID, it would basically just be a single bit mark to say remap or not when doing permission checks against this inode. It would follow some of the project id semantics (like inheritance from parent dir) > > I guess we should define the semantics for the required sub-tree > > marking, before we can talk about solutions. > > Good plan. So I've been thinking about how to do this without subtree marking and yet retain the subtree properties similar to project id. The advantage would be that if it can be done using only inode properties, then none of the permission prototypes need change. The only real subtree property we need is ability to bind into an unprivileged mount namespace, but we already have that. The gotcha about marking inodes is that they're all or nothing, so every subtree that gets access to the inode inherits the mark. This means that we cannot allow a user access to a marked inode without the cover of an unprivileged user namespace, but I think that's fixable in the permission check (basically if the inode is marked you *only* get access if you have a user_ns != init_user_ns and we do the permission shifts or you have user_ns == init_user_ns and you are admin capable). James
[toc] | [prev] | [next] | [standalone]
| From | Amir Goldstein <amir73il@gmail.com> |
|---|---|
| Date | 2017-02-08 08:00 +0100 |
| Message-ID | <t8uop-5m1-1@gated-at.bofh.it> |
| In reply to | #1576178 |
On Wed, Feb 8, 2017 at 1:42 AM, James Bottomley <James.Bottomley@hansenpartnership.com> wrote: > On Tue, 2017-02-07 at 14:25 -0800, Christoph Hellwig wrote: >> On Tue, Feb 07, 2017 at 11:01:29PM +0200, Amir Goldstein wrote: >> > Project id's are not exactly "subtree" semantic, but inheritance >> > semantics, >> > which is not the same when non empty directories get their project >> > id changed. >> > Here is a recap: >> > https://lwn.net/Articles/623835/ >> >> Yes - but if we abuse them for containers we could refine the >> semantics to simply not allow change of project ids from inside >> containers based on say capabilities. > You mean something like this: https://lwn.net/Articles/632917/ With the suggested protected_projects, projid 0 (also inside container) gets a special meaning, much like user 0, so we may do interesting things with the projid that is mapped to 0. > We can't really abuse projectid, it's part of the user namespace > mapping (for project quota). What we can do is have a new id that > behaves like it. > Perhaps we *can* use projid without abusing it. userns already maps projids, but there is no concept of "owning project" for a userns, nor does it make a lot of sense, because projid is not part of the credentials. But if we re-brand it as "container root projid", we can try to use it for defining semantics to grant unprivileged access to a subtree. The functionality you are trying to get with shiftfs mark does sounds a bit like "container root projid": - inodes with mapped projid MAY be uid/gid shifted - inodes with unmapped projid MAY NOT I realize this may be very raw, but its a start. If you like this direction we can try to develop it. > But like I said, we don't really need a ful ID, it would basically just > be a single bit mark to say remap or not when doing permission checks > against this inode. It would follow some of the project id semantics > (like inheritance from parent dir) > But a single bit would only work for single level of userns nesting won't it? >> > I guess we should define the semantics for the required sub-tree >> > marking, before we can talk about solutions. >> >> Good plan. > > So I've been thinking about how to do this without subtree marking and > yet retain the subtree properties similar to project id. The advantage > would be that if it can be done using only inode properties, then none > of the permission prototypes need change. The only real subtree > property we need is ability to bind into an unprivileged mount > namespace, but we already have that. The gotcha about marking inodes > is that they're all or nothing, so every subtree that gets access to > the inode inherits the mark. This means that we cannot allow a user > access to a marked inode without the cover of an unprivileged user > namespace, but I think that's fixable in the permission check > (basically if the inode is marked you *only* get access if you have a > user_ns != init_user_ns and we do the permission shifts or you have > user_ns == init_user_ns and you are admin capable). > I didn't follow, but it sounds like your proposed solutions is only good for single level of userns nesting. Do you think you can redefine it in terms of "container root projid".
[toc] | [prev] | [next] | [standalone]
| From | Konstantin Khlebnikov <khlebnikov@yandex-team.ru> |
|---|---|
| Date | 2017-02-08 13:10 +0100 |
| Message-ID | <t8zep-a8-7@gated-at.bofh.it> |
| In reply to | #1576304 |
On 08.02.2017 09:44, Amir Goldstein wrote: > On Wed, Feb 8, 2017 at 1:42 AM, James Bottomley > <James.Bottomley@hansenpartnership.com> wrote: >> On Tue, 2017-02-07 at 14:25 -0800, Christoph Hellwig wrote: >>> On Tue, Feb 07, 2017 at 11:01:29PM +0200, Amir Goldstein wrote: >>>> Project id's are not exactly "subtree" semantic, but inheritance >>>> semantics, >>>> which is not the same when non empty directories get their project >>>> id changed. >>>> Here is a recap: >>>> https://lwn.net/Articles/623835/ >>> >>> Yes - but if we abuse them for containers we could refine the >>> semantics to simply not allow change of project ids from inside >>> containers based on say capabilities. >> > > You mean something like this: > https://lwn.net/Articles/632917/ > > With the suggested protected_projects, projid 0 (also inside container) > gets a special meaning, much like user 0, so we may do interesting > things with the projid that is mapped to 0. > >> We can't really abuse projectid, it's part of the user namespace >> mapping (for project quota). What we can do is have a new id that >> behaves like it. >> > > Perhaps we *can* use projid without abusing it. > userns already maps projids, but there is no concept of "owning project" > for a userns, nor does it make a lot of sense, because projid is not > part of the credentials. > But if we re-brand it as "container root projid", we can try to use it > for defining semantics to grant unprivileged access to a subtree. > > The functionality you are trying to get with shiftfs mark does > sounds a bit like "container root projid": > - inodes with mapped projid MAY be uid/gid shifted > - inodes with unmapped projid MAY NOT > > I realize this may be very raw, but its a start. If you like this > direction we can try to develop it. > >> But like I said, we don't really need a ful ID, it would basically just >> be a single bit mark to say remap or not when doing permission checks >> against this inode. It would follow some of the project id semantics >> (like inheritance from parent dir) >> > > But a single bit would only work for single level of userns nesting won't it? > > >>>> I guess we should define the semantics for the required sub-tree >>>> marking, before we can talk about solutions. >>> >>> Good plan. >> >> So I've been thinking about how to do this without subtree marking and >> yet retain the subtree properties similar to project id. The advantage >> would be that if it can be done using only inode properties, then none >> of the permission prototypes need change. The only real subtree >> property we need is ability to bind into an unprivileged mount >> namespace, but we already have that. The gotcha about marking inodes >> is that they're all or nothing, so every subtree that gets access to >> the inode inherits the mark. This means that we cannot allow a user >> access to a marked inode without the cover of an unprivileged user >> namespace, but I think that's fixable in the permission check >> (basically if the inode is marked you *only* get access if you have a >> user_ns != init_user_ns and we do the permission shifts or you have >> user_ns == init_user_ns and you are admin capable). >> > > I didn't follow, but it sounds like your proposed solutions is only > good for single level of userns nesting. > Do you think you can redefine it in terms of "container root projid". > Looks like all this started from mangling uid/gid or some other metadata. As usual, I have to propose funny/insane solutions: proxify filesystem with fuse and mangle everything in userspace. Or add some kind of userspace-driver remapping/mangling into overlay, for example using BPF script (I see it everywhere nowdays).
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2017-02-08 16:00 +0100 |
| Message-ID | <t8BSW-1B1-5@gated-at.bofh.it> |
| In reply to | #1576304 |
On Wed, 2017-02-08 at 08:44 +0200, Amir Goldstein wrote: > On Wed, Feb 8, 2017 at 1:42 AM, James Bottomley > <James.Bottomley@hansenpartnership.com> wrote: > > On Tue, 2017-02-07 at 14:25 -0800, Christoph Hellwig wrote: > > > On Tue, Feb 07, 2017 at 11:01:29PM +0200, Amir Goldstein wrote: > > > > Project id's are not exactly "subtree" semantic, but > > > > inheritance semantics, > > > > which is not the same when non empty directories get their > > > > project > > > > id changed. > > > > Here is a recap: > > > > https://lwn.net/Articles/623835/ > > > > > > Yes - but if we abuse them for containers we could refine the > > > semantics to simply not allow change of project ids from inside > > > containers based on say capabilities. > > > > You mean something like this: > https://lwn.net/Articles/632917/ > > With the suggested protected_projects, projid 0 (also inside > container) gets a special meaning, much like user 0, so we may do > interesting things with the projid that is mapped to 0. > > > We can't really abuse projectid, it's part of the user namespace > > mapping (for project quota). What we can do is have a new id that > > behaves like it. > > > > Perhaps we *can* use projid without abusing it. userns already maps > projids, but there is no concept of "owning project" for a userns, > nor does it make a lot of sense, because projid is not part of the > credentials. But if we re-brand it as "container root projid", we can > try to use it for defining semantics to grant unprivileged access to > a subtree. > > The functionality you are trying to get with shiftfs mark does > sounds a bit like "container root projid": > - inodes with mapped projid MAY be uid/gid shifted > - inodes with unmapped projid MAY NOT > > I realize this may be very raw, but its a start. If you like this > direction we can try to develop it. So I don't think hijacking project id is the way to go. If we do that we interfere with using project quotas within containers. Now that project quotas work for both xfs and ext4, it's no longer really an xfs specific feature. I could see adding a shift on a per projectid basis, so project id still had its quota meaning, but you could get the uid/gid shift from a given project id. However, the big kicker is that the only filesystems you can actually set a projectid on (via the fsxattr) are ext4 and xfs. That's too few to make it work universally (we'd at least need btrfs and possibly a few others). However, that's just mechanism. We can begin with a volatile mark and work out how we want to store it later. I think following projectid properties is the important one, so the choice of whether to hijack, or attach to projectid is preserved but not mandated. James
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2017-02-08 16:30 +0100 |
| Message-ID | <t8ClX-20B-1@gated-at.bofh.it> |
| In reply to | #1576304 |
On Wed, 2017-02-08 at 08:44 +0200, Amir Goldstein wrote: > On Wed, Feb 8, 2017 at 1:42 AM, James Bottomley [...] > > So I've been thinking about how to do this without subtree marking > > and yet retain the subtree properties similar to project id. The > > advantage would be that if it can be done using only inode > > properties, then none of the permission prototypes need change. > > The only real subtree property we need is ability to bind into an > > unprivileged mount namespace, but we already have that. The gotcha > > about marking inodes is that they're all or nothing, so every > > subtree that gets access to the inode inherits the mark. This > > means that we cannot allow a user access to a marked inode without > > the cover of an unprivileged user namespace, but I think that's > > fixable in the permission check (basically if the inode is marked > > you *only* get access if you have a user_ns != init_user_ns and we > > do the permission shifts or you have user_ns == init_user_ns and > > you are admin capable). > > > > I didn't follow, but it sounds like your proposed solutions is only > good for single level of userns nesting. Do you think you can > redefine it in terms of "container root projid". I don't quite understand what you're getting at. user_ns mappings nest, but what we see depends on where you're trying to look at it. Let's take the kernel's view as the primary one. That's the kuid_t. The user has a different view, the uid_t and now we have the filesystem view (no actual type for this). The user view is produced by from the kernel view by chaining up all the maps from the current_user_ns and the filesystem view is produced by doing the same thing for the s_user_ns. So however many levels of user namespace nesting we have operating, we only have three views of what an id is: the user view, the kernel view and the filesystem view. All nesting does is change how those views are mapped but it doesn't alter the number of views. What the original shiftfs patches (not the ones that use s_user_ns) did was to introduce effectively an inode view and map between the kernel and the inode view using the shift mapping parameters; then the inode view would get mapped through the s_user_ns to become the filesystem view. In the s_user_ns version of shiftfs (the current patches), there's still an inode view, but we know that what we want to write to disk is the user view, so effectively the user view and the inode view become the same if the filesystem is marked otherwise the inode view and the kernel view are the same if it isn't. That's why I only need a single bit to tell me if I'm mapping or not and there are two separate regimes to check the permissions in: the user == inode view and the kernel == inode view. James
[toc] | [prev] | [next] | [standalone]
| From | Josh Triplett <josh@joshtriplett.org> |
|---|---|
| Date | 2017-02-08 03:00 +0100 |
| Message-ID | <t8pI5-2cw-17@gated-at.bofh.it> |
| In reply to | #1576006 |
On Tue, Feb 07, 2017 at 11:49:33AM -0800, Christoph Hellwig wrote: > On Tue, Feb 07, 2017 at 11:02:03AM -0800, James Bottomley wrote: > > > Another option would be to require something like a project as used > > > for project quotas as the root. This would also be conveniant as it > > > could storge the used remapping tables. > > > > So this would be like the current project quota except set on a > > subtree? I could see it being done that way but I don't see what > > advantage it has over using flags in the subtree itself (the mapping is > > known based on the mount namespace, so there's really only a single bit > > of information to store). > > projects (which are the underling concept for project quotas) are > per-subtree in practice - the flag is set on an inode and then > all directories and files underneath inherit the project ID, > hardlinking outside a project is prohinited. I'm interested in having a VFS-level way to do more than just a shift; I'd like to be able to arbitrarily remap IDs between what's on disk and the system IDs. If we're talking about developing a VFS-level solution for this, I'd like to avoid limiting it to just a shift. (A shift/range would definitely be the simplest solution for many common container cases, but not all.)
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2017-02-08 16:30 +0100 |
| Message-ID | <t8ClY-20B-17@gated-at.bofh.it> |
| In reply to | #1576225 |
On Tue, 2017-02-07 at 17:54 -0800, Josh Triplett wrote: > On Tue, Feb 07, 2017 at 11:49:33AM -0800, Christoph Hellwig wrote: > > On Tue, Feb 07, 2017 at 11:02:03AM -0800, James Bottomley wrote: > > > > Another option would be to require something like a project > > > > as used > > > > for project quotas as the root. This would also be conveniant > > > > as it > > > > could storge the used remapping tables. > > > > > > So this would be like the current project quota except set on a > > > subtree? I could see it being done that way but I don't see what > > > advantage it has over using flags in the subtree itself (the > > > mapping is > > > known based on the mount namespace, so there's really only a > > > single bit > > > of information to store). > > > > projects (which are the underling concept for project quotas) are > > per-subtree in practice - the flag is set on an inode and then > > all directories and files underneath inherit the project ID, > > hardlinking outside a project is prohinited. > > I'm interested in having a VFS-level way to do more than just a > shift; I'd like to be able to arbitrarily remap IDs between what's on > disk and the system IDs. OK, so the shift is effectively an arbitrary remap because it allows multiple ranges to be mapped (althought the userns currently imposes a maximum number of five extents but that limit is a bit arbitrary just to try to limit the amount of space the parametrisation takes). See kernel/user_namespace.c:map_id_up/down() > If we're talking about developing a VFS-level solution for this, > I'd like to avoid limiting it to just a shift. (A shift/range > would definitely be the simplest solution for many common container > cases, but not all.) I assume the above satisfies you on this point, but raises the question: do you want an arbitrary shift not parametrised by a user namespace? If so how many such shifts do you want ... giving some details of the use case would be helpful. James
[toc] | [prev] | [next] | [standalone]
| From | Josh Triplett <josh@joshtriplett.org> |
|---|---|
| Date | 2017-02-09 11:40 +0100 |
| Message-ID | <t8UiS-4ZJ-21@gated-at.bofh.it> |
| In reply to | #1576653 |
On Wed, Feb 08, 2017 at 07:22:45AM -0800, James Bottomley wrote: > On Tue, 2017-02-07 at 17:54 -0800, Josh Triplett wrote: > > On Tue, Feb 07, 2017 at 11:49:33AM -0800, Christoph Hellwig wrote: > > > On Tue, Feb 07, 2017 at 11:02:03AM -0800, James Bottomley wrote: > > > > > Another option would be to require something like a project > > > > > as used > > > > > for project quotas as the root. This would also be conveniant > > > > > as it > > > > > could storge the used remapping tables. > > > > > > > > So this would be like the current project quota except set on a > > > > subtree? I could see it being done that way but I don't see what > > > > advantage it has over using flags in the subtree itself (the > > > > mapping is > > > > known based on the mount namespace, so there's really only a > > > > single bit > > > > of information to store). > > > > > > projects (which are the underling concept for project quotas) are > > > per-subtree in practice - the flag is set on an inode and then > > > all directories and files underneath inherit the project ID, > > > hardlinking outside a project is prohinited. > > > > I'm interested in having a VFS-level way to do more than just a > > shift; I'd like to be able to arbitrarily remap IDs between what's on > > disk and the system IDs. > > OK, so the shift is effectively an arbitrary remap because it allows > multiple ranges to be mapped (althought the userns currently imposes a > maximum number of five extents but that limit is a bit arbitrary just > to try to limit the amount of space the parametrisation takes). See > kernel/user_namespace.c:map_id_up/down() > > > If we're talking about developing a VFS-level solution for this, > > I'd like to avoid limiting it to just a shift. (A shift/range > > would definitely be the simplest solution for many common container > > cases, but not all.) > > I assume the above satisfies you on this point, but raises the > question: do you want an arbitrary shift not parametrised by a user > namespace? If so how many such shifts do you want ... giving some > details of the use case would be helpful. The limit of five extents means this may not work in the most general case, no. One use case: given an on-disk filesystem, its name-to-number mapping, and your host name-to-number mapping, mount the filesystem with all the UIDs bidirectionally mapped to those on your host system. Another use case: given an on-disk filesystem with potentially arbitrary UIDs (not necessarily in a clean contiguous block), and a pile of unprivileged UIDs, mount the filesystem such that every on-disk UID gets a unique unprivileged UID. (I have some additional use cases, but they would require the ability to extend the mapping on the fly without remounting.)
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2017-02-09 17:40 +0100 |
| Message-ID | <t8ZVg-8tc-17@gated-at.bofh.it> |
| In reply to | #1577463 |
On Thu, 2017-02-09 at 02:36 -0800, Josh Triplett wrote: > On Wed, Feb 08, 2017 at 07:22:45AM -0800, James Bottomley wrote: > > On Tue, 2017-02-07 at 17:54 -0800, Josh Triplett wrote: > > > On Tue, Feb 07, 2017 at 11:49:33AM -0800, Christoph Hellwig > > > wrote: > > > > On Tue, Feb 07, 2017 at 11:02:03AM -0800, James Bottomley > > > > wrote: > > > > > > Another option would be to require something like a > > > > > > project as used for project quotas as the root. This would > > > > > > also be conveniant as it could storge the used remapping > > > > > > tables. > > > > > > > > > > So this would be like the current project quota except set on > > > > > a subtree? I could see it being done that way but I don't > > > > > see what advantage it has over using flags in the subtree > > > > > itself (the mapping is known based on the mount namespace, so > > > > > there's really only a single bit of information to store). > > > > > > > > projects (which are the underling concept for project quotas) > > > > are per-subtree in practice - the flag is set on an inode and > > > > then all directories and files underneath inherit the project > > > > ID, hardlinking outside a project is prohinited. > > > > > > I'm interested in having a VFS-level way to do more than just a > > > shift; I'd like to be able to arbitrarily remap IDs between > > > what's on disk and the system IDs. > > > > OK, so the shift is effectively an arbitrary remap because it > > allows multiple ranges to be mapped (althought the userns currently > > imposes a maximum number of five extents but that limit is a bit > > arbitrary just to try to limit the amount of space the > > parametrisation takes). See > > kernel/user_namespace.c:map_id_up/down() > > > > > If we're talking about developing a VFS-level solution for > > > this, I'd like to avoid limiting it to just a shift. (A > > > shift/range would definitely be the simplest solution for many > > > common container cases, but not all.) > > > > I assume the above satisfies you on this point, but raises the > > question: do you want an arbitrary shift not parametrised by a user > > namespace? If so how many such shifts do you want ... giving some > > details of the use case would be helpful. > > The limit of five extents means this may not work in the most general > case, no. That's not an API limit, so it can be changed if there's a need. The problem was merely how to parametrise a mapping without taking too much space. > One use case: given an on-disk filesystem, its name-to-number > mapping, and your host name-to-number mapping, mount the filesystem > with all the UIDs bidirectionally mapped to those on your host > system. This is pretty much what the s_user_ns does. > Another use case: given an on-disk filesystem with potentially > arbitrary UIDs (not necessarily in a clean contiguous block), and a > pile of unprivileged UIDs, mount the filesystem such that every on > -disk UID gets a unique unprivileged UID. So is this. Basically anything that begins by mounting gets a super block and can use the s_user_ns to map from the filesystem view to the kernel view of ids. Apart from greater sophistication in the parametrisation, it sounds like we have all the machinery you need. I'm sure the containers people will consider reasonable patches to change this. James > (I have some additional use cases, but they would require the ability > to extend the mapping on the fly without remounting.) >
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2017-02-13 11:30 +0100 |
| Message-ID | <tam3o-2qv-17@gated-at.bofh.it> |
| In reply to | #1577751 |
James Bottomley <James.Bottomley@HansenPartnership.com> writes: > On Thu, 2017-02-09 at 02:36 -0800, Josh Triplett wrote: >> On Wed, Feb 08, 2017 at 07:22:45AM -0800, James Bottomley wrote: >> > On Tue, 2017-02-07 at 17:54 -0800, Josh Triplett wrote: >> > > On Tue, Feb 07, 2017 at 11:49:33AM -0800, Christoph Hellwig >> > > wrote: >> > > > On Tue, Feb 07, 2017 at 11:02:03AM -0800, James Bottomley >> > > > wrote: >> > > > > > Another option would be to require something like a >> > > > > > project as used for project quotas as the root. This would >> > > > > > also be conveniant as it could storge the used remapping >> > > > > > tables. >> > > > > >> > > > > So this would be like the current project quota except set on >> > > > > a subtree? I could see it being done that way but I don't >> > > > > see what advantage it has over using flags in the subtree >> > > > > itself (the mapping is known based on the mount namespace, so >> > > > > there's really only a single bit of information to store). >> > > > >> > > > projects (which are the underling concept for project quotas) >> > > > are per-subtree in practice - the flag is set on an inode and >> > > > then all directories and files underneath inherit the project >> > > > ID, hardlinking outside a project is prohinited. >> > > >> > > I'm interested in having a VFS-level way to do more than just a >> > > shift; I'd like to be able to arbitrarily remap IDs between >> > > what's on disk and the system IDs. >> > >> > OK, so the shift is effectively an arbitrary remap because it >> > allows multiple ranges to be mapped (althought the userns currently >> > imposes a maximum number of five extents but that limit is a bit >> > arbitrary just to try to limit the amount of space the >> > parametrisation takes). See >> > kernel/user_namespace.c:map_id_up/down() >> > >> > > If we're talking about developing a VFS-level solution for >> > > this, I'd like to avoid limiting it to just a shift. (A >> > > shift/range would definitely be the simplest solution for many >> > > common container cases, but not all.) >> > >> > I assume the above satisfies you on this point, but raises the >> > question: do you want an arbitrary shift not parametrised by a user >> > namespace? If so how many such shifts do you want ... giving some >> > details of the use case would be helpful. >> >> The limit of five extents means this may not work in the most general >> case, no. > > That's not an API limit, so it can be changed if there's a need. The > problem was merely how to parametrise a mapping without taking too much > space. > >> One use case: given an on-disk filesystem, its name-to-number >> mapping, and your host name-to-number mapping, mount the filesystem >> with all the UIDs bidirectionally mapped to those on your host >> system. > > This is pretty much what the s_user_ns does. > >> Another use case: given an on-disk filesystem with potentially >> arbitrary UIDs (not necessarily in a clean contiguous block), and a >> pile of unprivileged UIDs, mount the filesystem such that every on >> -disk UID gets a unique unprivileged UID. > > So is this. Basically anything that begins by mounting gets a super > block and can use the s_user_ns to map from the filesystem view to the > kernel view of ids. Apart from greater sophistication in the > parametrisation, it sounds like we have all the machinery you need. > I'm sure the containers people will consider reasonable patches to > change this. Yes. And to be clear we have all of that merged now and mostly present and hooked up in all filesystems without any shiftfs like changes needed. To use this with a filesystem a last pass needs to be had to verify that the cases where something does not map are handled cleanly. Eric
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web