Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1394399 > unrolled thread
| Started by | Djalal Harouni <tixxdz@gmail.com> |
|---|---|
| First post | 2016-05-04 16:50 +0200 |
| Last post | 2016-05-12 21:50 +0200 |
| Articles | 20 on this page of 39 — 9 participants |
Back to article view | Back to linux.kernel
[RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-04 16:50 +0200
[RFC v2 PATCH 1/8] VFS: add CLONE_MNTNS_SHIFT_UIDGID flag to allow mounts to shift their UIDs/GIDs Djalal Harouni <tixxdz@gmail.com> - 2016-05-04 16:50 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Josh Triplett <josh@joshtriplett.org> - 2016-05-04 18:40 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-05-04 23:10 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-05 09:40 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-05-05 14:00 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-06 00:00 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-05-06 00:10 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-05-11 01:40 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Al Viro <viro@ZenIV.linux.org.uk> - 2016-05-11 02:40 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Al Viro <viro@ZenIV.linux.org.uk> - 2016-05-11 03:00 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-05-11 05:50 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-11 18:50 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-05-11 20:40 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-12 22:00 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-05-13 00:30 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-14 12:00 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-05-14 15:50 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-05-16 04:10 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems ebiederm@xmission.com (Eric W. Biederman) - 2016-05-16 04:10 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Seth Forshee <seth.forshee@canonical.com> - 2016-05-16 16:20 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems ebiederm@xmission.com (Eric W. Biederman) - 2016-05-16 19:00 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Seth Forshee <seth.forshee@canonical.com> - 2016-05-16 20:30 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems James Bottomley <James.Bottomley@HansenPartnership.com> - 2016-05-16 21:20 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems ebiederm@xmission.com (Eric W. Biederman) - 2016-05-18 01:00 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-17 13:50 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-17 17:50 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Serge Hallyn <serge.hallyn@ubuntu.com> - 2016-05-05 01:40 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-06 16:40 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Serge Hallyn <serge.hallyn@ubuntu.com> - 2016-05-09 18:30 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-10 12:40 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Dave Chinner <david@fromorbit.com> - 2016-05-05 02:30 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Andy Lutomirski <luto@amacapital.net> - 2016-05-05 03:50 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Dave Chinner <david@fromorbit.com> - 2016-05-05 04:30 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Andy Lutomirski <luto@amacapital.net> - 2016-05-05 05:40 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-06 00:40 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-06 00:30 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Dave Chinner <david@fromorbit.com> - 2016-05-06 05:00 +0200
Re: [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems Djalal Harouni <tixxdz@gmail.com> - 2016-05-12 21:50 +0200
Page 1 of 2 [1] 2 Next page →
| From | Djalal Harouni <tixxdz@gmail.com> |
|---|---|
| Date | 2016-05-04 16:50 +0200 |
| Subject | [RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems |
| Message-ID | <rv5Ym-5ES-11@gated-at.bofh.it> |
This is version 2 of the VFS:userns support portable root filesystems
RFC. Changes since version 1:
* Update documentation and remove some ambiguity about the feature.
Based on Josh Triplett comments.
* Use a new email address to send the RFC :-)
This RFC tries to explore how to support filesystem operations inside
user namespace using only VFS and a per mount namespace solution. This
allows to take advantage of user namespace separations without
introducing any change at the filesystems level. All this is handled
with the virtual view of mount namespaces.
1) Presentation:
================
The main aim is to support portable root filesystems and allow containers,
virtual machines and other cases to use the same root filesystem.
Due to security reasons, filesystems can't be mounted inside user
namespaces, and mounting them outside will not solve the problem since
they will show up with the wrong UIDs/GIDs. Read and write operations
will also fail and so on.
The current userspace solution is to automatically chown the whole root
filesystem before starting a container, example:
(host) init_user_ns 1000000:1065536 => (container) user_ns_X1 0:65535
(host) init_user_ns 2000000:2065536 => (container) user_ns_Y1 0:65535
(host) init_user_ns 3000000:3065536 => (container) user_ns_Z1 0:65535
...
Every time a chown is called, files are changed and so on... This
prevents to have portable filesystems where you can throw anywhere
and boot. Having an extra step to adapt the filesystem to the current
mapping and persist it will not allow to verify its integrity, it makes
snapshots and migration a bit harder, and probably other limitations...
It seems that there are multiple ways to allow user namespaces combine
nicely with filesystems, but none of them is that easy. The bind mount
and pin the user namespace during mount time will not work, bind mounts
share the same super block, hence you may endup working on the wrong
vfsmount context and there is no easy way to get out of that...
Using the user namespace in the super block seems the way to go, and
there is the "Support fuse mounts in user namespaces" [1] patches which
seem nice but perhaps too complex!? there is also the overlayfs solution,
and finaly the VFS layer solution.
We present here a simple VFS solution, everything is packed inside VFS,
filesystems don't need to know anything (except probably XFS, and special
operations inside union filesystems). Currently it supports ext4, btrfs
and overlayfs. Changes into filesystems are small, just parse the
vfs_shift_uids and vfs_shift_gids options during mount and set the
appropriate flags into the super_block structure.
1) Filesystems don't need the FS_USERNS_MOUNT flag, so no user
namespace mounting, they stay secure, nothing changes.
2) The solution is based on VFS and mount namespaces, we use the user
namespace of the containing mount namespace to check if we should shift
UIDs/GIDs from/to virtual <=> on-disk view.
If a filesystem was mounted with "vfs_shift_uids" and "vfs_shift_gids"
options, and if it shows up inside a mount namespace that supports VFS
UIDs/GIDs shifts then during each access we will remap UID/GID either
to virtual or to on-disk view using simple helper functions to allow the
access. In case the mount or current mount namespace do not support VFS
UID/GID shifts, we fallback to the old behaviour, no shift is performed.
3) The existing user namespace interface is the one used to do the
translation from virtual to on-disk mapping.
3) inodes will always keep their original values which reflect the
mapping inside init_user_ns which we consider the on-disk mapping.
3.1) During access we map to the virtual view, and if the
inode->{i_uid|i_gid} do not have a mapping in the mount namespace
we construct one for them.
3.2) For on-disk write we construct the appropriate kuid/kgid that
should be stored on-disk. If they have a mapping in the mount
namespace we use the corresponding uid_t/gid_t values of that
mapping inside the mount namespace and construct the kuid from
the pair init_user_ns and uid_t. This covers cases where the
mapping inside should be the one stored into on-disk. Now If they
don't have a mapping in the mount namespace, we fallback to the
old behaviour, the global kuid inside init_user_ns is the one
used to update the inode->i_uid.
As an example if the mapping 0:65535 inside mount namespace and outside
is 1000000:1065536, then 0:65535 will be the range that we use to
construct UIDs/GIDs mapping into init_user_ns and use it for on-disk
data. They represent the persistent values that we want to write to the
disk. Therefore, we don't keep track of any UID/GID shift that was applied
before, it gives portability and allows to use the previous mapping
which was freed for another root filesystem...
If the mapping inside the mount namespace is 1000:65535 and outside
is 2000:65535 then the range used to construct UIDs/GIDs mapping to
update inode->{i_uid|i_gid} will be the one inside the container, we
always use that one to construct the kuid/kgid from uid_t/gid_t and
init_user_ns.
$ cat /proc/self/uid_map
1000 2000 65536
$ stat -c '%u:%g' mountpoint/etc/fedora-release
65534:65534
$ stat -c '%u:%g' mountpoint/home/tixxdz/
1000:1000
$ touch mountpoint/newuser
touch: cannot touch ‘mountpoint/newuser’: Permission denied
$ stat -c '%u:%g' mountpoint/home/tixxdz/newuser
1000:1000
[ outside of namespaces] $ stat -c '%u:%g' mountpoint/home/tixxdz/newuser
1000:1000
Please note that the range here is not hardcoded to 65535, it can be any
value set by the creator of the user namespace. These patches use the
only interface user namespaces provide. 2**16 was used here to just show
how filesystems can be made portable by making the most used UIDs/GIDs
available inside containers.
Simple demo overlayfs, and btrfs mounted with vfs_shift_uids and
vfs_shift_gids. The overlayfs mounts will share the same upperdir. We
create two user namesapces every one with its own mapping and where
container-uid-2000000 will pull changes from container-uid-1000000
upperdir automatically.
[tixxdz@fedora-kvm btrfs_root]$ mount | grep btrfs
/dev/mapper/fedora-btrfs_root on /mnt/btrfs_root type btrfs (rw,relatime,seclabel,space_cache,vfs_shift_uids,vfs_shift_gids,subvolid=5,subvol=/)
[tixxdz@fedora-kvm btrfs_root]$ sudo mount -t overlay overlay \
-o,lowerdir=/mnt/btrfs_root/rootfs/fedora-tree,upperdir=/mnt/btrfs_root/container-uid-1000000/upperdir,workdir=/mnt/btrfs_root/container-uid-1000000/workdir \
/mnt/btrfs_root/container-uid-1000000/merged
[tixxdz@fedora-kvm btrfs_root]$ sudo mount -t overlay overlay \
-o,lowerdir=/mnt/btrfs_root/rootfs/fedora-tree,upperdir=/mnt/btrfs_root/container-uid-1000000/upperdir,workdir=/mnt/btrfs_root/container-uid-2000000/workdir \
/mnt/btrfs_root/container-uid-2000000/merged
[tixxdz@fedora-kvm btrfs_root]$ sudo chown -R 1000000.1000000 /mnt/btrfs_root/container-uid-1000000/workdir/work/
[tixxdz@fedora-kvm btrfs_root]$ sudo chown -R 2000000.2000000 /mnt/btrfs_root/container-uid-2000000/workdir/work/
[ Term 1 ]
[tixxdz@fedora-kvm container-uid-1000000]$ sudo ~/bin/mountns-uidshift -u 1000000
bash: /root/.bashrc: Permission denied
bash-4.3# cat /proc/self/uid_map
0 1000000 65536
bash-4.3# touch container-uid-1000000/merged/rootfile
bash-4.3# stat -c '%u:%g' container-uid-1000000/merged/rootfile
0:0
[ Term 2 ]
[tixxdz@fedora-kvm btrfs_root]$ sudo ~/bin/mountns-uidshift -u 2000000
[sudo] password for tixxdz:
bash: /root/.bashrc: Permission denied
bash-4.3# cat /proc/self/uid_map
0 2000000 65536
bash-4.3# stat -c '%u:%g' container-uid-2000000/merged/rootfile
0:0
[ Term 3 ] (outside of all namespaces)
[tixxdz@fedora-kvm btrfs_root]$ stat -c '%u:%g' container-uid-1000000/upperdir/rootfile
0:0
This means that root in user namespace or inside containers is able to
write inodes with uid/gid == 0 into disk. This may sound strange and
dangerous, yes of course, care must be taken, this way we have added
the following:
1) Filesystems when mounted must explicitly support "vfs_shift_uids"
and "vfs_shift_gids", we don't require mounting inside user namespaces.
2) Containers or mounts can have their parent directory as 0700, and
even before mounting clean the mount namespace, set the appropriate
propagation flags and so on...
3) To be able to set the CLONE_MNTNS_SHIFT_UIDGID flag on the new mount
namespace either caller has to be real root in init_user_ns, or the parent
of the new mount namespace has already that flag set. This allows
nesting which I discussed briefly with Serge Hallyn, and he suggested
that this should be supported. Preventing nesting is doomed to fail. This
way we have security and nesting at the same time. Of course if you clean
that flag you won't be able to set it next time only if you are capable
in init_user_ns.
4) If the mount namespaces has the flag CLONE_MNTNS_SHIFT_UIDGID set but
the filesystem was mounted without "vfs_shift_uids" and "vfs_shift_gids"
or does not support these options, then no shifting is performed. You
have to meet the two conditions at each access, otherwise we fallback to
current behaviour.
5) Only the creator of the mount namespace or one with similar
privileges is able to change the mapping rules of the user namespace of
that mount namespace. This ensures that only a more privileged is able
to change the mapping and at the same time it gives some flexibility
since the rules can be changed, and we never persist the virtual
UIDs/GIDs into disk, only the view in init_user_ns is always stored into
disk.
To complete this solution the current blocker is: since we need a way to
control mount namespaces we need a new flag, however all flags of current
clone() syscall are consumed, yes 32bits no luck! In this RFC I didn't
include a new syscall clone4() [2] which was already requested in the past,
and the patches for a new clone4() are already there. This way this RFC
stays minimal.
The flag we use here is just for demonstration, please see patch 0001
and the program mountns-uidshift.c [3] for that. Future versions
will include the new clone4() syscall.
2) TEST:
========
Apply on top of Linux 4.6-rc6 HEAD 04974df8049fc4240d2275, and use this
program mountns-uidshift.c to test the shifted mount namespaces.
https://raw.githubusercontent.com/OpenDZ/research/master/kernel/mountns-uidshift.c
With current mapping rules init_user_ns:
[1000000:1065536] => new_user_ns: [0:65536]
# cat /proc/self/uid_map
0 1000000 65536
# cat /proc/self/gid_map
0 1000000 65536
2.1) ext4:
==========
Setup:
/ on ext4 without vfs_shift_uids, vfs_shift_gids
/mnt/ext4_root on ext4 with vfs_shift_uids, vfs_shift_gids
/mnt/ext4_root/rootfs/fedore-tree (Another fedora rootfs)
/mnt/ext4_root/container-uid-1000000 (container files with uid 1000000)
/mnt/ext4_root/container-uid-1000000/mountpoint (bind mount of fedora-tree)
$ sudo mount -t ext4 -ovfs_shift_uids,vfs_shift_gids \
/dev/fedora/ext4_root /mnt/ext4_root/
$ mount | grep ext4 -
/dev/mapper/fedora-root on / type ext4 (rw,relatime,seclabel,data=ordered)
/dev/sda1 on /boot type ext4 (rw,relatime,seclabel,data=ordered)
/dev/mapper/fedora-ext4_root on /mnt/ext4_root type ext4 (rw,relatime,seclabel,data=ordered,vfs_shift_uids,vfs_shift_gids)
$ sudo mkdir /mnt/ext4_root/rootfs/
$ sudo yum -y --releasever=23 --installroot=/mnt/ext4_root/rootfs/fedora-tree \
--disablerepo='*' --enablerepo=fedora install systemd passwd yum fedora-release vim
$ sudo mkdir /mnt/ext4_root/container-uid-1000000/
$ sudo mkdir /mnt/ext4_root/container-uid-1000000/mountpoint
$ sudo chown -R 1000000.1000000 /mnt/ext4_root/container-uid-1000000/
$ sudo mount --bind -ovfs_shift_uids,vfs_shift_gids \
/mnt/ext4_root/rootfs/fedora-tree/ mountpoint/
$ mount | grep vfs_shift_uids -
/dev/mapper/fedora-ext4_root on /mnt/ext4_root type ext4 (rw,relatime,seclabel,data=ordered,vfs_shift_uids,vfs_shift_gids)
/dev/mapper/fedora-ext4_root on /mnt/ext4_root/container-uid-1000000/mountpoint type ext4 (rw,relatime,seclabel,data=ordered,vfs_shift_uids,vfs_shift_gids)
$ sudo ~/bin/mountns-uidshift -u 1000000
...
bash-4.3# id
uid=0(root) gid=0(root) groups=0(root) context=unconfined_u:unconfined_r:unconfined_t:s0-s0:c0.c1023
bash-4.3# cat /proc/self/uid_map
0 1000000 65536
bash-4.3# mount | grep shift -
/dev/mapper/fedora-ext4_root on /mnt/ext4_root type ext4 (rw,relatime,seclabel,data=ordered,vfs_shift_uids,vfs_shift_gids)
/dev/mapper/fedora-ext4_root on /mnt/ext4_root/container-uid-1000000/mountpoint type ext4 (rw,relatime,seclabel,data=ordered,vfs_shift_uids,vfs_shift_gids)
bash-4.3# stat -c '%u:%g' /etc/motd
65534:65534
bash-4.3# stat -c '%u:%g' /mnt/ext4_root/rootfs/fedora-tree/etc/motd
0:0
bash-4.3# stat -c '%u:%g' mountpoint/etc/motd
0:0
bash-4.3# stat -c '%u:%g' /etc/machine-id
65534:65534
bash-4.3# echo "blabla" > /etc/machine-id
bash: /etc/machine-id: Permission denied
bash-4.3# stat -c '%u:%g' mountpoint/etc/machine-id
0:0
bash-4.3# sha1sum mountpoint/etc/machine-id
edb24591988f0f003cd397704f49e92208b3015f mountpoint/etc/machine-id
bash-4.3# m=$(cat /dev/urandom | tr -cd 'a-f0-9' | head -c 32); echo $m | sha1sum; echo $m > mountpoint/etc/machine-id
f256a796b1f2ed09c4107f1f5aff2568fb2d79cc -
bash-4.3# sha1sum mountpoint/etc/machine-id
f256a796b1f2ed09c4107f1f5aff2568fb2d79cc mountpoint/etc/machine-id
bash-4.3# stat -c '%u:%g' mountpoint/etc/machine-id
0:0
[outside of namespaces]$ stat -c '%u:%g' /mnt/ext4_root/container-uid-1000000/mountpoint/etc/machine-id
0:0
Test with unprivileged user inside the new mount and user namespaces:
---------------------------------------------------------------------
Test with uid tixxdz == 1000, the user exists on both:
(1) /
(2) /mnt/ext4_root/rootfs/fedore-tree which is bind mounted into /mnt/ext4_root/container-uid-1000000/mountpoint
-bash-4.3$ touch /home/tixxdz/newfile
touch: cannot touch /home/tixxdz/newfile: Permission denied
-bash-4.3$ stat -c '%u:%g' /home/tixxdz/
65534:65534
-bash-4.3$ stat -c '%u:%g' mountpoint/home/tixxdz/
1000:1000
-bash-4.3$ touch mountpoint/home/tixxdz/newfile
-bash-4.3$ stat -c '%u:%g' mountpoint/home/tixxdz/newfile
1000:1000
[outside of namespaces]$ stat -c '%u:%g' /mnt/ext4_root/container-uid-1000000/mountpoint/home/tixxdz/newfile
1000:1000
2.2) btrfs:
===========
Same steps as ext4.
2.3) overlayfs:
===============
2.3.1) Native support using VFS:
Overlayfs is natively supported if lowerdir, upperdir and workdir are all
on a mount that supports vfs_shift_uids and vfs_shift_gids flags and we
are in a mount namespace that also supports that.
$ mount | grep btrfs
/dev/mapper/fedora-btrfs_root on /mnt/btrfs_root type btrfs (rw,relatime,seclabel,space_cache,vfs_shift_uids,vfs_shift_gids,subvolid=5,subvol=/)
$ cd /mnt/btrfs_root/
$ sudo mkdir -p container-uid-2000000/{upperdir,workdir,merged}
$ sudo chown -R 2000000.2000000 container-uid-2000000/
$ cd container-uid-2000000/
$ sudo mount -t overlay overlay -o,lowerdir=/mnt/btrfs_root/rootfs/fedora-tree,upperdir=upperdir,workdir=workdir merged
$ sudo chown -R 2000000.2000000 workdir/work/
$ sudo ~/bin/mountns-uidshift -u 2000000
...
bash-4.3# stat -c '%u:%g' merged/etc/passwd
0:0
bash-4.3# touch merged/overlayfs-file
bash-4.3# stat -c '%u:%g' merged/overlayfs-file
0:0
[outside of namespaces]# stat -c '%u:%g' /mnt/btrfs_root/container-uid-2000000/merged/overlayfs-file
0:0
[outside of namespaces]# stat -c '%u:%g' /mnt/btrfs_root/container-uid-2000000/upperdir/overlayfs-file
0:0
2.3.2) Complex support or union filesystems:
If overlayfs lowerdir and upperdir are not on a filesystem that supports
natively vfs_shift_uids and vfs_shift_gids then to support VFS UID/GID
shifts, we must adapt the helper functions that where introduced in this
series to take also a super_block struct and test if the appropriate flags
where set into overlayfs instead of the other filesystem which the inode
belongs to. The translation on-disk <=> virtual should happen then inside
overlayfs.
I think this will always be the case of union mounts which fetch an inode
from another mount. I think that solution (2.3.2) can also be implemented,
I had some ugly patches to implement this on top of overlayfs, but not
sure, better see what others think about VFS UID/GID shifts first.
IMO solution (2.3.1) if done correctly is the way to go, in the end all
this relates to the virtual view of UID/GID inside the kernel, and how
resources are translated to them, it's not related to overlayfs.
3) ROADMAP:
===========
* Confirm current design, and make sure that the mapping is done
correctly.
* Add clone4() syscall [2]
* Investigate if current setns() checks to enter new mount namespaces
are sufficient ?
* Add POSIX ACL support ?
* Check if all filesystem operations are correctly supported and recheck
permissions access.
* Do filesystems provide some operations to control disk or host resources ?
in other words are there some inodes on filesystems that allow to access
host resources, if so then maybe these inodes either should be marked only
safe in init_user_ns or get the appropriate capable() in init_user_ns if
missing. Needs investigation.
* Add XFS support.
References:
===========
[1] https://www.redhat.com/archives/dm-devel/2016-April/msg00368.html
[2] https://lkml.org/lkml/2015/3/15/10
[3] https://raw.githubusercontent.com/OpenDZ/research/master/kernel/mountns-uidshift.c
Thanks!
Patches:
[RFC v2 PATCH 0/8] VFS:userns: support portable root filesystems
[RFC v2 PATCH 1/8] VFS: add CLONE_MNTNS_SHIFT_UIDGID flag to allow mounts to shift their UIDs/GIDs
[RFC v2 PATCH 2/8] VFS:uidshift: add flags and helpers to shift UIDs and GIDs to virtual view
[RFC v2 PATCH 3/8] fs: Treat foreign mounts as nosuid
[RFC v2 PATCH 4/8] VFS:userns: shift UID/GID to virtual view during permission access
[RFC v2 PATCH 5/8] VFS:userns: add helpers to shift UIDs and GIDs into on-disk view
[RFC v2 PATCH 6/8] VFS:userns: shift UID/GID to on-disk view before any write to disk
[RFC v2 PATCH 7/8] ext4: add support for vfs_shift_uids and vfs_shift_gids mount options
[RFC v2 PATCH 8/8] btrfs: add support for vfs_shift_uids and vfs_shift_gids mount options
Diffstat for this RFC
fs/attr.c | 44 +++++++++++++++++++++++--------
fs/btrfs/super.c | 15 ++++++++++-
fs/exec.c | 2 +-
fs/ext4/super.c | 14 ++++++++++
fs/inode.c | 9 ++++---
fs/mount.h | 1 +
fs/namei.c | 6 +++--
fs/namespace.c | 190 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
fs/stat.c | 4 +--
include/linux/fs.h | 14 ++++++++++
include/linux/mount.h | 1 +
include/linux/user_namespace.h | 8 ++++++
include/uapi/linux/sched.h | 1 +
kernel/capability.c | 14 ++++++++--
kernel/fork.c | 4 +++
kernel/user_namespace.c | 13 ++++++++++
security/commoncap.c | 2 +-
security/selinux/hooks.c | 2 +-
18 files changed, 319 insertions(+), 25 deletions(-)
[toc] | [next] | [standalone]
| From | Djalal Harouni <tixxdz@gmail.com> |
|---|---|
| Date | 2016-05-04 16:50 +0200 |
| Subject | [RFC v2 PATCH 1/8] VFS: add CLONE_MNTNS_SHIFT_UIDGID flag to allow mounts to shift their UIDs/GIDs |
| Message-ID | <rv6hK-5S0-63@gated-at.bofh.it> |
| In reply to | #1394399 |
Add CLONE_MNTNS_SHIFT_UIDGID flag which is a mount namespace flag when
set mount points on filesystems that support UID/GID shifts will have
their UIDs and GIDs shifted by the VFS. The UID and GID mapping rules are per
mount namespace, they follow the rules of the user namespace of the containing
mount namespace. The UID/GID of inodes are supposed to always contain
the on-disk values, hence, the shift will be done inside VFS and it's a read
shift when we access the inodes.
This is a preparation patch.
Goal:
/* (1) */
clone4(CLONE_NEWNS|CLONE_MNTNS_SHIFT_UIDGID, ...)
/*
Setup container base mount namespace, rootfs and mount all
necessary mount points and filesystems that can't be mounted
in user namespaces. Filesystems that support uid/gid shifts
should set the mount parameters.
mount(..., mount_options=[vfs_shift_uids, vfs_shift_gids])
*/
/* (2) */
/*
Setup new mount and user namespaces and inherit the
CLONE_MNTNS_SHIFT_UIDGID flag from (1) into the new mount
namespace (2).
*/
clone4(CLONE_NEWUSER|CLONE_NEWNS|CLONE_MNTNS_SHIFT_UIDGID, ...)
/*
inodes of mount points here that support UID/GID shifts will have
automatically their UID/GID shifted according to the user
namespace rules of the current mount namespace (2).
*/
We create the new user and mount namespaces where:
1) The mount namespace allows mounts inside it that support UID and GID
shifting to perform the shifts if the CLONE_MNTNS_SHIFT_UIDGID is set
in the current mount namespace.
2) The UID and GID mapping is done according to the rules of the user
namespace of the containing mount namespace. The CLONE_MNTNS_SHIFT_UIDGID
follows the CLONE_NEWUSER|CLONE_NEWNS combination. This ensures that
only the creator of the mount namespace is able to adjust the user
namespace mapping rules.
The flag CLONE_MNTNS_SHIFT_UIDGID can be set on the mount namespace
only if:
1) The parent namespace has already CLONE_MNTNS_SHIFT_UIDGID set on
its mount namespace.
2) The caller has CAP_SYS_ADMIN in the init_user_ns namespace, since we
start from that namespace and we inherit some mount points we have to
protect files from privileged userns doing:
clone(CLONE_NEWUSER|CLONE_NEWNS|CLONE_MNTNS_SHIFT_UIDGID...)
This is blocked.
If a filesystem was mounted with "vfs_shift_uids" and "vfs_shift_gids"
and shows up in a mount namespace that does not include the
CLONE_MNTNS_SHIFT_UIDGID, then no shift is done. UIDs and GIDs will
not be changed at all, and things will continue to work as they are now.
Signed-off-by: Dongsu Park <dongsu@endocode.com>
Signed-off-by: Djalal Harouni <tixxdz@opendz.org>
---
fs/mount.h | 1 +
fs/namespace.c | 20 ++++++++++++++++++++
include/uapi/linux/sched.h | 1 +
kernel/fork.c | 4 ++++
4 files changed, 26 insertions(+)
diff --git a/fs/mount.h b/fs/mount.h
index 14db05d..1e317eb 100644
--- a/fs/mount.h
+++ b/fs/mount.h
@@ -6,6 +6,7 @@
struct mnt_namespace {
atomic_t count;
+ int flags;
struct ns_common ns;
struct mount * root;
struct list_head list;
diff --git a/fs/namespace.c b/fs/namespace.c
index 4fb1691..940ecfc 100644
--- a/fs/namespace.c
+++ b/fs/namespace.c
@@ -2774,6 +2774,7 @@ static struct mnt_namespace *alloc_mnt_ns(struct user_namespace *user_ns)
INIT_LIST_HEAD(&new_ns->list);
init_waitqueue_head(&new_ns->poll);
new_ns->event = 0;
+ new_ns->flags = 0;
new_ns->user_ns = get_user_ns(user_ns);
return new_ns;
}
@@ -2801,6 +2802,25 @@ struct mnt_namespace *copy_mnt_ns(unsigned long flags, struct mnt_namespace *ns,
if (IS_ERR(new_ns))
return new_ns;
+ if (flags & CLONE_MNTNS_SHIFT_UIDGID) {
+ /*
+ * If parent has the CLONE_MNTNS_SHIFT_UIDGID flag set
+ * or current is capable in init_user_ns, then we set the
+ * CLONE_MNTNS_SHIFT_UIDGID flag and allow mounts inside
+ * this namespace to shift their UID and GID.
+ *
+ * We check the init_user_ns here since we always start from
+ * that user namespace and mounts are by default available to all
+ * users. In this regard, only CAP_SYS_ADMIN in init_user_ns is
+ * allowed to start and propagate the CLONE_MNTNS_SHIFT_UIDGID
+ * flag to new mount namespaces.
+ */
+ if ((ns->flags & CLONE_MNTNS_SHIFT_UIDGID) || capable(CAP_SYS_ADMIN))
+ new_ns->flags |= CLONE_MNTNS_SHIFT_UIDGID;
+ else
+ return ERR_PTR(-EPERM);
+ }
+
namespace_lock();
/* First pass: copy the tree topology */
copy_flags = CL_COPY_UNBINDABLE | CL_EXPIRE;
diff --git a/include/uapi/linux/sched.h b/include/uapi/linux/sched.h
index 5f0fe01..9ba2124 100644
--- a/include/uapi/linux/sched.h
+++ b/include/uapi/linux/sched.h
@@ -19,6 +19,7 @@
#define CLONE_PARENT_SETTID 0x00100000 /* set the TID in the parent */
#define CLONE_CHILD_CLEARTID 0x00200000 /* clear the TID in the child */
#define CLONE_DETACHED 0x00400000 /* Unused, ignored */
+#define CLONE_MNTNS_SHIFT_UIDGID 0x00400000 /* If set allows to shift UID and GID for mounts that support it */
#define CLONE_UNTRACED 0x00800000 /* set if the tracing process can't force CLONE_PTRACE on this clone */
#define CLONE_CHILD_SETTID 0x01000000 /* set the TID in the child */
#define CLONE_NEWCGROUP 0x02000000 /* New cgroup namespace */
diff --git a/kernel/fork.c b/kernel/fork.c
index d277e83..41223cd 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -1264,6 +1264,10 @@ static struct task_struct *copy_process(unsigned long clone_flags,
if ((clone_flags & (CLONE_NEWNS|CLONE_FS)) == (CLONE_NEWNS|CLONE_FS))
return ERR_PTR(-EINVAL);
+ if ((clone_flags & CLONE_MNTNS_SHIFT_UIDGID) &&
+ !(clone_flags & CLONE_NEWNS))
+ return ERR_PTR(-EINVAL);
+
if ((clone_flags & (CLONE_NEWUSER|CLONE_FS)) == (CLONE_NEWUSER|CLONE_FS))
return ERR_PTR(-EINVAL);
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | Josh Triplett <josh@joshtriplett.org> |
|---|---|
| Date | 2016-05-04 18:40 +0200 |
| Message-ID | <rv80a-7v5-35@gated-at.bofh.it> |
| In reply to | #1394399 |
On Wed, May 04, 2016 at 04:26:46PM +0200, Djalal Harouni wrote: > This is version 2 of the VFS:userns support portable root filesystems > RFC. Changes since version 1: > > * Update documentation and remove some ambiguity about the feature. > Based on Josh Triplett comments. Thanks for the clarifications. > 3) The existing user namespace interface is the one used to do the > translation from virtual to on-disk mapping. This makes sense. Even if in the future we had a way to supply an arbitrary VFS UID/GID mapping for a mount, independent of the userns, what you've proposed would still make sense as a shorthand for the common case of using the same mapping for both userns and VFS. - Josh Triplett
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2016-05-04 23:10 +0200 |
| Message-ID | <rvcds-30d-27@gated-at.bofh.it> |
| In reply to | #1394399 |
On Wed, 2016-05-04 at 16:26 +0200, Djalal Harouni wrote: > This is version 2 of the VFS:userns support portable root filesystems > RFC. Changes since version 1: > > * Update documentation and remove some ambiguity about the feature. > Based on Josh Triplett comments. > * Use a new email address to send the RFC :-) > > > This RFC tries to explore how to support filesystem operations inside > user namespace using only VFS and a per mount namespace solution. > This > allows to take advantage of user namespace separations without > introducing any change at the filesystems level. All this is handled > with the virtual view of mount namespaces. > > > 1) Presentation: > ================ > > The main aim is to support portable root filesystems and allow > containers, virtual machines and other cases to use the same root > filesystem. Due to security reasons, filesystems can't be mounted > inside user namespaces, and mounting them outside will not solve the > problem since they will show up with the wrong UIDs/GIDs. Read and > write operations will also fail and so on. > > The current userspace solution is to automatically chown the whole > root filesystem before starting a container, example: > (host) init_user_ns 1000000:1065536 => (container) user_ns_X1 > 0:65535 > (host) init_user_ns 2000000:2065536 => (container) user_ns_Y1 > 0:65535 > (host) init_user_ns 3000000:3065536 => (container) user_ns_Z1 > 0:65535 > ... > > Every time a chown is called, files are changed and so on... This > prevents to have portable filesystems where you can throw anywhere > and boot. Having an extra step to adapt the filesystem to the current > mapping and persist it will not allow to verify its integrity, it > makes snapshots and migration a bit harder, and probably other > limitations... > > It seems that there are multiple ways to allow user namespaces > combine nicely with filesystems, but none of them is that easy. The > bind mount and pin the user namespace during mount time will not > work, bind mounts share the same super block, hence you may endup > working on the wrong vfsmount context and there is no easy way to get > out of that... So this option was discussed at the recent LSF/MM summit. The most supported suggestion was that you'd use a new internal fs type that had a struct mount with a new superblock and would copy the underlying inodes but substitute it's own with modified ->getatrr/->setattr calls that did the uid shift. In many ways it would be a remapping bind which would look similar to overlayfs but be a lot simpler. > Using the user namespace in the super block seems the way to go, and > there is the "Support fuse mounts in user namespaces" [1] patches > which seem nice but perhaps too complex!? So I don't think that does what you want. The fuse project I've used before to do uid/gid shifts for build containers is bindfs https://github.com/mpartel/bindfs/ It allows a --map argument where you specify pairs of uids/gids to map (tedious for large ranges, but the map can be fixed to use uid:range instead of individual). > there is also the overlayfs solution, and finaly the VFS layer > solution. > > We present here a simple VFS solution, everything is packed inside > VFS, filesystems don't need to know anything (except probably XFS, > and special operations inside union filesystems). Currently it > supports ext4, btrfs and overlayfs. Changes into filesystems are > small, just parse the vfs_shift_uids and vfs_shift_gids options > during mount and set the appropriate flags into the super_block > structure. So this looks a little daunting. It sprays the VFS with knowledge about the shifts and requires support from every underlying filesystem. A simple remapping bind filesystem would be a lot simpler and require no underlying filesystem support. James
[toc] | [prev] | [next] | [standalone]
| From | Djalal Harouni <tixxdz@gmail.com> |
|---|---|
| Date | 2016-05-05 09:40 +0200 |
| Message-ID | <rvm38-3Ek-7@gated-at.bofh.it> |
| In reply to | #1394722 |
On Wed, May 04, 2016 at 05:06:19PM -0400, James Bottomley wrote: > On Wed, 2016-05-04 at 16:26 +0200, Djalal Harouni wrote: > > This is version 2 of the VFS:userns support portable root filesystems > > RFC. Changes since version 1: > > > > * Update documentation and remove some ambiguity about the feature. > > Based on Josh Triplett comments. > > * Use a new email address to send the RFC :-) > > > > > > This RFC tries to explore how to support filesystem operations inside > > user namespace using only VFS and a per mount namespace solution. > > This > > allows to take advantage of user namespace separations without > > introducing any change at the filesystems level. All this is handled > > with the virtual view of mount namespaces. > > > > > > 1) Presentation: > > ================ > > > > The main aim is to support portable root filesystems and allow > > containers, virtual machines and other cases to use the same root > > filesystem. Due to security reasons, filesystems can't be mounted > > inside user namespaces, and mounting them outside will not solve the > > problem since they will show up with the wrong UIDs/GIDs. Read and > > write operations will also fail and so on. > > > > The current userspace solution is to automatically chown the whole > > root filesystem before starting a container, example: > > (host) init_user_ns 1000000:1065536 => (container) user_ns_X1 > > 0:65535 > > (host) init_user_ns 2000000:2065536 => (container) user_ns_Y1 > > 0:65535 > > (host) init_user_ns 3000000:3065536 => (container) user_ns_Z1 > > 0:65535 > > ... > > > > Every time a chown is called, files are changed and so on... This > > prevents to have portable filesystems where you can throw anywhere > > and boot. Having an extra step to adapt the filesystem to the current > > mapping and persist it will not allow to verify its integrity, it > > makes snapshots and migration a bit harder, and probably other > > limitations... > > > > It seems that there are multiple ways to allow user namespaces > > combine nicely with filesystems, but none of them is that easy. The > > bind mount and pin the user namespace during mount time will not > > work, bind mounts share the same super block, hence you may endup > > working on the wrong vfsmount context and there is no easy way to get > > out of that... > > So this option was discussed at the recent LSF/MM summit. The most > supported suggestion was that you'd use a new internal fs type that had > a struct mount with a new superblock and would copy the underlying > inodes but substitute it's own with modified ->getatrr/->setattr calls > that did the uid shift. In many ways it would be a remapping bind > which would look similar to overlayfs but be a lot simpler. Hmm, it's not only about ->getattr and ->setattr, you have all the other file system operations that need access too... which brings two points: 1) This new internal fs may end up doing what this RFC does... 2) or by quoting "new internal fs + its own super block + copy underlying inodes..." it seems like another overlayfs where you also need some decisions to copy what, etc. So, will this be really that light compared to current overlayfs ? not to mention that you need to hook up basically the same logic or something else inside overlayfs.. > > Using the user namespace in the super block seems the way to go, and > > there is the "Support fuse mounts in user namespaces" [1] patches > > which seem nice but perhaps too complex!? > > So I don't think that does what you want. The fuse project I've used > before to do uid/gid shifts for build containers is bindfs > > https://github.com/mpartel/bindfs/ > > It allows a --map argument where you specify pairs of uids/gids to map > (tedious for large ranges, but the map can be fixed to use uid:range > instead of individual). Ok, thanks for the link, will try to take a deep look but bindfs seem really big! > > there is also the overlayfs solution, and finaly the VFS layer > > solution. > > > > We present here a simple VFS solution, everything is packed inside > > VFS, filesystems don't need to know anything (except probably XFS, > > and special operations inside union filesystems). Currently it > > supports ext4, btrfs and overlayfs. Changes into filesystems are > > small, just parse the vfs_shift_uids and vfs_shift_gids options > > during mount and set the appropriate flags into the super_block > > structure. > > So this looks a little daunting. It sprays the VFS with knowledge > about the shifts and requires support from every underlying filesystem. Well, from my angle, shifts are just user namespace mappings which follow certain rules, and currently VFS and all filesystems are *already* doing some kind of shifting... This RFC uses mount namespaces which are the standard way to deal with mounts, now the mapping inside mount namespace can just be "inside: 0:1000" => "outside: 0:1000" and current implementation will just use it, at the same time I'm not sure if this mapping qualifies to be named "shift". I think that some folks here came up with the "shift" name to describe one of the use cases from a user interface that's it... maybe I should do s/vfs_shift_*/vfs_remap_*/ ? > A simple remapping bind filesystem would be a lot simpler and require > no underlying filesystem support. Yes probably, you still need to parse parameters but not at the filesystem level, and sure this RFC can do the same of course, but maybe it's not safe to shift/remap filesystems and their inodes on behalf of filesystems... and virtual filesystems which can share inodes ? > James > Thank you! -- Djalal Harouni http://opendz.org
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2016-05-05 14:00 +0200 |
| Message-ID | <rvq6K-7IA-17@gated-at.bofh.it> |
| In reply to | #1394910 |
On Thu, 2016-05-05 at 08:36 +0100, Djalal Harouni wrote: > On Wed, May 04, 2016 at 05:06:19PM -0400, James Bottomley wrote: > > On Wed, 2016-05-04 at 16:26 +0200, Djalal Harouni wrote: > > > This is version 2 of the VFS:userns support portable root > > > filesystems > > > RFC. Changes since version 1: > > > > > > * Update documentation and remove some ambiguity about the > > > feature. > > > Based on Josh Triplett comments. > > > * Use a new email address to send the RFC :-) > > > > > > > > > This RFC tries to explore how to support filesystem operations > > > inside user namespace using only VFS and a per mount namespace > > > solution. This allows to take advantage of user namespace > > > separations without introducing any change at the filesystems > > > level. All this is handled with the virtual view of mount > > > namespaces. > > > > > > > > > 1) Presentation: > > > ================ > > > > > > The main aim is to support portable root filesystems and allow > > > containers, virtual machines and other cases to use the same root > > > filesystem. Due to security reasons, filesystems can't be mounted > > > inside user namespaces, and mounting them outside will not solve > > > the problem since they will show up with the wrong UIDs/GIDs. > > > Read and write operations will also fail and so on. > > > > > > The current userspace solution is to automatically chown the > > > whole root filesystem before starting a container, example: > > > (host) init_user_ns 1000000:1065536 => (container) user_ns_X1 > > > 0:65535 > > > (host) init_user_ns 2000000:2065536 => (container) user_ns_Y1 > > > 0:65535 > > > (host) init_user_ns 3000000:3065536 => (container) user_ns_Z1 > > > 0:65535 > > > ... > > > > > > Every time a chown is called, files are changed and so on... This > > > prevents to have portable filesystems where you can throw > > > anywhere and boot. Having an extra step to adapt the filesystem > > > to the current mapping and persist it will not allow to verify > > > its integrity, it makes snapshots and migration a bit harder, and > > > probably other limitations... > > > > > > It seems that there are multiple ways to allow user namespaces > > > combine nicely with filesystems, but none of them is that easy. > > > The bind mount and pin the user namespace during mount time will > > > not work, bind mounts share the same super block, hence you may > > > endup working on the wrong vfsmount context and there is no easy > > > way to get out of that... > > > > So this option was discussed at the recent LSF/MM summit. The most > > supported suggestion was that you'd use a new internal fs type that > > had a struct mount with a new superblock and would copy the > > underlying inodes but substitute it's own with modified ->getatrr/ > > ->setattr calls that did the uid shift. In many ways it would be a > > remapping bind which would look similar to overlayfs but be a lot > > simpler. > > Hmm, it's not only about ->getattr and ->setattr, you have all the > other file system operations that need access too... Why? Or perhaps we should more cogently define the actual problem. My problem is simply mounting image volumes that were created with real uids at user namespace shifted uids because I'm downshifting the privileged ids in the container. I actually *only* need the uid/gids on the attributes shifted because that's what I need to manipulate the volumes. I actually think that other operations, like the file ioctl ones should, for security reasons, not be uid shifted. For instance with xfs you could set the panic mask and error tags and bring down the whole host. What extra things do you need access to and why? > which brings two points: > > 1) This new internal fs may end up doing what this RFC does... Well that was why I brought it up, yes. > 2) or by quoting "new internal fs + its own super block + copy > underlying inodes..." it seems like another overlayfs where you also > need some decisions to copy what, etc. So, will this be really > that light compared to current overlayfs ? not to mention that you > need to hook up basically the same logic or something else inside > overlayfs.. OK, so forget overlayfs, perhaps that was a bad example. It's like a uid shifting bind. The way it works is to use shadow inodes (unlike bind, but because you have to intercept the operations, so it's not a simple subtree operation) but there's no file copying. The shadow points to the real inode. > > > Using the user namespace in the super block seems the way to go, > > > and there is the "Support fuse mounts in user namespaces" [1] > > > patches which seem nice but perhaps too complex!? > > > > So I don't think that does what you want. The fuse project I've > > used before to do uid/gid shifts for build containers is bindfs > > > > https://github.com/mpartel/bindfs/ > > > > It allows a --map argument where you specify pairs of uids/gids to > > map (tedious for large ranges, but the map can be fixed to use > > uid:range instead of individual). > > Ok, thanks for the link, will try to take a deep look but bindfs seem > really big! Well, it does a lot more than just uid/gid shift. > > > there is also the overlayfs solution, and finaly the VFS layer > > > solution. > > > > > > We present here a simple VFS solution, everything is packed > > > inside VFS, filesystems don't need to know anything (except > > > probably XFS, and special operations inside union filesystems). > > > Currently it supports ext4, btrfs and overlayfs. Changes into > > > filesystems are small, just parse the vfs_shift_uids and > > > vfs_shift_gids options during mount and set the appropriate flags > > > into the super_block structure. > > > > So this looks a little daunting. It sprays the VFS with knowledge > > about the shifts and requires support from every underlying > > filesystem. > Well, from my angle, shifts are just user namespace mappings which > follow certain rules, and currently VFS and all filesystems are > *already* doing some kind of shifting... This RFC uses mount > namespaces which are the standard way to deal with mounts, now the > mapping inside mount namespace can just be "inside: 0:1000" => > "outside: 0:1000" and current implementation will just use it, at the > same time I'm not sure if this mapping qualifies to be named "shift". > I think that some folks here came up with the "shift" name to > describe one of the use cases from a user interface that's it... > maybe I should do s/vfs_shift_*/vfs_remap_*/ ? I don't think the naming is the issue ... it's the spread inside the vfs code (and in the underlying fs code). The vfs is very well layered, so touching all that code makes it look like there's a layering problem with the patch. Touching the underlying fs code looks even more problematic, but that may be necessary if you have a reason for wanting the file ioctls, because they're pass through and usually where the from_kuid() calls are in filesystems. > > A simple remapping bind filesystem would be a lot simpler and > > require no underlying filesystem support. > > Yes probably, you still need to parse parameters but not at the > filesystem level, They'd just be mount options. Basically instead of mount --bind source target, you'd do mount -t uidshift -o <shift options> source target. > and sure this RFC can do the same of course, but maybe it's not safe > to shift/remap filesystems and their inodes on behalf of > filesystems... and virtual filesystems which can share inodes ? That depends who you allow to do the shift. Each fstype in the kernel decides access to mount. For the uidshift, I was planning to allow only a capable admin in the initial namespace, meaning that only the admin in the host could set up the shifts. As long as the shifted filesystem is present, the container can then bind it wherever it wants in its mount namespace. James
[toc] | [prev] | [next] | [standalone]
| From | Djalal Harouni <tixxdz@gmail.com> |
|---|---|
| Date | 2016-05-06 00:00 +0200 |
| Message-ID | <rvzto-86w-15@gated-at.bofh.it> |
| In reply to | #1395060 |
On Thu, May 05, 2016 at 07:56:28AM -0400, James Bottomley wrote: > On Thu, 2016-05-05 at 08:36 +0100, Djalal Harouni wrote: > > On Wed, May 04, 2016 at 05:06:19PM -0400, James Bottomley wrote: > > > On Wed, 2016-05-04 at 16:26 +0200, Djalal Harouni wrote: > > > > This is version 2 of the VFS:userns support portable root > > > > filesystems > > > > RFC. Changes since version 1: > > > > > > > > * Update documentation and remove some ambiguity about the > > > > feature. > > > > Based on Josh Triplett comments. > > > > * Use a new email address to send the RFC :-) > > > > > > > > > > > > This RFC tries to explore how to support filesystem operations > > > > inside user namespace using only VFS and a per mount namespace > > > > solution. This allows to take advantage of user namespace > > > > separations without introducing any change at the filesystems > > > > level. All this is handled with the virtual view of mount > > > > namespaces. > > > > > > > > > > > > 1) Presentation: > > > > ================ > > > > > > > > The main aim is to support portable root filesystems and allow > > > > containers, virtual machines and other cases to use the same root > > > > filesystem. Due to security reasons, filesystems can't be mounted > > > > inside user namespaces, and mounting them outside will not solve > > > > the problem since they will show up with the wrong UIDs/GIDs. > > > > Read and write operations will also fail and so on. > > > > > > > > The current userspace solution is to automatically chown the > > > > whole root filesystem before starting a container, example: > > > > (host) init_user_ns 1000000:1065536 => (container) user_ns_X1 > > > > 0:65535 > > > > (host) init_user_ns 2000000:2065536 => (container) user_ns_Y1 > > > > 0:65535 > > > > (host) init_user_ns 3000000:3065536 => (container) user_ns_Z1 > > > > 0:65535 > > > > ... > > > > > > > > Every time a chown is called, files are changed and so on... This > > > > prevents to have portable filesystems where you can throw > > > > anywhere and boot. Having an extra step to adapt the filesystem > > > > to the current mapping and persist it will not allow to verify > > > > its integrity, it makes snapshots and migration a bit harder, and > > > > probably other limitations... > > > > > > > > It seems that there are multiple ways to allow user namespaces > > > > combine nicely with filesystems, but none of them is that easy. > > > > The bind mount and pin the user namespace during mount time will > > > > not work, bind mounts share the same super block, hence you may > > > > endup working on the wrong vfsmount context and there is no easy > > > > way to get out of that... > > > > > > So this option was discussed at the recent LSF/MM summit. The most > > > supported suggestion was that you'd use a new internal fs type that > > > had a struct mount with a new superblock and would copy the > > > underlying inodes but substitute it's own with modified ->getatrr/ > > > ->setattr calls that did the uid shift. In many ways it would be a > > > remapping bind which would look similar to overlayfs but be a lot > > > simpler. > > > > Hmm, it's not only about ->getattr and ->setattr, you have all the > > other file system operations that need access too... > > Why? Or perhaps we should more cogently define the actual problem. My > problem is simply mounting image volumes that were created with real > uids at user namespace shifted uids because I'm downshifting the > privileged ids in the container. I actually *only* need the uid/gids > on the attributes shifted because that's what I need to manipulate the We need them obviously for read/write/creation... ?! We want to handle also stock filesystems that were never edited without depending on any module or third party solution, mounting them outside user namespaces, and access inside. > volumes. I actually think that other operations, like the file ioctl > ones should, for security reasons, not be uid shifted. For instance > with xfs you could set the panic mask and error tags and bring down the > whole host. What extra things do you need access to and why? That's why precisely I said that mounting options not *inside* filesystems which means on their back, and on behalf of container managers, etc then you are exposed to such scenarios... some virtual file systems can also be mounted by unprivileged, how you will deal with something like a bind mount on them ? > > which brings two points: > > > > 1) This new internal fs may end up doing what this RFC does... > > Well that was why I brought it up, yes. yes but *with* extra code! that was my point. I'm not sure we need to bother with any *new* internal fs type nor hack around dir, file operations... yet that has to be shown, defined, coded ... ? > > 2) or by quoting "new internal fs + its own super block + copy > > underlying inodes..." it seems like another overlayfs where you also > > need some decisions to copy what, etc. So, will this be really > > that light compared to current overlayfs ? not to mention that you > > need to hook up basically the same logic or something else inside > > overlayfs.. > > OK, so forget overlayfs, perhaps that was a bad example. It's like a > uid shifting bind. The way it works is to use shadow inodes (unlike > bind, but because you have to intercept the operations, so it's not a > simple subtree operation) but there's no file copying. The shadow > points to the real inode. For that you need a super block struct for every mount... now if you also need a new internal fs + super block + shadowing inodes... it seems like you are going into overlayfs direction... I'm taking overlayfs as an example here, cause it's just nice and really dead simple! At the same time I'm not at all sure about what you are describing! and how you will deal with current mount and bind mounts tree and all the internals... > > > > Using the user namespace in the super block seems the way to go, > > > > and there is the "Support fuse mounts in user namespaces" [1] > > > > patches which seem nice but perhaps too complex!? > > > > > > So I don't think that does what you want. The fuse project I've > > > used before to do uid/gid shifts for build containers is bindfs > > > > > > https://github.com/mpartel/bindfs/ > > > > > > It allows a --map argument where you specify pairs of uids/gids to > > > map (tedious for large ranges, but the map can be fixed to use > > > uid:range instead of individual). > > > > Ok, thanks for the link, will try to take a deep look but bindfs seem > > really big! > > Well, it does a lot more than just uid/gid shift. > > > > > there is also the overlayfs solution, and finaly the VFS layer > > > > solution. > > > > > > > > We present here a simple VFS solution, everything is packed > > > > inside VFS, filesystems don't need to know anything (except > > > > probably XFS, and special operations inside union filesystems). > > > > Currently it supports ext4, btrfs and overlayfs. Changes into > > > > filesystems are small, just parse the vfs_shift_uids and > > > > vfs_shift_gids options during mount and set the appropriate flags > > > > into the super_block structure. > > > > > > So this looks a little daunting. It sprays the VFS with knowledge > > > about the shifts and requires support from every underlying > > > filesystem. > > > Well, from my angle, shifts are just user namespace mappings which > > follow certain rules, and currently VFS and all filesystems are > > *already* doing some kind of shifting... This RFC uses mount > > namespaces which are the standard way to deal with mounts, now the > > mapping inside mount namespace can just be "inside: 0:1000" => > > "outside: 0:1000" and current implementation will just use it, at the > > same time I'm not sure if this mapping qualifies to be named "shift". > > I think that some folks here came up with the "shift" name to > > describe one of the use cases from a user interface that's it... > > maybe I should do s/vfs_shift_*/vfs_remap_*/ ? > > I don't think the naming is the issue ... it's the spread inside the > vfs code (and in the underlying fs code). The vfs is very well Currently the underlying file systems just parse vfs_shift_uids and vfs_shif_gids > layered, so touching all that code makes it look like there's a > layering problem with the patch. Touching the underlying fs code looks Hmm, not sure I follow here ? We make use of the mount namespace which is part of the whole layer. Actually it's the *standard* way to control mounts. What do you mean here please ? > even more problematic, but that may be necessary if you have a reason > for wanting the file ioctls, because they're pass through and usually > where the from_kuid() calls are in filesystems. Hmm sorry, I'm not sure I'm following you here ? > > > A simple remapping bind filesystem would be a lot simpler and > > > require no underlying filesystem support. > > > > Yes probably, you still need to parse parameters but not at the > > filesystem level, > > They'd just be mount options. Basically instead of mount --bind source > target, you'd do mount -t uidshift -o <shift options> source target. > > > and sure this RFC can do the same of course, but maybe it's not safe > > to shift/remap filesystems and their inodes on behalf of > > filesystems... and virtual filesystems which can share inodes ? > > That depends who you allow to do the shift. Each fstype in the kernel > decides access to mount. For the uidshift, I was planning to allow > only a capable admin in the initial namespace, meaning that only the > admin in the host could set up the shifts. As long as the shifted > filesystem is present, the container can then bind it wherever it wants > in its mount namespace. Ah I see admin in initial namespace, yes sounds reasonable for security reasons, and how you will be able to achieve the user namespace shift ? > James > Thank you! -- Djalal Harouni http://opendz.org
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2016-05-06 00:10 +0200 |
| Message-ID | <rvzD3-9z-1@gated-at.bofh.it> |
| In reply to | #1395393 |
On Thu, 2016-05-05 at 22:49 +0100, Djalal Harouni wrote: > On Thu, May 05, 2016 at 07:56:28AM -0400, James Bottomley wrote: > > On Thu, 2016-05-05 at 08:36 +0100, Djalal Harouni wrote: > > > On Wed, May 04, 2016 at 05:06:19PM -0400, James Bottomley wrote: > > > > On Wed, 2016-05-04 at 16:26 +0200, Djalal Harouni wrote: > > > > > This is version 2 of the VFS:userns support portable root > > > > > filesystems > > > > > RFC. Changes since version 1: > > > > > > > > > > * Update documentation and remove some ambiguity about the > > > > > feature. Based on Josh Triplett comments. > > > > > * Use a new email address to send the RFC :-) > > > > > > > > > > > > > > > This RFC tries to explore how to support filesystem > > > > > operations inside user namespace using only VFS and a per > > > > > mount namespace solution. This allows to take advantage of > > > > > user namespace separations without introducing any change at > > > > > the filesystems level. All this is handled with the virtual > > > > > view of mount namespaces. > > > > > > > > > > > > > > > 1) Presentation: > > > > > ================ > > > > > > > > > > The main aim is to support portable root filesystems and > > > > > allow containers, virtual machines and other cases to use the > > > > > same root filesystem. Due to security reasons, filesystems > > > > > can't be mounted inside user namespaces, and mounting them > > > > > outside will not solve the problem since they will show up > > > > > with the wrong UIDs/GIDs. Read and write operations will also > > > > > fail and so on. > > > > > > > > > > The current userspace solution is to automatically chown the > > > > > whole root filesystem before starting a container, example: > > > > > (host) init_user_ns 1000000:1065536 => (container) > > > > > user_ns_X1 > > > > > 0:65535 > > > > > (host) init_user_ns 2000000:2065536 => (container) > > > > > user_ns_Y1 > > > > > 0:65535 > > > > > (host) init_user_ns 3000000:3065536 => (container) > > > > > user_ns_Z1 > > > > > 0:65535 > > > > > ... > > > > > > > > > > Every time a chown is called, files are changed and so on... > > > > > This prevents to have portable filesystems where you can > > > > > throw anywhere and boot. Having an extra step to adapt the > > > > > filesystem to the current mapping and persist it will not > > > > > allow to verify its integrity, it makes snapshots and > > > > > migration a bit harder, and probably other limitations... > > > > > > > > > > It seems that there are multiple ways to allow user > > > > > namespaces combine nicely with filesystems, but none of them > > > > > is that easy. The bind mount and pin the user namespace > > > > > during mount time will not work, bind mounts share the same > > > > > super block, hence you may endup working on the wrong > > > > > vfsmount context and there is no easy way to get out of > > > > > that... > > > > > > > > So this option was discussed at the recent LSF/MM summit. The > > > > most supported suggestion was that you'd use a new internal fs > > > > type that had a struct mount with a new superblock and would > > > > copy the underlying inodes but substitute it's own with > > > > modified ->getatrr/->setattr calls that did the uid shift. > > > > In many ways it would be a remapping bind which would look > > > > similar to overlayfs but be a lot simpler. > > > > > > Hmm, it's not only about ->getattr and ->setattr, you have all > > > the other file system operations that need access too... > > > > Why? Or perhaps we should more cogently define the actual problem. > > My problem is simply mounting image volumes that were created > > with real uids at user namespace shifted uids because I'm > > downshifting the privileged ids in the container. I actually > > *only* need the uid/gids on the attributes shifted because that's > > what I need to manipulate the > > > We need them obviously for read/write/creation... ?! OK, so the way attributes are populated on an inode is via getattr. You intercept that, you change the inode owner and group that are installed on the inode. That means that when you list the directory, you see the shift and the shifted uid/gid are used to check permissions for vfs_open(). > We want to handle also stock filesystems that were never edited > without depending on any module or third party solution, mounting > them outside user namespaces, and access inside. OK, but that's basically my requirements ... you didn't mention any of the esoteric filesystem ioctls, so I assume from the below you're not interested in shifting the uids there either? > > volumes. I actually think that other operations, like the file > > ioctl ones should, for security reasons, not be uid shifted. For > > instance with xfs you could set the panic mask and error tags and > > bring down the whole host. What extra things do you need access to > > and why? > > That's why precisely I said that mounting options not *inside* > filesystems which means on their back, and on behalf of container > managers, etc then you are exposed to such scenarios... some virtual > file systems can also be mounted by unprivileged, how you will deal > with something like a bind mount on them ? > > > > > which brings two points: > > > > > > 1) This new internal fs may end up doing what this RFC does... > > > > Well that was why I brought it up, yes. > > yes but *with* extra code! that was my point. I'm not sure we need to > bother with any *new* internal fs type nor hack around dir, file > operations... yet that has to be shown, defined, coded ... ? Either way requires patching the kernel. The question I was asking is is it better to confine the patch to a new fs type or directly change the vfs. > > > 2) or by quoting "new internal fs + its own super block + copy > > > underlying inodes..." it seems like another overlayfs where you > > > also need some decisions to copy what, etc. So, will this be > > > really that light compared to current overlayfs ? not to mention > > > that you need to hook up basically the same logic or something > > > else inside overlayfs.. > > > > OK, so forget overlayfs, perhaps that was a bad example. It's like > > a uid shifting bind. The way it works is to use shadow inodes > > (unlike bind, but because you have to intercept the operations, so > > it's not a simple subtree operation) but there's no file copying. > > The shadow points to the real inode. > > For that you need a super block struct for every mount... now if you > also need a new internal fs + super block + shadowing inodes... it > seems like you are going into overlayfs direction... Well, that's the way you build a shadowing fs. I'm not sure you need one sb per struct vfs mount, but it's certainly possible to code it that way. > I'm taking overlayfs as an example here, cause it's just nice and > really dead simple! > > At the same time I'm not at all sure about what you are describing! > and how you will deal with current mount and bind mounts tree and all > the internals... You mean would MS_REC functionality be supported? There's no reason why not, but there's no reason you have to either (it could even be optional, like it is for bind). > > > > Using the user namespace in the super block seems the way to > > > > go, and there is the "Support fuse mounts in user namespaces" > > > > [1] patches which seem nice but perhaps too complex!? > > > > > > > > So I don't think that does what you want. The fuse project > > > > I've used before to do uid/gid shifts for build containers is > > > > bindfs https://github.com/mpartel/bindfs/ > > > > > > > > It allows a --map argument where you specify pairs of uids/gids > > > > to map (tedious for large ranges, but the map can be fixed to > > > > use uid:range instead of individual). > > > > > > Ok, thanks for the link, will try to take a deep look but bindfs > > > seem really big! > > > > Well, it does a lot more than just uid/gid shift. > > > > > > > there is also the overlayfs solution, and finaly the VFS > > > > > layer solution. > > > > > > > > > > We present here a simple VFS solution, everything is packed > > > > > inside VFS, filesystems don't need to know anything (except > > > > > probably XFS, and special operations inside union > > > > > filesystems). Currently it supports ext4, btrfs and > > > > > overlayfs. Changes into filesystems are small, just parse the > > > > > vfs_shift_uids and vfs_shift_gids options during mount and > > > > > set the appropriate flags into the super_block structure. > > > > > > > > So this looks a little daunting. It sprays the VFS with > > > > knowledge about the shifts and requires support from every > > > > underlying filesystem. > > > > > Well, from my angle, shifts are just user namespace mappings > > > which follow certain rules, and currently VFS and all filesystems > > > are *already* doing some kind of shifting... This RFC uses mount > > > namespaces which are the standard way to deal with mounts, now > > > the mapping inside mount namespace can just be "inside: 0:1000" > > > => "outside: 0:1000" and current implementation will just use it, > > > at the same time I'm not sure if this mapping qualifies to be > > > named "shift". I think that some folks here came up with the > > > "shift" name to describe one of the use cases from a user > > > interface that's it... maybe I should do > > > s/vfs_shift_*/vfs_remap_*/ ? > > > > I don't think the naming is the issue ... it's the spread inside > > the vfs code (and in the underlying fs code). The vfs is very well > > Currently the underlying file systems just parse vfs_shift_uids and > vfs_shif_gids > > > layered, so touching all that code makes it look like there's a > > layering problem with the patch. Touching the underlying fs code > > looks > > > Hmm, not sure I follow here ? We make use of the mount namespace > which is part of the whole layer. Actually it's the *standard* way to > control mounts. What do you mean here please ? The patch touches a lot of the vfs. > > even more problematic, but that may be necessary if you have a > > reason for wanting the file ioctls, because they're pass through > > and usually where the from_kuid() calls are in filesystems. > > Hmm sorry, I'm not sure I'm following you here ? An ideal solution, given both our requirements, shouldn't require touching any underlying fs code. > > > > A simple remapping bind filesystem would be a lot simpler and > > > > require no underlying filesystem support. > > > > > > Yes probably, you still need to parse parameters but not at the > > > filesystem level, > > > > They'd just be mount options. Basically instead of mount --bind > > source target, you'd do mount -t uidshift -o <shift options> source > > target. > > > > > and sure this RFC can do the same of course, but maybe it's not > > > safe to shift/remap filesystems and their inodes on behalf of > > > filesystems... and virtual filesystems which can share inodes ? > > > > That depends who you allow to do the shift. Each fstype in the > > kernel decides access to mount. For the uidshift, I was planning > > to allow only a capable admin in the initial namespace, meaning > > that only the admin in the host could set up the shifts. As long > > as the shifted filesystem is present, the container can then bind > > it wherever it wants in its mount namespace. > > Ah I see admin in initial namespace, yes sounds reasonable for > security reasons, and how you will be able to achieve the user > namespace shift? As I said, it would be in the mount options of the command. The <shift options> above. Probably parametrised the way uid_map and gid_map are today. James
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2016-05-11 01:40 +0200 |
| Message-ID | <rxppU-4nc-15@gated-at.bofh.it> |
| In reply to | #1395396 |
On Thu, 2016-05-05 at 18:08 -0400, James Bottomley wrote:
> On Thu, 2016-05-05 at 22:49 +0100, Djalal Harouni wrote:
> > On Thu, May 05, 2016 at 07:56:28AM -0400, James Bottomley wrote:
> > > On Thu, 2016-05-05 at 08:36 +0100, Djalal Harouni wrote:
> > > > On Wed, May 04, 2016 at 05:06:19PM -0400, James Bottomley
> > > > wrote:
[...]
> > > > > So this option was discussed at the recent LSF/MM summit.
> > > > > The
> > > > > most supported suggestion was that you'd use a new internal
> > > > > fs
> > > > > type that had a struct mount with a new superblock and would
> > > > > copy the underlying inodes but substitute it's own with
> > > > > modified ->getatrr/->setattr calls that did the uid shift.
> > > > > In many ways it would be a remapping bind which would look
> > > > > similar to overlayfs but be a lot simpler.
> > > >
> > > > Hmm, it's not only about ->getattr and ->setattr, you have all
> > > > the other file system operations that need access too...
> > >
> > > Why? Or perhaps we should more cogently define the actual
> > > problem. My problem is simply mounting image volumes that were
> > > created with real uids at user namespace shifted uids because I'm
> > > downshifting the privileged ids in the container. I actually
> > > *only* need the uid/gids on the attributes shifted because that's
> > > what I need to manipulate the
> > >
> > We need them obviously for read/write/creation... ?!
>
> OK, so the way attributes are populated on an inode is via getattr.
> You intercept that, you change the inode owner and group that are
> installed on the inode. That means that when you list the directory,
> you see the shift and the shifted uid/gid are used to check
> permissions for vfs_open().
Just to illustrate how this could be done, here's a functional proof of
concept for a uid/gid shifting bind mount equivalent. It's not
actually a proper bind mount because it has to manufacture its own
inodes. As you can see, it can only be used by root, it will shift all
the uid/gid bits as well as the permission comparisons. It operates on
subtrees, so it can shift the uids/gids on any filesystem or part of
one and because the shifts are per superblock, it could actually shift
the same subtree for multiple users on different shifts. Best of all,
it requires no vfs changes at all, being entirely implemented inside
its own filesystem type.
You use it just like bind mount:
mount -t shiftfs <source> <target>
except that it takes uidshift=x:y:z and gidshift=x:y:z multiple times
as options. It's currently not recursive and it definitely needs
polishing to show things like mount options and be properly Kconfig
using.
There's a bit of an open question of whether it should have vfs
changes: the way the struct file f_inode and f_ops are hijacked is a
bit nasty and perhaps d_select_inode() could be made a bit cleverer to
help us here instead.
James
---
fs/Makefile | 1 +
fs/shiftfs.c | 790 +++++++++++++++++++++++++++++++++++++++++++++
include/uapi/linux/magic.h | 2 +
3 files changed, 793 insertions(+)
diff --git a/fs/Makefile b/fs/Makefile
index 85b6e13..bad03b2 100644
--- a/fs/Makefile
+++ b/fs/Makefile
@@ -128,3 +128,4 @@ obj-y += exofs/ # Multiple modules
obj-$(CONFIG_CEPH_FS) += ceph/
obj-$(CONFIG_PSTORE) += pstore/
obj-$(CONFIG_EFIVAR_FS) += efivarfs/
+obj-m += shiftfs.o
diff --git a/fs/shiftfs.c b/fs/shiftfs.c
new file mode 100644
index 0000000..b40cdfe
--- /dev/null
+++ b/fs/shiftfs.c
@@ -0,0 +1,790 @@
+#include <linux/cred.h>
+#include <linux/mount.h>
+#include <linux/file.h>
+#include <linux/fs.h>
+#include <linux/namei.h>
+#include <linux/module.h>
+#include <linux/kernel.h>
+#include <linux/magic.h>
+#include <linux/parser.h>
+#include <linux/statfs.h>
+#include <linux/slab.h>
+#include <linux/user_namespace.h>
+#include <linux/uidgid.h>
+
+struct shiftfs_super_info {
+ struct vfsmount *mnt;
+ struct uid_gid_map uid_map, gid_map;
+};
+
+static struct inode *shiftfs_new_inode(struct super_block *sb, umode_t mode,
+ struct dentry *dentry);
+
+enum {
+ OPT_UIDMAP,
+ OPT_GIDMAP,
+ OPT_LAST,
+};
+
+/* global filesystem options */
+static const match_table_t tokens = {
+ { OPT_UIDMAP, "uidmap=%u:%u:%u" },
+ { OPT_GIDMAP, "gidmap=%u:%u:%u" },
+ { OPT_LAST, NULL }
+};
+
+/*
+ * code stolen from user_namespace.c ... except that these functions
+ * return the same id back if unmapped ... should probably have a
+ * library?
+ */
+static u32 map_id_down(struct uid_gid_map *map, u32 id)
+{
+ unsigned idx, extents;
+ u32 first, last;
+
+ /* Find the matching extent */
+ extents = map->nr_extents;
+ smp_rmb();
+ for (idx = 0; idx < extents; idx++) {
+ first = map->extent[idx].first;
+ last = first + map->extent[idx].count - 1;
+ if (id >= first && id <= last)
+ break;
+ }
+ /* Map the id or note failure */
+ if (idx < extents)
+ id = (id - first) + map->extent[idx].lower_first;
+
+ return id;
+}
+
+static u32 map_id_up(struct uid_gid_map *map, u32 id)
+{
+ unsigned idx, extents;
+ u32 first, last;
+
+ /* Find the matching extent */
+ extents = map->nr_extents;
+ smp_rmb();
+ for (idx = 0; idx < extents; idx++) {
+ first = map->extent[idx].lower_first;
+ last = first + map->extent[idx].count - 1;
+ if (id >= first && id <= last)
+ break;
+ }
+ /* Map the id or note failure */
+ if (idx < extents)
+ id = (id - first) + map->extent[idx].first;
+
+ return id;
+}
+
+static bool mappings_overlap(struct uid_gid_map *new_map,
+ struct uid_gid_extent *extent)
+{
+ u32 upper_first, lower_first, upper_last, lower_last;
+ unsigned idx;
+
+ upper_first = extent->first;
+ lower_first = extent->lower_first;
+ upper_last = upper_first + extent->count - 1;
+ lower_last = lower_first + extent->count - 1;
+
+ for (idx = 0; idx < new_map->nr_extents; idx++) {
+ u32 prev_upper_first, prev_lower_first;
+ u32 prev_upper_last, prev_lower_last;
+ struct uid_gid_extent *prev;
+
+ prev = &new_map->extent[idx];
+
+ prev_upper_first = prev->first;
+ prev_lower_first = prev->lower_first;
+ prev_upper_last = prev_upper_first + prev->count - 1;
+ prev_lower_last = prev_lower_first + prev->count - 1;
+
+ /* Does the upper range intersect a previous extent? */
+ if ((prev_upper_first <= upper_last) &&
+ (prev_upper_last >= upper_first))
+ return true;
+
+ /* Does the lower range intersect a previous extent? */
+ if ((prev_lower_first <= lower_last) &&
+ (prev_lower_last >= lower_first))
+ return true;
+ }
+ return false;
+}
+/* end code stolen from user_namespace.c */
+
+static const struct cred *shiftfs_get_up_creds(struct super_block *sb)
+{
+ struct cred *cred = prepare_creds();
+ struct shiftfs_super_info *ssi = sb->s_fs_info;
+
+ if (!cred)
+ return NULL;
+
+ cred->fsuid = KUIDT_INIT(map_id_up(&ssi->uid_map, __kuid_val(cred->fsuid)));
+ cred->fsgid = KGIDT_INIT(map_id_up(&ssi->gid_map, __kgid_val(cred->fsgid)));
+
+ return cred;
+}
+
+static const struct cred *shiftfs_new_creds(const struct cred **newcred,
+ struct super_block *sb)
+{
+ const struct cred *cred = shiftfs_get_up_creds(sb);
+
+ *newcred = cred;
+
+ if (cred)
+ cred = override_creds(cred);
+ else
+ printk(KERN_ERR "Credential override failed: no memory\n");
+
+ return cred;
+}
+
+static void shiftfs_old_creds(const struct cred *oldcred,
+ const struct cred **newcred)
+{
+ if (!*newcred)
+ return;
+
+ revert_creds(oldcred);
+ put_cred(*newcred);
+}
+
+static int shiftfs_parse_options(struct shiftfs_super_info *ssi, char *options)
+{
+ char *p;
+ substring_t args[MAX_OPT_ARGS];
+ int from, to, count;
+ struct uid_gid_map *map, *maps[2] = {
+ [OPT_UIDMAP] = &ssi->uid_map,
+ [OPT_GIDMAP] = &ssi->gid_map,
+ };
+
+ while ((p = strsep(&options, ",")) != NULL) {
+ int token;
+ struct uid_gid_extent ext;
+
+ if (!*p)
+ continue;
+
+ token = match_token(p, tokens, args);
+ if (token != OPT_UIDMAP && token != OPT_GIDMAP)
+ return -EINVAL;
+ if (match_int(&args[0], &from) ||
+ match_int(&args[1], &to) ||
+ match_int(&args[2], &count))
+ return -EINVAL;
+ map = maps[token];
+ if (map->nr_extents >= UID_GID_MAP_MAX_EXTENTS)
+ return -EINVAL;
+ ext.first = from;
+ ext.lower_first = to;
+ ext.count = count;
+ if (mappings_overlap(map, &ext))
+ return -EINVAL;
+ map->extent[map->nr_extents++] = ext;
+ }
+ return 0;
+}
+
+static void shiftfs_d_iput(struct dentry *dentry, struct inode *inode)
+{
+ struct dentry *real = inode->i_private;
+
+ dput(real);
+ iput(inode);
+}
+
+static const struct dentry_operations shiftfs_dentry_ops = {
+ .d_iput = shiftfs_d_iput,
+};
+
+static int shiftfs_readlink(struct dentry *dentry, char __user *data,
+ int flags)
+{
+ struct dentry *real = dentry->d_inode->i_private;
+ const struct inode_operations *iop = real->d_inode->i_op;
+
+ if (iop->readlink)
+ return iop->readlink(real, data, flags);
+
+ return -EINVAL;
+}
+
+static const char *shiftfs_get_link(struct dentry *dentry, struct inode *inode,
+ struct delayed_call *done)
+{
+ if (dentry) {
+ struct dentry *real = dentry->d_inode->i_private;
+ struct inode *reali = real->d_inode;
+ const struct inode_operations *iop = reali->i_op;
+ const char *res = ERR_PTR(-EPERM);
+
+ if (iop->get_link)
+ res = iop->get_link(real, reali, done);
+
+ return res;
+ } else {
+ /* RCU lookup not supported */
+ return ERR_PTR(-ECHILD);
+ }
+}
+
+static int shiftfs_setxattr(struct dentry *dentry, const char *name,
+ const void *value, size_t size, int flags)
+{
+ struct dentry *real = dentry->d_inode->i_private;
+ const struct inode_operations *iop = real->d_inode->i_op;
+ int err = -EOPNOTSUPP;
+
+ if (iop->setxattr) {
+ const struct cred *oldcred, *newcred;
+
+ oldcred = shiftfs_new_creds(&newcred, dentry->d_sb);
+ err = iop->setxattr(real, name, value, size, flags);
+ shiftfs_old_creds(oldcred, &newcred);
+ }
+
+ return err;
+}
+
+static ssize_t shiftfs_getxattr(struct dentry *dentry, const char *name,
+ void *value, size_t size)
+{
+ struct dentry *real = dentry->d_inode->i_private;
+ const struct inode_operations *iop = real->d_inode->i_op;
+ int err = -EOPNOTSUPP;
+
+ if (iop->getxattr) {
+ const struct cred *oldcred, *newcred;
+
+ oldcred = shiftfs_new_creds(&newcred, dentry->d_sb);
+ err = iop->getxattr(real, name, value, size);
+ shiftfs_old_creds(oldcred, &newcred);
+ }
+
+ return err;
+}
+
+static ssize_t shiftfs_listxattr(struct dentry *dentry, char *list,
+ size_t size)
+{
+ struct dentry *real = dentry->d_inode->i_private;
+ const struct inode_operations *iop = real->d_inode->i_op;
+
+ if (iop->listxattr)
+ return iop->listxattr(real, list, size);
+
+ return -EINVAL;
+}
+
+static int shiftfs_removexattr(struct dentry *dentry, const char *name)
+{
+ struct dentry *real = dentry->d_inode->i_private;
+ const struct inode_operations *iop = real->d_inode->i_op;
+
+ if (iop->removexattr)
+ return iop->removexattr(real, name);
+
+ return -EINVAL;
+}
+
+static void shiftfs_fill_inode(struct inode *inode, struct dentry *dentry)
+{
+ struct inode *reali;
+ struct shiftfs_super_info *ssi = inode->i_sb->s_fs_info;
+
+ if (!dentry)
+ return;
+
+ reali = dentry->d_inode;
+
+ if (!reali->i_op->get_link)
+ inode->i_opflags |= IOP_NOFOLLOW;
+
+ inode->i_mapping = reali->i_mapping;
+ inode->i_private = dentry;
+
+ inode->i_uid = KUIDT_INIT(map_id_down(&ssi->uid_map, __kuid_val(reali->i_uid)));
+ inode->i_gid = KGIDT_INIT(map_id_down(&ssi->gid_map, __kgid_val(reali->i_gid)));
+}
+
+static int shiftfs_make_object(struct inode *dir, struct dentry *dentry,
+ umode_t mode, const char *symlink,
+ struct dentry *hardlink, bool excl)
+{
+ struct dentry *real = dir->i_private, *new;
+ struct inode *reali = real->d_inode, *newi;
+ const struct inode_operations *iop = reali->i_op;
+ int err;
+ const struct cred *oldcred, *newcred;
+ bool op_ok = false;
+
+ if (hardlink) {
+ op_ok = iop->link;
+ } else {
+ switch (mode & S_IFMT) {
+ case S_IFDIR:
+ op_ok = iop->mkdir;
+ break;
+ case S_IFREG:
+ op_ok = iop->create;
+ break;
+ case S_IFLNK:
+ op_ok = iop->symlink;
+ }
+ }
+ if (!op_ok)
+ return -EINVAL;
+
+
+ newi = shiftfs_new_inode(dentry->d_sb, mode, NULL);
+ if (!newi)
+ return -ENOMEM;
+
+ oldcred = shiftfs_new_creds(&newcred, dentry->d_sb);
+
+ inode_lock_nested(reali, I_MUTEX_PARENT);
+ new = lookup_one_len(dentry->d_name.name, real, dentry->d_name.len);
+ err = PTR_ERR(new);
+ if (IS_ERR(new))
+ goto out_unlock;
+
+ if (hardlink) {
+ struct dentry *realhardlink = hardlink->d_inode->i_private;
+
+ err = vfs_link(new, reali, realhardlink, NULL);
+ } else {
+ switch (mode & S_IFMT) {
+ case S_IFDIR:
+ err = vfs_mkdir(reali, new, mode);
+ break;
+ case S_IFREG:
+ err = vfs_create(reali, new, mode, excl);
+ break;
+ case S_IFLNK:
+ err = vfs_symlink(reali, new, symlink);
+ }
+ }
+
+ shiftfs_old_creds(oldcred, &newcred);
+
+ if (err)
+ goto out_dput;
+
+ shiftfs_fill_inode(newi, new);
+
+ d_instantiate(dentry, newi);
+
+ new = NULL;
+ newi = NULL;
+
+ out_dput:
+ dput(new);
+ out_unlock:
+ iput(newi);
+ inode_unlock(reali);
+
+ return err;
+}
+
+static int shiftfs_create(struct inode *dir, struct dentry *dentry,
+ umode_t mode, bool excl)
+{
+ mode |= S_IFREG;
+
+ return shiftfs_make_object(dir, dentry, mode, NULL, NULL, excl);
+}
+
+static int shiftfs_mkdir(struct inode *dir, struct dentry *dentry,
+ umode_t mode)
+{
+ mode |= S_IFDIR;
+
+ return shiftfs_make_object(dir, dentry, mode, NULL, NULL, false);
+}
+
+static int shiftfs_link(struct dentry *dentry, struct inode *dir,
+ struct dentry *hardlink)
+{
+ return shiftfs_make_object(dir, dentry, 0, NULL, hardlink, false);
+}
+
+static int shiftfs_symlink(struct inode *dir, struct dentry *dentry,
+ const char *symlink)
+{
+ return shiftfs_make_object(dir, dentry, S_IFLNK, symlink, NULL, false);
+}
+
+static int shiftfs_rm(struct inode *dir, struct dentry *dentry, bool rmdir)
+{
+ struct dentry *real = dir->i_private, *new;
+ struct inode *reali = real->d_inode;
+ int err;
+ const struct cred *oldcred, *newcred;
+
+ inode_lock_nested(reali, I_MUTEX_PARENT);
+
+ oldcred = shiftfs_new_creds(&newcred, dentry->d_sb);
+
+ new = lookup_one_len(dentry->d_name.name, real, dentry->d_name.len);
+ err = PTR_ERR(new);
+ if (IS_ERR(new))
+ goto out_unlock;
+
+ if (rmdir)
+ err = vfs_rmdir(reali, new);
+ else
+ err = vfs_unlink(reali, new, NULL);
+
+ dput(new);
+
+ out_unlock:
+ shiftfs_old_creds(oldcred, &newcred);
+ inode_unlock(reali);
+
+ return err;
+}
+
+static int shiftfs_unlink(struct inode *dir, struct dentry *dentry)
+{
+ return shiftfs_rm(dir, dentry, false);
+}
+
+static int shiftfs_rmdir(struct inode *dir, struct dentry *dentry)
+{
+ return shiftfs_rm(dir, dentry, true);
+}
+
+static int shiftfs_rename2(struct inode *olddir, struct dentry *old,
+ struct inode *newdir, struct dentry *new,
+ unsigned int flags)
+{
+ struct dentry *rodd = olddir->i_private, *rndd = newdir->i_private,
+ *realold = old->d_inode->i_private,
+ *realnew = new->d_inode->i_private;
+ struct inode *realolddir = rodd->d_inode, *realnewdir = rndd->d_inode;
+ const struct inode_operations *iop = realolddir->i_op;
+ int err;
+ const struct cred *oldcred, *newcred;
+
+ oldcred = shiftfs_new_creds(&newcred, old->d_sb);
+ err = iop->rename2(realolddir, realold, realnewdir, realnew, flags);
+ shiftfs_old_creds(oldcred, &newcred);
+
+ return err;
+}
+
+static struct dentry *shiftfs_lookup(struct inode *dir, struct dentry *dentry,
+ unsigned int flags)
+{
+ struct dentry *real = dir->i_private, *new;
+ struct inode *reali = real->d_inode, *newi;
+ const struct cred *oldcred, *newcred;
+
+ /* note: violation of usual fs rules here: dentries are never
+ * added with d_add. This is because we want no dentry cache
+ * for shiftfs. All lookups proceed through the dentry cache
+ * of the underlying filesystem, meaning we always see any
+ * changes in the underlying */
+
+ inode_lock(reali);
+ oldcred = shiftfs_new_creds(&newcred, dentry->d_sb);
+ new = lookup_one_len(dentry->d_name.name, real, dentry->d_name.len);
+ shiftfs_old_creds(oldcred, &newcred);
+ inode_unlock(reali);
+
+ if (IS_ERR(new) || !new)
+ return new;
+
+ if (!new->d_inode) {
+ dput(new);
+ return NULL;
+ }
+
+ newi = shiftfs_new_inode(dentry->d_sb, new->d_inode->i_mode, new);
+ if (!newi) {
+ dput(new);
+ return ERR_PTR(-ENOMEM);
+ }
+
+ d_instantiate(dentry, newi);
+
+ return NULL;
+}
+
+static int shiftfs_permission(struct inode *inode, int mask)
+{
+ struct dentry *real = inode->i_private;
+ struct inode *reali = real->d_inode;
+ const struct inode_operations *iop = reali->i_op;
+ int err;
+ const struct cred *oldcred, *newcred;
+
+ oldcred = shiftfs_new_creds(&newcred, inode->i_sb);
+ if (iop->permission)
+ err = iop->permission(reali, mask);
+ else
+ err = generic_permission(reali, mask);
+ shiftfs_old_creds(oldcred, &newcred);
+
+ return err;
+}
+
+static int shiftfs_setattr(struct dentry *dentry, struct iattr *attr)
+{
+ struct dentry *real = dentry->d_inode->i_private;
+ struct inode *reali = real->d_inode;
+ const struct inode_operations *iop = reali->i_op;
+ struct iattr newattr = *attr;
+ const struct cred *oldcred, *newcred;
+ struct shiftfs_super_info *ssi = dentry->d_sb->s_fs_info;
+ int err;
+
+ newattr.ia_uid = KUIDT_INIT(map_id_up(&ssi->uid_map, __kuid_val(attr->ia_uid)));
+ newattr.ia_gid = KGIDT_INIT(map_id_up(&ssi->gid_map, __kgid_val(attr->ia_gid)));
+
+ oldcred = shiftfs_new_creds(&newcred, dentry->d_sb);
+ if (iop->setattr)
+ err = iop->setattr(real, &newattr);
+ else
+ err = simple_setattr(real, &newattr);
+ shiftfs_old_creds(oldcred, &newcred);
+
+ return err;
+}
+
+static int shiftfs_getattr(struct vfsmount *mnt, struct dentry *dentry,
+ struct kstat *stat)
+{
+ struct inode *inode = dentry->d_inode;
+ struct dentry *real = inode->i_private;
+ struct inode *reali = real->d_inode;
+ const struct inode_operations *iop = reali->i_op;
+ int err = 0;
+
+ mnt = dentry->d_sb->s_fs_info;
+
+ if (iop->getattr)
+ err = iop->getattr(mnt, real, stat);
+ else
+ generic_fillattr(reali, stat);
+
+ if (err)
+ return err;
+
+ stat->uid = inode->i_uid;
+ stat->gid = inode->i_gid;
+ return 0;
+}
+
+struct shiftfs_fop_carrier {
+ struct inode *inode;
+ int (*release)(struct inode *, struct file *);
+ struct file_operations fop;
+};
+
+static int shiftfs_release(struct inode *inode, struct file *file)
+{
+ struct shiftfs_fop_carrier *sfc;
+ int err = 0;
+
+ sfc = container_of(file->f_op, struct shiftfs_fop_carrier, fop);
+
+ if (sfc->release)
+ err = sfc->release(inode, file);
+
+ file->f_inode = sfc->inode;
+ file->f_op = sfc->inode->i_fop;
+
+ kfree(sfc);
+
+ return err;
+}
+
+static int shiftfs_open(struct inode *inode, struct file *file)
+{
+ struct dentry *real = inode->i_private;
+ struct inode *reali = real->d_inode;
+ const struct file_operations *fop;
+ struct shiftfs_fop_carrier *sfc;
+ int err = 0;
+
+ sfc = kmalloc(sizeof(*sfc), GFP_KERNEL);
+ if (!sfc)
+ return -ENOMEM;
+
+ if (real->d_flags & DCACHE_OP_SELECT_INODE)
+ reali = real->d_op->d_select_inode(real, file->f_flags);
+
+ fop = reali->i_fop;
+ sfc->inode = inode;
+ memcpy(&sfc->fop, fop, sizeof(*fop));
+ sfc->release = sfc->fop.release;
+ sfc->fop.release = shiftfs_release;
+
+ file->f_op = &sfc->fop;
+ file->f_inode = reali;
+
+ if (fop->open)
+ err = fop->open(reali, file);
+
+ return err;
+}
+
+static const struct inode_operations shiftfs_inode_ops = {
+ /* intercepted */
+ .lookup = shiftfs_lookup,
+ .getattr = shiftfs_getattr,
+ .setattr = shiftfs_setattr,
+ .permission = shiftfs_permission,
+
+ /*pass though */
+ .mkdir = shiftfs_mkdir,
+ .symlink = shiftfs_symlink,
+ .get_link = shiftfs_get_link,
+ .readlink = shiftfs_readlink,
+ .unlink = shiftfs_unlink,
+ .rmdir = shiftfs_rmdir,
+ .rename2 = shiftfs_rename2,
+ .link = shiftfs_link,
+ .create = shiftfs_create,
+ .mknod = NULL, /* no special files currently */
+ .setxattr = shiftfs_setxattr,
+ .getxattr = shiftfs_getxattr,
+ .listxattr = shiftfs_listxattr,
+ .removexattr = shiftfs_removexattr,
+};
+
+static const struct file_operations shiftfs_file_ops = {
+ .open = shiftfs_open,
+};
+
+static struct inode *shiftfs_new_inode(struct super_block *sb, umode_t mode,
+ struct dentry *dentry)
+{
+ struct inode *inode;
+
+ inode = new_inode(sb);
+ if (!inode)
+ return NULL;
+
+ mode &= S_IFMT;
+
+ inode->i_ino = get_next_ino();
+ inode->i_mode = mode;
+ inode->i_flags |= S_NOATIME | S_NOCMTIME;
+
+ inode->i_op = &shiftfs_inode_ops;
+ inode->i_fop = &shiftfs_file_ops;
+
+ shiftfs_fill_inode(inode, dentry);
+
+ return inode;
+}
+
+static void shiftfs_put_super(struct super_block *sb)
+{
+ struct shiftfs_super_info *ssi = sb->s_fs_info;
+
+ mntput(ssi->mnt);
+ kfree(ssi);
+}
+
+static const struct super_operations shiftfs_super_ops = {
+ .put_super = shiftfs_put_super,
+};
+
+struct shiftfs_data {
+ void *data;
+ const char *path;
+};
+
+static int shiftfs_fill_super(struct super_block *sb, void *raw_data,
+ int silent)
+{
+ struct shiftfs_data *data = raw_data;
+ char *name = kstrdup(data->path, GFP_KERNEL);
+ int err = -ENOMEM;
+ struct shiftfs_super_info *ssi = NULL;
+ struct path path;
+
+ if (!name)
+ goto out;
+
+ ssi = kzalloc(sizeof(*ssi), GFP_KERNEL);
+ if (!ssi)
+ goto out;
+
+ err = -EPERM;
+ if (!capable(CAP_SYS_ADMIN))
+ goto out;
+
+ err = shiftfs_parse_options(ssi, data->data);
+ if (err)
+ goto out;
+
+ err = kern_path(name, LOOKUP_FOLLOW, &path);
+ if (err)
+ goto out;
+
+ if (!S_ISDIR(path.dentry->d_inode->i_mode)) {
+ err = -ENOTDIR;
+ goto out_put;
+ }
+ ssi->mnt = path.mnt;
+
+ sb->s_fs_info = ssi;
+ sb->s_magic = SHIFTFS_MAGIC;
+ sb->s_op = &shiftfs_super_ops;
+ sb->s_d_op = &shiftfs_dentry_ops;
+ sb->s_root = d_make_root(shiftfs_new_inode(sb, S_IFDIR, path.dentry));
+
+ return 0;
+
+ out_put:
+ path_put(&path);
+ out:
+ kfree(name);
+ if (err)
+ kfree(ssi);
+ return err;
+}
+
+static struct dentry *shiftfs_mount(struct file_system_type *fs_type,
+ int flags, const char *dev_name, void *data)
+{
+ struct shiftfs_data d = { data, dev_name };
+
+ return mount_nodev(fs_type, flags, &d, shiftfs_fill_super);
+}
+
+static struct file_system_type shiftfs_type = {
+ .owner = THIS_MODULE,
+ .name = "shiftfs",
+ .mount = shiftfs_mount,
+ .kill_sb = kill_anon_super,
+};
+
+static int __init shiftfs_init(void)
+{
+ return register_filesystem(&shiftfs_type);
+}
+
+static void __exit shiftfs_exit(void)
+{
+ unregister_filesystem(&shiftfs_type);
+}
+
+MODULE_ALIAS_FS("shiftfs");
+MODULE_AUTHOR("James Bottomley");
+MODULE_DESCRIPTION("uid/gid shifting bind filesystem");
+MODULE_LICENSE("GPL v2");
+module_init(shiftfs_init)
+module_exit(shiftfs_exit)
diff --git a/include/uapi/linux/magic.h b/include/uapi/linux/magic.h
index 0de181a..d7992f5 100644
--- a/include/uapi/linux/magic.h
+++ b/include/uapi/linux/magic.h
@@ -79,4 +79,6 @@
#define NSFS_MAGIC 0x6e736673
#define BPF_FS_MAGIC 0xcafe4a11
+#define SHIFTFS_MAGIC 0x6a656a62
+
#endif /* __LINUX_MAGIC_H__ */
[toc] | [prev] | [next] | [standalone]
| From | Al Viro <viro@ZenIV.linux.org.uk> |
|---|---|
| Date | 2016-05-11 02:40 +0200 |
| Message-ID | <rxqlY-5pF-7@gated-at.bofh.it> |
| In reply to | #1398590 |
On Tue, May 10, 2016 at 04:36:56PM -0700, James Bottomley wrote: > mount -t shiftfs <source> <target> Note to self: do not eat while reading l-k...
[toc] | [prev] | [next] | [standalone]
| From | Al Viro <viro@ZenIV.linux.org.uk> |
|---|---|
| Date | 2016-05-11 03:00 +0200 |
| Message-ID | <rxqFk-5B4-5@gated-at.bofh.it> |
| In reply to | #1398590 |
On Tue, May 10, 2016 at 04:36:56PM -0700, James Bottomley wrote:
> +static int shiftfs_rename2(struct inode *olddir, struct dentry *old,
> + struct inode *newdir, struct dentry *new,
> + unsigned int flags)
> +{
> + struct dentry *rodd = olddir->i_private, *rndd = newdir->i_private,
> + *realold = old->d_inode->i_private,
> + *realnew = new->d_inode->i_private;
> + struct inode *realolddir = rodd->d_inode, *realnewdir = rndd->d_inode;
> + const struct inode_operations *iop = realolddir->i_op;
> + int err;
> + const struct cred *oldcred, *newcred;
> +
> + oldcred = shiftfs_new_creds(&newcred, old->d_sb);
> + err = iop->rename2(realolddir, realold, realnewdir, realnew, flags);
> + shiftfs_old_creds(oldcred, &newcred);
... and you've just violated all locking rules for ->rename2().
> +static struct dentry *shiftfs_lookup(struct inode *dir, struct dentry *dentry,
> + unsigned int flags)
> +{
> + struct dentry *real = dir->i_private, *new;
> + struct inode *reali = real->d_inode, *newi;
> + const struct cred *oldcred, *newcred;
> +
> + /* note: violation of usual fs rules here: dentries are never
> + * added with d_add. This is because we want no dentry cache
> + * for shiftfs. All lookups proceed through the dentry cache
> + * of the underlying filesystem, meaning we always see any
> + * changes in the underlying */
Bloody wonderful. So
* we lose caching the negative lookups
* we've got buggered hardlinks (different inodes for those)
* it has never, ever been tried on -next (would do rather nasty
things on that d_instantiate())
> +
> + kfree(sfc);
> +
> + return err;
> +}
> + file->f_op = &sfc->fop;
Lovely - now try that with underlying fs something built modular.
Or try to use it on top of something with non-trivial dentry_operations
(hell, on top of itself, for starters).
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2016-05-11 05:50 +0200 |
| Message-ID | <rxtjQ-8qN-3@gated-at.bofh.it> |
| In reply to | #1398619 |
On Wed, 2016-05-11 at 01:53 +0100, Al Viro wrote:
> On Tue, May 10, 2016 at 04:36:56PM -0700, James Bottomley wrote:
> > +static int shiftfs_rename2(struct inode *olddir, struct dentry
> > *old,
> > + struct inode *newdir, struct dentry
> > *new,
> > + unsigned int flags)
> > +{
> > + struct dentry *rodd = olddir->i_private, *rndd = newdir
> > ->i_private,
> > + *realold = old->d_inode->i_private,
> > + *realnew = new->d_inode->i_private;
> > + struct inode *realolddir = rodd->d_inode, *realnewdir =
> > rndd->d_inode;
> > + const struct inode_operations *iop = realolddir->i_op;
> > + int err;
> > + const struct cred *oldcred, *newcred;
> > +
> > + oldcred = shiftfs_new_creds(&newcred, old->d_sb);
> > + err = iop->rename2(realolddir, realold, realnewdir,
> > realnew, flags);
> > + shiftfs_old_creds(oldcred, &newcred);
>
> ... and you've just violated all locking rules for ->rename2().
Yes, sorry, somehow I missed that when I converted everything else to
the vfs_ functions.
> > +static struct dentry *shiftfs_lookup(struct inode *dir, struct
> > dentry *dentry,
> > + unsigned int flags)
> > +{
> > + struct dentry *real = dir->i_private, *new;
> > + struct inode *reali = real->d_inode, *newi;
> > + const struct cred *oldcred, *newcred;
> > +
> > + /* note: violation of usual fs rules here: dentries are
> > never
> > + * added with d_add. This is because we want no dentry
> > cache
> > + * for shiftfs. All lookups proceed through the dentry
> > cache
> > + * of the underlying filesystem, meaning we always see any
> > + * changes in the underlying */
>
> Bloody wonderful. So
> * we lose caching the negative lookups
We do? They should be cached in the underlying layer's dcache. If
that's not enough, I can hash them, but I was trying to avoid doubling
the dcache size.
> * we've got buggered hardlinks (different inodes for those)
Yes, had a note to do the lookup, but forgot.
> * it has never, ever been tried on -next (would do rather nasty
> things on that d_instantiate())
So this is just a proof of concept; I figured it was best to do it
against current rather than have people who wanted to try it pull in
your tree. I can respin it after the merge window closes.
>
> > +
> > + kfree(sfc);
> > +
> > + return err;
> > +}
>
> > + file->f_op = &sfc->fop;
>
> Lovely - now try that with underlying fs something built modular.
>
> Or try to use it on top of something with non-trivial
> dentry_operations
> (hell, on top of itself, for starters).
So if I add the missing fops_get/put, you're happy with the way this
hijacks f_op and f_inode?
James
[toc] | [prev] | [next] | [standalone]
| From | Djalal Harouni <tixxdz@gmail.com> |
|---|---|
| Date | 2016-05-11 18:50 +0200 |
| Message-ID | <rxFuG-3yH-7@gated-at.bofh.it> |
| In reply to | #1398590 |
On Tue, May 10, 2016 at 04:36:56PM -0700, James Bottomley wrote: > On Thu, 2016-05-05 at 18:08 -0400, James Bottomley wrote: [...] > > > > OK, so the way attributes are populated on an inode is via getattr. > > You intercept that, you change the inode owner and group that are > > installed on the inode. That means that when you list the directory, > > you see the shift and the shifted uid/gid are used to check > > permissions for vfs_open(). > > Just to illustrate how this could be done, here's a functional proof of > concept for a uid/gid shifting bind mount equivalent. It's not > actually a proper bind mount because it has to manufacture its own > inodes. As you can see, it can only be used by root, it will shift all > the uid/gid bits as well as the permission comparisons. It operates on > subtrees, so it can shift the uids/gids on any filesystem or part of > one and because the shifts are per superblock, it could actually shift > the same subtree for multiple users on different shifts. Best of all, > it requires no vfs changes at all, being entirely implemented inside > its own filesystem type. First, I guess this should be in a separate thread.. this way this RFC was just hijacked! Obviously as you say later in your response it may require a VFS change... You have just consumed all inodes... what about containers or small apps that are spawned quickly... it can even used maybe as a DoS... maybe you endup reporting different inode numbers... ? > You use it just like bind mount: > > mount -t shiftfs <source> <target> > > except that it takes uidshift=x:y:z and gidshift=x:y:z multiple times > as options. It's currently not recursive and it definitely needs > polishing to show things like mount options and be properly Kconfig > using. why it's not recursive ? and what if you have circular bind mounts ? Hmm anyway you are mounting this on behalf of filesystems, so if you add the recursive thing, you will just probably make everything worse, by making any /proc, /sys dentry that's under that path shiftable, and unprivileged users can just create user namespaces and read /proc/* and all the other stuff that doesn't have capable() related to the init_user_ns host... what if you have paths like /filesystem0/uidshiftedY/dir, /filesystem0/uidshiftedX/dir , /filesystem0/notshifted/dir where some of them are also bind mounts that point to same dentry ? Also, you create a totally new user namespace interface here! by making your own new interface we just lose the notion of init_user_ns and its children and mapping ? I'm not sure of the implication of all this... your user namespace mapping is not related at all to init_user_ns! it seems that it has its own init_user_ns ? does a capable() check now on a shifted filesystem relates to that and hence to your mapping or to the real init_user_ns ? > There's a bit of an open question of whether it should have vfs > changes: the way the struct file f_inode and f_ops are hijacked is a > bit nasty and perhaps d_select_inode() could be made a bit cleverer to > help us here instead. I'm not sure if this PoC works... but you sure you didn't introduce a serious vulnerability here ? you use a new mapping and you update current_fsuid() creds up, which is global on any fs operation, so may be: lets operate on any inode, update our current_fsuid()... and access the rest of *unshifted filesystems*... !? The worst thing is that current_fsuid() does not follow now the /proc/self/uid_map interface! this is a serious vulnerability and a mix of the current semantics... it's updated but using other rules...? For overlayfs I did write an expriment but for me it's not an overlayfs or another new filesystem problem... we are manipulating UID/GID identities... It would have been better if you did send this as a separate thread. It was a vfs:userns RFC fix which if we continue we turn it into a complicated thing! implement another new light filesystem with userns... (overlayfs...) Will follow up if the appropriate thread is created, not here, I guess it's ok ? > James > Thank you for your feedback! -- Djalal Harouni http://opendz.org
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2016-05-11 20:40 +0200 |
| Message-ID | <rxHd8-5ub-15@gated-at.bofh.it> |
| In reply to | #1399352 |
On Wed, 2016-05-11 at 17:42 +0100, Djalal Harouni wrote: > On Tue, May 10, 2016 at 04:36:56PM -0700, James Bottomley wrote: > > On Thu, 2016-05-05 at 18:08 -0400, James Bottomley wrote: > [...] > > > > > > OK, so the way attributes are populated on an inode is via > > > getattr. You intercept that, you change the inode owner and > > > group that are installed on the inode. That means that when you > > > list the directory, you see the shift and the shifted uid/gid are > > > used to check permissions for vfs_open(). > > > > Just to illustrate how this could be done, here's a functional > > proof of concept for a uid/gid shifting bind mount equivalent. > > It's not actually a proper bind mount because it has to > > manufacture its own inodes. As you can see, it can only be used by > > root, it will shift all the uid/gid bits as well as the permission > > comparisons. It operates on subtrees, so it can shift the > > uids/gids on any filesystem or part of one and because the shifts > > are per superblock, it could actually shift the same subtree for > > multiple users on different shifts. Best of all, it requires no > > vfs changes at all, being entirely implemented inside its own > > filesystem type. > > First, I guess this should be in a separate thread.. this way this > RFC was just hijacked! > > Obviously as you say later in your response it may require a VFS > change... I thought it may but viro didn't rip my head off for shifting the file operations and inode, so perhaps it's OK as is. > You have just consumed all inodes... what about containers or small > apps that are spawned quickly... it can even used maybe as a DoS... > maybe you endup reporting different inode numbers... ? Please explain? Shiftfs deliberately doesn't populate its dentry cache, so it basically has the same number inodes and dentries in use as the lower filesystem would ordinarily have. > > > You use it just like bind mount: > > > > mount -t shiftfs <source> <target> > > > > except that it takes uidshift=x:y:z and gidshift=x:y:z multiple > > times > > as options. It's currently not recursive and it definitely needs > > polishing to show things like mount options and be properly Kconfig > > using. > > why it's not recursive ? and what if you have circular bind mounts ? Because, as I said, it's a proof of concept. It can easily have MS_REC semantics added. > Hmm anyway you are mounting this on behalf of filesystems, so if you > add the recursive thing, you will just probably make everything > worse, by making any /proc, /sys dentry that's under that path > shiftable, and unprivileged users can just create user namespaces and > read /proc/* and all the other stuff that doesn't have capable() > related to the init_user_ns host... That's up to the admin who does the shifting. Recursive would be an option if added. > what if you have paths like /filesystem0/uidshiftedY/dir, > /filesystem0/uidshiftedX/dir , /filesystem0/notshifted/dir > where some of them are also bind mounts that point to same dentry ? Without recursive semantics, you see the underlying inode. With them, you see the upper vfsmnts. Shiftfs isn't idempotent, so you would need to be careful about nesting. However, that's an admin problem. > Also, you create a totally new user namespace interface here! by > making your own new interface we just lose the notion of init_user_ns > and its children and mapping ? I don't quite understand this; the only use of the init_user_ns is the capable(CAP_SYS_ADMIN) in fill_super which is how only the real admin can mount at a shifted uid/gid. Otherwise, there's no need to see into the userns because filesystems see the kuid_t/kgid_t which is what I'm shifting. > I'm not sure of the implication of all this... your user namespace > mapping is not related at all to init_user_ns! it seems that it has > its own init_user_ns ? does a capable() check now on a shifted > filesystem relates to that and hence to your mapping or to the real > init_user_ns ? capable(CAP_SYS_ADMIN) == ns_capable(&init_user_ns, CAP_SYS_ADMIN) Or is there a misunderstanding here about how user namespaces work inside the kernel? The design is that the ID shift is done as you cross the kernel boundary, so a filesystem, being usually all in-kernel operating via the VFS interfaces, ideally never needs to make any from_kuid/make_kuid calls. However, there are ways filesystems can send data across the kernel/user bounary outside of the usual vfs interfaces (ioctls being the most usual one) so in that specific code, they have to do the kuid_t to uid_t changes themselves. Shiftfs never sends data to the user outside of the VFS so it never needs to do this and can operate entirely on kuid_ts. > > There's a bit of an open question of whether it should have vfs > > changes: the way the struct file f_inode and f_ops are hijacked is > > a bit nasty and perhaps d_select_inode() could be made a bit > > cleverer to help us here instead. > > I'm not sure if this PoC works... but you sure you didn't introduce > a serious vulnerability here ? you use a new mapping and you update > current_fsuid() creds up, which is global on any fs operation, so may > be: lets operate on any inode, update our current_fsuid()... and > access the rest of *unshifted filesystems*... !? The credentials are per thread, so it's a standard way of doing credential shifting and no other threads of execution in the same task get access. As long as you bound the override_creds()/revert_creds() pairs within the kernel, you're safe. > The worst thing is that current_fsuid() does not follow now the > /proc/self/uid_map interface! this is a serious vulnerability and a > mix of the current semantics... it's updated but using other > rules...? current_fsuid() is aready mapped via the userns; it's already a kuid_t at its final value. Shifting that is what you want to remap underlying volume uid/gid's. The uidmap/gidmap inputs to this are shifts on the final underlying uid/gids. So, if I've got a uid_map in a userns of 0:100000:1000 which remaps all the privileged ids down to 100000, but I have a volume which still has realids, I can mount that volume using shiftfs with uidmap=0:100000:1000 and it will allow this userns to read and write the volume through its remapped ids. > For overlayfs I did write an expriment but for me it's not an > overlayfs or another new filesystem problem... we are manipulating > UID/GID identities... > > It would have been better if you did send this as a separate thread. > It was a vfs:userns RFC fix which if we continue we turn it into a > complicated thing! implement another new light filesystem with > userns... (overlayfs...) > > Will follow up if the appropriate thread is created, not here, I > guess it's ok ? Well, I can resend the patch as a separate thread when I've fixed some of the problems viro pointed out. James > > James > > > > Thank you for your feedback! > >
[toc] | [prev] | [next] | [standalone]
| From | Djalal Harouni <tixxdz@gmail.com> |
|---|---|
| Date | 2016-05-12 22:00 +0200 |
| Message-ID | <ry4W6-3UF-29@gated-at.bofh.it> |
| In reply to | #1399413 |
On Wed, May 11, 2016 at 11:33:38AM -0700, James Bottomley wrote: > On Wed, 2016-05-11 at 17:42 +0100, Djalal Harouni wrote: > > On Tue, May 10, 2016 at 04:36:56PM -0700, James Bottomley wrote: [...] > > Hmm anyway you are mounting this on behalf of filesystems, so if you > > add the recursive thing, you will just probably make everything > > worse, by making any /proc, /sys dentry that's under that path > > shiftable, and unprivileged users can just create user namespaces and > > read /proc/* and all the other stuff that doesn't have capable() > > related to the init_user_ns host... > > That's up to the admin who does the shifting. Recursive would be an > option if added. Hmm, not sure if you get my point... you just made it an admin problem where admins want to mount an image downloaded verify it and use it for their container with /proc...! that's another problem! > > what if you have paths like /filesystem0/uidshiftedY/dir, > > /filesystem0/uidshiftedX/dir , /filesystem0/notshifted/dir > > where some of them are also bind mounts that point to same dentry ? > > Without recursive semantics, you see the underlying inode. With them, > you see the upper vfsmnts. Shiftfs isn't idempotent, so you would need > to be careful about nesting. However, that's an admin problem. > > > Also, you create a totally new user namespace interface here! by > > making your own new interface we just lose the notion of init_user_ns > > and its children and mapping ? > > I don't quite understand this; the only use of the init_user_ns is the > capable(CAP_SYS_ADMIN) in fill_super which is how only the real admin > can mount at a shifted uid/gid. Otherwise, there's no need to see into > the userns because filesystems see the kuid_t/kgid_t which is what I'm > shifting. > > > I'm not sure of the implication of all this... your user namespace > > mapping is not related at all to init_user_ns! it seems that it has > > its own init_user_ns ? does a capable() check now on a shifted > > filesystem relates to that and hence to your mapping or to the real > > init_user_ns ? > > capable(CAP_SYS_ADMIN) == ns_capable(&init_user_ns, CAP_SYS_ADMIN) > > Or is there a misunderstanding here about how user namespaces work > inside the kernel? The design is that the ID shift is done as you > cross the kernel boundary, so a filesystem, being usually all in-kernel > operating via the VFS interfaces, ideally never needs to make any > from_kuid/make_kuid calls. However, there are ways filesystems can > send data across the kernel/user bounary outside of the usual vfs > interfaces (ioctls being the most usual one) so in that specific code, > they have to do the kuid_t to uid_t changes themselves. Shiftfs never > sends data to the user outside of the VFS so it never needs to do this > and can operate entirely on kuid_ts. > > > > There's a bit of an open question of whether it should have vfs > > > changes: the way the struct file f_inode and f_ops are hijacked is > > > a bit nasty and perhaps d_select_inode() could be made a bit > > > cleverer to help us here instead. > > > > I'm not sure if this PoC works... but you sure you didn't introduce > > a serious vulnerability here ? you use a new mapping and you update > > current_fsuid() creds up, which is global on any fs operation, so may > > be: lets operate on any inode, update our current_fsuid()... and > > access the rest of *unshifted filesystems*... !? > > The credentials are per thread, so it's a standard way of doing > credential shifting and no other threads of execution in the same task > get access. As long as you bound the override_creds()/revert_creds() > pairs within the kernel, you're safe. No, and here sorry I mean shifted. current_fsuid() is global through all fs operations which means it crosses user namespaces... it was safe the days of only init_user_ns, not anymore... You give a mapping inside containers to fsuid where they don't want to have it... this allows to operate on inodes inside other containers... update current_fsuid() even if we want that user to be nobody inside the container... and later it can access the inodes of the shifted fs... and by same current of course... > > The worst thing is that current_fsuid() does not follow now the > > /proc/self/uid_map interface! this is a serious vulnerability and a > > mix of the current semantics... it's updated but using other > > rules...? > > current_fsuid() is aready mapped via the userns; it's already a kuid_t > at its final value. Shifting that is what you want to remap underlying > volume uid/gid's. The uidmap/gidmap inputs to this are shifts on the > final underlying uid/gids. => some points: Changing setfsuid() its interfaces and rules... or an indrect way to break another syscall... The userns used for *mapping* is totatly different and not standard... losing "init_user_ns and its decendents userns *semantics*...", a yet a totatly unlinked mapping... Breaking current_uid(),current_euid(),current_fsuid() which are mapped but in *different* user namespaces... hence different values inside namespaces... you can change your userns mapping but that current_fsuid specific one will always be remapped to some other value inside even if you don't want it... It crosses user namespaces... uid and euid are remapped according to /proc/self/uid_map, fsuid is remapped according to this new interface... Hard coding the mapping, nested containers/apps may *share* fsuid and can't get rid of it even if they change the inside userns mapping to disable, split, reduce mapped users or offer better isolation they can't... no way to make private inodes inside containers if they share the final fsuid, inside container mapping is ignored... ... > the privileged ids down to 100000, but I have a volume which still has > realids, I can mount that volume using shiftfs with > uidmap=0:100000:1000 and it will allow this userns to read and write > the volume through its remapped ids. > > > For overlayfs I did write an expriment but for me it's not an > > overlayfs or another new filesystem problem... we are manipulating > > UID/GID identities... > > > > It would have been better if you did send this as a separate thread. > > It was a vfs:userns RFC fix which if we continue we turn it into a > > complicated thing! implement another new light filesystem with > > userns... (overlayfs...) > > > > Will follow up if the appropriate thread is created, not here, I > > guess it's ok ? > > Well, I can resend the patch as a separate thread when I've fixed some > of the problems viro pointed out. > > James > > > > James > > > > > > > Thank you for your feedback! > > > > > Thanks! -- Djalal Harouni http://opendz.org
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2016-05-13 00:30 +0200 |
| Message-ID | <ry7hl-6XB-7@gated-at.bofh.it> |
| In reply to | #1400325 |
On Thu, 2016-05-12 at 20:55 +0100, Djalal Harouni wrote: > On Wed, May 11, 2016 at 11:33:38AM -0700, James Bottomley wrote: > > On Wed, 2016-05-11 at 17:42 +0100, Djalal Harouni wrote: > > > On Tue, May 10, 2016 at 04:36:56PM -0700, James Bottomley wrote: > [...] > > > Hmm anyway you are mounting this on behalf of filesystems, so if > > > you > > > add the recursive thing, you will just probably make everything > > > worse, by making any /proc, /sys dentry that's under that path > > > shiftable, and unprivileged users can just create user namespaces > > > and > > > read /proc/* and all the other stuff that doesn't have capable() > > > related to the init_user_ns host... > > > > That's up to the admin who does the shifting. Recursive would be > > an > > option if added. > > Hmm, not sure if you get my point... you just made it an admin > problem where admins want to mount an image downloaded verify it and > use it for their container with /proc...! that's another problem! You can't allow unprivileged containers to shift uids on arbitrary filesystems, so the admin always has to do something for the initial setup. > > > what if you have paths like /filesystem0/uidshiftedY/dir, > > > /filesystem0/uidshiftedX/dir , /filesystem0/notshifted/dir > > > where some of them are also bind mounts that point to same dentry > > > ? > > > > Without recursive semantics, you see the underlying inode. With > > them, you see the upper vfsmnts. Shiftfs isn't idempotent, so you > > would need to be careful about nesting. However, that's an admin > > problem. > > > > > Also, you create a totally new user namespace interface here! by > > > making your own new interface we just lose the notion of > > > init_user_ns and its children and mapping ? > > > > I don't quite understand this; the only use of the init_user_ns is > > the capable(CAP_SYS_ADMIN) in fill_super which is how only the real > > admin can mount at a shifted uid/gid. Otherwise, there's no need > > to see into the userns because filesystems see the kuid_t/kgid_t > > which is what I'm shifting. > > > > > I'm not sure of the implication of all this... your user > > > namespace mapping is not related at all to init_user_ns! it seems > > > that it has its own init_user_ns ? does a capable() check now > > > on a shifted filesystem relates to that and hence to your mapping > > > or to the real init_user_ns ? > > > > capable(CAP_SYS_ADMIN) == ns_capable(&init_user_ns, CAP_SYS_ADMIN) > > > > Or is there a misunderstanding here about how user namespaces work > > inside the kernel? The design is that the ID shift is done as you > > cross the kernel boundary, so a filesystem, being usually all in > > -kernel operating via the VFS interfaces, ideally never needs to > > make any from_kuid/make_kuid calls. However, there are ways > > filesystems can send data across the kernel/user bounary outside of > > the usual vfs interfaces (ioctls being the most usual one) so in > > that specific code, they have to do the kuid_t to uid_t changes > > themselves. Shiftfs never sends data to the user outside of the > > VFS so it never needs to do this and can operate entirely on > > kuid_ts. > > > > > > There's a bit of an open question of whether it should have vfs > > > > changes: the way the struct file f_inode and f_ops are hijacked > > > > is a bit nasty and perhaps d_select_inode() could be made a bit > > > > cleverer to help us here instead. > > > > > > I'm not sure if this PoC works... but you sure you didn't > > > introduce a serious vulnerability here ? you use a new mapping > > > and you update current_fsuid() creds up, which is global on any > > > fs operation, so may be: lets operate on any inode, update our > > > current_fsuid()... and access the rest of *unshifted filesystems* > > > ... !? > > > > The credentials are per thread, so it's a standard way of doing > > credential shifting and no other threads of execution in the same > > task get access. As long as you bound the override_creds()/revert_c > > reds() pairs within the kernel, you're safe. > > No, and here sorry I mean shifted. > > current_fsuid() is global through all fs operations which means it > crosses user namespaces... it was safe the days of only init_user_ns, > not anymore... You give a mapping inside containers to fsuid where > they don't want to have it... this allows to operate on inodes inside > other containers... update current_fsuid() even if we want that user > to be nobody inside the container... and later it can access the > inodes of the shifted fs... and by same current of course... OK, I still don't understand what you're getting at. There are three per-thread uids: uid, euid and fsuid (real, effective and filesystem). They're all either settable via syscall or inherited on fork. They're all kernel side, meaning they're kuid_t. Their values stay invariant as you move through namespaces. They change (and get mapped according to the current user namespace setting) when you call set[fe]uid() So when I enter a user namespace with mapping 0 100000 1000 and call setuid(0) (which sets all three). they all pick up the kuid_t of 100000. This means that writing a file inside the user namespace after calling setuid(0) appears as real uid 100000 on the medium even though if I call getuid() from the namespace, I get back 0. What shiftfs does is hijack temporarily the kernel fsuid/fsgid for permission checks, so you can remap to any old uid on the medium (although usually you'd pass in uidmap=0:100000:1000") it maps back from kuid_t 100000 to kuid_t 0, which is why the container can now read and write the underlying medium at on-media id 0 even through root inside the container has kuid_t 100000. There's no permanent change of fsuid and it stays at its invariant value for the thread except as a temporary measure to do the permission checks on the underlying of the shifted filesystem. > > > The worst thing is that current_fsuid() does not follow now the > > > /proc/self/uid_map interface! this is a serious vulnerability and > > > a mix of the current semantics... it's updated but using other > > > rules...? > > > > current_fsuid() is aready mapped via the userns; it's already a > > kuid_t at its final value. Shifting that is what you want to remap > > underlying volume uid/gid's. The uidmap/gidmap inputs to this are > > shifts on the final underlying uid/gids. > > => some points: > Changing setfsuid() its interfaces and rules... or an indrect way to > break another syscall... There is no change to setfsuid(). > The userns used for *mapping* is totatly different and not standard.. > . losing "init_user_ns and its decendents userns *semantics*...", a > yet a totatly unlinked mapping... There is no user namespace mapping at all. This is a simple shift, kernel side, of uids and gids at their kuid_t values. > Breaking current_uid(),current_euid(),current_fsuid() which are > mapped but in *different* user namespaces... hence different values > inside namespaces... you can change your userns mapping but that > current_fsuid specific one will always be remapped to some other > value inside even if you don't want it... It crosses user > namespaces... uid and euid are remapped according to /proc/self/uid_ > map, fsuid is remapped according to this new interface... > > Hard coding the mapping, nested containers/apps may *share* fsuid and > can't get rid of it even if they change the inside userns mapping to > disable, split, reduce mapped users or offer better isolation they > can't... no way to make private inodes inside containers if they > share the final fsuid, inside container mapping is ignored... > ... OK, I think there's a misunderstanding about how credential overrides work. They're not permanent changes to the credentials, they're temporary ones to get stuff done within the kernel at a temporary privilege. You can make credentials permanent if you go through prepare_creds()/commit_creds(), but for making them temporary you do prepare_creds()/override_creds() and then revert_creds() once you're done using them. If you want to see a current use of this, try fs/open.c:faccessat. What it's doing is temporarily overriding fsuid with the real uid to check the permissions before reverting the credentials and returning to the user. James > > the privileged ids down to 100000, but I have a volume which still > > has realids, I can mount that volume using shiftfs with > > uidmap=0:100000:1000 and it will allow this userns to read and > > write the volume through its remapped ids. > > > > > For overlayfs I did write an expriment but for me it's not an > > > overlayfs or another new filesystem problem... we are > > > manipulating UID/GID identities... > > > > > > It would have been better if you did send this as a separate > > > thread. It was a vfs:userns RFC fix which if we continue we turn > > > it into a complicated thing! implement another new light > > > filesystem with userns... (overlayfs...) > > > > > > Will follow up if the appropriate thread is created, not here, I > > > guess it's ok ? > > > > Well, I can resend the patch as a separate thread when I've fixed > > some of the problems viro pointed out. > > > > James > > > > > > James > > > > > > > > > > Thank you for your feedback! > > > > > > > > > > Thanks! >
[toc] | [prev] | [next] | [standalone]
| From | Djalal Harouni <tixxdz@gmail.com> |
|---|---|
| Date | 2016-05-14 12:00 +0200 |
| Message-ID | <ryEwy-6u9-9@gated-at.bofh.it> |
| In reply to | #1400395 |
On Thu, May 12, 2016 at 03:24:12PM -0700, James Bottomley wrote: > On Thu, 2016-05-12 at 20:55 +0100, Djalal Harouni wrote: > > On Wed, May 11, 2016 at 11:33:38AM -0700, James Bottomley wrote: > > > On Wed, 2016-05-11 at 17:42 +0100, Djalal Harouni wrote: [...] > > > > > > The credentials are per thread, so it's a standard way of doing > > > credential shifting and no other threads of execution in the same > > > task get access. As long as you bound the override_creds()/revert_c > > > reds() pairs within the kernel, you're safe. > > > > No, and here sorry I mean shifted. > > > > current_fsuid() is global through all fs operations which means it > > crosses user namespaces... it was safe the days of only init_user_ns, > > not anymore... You give a mapping inside containers to fsuid where > > they don't want to have it... this allows to operate on inodes inside > > other containers... update current_fsuid() even if we want that user > > to be nobody inside the container... and later it can access the > > inodes of the shifted fs... and by same current of course... > > OK, I still don't understand what you're getting at. There are three > per-thread uids: uid, euid and fsuid (real, effective and filesystem). > They're all either settable via syscall or inherited on fork. They're > all kernel side, meaning they're kuid_t. Their values stay invariant > as you move through namespaces. They change (and get mapped according > to the current user namespace setting) when you call set[fe]uid() So > when I enter a user namespace with mapping > > 0 100000 1000 > > and call setuid(0) (which sets all three). they all pick up the kuid_t > of 100000. This means that writing a file inside the user namespace > after calling setuid(0) appears as real uid 100000 on the medium even > though if I call getuid() from the namespace, I get back 0. What > shiftfs does is hijack temporarily the kernel fsuid/fsgid for > permission checks, so you can remap to any old uid on the medium > (although usually you'd pass in uidmap=0:100000:1000") it maps back > from kuid_t 100000 to kuid_t 0, which is why the container can now read > and write the underlying medium at on-media id 0 even through root > inside the container has kuid_t 100000. There's no permanent change of > fsuid and it stays at its invariant value for the thread except as a > temporary measure to do the permission checks on the underlying of the > shifted filesystem. > > > > > The worst thing is that current_fsuid() does not follow now the > > > > /proc/self/uid_map interface! this is a serious vulnerability and > > > > a mix of the current semantics... it's updated but using other > > > > rules...? > > > > > > current_fsuid() is aready mapped via the userns; it's already a > > > kuid_t at its final value. Shifting that is what you want to remap > > > underlying volume uid/gid's. The uidmap/gidmap inputs to this are > > > shifts on the final underlying uid/gids. > > > > => some points: > > Changing setfsuid() its interfaces and rules... or an indrect way to > > break another syscall... > > There is no change to setfsuid(). > > > The userns used for *mapping* is totatly different and not standard.. > > . losing "init_user_ns and its decendents userns *semantics*...", a > > yet a totatly unlinked mapping... > > There is no user namespace mapping at all. This is a simple shift, > kernel side, of uids and gids at their kuid_t values. > > > Breaking current_uid(),current_euid(),current_fsuid() which are > > mapped but in *different* user namespaces... hence different values > > inside namespaces... you can change your userns mapping but that > > current_fsuid specific one will always be remapped to some other > > value inside even if you don't want it... It crosses user > > namespaces... uid and euid are remapped according to /proc/self/uid_ > > map, fsuid is remapped according to this new interface... > > > > Hard coding the mapping, nested containers/apps may *share* fsuid and > > can't get rid of it even if they change the inside userns mapping to > > disable, split, reduce mapped users or offer better isolation they > > can't... no way to make private inodes inside containers if they > > share the final fsuid, inside container mapping is ignored... > > ... > > OK, I think there's a misunderstanding about how credential overrides > work. They're not permanent changes to the credentials, they're > temporary ones to get stuff done within the kernel at a temporary > privilege. You can make credentials permanent if you go through > prepare_creds()/commit_creds(), but for making them temporary you do > prepare_creds()/override_creds() and then revert_creds() once you're > done using them. > > If you want to see a current use of this, try fs/open.c:faccessat. > What it's doing is temporarily overriding fsuid with the real uid to > check the permissions before reverting the credentials and returning to > the user. Thank you for explaining things, but I think you should take the time to read this RFC and understand some problems. This is a quick dump of some problems that it avoids...: In this series we don't hijack setfsuid() in an indirect way, setfsuid maps UIDs into current userns according to rules set by parent. Changing current_fsuid() to some other mapping is a way to allow processes to bypass that and use it to access other inodes... This should not change and fsuid should continue to follow these rules... A cred->fsuid solution is safe or used to be safe only inside init_user_ns where there is always a mapping or in context of current user namespace. In an other user namespace with 0:1000:1 mapping, you can't set it to arbitrary mapping like 0:4000:1... It will give confined processes access to inodes that satisfy the kuid_t 4000 mapping and which the app/container wants to deny, they only want 0:1000:1. .. We don't cross user namespaces, we don't use different mappings for cred->uid, cred->fsuid... A clean solution is to shift inodes UID/GID and not change fsuid to cross namespaces. Not to mention how it may interact with capabilities... We follow user namespace rules and we keep "the parent defines a range that the children can't escape" semantics. There is a clear relation between user namespaces that should not be broken. We explicitly don't define a new user namespace mapping nor add a new interface for the simple reason it's: *too complicated*. We can do that, but no thanks! May be in future if there is a real need or things are clear... The current user namespace interface is getting standard and stable, so we just keep it that way and make it consistant inside VFS. We give VFS control of that, and we make mount namespaces the central part of this whole logic. We make admins life easier where they can pull container images, root filesystems from containers/apps hubs... verify the signature and start them with different mappings according to host resources... We don't want them to do anything. The design was planned to make it easier for users, it should work out of the box, and it can be used to handle complex stuff too, since it's flexible. Able to support most filesystems including on-disk filesystems natively. Able to support disk quota according to the shifted UID/GID on-disk values. Especially during inode creation... Able to support ACL if requested. The user namespace mapping is kept a runtime configure option, we don't pin a special mapping at any time, and of course parent creator of user namespace is the one that can manipulate it, at the same time the mapping is restricted according to grandpa rules and so on... It allows unprivileged to use the VFS UID/GID shift without the intervention of a privileged process each time. The real privileged process sets the filesystem and the mount namespace the first time, then it should work for all nested namespaces and containers. It does not need the intervation of init_user_ns root to set the mapping and make it work, you don't have to go in and go out to setup the thing, etc. We don't do this on behalf of filesystems, they should explicitly support it. procfs and other host resource virtual filesystems are safe and currently they don't need shifting. We try to fix the problem where it should be fixed, and not hide it... -- Djalal Harouni http://opendz.org
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2016-05-14 15:50 +0200 |
| Message-ID | <ryI77-1uw-1@gated-at.bofh.it> |
| In reply to | #1401079 |
On Sat, 2016-05-14 at 10:53 +0100, Djalal Harouni wrote: > On Thu, May 12, 2016 at 03:24:12PM -0700, James Bottomley wrote: > > On Thu, 2016-05-12 at 20:55 +0100, Djalal Harouni wrote: > > > On Wed, May 11, 2016 at 11:33:38AM -0700, James Bottomley wrote: > > > > On Wed, 2016-05-11 at 17:42 +0100, Djalal Harouni wrote: > [...] > > > > > > > > > The credentials are per thread, so it's a standard way of doing > > > > credential shifting and no other threads of execution in the > > > > same > > > > task get access. As long as you bound the > > > > override_creds()/revert_c > > > > reds() pairs within the kernel, you're safe. > > > > > > No, and here sorry I mean shifted. > > > > > > current_fsuid() is global through all fs operations which means > > > it > > > crosses user namespaces... it was safe the days of only > > > init_user_ns, > > > not anymore... You give a mapping inside containers to fsuid > > > where > > > they don't want to have it... this allows to operate on inodes > > > inside > > > other containers... update current_fsuid() even if we want that > > > user > > > to be nobody inside the container... and later it can access the > > > inodes of the shifted fs... and by same current of course... > > > > OK, I still don't understand what you're getting at. There are > > three > > per-thread uids: uid, euid and fsuid (real, effective and > > filesystem). > > They're all either settable via syscall or inherited on fork. > > They're > > all kernel side, meaning they're kuid_t. Their values stay > > invariant > > as you move through namespaces. They change (and get mapped > > according > > to the current user namespace setting) when you call set[fe]uid() > > So > > when I enter a user namespace with mapping > > > > 0 100000 1000 > > > > and call setuid(0) (which sets all three). they all pick up the > > kuid_t > > of 100000. This means that writing a file inside the user > > namespace > > after calling setuid(0) appears as real uid 100000 on the medium > > even > > though if I call getuid() from the namespace, I get back 0. What > > shiftfs does is hijack temporarily the kernel fsuid/fsgid for > > permission checks, so you can remap to any old uid on the medium > > (although usually you'd pass in uidmap=0:100000:1000") it maps back > > from kuid_t 100000 to kuid_t 0, which is why the container can now > > read > > and write the underlying medium at on-media id 0 even through root > > inside the container has kuid_t 100000. There's no permanent > > change of > > fsuid and it stays at its invariant value for the thread except as > > a > > temporary measure to do the permission checks on the underlying of > > the > > shifted filesystem. > > > > > > > The worst thing is that current_fsuid() does not follow now > > > > > the > > > > > /proc/self/uid_map interface! this is a serious vulnerability > > > > > and > > > > > a mix of the current semantics... it's updated but using > > > > > other > > > > > rules...? > > > > > > > > current_fsuid() is aready mapped via the userns; it's already a > > > > kuid_t at its final value. Shifting that is what you want to > > > > remap > > > > underlying volume uid/gid's. The uidmap/gidmap inputs to this > > > > are > > > > shifts on the final underlying uid/gids. > > > > > > => some points: > > > Changing setfsuid() its interfaces and rules... or an indrect way > > > to > > > break another syscall... > > > > There is no change to setfsuid(). > > > > > The userns used for *mapping* is totatly different and not > > > standard.. > > > . losing "init_user_ns and its decendents userns *semantics*...", > > > a > > > yet a totatly unlinked mapping... > > > > There is no user namespace mapping at all. This is a simple shift, > > kernel side, of uids and gids at their kuid_t values. > > > > > Breaking current_uid(),current_euid(),current_fsuid() which are > > > mapped but in *different* user namespaces... hence different > > > values > > > inside namespaces... you can change your userns mapping but that > > > current_fsuid specific one will always be remapped to some other > > > value inside even if you don't want it... It crosses user > > > namespaces... uid and euid are remapped according to > > > /proc/self/uid_ > > > map, fsuid is remapped according to this new interface... > > > > > > Hard coding the mapping, nested containers/apps may *share* fsuid > > > and > > > can't get rid of it even if they change the inside userns mapping > > > to > > > disable, split, reduce mapped users or offer better isolation > > > they > > > can't... no way to make private inodes inside containers if they > > > share the final fsuid, inside container mapping is ignored... > > > ... > > > > OK, I think there's a misunderstanding about how credential > > overrides > > work. They're not permanent changes to the credentials, they're > > temporary ones to get stuff done within the kernel at a temporary > > privilege. You can make credentials permanent if you go through > > prepare_creds()/commit_creds(), but for making them temporary you > > do > > prepare_creds()/override_creds() and then revert_creds() once > > you're > > done using them. > > > > If you want to see a current use of this, try fs/open.c:faccessat. > > What it's doing is temporarily overriding fsuid with the real uid > > to > > check the permissions before reverting the credentials and > > returning to > > the user. > > Thank you for explaining things, but I think you should take the time > to read this RFC and understand some problems. This is a quick dump > of some problems that it avoids...: I did. The problem is how to get the userns to read and write files at the interior not the exterior id. Your solution is to thread the mapping through the VFS and even on to the filesystems themselves to get the mount option. I already commented that this is a bit ugly and couldn't it be encapsulated in a filesystem. The way I approached the problem is from the base that I do have build container roots with shifted uid/gids because I installed them that way. So, if it already works, one possible solution is to have a filesystem which does the shift and mounts the shifted root somewhere in the mount tree for the namespace to access. The point about doing it this way is that the filesystem that does it needs no user namespace knowledge. All it does is remap from one on disk id to another using a map function. How it gets the map was left up to the admin in the implementation. > In this series we don't hijack setfsuid() in an indirect way, > setfsuid maps UIDs into current userns according to rules set by > parent. Changing current_fsuid() to some other mapping is a way to > allow processes to bypass that and use it to access other inodes... > This should not change and fsuid should continue to follow these > rules... Both solutions do this > A cred->fsuid solution is safe or used to be safe only inside > init_user_ns where there is always a mapping or in context of current > user namespace. In an other user namespace with 0:1000:1 mapping, > you can't set it to arbitrary mapping like 0:4000:1... It will give > confined processes access to inodes that satisfy the kuid_t 4000 > mapping and which the app/container wants to deny, they only want > 0:1000:1. .. OK, so both solutions are safe here too. Your safety comes from only remapping in the userns; mine comes from the normal filesystem acl rules: either the userns for different users all have disjoint ids regulated by /etc/subuidmap or they're all using the same one (like docker 1.10) in either case, you could regulate by having the mount under a directory which is accessible only to the userns owner. > We don't cross user namespaces, we don't use different mappings for > cred->uid, cred->fsuid... A clean solution is to shift inodes > UID/GID and not change fsuid to cross namespaces. Not to mention how > it may interact with capabilities... This is a subjective question on what constitutes "clean". I think we both think the other solution isn't clean, so that's for others to adjudicate. > We follow user namespace rules and we keep "the parent defines a > range that the children can't escape" semantics. There is a clear > relation between user namespaces that should not be broken. OK, so I separated the problem into a userns one, which remaps for the processes in user space, and a vfs one which remaps the on-disk id. However, they could be combined by allowing the userns to mount shiftfs but only on designated filesystems and setting the uidmappings to the same ones as the userns. > We explicitly don't define a new user namespace mapping nor add a new > interface for the simple reason it's: *too complicated*. We can do > that, but no thanks! May be in future if there is a real need or > things are clear... The current user namespace interface is getting > standard and stable, so we just keep it that way and make it > consistant inside VFS. I don't accept the too complicated point. For fully unprivileged containers, the host admin already has to set up the subuid/subgid map files which is most of the complexity. Once that's done, the same maps can be used to shift mount. Once it's all set up, no further intervention is required. > We give VFS control of that, and we make mount namespaces the central > part of this whole logic. Right, that's what causes the logic to thread throughout the entire vfs and into the fs layer. The fundamental point of difference is that I'd like a solution which encapsulates the problem rather than exposing it to the vfs. > We make admins life easier where they can pull container images, root > filesystems from containers/apps hubs... verify the signature and > start them with different mappings according to host resources... We > don't want them to do anything. The design was planned to make it > easier for users, it should work out of the box, and it can be used > to handle complex stuff too, since it's flexible. Either works easily for users. Setting stuff up is always the job of the admin in both solutions. > Able to support most filesystems including on-disk filesystems > natively. Shiftfs does this. More importantly it supports subtrees, so I can unpack an image root on to an existing filesystem and remap it into a container. > Able to support disk quota according to the shifted UID/GID on-disk > values. Especially during inode creation... Quota can be shifted, I just wasn't sure it was necessary. If the usual use case is for unpacked roots, chances are you want the remapping to use the group quota of the userns owner, which they'd get naturally so, while it's possible to remap projid, I didn't think it needed to be done. > Able to support ACL if requested. Both do this. > The user namespace mapping is kept a runtime configure option, we > don't pin a special mapping at any time, and of course parent creator > of user namespace is the one that can manipulate it, at the same time > the mapping is restricted according to grandpa rules and so on... > > It allows unprivileged to use the VFS UID/GID shift without the > intervention of a privileged process each time. The real privileged > process sets the filesystem and the mount namespace the first time, > then it should work for all nested namespaces and containers. It does > not need the intervation of init_user_ns root to set the mapping and > make it work, you don't have to go in and go out to setup the thing, > etc. Both solutions work like this. When I use this for shifted roots of emulation containers, it's set up once at start of day. I then build the containers unprivileged using newsubuid/newsubgid as I'm using them. Once the shifts are done at start of day, no other admin support is required. James > We don't do this on behalf of filesystems, they should explicitly > support it. procfs and other host resource virtual filesystems are > safe and currently they don't need shifting. > > We try to fix the problem where it should be fixed, and not hide > it...
[toc] | [prev] | [next] | [standalone]
| From | James Bottomley <James.Bottomley@HansenPartnership.com> |
|---|---|
| Date | 2016-05-16 04:10 +0200 |
| Message-ID | <rzg8W-u7-239@gated-at.bofh.it> |
| In reply to | #1401087 |
On Sat, 2016-05-14 at 21:21 -0500, Eric W. Biederman wrote: > James if you could see shiftfs with a different set of merits than > what to Djalal is doing I think that would be useful. As it would > allow everyone to concentrate on getting the bugs out of their > solutions. Just to reply to this specific point. Djalal's patches can't actually work for me because I use subtree based roots rather than whole fs roots ... it's mostly because I work with image directories, not the full mounted images themselves. For stuff I unpack into /home, I could see having /home on a separate directory and adding the vfs_shift_ flags. however, I'm not doing (and it would be really unsafe to do) that for / to get my images that unpack in /var/tmp (like the obs build roots). However, half the ugliness of the patch set is that it needs lower layer FS support because vfs_shift_ are mount flags in the superblock. If they were made subtree flags instead (so MNT_ flags), I think you could eliminate the need to modify any underlying filesystems and they would allow us to mark subtrees for shifting. the mount command would need modifying to add them (like it was for --shared and --private) so we'd need an additional --vfs-shift --ufs-shift to mark the subtree but then the series would work for bind mounting subtrees, which is what I need. And they would work for *any* filesystem without modification. This would probably be the better of both worlds because it will work for the docker case as well. James
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2016-05-16 04:10 +0200 |
| Message-ID | <rzg8X-u7-241@gated-at.bofh.it> |
| In reply to | #1401087 |
James Bottomley <James.Bottomley@HansenPartnership.com> writes: > On Sat, 2016-05-14 at 10:53 +0100, Djalal Harouni wrote: Just a couple of quick comments from a very high level design point. - I think a shiftfs is valuable in the same way that overlayfs is valuable. Esepcially in the Docker case where a lot of containers want a shared base image (for efficiency), but it is desirable to run those containers in different user namespaces for safety. - It is also the plan to make it possible to mount a filesystem where the uids and gids of that filesystem on disk do not have a one to one mapping to kernel uids and gids. 99% of the work has already be done, for all filesystem except XFS. That said there are some significant issues to work through, before something like that can be enabled. * Handling of uids/gids on disk that don't map into a kuid/kgid. * Safety from poisoned filesystem images. I have slowly been working with Seth Forshee on these issues as the last thing I want is to introduce more security bugs right now. Seth being a braver man than I am has already merged his changes into the Ubuntu kernel. Right now we are targeting fuse, because fuse is already designed to handle poisoned filesystem images. So to safely enable this kind of mapping for fuse is not a giant step. The big thing from my point of view is to get the VFS interfaces correct so that the VFS handles all of the weird cases that come up with uids and gids that don't map, and any other weird cases. Keeping the weird bits out of the filesystems. James, Djalal I regert I have not been able to read through either of your patches cloesely yet. From a high level view I believe there are use cases for both approaches, and the use cases do not necessarily overlap. Djalal I think you are seeing the upsides and not the practical dangers of poisoned filesystem images. James I think you are missing the fact that all filesystems already have the make_kuid and make_kgid calls right where the data comes off disk, and the from_kuid and from_kgid calls right where the on-disk data is being created just before it goes on disk. Which means that the actual impact on filesystems of the translation is trivial. Where the actual impact of filesystems is much higher is the infrastructure needed to ensure poisoned filesystem images do not cause a kernel compromise. That extends to the filesystem testing and code review process beyond and is more than just a kernel problem. Hardening that attack surface of the disk side of filesystems is difficult especially when not impacting filesystem performance. So I don't think it makes sense to frame this as an either/or situation. I think there is a need for both solutions. Djalal if you could work with Seth I think that would be very useful. I know I am dragging my heels there but I really hope I can dig in and get everything reviewed and merged soonish. James if you could see shiftfs with a different set of merits than what to Djalal is doing I think that would be useful. As it would allow everyone to concentrate on getting the bugs out of their solutions. That said I am not certain shiftfs makes sense without Seth's patches to handle the weird cases at the VFS level. What do you do with uids and gids that don't map? You can reinvent how to handle the strange cases in shfitfs or we can work on solving this problem at the VFS level so people don't have to go through the error prone work of reinventing solutions. The big ugly nasty in all of this is that we are fundamentally dealing with uids and gids which are security identifiers. Practically any bug is exploitable and CVE worthy. So it make sense to tread very carefully. Even with care it can takes months if not years to get the number of bugs down to a level where you are not the favorite target of people looking for exploitable kernel bugs. Eric
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web