Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1250013 > unrolled thread
| Started by | Alexei Starovoitov <ast@plumgrid.com> |
|---|---|
| First post | 2015-10-18 06:40 +0200 |
| Last post | 2015-10-20 20:50 +0200 |
| Articles | 10 on this page of 30 — 5 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Alexei Starovoitov <ast@plumgrid.com> - 2015-10-18 06:40 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Daniel Borkmann <daniel@iogearbox.net> - 2015-10-18 17:10 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Daniel Borkmann <daniel@iogearbox.net> - 2015-10-18 18:50 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Alexei Starovoitov <ast@plumgrid.com> - 2015-10-18 23:00 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Hannes Frederic Sowa <hannes@stressinduktion.org> - 2015-10-19 09:40 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Daniel Borkmann <daniel@iogearbox.net> - 2015-10-19 12:00 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Daniel Borkmann <daniel@iogearbox.net> - 2015-10-19 16:30 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Alexei Starovoitov <ast@plumgrid.com> - 2015-10-19 18:30 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Daniel Borkmann <daniel@iogearbox.net> - 2015-10-19 19:40 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Alexei Starovoitov <ast@plumgrid.com> - 2015-10-19 20:20 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Hannes Frederic Sowa <hannes@stressinduktion.org> - 2015-10-19 20:50 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Alexei Starovoitov <ast@plumgrid.com> - 2015-10-19 21:40 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Hannes Frederic Sowa <hannes@stressinduktion.org> - 2015-10-19 22:10 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Alexei Starovoitov <ast@plumgrid.com> - 2015-10-19 22:50 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Daniel Borkmann <daniel@iogearbox.net> - 2015-10-20 00:20 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Alexei Starovoitov <ast@plumgrid.com> - 2015-10-20 02:40 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Daniel Borkmann <daniel@iogearbox.net> - 2015-10-20 10:50 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Alexei Starovoitov <ast@plumgrid.com> - 2015-10-20 20:00 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs ebiederm@xmission.com (Eric W. Biederman) - 2015-10-20 21:10 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Daniel Borkmann <daniel@iogearbox.net> - 2015-10-21 17:20 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Thomas Graf <tgraf@suug.ch> - 2015-10-21 20:40 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Alexei Starovoitov <ast@plumgrid.com> - 2015-10-22 00:50 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Daniel Borkmann <daniel@iogearbox.net> - 2015-10-22 15:30 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs ebiederm@xmission.com (Eric W. Biederman) - 2015-10-22 21:50 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Daniel Borkmann <daniel@iogearbox.net> - 2015-10-23 15:50 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Hannes Frederic Sowa <hannes@stressinduktion.org> - 2015-10-20 11:50 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Hannes Frederic Sowa <hannes@stressinduktion.org> - 2015-10-20 01:10 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Alexei Starovoitov <ast@plumgrid.com> - 2015-10-20 03:10 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Hannes Frederic Sowa <hannes@stressinduktion.org> - 2015-10-20 12:10 +0200
Re: [PATCH net-next 3/4] bpf: add support for persistent maps/progs Alexei Starovoitov <ast@plumgrid.com> - 2015-10-20 20:50 +0200
Page 2 of 2 — ← Prev page 1 [2]
| From | Thomas Graf <tgraf@suug.ch> |
|---|---|
| Date | 2015-10-21 20:40 +0200 |
| Message-ID | <qm6sN-2SJ-3@gated-at.bofh.it> |
| In reply to | #1252939 |
On 10/21/15 at 05:17pm, Daniel Borkmann wrote:
> On 10/20/2015 08:56 PM, Eric W. Biederman wrote:
> ...
> >Just FYI: Using a device for this kind of interface is pretty
> >much a non-starter as that quickly gets you into situations where
> >things do not work in containers. If someone gets a version of device
> >namespaces past GregKH it might be up for discussion to use character
> >devices.
>
> Okay, you are referring to this discussion here:
>
> http://thread.gmane.org/gmane.linux.kernel.containers/26760
>
> What had been mentioned earlier in this thread was to have a namespace
> pass-through facility enforced by device cgroups we have in the kernel,
> which is one out of various means used to enforce policy today by
> deployment systems such as docker, for example. But more below.
>
> I think this all depends on the kind of expectations we have, where all
> this is going. In the original proposal, it was agreed to have the
> operation that creates a node as 'capable(CAP_SYS_ADMIN)'-only (in the
> way like most of the rest of eBPF is restricted), and based on the use
> case we distribute such objects to unprivileged applications. But I
> understand that it seems the trend lately to lift eBPF restrictions at
> some point anyway, and thus the CAP_SYS_ADMIN is suddenly irrelevant
> again. Fair enough.
>
> Don't get me wrong, I really don't mind if it will be some version of
> this fs patch or whatever architecture else we find consensus on, I
> think this discussion is merely trying to evaluate/discuss on what seems
> to be a good fit, also in terms of future requirements and integration.
>
> So far, during this discussion, it was proposed to modify the file system
> to a single-mount one and to stick this under /sys/kernel/bpf/. This
> will not have "real" namespace support either, but it was proposed to
> have a following structure:
>
> /sys/kernel/bpf/username/<optional_dirs_mkdir_by_user>/progX
This would probably work as you would typically map the ebpf map
using -v like this to give a stable path:
docker run -v /sys/kernel/bpf/foo/maps/progX:/map proX
> So, the file system will have kind of a user home-directory for each user
> to isolate through permissions, if I understood correctly.
>
> If we really want to go this route, then I think there are no big stones
> in the way for the other model either. It should look roughly drafted like
> the below.
>
> Together with device cgroups for containers, it would allow scenarios where
> you can have:
>
> * eBPF (map/prog) device pass-through so a map/prog could even be shared out
> from the initial namespace into individual ones/all (one could possibly
> extend such maps as read-only for these consumers).
> * eBPF device creation for unprivileged users with permissions being set
> accordingly (as in fs case).
> * Since cgroup controller can also do wildcards on major/minors, we could
> make that further fine-grained.
> * eBPF device creation can also be enforced by the cgroup controller to be
> entirely disallowed for a specific container.
>
> (An admin can determine the dynamically created major f.e. under /proc/devices.)
I've read the discussion passively and my take away is that, frankly,
I think the differences are somewhat minor. Both architectures can
scale to what we need. Both will do the job. I'm slightly worried about
exposing uAPI as a FS, I think that didn't work too well for sysfs. It's
pretty much a define the format once and never touch it again kind of
deal.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alexei Starovoitov <ast@plumgrid.com> |
|---|---|
| Date | 2015-10-22 00:50 +0200 |
| Message-ID | <qmamK-8tQ-9@gated-at.bofh.it> |
| In reply to | #1253115 |
On 10/21/15 11:34 AM, Thomas Graf wrote: >> So far, during this discussion, it was proposed to modify the file system >> >to a single-mount one and to stick this under/sys/kernel/bpf/. This >> >will not have "real" namespace support either, but it was proposed to >> >have a following structure: >> > >> > /sys/kernel/bpf/username/<optional_dirs_mkdir_by_user>/progX > This would probably work as you would typically map the ebpf map > using -v like this to give a stable path: > > docker run -v /sys/kernel/bpf/foo/maps/progX:/map proX yep tracefs works inside docker the same way. May be we should let users pick names similar to this fs patch to make the above easier to use. Also from bpf syscall point of the user shouldn't see /sys/kernel/bpf/user/ prefix. Only 'optional_dirs_mkdir_by_user/name' when doing pin/new_fd. May be prog type should be a fixed part of the path as well. >> >Together with device cgroups for containers, it would allow scenarios where >> >you can have: >> > >> > * eBPF (map/prog) device pass-through so a map/prog could even be shared out >> > from the initial namespace into individual ones/all (one could possibly >> > extend such maps as read-only for these consumers). >> > * eBPF device creation for unprivileged users with permissions being set >> > accordingly (as in fs case). >> > * Since cgroup controller can also do wildcards on major/minors, we could >> > make that further fine-grained. >> > * eBPF device creation can also be enforced by the cgroup controller to be >> > entirely disallowed for a specific container. none of the above is practical. It can be demoed in a canned environment, but it's a complete mismatch of apis. cgroup/dev is a static config, whereas bpf-cdev is dynamic (with minors out of idr for all users) When you have to hack drivers/base/core.c to get there it should have been a warning sign that something is wrong with this cdev approach. > I've read the discussion passively and my take away is that, frankly, > I think the differences are somewhat minor. Both architectures can > scale to what we need. Both will do the job. I'm slightly worried about > exposing uAPI as a FS, I think that didn't work too well for sysfs. It's > pretty much a define the format once and never touch it again kind of > deal. It's even worse in cdev style since it piggy backs on sysfs. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Daniel Borkmann <daniel@iogearbox.net> |
|---|---|
| Date | 2015-10-22 15:30 +0200 |
| Message-ID | <qmo6m-3GO-7@gated-at.bofh.it> |
| In reply to | #1253322 |
On 10/22/2015 12:44 AM, Alexei Starovoitov wrote: ... > all users) When you have to hack drivers/base/core.c to get there it > should have been a warning sign that something is wrong with > this cdev approach. Hmm, you know, this had nothing to do with it, merely to save ~20 LoC that I can do just as well inside BPF framework. No changes in driver API needed. >> I've read the discussion passively and my take away is that, frankly, >> I think the differences are somewhat minor. Both architectures can >> scale to what we need. Both will do the job. I'm slightly worried about >> exposing uAPI as a FS, I think that didn't work too well for sysfs. It's >> pretty much a define the format once and never touch it again kind of >> deal. > > It's even worse in cdev style since it piggy backs on sysfs. I don't mind with what approach we're going in the end, but this kind of discussion is really tiring, and not going anywhere. Lets just make a beer call, so we can hash out a way forward that works for everyone. On that note: cheers! ;) Daniel -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | ebiederm@xmission.com (Eric W. Biederman) |
|---|---|
| Date | 2015-10-22 21:50 +0200 |
| Message-ID | <qmu25-3Sv-13@gated-at.bofh.it> |
| In reply to | #1252939 |
Daniel Borkmann <daniel@iogearbox.net> writes:
> On 10/20/2015 08:56 PM, Eric W. Biederman wrote:
> ...
>> Just FYI: Using a device for this kind of interface is pretty
>> much a non-starter as that quickly gets you into situations where
>> things do not work in containers. If someone gets a version of device
>> namespaces past GregKH it might be up for discussion to use character
>> devices.
>
> Okay, you are referring to this discussion here:
>
> http://thread.gmane.org/gmane.linux.kernel.containers/26760
That is a piece of it. It is an old old discussion (which generally has
been handled poorly). For the forseeable future device namespaces have
a firm NACK by GregKH. Which means that dynamic character device based
interfaces do not work in containers. Which means if you are not
talking about physical hardware, character devices are a poor fit.
Making a character based interface for eBPF not workable.
Eric
p.s. There are plenty of reasons (even if privilege remains a
requirement) to ask how can this functionality be used in a
container. If for no other reason than sandboxing privileged
applications is typically a good idea.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Daniel Borkmann <daniel@iogearbox.net> |
|---|---|
| Date | 2015-10-23 15:50 +0200 |
| Message-ID | <qmKTf-32t-1@gated-at.bofh.it> |
| In reply to | #1254136 |
On 10/22/2015 09:35 PM, Eric W. Biederman wrote: > Daniel Borkmann <daniel@iogearbox.net> writes: >> On 10/20/2015 08:56 PM, Eric W. Biederman wrote: >> ... >>> Just FYI: Using a device for this kind of interface is pretty >>> much a non-starter as that quickly gets you into situations where >>> things do not work in containers. If someone gets a version of device >>> namespaces past GregKH it might be up for discussion to use character >>> devices. >> >> Okay, you are referring to this discussion here: >> >> http://thread.gmane.org/gmane.linux.kernel.containers/26760 > > That is a piece of it. It is an old old discussion (which generally has > been handled poorly). For the forseeable future device namespaces have > a firm NACK by GregKH. Which means that dynamic character device based > interfaces do not work in containers. Which means if you are not > talking about physical hardware, character devices are a poor fit. Yes, it breaks down with real namespace support. Reworking the set with an improved version of the fs code is already in progress. Thanks, Daniel -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Hannes Frederic Sowa <hannes@stressinduktion.org> |
|---|---|
| Date | 2015-10-20 11:50 +0200 |
| Message-ID | <qlBIn-8jI-17@gated-at.bofh.it> |
| In reply to | #1251172 |
Hey Alexei, On Tue, Oct 20, 2015, at 02:30, Alexei Starovoitov wrote: > On 10/19/15 3:17 PM, Daniel Borkmann wrote: > > On 10/19/2015 10:48 PM, Alexei Starovoitov wrote: > >> On 10/19/15 1:03 PM, Hannes Frederic Sowa wrote: > >>> > >>> I doubt it will stay a lightweight feature as it should not be in the > >>> responsibility of user space to provide those debug facilities. > >> > >> It feels we're talking past each other. > >> I want to solve 'persistent map' problem. > >> debugging of maps/progs, hierarchy, etc are all nice to have, > >> but different issues. > > > > Ok, so you are saying that this file system will have *only* regular > > files that are to be considered nodes for eBPF maps and progs. Nothing > > else will ever be added to it? Yet, eBPF map and prog nodes are *not* > > *regular files* to me, this seems odd. > > > > As soon as you are starting to add additional folders that contain > > files dumping additional meta data, etc, you basically end up on a > > bit of similar concept to sysfs, no? > > as we discussed in this thread and earlier during plumbers I think > it would be good to expose key/values somehow in this fs. > 'how' is a big question. > But regardless which path we take, sysfs is too rigid. > For the sake of argument say we do every key as a new file in bpffs. > It's not very scalable, but comparing to sysfs it's better > (resource wise). > If we decide to add bpf syscall command to expose map details we > can provide pretty printer to it in a form of printk-like string > or via some schema passed to syscall, so that keys(file names) can > look properly in bpffs. > The above and other ideas all possible in bpffs, but not possible > in sysfs. If you come up with a layout (albeit I don't understand how to enforce it, later on this more), it is kind of an uapi. It is not possible to just change the layout of the filesystem in every kernel version. Take f.e. bpffs like Daniel proposed it, how would you extend it to a key/value store like of map. Really, my IPv6 addresses very often have '\0' inside them, I don't see a way, if they are part of a key, how to represent them in a filename. That is forbidden by all means. Adding pretty printer for such a filesystem is something user space should do. How do you want to handle that? Make printk available from ebpf programs and ebpf programs bring their own pretty-printer for their key and values? This really looks like a security hazard. User space is so much easier and fuse, if you want to have a key/value representation of your map, why not in user space? In any way, in case a key/value filesystem is needed I certainly want to go with Eric Biederman and have one mount point per map. Otherwise I also can't see how this should work in terms of permissions. > >> In case of persistent maps I imagine unprivileged process would want > >> to use it eventually as well, so this requirement already kills cdev > >> approach for me, since I don't think we ever let unprivileged apps > >> create cdev with syscall. > > > > Hmm, I see. So far the discussion was only about having this for privileged > > users (also in this fs patch). F.e. privileged system daemons could setup > > and distribute progs/maps to consumers, etc (f.e. seccomp and tc case). > > It completely makes sense to restrict it to admin today, but design > should not prevent relaxing it in the future. Even today lot's of unprivileged devices are used by users from day to day. Soundcards, terminals etc. This is also possible with cdevs. > > When we start lifting this, eBPF maps by its own will become a real kernel > > IPC facility for unprivileged Linux applications (independently whether > > they are connected to an actual eBPF program). Those kernel IPC facilities > > that are anchored in the file system like named pipes and Unix domain > > sockets > > are indicated as such as special files, no? > > not everything in unix is a model that should be followed. > af_unix with name[0]!=0 is a bad api that wasn't thought through. > Thankfully Linux improved it with abstract names that don't use > special files. > bpf maps obviously is not an IPC (either pinned or not). Sure, it is, IPC between ebpf programs and some kind of user space controlling application. > >> sure, then we can force all bpffs to have the same hierarchy and mounted > >> in /sys/kernel/bpf location. That would be the same. > > > > That would imply to have a mount_single() file system (like f.e. tracefs > > and > > securityfs), right? > > Probably. I'm not sure whether it should be single fs or we allow > multiple mount points. There are pro and con for both. Multiple mount points actually seem dangerous to me. Especially if fds get pinned into multiple of those filesystems. > > So you'd loose having various mounts in different namespaces. And if you > > allow various mount points, how would that /facilitate/ to an admin to > > identify all eBPF objects/resources currently present in the system? > > if it's single mount point there are no issues, but would be nice > to separate users and namespaces somehow. > > > Or to an application developer finding possible mount points for his own > > application so that bpf(2) syscall to create these nodes succeeds? Would > > you make mounting also unprivileged? What if various distros have these > > mount points at different locations? What should unprivileged applications > > do to know that they can use these locations for themselves? > > mounting is root only of course and having standard location answers > all of these questions. Then we miss the separation of users and namespaces we were talking about above. This also needs some way of governing user space application, I don't believe users will conform to the "you should first add a directory for your user then for your program"-convention. cgroups layout kind of got first "standardized" by systemd. > > Also, since they are only regular files, one can only try and find these > > objects based on their naming schemes, which seems to get a bit odd in case > > this file system also carries (perhaps future?) other regular files that > > are not eBPF map and program-special. > > not sure what you meant. If names are given by kernel, there are no > problem finding them. If by user, we'd file attributes like the way > you did with xattr. > > >> It feels you're pushing for cdev only because of that potential > >> debugging need. Did you actually face that need? I didn't and > >> don't like to add 'nice to have' feature until real need comes. > > > > I think this discussion arose, because the question of how flexible we are > > in future to extend this facility. Nothing more. > > Exactly and cdev style pushes us into the corner of traditional > cdev with ioctl which I don't think is flexible enough. But we don't need ioctls! ;) Bye, Hannes -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Hannes Frederic Sowa <hannes@stressinduktion.org> |
|---|---|
| Date | 2015-10-20 01:10 +0200 |
| Message-ID | <qlrJ0-2d8-13@gated-at.bofh.it> |
| In reply to | #1251062 |
Hello Alexei, On Mon, Oct 19, 2015, at 22:48, Alexei Starovoitov wrote: > On 10/19/15 1:03 PM, Hannes Frederic Sowa wrote: > > > > I doubt it will stay a lightweight feature as it should not be in the > > responsibility of user space to provide those debug facilities. > > It feels we're talking past each other. > I want to solve 'persistent map' problem. > debugging of maps/progs, hierarchy, etc are all nice to have, > but different issues. I understand that. Main problem is to persist fds with attached maps, sure. I am not a big fan of the kernel referencing or persisting resources on behalf of its own. IMHO they should be attached to user space programs. So I am still in favor for user space, but I see that most people are not comfortable with that. So a way to make this kind of persistence in the kernel as introspectable as possible is the main goal of the idea we are sharing here. I bet commercial software will make use of this ebpf framework, too. And the kernel always helped me and gave me a way to see what is going on, debug which part of my operating system universe interacts with which other part. Merely dropping file descriptors with data attached to them in an filesystem seems not to fulfill my need at all. I would love to see where resources are referenced and why, like I am nowadays. Btw.: has anybody had a look at kdbus if it allows user space to much more easily handle those file descriptors (as an alternative to af_unix). I haven't, yet. > In case of persistent maps I imagine unprivileged process would want > to use it eventually as well, so this requirement already kills cdev > approach for me, since I don't think we ever let unprivileged apps > create cdev with syscall. TTY code is creating nodes on behalf of users, I check if this could work for bpf cdevs as well. > > The bpf syscall is still used to create the pseudo nodes. If they should > > be persistent they just get registered in the sysfs class hierarchy. > > nope. they should not. sysfs is debugging/tunning facility. > There is absolutely no need for bpf to plug into sysfs. > > >> Doing 'resource stats' via sysfs requires bpf to add to sysfs, which > >> is not this cdev approach. > > > > This is not yet part of the patch, but I think this would be added. > > Daniel? > > please don't. I'm strongly against adding unnecessary bloat. > > > I don't think there are broad differences. But in case a namespaces uses > > huge number of maps with tons of data, the admin in the initial > > namespace might want to debug that without searching all mountpoints and > > find dependencies between processes etc. IMHO sysfs approach can be > > better extended here. > > sure, then we can force all bpffs to have the same hierarchy and mounted > in /sys/kernel/bpf location. That would be the same. > > It feels you're pushing for cdev only because of that potential > debugging need. Did you actually face that need? I didn't and > don't like to add 'nice to have' feature until real need comes. Given that we want to monitor the load of a hashmap for graphing purposes. Or liberate some hashmaps from its restriction on number of keys and make upper bounds configurable by admins who know the dimensions of their systems and not some software deep down buried in the bpf syscall where I might not have access to source code. In tc force e.g. hashmaps to do garbage collection because we cannot be sure that under DoS attacks user space clean up gets scheduled early enough if ebpf adds flows to hashtables. I do see need to expand and implement some kind of policy in the future. > >> Also I don't buy the point of reinventing sysfs. bpffs is not doing > >> sysfs. I don't want to see _every_ bpf object in sysfs. It's way too > >> much overhead. Classic doesn't have sysfs and everyone have been > >> using it just fine. > > > > But classic bpf does not have persistence for maps and data. ;) There is > > a 1:1 relationship between socket and bpf_prog for example. > > single task in seccomp can have a chain of bpf progs, so hierarchy > is already there. And it would be great to inspect them. > > But how can the filesystem be extended in terms of tunables and > > information? File attributes? Wouldn't it need the same infrastructure > > otherwise as sysfs? Some third-party lookup filesystem or ioctl? This > > char dev approach also pins maps and progs while giving more policy in > > hand of central user space programs we are currently using (udev, > > systemd, whatever, etc.). > > tunables for bpf maps? There are no such things today. > I think you're implying that we can add rhashtable type of map, so > admin can tune thresholds ? Ouch. I think if we add it, its parameters > will be specified by the user that is creating the map only. There will > be no tunables exposed to sysfs and there should be no way of creating > maps via sysfs. I am fine with creating maps only by bpf syscall. But to hide configuration details or at least not be really able to query them easily seems odd to me. If we go with the ebpffs how could those attributes be added? Thanks, Hannes -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alexei Starovoitov <ast@plumgrid.com> |
|---|---|
| Date | 2015-10-20 03:10 +0200 |
| Message-ID | <qltB9-4Yz-41@gated-at.bofh.it> |
| In reply to | #1251148 |
On 10/19/15 4:02 PM, Hannes Frederic Sowa wrote: > I bet commercial software will make use of this ebpf framework, too. And > the kernel always helped me and gave me a way to see what is going on, > debug which part of my operating system universe interacts with which > other part. Merely dropping file descriptors with data attached to them > in an filesystem seems not to fulfill my need at all. I would love to > see where resources are referenced and why, like I am nowadays. agree. common fs with hierarchy will give this visibility in one place. >> >It feels you're pushing for cdev only because of that potential >> >debugging need. Did you actually face that need? I didn't and >> >don't like to add 'nice to have' feature until real need comes. > Given that we want to monitor the load of a hashmap for graphing > purposes. Or liberate some hashmaps from its restriction on number of > keys and make upper bounds configurable by admins who know the > dimensions of their systems and not some software deep down buried in > the bpf syscall where I might not have access to source code. In tc > force e.g. hashmaps to do garbage collection because we cannot be sure > that under DoS attacks user space clean up gets scheduled early enough > if ebpf adds flows to hashtables. I do see need to expand and implement > some kind of policy in the future. disagree here. admin should not interfere with map parameters. What you proposing above sounds very very dangerous. Admins to configure GC of maps? What do you think the programs will do with such sophisticated maps? What kind of networking app you have in mind? Anyway that's a bit off-topic. I'm very curious though. >> >single task in seccomp can have a chain of bpf progs, so hierarchy >> >is already there. > And it would be great to inspect them. again let's not mix criu and lsof-like requirements with 'pin fd'. For visibility of normal maps we can add fdinfo and lsof can pick it up without any fs or any cdevs. > I am fine with creating maps only by bpf syscall. But to hide > configuration details or at least not be really able to query them > easily seems odd to me. If we go with the ebpffs how could those > attributes be added? I'm not advocating to hide details. Most of the time maps will not be pinned, so fdinfo seems the easiest way to show things like key_size, value_size, max_entries, type. Even if we decide to do it some other way, it's not related to 'pin fd' discussion, since debugging/visibility is nice to have for all bpf objects. Note that walking of key/value without pretty-printers provided by the app is meaningless for admin, so only things like 'how much memory this map is using' are useful. May be we should try to draft the hierarchy of this common fs. How about: /sys/kernel/bpf/username/optional_dirs_mkdir_by_user/progX and 'cat' of it will print the same as fdinfo for normal maps, so admin can see what maps were pinned by user and its cost. Inside 'fdinfo' output we can provide pointers to which progs are using which maps as # cat /sys/kernel/bpf/.../mapX key_size: 4 used_by: /proc/xxx/fd/5 # cat /sys/kernel/bpf/.../progY type: socket using: /proc/xxx/fd/6 using: /sys/kernel/bpf/.../mapZ and similar for cat /proc/xxx/fdinfo/6 but showing hierarchy as directories is non starter, since it's no a tree. All of these would be nice, but doesn't have to be implemented along with 'pin fd' feature. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Hannes Frederic Sowa <hannes@stressinduktion.org> |
|---|---|
| Date | 2015-10-20 12:10 +0200 |
| Message-ID | <qlC1J-uk-23@gated-at.bofh.it> |
| In reply to | #1251248 |
Hello Alexei, On Tue, Oct 20, 2015, at 03:09, Alexei Starovoitov wrote: > On 10/19/15 4:02 PM, Hannes Frederic Sowa wrote: > > I bet commercial software will make use of this ebpf framework, too. And > > the kernel always helped me and gave me a way to see what is going on, > > debug which part of my operating system universe interacts with which > > other part. Merely dropping file descriptors with data attached to them > > in an filesystem seems not to fulfill my need at all. I would love to > > see where resources are referenced and why, like I am nowadays. > > agree. common fs with hierarchy will give this visibility in > one place. > > >> >It feels you're pushing for cdev only because of that potential > >> >debugging need. Did you actually face that need? I didn't and > >> >don't like to add 'nice to have' feature until real need comes. > > Given that we want to monitor the load of a hashmap for graphing > > purposes. Or liberate some hashmaps from its restriction on number of > > keys and make upper bounds configurable by admins who know the > > dimensions of their systems and not some software deep down buried in > > the bpf syscall where I might not have access to source code. In tc > > force e.g. hashmaps to do garbage collection because we cannot be sure > > that under DoS attacks user space clean up gets scheduled early enough > > if ebpf adds flows to hashtables. I do see need to expand and implement > > some kind of policy in the future. > > disagree here. admin should not interfere with map parameters. > What you proposing above sounds very very dangerous. > Admins to configure GC of maps? What do you think the programs will do > with such sophisticated maps? What kind of networking app you have > in mind? Anyway that's a bit off-topic. I'm very curious though. <off-topic> Just a pretty obvious idea is accurate sampling of flows. </off-topic> > >> >single task in seccomp can have a chain of bpf progs, so hierarchy > >> >is already there. > > And it would be great to inspect them. > > again let's not mix criu and lsof-like requirements with 'pin fd'. > For visibility of normal maps we can add fdinfo and lsof > can pick it up without any fs or any cdevs. fdinfo tells me where my position in a file is and which locks the file have? Nothing like that is supposed to work on bpf file descriptors, because they are kind of special. A new hierarchy has to be installed alongside fdinfo/. > > I am fine with creating maps only by bpf syscall. But to hide > > configuration details or at least not be really able to query them > > easily seems odd to me. If we go with the ebpffs how could those > > attributes be added? > > I'm not advocating to hide details. Most of the time maps will not be > pinned, so fdinfo seems the easiest way to show things like key_size, > value_size, max_entries, type. This is an argument in favor of the "fdinfo-like" approach. So far, if someone wants to delve into the details of a map my approach would be to take the file descriptor and make it persistence. I have to think about that some more. > Even if we decide to do it some other way, it's not related to 'pin fd' > discussion, since debugging/visibility is nice to have for all bpf > objects. Note that walking of key/value without pretty-printers > provided by the app is meaningless for admin, so only things > like 'how much memory this map is using' are useful. Yes, absolutely and I am absolutely against pretty printing key values in kernel domain. > May be we should try to draft the hierarchy of this common fs. > How about: > /sys/kernel/bpf/username/optional_dirs_mkdir_by_user/progX > and 'cat' of it will print the same as fdinfo for normal maps, > so admin can see what maps were pinned by user and its cost. So cat-ing them will produce text output with some details about the map? This is what I wanted to avoid. The concept with symlinks and small files seems much cleaner and nicer to me. Also you cannot add writable attributes to this filesystem or you overload stuff heavily? > Inside 'fdinfo' output we can provide pointers to which progs > are using which maps as > # cat /sys/kernel/bpf/.../mapX > key_size: 4 > used_by: /proc/xxx/fd/5 > # cat /sys/kernel/bpf/.../progY > type: socket > using: /proc/xxx/fd/6 > using: /sys/kernel/bpf/.../mapZ > and similar for cat /proc/xxx/fdinfo/6 > but showing hierarchy as directories is non starter, since > it's no a tree. It is not a tree but a graph, sure, that's why sysfs allows to break the cyclic dependencies and create symlinks (see holders/ directories). ;) And if you implement the same set of features IMHO you basically re-implement sysfs. In the beginning we just expose the basic maps and there won't be any features in sysfs, but it will be cheap to have read/write flags on maps etc. etc. (I don't know what people will come up with, yet.). In my opinion those are clearly attributes of a map and should be defined and managed alongside with their holders. > All of these would be nice, but doesn't have to be implemented > along with 'pin fd' feature. The pinfd feature will provide the future infrastructure alongside to make this usable, so I think it is worth spending time to think about it. Thanks, Hannes -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alexei Starovoitov <ast@plumgrid.com> |
|---|---|
| Date | 2015-10-20 20:50 +0200 |
| Message-ID | <qlK8W-3Ji-11@gated-at.bofh.it> |
| In reply to | #1251558 |
On 10/20/15 3:07 AM, Hannes Frederic Sowa wrote: > <off-topic> > Just a pretty obvious idea is accurate sampling of flows. > </off-topic> ok, so you want to time out flows. Makes sense, but it should be done by user space with little or none help from the kernel. > fdinfo tells me where my position in a file is and which locks the file > have? obviously not. see the example fdinfo from the other email. > So far, if someone wants to delve into the details of a map my approach > would be to take the file descriptor and make it persistence. I have to > think about that some more. nope. you cannot do that. admin should never interfere with running process this way. > Yes, absolutely and I am absolutely against pretty printing key values > in kernel domain. let's table that part. I think it can be useful, but it's irrelevant for this discussion. > So cat-ing them will produce text output with some details about the > map? This is what I wanted to avoid. The concept with symlinks and small > files seems much cleaner and nicer to me. Also you cannot add writable > attributes to this filesystem or you overload stuff heavily? nope. no writeable stuff. fdinfo is read-only. > It is not a tree but a graph, sure, that's why sysfs allows to break the > cyclic dependencies and create symlinks (see holders/ directories). ;) that's an obvious example of another resource waste. You can do that for real devices, but for thousands of maps and programs it is really a waste. > And if you implement the same set of features IMHO you basically > re-implement sysfs. In the beginning we just expose the basic maps and > there won't be any features in sysfs, but it will be cheap to have > read/write flags on maps etc. etc. (I don't know what people will come > up with, yet.). In my opinion those are clearly attributes of a map and > should be defined and managed alongside with their holders. nope. bpf syscall is the only interface to access maps. if we expose them in bpffs it will be read-only for debugging only. > The pinfd feature will provide the future infrastructure alongside to > make this usable, so I think it is worth spending time to think about > it. yes. but since we're going in circles, let's have a 'beer call' to resolve it :) -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web