Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1308974

Re: [PATCH v2 net-next 3/4] bpf: bpf_htab: Add syscall to iterate percpu value of a key

Path csiph.com!aioe.org!bofh.it!news.nic.it!robomod
From Ming Lei <tom.leiming@gmail.com>
Newsgroups linux.kernel
Subject Re: [PATCH v2 net-next 3/4] bpf: bpf_htab: Add syscall to iterate percpu value of a key
Date Thu, 14 Jan 2016 04:30:02 +0100
Message-ID <qQGLM-130-5@gated-at.bofh.it> (permalink)
References <qQ2v0-6Gt-3@gated-at.bofh.it> <qQ2v0-6Gt-15@gated-at.bofh.it> <qQjFw-1z9-21@gated-at.bofh.it> <qQmal-3tP-9@gated-at.bofh.it> <qQvQn-1KI-33@gated-at.bofh.it> <qQETE-83s-19@gated-at.bofh.it>
Dkim-Signature v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113; h=mime-version:in-reply-to:references:date:message-id:subject:from:to :cc:content-type; bh=JDr9G3/nuud1ItCX4CTaCSLj1/8m/XSqb1gdD+wvonE=; b=0GlixWBpAqfHDH+jJW8HF9PEPHCWktb0MLgMsOJmRDO3cMU6sksThNjQHmofBfGklJ v9R0LXKYHJ0g2Z1VXyIuJXmXIdLIHRYytZYKWrIoYFrTJhoS3DGxdwE2/EHBeUe8ZRgm n/PUmGsStuXaIF9r/zZ8h6rXP4qOQ6XzgSM+Y+uYIK8z8Jic44F53AJ8wo1zAU6cKc2v 3xk2UMZ5jPQNv3fwkBXhDeC68GyztBBYDsyfd2V2U5gV2b5USEux8zJBlog0S5qK3h5P +c4Ieva5KSycMdTLs9YEI25oBUpvmrE1zIIV3cQAYk6Ed55TDTBii4y5ndctwrNBZaEB 2g+Q==
X-Google-Dkim-Signature v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20130820; h=x-gm-message-state:mime-version:in-reply-to:references:date :message-id:subject:from:to:cc:content-type; bh=JDr9G3/nuud1ItCX4CTaCSLj1/8m/XSqb1gdD+wvonE=; b=dnlucvzOezLIFN9qMUXSmVXSgBMGPW+h6+1DohbiQyaMnJtHwSuC6ydqUuSuXXJnqO NGwG+XaHIg76jIjxeouopabMxqyih4/de7J0rQ1noTDOh2w5Vss8IUeOCMktvPdVEtWF sKla6Oo+yiQmd9e8vpTgSwJ+L0K1SsEcdSUJJHYRFKGlpyUfxdSvsJPCgLPh6ySQy246 x48UzqghRhGOu+uwcm9vp+QMQ9ry+tROqDMLvXjJnyfvwARSGAyCKHwPzEAWATAaZerA 10wAWREuLEf662rTY4xS8CZq63K6pfzXaMpr3RZwk5iYC5C/W5AEqzWHXygiXxWAuoeK UWGg==
X-Gm-Message-State ALoCoQkX2HbZr8WMPtd3aRmI4BwUasnAnQ/kbIHiniKPezS5TxVPPi6x7SnQoJle0OBPiOV+e9ruV+wCQIbbZ9nz2kmQxqGcng==
MIME-Version 1.0
X-Received by 10.129.103.195 with SMTP id b186mr1546164ywc.197.1452741805454; Wed, 13 Jan 2016 19:23:25 -0800 (PST)
Content-Type text/plain; charset=UTF-8
Sender robomod@news.nic.it
List-ID <linux-kernel.vger.kernel.org>
X-Mailing-List linux-kernel@vger.kernel.org
Approved robomod@news.nic.it
Lines 101
Organization linux.* mail to news gateway
X-Original-Cc Martin KaFai Lau <kafai@fb.com>, Network Development <netdev@vger.kernel.org>, Linux Kernel Mailing List <linux-kernel@vger.kernel.org>, FB Kernel Team <kernel-team@fb.com>
X-Original-Date Thu, 14 Jan 2016 11:23:25 +0800
X-Original-Message-ID <CACVXFVNuj_u+sFsSUMa2GKEyhDJr2r+TDDAxGXvbbMj5NuV2AA@mail.gmail.com>
X-Original-References <1452586852-1575604-1-git-send-email-kafai@fb.com> <1452586852-1575604-4-git-send-email-kafai@fb.com> <CACVXFVPAR52e01vVqtZ5LF5j7zscjAVgOccW1erj-a2VYScJrg@mail.gmail.com> <20160113052341.GB37858@ast-mbp.thefacebook.com> <CACVXFVMMPkekaNvo+ByNGcryB7Mum91xT7nR1YkK+kKu-J1-5w@mail.gmail.com> <20160114012433.GB43324@ast-mbp.thefacebook.com>
X-Original-Sender linux-kernel-owner@vger.kernel.org
Xref csiph.com linux.kernel:1308974

Show key headers only | View raw


On Thu, Jan 14, 2016 at 9:24 AM, Alexei Starovoitov
<alexei.starovoitov@gmail.com> wrote:
> On Wed, Jan 13, 2016 at 11:43:50PM +0800, Ming Lei wrote:
>> On Wed, Jan 13, 2016 at 1:23 PM, Alexei Starovoitov
>> <alexei.starovoitov@gmail.com> wrote:
>> > On Wed, Jan 13, 2016 at 10:42:49AM +0800, Ming Lei wrote:
>> >>
>> >> So I don't think it is good to retrieve value from one CPU via one
>> >> single system call, and accumulate them finally in userspace.
>> >>
>> >> One approach I thought of is to define the function(or sort of)
>> >>
>> >>  handle_cpu_value(u32 cpu, void *val_cpu, void *val_total)
>> >>
>> >> in bpf kernel code for collecting value from each cpu and
>> >> accumulating them into 'val_total', and most of situations, the
>> >> function can be implemented without loop most of situations.
>> >> kernel can call this function directly, and the total value can be
>> >> return to userspace by one single syscall.
>> >>
>> >> Alexei and anyone, could you comment on this draft idea for
>> >> perpcu map?
>> >
>> > I'm not sure how you expect user space to specify such callback.
>> > Kernel cannot execute user code.
>>
>> I mean the above callback function can be built into bpf code and then
>> run from kernel after loading like in packet filter case by tcpdump, maybe
>> one new prog type is needed. It is doable in theroy. I need to investigate
>> a bit to understand how it can be called from kernel, and it might be OK
>> to call it via kprobe, but not elegent just for accumulating value from each
>> CPU.
>
> that would be a total overkill.

With current bpf framework, it isn't difficult to implement the prog type
for percpu, and it is still reasonable because accumulating values
from each CPU is one policy and should have been defined/provide
from bpf code in case of percpu map.

>
>> > Also syscall/malloc/etc is a noise comparing to ipi and it
>> > will still be there, so
>> > for(all cpus) { syscall+ipi;} will have the same speed.
>>
>> In the syscall path, lots of slow things, and finally the accumulated
>> value is often stale and may not reprensent accurate number at any
>> time, and can be thought as invalid.
>
> no. stale != invalid.
> Some analytics/monitor applications are good with ball park numbers
> and for them regular hash map with non-atomic increment is good enough,
> but others need accurate numbers. Even though they may be seconds stale.

Let me make it clearly, and the following is one case for reading packet length
on each CPU.

time: --------T0--------T1----------T2---------T3-------T4-----------------
CPU0:               E0(0)  E0(1)  .... E0(2M)
CPU1:                               E1(0)  E1(2) .... E1(1M)
CPU2:                                                  E2(0) ....
E2(10K)
CPU3:
E3(0)...E3(1K)

                    R0
                                    R1
                                                      R2
                                                                       R3

               A4

1) suppose the system has 4 cpu cores

2) There is one single event(suppose it is packets received) we are
interetested in, and if the event happens on CPU(i) we will add the
received packet's length into the value of CPU(i)

2) E0(i) reprents the ith event happened on CPU0 from T0, which will
be recored into the value of CPU(0), E1(i) represents the ith event on
CPU(1) from T1, ....

2) R0 represents reading value of CPU0 at the time T0, and R1 represents
reading value of CPU1 at the time T1, .....

3) A4 represents accumulating all percpu value into one totoal value at time T4,
suppose value(A4) = value(R0) + value(R1) + value(R2) + value(R3).

3) if we use syscall to implement Ri(i=1...3), the period between T(i)
and T(i+1)
can become quite big, for example dozens of seconds, so the accumulated value
in A4 can't represent the actual/correct value(counter) at any time between T0
and T4, and the value is wrong actually, and all events in above diagram
(E0(0)~E0(2M), E1(0)~E1(1M),  E2(0) .... E2(10K), ...) aren't counted at all,
and the missed number can be quite huge.

So does the value got by A4 make sense for user?


-- 
Ming Lei

Back to linux.kernel | Previous | Next — Previous in thread | Find similar | Unroll thread


Thread

[PATCH v2 net-next 3/4] bpf: bpf_htab: Add syscall to iterate percpu value of a key Martin KaFai Lau <kafai@fb.com> - 2016-01-12 09:30 +0100
  Re: [PATCH v2 net-next 3/4] bpf: bpf_htab: Add syscall to iterate  percpu value of a key Eric Dumazet <eric.dumazet@gmail.com> - 2016-01-13 03:30 +0100
  Re: [PATCH v2 net-next 3/4] bpf: bpf_htab: Add syscall to iterate  percpu value of a key Ming Lei <tom.leiming@gmail.com> - 2016-01-13 03:50 +0100
    Re: [PATCH v2 net-next 3/4] bpf: bpf_htab: Add syscall to iterate  percpu value of a key Alexei Starovoitov <alexei.starovoitov@gmail.com> - 2016-01-13 06:30 +0100
      Re: [PATCH v2 net-next 3/4] bpf: bpf_htab: Add syscall to iterate  percpu value of a key Ming Lei <tom.leiming@gmail.com> - 2016-01-13 16:50 +0100
        Re: [PATCH v2 net-next 3/4] bpf: bpf_htab: Add syscall to iterate  percpu value of a key Alexei Starovoitov <alexei.starovoitov@gmail.com> - 2016-01-14 02:30 +0100
          Re: [PATCH v2 net-next 3/4] bpf: bpf_htab: Add syscall to iterate  percpu value of a key Ming Lei <tom.leiming@gmail.com> - 2016-01-14 04:30 +0100

csiph-web