Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1194994 > unrolled thread
| Started by | He Kuang <hekuang@huawei.com> |
|---|---|
| First post | 2015-07-29 11:40 +0200 |
| Last post | 2015-08-06 09:00 +0200 |
| Articles | 20 on this page of 21 — 6 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event He Kuang <hekuang@huawei.com> - 2015-07-29 11:40 +0200
Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event pi3orama <pi3orama@163.com> - 2015-07-29 22:10 +0200
Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event Alexei Starovoitov <ast@plumgrid.com> - 2015-08-03 21:50 +0200
Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event "Wangnan (F)" <wangnan0@huawei.com> - 2015-08-04 11:10 +0200
Re: Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event "Wangnan (F)" <wangnan0@huawei.com> - 2015-08-05 04:00 +0200
Re: Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event "Wangnan (F)" <wangnan0@huawei.com> - 2015-08-05 04:10 +0200
Re: [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event "Wangnan (F)" <wangnan0@huawei.com> - 2015-08-05 09:00 +0200
Re: [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event Alexei Starovoitov <ast@plumgrid.com> - 2015-08-05 09:20 +0200
Re: [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event "Wangnan (F)" <wangnan0@huawei.com> - 2015-08-05 10:40 +0200
Re: [llvm-dev] [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event Alexei Starovoitov <alexei.starovoitov@gmail.com> - 2015-08-06 05:30 +0200
Re: [llvm-dev] [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event "Wangnan (F)" <wangnan0@huawei.com> - 2015-08-06 06:40 +0200
Re: [llvm-dev] [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event Alexei Starovoitov <alexei.starovoitov@gmail.com> - 2015-08-06 09:00 +0200
Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event "Wangnan (F)" <wangnan0@huawei.com> - 2015-08-12 04:40 +0200
Re: [llvm-dev] llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event Alexei Starovoitov <alexei.starovoitov@gmail.com> - 2015-08-12 07:00 +0200
Re: [llvm-dev] llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event "Wangnan (F)" <wangnan0@huawei.com> - 2015-08-12 07:50 +0200
Re: [llvm-dev] llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event Brenden Blanco <bblanco@gmail.com> - 2015-08-12 15:20 +0200
Re: [llvm-dev] llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event "Wangnan (F)" <wangnan0@huawei.com> - 2015-08-13 08:30 +0200
Re: [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event He Kuang <hekuang@huawei.com> - 2015-08-05 11:00 +0200
Re: [llvm-dev] [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event Alexei Starovoitov <alexei.starovoitov@gmail.com> - 2015-08-06 05:50 +0200
Re: [llvm-dev] [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event "Wangnan (F)" <wangnan0@huawei.com> - 2015-08-06 06:40 +0200
Re: [llvm-dev] [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event Alexei Starovoitov <alexei.starovoitov@gmail.com> - 2015-08-06 09:00 +0200
Page 1 of 2 [1] 2 Next page →
| From | He Kuang <hekuang@huawei.com> |
|---|---|
| Date | 2015-07-29 11:40 +0200 |
| Subject | Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pRw0b-4bi-19@gated-at.bofh.it> |
Hi, Alexei
On 2015/7/28 10:18, Alexei Starovoitov wrote:
> On 7/25/15 3:04 AM, He Kuang wrote:
>> I noticed that for 64-bit elf format, the reloc sections have
>> 'Addend' in the entry, but there's no 'Addend' info in bpf elf
>> file(64bit). I think there must be something wrong in the process
>> of .s -> .o, which related to 64bit/32bit. Anyway, we can parse out the
>> AT_name now, DW_AT_LOCATION still missed and need your help.
>
> looks like objdump/llvm-dwarfdump can only read known EM,
> but that that shouldn't be the problem for your dwarf reader right?
> It should be able to recognize id-s of ELF::R_X86_64_* relo used right?
> As far as AT_location for testprog.c it seems there is no info for
> local variables because they were optimized away.
> With -O0 I see AT_location being emitted.
> Also line number info seems to be good in both cases.
> But in our case, we don't need this anyway, no? we need to see
> the types of structs mainly or you have some other use cases?
> I think line number info would be great to correlate the error reported
> by verifier into specific line in C.
>
Yes, without AT_location, we can lookup the user output data type
by line number, but there're some issues when we look deep.
There're two steps of work that should be done in user space,
first we embed data type into bpf output record, then we use this
type, or index or some other identifier to lookup the type from
dwarf info, so we got a few plans.
* Plan A. Use line number to identify the user data type
Predefined macros:
#define DEFINE_BPF_OUTPUT_DATA(type, var) \
const int BPF_OUTPUT_LINE__##var = __LINE__; type var;
#define BPF_OUTPUT_TRACE_DATA(data, size) \
__bpf_output_trace_data(BPF_OUTPUT_LINE__##data, &data, size)
User defined BPF code:
struct user_define_struct {
...
};
int testprog(int myvar_a, long myvar_b)
{
DEFINE_BPF_OUTPUT_DATA(struct user_define_struct, myvar_c);
BPF_OUTPUT_TRACE_DATA(myvar_c, sizeof(myvar_c));
...
We use macros to embed linenum implicitly, which leads an extra
restriction that user should not define multiple variables in the
same line and not split the macro over multiple lines, like this:
22 DEFINE_BPF_OUTPUT_DATA(struct xxtype, a); DEFINE_BPF_OUTPUT_DATA(struct xxtype, b);
Or
22 DEFINE_BPF_OUTPUT_DATA(struct user_define_struct,
23 myvar_c);
DW_AT_decl_line = 22, while __LINE__ = 23
So we should add verifier in the llvm BPF backend to warn on the
above codes.
* Plan B. Lookup variable type from dwarf AT_location info
We can make use of the output data variable's address, for bpf is
a minus offset to frame base. Then lookup matched offset from
location info(e.g. "DW_OP_fbreg: -32") to identify the variable
type.
For getting the frame base address, we can use builtin functions
like __builtin_frame_base() and __builtin_dwarf_cfa() which
returns the call frame base address. Currently those builtin
functions are not implemented in BPF lower operation yet, so we
tested our bpf program by using a variable tag on frame base, as
following:
struct user_define_struct {
...
};
typedef struct {} frame_base_tag;
int testprog(void)
{
frame_base_tag BPF_FRAME_BASE;
struct user_define_struct myvar_a;
__bpf_trace_output_data((void *)&myvar_a - (void *)&BPF_FRAME_BASE,
&myvar_a, sizeof(myvar_a));
...
The first argument of __bpf_trace_output_data() will be caculated
and it's easy to traverse the variable DIEs in dwarf info and
check each DW_AT_location attribute to find the corresponding
variable type.
The things let us worry about is the opimization may reuse the
stack space which can cause different variables share the same
address, by some rough tests that kind of optimization does not
appear.
* Comparison
Plan A needs less effort and easy to implement, but requires more
check to ensure user not use multiple definition in the same line
and not use macro cross lines.
The advantages of plan B is that we do not need introduce macros
showed in above example and all the things are done implicitly,
but the AT_location info is the prerequisite of this plan, I'm
not sure whether we can guarantee this info in dwarf or not.
Another way we can think of is adding new builtin functions to
indicate the compilier to generate codes return the dwarf type
index directly:
__bpf_trace_output_data(__builtin_dwarf_type(myvar_a), &myvar_a, size);
What's your opinion on those plans, and do you have more
suggestion?
Thank you.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | pi3orama <pi3orama@163.com> |
|---|---|
| Date | 2015-07-29 22:10 +0200 |
| Subject | Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pRFPQ-1zd-21@gated-at.bofh.it> |
| In reply to | #1194994 |
发自我的 iPhone
> 在 2015年7月30日,上午1:13,Alexei Starovoitov <ast@plumgrid.com> 写道:
>
>> On 7/29/15 2:38 AM, He Kuang wrote:
>> Hi, Alexei
>>
>>> On 2015/7/28 10:18, Alexei Starovoitov wrote:
>>>> On 7/25/15 3:04 AM, He Kuang wrote:
>>>> I noticed that for 64-bit elf format, the reloc sections have
>>>> 'Addend' in the entry, but there's no 'Addend' info in bpf elf
>>>> file(64bit). I think there must be something wrong in the process
>>>> of .s -> .o, which related to 64bit/32bit. Anyway, we can parse out the
>>>> AT_name now, DW_AT_LOCATION still missed and need your help.
>>>
>>> looks like objdump/llvm-dwarfdump can only read known EM,
>>> but that that shouldn't be the problem for your dwarf reader right?
>>> It should be able to recognize id-s of ELF::R_X86_64_* relo used right?
>>> As far as AT_location for testprog.c it seems there is no info for
>>> local variables because they were optimized away.
>>> With -O0 I see AT_location being emitted.
>>> Also line number info seems to be good in both cases.
>>> But in our case, we don't need this anyway, no? we need to see
>>> the types of structs mainly or you have some other use cases?
>>> I think line number info would be great to correlate the error reported
>>> by verifier into specific line in C.
>>
>> Yes, without AT_location, we can lookup the user output data type
>> by line number, but there're some issues when we look deep.
>>
>> There're two steps of work that should be done in user space,
>> first we embed data type into bpf output record, then we use this
>> type, or index or some other identifier to lookup the type from
>> dwarf info, so we got a few plans.
>>
>> * Plan A. Use line number to identify the user data type
>>
>> Predefined macros:
>>
>> #define DEFINE_BPF_OUTPUT_DATA(type,
>> var) \
>> const int BPF_OUTPUT_LINE__##var = __LINE__; type var;
>>
>> #define BPF_OUTPUT_TRACE_DATA(data, size) \
>> __bpf_output_trace_data(BPF_OUTPUT_LINE__##data, &data, size)
>>
>> User defined BPF code:
>>
>> struct user_define_struct {
>> ...
>> };
>>
>> int testprog(int myvar_a, long myvar_b)
>> {
>> DEFINE_BPF_OUTPUT_DATA(struct user_define_struct, myvar_c);
>>
>> BPF_OUTPUT_TRACE_DATA(myvar_c, sizeof(myvar_c));
>>
>> ...
>>
>> We use macros to embed linenum implicitly, which leads an extra
>> restriction that user should not define multiple variables in the
>> same line and not split the macro over multiple lines, like this:
>>
>> 22 DEFINE_BPF_OUTPUT_DATA(struct xxtype, a);
>> DEFINE_BPF_OUTPUT_DATA(struct xxtype, b);
>>
>> Or
>>
>> 22 DEFINE_BPF_OUTPUT_DATA(struct user_define_struct,
>> 23 myvar_c);
>>
>> DW_AT_decl_line = 22, while __LINE__ = 23
>>
>> So we should add verifier in the llvm BPF backend to warn on the
>> above codes.
>>
>> * Plan B. Lookup variable type from dwarf AT_location info
>>
>> We can make use of the output data variable's address, for bpf is
>> a minus offset to frame base. Then lookup matched offset from
>> location info(e.g. "DW_OP_fbreg: -32") to identify the variable
>> type.
>>
>> For getting the frame base address, we can use builtin functions
>> like __builtin_frame_base() and __builtin_dwarf_cfa() which
>> returns the call frame base address. Currently those builtin
>> functions are not implemented in BPF lower operation yet, so we
>> tested our bpf program by using a variable tag on frame base, as
>> following:
>>
>> struct user_define_struct {
>> ...
>> };
>>
>> typedef struct {} frame_base_tag;
>>
>> int testprog(void)
>> {
>> frame_base_tag BPF_FRAME_BASE;
>> struct user_define_struct myvar_a;
>>
>> __bpf_trace_output_data((void *)&myvar_a - (void *)&BPF_FRAME_BASE,
>> &myvar_a, sizeof(myvar_a));
>> ...
>>
>> The first argument of __bpf_trace_output_data() will be caculated
>> and it's easy to traverse the variable DIEs in dwarf info and
>> check each DW_AT_location attribute to find the corresponding
>> variable type.
>>
>> The things let us worry about is the opimization may reuse the
>> stack space which can cause different variables share the same
>> address, by some rough tests that kind of optimization does not
>> appear.
>>
>> * Comparison
>>
>> Plan A needs less effort and easy to implement, but requires more
>> check to ensure user not use multiple definition in the same line
>> and not use macro cross lines.
>>
>> The advantages of plan B is that we do not need introduce macros
>> showed in above example and all the things are done implicitly,
>> but the AT_location info is the prerequisite of this plan, I'm
>> not sure whether we can guarantee this info in dwarf or not.
>>
>> Another way we can think of is adding new builtin functions to
>> indicate the compilier to generate codes return the dwarf type
>> index directly:
>>
>> __bpf_trace_output_data(__builtin_dwarf_type(myvar_a), &myvar_a, size);
>
> probably both A and B won't really work when programs get bigger
> and optimizations will start moving lines around.
> the builtin_dwarf_type idea is actually quite interesting.
> Potentially that builtin can stringify type name and later we can
> search it in dwarf. Please take a look how to add such builtin.
> There are few similar builtins that deal with exception handling
> and need type info. May be they can be reused. Like:
> int_eh_typeid_for and int_eh_dwarf_cfa
>
Hi Alexei,
I was wondering if you could give us a hint on adding BPF specific builtins?
Doesn't like other machines, currently there's no Builtins.def for BPF to hold
builtins for that specific target. If we start creating such builtins, we can bring
more there. One builtins I want to see should be __builtin_bpf_strcmp(char*, char*),
with that we can filter events base on name of a task. What I really need is filtering
events based on comm of the main thread, which should be useful on Android
since in Android name of working threads are always something like "Binder_?",
only the main threads names are meaningful. But let's work on strcmp first.
Thank you.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alexei Starovoitov <ast@plumgrid.com> |
|---|---|
| Date | 2015-08-03 21:50 +0200 |
| Message-ID | <pTtUe-4n6-27@gated-at.bofh.it> |
| In reply to | #1194994 |
On 7/31/15 3:18 AM, Wangnan (F) wrote:
>
> However with the newest llvm + clang the DWARF info is still incorrect:
>
> $ objdump --dwarf=info ./out.o
> ...
> <1><3f>: Abbrev Number: 3 (DW_TAG_structure_type)
> <40> DW_AT_name : (indirect string, offset: 0x0): clang
> version 3.8.0 (http://llvm.org/git/clang.git
> f0fcd3432cbed83500df70c18f275d8affb89e5e) (http://llvm.org/git/llvm.git
> c8ccd78d31d4949fa1c14e954ccb06253e18cf37)
> <44> DW_AT_byte_size : 8
> <45> DW_AT_decl_file : 1
> <46> DW_AT_decl_line : 4
> <2><47>: Abbrev Number: 4 (DW_TAG_member)
> <48> DW_AT_name : (indirect string, offset: 0x0): clang
> version 3.8.0 (http://llvm.org/git/clang.git
> f0fcd3432cbed83500df70c18f275d8affb89e5e) (http://llvm.org/git/llvm.git
> c8ccd78d31d4949fa1c14e954ccb06253e18cf37)
> <4c> DW_AT_type : <0x60>
> <50> DW_AT_decl_file : 1
> <51> DW_AT_decl_line : 5
> <52> DW_AT_data_member_location: 0
> <2><53>: Abbrev Number: 4 (DW_TAG_member)
> <54> DW_AT_name : (indirect string, offset: 0x0): clang
> version 3.8.0 (http://llvm.org/git/clang.git
> f0fcd3432cbed83500df70c18f275d8affb89e5e) (http://llvm.org/git/llvm.git
> c8ccd78d31d4949fa1c14e954ccb06253e18cf37)
> <58> DW_AT_type : <0x60>
> <5c> DW_AT_decl_file : 1
> <5d> DW_AT_decl_line : 6
> <5e> DW_AT_data_member_location: 4
> ...
>
> The DW_AT_name is missing.
didn't have time to look at it.
from your llvm patches looks like you've got quite experienced
with it already :)
> I'll post 2 LLVM patches by replying this mail. Please have a look and
> help me
> send them to LLVM if you think my code is correct.
patch 1:
I don't quite understand the purpose of builtin_dwarf_cfa
returning R11. It's a special register seen inside llvm codegen
only. It doesn't have kernel meaning.
patch 2:
do we really need to hack clang?
Can you just define a function that aliases to intrinsic,
like we do for ld_abs/ld_ind ?
void bpf_store_half(void *skb, u64 off, u64 val) asm("llvm.bpf.store.half");
then no extra patches necessary.
> struct my_str {
> int x;
> int y;
> };
> struct my_str __str_my_str;
>
> struct my_str2 {
> int x;
> int y;
> int z;
> };
> struct my_str2 __str_my_str2;
>
> test_func(__builtin_bpf_typeid(&__str_my_str));
> test_func(__builtin_bpf_typeid(&__str_my_str2));
> mov r1, 1
> call 4660
> mov r1, 2
> call 4660
this part looks good. I think it's usable.
> 1. llvm.eh_typeid_for can be used on global variables only. So for each
> output
> structure we have to define a global varable.
why? I think it should work with local pointers as well.
> 2. We still need to find a way to connect the fetchd typeid with DWARF
> info.
> Inserting that ID into DWARF may workable?
hmm, that id should be the same id we're seeing in dwarf, right?
I think it's used in exception handling which is reusing some of
the dwarf stuff for this, so the must be a way to connect that id
to actual type info. Though last time I looked at EH was
during g++ hacking days. No idea how llvm does it exactly, but
I'm assuming the logic for rtti should be similar.
btw, for any deep llvm stuff you may need to move the thread to llvmdev.
May be folks there will have other ideas.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Wangnan (F)" <wangnan0@huawei.com> |
|---|---|
| Date | 2015-08-04 11:10 +0200 |
| Subject | Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pTGop-60Z-3@gated-at.bofh.it> |
| In reply to | #1199229 |
For people who in llvmdev:
This mail is belong to a thread in linux kernel mailing list, the first
message
can be retrived from:
http://lkml.kernel.org/r/55B1535E.8090406@plumgrid.com
Our goal is to fild a way to make BPF program get an unique ID for each type
so it can pass the ID to other part of kernel, then we can retrive the
type and
decode the structure using DWARF information. Currently we have two problem
needs to solve:
1. Dwarf information generated by BPF backend lost all DW_AT_name field;
2. How to get typeid from local variable? I tried llvm.eh_typeid_for
but it support global variable only.
Following is my response to Alexei.
On 2015/8/4 3:44, Alexei Starovoitov wrote:
> On 7/31/15 3:18 AM, Wangnan (F) wrote:
>
[SNIP]
> didn't have time to look at it.
> from your llvm patches looks like you've got quite experienced
> with it already :)
>
>> I'll post 2 LLVM patches by replying this mail. Please have a look and
>> help me
>> send them to LLVM if you think my code is correct.
>
> patch 1:
> I don't quite understand the purpose of builtin_dwarf_cfa
> returning R11. It's a special register seen inside llvm codegen
> only. It doesn't have kernel meaning.
>
Kernel side verifier allows us to do arithmetic computation using two
local variable
address or local variable address and R11. Therefore, we can compute the
location
of a local variable using:
mark = &my_var_a - __builtin_frame_address(0);
If the stack allocation is fixed (if the location is never reused), the
above 'mark'
can be uniquely identify a local variable. That's why I'm interesting in
it. However
I'm not sure whether the prerequestion is hold.
> patch 2:
> do we really need to hack clang?
> Can you just define a function that aliases to intrinsic,
> like we do for ld_abs/ld_ind ?
> void bpf_store_half(void *skb, u64 off, u64 val)
> asm("llvm.bpf.store.half");
> then no extra patches necessary.
>
>> struct my_str {
>> int x;
>> int y;
>> };
>> struct my_str __str_my_str;
>>
>> struct my_str2 {
>> int x;
>> int y;
>> int z;
>> };
>> struct my_str2 __str_my_str2;
>>
>> test_func(__builtin_bpf_typeid(&__str_my_str));
>> test_func(__builtin_bpf_typeid(&__str_my_str2));
>> mov r1, 1
>> call 4660
>> mov r1, 2
>> call 4660
>
> this part looks good. I think it's usable.
>
> > 1. llvm.eh_typeid_for can be used on global variables only. So for each
> > output
> > structure we have to define a global varable.
>
> why? I think it should work with local pointers as well.
>
It is defined by LLVM, in lib/CodeGen/Analysis.cpp:
/// ExtractTypeInfo - Returns the type info, possibly bitcast, encoded in V.
GlobalValue *llvm::ExtractTypeInfo(Value *V) {
...
assert((GV || isa<ConstantPointerNull>(V)) &&
"TypeInfo must be a global variable or NULL"); <-- we can
use only constant pointers
return GV;
}
So from llvm::Intrinsic::eh_typeid_for we can get type of global
variables only.
We may need a new intrinsic for that.
> > 2. We still need to find a way to connect the fetchd typeid with DWARF
> > info.
> > Inserting that ID into DWARF may workable?
>
> hmm, that id should be the same id we're seeing in dwarf, right?
There's no 'typeid' field in dwarf originally. I'm still looking for a way
to inject this ID into dwarf infromation.
> I think it's used in exception handling which is reusing some of
> the dwarf stuff for this, so the must be a way to connect that id
> to actual type info. Though last time I looked at EH was
> during g++ hacking days. No idea how llvm does it exactly, but
> I'm assuming the logic for rtti should be similar.
>
I'm not sure whether RTTI use dwarf to deduce type information. I think not,
because dwarf infos can be stripped out.
> btw, for any deep llvm stuff you may need to move the thread to llvmdev.
> May be folks there will have other ideas.
>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Wangnan (F)" <wangnan0@huawei.com> |
|---|---|
| Date | 2015-08-05 04:00 +0200 |
| Subject | Re: Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pTW9P-3AW-1@gated-at.bofh.it> |
| In reply to | #1199571 |
On 2015/8/4 17:01, Wangnan (F) wrote:
> For people who in llvmdev:
>
> This mail is belong to a thread in linux kernel mailing list, the
> first message
> can be retrived from:
>
> http://lkml.kernel.org/r/55B1535E.8090406@plumgrid.com
>
> Our goal is to fild a way to make BPF program get an unique ID for
> each type
> so it can pass the ID to other part of kernel, then we can retrive the
> type and
> decode the structure using DWARF information. Currently we have two
> problem
> needs to solve:
>
> 1. Dwarf information generated by BPF backend lost all DW_AT_name field;
>
> 2. How to get typeid from local variable? I tried llvm.eh_typeid_for
> but it support global variable only.
>
> Following is my response to Alexei.
>
> On 2015/8/4 3:44, Alexei Starovoitov wrote:
>> On 7/31/15 3:18 AM, Wangnan (F) wrote:
>>
>
> [SNIP]
>
>> didn't have time to look at it.
>> from your llvm patches looks like you've got quite experienced
>> with it already :)
>>
>>> I'll post 2 LLVM patches by replying this mail. Please have a look and
>>> help me
>>> send them to LLVM if you think my code is correct.
>>
>> patch 1:
>> I don't quite understand the purpose of builtin_dwarf_cfa
>> returning R11. It's a special register seen inside llvm codegen
>> only. It doesn't have kernel meaning.
>>
>
> Kernel side verifier allows us to do arithmetic computation using two
> local variable
> address or local variable address and R11. Therefore, we can compute
> the location
> of a local variable using:
>
> mark = &my_var_a - __builtin_frame_address(0);
>
> If the stack allocation is fixed (if the location is never reused),
> the above 'mark'
> can be uniquely identify a local variable. That's why I'm interesting
> in it. However
> I'm not sure whether the prerequestion is hold.
>
>> patch 2:
>> do we really need to hack clang?
>> Can you just define a function that aliases to intrinsic,
>> like we do for ld_abs/ld_ind ?
>> void bpf_store_half(void *skb, u64 off, u64 val)
>> asm("llvm.bpf.store.half");
>> then no extra patches necessary.
>>
>>> struct my_str {
>>> int x;
>>> int y;
>>> };
>>> struct my_str __str_my_str;
>>>
>>> struct my_str2 {
>>> int x;
>>> int y;
>>> int z;
>>> };
>>> struct my_str2 __str_my_str2;
>>>
>>> test_func(__builtin_bpf_typeid(&__str_my_str));
>>> test_func(__builtin_bpf_typeid(&__str_my_str2));
>>> mov r1, 1
>>> call 4660
>>> mov r1, 2
>>> call 4660
>>
>> this part looks good. I think it's usable.
>>
>> > 1. llvm.eh_typeid_for can be used on global variables only. So for
>> each
>> > output
>> > structure we have to define a global varable.
>>
>> why? I think it should work with local pointers as well.
>>
>
> It is defined by LLVM, in lib/CodeGen/Analysis.cpp:
>
> /// ExtractTypeInfo - Returns the type info, possibly bitcast, encoded
> in V.
> GlobalValue *llvm::ExtractTypeInfo(Value *V) {
> ...
> assert((GV || isa<ConstantPointerNull>(V)) &&
> "TypeInfo must be a global variable or NULL"); <-- we can
> use only constant pointers
> return GV;
> }
>
> So from llvm::Intrinsic::eh_typeid_for we can get type of global
> variables only.
>
> We may need a new intrinsic for that.
>
>
>> > 2. We still need to find a way to connect the fetchd typeid with DWARF
>> > info.
>> > Inserting that ID into DWARF may workable?
>>
>> hmm, that id should be the same id we're seeing in dwarf, right?
>
> There's no 'typeid' field in dwarf originally. I'm still looking for a
> way
> to inject this ID into dwarf infromation.
>
>> I think it's used in exception handling which is reusing some of
>> the dwarf stuff for this, so the must be a way to connect that id
>> to actual type info. Though last time I looked at EH was
>> during g++ hacking days. No idea how llvm does it exactly, but
>> I'm assuming the logic for rtti should be similar.
>>
>
> I'm not sure whether RTTI use dwarf to deduce type information. I
> think not,
> because dwarf infos can be stripped out.
>
Hi Alexei,
Just found that llvm::Intrinsic::eh_typeid_for is function specific. ID
of same type in
different functions may be different. Here is an example:
static int (*bpf_output_event)(unsigned long, void *buf, int size) =
(void *) 0x1234;
struct my_str {
int x;
int y;
};
struct my_str __str_my_str;
struct my_str2 {
int x;
int y;
int z;
};
struct my_str2 __str_my_str2;
int func(int *ctx)
{
struct my_str var_a;
struct my_str2 var_b;
bpf_output_event(__builtin_bpf_typeid(&__str_my_str), &var_a,
sizeof(var_a));
bpf_output_event(__builtin_bpf_typeid(&__str_my_str2), &var_b,
sizeof(var_b));
return 0;
}
int func2(int *ctx)
{
struct my_str var_a;
struct my_str2 var_b;
/* change order here */
bpf_output_event(__builtin_bpf_typeid(&__str_my_str2), &var_b,
sizeof(var_b));
bpf_output_event(__builtin_bpf_typeid(&__str_my_str), &var_a,
sizeof(var_a))
return 0;
}
This program uses __builtin_bpf_typeid(llvm::Intrinsic::eh_typeid_for)
in func and func2
for same two types but in different order. We expect same type get same ID.
Compiled with:
$ clang -target bpf -S -O2 -c ./test_bpf_typeid.c
The result is:
.text
.globl func
.align 8
func: # @func
# BB#0: # %entry
mov r2, r10
addi r2, -8
mov r1, 1
mov r3, 8
call 4660
mov r2, r10
addi r2, -24
mov r1, 2
mov r3, 12
call 4660
mov r0, 0
ret
.globl func2
.align 8
func2: # @func2
# BB#0: # %entry
mov r2, r10
addi r2, -24
mov r1, 1 <--- we want 2 here.
mov r3, 12
call 4660
mov r2, r10
addi r2, -8
mov r1, 2 <--- we want 1 here.
mov r3, 8
call 4660
mov r0, 0
ret
.comm __str_my_str,8,4 # @__str_my_str
.comm __str_my_str2,12,4 # @__str_my_str2
Conclusion: llvm::Intrinsic::eh_typeid_for is not on the right direction...
Thank you.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Wangnan (F)" <wangnan0@huawei.com> |
|---|---|
| Date | 2015-08-05 04:10 +0200 |
| Subject | Re: Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pTWjx-41F-43@gated-at.bofh.it> |
| In reply to | #1200351 |
Send again since llvmdev is moved to llvm-dev@lists.llvm.org
On 2015/8/5 9:58, Wangnan (F) wrote:
>
>
> On 2015/8/4 17:01, Wangnan (F) wrote:
>> For people who in llvmdev:
>>
>> This mail is belong to a thread in linux kernel mailing list, the
>> first message
>> can be retrived from:
>>
>> http://lkml.kernel.org/r/55B1535E.8090406@plumgrid.com
>>
>> Our goal is to fild a way to make BPF program get an unique ID for
>> each type
>> so it can pass the ID to other part of kernel, then we can retrive
>> the type and
>> decode the structure using DWARF information. Currently we have two
>> problem
>> needs to solve:
>>
>> 1. Dwarf information generated by BPF backend lost all DW_AT_name field;
>>
>> 2. How to get typeid from local variable? I tried llvm.eh_typeid_for
>> but it support global variable only.
>>
>> Following is my response to Alexei.
>>
>> On 2015/8/4 3:44, Alexei Starovoitov wrote:
>>> On 7/31/15 3:18 AM, Wangnan (F) wrote:
>>>
>>
>> [SNIP]
>>
>>> didn't have time to look at it.
>>> from your llvm patches looks like you've got quite experienced
>>> with it already :)
>>>
>>>> I'll post 2 LLVM patches by replying this mail. Please have a look and
>>>> help me
>>>> send them to LLVM if you think my code is correct.
>>>
>>> patch 1:
>>> I don't quite understand the purpose of builtin_dwarf_cfa
>>> returning R11. It's a special register seen inside llvm codegen
>>> only. It doesn't have kernel meaning.
>>>
>>
>> Kernel side verifier allows us to do arithmetic computation using two
>> local variable
>> address or local variable address and R11. Therefore, we can compute
>> the location
>> of a local variable using:
>>
>> mark = &my_var_a - __builtin_frame_address(0);
>>
>> If the stack allocation is fixed (if the location is never reused),
>> the above 'mark'
>> can be uniquely identify a local variable. That's why I'm interesting
>> in it. However
>> I'm not sure whether the prerequestion is hold.
>>
>>> patch 2:
>>> do we really need to hack clang?
>>> Can you just define a function that aliases to intrinsic,
>>> like we do for ld_abs/ld_ind ?
>>> void bpf_store_half(void *skb, u64 off, u64 val)
>>> asm("llvm.bpf.store.half");
>>> then no extra patches necessary.
>>>
>>>> struct my_str {
>>>> int x;
>>>> int y;
>>>> };
>>>> struct my_str __str_my_str;
>>>>
>>>> struct my_str2 {
>>>> int x;
>>>> int y;
>>>> int z;
>>>> };
>>>> struct my_str2 __str_my_str2;
>>>>
>>>> test_func(__builtin_bpf_typeid(&__str_my_str));
>>>> test_func(__builtin_bpf_typeid(&__str_my_str2));
>>>> mov r1, 1
>>>> call 4660
>>>> mov r1, 2
>>>> call 4660
>>>
>>> this part looks good. I think it's usable.
>>>
>>> > 1. llvm.eh_typeid_for can be used on global variables only. So for
>>> each
>>> > output
>>> > structure we have to define a global varable.
>>>
>>> why? I think it should work with local pointers as well.
>>>
>>
>> It is defined by LLVM, in lib/CodeGen/Analysis.cpp:
>>
>> /// ExtractTypeInfo - Returns the type info, possibly bitcast,
>> encoded in V.
>> GlobalValue *llvm::ExtractTypeInfo(Value *V) {
>> ...
>> assert((GV || isa<ConstantPointerNull>(V)) &&
>> "TypeInfo must be a global variable or NULL"); <-- we can
>> use only constant pointers
>> return GV;
>> }
>>
>> So from llvm::Intrinsic::eh_typeid_for we can get type of global
>> variables only.
>>
>> We may need a new intrinsic for that.
>>
>>
>>> > 2. We still need to find a way to connect the fetchd typeid with
>>> DWARF
>>> > info.
>>> > Inserting that ID into DWARF may workable?
>>>
>>> hmm, that id should be the same id we're seeing in dwarf, right?
>>
>> There's no 'typeid' field in dwarf originally. I'm still looking for
>> a way
>> to inject this ID into dwarf infromation.
>>
>>> I think it's used in exception handling which is reusing some of
>>> the dwarf stuff for this, so the must be a way to connect that id
>>> to actual type info. Though last time I looked at EH was
>>> during g++ hacking days. No idea how llvm does it exactly, but
>>> I'm assuming the logic for rtti should be similar.
>>>
>>
>> I'm not sure whether RTTI use dwarf to deduce type information. I
>> think not,
>> because dwarf infos can be stripped out.
>>
>
> Hi Alexei,
>
> Just found that llvm::Intrinsic::eh_typeid_for is function specific.
> ID of same type in
> different functions may be different. Here is an example:
>
> static int (*bpf_output_event)(unsigned long, void *buf, int size) =
> (void *) 0x1234;
>
> struct my_str {
> int x;
> int y;
> };
> struct my_str __str_my_str;
>
> struct my_str2 {
> int x;
> int y;
> int z;
> };
> struct my_str2 __str_my_str2;
>
> int func(int *ctx)
> {
> struct my_str var_a;
> struct my_str2 var_b;
> bpf_output_event(__builtin_bpf_typeid(&__str_my_str), &var_a,
> sizeof(var_a));
> bpf_output_event(__builtin_bpf_typeid(&__str_my_str2), &var_b,
> sizeof(var_b));
> return 0;
> }
>
> int func2(int *ctx)
> {
> struct my_str var_a;
> struct my_str2 var_b;
>
> /* change order here */
> bpf_output_event(__builtin_bpf_typeid(&__str_my_str2), &var_b,
> sizeof(var_b));
> bpf_output_event(__builtin_bpf_typeid(&__str_my_str), &var_a,
> sizeof(var_a))
> return 0;
> }
>
> This program uses __builtin_bpf_typeid(llvm::Intrinsic::eh_typeid_for)
> in func and func2
> for same two types but in different order. We expect same type get
> same ID.
>
> Compiled with:
>
> $ clang -target bpf -S -O2 -c ./test_bpf_typeid.c
>
> The result is:
>
> .text
> .globl func
> .align 8
> func: # @func
> # BB#0: # %entry
> mov r2, r10
> addi r2, -8
> mov r1, 1
> mov r3, 8
> call 4660
> mov r2, r10
> addi r2, -24
> mov r1, 2
> mov r3, 12
> call 4660
> mov r0, 0
> ret
>
> .globl func2
> .align 8
> func2: # @func2
> # BB#0: # %entry
> mov r2, r10
> addi r2, -24
> mov r1, 1 <--- we want 2 here.
> mov r3, 12
> call 4660
> mov r2, r10
> addi r2, -8
> mov r1, 2 <--- we want 1 here.
> mov r3, 8
> call 4660
> mov r0, 0
> ret
>
> .comm __str_my_str,8,4 # @__str_my_str
> .comm __str_my_str2,12,4 # @__str_my_str2
>
>
> Conclusion: llvm::Intrinsic::eh_typeid_for is not on the right
> direction...
>
> Thank you.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Wangnan (F)" <wangnan0@huawei.com> |
|---|---|
| Date | 2015-08-05 09:00 +0200 |
| Subject | Re: [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pU0Qa-22l-17@gated-at.bofh.it> |
| In reply to | #1200356 |
On 2015/8/5 10:05, Wangnan (F) wrote:
> Send again since llvmdev is moved to llvm-dev@lists.llvm.org
>
> On 2015/8/5 9:58, Wangnan (F) wrote:
>>
>>>
>>> On 2015/8/4 3:44, Alexei Starovoitov wrote:
>>>> On 7/31/15 3:18 AM, Wangnan (F) wrote:
>>>>
>>>
>>> [SNIP]
>>>
>>>> didn't have time to look at it.
>>>> from your llvm patches looks like you've got quite experienced
>>>> with it already :)
>>>>
>>>>> I'll post 2 LLVM patches by replying this mail. Please have a look
>>>>> and
>>>>> help me
>>>>> send them to LLVM if you think my code is correct.
>>>>
>>>> patch 1:
>>>> I don't quite understand the purpose of builtin_dwarf_cfa
>>>> returning R11. It's a special register seen inside llvm codegen
>>>> only. It doesn't have kernel meaning.
>>>>
>>>
>>> Kernel side verifier allows us to do arithmetic computation using
>>> two local variable
>>> address or local variable address and R11. Therefore, we can compute
>>> the location
>>> of a local variable using:
>>>
>>> mark = &my_var_a - __builtin_frame_address(0);
>>>
>>> If the stack allocation is fixed (if the location is never reused),
>>> the above 'mark'
>>> can be uniquely identify a local variable. That's why I'm
>>> interesting in it. However
>>> I'm not sure whether the prerequestion is hold.
>>>
>>>> patch 2:
>>>> do we really need to hack clang?
>>>> Can you just define a function that aliases to intrinsic,
>>>> like we do for ld_abs/ld_ind ?
>>>> void bpf_store_half(void *skb, u64 off, u64 val)
>>>> asm("llvm.bpf.store.half");
>>>> then no extra patches necessary.
>>>>
And for this:
I tried this test function:
void bpf_store_half(void *skb, int off, int val) asm("llvm.bpf.store.half");
int func()
{
bpf_store_half(0, 0, 0);
return 0;
}
Compiled with:
$ clang -g -target bpf -O2 -S -c test.c
And get this:
.text
.globl func
.align 8
func: # @func
# BB#0: # %entry
mov r1, 0
mov r2, 0
mov r3, 0
call llvm.bpf.store.half
mov r0, 0
ret
Without -S, it generate a function relocation:
$ objdump -r ./test.o
./test.o: file format elf64-little
RELOCATION RECORDS FOR [.text]:
OFFSET TYPE VALUE
0000000000000018 UNKNOWN llvm.bpf.store.half
It doesn't work as you suggestion. I think we still need to do something
in clang frontend, or it can only be used in '.ll'.
Thank you.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alexei Starovoitov <ast@plumgrid.com> |
|---|---|
| Date | 2015-08-05 09:20 +0200 |
| Subject | Re: [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pU19w-2FR-1@gated-at.bofh.it> |
| In reply to | #1200462 |
On 8/4/15 11:51 PM, Wangnan (F) wrote:
> void bpf_store_half(void *skb, int off, int val)
> asm("llvm.bpf.store.half");
> int func()
> {
> bpf_store_half(0, 0, 0);
> return 0;
> }
>
> Compiled with:
>
> $ clang -g -target bpf -O2 -S -c test.c
>
> And get this:
>
> .text
> .globl func
> .align 8
> func: # @func
> # BB#0: # %entry
> mov r1, 0
> mov r2, 0
> mov r3, 0
> call llvm.bpf.store.half
it didn't work because number and types of args were incompatible.
Every samples/bpf/sockex[0-9]_kern.c is using llvm.bpf.load.* intrinsics.
the typeid changing ids with order is surprising.
I think the assertion in ExtractTypeInfo() is not hard.
Just there were no such use cases. May be we can do something
similar to what LowerIntrinsicCall() does and lower it differently
in the backend.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Wangnan (F)" <wangnan0@huawei.com> |
|---|---|
| Date | 2015-08-05 10:40 +0200 |
| Subject | Re: [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pU2oV-4nL-1@gated-at.bofh.it> |
| In reply to | #1200468 |
On 2015/8/5 15:11, Alexei Starovoitov wrote:
> On 8/4/15 11:51 PM, Wangnan (F) wrote:
>> void bpf_store_half(void *skb, int off, int val)
>> asm("llvm.bpf.store.half");
>> int func()
>> {
>> bpf_store_half(0, 0, 0);
>> return 0;
>> }
>>
>> Compiled with:
>>
>> $ clang -g -target bpf -O2 -S -c test.c
>>
>> And get this:
>>
>> .text
>> .globl func
>> .align 8
>> func: # @func
>> # BB#0: # %entry
>> mov r1, 0
>> mov r2, 0
>> mov r3, 0
>> call llvm.bpf.store.half
>
> it didn't work because number and types of args were incompatible.
> Every samples/bpf/sockex[0-9]_kern.c is using llvm.bpf.load.* intrinsics.
>
Make it work. Thank you.
It doesn't work for me at first since in my llvm there's only
llvm.bpf.load.*.
I think llvm.bpf.store.* belone to some patches you haven't posted yet?
> the typeid changing ids with order is surprising.
> I think the assertion in ExtractTypeInfo() is not hard.
> Just there were no such use cases. May be we can do something
> similar to what LowerIntrinsicCall() does and lower it differently
> in the backend.
>
But in backend can we still get type information? I thought type is
meaningful in frontend only, and backend behaviors is unable to affect
DWARF generation, right?
Thank you.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alexei Starovoitov <alexei.starovoitov@gmail.com> |
|---|---|
| Date | 2015-08-06 05:30 +0200 |
| Subject | Re: [llvm-dev] [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pUk2t-4Ik-5@gated-at.bofh.it> |
| In reply to | #1200528 |
On Wed, Aug 05, 2015 at 04:28:13PM +0800, Wangnan (F) wrote: > > It doesn't work for me at first since in my llvm there's only > llvm.bpf.load.*. > > I think llvm.bpf.store.* belone to some patches you haven't posted yet? nope. only loads have special instructions ld_abs/ld_ind which are represented by these intrinsics. stores, so far, are done via single bpf_store_bytes() helper function. > >the typeid changing ids with order is surprising. > >I think the assertion in ExtractTypeInfo() is not hard. > >Just there were no such use cases. May be we can do something > >similar to what LowerIntrinsicCall() does and lower it differently > >in the backend. > > > But in backend can we still get type information? I thought type is > meaningful in frontend only, and backend behaviors is unable to affect > DWARF generation, right? why do we need to affect type generation? we just need to know dwarf type id in the backend, so we can emit it as a constant. I still think lowering eh_typeid_for differently may work. Like instead of doing GV = ExtractTypeInfo(I.getArgOperand(0)) followed by getMachineFunction().getMMI().getTypeIDFor(GV) we can get dwarf type id from I.getArgOperand(0) if it's any pointer to struct type. I'm not familiar with dwarf handling part of llvm, but feels possible. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Wangnan (F)" <wangnan0@huawei.com> |
|---|---|
| Date | 2015-08-06 06:40 +0200 |
| Subject | Re: [llvm-dev] [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pUl8d-6gt-3@gated-at.bofh.it> |
| In reply to | #1201408 |
On 2015/8/6 11:22, Alexei Starovoitov wrote:
> On Wed, Aug 05, 2015 at 04:28:13PM +0800, Wangnan (F) wrote:
>> It doesn't work for me at first since in my llvm there's only
>> llvm.bpf.load.*.
>>
>> I think llvm.bpf.store.* belone to some patches you haven't posted yet?
> nope. only loads have special instructions ld_abs/ld_ind
> which are represented by these intrinsics.
> stores, so far, are done via single bpf_store_bytes() helper function.
>
>>> the typeid changing ids with order is surprising.
>>> I think the assertion in ExtractTypeInfo() is not hard.
>>> Just there were no such use cases. May be we can do something
>>> similar to what LowerIntrinsicCall() does and lower it differently
>>> in the backend.
>>>
>> But in backend can we still get type information? I thought type is
>> meaningful in frontend only, and backend behaviors is unable to affect
>> DWARF generation, right?
> why do we need to affect type generation? we just need to know dwarf
> type id in the backend, so we can emit it as a constant.
> I still think lowering eh_typeid_for differently may work.
> Like instead of doing
> GV = ExtractTypeInfo(I.getArgOperand(0)) followed by
> getMachineFunction().getMMI().getTypeIDFor(GV)
> we can get dwarf type id from I.getArgOperand(0) if it's
> any pointer to struct type.
I have a bad news to tell:
#include <stdio.h>
struct my_str {
int x;
int y;
} __gv_my_str;
struct my_str __gv_my_str_;
struct my_str2 {
int x;
int y;
} __gv_my_str2;
int typeid(void *p) asm("llvm.eh.typeid.for");
int main()
{
printf("%d\n", typeid(&__gv_my_str));
printf("%d\n", typeid(&__gv_my_str_));
printf("%d\n", typeid(&__gv_my_str2));
return 0;
}
Compiled with clang into x86 executable, then:
$ ./a.out
3
2
1
See? I have two types but reported 3 IDs.
And here is the implementation of getTypeIDFor, in
lib/CodeGen/MachineModuleInfo.cpp:
unsigned MachineModuleInfo::getTypeIDFor(const GlobalValue *TI) {
for (unsigned i = 0, N = TypeInfos.size(); i != N; ++i)
if (TypeInfos[i] == TI) return i + 1;
TypeInfos.push_back(TI);
return TypeInfos.size();
}
It only checks value in a stupid way.
Now the dwarf side becomes clear (see my other response), but the
frontend may require
totally reconsidering.
Do you know someone in LLVM-dev who can help us?
Thank you.
> I'm not familiar with dwarf handling part of llvm, but feels possible.
>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alexei Starovoitov <alexei.starovoitov@gmail.com> |
|---|---|
| Date | 2015-08-06 09:00 +0200 |
| Subject | Re: [llvm-dev] [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pUnjI-1ck-9@gated-at.bofh.it> |
| In reply to | #1201424 |
On Thu, Aug 06, 2015 at 12:35:30PM +0800, Wangnan (F) wrote:
>
>
> On 2015/8/6 11:22, Alexei Starovoitov wrote:
> >On Wed, Aug 05, 2015 at 04:28:13PM +0800, Wangnan (F) wrote:
> >>It doesn't work for me at first since in my llvm there's only
> >>llvm.bpf.load.*.
> >>
> >>I think llvm.bpf.store.* belone to some patches you haven't posted yet?
> >nope. only loads have special instructions ld_abs/ld_ind
> >which are represented by these intrinsics.
> >stores, so far, are done via single bpf_store_bytes() helper function.
> >
> >>>the typeid changing ids with order is surprising.
> >>>I think the assertion in ExtractTypeInfo() is not hard.
> >>>Just there were no such use cases. May be we can do something
> >>>similar to what LowerIntrinsicCall() does and lower it differently
> >>>in the backend.
> >>>
> >>But in backend can we still get type information? I thought type is
> >>meaningful in frontend only, and backend behaviors is unable to affect
> >>DWARF generation, right?
> >why do we need to affect type generation? we just need to know dwarf
> >type id in the backend, so we can emit it as a constant.
> >I still think lowering eh_typeid_for differently may work.
> >Like instead of doing
> >GV = ExtractTypeInfo(I.getArgOperand(0)) followed by
> >getMachineFunction().getMMI().getTypeIDFor(GV)
> >we can get dwarf type id from I.getArgOperand(0) if it's
> >any pointer to struct type.
>
> I have a bad news to tell:
>
> #include <stdio.h>
> struct my_str {
> int x;
> int y;
> } __gv_my_str;
> struct my_str __gv_my_str_;
>
> struct my_str2 {
> int x;
> int y;
> } __gv_my_str2;
>
> int typeid(void *p) asm("llvm.eh.typeid.for");
>
> int main()
> {
> printf("%d\n", typeid(&__gv_my_str));
> printf("%d\n", typeid(&__gv_my_str_));
> printf("%d\n", typeid(&__gv_my_str2));
> return 0;
> }
>
> Compiled with clang into x86 executable, then:
>
> $ ./a.out
> 3
> 2
> 1
>
> See? I have two types but reported 3 IDs.
that's expected. We don't have to use default lowering
of typeid_for with getTypeIDFor. bpf backend specific
lowering can be different, though in this case it's odd
that id for __gv_my_str and __gv_my_str_ are different.
__gv_my_str and __gv_my_str2 should be different.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Wangnan (F)" <wangnan0@huawei.com> |
|---|---|
| Date | 2015-08-12 04:40 +0200 |
| Subject | Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pWu7n-7pH-3@gated-at.bofh.it> |
| In reply to | #1199229 |
On 2015/8/4 3:44, Alexei Starovoitov wrote:
[SNIP]
>> I'll post 2 LLVM patches by replying this mail. Please have a look and
>> help me
>> send them to LLVM if you think my code is correct.
>
>
[SNIP]
> patch 2:
> do we really need to hack clang?
> Can you just define a function that aliases to intrinsic,
> like we do for ld_abs/ld_ind ?
> void bpf_store_half(void *skb, u64 off, u64 val)
> asm("llvm.bpf.store.half");
> then no extra patches necessary.
Hi Alexei,
By two weeks researching, I have to give you a sad answer that:
target specific intrinsic is not work.
I tried target specific intrinsic. However, LLVM isolates backend and
frontend, and there's no way to pass language level type information
to backend code.
Think about a program like this:
struct strA { int a; }
struct strB { int b; }
int func() {
struct strA a;
struct strB b;
a.a = 1;
b.b = 2;
bpf_output(gettype(a), &a);
bpf_output(gettype(b), &b);
return 0;
}
BPF backend can't (and needn't) tell the difference between local
variables a and b in theory. In LLVM implementation, it filters type
information out using ComputeValueVTs(). Please have a look at
SelectionDAGBuilder::visitIntrinsicCall in
lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp and
SelectionDAGBuilder::visitTargetIntrinsic in the same file. in
visitTargetIntrinsic, ComputeValueVTs acts as a barrier which strips
type information out from CallInst ("I"), and leave SDValue and SDVTList
("Ops" and "VTs") to target code. SDValue and SDVTList are wrappers of
EVT and MVT, all information we concern won't be passed here.
I think now we have 2 choices:
1. Hacking into clang, implement target specific builtin function. Now I
have worked out a ugly but workable patch which setup a builtin
function:
__builtin_bpf_typeid(), which accepts local or global variable then
returns different constant for different types.
2. Implementing an LLVM intrinsic call (llvm.typeid), make it be
processed in
visitIntrinsicCall(). I think we can get something useful if it is
processed
with that function.
The next thing should be generating debug information to map type and
constants which issued by __builtin_bpf_typeid() or llvm.typeid. Now we
have a crazy idea that, if we limit the name of the structure to 8 bytes,
we can insert the name into a u64, then there would be no need to consider
type information in DWARF. For example, in the above sample code, gettype(a)
will issue 0x0000000041727473 because its type is "strA". What do you think?
Thank you.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alexei Starovoitov <alexei.starovoitov@gmail.com> |
|---|---|
| Date | 2015-08-12 07:00 +0200 |
| Subject | Re: [llvm-dev] llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pWwiS-24c-7@gated-at.bofh.it> |
| In reply to | #1205558 |
On Wed, Aug 12, 2015 at 10:34:43AM +0800, Wangnan (F) via llvm-dev wrote:
>
> Think about a program like this:
>
> struct strA { int a; }
> struct strB { int b; }
> int func() {
> struct strA a;
> struct strB b;
>
> a.a = 1;
> b.b = 2;
> bpf_output(gettype(a), &a);
> bpf_output(gettype(b), &b);
> return 0;
> }
>
> BPF backend can't (and needn't) tell the difference between local
> variables a and b in theory. In LLVM implementation, it filters type
> information out using ComputeValueVTs(). Please have a look at
> SelectionDAGBuilder::visitIntrinsicCall in
> lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp and
> SelectionDAGBuilder::visitTargetIntrinsic in the same file. in
> visitTargetIntrinsic, ComputeValueVTs acts as a barrier which strips
> type information out from CallInst ("I"), and leave SDValue and SDVTList
> ("Ops" and "VTs") to target code. SDValue and SDVTList are wrappers of
> EVT and MVT, all information we concern won't be passed here.
>
> I think now we have 2 choices:
>
> 1. Hacking into clang, implement target specific builtin function. Now I
> have worked out a ugly but workable patch which setup a builtin function:
> __builtin_bpf_typeid(), which accepts local or global variable then
> returns different constant for different types.
>
> 2. Implementing an LLVM intrinsic call (llvm.typeid), make it be processed
> in
> visitIntrinsicCall(). I think we can get something useful if it is
> processed
> with that function.
Yeah. You're right about pure target intrinsics.
I think llvm.typeid might work. imo it's cleaner than
doing it at clang level.
> The next thing should be generating debug information to map type and
> constants which issued by __builtin_bpf_typeid() or llvm.typeid. Now we
> have a crazy idea that, if we limit the name of the structure to 8 bytes,
> we can insert the name into a u64, then there would be no need to consider
> type information in DWARF. For example, in the above sample code, gettype(a)
> will issue 0x0000000041727473 because its type is "strA". What do you think?
that's way too hacky.
I was thinking when compiling we can keep llvm ir along with .o
instead of dwarf and extract type info from there.
dwarf has names and other things that we don't need. We only
care about actual field layout of the structs.
But it probably won't be easy to parse llvm ir on perf side
instead of dwarf.
btw, if you haven't looked at iovisor/bcc, there we're solving
similar problem differently. There we use clang rewriter, so all
structs fields are visible at this level, then we use bpf backend
in JIT mode and push bpf instructions into the kernel on the fly
completely skipping ELF and .o
For example in:
https://github.com/iovisor/bcc/blob/master/examples/distributed_bridge/tunnel.c
when you see
struct ethernet_t {
unsigned long long dst:48;
unsigned long long src:48;
unsigned int type:16;
} BPF_PACKET_HEADER;
struct ethernet_t *ethernet = cursor_advance(cursor, sizeof(*ethernet));
... ethernet->src ...
is recognized by clang rewriter and ->src is converted to a different
C code that is sent again into clang.
So there is no need to use dwarf or patch clang/llvm. clang rewriter
has all the info.
I'm not sure you can live with clang/llvm on the host where you
want to run the tracing bits, but if you can that's an easier option.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Wangnan (F)" <wangnan0@huawei.com> |
|---|---|
| Date | 2015-08-12 07:50 +0200 |
| Subject | Re: [llvm-dev] llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pWx5g-3D9-3@gated-at.bofh.it> |
| In reply to | #1205584 |
On 2015/8/12 12:57, Alexei Starovoitov wrote:
> On Wed, Aug 12, 2015 at 10:34:43AM +0800, Wangnan (F) via llvm-dev wrote:
>> Think about a program like this:
>>
>> struct strA { int a; }
>> struct strB { int b; }
>> int func() {
>> struct strA a;
>> struct strB b;
>>
>> a.a = 1;
>> b.b = 2;
>> bpf_output(gettype(a), &a);
>> bpf_output(gettype(b), &b);
>> return 0;
>> }
>>
>> BPF backend can't (and needn't) tell the difference between local
>> variables a and b in theory. In LLVM implementation, it filters type
>> information out using ComputeValueVTs(). Please have a look at
>> SelectionDAGBuilder::visitIntrinsicCall in
>> lib/CodeGen/SelectionDAG/SelectionDAGBuilder.cpp and
>> SelectionDAGBuilder::visitTargetIntrinsic in the same file. in
>> visitTargetIntrinsic, ComputeValueVTs acts as a barrier which strips
>> type information out from CallInst ("I"), and leave SDValue and SDVTList
>> ("Ops" and "VTs") to target code. SDValue and SDVTList are wrappers of
>> EVT and MVT, all information we concern won't be passed here.
>>
>> I think now we have 2 choices:
>>
>> 1. Hacking into clang, implement target specific builtin function. Now I
>> have worked out a ugly but workable patch which setup a builtin function:
>> __builtin_bpf_typeid(), which accepts local or global variable then
>> returns different constant for different types.
>>
>> 2. Implementing an LLVM intrinsic call (llvm.typeid), make it be processed
>> in
>> visitIntrinsicCall(). I think we can get something useful if it is
>> processed
>> with that function.
> Yeah. You're right about pure target intrinsics.
> I think llvm.typeid might work. imo it's cleaner than
> doing it at clang level.
>
>> The next thing should be generating debug information to map type and
>> constants which issued by __builtin_bpf_typeid() or llvm.typeid. Now we
>> have a crazy idea that, if we limit the name of the structure to 8 bytes,
>> we can insert the name into a u64, then there would be no need to consider
>> type information in DWARF. For example, in the above sample code, gettype(a)
>> will issue 0x0000000041727473 because its type is "strA". What do you think?
> that's way too hacky.
> I was thinking when compiling we can keep llvm ir along with .o
> instead of dwarf and extract type info from there.
> dwarf has names and other things that we don't need. We only
> care about actual field layout of the structs.
> But it probably won't be easy to parse llvm ir on perf side
> instead of dwarf.
Shipping both llvm IR and .o to perf makes it harder to use. I'm
not sure whether it is a good idea. If we are unable to encode the
structure using a u64, let's still dig into dwarf.
We have another idea that we can utilize dwarf's existing feature.
For example, when __buildin_bpf_typeid() get called, define an enumerate
type in dwarf info, so you'll find:
<1><2a>: Abbrev Number: 2 (DW_TAG_enumeration_type)
<2b> DW_AT_name : (indirect string, offset: 0xec): TYPEINFO
<2f> DW_AT_byte_size : 4
<30> DW_AT_decl_file : 1
<31> DW_AT_decl_line : 3
<2><32>: Abbrev Number: 3 (DW_TAG_enumerator)
<33> DW_AT_name : (indirect string, offset: 0xcc):
__typeinfo_strA
<37> DW_AT_const_value : 2
<2><38>: Abbrev Number: 3 (DW_TAG_enumerator)
<39> DW_AT_name : (indirect string, offset: 0xdc):
__typeinfo_strB
<3d> DW_AT_const_value : 3
or this:
<3><54>: Abbrev Number: 4 (DW_TAG_variable)
<55> DW_AT_const_value : 2
<66> DW_AT_name : (indirect string, offset: 0x1e):
__typeinfo_strA
<6a> DW_AT_decl_file : 1
<6b> DW_AT_decl_line : 29
<6c> DW_AT_type : <0x72>
then from DW_AT_name and DW_AT_const_value we can do the mapping.
Drawback is that
all __typeinfo_ prefixed names become reserved.
> btw, if you haven't looked at iovisor/bcc, there we're solving
> similar problem differently. There we use clang rewriter, so all
> structs fields are visible at this level, then we use bpf backend
> in JIT mode and push bpf instructions into the kernel on the fly
> completely skipping ELF and .o
> For example in:
> https://github.com/iovisor/bcc/blob/master/examples/distributed_bridge/tunnel.c
> when you see
> struct ethernet_t {
> unsigned long long dst:48;
> unsigned long long src:48;
> unsigned int type:16;
> } BPF_PACKET_HEADER;
> struct ethernet_t *ethernet = cursor_advance(cursor, sizeof(*ethernet));
> ... ethernet->src ...
> is recognized by clang rewriter and ->src is converted to a different
> C code that is sent again into clang.
> So there is no need to use dwarf or patch clang/llvm. clang rewriter
> has all the info.
Could you please give us further information about your clang rewriter?
I guess you need a new .so when injecting those code into kernel?
> I'm not sure you can live with clang/llvm on the host where you
> want to run the tracing bits, but if you can that's an easier option.
>
I'm not sure. Our target platform should be embedded devices like
smartphone.
Bringing full clang/llvm environment there is not acceptable.
Thank you.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Brenden Blanco <bblanco@gmail.com> |
|---|---|
| Date | 2015-08-12 15:20 +0200 |
| Subject | Re: [llvm-dev] llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pWE6K-5sr-17@gated-at.bofh.it> |
| In reply to | #1205597 |
Hi Wangnan, I've been authoring the BCC development, so I'll answer those specific questions. > > > Could you please give us further information about your clang rewriter? > I guess you need a new .so when injecting those code into kernel? The rewriter runs all of its passes in a single process, creating no files on disk and having no external dependencies in terms of toolchain. 1. Entry point: bpf_module_create() - C API call to create module, can take filename or directly a c string with the full contents of the program 2. Convert contents into a clang memory buffer 3. Set up a clang driver::CompilerInvocation in the style of the clang interpreter example 4. Run a rewriter pass over the memory buffer file, annotating and/or doing BPF specific magic on the input source a. Open BPF maps with a call to bpf_create_map directly b. Convert references to map operations with the specific FD of the new map c. Convert arguments to bpf_probe_read calls as needed d. Collect the externed function names to avoid section() hack in the language 5. Re-run the CompilerInvocation on the modified sources 6. JIT the llvm::Module to bpf arch 7. Load the resulting in-memory ".o" to bpf_prog_load, keeping the FD alive in the compiler process 8. Attach the FD as necessary to perf events, socket, tc, etc. 9. goto 1 The above steps are captured in the BCC github repo in src/cc, with the clang specific bits inside of the frontends/clang subdirectory. > I'm not sure. Our target platform should be embedded devices like > smartphone. > Bringing full clang/llvm environment there is not acceptable. The artifact from the build process of BCC is a shared library, which has the clang/llvm .a embedded within them. It is not yet a single binary, but not unfeasible to make it so. The clang toolchain itself does not need to exist on the target. I have not attempted to cross-compile BCC to any architecture, currently x86_64 only. If you have more BCC specific questions not involving clang/llvm, perhaps you can ping Alexei/myself off of the llvm-dev list, in case this discussion is not relevant to them. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Wangnan (F)" <wangnan0@huawei.com> |
|---|---|
| Date | 2015-08-13 08:30 +0200 |
| Subject | Re: [llvm-dev] llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pWUbv-3xi-3@gated-at.bofh.it> |
| In reply to | #1206075 |
Thank you for your reply. Add He Kuang to CC list. On 2015/8/12 21:15, Brenden Blanco wrote: > Hi Wangnan, I've been authoring the BCC development, so I'll answer > those specific questions. >> >> Could you please give us further information about your clang rewriter? >> I guess you need a new .so when injecting those code into kernel? > The rewriter runs all of its passes in a single process, creating no > files on disk and having no external dependencies in terms of > toolchain. > 1. Entry point: bpf_module_create() - C API call to create module, can > take filename or directly a c string with the full contents of the > program > 2. Convert contents into a clang memory buffer > 3. Set up a clang driver::CompilerInvocation in the style of the clang > interpreter example > 4. Run a rewriter pass over the memory buffer file, annotating and/or > doing BPF specific magic on the input source > a. Open BPF maps with a call to bpf_create_map directly > b. Convert references to map operations with the specific FD of the new map > c. Convert arguments to bpf_probe_read calls as needed > d. Collect the externed function names to avoid section() hack in the language > 5. Re-run the CompilerInvocation on the modified sources > 6. JIT the llvm::Module to bpf arch > 7. Load the resulting in-memory ".o" to bpf_prog_load, keeping the FD > alive in the compiler process > 8. Attach the FD as necessary to perf events, socket, tc, etc. > 9. goto 1 > > The above steps are captured in the BCC github repo in src/cc, with > the clang specific bits inside of the frontends/clang subdirectory. > >> I'm not sure. Our target platform should be embedded devices like >> smartphone. >> Bringing full clang/llvm environment there is not acceptable. > The artifact from the build process of BCC is a shared library, which > has the clang/llvm .a embedded within them. It is not yet a single > binary, but not unfeasible to make it so. The clang toolchain itself > does not need to exist on the target. I have not attempted to > cross-compile BCC to any architecture, currently x86_64 only. > > If you have more BCC specific questions not involving clang/llvm, > perhaps you can ping Alexei/myself off of the llvm-dev list, in case > this discussion is not relevant to them. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | He Kuang <hekuang@huawei.com> |
|---|---|
| Date | 2015-08-05 11:00 +0200 |
| Subject | Re: [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pU2Ij-4Ky-25@gated-at.bofh.it> |
| In reply to | #1194994 |
Hi, Alexei On 2015/7/30 1:13, Alexei Starovoitov wrote: > On 7/29/15 2:38 AM, He Kuang wrote: >> Hi, Alexei >> >> On 2015/7/28 10:18, Alexei Starovoitov wrote: >>> On 7/25/15 3:04 AM, He Kuang wrote: >>>> I noticed that for 64-bit elf format, the reloc sections have >>>> 'Addend' in the entry, but there's no 'Addend' info in bpf elf >>>> file(64bit). I think there must be something wrong in the process >>>> of .s -> .o, which related to 64bit/32bit. Anyway, we can parse out the >>>> AT_name now, DW_AT_LOCATION still missed and need your help. >>> Another thing about DW_AT_name, we've already found that the name string is stored indirectly and needs relocation which is architecture specific, while the e_machine info in bpf obj file is "unknown", both objdump and libdw cannot parse DW_AT_name correctly. Should we just use a known architeture for bpf object file instead of "unknown"? If so, we can use the existing relocation codes in libdw and get DIE name by simply invoking dwarf_diename(). The drawback of this method is that, e.g. we use "x86-64" instead, is hard to distinguish bpf obj file with x86-64 elf file. Do you think this is ok? Otherwise, for not touching libdw, we should reimplement the relocation codes already in libdw for bpf elf file with "unknown" machine info specially in perf. I wonder whether it is worth doing this and what's your opinion? Thank you. >> index directly: >> >> __bpf_trace_output_data(__builtin_dwarf_type(myvar_a), &myvar_a, size); >> > > probably both A and B won't really work when programs get bigger > and optimizations will start moving lines around. > the builtin_dwarf_type idea is actually quite interesting. > Potentially that builtin can stringify type name and later we can > search it in dwarf. Please take a look how to add such builtin. > There are few similar builtins that deal with exception handling > and need type info. May be they can be reused. Like: > int_eh_typeid_for and int_eh_dwarf_cfa > > -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alexei Starovoitov <alexei.starovoitov@gmail.com> |
|---|---|
| Date | 2015-08-06 05:50 +0200 |
| Subject | Re: [llvm-dev] [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pUklP-54X-11@gated-at.bofh.it> |
| In reply to | #1200548 |
On Wed, Aug 05, 2015 at 04:59:01PM +0800, He Kuang wrote: > Hi, Alexei > > On 2015/7/30 1:13, Alexei Starovoitov wrote: > >On 7/29/15 2:38 AM, He Kuang wrote: > >>Hi, Alexei > >> > >>On 2015/7/28 10:18, Alexei Starovoitov wrote: > >>>On 7/25/15 3:04 AM, He Kuang wrote: > >>>>I noticed that for 64-bit elf format, the reloc sections have > >>>>'Addend' in the entry, but there's no 'Addend' info in bpf elf > >>>>file(64bit). I think there must be something wrong in the process > >>>>of .s -> .o, which related to 64bit/32bit. Anyway, we can parse out the > >>>>AT_name now, DW_AT_LOCATION still missed and need your help. > >>> > > Another thing about DW_AT_name, we've already found that the name > string is stored indirectly and needs relocation which is > architecture specific, while the e_machine info in bpf obj file > is "unknown", both objdump and libdw cannot parse DW_AT_name > correctly. > > Should we just use a known architeture for bpf object file > instead of "unknown"? If so, we can use the existing relocation > codes in libdw and get DIE name by simply invoking > dwarf_diename(). The drawback of this method is that, e.g. we > use "x86-64" instead, is hard to distinguish bpf obj file with > x86-64 elf file. Do you think this is ok? The only clean way would be to register bpf as an architecture with elf standards committee. I have no idea who is doing that and how much such new e_machine registration may cost. So far using EM_NONE is a hack to avoid bureaucracy. Are dwarf relocation processor specific? Then simple hack to elfutils/libdw to treat EM_NONE as X64 should do the trick, right? If that indeed works, we can tweak bpf backend to use EM_X86_64, but then the danger that such .o file will be wrongly recognized by elf utils. imo it's safer to keep it as EM_NONE until real number is assigned, but even after it's assigned it will take time to propagate that value. So for now I would try to find a solution keeping EM_NONE hack. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Wangnan (F)" <wangnan0@huawei.com> |
|---|---|
| Date | 2015-08-06 06:40 +0200 |
| Subject | Re: [llvm-dev] [LLVMdev] Cc llvmdev: Re: llvm bpf debug info. Re: [RFC PATCH v4 3/3] bpf: Introduce function for outputing data to perf event |
| Message-ID | <pUl8d-6gt-5@gated-at.bofh.it> |
| In reply to | #1201414 |
On 2015/8/6 11:41, Alexei Starovoitov wrote: > On Wed, Aug 05, 2015 at 04:59:01PM +0800, He Kuang wrote: >> Hi, Alexei >> >> On 2015/7/30 1:13, Alexei Starovoitov wrote: >>> On 7/29/15 2:38 AM, He Kuang wrote: >>>> Hi, Alexei >>>> >>>> On 2015/7/28 10:18, Alexei Starovoitov wrote: >>>>> On 7/25/15 3:04 AM, He Kuang wrote: >>>>>> I noticed that for 64-bit elf format, the reloc sections have >>>>>> 'Addend' in the entry, but there's no 'Addend' info in bpf elf >>>>>> file(64bit). I think there must be something wrong in the process >>>>>> of .s -> .o, which related to 64bit/32bit. Anyway, we can parse out the >>>>>> AT_name now, DW_AT_LOCATION still missed and need your help. >> Another thing about DW_AT_name, we've already found that the name >> string is stored indirectly and needs relocation which is >> architecture specific, while the e_machine info in bpf obj file >> is "unknown", both objdump and libdw cannot parse DW_AT_name >> correctly. >> >> Should we just use a known architeture for bpf object file >> instead of "unknown"? If so, we can use the existing relocation >> codes in libdw and get DIE name by simply invoking >> dwarf_diename(). The drawback of this method is that, e.g. we >> use "x86-64" instead, is hard to distinguish bpf obj file with >> x86-64 elf file. Do you think this is ok? > The only clean way would be to register bpf as an architecture > with elf standards committee. I have no idea who is doing that and > how much such new e_machine registration may cost. > So far using EM_NONE is a hack to avoid bureaucracy. > Are dwarf relocation processor specific? > Then simple hack to elfutils/libdw to treat EM_NONE as X64 > should do the trick, right? > If that indeed works, we can tweak bpf backend to use EM_X86_64, > but then the danger that such .o file will be wrongly > recognized by elf utils. imo it's safer to keep it as EM_NONE > until real number is assigned, but even after it's assigned it > will take time to propagate that value. So for now I would try > to find a solution keeping EM_NONE hack. > What about hacking ELF binary in memory? 1. load the object into memory; 2. twist the machine code to EM_X86_64; 3. load it using elf_begin; 4. return the twested elf memory image using libdwfl's find_elf callback. Then libdw will recognise BPF's object file as a X86_64 object file. If required, relocation sections can also be twisted in this way. Should not very hard since we can only consider one relocation type. Then let's start thinking how to introduce EM_BPF. We can rely on the hacking until EM_BPF symbol reaches elfutils in perf. What do you think? Thank you. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web