Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #188871 > unrolled thread

codesearch across lines

Started bykamaraju kusumanchi <raju.mailinglists@gmail.com>
First post2017-11-11 05:20 +0100
Last post2017-11-12 16:20 +0100
Articles 9 — 5 participants

Back to article view | Back to linux.debian.user


Contents

  codesearch across lines kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2017-11-11 05:20 +0100
    Re: codesearch across lines Roberto C. Sánchez <roberto@debian.org> - 2017-11-11 14:10 +0100
      Re: codesearch across lines kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2017-11-11 15:30 +0100
        Re: codesearch across lines Curt <curty@free.fr> - 2017-11-11 18:10 +0100
          Re: codesearch across lines kamaraju kusumanchi <raju.mailinglists@gmail.com> - 2017-11-12 00:00 +0100
            Re: codesearch across lines Curt <curty@free.fr> - 2017-11-12 12:00 +0100
              Re: codesearch across lines The Wanderer <wanderer@fastmail.fm> - 2017-11-12 14:20 +0100
                Re: codesearch across lines Curt <curty@free.fr> - 2017-11-12 15:30 +0100
                  Re: codesearch across lines Richard Owlett <rowlett@cloud85.net> - 2017-11-12 16:20 +0100

#188871 — codesearch across lines

Fromkamaraju kusumanchi <raju.mailinglists@gmail.com>
Date2017-11-11 05:20 +0100
Subjectcodesearch across lines
Message-ID<uKvqW-4Jq-1@gated-at.bofh.it>
In codesearch.debian.net , is it possible to search for multiple words
that may occur across different lines and not necessarily on the same
line? For example, there are no results when I search for

pandas str filetype:python

which probably happens because it looks for lines that contain
'pandas' and 'str'. But what I would like to see is files where both
words are present (but not necessarily on the same line). Any ideas on
how to achieve that?

thanks
raju
-- 
Kamaraju S Kusumanchi | http://raju.shoutwiki.com/wiki/Blog

[toc] | [next] | [standalone]


#188876

FromRoberto C. Sánchez <roberto@debian.org>
Date2017-11-11 14:10 +0100
Message-ID<uKDHP-28K-5@gated-at.bofh.it>
In reply to#188871
On Fri, Nov 10, 2017 at 08:18:17PM -0800, kamaraju kusumanchi wrote:
> In codesearch.debian.net , is it possible to search for multiple words
> that may occur across different lines and not necessarily on the same
> line? For example, there are no results when I search for
> 
The FAQ for codesearch.d.n states that it supportes the RE2 regular
expression syntax, as described here:

https://github.com/google/re2/blob/master/doc/syntax.txt

Here is the section on available flags:

Flags:
i	case-insensitive (default false)
m	multi-line mode: «^» and «$» match begin/end line in addition to begin/end text (default false)
s	let «.» match «\n» (default false)
U	ungreedy: swap meaning of «x*» and «x*?», «x+» and «x+?», etc (default false)
Flag syntax is «xyz» (set) or «-xyz» (clear) or «xy-z» (set «xy», clear «z»).

Regards,

-Roberto

-- 
Roberto C. Sánchez

[toc] | [prev] | [next] | [standalone]


#188880

Fromkamaraju kusumanchi <raju.mailinglists@gmail.com>
Date2017-11-11 15:30 +0100
Message-ID<uKEXf-2U0-3@gated-at.bofh.it>
In reply to#188876
On Sat, Nov 11, 2017 at 5:01 AM, Roberto C. Sánchez <roberto@debian.org> wrote:
> On Fri, Nov 10, 2017 at 08:18:17PM -0800, kamaraju kusumanchi wrote:
>> In codesearch.debian.net , is it possible to search for multiple words
>> that may occur across different lines and not necessarily on the same
>> line? For example, there are no results when I search for
>>
> The FAQ for codesearch.d.n states that it supportes the RE2 regular
> expression syntax, as described here:
>
> https://github.com/google/re2/blob/master/doc/syntax.txt
>
> Here is the section on available flags:
>
> Flags:
> i       case-insensitive (default false)
> m       multi-line mode: «^» and «$» match begin/end line in addition to begin/end text (default false)
> s       let «.» match «\n» (default false)
> U       ungreedy: swap meaning of «x*» and «x*?», «x+» and «x+?», etc (default false)
> Flag syntax is «xyz» (set) or «-xyz» (clear) or «xy-z» (set «xy», clear «z»).
>

Thanks. How can I specify the flag in the searches? So, in my example, I tried

pandas str /m

But that query is not finishing.

-- 
Kamaraju S Kusumanchi | http://raju.shoutwiki.com/wiki/Blog

[toc] | [prev] | [next] | [standalone]


#188887

FromCurt <curty@free.fr>
Date2017-11-11 18:10 +0100
Message-ID<uKHs6-4FR-7@gated-at.bofh.it>
In reply to#188880
On 2017-11-11, kamaraju kusumanchi <raju.mailinglists@gmail.com> wrote:
>
> Thanks. How can I specify the flag in the searches? So, in my example, I tried
>
> pandas str /m
>
> But that query is not finishing.
>

'(?i) PanDas' seemed to work.

-- 
"A simpering Bambi narcissist and a thieving, fanatical Albanian dwarf."
Christopher Hitchens, commenting shortly after the nearly concurrent deaths 
of Lady Diana and Mother Theresa.

[toc] | [prev] | [next] | [standalone]


#188891

Fromkamaraju kusumanchi <raju.mailinglists@gmail.com>
Date2017-11-12 00:00 +0100
Message-ID<uKMUN-86K-3@gated-at.bofh.it>
In reply to#188887
On Sat, Nov 11, 2017 at 9:08 AM, Curt <curty@free.fr> wrote:
> On 2017-11-11, kamaraju kusumanchi <raju.mailinglists@gmail.com> wrote:
>>
>> Thanks. How can I specify the flag in the searches? So, in my example, I tried
>>
>> pandas str /m
>>
>> But that query is not finishing.
>>
>
> '(?i) PanDas' seemed to work.
>
Ok. That works and does case insensitive search. But the corresponding
multiline option does not seem to achieve what I want. I tried

(?m) pandas str

and it dies not give any results.

-- 
Kamaraju S Kusumanchi | http://raju.shoutwiki.com/wiki/Blog

[toc] | [prev] | [next] | [standalone]


#188903

FromCurt <curty@free.fr>
Date2017-11-12 12:00 +0100
Message-ID<uKY9z-7tU-13@gated-at.bofh.it>
In reply to#188891
On 2017-11-11, kamaraju kusumanchi <raju.mailinglists@gmail.com> wrote:
> On Sat, Nov 11, 2017 at 9:08 AM, Curt <curty@free.fr> wrote:
>> On 2017-11-11, kamaraju kusumanchi <raju.mailinglists@gmail.com> wrote:
>>>
>>> Thanks. How can I specify the flag in the searches? So, in my example, I tried
>>>
>>> pandas str /m
>>>
>>> But that query is not finishing.
>>>
>>
>> '(?i) PanDas' seemed to work.
>>
> Ok. That works and does case insensitive search. But the corresponding
> multiline option does not seem to achieve what I want. I tried
>
> (?m) pandas str
>
> and it dies not give any results.
>

It doesn't die for me. It gives no results, though.

This "works" (for an arbritrary meaning of "work" because I am not cut
out for regular expressions or even irregular expressions and I feel a
migraine coming on):

 (?m)(\W|^)panda.*str(\W|$)

At least there's results and that's always gratifying (in a way).

-- 
"If you want to build a ship, don’t herd people together to collect
wood and don’t assign them tasks and work, but rather teach them to
long for the endless immensity of the sea."
— Antoine de Saint-Exupéry 

[toc] | [prev] | [next] | [standalone]


#188906

FromThe Wanderer <wanderer@fastmail.fm>
Date2017-11-12 14:20 +0100
Message-ID<uL0l4-K6-9@gated-at.bofh.it>
In reply to#188903

[Multipart message — attachments visible in raw view] — view raw

On 2017-11-12 at 05:57, Curt wrote:

> On 2017-11-11, kamaraju kusumanchi <raju.mailinglists@gmail.com>
> wrote:
> 
>> On Sat, Nov 11, 2017 at 9:08 AM, Curt <curty@free.fr> wrote:

>>> '(?i) PanDas' seemed to work.
>>> 
>> Ok. That works and does case insensitive search. But the
>> corresponding multiline option does not seem to achieve what I
>> want. I tried
>> 
>> (?m) pandas str
>> 
>> and it dies not give any results.

That would be because this is most likely searching for the literal
sequence ' pandas str' somewhere in the document. (Vs., previously,
searching for lines in the document which contain that literal sequence.)

> It doesn't die for me. It gives no results, though.
> 
> This "works" (for an arbritrary meaning of "work" because I am not
> cut out for regular expressions or even irregular expressions and I
> feel a migraine coming on):
> 
> (?m)(\W|^)panda.*str(\W|$)

That would be expected to find only documents containing 'panda'
followed by 'str'. To also find ones which contain 'str' followed by
'pandas' (and add the missing 's' back in), you'd probably want:

(?m)(\W|^)(pandas.*str|str.*pandas)(\W|$)

I have not tested this, but I use similar '(a.*b|b.*a)' regexes on a
semi-regular basis for searching one of my own text archives.

(Also, I'm not sure the '\W' bits are needed, but I don't know the field
of what's-being-searched-for well enough to be certain about why those
may have been added.)

-- 
   The Wanderer

The reasonable man adapts himself to the world; the unreasonable one
persists in trying to adapt the world to himself. Therefore all
progress depends on the unreasonable man.         -- George Bernard Shaw

[toc] | [prev] | [next] | [standalone]


#188907

FromCurt <curty@free.fr>
Date2017-11-12 15:30 +0100
Message-ID<uL1qN-1lZ-3@gated-at.bofh.it>
In reply to#188906
On 2017-11-12, The Wanderer <wanderer@fastmail.fm> wrote:
>
>> (?m)(\W|^)panda.*str(\W|$)
>
> That would be expected to find only documents containing 'panda'
> followed by 'str'. To also find ones which contain 'str' followed by
> 'pandas' (and add the missing 's' back in), you'd probably want:
>
> (?m)(\W|^)(pandas.*str|str.*pandas)(\W|$)
>
> I have not tested this, but I use similar '(a.*b|b.*a)' regexes on a
> semi-regular basis for searching one of my own text archives.

I tried that, actually, following the same logic, or thought I did
(there might have been a typo somewhere) yet it produced *less* results,
but trying the formula again now it seems to "work" (although the regex
is pretty useless because it matches reams and reams of stuff because
'anda' 'panda' 'pandas' 'expandable' 'str' 'struct' 'instruction' 'castration',
etc. are all matched).

This produces two hits (from the same file) only:

 (?m)(\W|^)\bpanda\b.*\bstr\b|\bstr\b.*\bpanda\b(\W|$)

I'm not certain how you're supposed to construct the formula to only match
the literal strings (if literal is indeed the term) "panda" and "str".

I also have no idea what the expected results might look like. To add
insult insult to injury, I know nothing about regexes either.

;-)


> (Also, I'm not sure the '\W' bits are needed, but I don't know the field
> of what's-being-searched-for well enough to be certain about why those
> may have been added.)
>

[toc] | [prev] | [next] | [standalone]


#188908

FromRichard Owlett <rowlett@cloud85.net>
Date2017-11-12 16:20 +0100
Message-ID<uL2db-1Rn-3@gated-at.bofh.it>
In reply to#188907
On 11/12/2017 08:23 AM, Curt wrote:
> On 2017-11-12, The Wanderer <wanderer@fastmail.fm> wrote:
>>
>>> (?m)(\W|^)panda.*str(\W|$)
>>
>> That would be expected to find only documents containing 'panda'
>> followed by 'str'. To also find ones which contain 'str' followed by
>> 'pandas' (and add the missing 's' back in), you'd probably want:
>>
>> (?m)(\W|^)(pandas.*str|str.*pandas)(\W|$)
>>
>> I have not tested this, but I use similar '(a.*b|b.*a)' regexes on a
>> semi-regular basis for searching one of my own text archives.
>
> I tried that, actually, following the same logic, or thought I did
> (there might have been a typo somewhere) yet it produced *less* results,
> but trying the formula again now it seems to "work" (although the regex
> is pretty useless because it matches reams and reams of stuff because
> 'anda' 'panda' 'pandas' 'expandable' 'str' 'struct' 'instruction' 'castration',
> etc. are all matched).
>
> This produces two hits (from the same file) only:
>
>  (?m)(\W|^)\bpanda\b.*\bstr\b|\bstr\b.*\bpanda\b(\W|$)
>
> I'm not certain how you're supposed to construct the formula to only match
> the literal strings (if literal is indeed the term) "panda" and "str".
>
> I also have no idea what the expected results might look like. To add
> insult insult to injury, I know nothing about regexes either.
>
> ;-)
>
>
>> (Also, I'm not sure the '\W' bits are needed, but I don't know the field
>> of what's-being-searched-for well enough to be certain about why those
>> may have been added.)
>>
>
>

To paraphrase, "A code fragment is worth a thousand bytes of descriptive 
text" ;/

Post a half dozen lines of code followed by the desired output.
The posted lines should be < 30 characters to prevent confusion caused 
by line wrap problems when displayed. The example meed not be valid code 
in language used -- only character sequences are of interest.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web