Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1659855 > unrolled thread

[PATCH 0/3] Bitmap optimisations

Started byMatthew Wilcox <willy@infradead.org>
First post2017-06-07 16:40 +0200
Last post2017-06-07 23:40 +0200
Articles 9 — 3 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH 0/3] Bitmap optimisations Matthew Wilcox <willy@infradead.org> - 2017-06-07 16:40 +0200
    [PATCH 1/3] bitmap: Optimise bitmap_set and bitmap_clear of a single bit Matthew Wilcox <willy@infradead.org> - 2017-06-07 16:40 +0200
    [PATCH 3/3] bitmap: Use memcmp optimisation in more situations Matthew Wilcox <willy@infradead.org> - 2017-06-07 16:40 +0200
      Re: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations Andy Shevchenko <andy.shevchenko@gmail.com> - 2017-06-08 03:50 +0200
        Re: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations Matthew Wilcox <willy@infradead.org> - 2017-06-08 05:00 +0200
          Re: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations Andy Shevchenko <andy.shevchenko@gmail.com> - 2017-06-08 14:40 +0200
            Re: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations Rasmus Villemoes <linux@rasmusvillemoes.dk> - 2017-06-08 15:50 +0200
              Re: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations Andy Shevchenko <andy.shevchenko@gmail.com> - 2017-06-08 16:50 +0200
    Re: [PATCH 0/3] Bitmap optimisations Rasmus Villemoes <linux@rasmusvillemoes.dk> - 2017-06-07 23:40 +0200

#1659855 — [PATCH 0/3] Bitmap optimisations

FromMatthew Wilcox <willy@infradead.org>
Date2017-06-07 16:40 +0200
Subject[PATCH 0/3] Bitmap optimisations
Message-ID<tPK8a-7Mn-21@gated-at.bofh.it>
From: Matthew Wilcox <mawilcox@microsoft.com>

These three bitmap patches use more efficient specialisations when the
compiler can figure out that it's safe to do so.

Matthew Wilcox (3):
  bitmap: Optimise bitmap_set and bitmap_clear of a single bit
  Turn bitmap_set and bitmap_clear into memset when possible
  bitmap: Use memcmp optimisation in more situations

 include/linux/bitmap.h | 33 +++++++++++++++++++++++++++------
 lib/bitmap.c           |  8 ++++----
 2 files changed, 31 insertions(+), 10 deletions(-)

-- 
2.11.0

[toc] | [next] | [standalone]


#1659859 — [PATCH 1/3] bitmap: Optimise bitmap_set and bitmap_clear of a single bit

FromMatthew Wilcox <willy@infradead.org>
Date2017-06-07 16:40 +0200
Subject[PATCH 1/3] bitmap: Optimise bitmap_set and bitmap_clear of a single bit
Message-ID<tPKhQ-7Rr-29@gated-at.bofh.it>
In reply to#1659855
From: Matthew Wilcox <mawilcox@microsoft.com>

We have eight users calling bitmap_clear for a single bit and seventeen
calling bitmap_set for a single bit.  Rather than fix all of them to call
__clear_bit or __set_bit, turn bitmap_clear and bitmap_set into inline
functions and make this special case efficient.

Signed-off-by: Matthew Wilcox <mawilcox@microsoft.com>
---
 include/linux/bitmap.h | 23 ++++++++++++++++++++---
 lib/bitmap.c           |  8 ++++----
 2 files changed, 24 insertions(+), 7 deletions(-)

diff --git a/include/linux/bitmap.h b/include/linux/bitmap.h
index 3b77588a9360..4e0f0c8167af 100644
--- a/include/linux/bitmap.h
+++ b/include/linux/bitmap.h
@@ -112,9 +112,8 @@ extern int __bitmap_intersects(const unsigned long *bitmap1,
 extern int __bitmap_subset(const unsigned long *bitmap1,
 			const unsigned long *bitmap2, unsigned int nbits);
 extern int __bitmap_weight(const unsigned long *bitmap, unsigned int nbits);
-
-extern void bitmap_set(unsigned long *map, unsigned int start, int len);
-extern void bitmap_clear(unsigned long *map, unsigned int start, int len);
+extern void __bitmap_set(unsigned long *map, unsigned int start, int len);
+extern void __bitmap_clear(unsigned long *map, unsigned int start, int len);
 
 extern unsigned long bitmap_find_next_zero_area_off(unsigned long *map,
 						    unsigned long size,
@@ -315,6 +314,24 @@ static __always_inline int bitmap_weight(const unsigned long *src, unsigned int
 	return __bitmap_weight(src, nbits);
 }
 
+static __always_inline void bitmap_set(unsigned long *map, unsigned int start,
+		unsigned int nbits)
+{
+	if (__builtin_constant_p(nbits) && nbits == 1)
+		__set_bit(start, map);
+	else
+		__bitmap_set(map, start, nbits);
+}
+
+static __always_inline void bitmap_clear(unsigned long *map, unsigned int start,
+		unsigned int nbits)
+{
+	if (__builtin_constant_p(nbits) && nbits == 1)
+		__clear_bit(start, map);
+	else
+		__bitmap_clear(map, start, nbits);
+}
+
 static inline void bitmap_shift_right(unsigned long *dst, const unsigned long *src,
 				unsigned int shift, int nbits)
 {
diff --git a/lib/bitmap.c b/lib/bitmap.c
index 0b66f0e5eb6b..6e354a6d0dbb 100644
--- a/lib/bitmap.c
+++ b/lib/bitmap.c
@@ -251,7 +251,7 @@ int __bitmap_weight(const unsigned long *bitmap, unsigned int bits)
 }
 EXPORT_SYMBOL(__bitmap_weight);
 
-void bitmap_set(unsigned long *map, unsigned int start, int len)
+void __bitmap_set(unsigned long *map, unsigned int start, int len)
 {
 	unsigned long *p = map + BIT_WORD(start);
 	const unsigned int size = start + len;
@@ -270,9 +270,9 @@ void bitmap_set(unsigned long *map, unsigned int start, int len)
 		*p |= mask_to_set;
 	}
 }
-EXPORT_SYMBOL(bitmap_set);
+EXPORT_SYMBOL(__bitmap_set);
 
-void bitmap_clear(unsigned long *map, unsigned int start, int len)
+void __bitmap_clear(unsigned long *map, unsigned int start, int len)
 {
 	unsigned long *p = map + BIT_WORD(start);
 	const unsigned int size = start + len;
@@ -291,7 +291,7 @@ void bitmap_clear(unsigned long *map, unsigned int start, int len)
 		*p &= ~mask_to_clear;
 	}
 }
-EXPORT_SYMBOL(bitmap_clear);
+EXPORT_SYMBOL(__bitmap_clear);
 
 /**
  * bitmap_find_next_zero_area_off - find a contiguous aligned zero area
-- 
2.11.0

[toc] | [prev] | [next] | [standalone]


#1659862 — [PATCH 3/3] bitmap: Use memcmp optimisation in more situations

FromMatthew Wilcox <willy@infradead.org>
Date2017-06-07 16:40 +0200
Subject[PATCH 3/3] bitmap: Use memcmp optimisation in more situations
Message-ID<tPKhR-7Rr-43@gated-at.bofh.it>
In reply to#1659855
From: Matthew Wilcox <mawilcox@microsoft.com>

Commit 7dd968163f ("bitmap: bitmap_equal memcmp optimization") was
rather more restrictive than necessary; we can use memcmp() to implement
bitmap_equal() as long as the number of bits can be proved to be a
multiple of 8.  And architectures other than s390 may be able to make
good use of this optimisation.

Signed-off-by: Matthew Wilcox <mawilcox@microsoft.com>
---
 include/linux/bitmap.h | 4 +---
 1 file changed, 1 insertion(+), 3 deletions(-)

diff --git a/include/linux/bitmap.h b/include/linux/bitmap.h
index 0b3e4452b054..26244e0098f0 100644
--- a/include/linux/bitmap.h
+++ b/include/linux/bitmap.h
@@ -266,10 +266,8 @@ static inline int bitmap_equal(const unsigned long *src1,
 {
 	if (small_const_nbits(nbits))
 		return !((*src1 ^ *src2) & BITMAP_LAST_WORD_MASK(nbits));
-#ifdef CONFIG_S390
-	if (__builtin_constant_p(nbits) && (nbits % BITS_PER_LONG) == 0)
+	if (__builtin_constant_p(nbits & 7) && IS_ALIGNED(nbits, 8))
 		return !memcmp(src1, src2, nbits / 8);
-#endif
 	return __bitmap_equal(src1, src2, nbits);
 }
 
-- 
2.11.0

[toc] | [prev] | [next] | [standalone]


#1660649 — Re: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations

FromAndy Shevchenko <andy.shevchenko@gmail.com>
Date2017-06-08 03:50 +0200
SubjectRe: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations
Message-ID<tPUKe-65e-7@gated-at.bofh.it>
In reply to#1659862
On Wed, Jun 7, 2017 at 5:29 PM, Matthew Wilcox <willy@infradead.org> wrote:

> Commit 7dd968163f ("bitmap: bitmap_equal memcmp optimization") was
> rather more restrictive than necessary; we can use memcmp() to implement
> bitmap_equal() as long as the number of bits can be proved to be a
> multiple of 8.  And architectures other than s390 may be able to make
> good use of this optimisation.

> -       if (__builtin_constant_p(nbits) && (nbits % BITS_PER_LONG) == 0)
> +       if (__builtin_constant_p(nbits & 7) && IS_ALIGNED(nbits, 8))
>                 return !memcmp(src1, src2, nbits / 8);

I'm not sure this is a fully correct change.
What exactly ' & 7' part does?
For me looks like you may just drop it.

-- 
With Best Regards,
Andy Shevchenko

[toc] | [prev] | [next] | [standalone]


#1660694 — Re: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations

FromMatthew Wilcox <willy@infradead.org>
Date2017-06-08 05:00 +0200
SubjectRe: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations
Message-ID<tPVPX-6Nn-9@gated-at.bofh.it>
In reply to#1660649
On Thu, Jun 08, 2017 at 04:48:04AM +0300, Andy Shevchenko wrote:
> > Commit 7dd968163f ("bitmap: bitmap_equal memcmp optimization") was
> > rather more restrictive than necessary; we can use memcmp() to implement
> > bitmap_equal() as long as the number of bits can be proved to be a
> > multiple of 8.  And architectures other than s390 may be able to make
> > good use of this optimisation.
> 
> > -       if (__builtin_constant_p(nbits) && (nbits % BITS_PER_LONG) == 0)
> > +       if (__builtin_constant_p(nbits & 7) && IS_ALIGNED(nbits, 8))
> >                 return !memcmp(src1, src2, nbits / 8);
> 
> I'm not sure this is a fully correct change.
> What exactly ' & 7' part does?
> For me looks like you may just drop it.

We only need to know if the bottom 3 bits are 0 to apply this optimisation.
For example, if we have a user which does this:

	nbits = 8;
	if (argle)
		nbits += 8;
	if (bitmap_equal(ptr1, ptr2, nbits))
		blah();

then we can use memcmp() because gcc can deduce that the bottom 3 bits
are never set (try it!  it works!).  We don't need nbits as a whole to
be const.

[toc] | [prev] | [next] | [standalone]


#1661130 — Re: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations

FromAndy Shevchenko <andy.shevchenko@gmail.com>
Date2017-06-08 14:40 +0200
SubjectRe: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations
Message-ID<tQ4Tg-4hK-37@gated-at.bofh.it>
In reply to#1660694
On Thu, Jun 8, 2017 at 5:55 AM, Matthew Wilcox <willy@infradead.org> wrote:
> On Thu, Jun 08, 2017 at 04:48:04AM +0300, Andy Shevchenko wrote:
>> > Commit 7dd968163f ("bitmap: bitmap_equal memcmp optimization") was
>> > rather more restrictive than necessary; we can use memcmp() to implement
>> > bitmap_equal() as long as the number of bits can be proved to be a
>> > multiple of 8.  And architectures other than s390 may be able to make
>> > good use of this optimisation.
>>
>> > -       if (__builtin_constant_p(nbits) && (nbits % BITS_PER_LONG) == 0)
>> > +       if (__builtin_constant_p(nbits & 7) && IS_ALIGNED(nbits, 8))
>> >                 return !memcmp(src1, src2, nbits / 8);
>>
>> I'm not sure this is a fully correct change.
>> What exactly ' & 7' part does?
>> For me looks like you may just drop it.
>
> We only need to know if the bottom 3 bits are 0 to apply this optimisation.
> For example, if we have a user which does this:
>
>         nbits = 8;
>         if (argle)
>                 nbits += 8;
>         if (bitmap_equal(ptr1, ptr2, nbits))
>                 blah();
>
> then we can use memcmp() because gcc can deduce that the bottom 3 bits
> are never set (try it!  it works!).  We don't need nbits as a whole to
> be const.

What I'm talking about is that by my opinion the both below are equivalent.
__builtin_constant_p(nbits)
__builtin_constant_p(nbits & 7)

Thus, again, what & 7 does there?

-- 
With Best Regards,
Andy Shevchenko

[toc] | [prev] | [next] | [standalone]


#1661212 — Re: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations

FromRasmus Villemoes <linux@rasmusvillemoes.dk>
Date2017-06-08 15:50 +0200
SubjectRe: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations
Message-ID<tQ5YZ-4WL-3@gated-at.bofh.it>
In reply to#1661130
On 8 June 2017 at 14:31, Andy Shevchenko <andy.shevchenko@gmail.com> wrote:
> On Thu, Jun 8, 2017 at 5:55 AM, Matthew Wilcox <willy@infradead.org> wrote:
>> We only need to know if the bottom 3 bits are 0 to apply this optimisation.
>> For example, if we have a user which does this:
>>
>>         nbits = 8;
>>         if (argle)
>>                 nbits += 8;
>>         if (bitmap_equal(ptr1, ptr2, nbits))
>>                 blah();
>>
>> then we can use memcmp() because gcc can deduce that the bottom 3 bits
>> are never set (try it!  it works!).  We don't need nbits as a whole to
>> be const.
>
> What I'm talking about is that by my opinion the both below are equivalent.
> __builtin_constant_p(nbits)
> __builtin_constant_p(nbits & 7)

They are not. Read Matthew's example again. Assuming that argle is
something non-constant (maybe an argument to the function), the value
of nbits at the time of the bitmap_equal call is _not_ a
compile-time-constant. However, if the compiler is smart (which at
least some versions of gcc are), the compiler may deduce that nbits is
either 8 or 16; there really are no other options. Hence it _is_
statically known that nbits is divisible by 8, so the expression
nbits&7 _is_ compile-time constant (0), so gcc can change the
bitmap_equal call to a memcmp call.

(It may then either pass a run-time value of nbits>>3 and emit a
single memcmp call, or it may decide to unroll the two options,
creating two memcmp calls with 1 and 2 as compile-time arguments;
these may or may not then in turn be "inlined" to code doing roughly
*(u8*)p1 == *(u8*)p2 and similarly for u16 casts).

[toc] | [prev] | [next] | [standalone]


#1661350 — Re: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations

FromAndy Shevchenko <andy.shevchenko@gmail.com>
Date2017-06-08 16:50 +0200
SubjectRe: [PATCH 3/3] bitmap: Use memcmp optimisation in more situations
Message-ID<tQ6V4-5vT-23@gated-at.bofh.it>
In reply to#1661212
On Thu, Jun 8, 2017 at 4:43 PM, Rasmus Villemoes
<linux@rasmusvillemoes.dk> wrote:
> On 8 June 2017 at 14:31, Andy Shevchenko <andy.shevchenko@gmail.com> wrote:
>> On Thu, Jun 8, 2017 at 5:55 AM, Matthew Wilcox <willy@infradead.org> wrote:
>>> We only need to know if the bottom 3 bits are 0 to apply this optimisation.
>>> For example, if we have a user which does this:
>>>
>>>         nbits = 8;
>>>         if (argle)
>>>                 nbits += 8;
>>>         if (bitmap_equal(ptr1, ptr2, nbits))
>>>                 blah();
>>>
>>> then we can use memcmp() because gcc can deduce that the bottom 3 bits
>>> are never set (try it!  it works!).  We don't need nbits as a whole to
>>> be const.
>>
>> What I'm talking about is that by my opinion the both below are equivalent.
>> __builtin_constant_p(nbits)
>> __builtin_constant_p(nbits & 7)
>
> They are not. Read Matthew's example again. Assuming that argle is
> something non-constant (maybe an argument to the function), the value
> of nbits at the time of the bitmap_equal call is _not_ a
> compile-time-constant. However, if the compiler is smart (which at
> least some versions of gcc are), the compiler may deduce that nbits is
> either 8 or 16; there really are no other options. Hence it _is_
> statically known that nbits is divisible by 8, so the expression
> nbits&7 _is_ compile-time constant (0), so gcc can change the
> bitmap_equal call to a memcmp call.

Yeah, thanks for detailed explanation.
So, basically what we do, we consider
1. 3 LSBs _is_ constant, *and*
2. They are equal to 0.

> (It may then either pass a run-time value of nbits>>3 and emit a
> single memcmp call, or it may decide to unroll the two options,
> creating two memcmp calls with 1 and 2 as compile-time arguments;
> these may or may not then in turn be "inlined" to code doing roughly
> *(u8*)p1 == *(u8*)p2 and similarly for u16 casts).

-- 
With Best Regards,
Andy Shevchenko

[toc] | [prev] | [next] | [standalone]


#1660261

FromRasmus Villemoes <linux@rasmusvillemoes.dk>
Date2017-06-07 23:40 +0200
Message-ID<tPQQh-3EL-17@gated-at.bofh.it>
In reply to#1659855
On Wed, Jun 07 2017, Matthew Wilcox <willy@infradead.org> wrote:

> From: Matthew Wilcox <mawilcox@microsoft.com>
>
> These three bitmap patches use more efficient specialisations when the
> compiler can figure out that it's safe to do so.
>

With the pointer arithmetic fixed,

Acked-by: Rasmus Villemoes <linux@rasmusvillemoes.dk>

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web