Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1618484 > unrolled thread

[PATCH 1/3] ptr_ring: batch ring zeroing

Started by"Michael S. Tsirkin" <mst@redhat.com>
First post2017-04-07 08:00 +0200
Last post2017-04-08 14:20 +0200
Articles 4 — 2 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH 1/3] ptr_ring: batch ring zeroing "Michael S. Tsirkin" <mst@redhat.com> - 2017-04-07 08:00 +0200
    [PATCH 3/3] ptr_ring: support testing different batching sizes "Michael S. Tsirkin" <mst@redhat.com> - 2017-04-07 08:00 +0200
    [PATCH 2/3] ringtest: support test specific parameters "Michael S. Tsirkin" <mst@redhat.com> - 2017-04-07 08:00 +0200
    Re: [PATCH 1/3] ptr_ring: batch ring zeroing Jesper Dangaard Brouer <brouer@redhat.com> - 2017-04-08 14:20 +0200

#1618484 — [PATCH 1/3] ptr_ring: batch ring zeroing

From"Michael S. Tsirkin" <mst@redhat.com>
Date2017-04-07 08:00 +0200
Subject[PATCH 1/3] ptr_ring: batch ring zeroing
Message-ID<ttv69-31Q-1@gated-at.bofh.it>
A known weakness in ptr_ring design is that it does not handle well the
situation when ring is almost full: as entries are consumed they are
immediately used again by the producer, so consumer and producer are
writing to a shared cache line.

To fix this, add batching to consume calls: as entries are
consumed do not write NULL into the ring until we get
a multiple (in current implementation 2x) of cache lines
away from the producer. At that point, write them all out.

We do the write out in the reverse order to keep
producer from sharing cache with consumer for as long
as possible.

Writeout also triggers when ring wraps around - there's
no special reason to do this but it helps keep the code
a bit simpler.

What should we do if getting away from producer by 2 cache lines
would mean we are keeping the ring moe than half empty?
Maybe we should reduce the batching in this case,
current patch simply reduces the batching.

Notes:
- it is no longer true that a call to consume guarantees
  that the following call to produce will succeed.
  No users seem to assume that.
- batching can also in theory reduce the signalling rate:
  users that would previously send interrups to the producer
  to wake it up after consuming each entry would now only
  need to do this once in a batch.
  Doing this would be easy by returning a flag to the caller.
  No users seem to do signalling on consume yet so this was not
  implemented yet.

Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
---

Jason, I am curious whether the following gives you some of
the performance boost that you see with vhost batching
patches. Is vhost batching on top still helpful?

 include/linux/ptr_ring.h | 63 +++++++++++++++++++++++++++++++++++++++++-------
 1 file changed, 54 insertions(+), 9 deletions(-)

diff --git a/include/linux/ptr_ring.h b/include/linux/ptr_ring.h
index 6c70444..6b2e0dd 100644
--- a/include/linux/ptr_ring.h
+++ b/include/linux/ptr_ring.h
@@ -34,11 +34,13 @@
 struct ptr_ring {
 	int producer ____cacheline_aligned_in_smp;
 	spinlock_t producer_lock;
-	int consumer ____cacheline_aligned_in_smp;
+	int consumer_head ____cacheline_aligned_in_smp; /* next valid entry */
+	int consumer_tail; /* next entry to invalidate */
 	spinlock_t consumer_lock;
 	/* Shared consumer/producer data */
 	/* Read-only by both the producer and the consumer */
 	int size ____cacheline_aligned_in_smp; /* max entries in queue */
+	int batch; /* number of entries to consume in a batch */
 	void **queue;
 };
 
@@ -170,7 +172,7 @@ static inline int ptr_ring_produce_bh(struct ptr_ring *r, void *ptr)
 static inline void *__ptr_ring_peek(struct ptr_ring *r)
 {
 	if (likely(r->size))
-		return r->queue[r->consumer];
+		return r->queue[r->consumer_head];
 	return NULL;
 }
 
@@ -231,9 +233,38 @@ static inline bool ptr_ring_empty_bh(struct ptr_ring *r)
 /* Must only be called after __ptr_ring_peek returned !NULL */
 static inline void __ptr_ring_discard_one(struct ptr_ring *r)
 {
-	r->queue[r->consumer++] = NULL;
-	if (unlikely(r->consumer >= r->size))
-		r->consumer = 0;
+	/* Fundamentally, what we want to do is update consumer
+	 * index and zero out the entry so producer can reuse it.
+	 * Doing it naively at each consume would be as simple as:
+	 *       r->queue[r->consumer++] = NULL;
+	 *       if (unlikely(r->consumer >= r->size))
+	 *               r->consumer = 0;
+	 * but that is suboptimal when the ring is full as producer is writing
+	 * out new entries in the same cache line.  Defer these updates until a
+	 * batch of entries has been consumed.
+	 */
+	int head = r->consumer_head++;
+
+	/* Once we have processed enough entries invalidate them in
+	 * the ring all at once so producer can reuse their space in the ring.
+	 * We also do this when we reach end of the ring - not mandatory
+	 * but helps keep the implementation simple.
+	 */
+	if (unlikely(r->consumer_head - r->consumer_tail >= r->batch ||
+		     r->consumer_head >= r->size)) {
+		/* Zero out entries in the reverse order: this way we touch the
+		 * cache line that producer might currently be reading the last;
+		 * producer won't make progress and touch other cache lines
+		 * besides the first one until we write out all entries.
+		 */
+		while (likely(head >= r->consumer_tail))
+			r->queue[head--] = NULL;
+		r->consumer_tail = r->consumer_head;
+	}
+	if (unlikely(r->consumer_head >= r->size)) {
+		r->consumer_head = 0;
+		r->consumer_tail = 0;
+	}
 }
 
 static inline void *__ptr_ring_consume(struct ptr_ring *r)
@@ -345,14 +376,27 @@ static inline void **__ptr_ring_init_queue_alloc(int size, gfp_t gfp)
 	return kzalloc(ALIGN(size * sizeof(void *), SMP_CACHE_BYTES), gfp);
 }
 
+static inline void __ptr_ring_set_size(struct ptr_ring *r, int size)
+{
+	r->size = size;
+	r->batch = SMP_CACHE_BYTES * 2 / sizeof(*(r->queue));
+	/* We need to set batch at least to 1 to make logic
+	 * in __ptr_ring_discard_one work correctly.
+	 * Batching too much (because ring is small) would cause a lot of
+	 * burstiness. Needs tuning, for now disable batching.
+	 */
+	if (r->batch > r->size / 2 || !r->batch)
+		r->batch = 1;
+}
+
 static inline int ptr_ring_init(struct ptr_ring *r, int size, gfp_t gfp)
 {
 	r->queue = __ptr_ring_init_queue_alloc(size, gfp);
 	if (!r->queue)
 		return -ENOMEM;
 
-	r->size = size;
-	r->producer = r->consumer = 0;
+	__ptr_ring_set_size(r, size);
+	r->producer = r->consumer_head = r->consumer_tail = 0;
 	spin_lock_init(&r->producer_lock);
 	spin_lock_init(&r->consumer_lock);
 
@@ -373,9 +417,10 @@ static inline void **__ptr_ring_swap_queue(struct ptr_ring *r, void **queue,
 		else if (destroy)
 			destroy(ptr);
 
-	r->size = size;
+	__ptr_ring_set_size(r, size);
 	r->producer = producer;
-	r->consumer = 0;
+	r->consumer_head = 0;
+	r->consumer_tail = 0;
 	old = r->queue;
 	r->queue = queue;
 
-- 
MST

[toc] | [next] | [standalone]


#1618485 — [PATCH 3/3] ptr_ring: support testing different batching sizes

From"Michael S. Tsirkin" <mst@redhat.com>
Date2017-04-07 08:00 +0200
Subject[PATCH 3/3] ptr_ring: support testing different batching sizes
Message-ID<ttv69-31Q-3@gated-at.bofh.it>
In reply to#1618484
Use the param flag for that.

Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
---
 tools/virtio/ringtest/ptr_ring.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/tools/virtio/ringtest/ptr_ring.c b/tools/virtio/ringtest/ptr_ring.c
index 635b07b..7b22f1b 100644
--- a/tools/virtio/ringtest/ptr_ring.c
+++ b/tools/virtio/ringtest/ptr_ring.c
@@ -97,6 +97,9 @@ void alloc_ring(void)
 {
 	int ret = ptr_ring_init(&array, ring_size, 0);
 	assert(!ret);
+	/* Hacky way to poke at ring internals. Useful for testing though. */
+	if (param)
+		array.batch = param;
 }
 
 /* guest side */
-- 
MST

[toc] | [prev] | [next] | [standalone]


#1618487 — [PATCH 2/3] ringtest: support test specific parameters

From"Michael S. Tsirkin" <mst@redhat.com>
Date2017-04-07 08:00 +0200
Subject[PATCH 2/3] ringtest: support test specific parameters
Message-ID<ttv69-31Q-9@gated-at.bofh.it>
In reply to#1618484
Add a new flag for passing test-specific parameters.

Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
---
 tools/virtio/ringtest/main.c | 13 +++++++++++++
 tools/virtio/ringtest/main.h |  2 ++
 2 files changed, 15 insertions(+)

diff --git a/tools/virtio/ringtest/main.c b/tools/virtio/ringtest/main.c
index f31353f..022ae95 100644
--- a/tools/virtio/ringtest/main.c
+++ b/tools/virtio/ringtest/main.c
@@ -20,6 +20,7 @@
 int runcycles = 10000000;
 int max_outstanding = INT_MAX;
 int batch = 1;
+int param = 0;
 
 bool do_sleep = false;
 bool do_relax = false;
@@ -247,6 +248,11 @@ static const struct option longopts[] = {
 		.val = 'b',
 	},
 	{
+		.name = "param",
+		.has_arg = required_argument,
+		.val = 'p',
+	},
+	{
 		.name = "sleep",
 		.has_arg = no_argument,
 		.val = 's',
@@ -274,6 +280,7 @@ static void help(void)
 		" [--run-cycles C (default: %d)]"
 		" [--batch b]"
 		" [--outstanding o]"
+		" [--param p]"
 		" [--sleep]"
 		" [--relax]"
 		" [--exit]"
@@ -328,6 +335,12 @@ int main(int argc, char **argv)
 			assert(c > 0 && c < INT_MAX);
 			max_outstanding = c;
 			break;
+		case 'p':
+			c = strtol(optarg, &endptr, 0);
+			assert(!*endptr);
+			assert(c > 0 && c < INT_MAX);
+			param = c;
+			break;
 		case 'b':
 			c = strtol(optarg, &endptr, 0);
 			assert(!*endptr);
diff --git a/tools/virtio/ringtest/main.h b/tools/virtio/ringtest/main.h
index 14142fa..90b0133 100644
--- a/tools/virtio/ringtest/main.h
+++ b/tools/virtio/ringtest/main.h
@@ -10,6 +10,8 @@
 
 #include <stdbool.h>
 
+extern int param;
+
 extern bool do_exit;
 
 #if defined(__x86_64__) || defined(__i386__)
-- 
MST

[toc] | [prev] | [next] | [standalone]


#1619270

FromJesper Dangaard Brouer <brouer@redhat.com>
Date2017-04-08 14:20 +0200
Message-ID<ttXvr-56J-1@gated-at.bofh.it>
In reply to#1618484
On Fri, 7 Apr 2017 08:49:57 +0300
"Michael S. Tsirkin" <mst@redhat.com> wrote:

> A known weakness in ptr_ring design is that it does not handle well the
> situation when ring is almost full: as entries are consumed they are
> immediately used again by the producer, so consumer and producer are
> writing to a shared cache line.
> 
> To fix this, add batching to consume calls: as entries are
> consumed do not write NULL into the ring until we get
> a multiple (in current implementation 2x) of cache lines
> away from the producer. At that point, write them all out.
> 
> We do the write out in the reverse order to keep
> producer from sharing cache with consumer for as long
> as possible.
> 
> Writeout also triggers when ring wraps around - there's
> no special reason to do this but it helps keep the code
> a bit simpler.
> 
> What should we do if getting away from producer by 2 cache lines
> would mean we are keeping the ring moe than half empty?
> Maybe we should reduce the batching in this case,
> current patch simply reduces the batching.
> 
> Notes:
> - it is no longer true that a call to consume guarantees
>   that the following call to produce will succeed.
>   No users seem to assume that.
> - batching can also in theory reduce the signalling rate:
>   users that would previously send interrups to the producer
>   to wake it up after consuming each entry would now only
>   need to do this once in a batch.
>   Doing this would be easy by returning a flag to the caller.
>   No users seem to do signalling on consume yet so this was not
>   implemented yet.
> 
> Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
> ---
> 
> Jason, I am curious whether the following gives you some of
> the performance boost that you see with vhost batching
> patches. Is vhost batching on top still helpful?
> 
>  include/linux/ptr_ring.h | 63 +++++++++++++++++++++++++++++++++++++++++-------
>  1 file changed, 54 insertions(+), 9 deletions(-)
> 
> diff --git a/include/linux/ptr_ring.h b/include/linux/ptr_ring.h
> index 6c70444..6b2e0dd 100644
> --- a/include/linux/ptr_ring.h
> +++ b/include/linux/ptr_ring.h
> @@ -34,11 +34,13 @@
>  struct ptr_ring {
>  	int producer ____cacheline_aligned_in_smp;
>  	spinlock_t producer_lock;
> -	int consumer ____cacheline_aligned_in_smp;
> +	int consumer_head ____cacheline_aligned_in_smp; /* next valid entry */
> +	int consumer_tail; /* next entry to invalidate */
>  	spinlock_t consumer_lock;
>  	/* Shared consumer/producer data */
>  	/* Read-only by both the producer and the consumer */
>  	int size ____cacheline_aligned_in_smp; /* max entries in queue */
> +	int batch; /* number of entries to consume in a batch */
>  	void **queue;
>  };
>  
> @@ -170,7 +172,7 @@ static inline int ptr_ring_produce_bh(struct ptr_ring *r, void *ptr)
>  static inline void *__ptr_ring_peek(struct ptr_ring *r)
>  {
>  	if (likely(r->size))
> -		return r->queue[r->consumer];
> +		return r->queue[r->consumer_head];
>  	return NULL;
>  }
>  
> @@ -231,9 +233,38 @@ static inline bool ptr_ring_empty_bh(struct ptr_ring *r)
>  /* Must only be called after __ptr_ring_peek returned !NULL */
>  static inline void __ptr_ring_discard_one(struct ptr_ring *r)
>  {
> -	r->queue[r->consumer++] = NULL;
> -	if (unlikely(r->consumer >= r->size))
> -		r->consumer = 0;
> +	/* Fundamentally, what we want to do is update consumer
> +	 * index and zero out the entry so producer can reuse it.
> +	 * Doing it naively at each consume would be as simple as:
> +	 *       r->queue[r->consumer++] = NULL;
> +	 *       if (unlikely(r->consumer >= r->size))
> +	 *               r->consumer = 0;
> +	 * but that is suboptimal when the ring is full as producer is writing
> +	 * out new entries in the same cache line.  Defer these updates until a
> +	 * batch of entries has been consumed.
> +	 */
> +	int head = r->consumer_head++;
> +
> +	/* Once we have processed enough entries invalidate them in
> +	 * the ring all at once so producer can reuse their space in the ring.
> +	 * We also do this when we reach end of the ring - not mandatory
> +	 * but helps keep the implementation simple.
> +	 */
> +	if (unlikely(r->consumer_head - r->consumer_tail >= r->batch ||
> +		     r->consumer_head >= r->size)) {
> +		/* Zero out entries in the reverse order: this way we touch the
> +		 * cache line that producer might currently be reading the last;
> +		 * producer won't make progress and touch other cache lines
> +		 * besides the first one until we write out all entries.
> +		 */
> +		while (likely(head >= r->consumer_tail))
> +			r->queue[head--] = NULL;
> +		r->consumer_tail = r->consumer_head;
> +	}
> +	if (unlikely(r->consumer_head >= r->size)) {
> +		r->consumer_head = 0;
> +		r->consumer_tail = 0;
> +	}
>  }

I love this idea.  Reviewed and discussed the idea in-person with MST
during netdevconf[1] at this laptop.  I promised I will also run it
through my micro-benchmarking[2] once I return home (hint ptr_ring gets
used in network stack as skb_array).

Reviewed-by: Jesper Dangaard Brouer <brouer@redhat.com>

[1] http://netdevconf.org/2.1/
[2] https://github.com/netoptimizer/prototype-kernel/blob/master/kernel/lib/skb_array_bench01.c
-- 
Best regards,
  Jesper Dangaard Brouer
  MSc.CS, Principal Kernel Engineer at Red Hat
  LinkedIn: http://www.linkedin.com/in/brouer

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web