Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1220598 > unrolled thread
| Started by | Joonsoo Kim <js1304@gmail.com> |
|---|---|
| First post | 2015-09-08 10:30 +0200 |
| Last post | 2015-09-21 13:00 +0200 |
| Articles | 4 — 3 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH 12/12] mm, page_alloc: Only enforce watermarks for order-0 allocations Joonsoo Kim <js1304@gmail.com> - 2015-09-08 10:30 +0200
Re: [PATCH 12/12] mm, page_alloc: Only enforce watermarks for order-0 allocations Mel Gorman <mgorman@techsingularity.net> - 2015-09-09 15:00 +0200
Re: [PATCH 12/12] mm, page_alloc: Only enforce watermarks for order-0 allocations Joonsoo Kim <iamjoonsoo.kim@lge.com> - 2015-09-18 09:00 +0200
Re: [PATCH 12/12] mm, page_alloc: Only enforce watermarks for order-0 allocations Mel Gorman <mgorman@techsingularity.net> - 2015-09-21 13:00 +0200
| From | Joonsoo Kim <js1304@gmail.com> |
|---|---|
| Date | 2015-09-08 10:30 +0200 |
| Subject | Re: [PATCH 12/12] mm, page_alloc: Only enforce watermarks for order-0 allocations |
| Message-ID | <q6mrU-7dD-19@gated-at.bofh.it> |
2015-08-24 21:30 GMT+09:00 Mel Gorman <mgorman@techsingularity.net>:
> The primary purpose of watermarks is to ensure that reclaim can always
> make forward progress in PF_MEMALLOC context (kswapd and direct reclaim).
> These assume that order-0 allocations are all that is necessary for
> forward progress.
>
> High-order watermarks serve a different purpose. Kswapd had no high-order
> awareness before they were introduced (https://lkml.org/lkml/2004/9/5/9).
> This was particularly important when there were high-order atomic requests.
> The watermarks both gave kswapd awareness and made a reserve for those
> atomic requests.
>
> There are two important side-effects of this. The most important is that
> a non-atomic high-order request can fail even though free pages are available
> and the order-0 watermarks are ok. The second is that high-order watermark
> checks are expensive as the free list counts up to the requested order must
> be examined.
>
> With the introduction of MIGRATE_HIGHATOMIC it is no longer necessary to
> have high-order watermarks. Kswapd and compaction still need high-order
> awareness which is handled by checking that at least one suitable high-order
> page is free.
I still don't think that this one suitable high-order page is enough.
If fragmentation happens, there would be no order-2 freepage. If kswapd
prepares only 1 order-2 freepage, one of two successive process forks
(AFAIK, fork in x86 and ARM require order 2 page) must go to direct reclaim
to make order-2 freepage. Kswapd cannot make order-2 freepage in that
short time. It causes latency to many high-order freepage requestor
in fragmented situation.
> With the patch applied, there was little difference in the allocation
> failure rates as the atomic reserves are small relative to the number of
> allocation attempts. The expected impact is that there will never be an
> allocation failure report that shows suitable pages on the free lists.
Due to highatomic pageblock and freepage count mismatch per allocation
flag, allocation failure with suitable pages can still be possible.
> The one potential side-effect of this is that in a vanilla kernel, the
> watermark checks may have kept a free page for an atomic allocation. Now,
> we are 100% relying on the HighAtomic reserves and an early allocation to
> have allocated them. If the first high-order atomic allocation is after
> the system is already heavily fragmented then it'll fail.
>
> Signed-off-by: Mel Gorman <mgorman@techsingularity.net>
> ---
> mm/page_alloc.c | 38 ++++++++++++++++++++++++--------------
> 1 file changed, 24 insertions(+), 14 deletions(-)
>
> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index 2415f882b89c..35dc578730d1 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -2280,8 +2280,10 @@ static inline bool should_fail_alloc_page(gfp_t gfp_mask, unsigned int order)
> #endif /* CONFIG_FAIL_PAGE_ALLOC */
>
> /*
> - * Return true if free pages are above 'mark'. This takes into account the order
> - * of the allocation.
> + * Return true if free base pages are above 'mark'. For high-order checks it
> + * will return true of the order-0 watermark is reached and there is at least
> + * one free page of a suitable size. Checking now avoids taking the zone lock
> + * to check in the allocation paths if no pages are free.
> */
> static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> unsigned long mark, int classzone_idx, int alloc_flags,
> @@ -2289,7 +2291,7 @@ static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> {
> long min = mark;
> int o;
> - long free_cma = 0;
> + const bool atomic = (alloc_flags & ALLOC_HARDER);
>
> /* free_pages may go negative - that's OK */
> free_pages -= (1 << order) - 1;
> @@ -2301,7 +2303,7 @@ static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> * If the caller is not atomic then discount the reserves. This will
> * over-estimate how the atomic reserve but it avoids a search
> */
> - if (likely(!(alloc_flags & ALLOC_HARDER)))
> + if (likely(!atomic))
> free_pages -= z->nr_reserved_highatomic;
> else
> min -= min / 4;
> @@ -2309,22 +2311,30 @@ static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> #ifdef CONFIG_CMA
> /* If allocation can't use CMA areas don't use free CMA pages */
> if (!(alloc_flags & ALLOC_CMA))
> - free_cma = zone_page_state(z, NR_FREE_CMA_PAGES);
> + free_pages -= zone_page_state(z, NR_FREE_CMA_PAGES);
> #endif
>
> - if (free_pages - free_cma <= min + z->lowmem_reserve[classzone_idx])
> + if (free_pages <= min + z->lowmem_reserve[classzone_idx])
> return false;
> - for (o = 0; o < order; o++) {
> - /* At the next order, this order's pages become unavailable */
> - free_pages -= z->free_area[o].nr_free << o;
>
> - /* Require fewer higher order pages to be free */
> - min >>= 1;
> + /* order-0 watermarks are ok */
> + if (!order)
> + return true;
> +
> + /* Check at least one high-order page is free */
> + for (o = order; o < MAX_ORDER; o++) {
> + struct free_area *area = &z->free_area[o];
> + int mt;
> +
> + if (atomic && area->nr_free)
> + return true;
How about checking area->nr_free first?
In both atomic and !atomic case, nr_free == 0 means
there is no appropriate pages.
So,
if (!area->nr_free)
continue;
if (atomic)
return true;
...
> - if (free_pages <= min)
> - return false;
> + for (mt = 0; mt < MIGRATE_PCPTYPES; mt++) {
> + if (!list_empty(&area->free_list[mt]))
> + return true;
> + }
I'm not sure this is really faster than previous.
We need to check three lists on each order.
Think about order-2 case. I guess order-2 is usually on movable
pageblock rather than unmovable pageblock. In this case,
we need to check three lists so cost is more.
And, if system is fragmented and has not enough order-2 freepage,
we need to check 3,4,..., MAX_ORDER-1 to find out that
there is no order-2 freepage. This would be more costly
than previous approach.
Thanks.
> }
> - return true;
> + return false;
> }
>
> bool zone_watermark_ok(struct zone *z, unsigned int order, unsigned long mark,
> --
> 2.4.6
>
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org. For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2015-09-09 15:00 +0200 |
| Subject | Re: [PATCH 12/12] mm, page_alloc: Only enforce watermarks for order-0 allocations |
| Message-ID | <q6N8L-3wZ-5@gated-at.bofh.it> |
| In reply to | #1220598 |
On Tue, Sep 08, 2015 at 05:26:13PM +0900, Joonsoo Kim wrote:
> 2015-08-24 21:30 GMT+09:00 Mel Gorman <mgorman@techsingularity.net>:
> > The primary purpose of watermarks is to ensure that reclaim can always
> > make forward progress in PF_MEMALLOC context (kswapd and direct reclaim).
> > These assume that order-0 allocations are all that is necessary for
> > forward progress.
> >
> > High-order watermarks serve a different purpose. Kswapd had no high-order
> > awareness before they were introduced (https://lkml.org/lkml/2004/9/5/9).
> > This was particularly important when there were high-order atomic requests.
> > The watermarks both gave kswapd awareness and made a reserve for those
> > atomic requests.
> >
> > There are two important side-effects of this. The most important is that
> > a non-atomic high-order request can fail even though free pages are available
> > and the order-0 watermarks are ok. The second is that high-order watermark
> > checks are expensive as the free list counts up to the requested order must
> > be examined.
> >
> > With the introduction of MIGRATE_HIGHATOMIC it is no longer necessary to
> > have high-order watermarks. Kswapd and compaction still need high-order
> > awareness which is handled by checking that at least one suitable high-order
> > page is free.
>
> I still don't think that this one suitable high-order page is enough.
> If fragmentation happens, there would be no order-2 freepage. If kswapd
> prepares only 1 order-2 freepage, one of two successive process forks
> (AFAIK, fork in x86 and ARM require order 2 page) must go to direct reclaim
> to make order-2 freepage. Kswapd cannot make order-2 freepage in that
> short time. It causes latency to many high-order freepage requestor
> in fragmented situation.
>
So what do you suggest instead? A fixed number, some other heuristic?
You have pushed several times now for the series to focus on the latency
of standard high-order allocations but again I will say that it is outside
the scope of this series. If you want to take steps to reduce the latency
of ordinary high-order allocation requests that can sleep then it should
be a separate series.
> > With the patch applied, there was little difference in the allocation
> > failure rates as the atomic reserves are small relative to the number of
> > allocation attempts. The expected impact is that there will never be an
> > allocation failure report that shows suitable pages on the free lists.
>
> Due to highatomic pageblock and freepage count mismatch per allocation
> flag, allocation failure with suitable pages can still be possible.
>
An allocation failure of this type would be a !atomic allocation that
cannot access the reserve. If such allocations requests can access the
reserve then it defeats the whole point of the pageblock type.
> > + * Return true if free base pages are above 'mark'. For high-order checks it
> > + * will return true of the order-0 watermark is reached and there is at least
> > + * one free page of a suitable size. Checking now avoids taking the zone lock
> > + * to check in the allocation paths if no pages are free.
> > */
> > static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> > unsigned long mark, int classzone_idx, int alloc_flags,
> > @@ -2289,7 +2291,7 @@ static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> > {
> > long min = mark;
> > int o;
> > - long free_cma = 0;
> > + const bool atomic = (alloc_flags & ALLOC_HARDER);
> >
> > /* free_pages may go negative - that's OK */
> > free_pages -= (1 << order) - 1;
> > @@ -2301,7 +2303,7 @@ static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> > * If the caller is not atomic then discount the reserves. This will
> > * over-estimate how the atomic reserve but it avoids a search
> > */
> > - if (likely(!(alloc_flags & ALLOC_HARDER)))
> > + if (likely(!atomic))
> > free_pages -= z->nr_reserved_highatomic;
> > else
> > min -= min / 4;
> > @@ -2309,22 +2311,30 @@ static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> > #ifdef CONFIG_CMA
> > /* If allocation can't use CMA areas don't use free CMA pages */
> > if (!(alloc_flags & ALLOC_CMA))
> > - free_cma = zone_page_state(z, NR_FREE_CMA_PAGES);
> > + free_pages -= zone_page_state(z, NR_FREE_CMA_PAGES);
> > #endif
> >
> > - if (free_pages - free_cma <= min + z->lowmem_reserve[classzone_idx])
> > + if (free_pages <= min + z->lowmem_reserve[classzone_idx])
> > return false;
> > - for (o = 0; o < order; o++) {
> > - /* At the next order, this order's pages become unavailable */
> > - free_pages -= z->free_area[o].nr_free << o;
> >
> > - /* Require fewer higher order pages to be free */
> > - min >>= 1;
> > + /* order-0 watermarks are ok */
> > + if (!order)
> > + return true;
> > +
> > + /* Check at least one high-order page is free */
> > + for (o = order; o < MAX_ORDER; o++) {
> > + struct free_area *area = &z->free_area[o];
> > + int mt;
> > +
> > + if (atomic && area->nr_free)
> > + return true;
>
> How about checking area->nr_free first?
> In both atomic and !atomic case, nr_free == 0 means
> there is no appropriate pages.
>
> So,
> if (!area->nr_free)
> continue;
> if (atomic)
> return true;
> ...
>
>
> > - if (free_pages <= min)
> > - return false;
> > + for (mt = 0; mt < MIGRATE_PCPTYPES; mt++) {
> > + if (!list_empty(&area->free_list[mt]))
> > + return true;
> > + }
>
> I'm not sure this is really faster than previous.
> We need to check three lists on each order.
>
> Think about order-2 case. I guess order-2 is usually on movable
> pageblock rather than unmovable pageblock. In this case,
> we need to check three lists so cost is more.
>
Ok, the extra check makes sense. Thanks.
--
Mel Gorman
SUSE Labs
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Joonsoo Kim <iamjoonsoo.kim@lge.com> |
|---|---|
| Date | 2015-09-18 09:00 +0200 |
| Subject | Re: [PATCH 12/12] mm, page_alloc: Only enforce watermarks for order-0 allocations |
| Message-ID | <q9XOh-3rU-7@gated-at.bofh.it> |
| In reply to | #1221401 |
On Wed, Sep 09, 2015 at 01:39:01PM +0100, Mel Gorman wrote:
> On Tue, Sep 08, 2015 at 05:26:13PM +0900, Joonsoo Kim wrote:
> > 2015-08-24 21:30 GMT+09:00 Mel Gorman <mgorman@techsingularity.net>:
> > > The primary purpose of watermarks is to ensure that reclaim can always
> > > make forward progress in PF_MEMALLOC context (kswapd and direct reclaim).
> > > These assume that order-0 allocations are all that is necessary for
> > > forward progress.
> > >
> > > High-order watermarks serve a different purpose. Kswapd had no high-order
> > > awareness before they were introduced (https://lkml.org/lkml/2004/9/5/9).
> > > This was particularly important when there were high-order atomic requests.
> > > The watermarks both gave kswapd awareness and made a reserve for those
> > > atomic requests.
> > >
> > > There are two important side-effects of this. The most important is that
> > > a non-atomic high-order request can fail even though free pages are available
> > > and the order-0 watermarks are ok. The second is that high-order watermark
> > > checks are expensive as the free list counts up to the requested order must
> > > be examined.
> > >
> > > With the introduction of MIGRATE_HIGHATOMIC it is no longer necessary to
> > > have high-order watermarks. Kswapd and compaction still need high-order
> > > awareness which is handled by checking that at least one suitable high-order
> > > page is free.
> >
> > I still don't think that this one suitable high-order page is enough.
> > If fragmentation happens, there would be no order-2 freepage. If kswapd
> > prepares only 1 order-2 freepage, one of two successive process forks
> > (AFAIK, fork in x86 and ARM require order 2 page) must go to direct reclaim
> > to make order-2 freepage. Kswapd cannot make order-2 freepage in that
> > short time. It causes latency to many high-order freepage requestor
> > in fragmented situation.
> >
>
> So what do you suggest instead? A fixed number, some other heuristic?
> You have pushed several times now for the series to focus on the latency
> of standard high-order allocations but again I will say that it is outside
> the scope of this series. If you want to take steps to reduce the latency
> of ordinary high-order allocation requests that can sleep then it should
> be a separate series.
I don't understand why you think it should be a separate series.
I don't know exact reason why high order watermark check is
introduced, but, based on your description, it is for high-order
allocation request in atomic context. And, it would accidently take care
about latency. It is used for a long time and your patch try to remove it
and it only takes care about success rate. That means that your patch
could cause regression. I think that if this happens actually, it is handled
in this patchset instead of separate series.
In review of previous version, I suggested that removing watermark
check only for higher than PAGE_ALLOC_COSTLY_ORDER. You didn't accept
that and I still don't agree with your approach. You can show me that
my concern is wrong via some number.
One candidate test for this is that making system fragmented and
run hackbench which uses a lot of high-order allocation and measure
elapsed-time.
Thanks.
>
> > > With the patch applied, there was little difference in the allocation
> > > failure rates as the atomic reserves are small relative to the number of
> > > allocation attempts. The expected impact is that there will never be an
> > > allocation failure report that shows suitable pages on the free lists.
> >
> > Due to highatomic pageblock and freepage count mismatch per allocation
> > flag, allocation failure with suitable pages can still be possible.
> >
>
> An allocation failure of this type would be a !atomic allocation that
> cannot access the reserve. If such allocations requests can access the
> reserve then it defeats the whole point of the pageblock type.
>
> > > + * Return true if free base pages are above 'mark'. For high-order checks it
> > > + * will return true of the order-0 watermark is reached and there is at least
> > > + * one free page of a suitable size. Checking now avoids taking the zone lock
> > > + * to check in the allocation paths if no pages are free.
> > > */
> > > static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> > > unsigned long mark, int classzone_idx, int alloc_flags,
> > > @@ -2289,7 +2291,7 @@ static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> > > {
> > > long min = mark;
> > > int o;
> > > - long free_cma = 0;
> > > + const bool atomic = (alloc_flags & ALLOC_HARDER);
> > >
> > > /* free_pages may go negative - that's OK */
> > > free_pages -= (1 << order) - 1;
> > > @@ -2301,7 +2303,7 @@ static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> > > * If the caller is not atomic then discount the reserves. This will
> > > * over-estimate how the atomic reserve but it avoids a search
> > > */
> > > - if (likely(!(alloc_flags & ALLOC_HARDER)))
> > > + if (likely(!atomic))
> > > free_pages -= z->nr_reserved_highatomic;
> > > else
> > > min -= min / 4;
> > > @@ -2309,22 +2311,30 @@ static bool __zone_watermark_ok(struct zone *z, unsigned int order,
> > > #ifdef CONFIG_CMA
> > > /* If allocation can't use CMA areas don't use free CMA pages */
> > > if (!(alloc_flags & ALLOC_CMA))
> > > - free_cma = zone_page_state(z, NR_FREE_CMA_PAGES);
> > > + free_pages -= zone_page_state(z, NR_FREE_CMA_PAGES);
> > > #endif
> > >
> > > - if (free_pages - free_cma <= min + z->lowmem_reserve[classzone_idx])
> > > + if (free_pages <= min + z->lowmem_reserve[classzone_idx])
> > > return false;
> > > - for (o = 0; o < order; o++) {
> > > - /* At the next order, this order's pages become unavailable */
> > > - free_pages -= z->free_area[o].nr_free << o;
> > >
> > > - /* Require fewer higher order pages to be free */
> > > - min >>= 1;
> > > + /* order-0 watermarks are ok */
> > > + if (!order)
> > > + return true;
> > > +
> > > + /* Check at least one high-order page is free */
> > > + for (o = order; o < MAX_ORDER; o++) {
> > > + struct free_area *area = &z->free_area[o];
> > > + int mt;
> > > +
> > > + if (atomic && area->nr_free)
> > > + return true;
> >
> > How about checking area->nr_free first?
> > In both atomic and !atomic case, nr_free == 0 means
> > there is no appropriate pages.
> >
> > So,
> > if (!area->nr_free)
> > continue;
> > if (atomic)
> > return true;
> > ...
> >
> >
> > > - if (free_pages <= min)
> > > - return false;
> > > + for (mt = 0; mt < MIGRATE_PCPTYPES; mt++) {
> > > + if (!list_empty(&area->free_list[mt]))
> > > + return true;
> > > + }
> >
> > I'm not sure this is really faster than previous.
> > We need to check three lists on each order.
> >
> > Think about order-2 case. I guess order-2 is usually on movable
> > pageblock rather than unmovable pageblock. In this case,
> > we need to check three lists so cost is more.
> >
>
> Ok, the extra check makes sense. Thanks.
>
> --
> Mel Gorman
> SUSE Labs
>
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org. For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2015-09-21 13:00 +0200 |
| Subject | Re: [PATCH 12/12] mm, page_alloc: Only enforce watermarks for order-0 allocations |
| Message-ID | <qb6Zd-4p2-39@gated-at.bofh.it> |
| In reply to | #1227601 |
On Fri, Sep 18, 2015 at 03:56:21PM +0900, Joonsoo Kim wrote: > On Wed, Sep 09, 2015 at 01:39:01PM +0100, Mel Gorman wrote: > > On Tue, Sep 08, 2015 at 05:26:13PM +0900, Joonsoo Kim wrote: > > > 2015-08-24 21:30 GMT+09:00 Mel Gorman <mgorman@techsingularity.net>: > > > > The primary purpose of watermarks is to ensure that reclaim can always > > > > make forward progress in PF_MEMALLOC context (kswapd and direct reclaim). > > > > These assume that order-0 allocations are all that is necessary for > > > > forward progress. > > > > > > > > High-order watermarks serve a different purpose. Kswapd had no high-order > > > > awareness before they were introduced (https://lkml.org/lkml/2004/9/5/9). > > > > This was particularly important when there were high-order atomic requests. > > > > The watermarks both gave kswapd awareness and made a reserve for those > > > > atomic requests. > > > > > > > > There are two important side-effects of this. The most important is that > > > > a non-atomic high-order request can fail even though free pages are available > > > > and the order-0 watermarks are ok. The second is that high-order watermark > > > > checks are expensive as the free list counts up to the requested order must > > > > be examined. > > > > > > > > With the introduction of MIGRATE_HIGHATOMIC it is no longer necessary to > > > > have high-order watermarks. Kswapd and compaction still need high-order > > > > awareness which is handled by checking that at least one suitable high-order > > > > page is free. > > > > > > I still don't think that this one suitable high-order page is enough. > > > If fragmentation happens, there would be no order-2 freepage. If kswapd > > > prepares only 1 order-2 freepage, one of two successive process forks > > > (AFAIK, fork in x86 and ARM require order 2 page) must go to direct reclaim > > > to make order-2 freepage. Kswapd cannot make order-2 freepage in that > > > short time. It causes latency to many high-order freepage requestor > > > in fragmented situation. > > > > > > > So what do you suggest instead? A fixed number, some other heuristic? > > You have pushed several times now for the series to focus on the latency > > of standard high-order allocations but again I will say that it is outside > > the scope of this series. If you want to take steps to reduce the latency > > of ordinary high-order allocation requests that can sleep then it should > > be a separate series. > > I don't understand why you think it should be a separate series. Because atomic high-order allocation success and normal high-order allocation stall latency are different problems. Atomic high-order allocation successes are about reserves, normal high-order allocations are about reclaim. > I don't know exact reason why high order watermark check is > introduced, but, based on your description, it is for high-order > allocation request in atomic context. Mostly yes, the initial motivation is described in the linked mail -- give kswapd high-order awareness because otherwise (higher-order && !wait) allocations that fail would wake kswapd but it would go back to sleep. > And, it would accidently take care > about latency. Except all it does is defer the problem. If kswapd frees N high-order pages then it disrupts the system to satisfy the request, potentially reclaiming hot pages for an allocation attempt that *may* occur that will stall if there are N+1 allocation requests. Kswapd reclaiming additional pages is definite system disruption and potentially increases thrashing *now* to help an event that *might* occur in the future. > It is used for a long time and your patch try to remove it > and it only takes care about success rate. That means that your patch > could cause regression. I think that if this happens actually, it is handled > in this patchset instead of separate series. > Except it doesn't really. Current situation o A high-order watermark check might fail for a normal high-order allocation request. On failure, stall to reclaim more pages which may or may not succeed o An atomic allocation may use a lower watermark but it can still fail even if there are free pages on the list Patched situation o A watermark check might fail for a normal high-order allocation request and cannot use one of the reserved pages. On failure, stall to reclaim more pages which may or may not succeed. Functionally, this is very similar to current behaviour o An atomic allocation may use the reserves so if a free page exists, it will be used Functionally, this is more reliable than current behaviour as there is still potential for disruption > In review of previous version, I suggested that removing watermark > check only for higher than PAGE_ALLOC_COSTLY_ORDER. It increases complexity for reasons that are not quantified. > You didn't accept > that and I still don't agree with your approach. You can show me that > my concern is wrong via some number. > > One candidate test for this is that making system fragmented and > run hackbench which uses a lot of high-order allocation and measure > elapsed-time. > o There is no difference in normal allocation high-order success rates with this series appied o With the series applied, such tests complete in approximately the same time o For the tests with parallel high-order allocation requests, there was no significant difference in the elapsed times although success rates were slightly higher Each time the full sets of tests take about 4 days to complete on this series and so far no problems of the type you describe have been found. If such a test case is found then there would a clear workload to justify either having kswapd reclaiming multiple pages or apply the old watermark scheme for lower orders. -- Mel Gorman SUSE Labs -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web