Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1435577 > unrolled thread
| Started by | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| First post | 2016-07-01 22:10 +0200 |
| Last post | 2016-07-07 13:00 +0200 |
| Articles | 7 — 2 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
[PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone Mel Gorman <mgorman@techsingularity.net> - 2016-07-01 22:10 +0200
Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone Minchan Kim <minchan@kernel.org> - 2016-07-05 08:20 +0200
Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone Mel Gorman <mgorman@techsingularity.net> - 2016-07-05 12:40 +0200
Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone Minchan Kim <minchan@kernel.org> - 2016-07-06 03:30 +0200
Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone Mel Gorman <mgorman@techsingularity.net> - 2016-07-06 10:50 +0200
Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone Minchan Kim <minchan@kernel.org> - 2016-07-07 08:30 +0200
Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone Mel Gorman <mgorman@techsingularity.net> - 2016-07-07 13:00 +0200
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2016-07-01 22:10 +0200 |
| Subject | [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone |
| Message-ID | <rQcVb-2Ax-21@gated-at.bofh.it> |
kswapd scans from highest to lowest for a zone that requires balancing.
This was necessary when reclaim was per-zone to fairly age pages on lower
zones. Now that we are reclaiming on a per-node basis, any eligible zone
can be used and pages will still be aged fairly. This patch avoids
reclaiming excessively unless buffer_heads are over the limit and it's
necessary to reclaim from a higher zone than requested by the waker of
kswapd to relieve low memory pressure.
[hillf.zj@alibaba-inc.com: Force kswapd reclaim no more than needed]
Link: http://lkml.kernel.org/r/1466518566-30034-12-git-send-email-mgorman@techsingularity.net
Signed-off-by: Mel Gorman <mgorman@techsingularity.net>
Signed-off-by: Hillf Danton <hillf.zj@alibaba-inc.com>
Acked-by: Vlastimil Babka <vbabka@suse.cz>
---
mm/vmscan.c | 56 ++++++++++++++++++++++++--------------------------------
1 file changed, 24 insertions(+), 32 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index 911142d25de2..2f898ba2ee2e 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -3141,31 +3141,36 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int classzone_idx)
sc.nr_reclaimed = 0;
- /* Scan from the highest requested zone to dma */
- for (i = classzone_idx; i >= 0; i--) {
- zone = pgdat->node_zones + i;
- if (!populated_zone(zone))
- continue;
-
- /*
- * If the number of buffer_heads in the machine
- * exceeds the maximum allowed level and this node
- * has a highmem zone, force kswapd to reclaim from
- * it to relieve lowmem pressure.
- */
- if (buffer_heads_over_limit && is_highmem_idx(i)) {
- classzone_idx = i;
- break;
- }
+ /*
+ * If the number of buffer_heads in the machine exceeds the
+ * maximum allowed level then reclaim from all zones. This is
+ * not specific to highmem as highmem may not exist but it is
+ * it is expected that buffer_heads are stripped in writeback.
+ */
+ if (buffer_heads_over_limit) {
+ for (i = MAX_NR_ZONES - 1; i >= 0; i--) {
+ zone = pgdat->node_zones + i;
+ if (!populated_zone(zone))
+ continue;
- if (!zone_balanced(zone, order, 0)) {
classzone_idx = i;
break;
}
}
- if (i < 0)
- goto out;
+ /*
+ * Only reclaim if there are no eligible zones. Check from
+ * high to low zone to avoid prematurely clearing pgdat
+ * congested state.
+ */
+ for (i = classzone_idx; i >= 0; i--) {
+ zone = pgdat->node_zones + i;
+ if (!populated_zone(zone))
+ continue;
+
+ if (zone_balanced(zone, sc.order, classzone_idx))
+ goto out;
+ }
/*
* Do some background aging of the anon list, to give
@@ -3211,19 +3216,6 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int classzone_idx)
break;
/*
- * Stop reclaiming if any eligible zone is balanced and clear
- * node writeback or congested.
- */
- for (i = 0; i <= classzone_idx; i++) {
- zone = pgdat->node_zones + i;
- if (!populated_zone(zone))
- continue;
-
- if (zone_balanced(zone, sc.order, classzone_idx))
- goto out;
- }
-
- /*
* Raise priority if scanning rate is too low or there was no
* progress in reclaiming pages
*/
--
2.6.4
[toc] | [next] | [standalone]
| From | Minchan Kim <minchan@kernel.org> |
|---|---|
| Date | 2016-07-05 08:20 +0200 |
| Subject | Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone |
| Message-ID | <rRrS9-7vt-3@gated-at.bofh.it> |
| In reply to | #1435577 |
On Fri, Jul 01, 2016 at 09:01:19PM +0100, Mel Gorman wrote:
> kswapd scans from highest to lowest for a zone that requires balancing.
> This was necessary when reclaim was per-zone to fairly age pages on lower
> zones. Now that we are reclaiming on a per-node basis, any eligible zone
> can be used and pages will still be aged fairly. This patch avoids
> reclaiming excessively unless buffer_heads are over the limit and it's
> necessary to reclaim from a higher zone than requested by the waker of
> kswapd to relieve low memory pressure.
>
> [hillf.zj@alibaba-inc.com: Force kswapd reclaim no more than needed]
> Link: http://lkml.kernel.org/r/1466518566-30034-12-git-send-email-mgorman@techsingularity.net
> Signed-off-by: Mel Gorman <mgorman@techsingularity.net>
> Signed-off-by: Hillf Danton <hillf.zj@alibaba-inc.com>
> Acked-by: Vlastimil Babka <vbabka@suse.cz>
> ---
> mm/vmscan.c | 56 ++++++++++++++++++++++++--------------------------------
> 1 file changed, 24 insertions(+), 32 deletions(-)
>
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index 911142d25de2..2f898ba2ee2e 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -3141,31 +3141,36 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int classzone_idx)
>
> sc.nr_reclaimed = 0;
>
> - /* Scan from the highest requested zone to dma */
> - for (i = classzone_idx; i >= 0; i--) {
> - zone = pgdat->node_zones + i;
> - if (!populated_zone(zone))
> - continue;
> -
> - /*
> - * If the number of buffer_heads in the machine
> - * exceeds the maximum allowed level and this node
> - * has a highmem zone, force kswapd to reclaim from
> - * it to relieve lowmem pressure.
> - */
> - if (buffer_heads_over_limit && is_highmem_idx(i)) {
> - classzone_idx = i;
> - break;
> - }
> + /*
> + * If the number of buffer_heads in the machine exceeds the
> + * maximum allowed level then reclaim from all zones. This is
> + * not specific to highmem as highmem may not exist but it is
> + * it is expected that buffer_heads are stripped in writeback.
> + */
> + if (buffer_heads_over_limit) {
> + for (i = MAX_NR_ZONES - 1; i >= 0; i--) {
> + zone = pgdat->node_zones + i;
> + if (!populated_zone(zone))
> + continue;
>
> - if (!zone_balanced(zone, order, 0)) {
> classzone_idx = i;
> break;
> }
> }
>
> - if (i < 0)
> - goto out;
> + /*
> + * Only reclaim if there are no eligible zones. Check from
> + * high to low zone to avoid prematurely clearing pgdat
> + * congested state.
I cannot understand "prematurely clearing pgdat congested state".
Could you add more words to clear it out?
> + */
> + for (i = classzone_idx; i >= 0; i--) {
> + zone = pgdat->node_zones + i;
> + if (!populated_zone(zone))
> + continue;
> +
> + if (zone_balanced(zone, sc.order, classzone_idx))
If buffer_head is over limit, old logic force to reclaim highmem but
this zone_balanced logic will prevent it.
> + goto out;
> + }
>
> /*
> * Do some background aging of the anon list, to give
> @@ -3211,19 +3216,6 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int classzone_idx)
> break;
>
> /*
> - * Stop reclaiming if any eligible zone is balanced and clear
> - * node writeback or congested.
> - */
> - for (i = 0; i <= classzone_idx; i++) {
> - zone = pgdat->node_zones + i;
> - if (!populated_zone(zone))
> - continue;
> -
> - if (zone_balanced(zone, sc.order, classzone_idx))
> - goto out;
> - }
> -
> - /*
> * Raise priority if scanning rate is too low or there was no
> * progress in reclaiming pages
> */
> --
> 2.6.4
>
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org. For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2016-07-05 12:40 +0200 |
| Subject | Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone |
| Message-ID | <rRvVL-1Aj-11@gated-at.bofh.it> |
| In reply to | #1436784 |
On Tue, Jul 05, 2016 at 03:11:17PM +0900, Minchan Kim wrote:
> > - if (i < 0)
> > - goto out;
> > + /*
> > + * Only reclaim if there are no eligible zones. Check from
> > + * high to low zone to avoid prematurely clearing pgdat
> > + * congested state.
>
> I cannot understand "prematurely clearing pgdat congested state".
> Could you add more words to clear it out?
>
It's surprisingly difficult to concisely explain. Is this any better?
/*
* Only reclaim if there are no eligible zones. Check from
* high to low zone as allocations prefer higher zones.
* Scanning from low to high zone would allow congestion to be
* cleared during a very small window when a small low
* zone was balanced even under extreme pressure when the
* overall node may be congested.
*/
> > + */
> > + for (i = classzone_idx; i >= 0; i--) {
> > + zone = pgdat->node_zones + i;
> > + if (!populated_zone(zone))
> > + continue;
> > +
> > + if (zone_balanced(zone, sc.order, classzone_idx))
>
> If buffer_head is over limit, old logic force to reclaim highmem but
> this zone_balanced logic will prevent it.
>
The old logic was always busted on 64-bit because is_highmem would always
be 0. The original intent appears to be that buffer_heads_over_limit
would release the buffers when pages went inactive. There are a number
of things we treated inconsistently that get fixed up in the series and
buffer_heads_over_limit is one of them.
--
Mel Gorman
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Minchan Kim <minchan@kernel.org> |
|---|---|
| Date | 2016-07-06 03:30 +0200 |
| Subject | Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone |
| Message-ID | <rRJP4-2sv-5@gated-at.bofh.it> |
| In reply to | #1436913 |
On Tue, Jul 05, 2016 at 11:38:06AM +0100, Mel Gorman wrote:
> On Tue, Jul 05, 2016 at 03:11:17PM +0900, Minchan Kim wrote:
> > > - if (i < 0)
> > > - goto out;
> > > + /*
> > > + * Only reclaim if there are no eligible zones. Check from
> > > + * high to low zone to avoid prematurely clearing pgdat
> > > + * congested state.
> >
> > I cannot understand "prematurely clearing pgdat congested state".
> > Could you add more words to clear it out?
> >
>
> It's surprisingly difficult to concisely explain. Is this any better?
>
> /*
> * Only reclaim if there are no eligible zones. Check from
> * high to low zone as allocations prefer higher zones.
> * Scanning from low to high zone would allow congestion to be
> * cleared during a very small window when a small low
> * zone was balanced even under extreme pressure when the
> * overall node may be congested.
> */
Surely, it's better. Thanks for the explaining.
I doubt we need such corner case logic at this moment and how it works well
without consistent scan from other callers of zone_balanced where scans
from low to high.
> > > + */
> > > + for (i = classzone_idx; i >= 0; i--) {
> > > + zone = pgdat->node_zones + i;
> > > + if (!populated_zone(zone))
> > > + continue;
> > > +
> > > + if (zone_balanced(zone, sc.order, classzone_idx))
> >
> > If buffer_head is over limit, old logic force to reclaim highmem but
> > this zone_balanced logic will prevent it.
> >
>
> The old logic was always busted on 64-bit because is_highmem would always
> be 0. The original intent appears to be that buffer_heads_over_limit
> would release the buffers when pages went inactive. There are a number
Yes but the difference is in old, it was handled both direct and background
reclaim once buffers_heads is over the limit but your change slightly
changs it so kswapd couldn't reclaim high zone if any eligible zone
is balanced. I don't know how big difference it can make but we saw
highmem buffer_head problems several times, IIRC. So, I just wanted
to notice it to you. whether it's handled or not, it's up to you.
> of things we treated inconsistently that get fixed up in the series and
> buffer_heads_over_limit is one of them.
>
> --
> Mel Gorman
> SUSE Labs
>
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org. For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2016-07-06 10:50 +0200 |
| Subject | Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone |
| Message-ID | <rRQGS-6SL-17@gated-at.bofh.it> |
| In reply to | #1437353 |
On Wed, Jul 06, 2016 at 10:25:54AM +0900, Minchan Kim wrote:
> On Tue, Jul 05, 2016 at 11:38:06AM +0100, Mel Gorman wrote:
> > On Tue, Jul 05, 2016 at 03:11:17PM +0900, Minchan Kim wrote:
> > > > - if (i < 0)
> > > > - goto out;
> > > > + /*
> > > > + * Only reclaim if there are no eligible zones. Check from
> > > > + * high to low zone to avoid prematurely clearing pgdat
> > > > + * congested state.
> > >
> > > I cannot understand "prematurely clearing pgdat congested state".
> > > Could you add more words to clear it out?
> > >
> >
> > It's surprisingly difficult to concisely explain. Is this any better?
> >
> > /*
> > * Only reclaim if there are no eligible zones. Check from
> > * high to low zone as allocations prefer higher zones.
> > * Scanning from low to high zone would allow congestion to be
> > * cleared during a very small window when a small low
> > * zone was balanced even under extreme pressure when the
> > * overall node may be congested.
> > */
>
> Surely, it's better. Thanks for the explaining.
>
> I doubt we need such corner case logic at this moment and how it works well
> without consistent scan from other callers of zone_balanced where scans
> from low to high.
>
I observed that if scanning from low to high here that under heavy memory
pressure that kswapd would scan much more aggressively but unable to reclaim
pages. Granted, part of the problem at the time was that kswapd was woken
based on the first zone in the zoneref instead of the highest zone allowed
by the allocation request which gets addressed by "mm, page_alloc: wake
kswapd based on the highest eligible zone".
> > > > + */
> > > > + for (i = classzone_idx; i >= 0; i--) {
> > > > + zone = pgdat->node_zones + i;
> > > > + if (!populated_zone(zone))
> > > > + continue;
> > > > +
> > > > + if (zone_balanced(zone, sc.order, classzone_idx))
> > >
> > > If buffer_head is over limit, old logic force to reclaim highmem but
> > > this zone_balanced logic will prevent it.
> > >
> >
> > The old logic was always busted on 64-bit because is_highmem would always
> > be 0. The original intent appears to be that buffer_heads_over_limit
> > would release the buffers when pages went inactive. There are a number
>
> Yes but the difference is in old, it was handled both direct and background
> reclaim once buffers_heads is over the limit but your change slightly
> changs it so kswapd couldn't reclaim high zone if any eligible zone
> is balanced. I don't know how big difference it can make but we saw
> highmem buffer_head problems several times, IIRC. So, I just wanted
> to notice it to you. whether it's handled or not, it's up to you.
>
The last time I remember buffer_heads_over_limit was an NTFS filesystem
using small sub-page block sizes with a large highmem:lowmem ratio. If a
similar situation is encountered then a test patch would be something like;
diff --git a/mm/vmscan.c b/mm/vmscan.c
index dc12af938a8d..a8ebd1871f16 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -3151,7 +3151,7 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int classzone_idx)
* zone was balanced even under extreme pressure when the
* overall node may be congested.
*/
- for (i = sc.reclaim_idx; i >= 0; i--) {
+ for (i = sc.reclaim_idx; i >= 0 && !buffer_heads_over_limit; i--) {
zone = pgdat->node_zones + i;
if (!populated_zone(zone))
continue;
I'm not going to go with it for now because buffer_heads_over_limit is not
necessarily a problem unless lowmem is factor. We don't want background
reclaim to go ahead unnecessarily just because buffer_heads_over_limit.
It could be distinguished by only forcing reclaim to go ahead on systems
with highmem.
--
Mel Gorman
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Minchan Kim <minchan@kernel.org> |
|---|---|
| Date | 2016-07-07 08:30 +0200 |
| Subject | Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone |
| Message-ID | <rSaYV-3ar-9@gated-at.bofh.it> |
| In reply to | #1437550 |
On Wed, Jul 06, 2016 at 09:42:00AM +0100, Mel Gorman wrote:
<snip>
> > > >
> > > > If buffer_head is over limit, old logic force to reclaim highmem but
> > > > this zone_balanced logic will prevent it.
> > > >
> > >
> > > The old logic was always busted on 64-bit because is_highmem would always
> > > be 0. The original intent appears to be that buffer_heads_over_limit
> > > would release the buffers when pages went inactive. There are a number
> >
> > Yes but the difference is in old, it was handled both direct and background
> > reclaim once buffers_heads is over the limit but your change slightly
> > changs it so kswapd couldn't reclaim high zone if any eligible zone
> > is balanced. I don't know how big difference it can make but we saw
> > highmem buffer_head problems several times, IIRC. So, I just wanted
> > to notice it to you. whether it's handled or not, it's up to you.
> >
>
> The last time I remember buffer_heads_over_limit was an NTFS filesystem
> using small sub-page block sizes with a large highmem:lowmem ratio. If a
> similar situation is encountered then a test patch would be something like;
>
> diff --git a/mm/vmscan.c b/mm/vmscan.c
> index dc12af938a8d..a8ebd1871f16 100644
> --- a/mm/vmscan.c
> +++ b/mm/vmscan.c
> @@ -3151,7 +3151,7 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int classzone_idx)
> * zone was balanced even under extreme pressure when the
> * overall node may be congested.
> */
> - for (i = sc.reclaim_idx; i >= 0; i--) {
> + for (i = sc.reclaim_idx; i >= 0 && !buffer_heads_over_limit; i--) {
> zone = pgdat->node_zones + i;
> if (!populated_zone(zone))
> continue;
>
> I'm not going to go with it for now because buffer_heads_over_limit is not
> necessarily a problem unless lowmem is factor. We don't want background
> reclaim to go ahead unnecessarily just because buffer_heads_over_limit.
> It could be distinguished by only forcing reclaim to go ahead on systems
> with highmem.
If you don't think it's a problem, I don't want to insist on it because I don't
have any report/workload right now. Instead, please write some comment in there
for others to understand why kswapd is okay to ignore buffer_heads_over_limit
unlike direct reclaim. Such non-symmetric behavior is really hard to follow
without any description.
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2016-07-07 13:00 +0200 |
| Subject | Re: [PATCH 11/31] mm: vmscan: do not reclaim from kswapd if there is any eligible zone |
| Message-ID | <rSfce-5Oc-11@gated-at.bofh.it> |
| In reply to | #1438193 |
On Thu, Jul 07, 2016 at 03:27:01PM +0900, Minchan Kim wrote: > > I'm not going to go with it for now because buffer_heads_over_limit is not > > necessarily a problem unless lowmem is factor. We don't want background > > reclaim to go ahead unnecessarily just because buffer_heads_over_limit. > > It could be distinguished by only forcing reclaim to go ahead on systems > > with highmem. > > If you don't think it's a problem, I don't want to insist on it because I don't > have any report/workload right now. Instead, please write some comment in there > for others to understand why kswapd is okay to ignore buffer_heads_over_limit > unlike direct reclaim. Such non-symmetric behavior is really hard to follow > without any description. Ok, I'll add a patch later in the series that addresses the issue. Currently it's called "mm, vmscan: Have kswapd reclaim from all zones if reclaiming and buffer_heads_over_limit". -- Mel Gorman SUSE Labs
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web