Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1651356 > unrolled thread
| Started by | Heiko Carstens <heiko.carstens@de.ibm.com> |
|---|---|
| First post | 2017-05-26 14:30 +0200 |
| Last post | 2017-06-01 09:20 +0200 |
| Articles | 12 — 2 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [-next] memory hotplug regression Heiko Carstens <heiko.carstens@de.ibm.com> - 2017-05-26 14:30 +0200
Re: [-next] memory hotplug regression Michal Hocko <mhocko@kernel.org> - 2017-05-29 11:00 +0200
Re: [-next] memory hotplug regression Heiko Carstens <heiko.carstens@de.ibm.com> - 2017-05-29 12:20 +0200
Re: [-next] memory hotplug regression Michal Hocko <mhocko@kernel.org> - 2017-05-29 12:50 +0200
Re: [-next] memory hotplug regression Michal Hocko <mhocko@kernel.org> - 2017-05-30 14:30 +0200
Re: [-next] memory hotplug regression Heiko Carstens <heiko.carstens@de.ibm.com> - 2017-05-30 14:40 +0200
Re: [-next] memory hotplug regression Michal Hocko <mhocko@kernel.org> - 2017-05-30 16:40 +0200
Re: [-next] memory hotplug regression Michal Hocko <mhocko@kernel.org> - 2017-05-30 17:10 +0200
Re: [-next] memory hotplug regression Heiko Carstens <heiko.carstens@de.ibm.com> - 2017-05-30 20:40 +0200
Re: [-next] memory hotplug regression Michal Hocko <mhocko@kernel.org> - 2017-05-31 08:30 +0200
Re: [-next] memory hotplug regression Heiko Carstens <heiko.carstens@de.ibm.com> - 2017-06-01 09:00 +0200
Re: [-next] memory hotplug regression Michal Hocko <mhocko@kernel.org> - 2017-06-01 09:20 +0200
| From | Heiko Carstens <heiko.carstens@de.ibm.com> |
|---|---|
| Date | 2017-05-26 14:30 +0200 |
| Subject | Re: [-next] memory hotplug regression |
| Message-ID | <tLmxr-5Rf-5@gated-at.bofh.it> |
On Wed, May 24, 2017 at 10:39:57AM +0200, Michal Hocko wrote: > On Wed 24-05-17 10:20:22, Heiko Carstens wrote: > > Having the ZONE_MOVABLE default was actually the only point why s390's > > arch_add_memory() was rather complex compared to other architectures. > > > > We always had this behaviour, since we always wanted to be able to offline > > memory after it was brought online. Given that back then "online_movable" > > did not exist, the initial s390 memory hotplug support simply added all > > additional memory to ZONE_MOVABLE. > > > > Keeping the default the same would be quite important. > > Hmm, that is really unfortunate because I would _really_ like to get rid > of the previous semantic which was really awkward. The whole point of > the rework is to get rid of the nasty zone shifting. > > Is it an option to use `online_movable' rather than `online' in your setup? > Btw. my long term plan is to remove the zone range constrains altogether > so you could online each memblock to the type you want. Would that be > sufficient for you in general? Why is it a problem to change the default for 'online'? As far as I can see that doesn't have too much to do with the order of zones, no? By the way: we played around a bit with the changes wrt memory hotplug. There are a two odd things: 1) With the new code I can generate overlapping zones for ZONE_DMA and ZONE_NORMAL: --- new code: DMA [mem 0x0000000000000000-0x000000007fffffff] Normal [mem 0x0000000080000000-0x000000017fffffff] # cat /sys/devices/system/memory/block_size_bytes 10000000 # cat /sys/devices/system/memory/memory5/valid_zones DMA # echo 0 > /sys/devices/system/memory/memory5/online # cat /sys/devices/system/memory/memory5/valid_zones Normal # echo 1 > /sys/devices/system/memory/memory5/online Normal # cat /proc/zoneinfo Node 0, zone DMA spanned 524288 <----- present 458752 managed 455078 start_pfn: 0 <----- Node 0, zone Normal spanned 720896 present 589824 managed 571648 start_pfn: 327680 <----- So ZONE_DMA ends within ZONE_NORMAL. This shouldn't be possible, unless this restriction is gone? --- old code: # echo 0 > /sys/devices/system/memory/memory5/online # cat /sys/devices/system/memory/memory5/valid_zones DMA # echo online_movable > /sys/devices/system/memory/memory5/state -bash: echo: write error: Invalid argument # echo online_kernel > /sys/devices/system/memory/memory5/state -bash: echo: write error: Invalid argument # echo online > /sys/devices/system/memory/memory5/state # cat /sys/devices/system/memory/memory5/valid_zones DMA 2) Another oddity is that after a memory block was brought online it's association to ZONE_NORMAL or ZONE_MOVABLE seems to be fixed. Even if it is brought offline afterwards: # cat /sys/devices/system/memory/memory16/valid_zones Normal Movable # echo online_movable > /sys/devices/system/memory/memory16/state # echo offline > /sys/devices/system/memory/memory16/state # cat /sys/devices/system/memory/memory16/valid_zones Movable <---- should be "Normal Movable" I assume this happens because start_pfn and spanned pages of the zones aren't updated if a memory block at the beginning or end of a zone is brought offline.
[toc] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-05-29 11:00 +0200 |
| Message-ID | <tMoGS-5NX-9@gated-at.bofh.it> |
| In reply to | #1651356 |
On Fri 26-05-17 14:25:09, Heiko Carstens wrote:
> On Wed, May 24, 2017 at 10:39:57AM +0200, Michal Hocko wrote:
> > On Wed 24-05-17 10:20:22, Heiko Carstens wrote:
> > > Having the ZONE_MOVABLE default was actually the only point why s390's
> > > arch_add_memory() was rather complex compared to other architectures.
> > >
> > > We always had this behaviour, since we always wanted to be able to offline
> > > memory after it was brought online. Given that back then "online_movable"
> > > did not exist, the initial s390 memory hotplug support simply added all
> > > additional memory to ZONE_MOVABLE.
> > >
> > > Keeping the default the same would be quite important.
> >
> > Hmm, that is really unfortunate because I would _really_ like to get rid
> > of the previous semantic which was really awkward. The whole point of
> > the rework is to get rid of the nasty zone shifting.
> >
> > Is it an option to use `online_movable' rather than `online' in your setup?
> > Btw. my long term plan is to remove the zone range constrains altogether
> > so you could online each memblock to the type you want. Would that be
> > sufficient for you in general?
>
> Why is it a problem to change the default for 'online'? As far as I can see
> that doesn't have too much to do with the order of zones, no?
`online' (aka MMOP_ONLINE_KEEP) should always inherit its current zone.
The previous implementation made an exception to allow to shift to
another zone if it is on the border of two zones. This is what I wanted
to get rid of because it is just too ugly to live.
But now I am not really sure what is the usecase here. I assume you know
how to online the memoery. That's why you had to play tricks with the
zones previously. All you need now is to use the proper MMOP_ONLINE*
> By the way: we played around a bit with the changes wrt memory
> hotplug. There are a two odd things:
>
> 1) With the new code I can generate overlapping zones for ZONE_DMA and
> ZONE_NORMAL:
>
> --- new code:
>
> DMA [mem 0x0000000000000000-0x000000007fffffff]
> Normal [mem 0x0000000080000000-0x000000017fffffff]
>
> # cat /sys/devices/system/memory/block_size_bytes
> 10000000
> # cat /sys/devices/system/memory/memory5/valid_zones
> DMA
> # echo 0 > /sys/devices/system/memory/memory5/online
> # cat /sys/devices/system/memory/memory5/valid_zones
> Normal
> # echo 1 > /sys/devices/system/memory/memory5/online
> Normal
OK, interesting. I will double check the code.
> # cat /proc/zoneinfo
> Node 0, zone DMA
> spanned 524288 <-----
> present 458752
> managed 455078
> start_pfn: 0 <-----
>
> Node 0, zone Normal
> spanned 720896
> present 589824
> managed 571648
> start_pfn: 327680 <-----
>
> So ZONE_DMA ends within ZONE_NORMAL. This shouldn't be possible, unless
> this restriction is gone?
>
> --- old code:
>
> # echo 0 > /sys/devices/system/memory/memory5/online
> # cat /sys/devices/system/memory/memory5/valid_zones
> DMA
> # echo online_movable > /sys/devices/system/memory/memory5/state
> -bash: echo: write error: Invalid argument
> # echo online_kernel > /sys/devices/system/memory/memory5/state
> -bash: echo: write error: Invalid argument
> # echo online > /sys/devices/system/memory/memory5/state
> # cat /sys/devices/system/memory/memory5/valid_zones
> DMA
>
>
> 2) Another oddity is that after a memory block was brought online it's
> association to ZONE_NORMAL or ZONE_MOVABLE seems to be fixed. Even if it
> is brought offline afterwards:
This is intended behavior because I got rid of the tricky&ugly zone
shifting code. Ultimately I would like to allow for overlapping zones
so the explicit online_{movable,kernel} will _always_ work.
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Heiko Carstens <heiko.carstens@de.ibm.com> |
|---|---|
| Date | 2017-05-29 12:20 +0200 |
| Message-ID | <tMpWh-6QO-5@gated-at.bofh.it> |
| In reply to | #1652311 |
On Mon, May 29, 2017 at 10:52:31AM +0200, Michal Hocko wrote:
> > Why is it a problem to change the default for 'online'? As far as I can see
> > that doesn't have too much to do with the order of zones, no?
>
> `online' (aka MMOP_ONLINE_KEEP) should always inherit its current zone.
> The previous implementation made an exception to allow to shift to
> another zone if it is on the border of two zones. This is what I wanted
> to get rid of because it is just too ugly to live.
>
> But now I am not really sure what is the usecase here. I assume you know
> how to online the memoery. That's why you had to play tricks with the
> zones previously. All you need now is to use the proper MMOP_ONLINE*
Yes, however that implies that existing user space has to be changed to
achieve the same semantics as before. That's the usecase I'm talking about.
On the other hand this change would finally make s390 behave like all other
architectures, which is certainly not a bad thing. So, while thinking again
I think you convinced me to agree with this change.
> > 2) Another oddity is that after a memory block was brought online it's
> > association to ZONE_NORMAL or ZONE_MOVABLE seems to be fixed. Even if it
> > is brought offline afterwards:
>
> This is intended behavior because I got rid of the tricky&ugly zone
> shifting code. Ultimately I would like to allow for overlapping zones
> so the explicit online_{movable,kernel} will _always_ work.
Ok, I see. This change (fixed memory block to zone mapping after first
online) is a bit surprising. On the other hand I can't think of a sane
usecase why one wants to change the zone a memory block belongs to.
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-05-29 12:50 +0200 |
| Message-ID | <tMqpk-72M-23@gated-at.bofh.it> |
| In reply to | #1652398 |
On Mon 29-05-17 12:11:28, Heiko Carstens wrote:
> On Mon, May 29, 2017 at 10:52:31AM +0200, Michal Hocko wrote:
> > > Why is it a problem to change the default for 'online'? As far as I can see
> > > that doesn't have too much to do with the order of zones, no?
> >
> > `online' (aka MMOP_ONLINE_KEEP) should always inherit its current zone.
> > The previous implementation made an exception to allow to shift to
> > another zone if it is on the border of two zones. This is what I wanted
> > to get rid of because it is just too ugly to live.
> >
> > But now I am not really sure what is the usecase here. I assume you know
> > how to online the memoery. That's why you had to play tricks with the
> > zones previously. All you need now is to use the proper MMOP_ONLINE*
>
> Yes, however that implies that existing user space has to be changed to
> achieve the same semantics as before. That's the usecase I'm talking about.
Yes that is really unfortunate. It is even more unfortunate how the
original behavior got merged without a deeper consideration.
> On the other hand this change would finally make s390 behave like all other
> architectures, which is certainly not a bad thing. So, while thinking again
> I think you convinced me to agree with this change.
That is definitely good to hear. Btw. I plan to change the semantic even
further. MMOP_ONLINE_KEEP currently ignores movable_node setting and I
plan to change that. Hopefully this won't break more userspace...
> > > 2) Another oddity is that after a memory block was brought online it's
> > > association to ZONE_NORMAL or ZONE_MOVABLE seems to be fixed. Even if it
> > > is brought offline afterwards:
> >
> > This is intended behavior because I got rid of the tricky&ugly zone
> > shifting code. Ultimately I would like to allow for overlapping zones
> > so the explicit online_{movable,kernel} will _always_ work.
>
> Ok, I see. This change (fixed memory block to zone mapping after first
> online) is a bit surprising. On the other hand I can't think of a sane
> usecase why one wants to change the zone a memory block belongs to.
Longeterm I would really like to remove any constrains on where to
online movable or kernel memory. So even if this will be problem it will
be only temporary.
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-05-30 14:30 +0200 |
| Message-ID | <tMOrE-7i1-11@gated-at.bofh.it> |
| In reply to | #1651356 |
On Fri 26-05-17 14:25:09, Heiko Carstens wrote:
[...]
> 1) With the new code I can generate overlapping zones for ZONE_DMA and
> ZONE_NORMAL:
>
> --- new code:
>
> DMA [mem 0x0000000000000000-0x000000007fffffff]
> Normal [mem 0x0000000080000000-0x000000017fffffff]
>
> # cat /sys/devices/system/memory/block_size_bytes
> 10000000
> # cat /sys/devices/system/memory/memory5/valid_zones
> DMA
> # echo 0 > /sys/devices/system/memory/memory5/online
> # cat /sys/devices/system/memory/memory5/valid_zones
> Normal
> # echo 1 > /sys/devices/system/memory/memory5/online
> Normal
>
> # cat /proc/zoneinfo
> Node 0, zone DMA
> spanned 524288 <-----
> present 458752
> managed 455078
> start_pfn: 0 <-----
>
> Node 0, zone Normal
> spanned 720896
> present 589824
> managed 571648
> start_pfn: 327680 <-----
>
> So ZONE_DMA ends within ZONE_NORMAL. This shouldn't be possible, unless
> this restriction is gone?
The patch below should help.
> --- old code:
>
> # echo 0 > /sys/devices/system/memory/memory5/online
> # cat /sys/devices/system/memory/memory5/valid_zones
> DMA
> # echo online_movable > /sys/devices/system/memory/memory5/state
> -bash: echo: write error: Invalid argument
> # echo online_kernel > /sys/devices/system/memory/memory5/state
> -bash: echo: write error: Invalid argument
this error doesn't make any sense. Because we we want to online kernel
memory and DMA is pretty much the kernel memory
> # echo online > /sys/devices/system/memory/memory5/state
> # cat /sys/devices/system/memory/memory5/valid_zones
> DMA
---
From 91a432ceb6af9a8f3791d97b6731d2010cbd5b47 Mon Sep 17 00:00:00 2001
From: Michal Hocko <mhocko@suse.com>
Date: Tue, 30 May 2017 13:56:23 +0200
Subject: [PATCH] mm, memory_hotplug: do not assume ZONE_NORMAL is default
kernel zone
Heiko Carstens has noticed that he can generate overlapping zones for
ZONE_DMA and ZONE_NORMAL:
DMA [mem 0x0000000000000000-0x000000007fffffff]
Normal [mem 0x0000000080000000-0x000000017fffffff]
$ cat /sys/devices/system/memory/block_size_bytes
10000000
$ cat /sys/devices/system/memory/memory5/valid_zones
DMA
$ echo 0 > /sys/devices/system/memory/memory5/online
$ cat /sys/devices/system/memory/memory5/valid_zones
Normal
$ echo 1 > /sys/devices/system/memory/memory5/online
Normal
$ cat /proc/zoneinfo
Node 0, zone DMA
spanned 524288 <-----
present 458752
managed 455078
start_pfn: 0 <-----
Node 0, zone Normal
spanned 720896
present 589824
managed 571648
start_pfn: 327680 <-----
The reason is that we assume that the default zone for kernel onlining
is ZONE_NORMAL. This was a simplification introduced by the memory
hotplug rework and it is easily fixable by checking the range overlap in
the zone order and considering the first matching zone as the default
one. If there is no such zone then assume ZONE_NORMAL as we have been
doing so far.
Fixes: "mm, memory_hotplug: do not associate hotadded memory to zones until online"
Reported-by: Heiko Carstens <heiko.carstens@de.ibm.com>
Signed-off-by: Michal Hocko <mhocko@suse.com>
---
drivers/base/memory.c | 2 +-
include/linux/memory_hotplug.h | 2 ++
mm/memory_hotplug.c | 22 +++++++++++++++++++---
3 files changed, 22 insertions(+), 4 deletions(-)
diff --git a/drivers/base/memory.c b/drivers/base/memory.c
index b86fda30ce62..c7c4e0325cdb 100644
--- a/drivers/base/memory.c
+++ b/drivers/base/memory.c
@@ -419,7 +419,7 @@ static ssize_t show_valid_zones(struct device *dev,
nid = pfn_to_nid(start_pfn);
if (allow_online_pfn_range(nid, start_pfn, nr_pages, MMOP_ONLINE_KERNEL)) {
- strcat(buf, NODE_DATA(nid)->node_zones[ZONE_NORMAL].name);
+ strcat(buf, default_zone_for_pfn(nid, start_pfn, nr_pages)->name);
append = true;
}
diff --git a/include/linux/memory_hotplug.h b/include/linux/memory_hotplug.h
index 9e0249d0f5e4..ed167541e4fc 100644
--- a/include/linux/memory_hotplug.h
+++ b/include/linux/memory_hotplug.h
@@ -309,4 +309,6 @@ extern struct page *sparse_decode_mem_map(unsigned long coded_mem_map,
unsigned long pnum);
extern bool allow_online_pfn_range(int nid, unsigned long pfn, unsigned long nr_pages,
int online_type);
+extern struct zone *default_zone_for_pfn(int nid, unsigned long pfn,
+ unsigned long nr_pages);
#endif /* __LINUX_MEMORY_HOTPLUG_H */
diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
index 0a895df2397e..792c098e0e5f 100644
--- a/mm/memory_hotplug.c
+++ b/mm/memory_hotplug.c
@@ -858,7 +858,7 @@ bool allow_online_pfn_range(int nid, unsigned long pfn, unsigned long nr_pages,
{
struct pglist_data *pgdat = NODE_DATA(nid);
struct zone *movable_zone = &pgdat->node_zones[ZONE_MOVABLE];
- struct zone *normal_zone = &pgdat->node_zones[ZONE_NORMAL];
+ struct zone *default_zone = default_zone_for_pfn(nid, pfn, nr_pages);
/*
* TODO there shouldn't be any inherent reason to have ZONE_NORMAL
@@ -872,7 +872,7 @@ bool allow_online_pfn_range(int nid, unsigned long pfn, unsigned long nr_pages,
return true;
return movable_zone->zone_start_pfn >= pfn + nr_pages;
} else if (online_type == MMOP_ONLINE_MOVABLE) {
- return zone_end_pfn(normal_zone) <= pfn;
+ return zone_end_pfn(default_zone) <= pfn;
}
/* MMOP_ONLINE_KEEP will always succeed and inherits the current zone */
@@ -937,6 +937,22 @@ void __ref move_pfn_range_to_zone(struct zone *zone,
set_zone_contiguous(zone);
}
+struct zone *default_zone_for_pfn(int nid, unsigned long start_pfn,
+ unsigned long nr_pages)
+{
+ struct pglist_data *pgdat = NODE_DATA(nid);
+ int zid;
+
+ for (zid = 0; zid < MAX_NR_ZONES; zid++) {
+ struct zone *zone = &pgdat->node_zones[zid];
+
+ if (zone_intersects(zone, start_pfn, nr_pages))
+ return zone;
+ }
+
+ return &pgdat->node_zones[ZONE_NORMAL];
+}
+
/*
* Associates the given pfn range with the given node and the zone appropriate
* for the given online type.
@@ -945,7 +961,7 @@ static struct zone * __meminit move_pfn_range(int online_type, int nid,
unsigned long start_pfn, unsigned long nr_pages)
{
struct pglist_data *pgdat = NODE_DATA(nid);
- struct zone *zone = &pgdat->node_zones[ZONE_NORMAL];
+ struct zone *zone = default_zone_for_pfn(nid, start_pfn, nr_pages);
if (online_type == MMOP_ONLINE_KEEP) {
struct zone *movable_zone = &pgdat->node_zones[ZONE_MOVABLE];
--
2.11.0
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Heiko Carstens <heiko.carstens@de.ibm.com> |
|---|---|
| Date | 2017-05-30 14:40 +0200 |
| Message-ID | <tMOBk-7lm-29@gated-at.bofh.it> |
| In reply to | #1653177 |
On Tue, May 30, 2017 at 02:18:06PM +0200, Michal Hocko wrote:
> > So ZONE_DMA ends within ZONE_NORMAL. This shouldn't be possible, unless
> > this restriction is gone?
>
> The patch below should help.
It does fix this specific problem, but introduces a new one:
# echo online_movable > /sys/devices/system/memory/memory16/state
# cat /sys/devices/system/memory/memory16/valid_zones
Movable
# echo offline > /sys/devices/system/memory/memory16/state
# cat /sys/devices/system/memory/memory16/valid_zones
<--- no output
Memory block 16 is the only one I onlined and offlineto ZONE_MOVABLE.
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-05-30 16:40 +0200 |
| Message-ID | <tMQtr-7f-15@gated-at.bofh.it> |
| In reply to | #1653189 |
On Tue 30-05-17 14:37:24, Heiko Carstens wrote:
> On Tue, May 30, 2017 at 02:18:06PM +0200, Michal Hocko wrote:
> > > So ZONE_DMA ends within ZONE_NORMAL. This shouldn't be possible, unless
> > > this restriction is gone?
> >
> > The patch below should help.
>
> It does fix this specific problem, but introduces a new one:
>
> # echo online_movable > /sys/devices/system/memory/memory16/state
> # cat /sys/devices/system/memory/memory16/valid_zones
> Movable
> # echo offline > /sys/devices/system/memory/memory16/state
> # cat /sys/devices/system/memory/memory16/valid_zones
> <--- no output
>
> Memory block 16 is the only one I onlined and offlineto ZONE_MOVABLE.
Could you test the this on top please?
---
diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
index 792c098e0e5f..a26f9f8e6365 100644
--- a/mm/memory_hotplug.c
+++ b/mm/memory_hotplug.c
@@ -937,13 +937,18 @@ void __ref move_pfn_range_to_zone(struct zone *zone,
set_zone_contiguous(zone);
}
+/*
+ * Returns a default kernel memory zone for the given pfn range.
+ * If no kernel zone covers this pfn range it will automatically go
+ * to the ZONE_NORMAL.
+ */
struct zone *default_zone_for_pfn(int nid, unsigned long start_pfn,
unsigned long nr_pages)
{
struct pglist_data *pgdat = NODE_DATA(nid);
int zid;
- for (zid = 0; zid < MAX_NR_ZONES; zid++) {
+ for (zid = 0; zid <= ZONE_NORMAL; zid++) {
struct zone *zone = &pgdat->node_zones[zid];
if (zone_intersects(zone, start_pfn, nr_pages))
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-05-30 17:10 +0200 |
| Message-ID | <tMQWu-wG-23@gated-at.bofh.it> |
| In reply to | #1653286 |
On Tue 30-05-17 16:55:01, Heiko Carstens wrote:
> On Tue, May 30, 2017 at 04:32:47PM +0200, Michal Hocko wrote:
> > On Tue 30-05-17 14:37:24, Heiko Carstens wrote:
> > > On Tue, May 30, 2017 at 02:18:06PM +0200, Michal Hocko wrote:
> > > > > So ZONE_DMA ends within ZONE_NORMAL. This shouldn't be possible, unless
> > > > > this restriction is gone?
> > > >
> > > > The patch below should help.
> > >
> > > It does fix this specific problem, but introduces a new one:
> > >
> > > # echo online_movable > /sys/devices/system/memory/memory16/state
> > > # cat /sys/devices/system/memory/memory16/valid_zones
> > > Movable
> > > # echo offline > /sys/devices/system/memory/memory16/state
> > > # cat /sys/devices/system/memory/memory16/valid_zones
> > > <--- no output
> > >
> > > Memory block 16 is the only one I onlined and offlineto ZONE_MOVABLE.
> >
> > Could you test the this on top please?
> > ---
> > diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
> > index 792c098e0e5f..a26f9f8e6365 100644
> > --- a/mm/memory_hotplug.c
> > +++ b/mm/memory_hotplug.c
> > @@ -937,13 +937,18 @@ void __ref move_pfn_range_to_zone(struct zone *zone,
> > set_zone_contiguous(zone);
> > }
> >
> > +/*
> > + * Returns a default kernel memory zone for the given pfn range.
> > + * If no kernel zone covers this pfn range it will automatically go
> > + * to the ZONE_NORMAL.
> > + */
> > struct zone *default_zone_for_pfn(int nid, unsigned long start_pfn,
> > unsigned long nr_pages)
> > {
> > struct pglist_data *pgdat = NODE_DATA(nid);
> > int zid;
> >
> > - for (zid = 0; zid < MAX_NR_ZONES; zid++) {
> > + for (zid = 0; zid <= ZONE_NORMAL; zid++) {
> > struct zone *zone = &pgdat->node_zones[zid];
> >
> > if (zone_intersects(zone, start_pfn, nr_pages))
>
> Still broken, but in different way(s):
>
> # cat /sys/devices/system/memory/memory16/valid_zones
> Normal Movable
> # echo online_movable > /sys/devices/system/memory/memory16/state
> # cat /sys/devices/system/memory/memory16/valid_zones
> Movable
> # cat /sys/devices/system/memory/memory18/valid_zones
> Movable
> # echo online > /sys/devices/system/memory/memory18/state
> # cat /sys/devices/system/memory/memory18/valid_zones
> Normal <--- should be Movable
> # cat /sys/devices/system/memory/memory17/valid_zones
> <--- no output
OK, I will sit on this tomorrow with a clean head without doing 10
things at the same time. Sorry about your wasted time!
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Heiko Carstens <heiko.carstens@de.ibm.com> |
|---|---|
| Date | 2017-05-30 20:40 +0200 |
| Message-ID | <tMQWu-wG-25@gated-at.bofh.it> |
| In reply to | #1653286 |
On Tue, May 30, 2017 at 04:32:47PM +0200, Michal Hocko wrote:
> On Tue 30-05-17 14:37:24, Heiko Carstens wrote:
> > On Tue, May 30, 2017 at 02:18:06PM +0200, Michal Hocko wrote:
> > > > So ZONE_DMA ends within ZONE_NORMAL. This shouldn't be possible, unless
> > > > this restriction is gone?
> > >
> > > The patch below should help.
> >
> > It does fix this specific problem, but introduces a new one:
> >
> > # echo online_movable > /sys/devices/system/memory/memory16/state
> > # cat /sys/devices/system/memory/memory16/valid_zones
> > Movable
> > # echo offline > /sys/devices/system/memory/memory16/state
> > # cat /sys/devices/system/memory/memory16/valid_zones
> > <--- no output
> >
> > Memory block 16 is the only one I onlined and offlineto ZONE_MOVABLE.
>
> Could you test the this on top please?
> ---
> diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
> index 792c098e0e5f..a26f9f8e6365 100644
> --- a/mm/memory_hotplug.c
> +++ b/mm/memory_hotplug.c
> @@ -937,13 +937,18 @@ void __ref move_pfn_range_to_zone(struct zone *zone,
> set_zone_contiguous(zone);
> }
>
> +/*
> + * Returns a default kernel memory zone for the given pfn range.
> + * If no kernel zone covers this pfn range it will automatically go
> + * to the ZONE_NORMAL.
> + */
> struct zone *default_zone_for_pfn(int nid, unsigned long start_pfn,
> unsigned long nr_pages)
> {
> struct pglist_data *pgdat = NODE_DATA(nid);
> int zid;
>
> - for (zid = 0; zid < MAX_NR_ZONES; zid++) {
> + for (zid = 0; zid <= ZONE_NORMAL; zid++) {
> struct zone *zone = &pgdat->node_zones[zid];
>
> if (zone_intersects(zone, start_pfn, nr_pages))
Still broken, but in different way(s):
# cat /sys/devices/system/memory/memory16/valid_zones
Normal Movable
# echo online_movable > /sys/devices/system/memory/memory16/state
# cat /sys/devices/system/memory/memory16/valid_zones
Movable
# cat /sys/devices/system/memory/memory18/valid_zones
Movable
# echo online > /sys/devices/system/memory/memory18/state
# cat /sys/devices/system/memory/memory18/valid_zones
Normal <--- should be Movable
# cat /sys/devices/system/memory/memory17/valid_zones
<--- no output
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-05-31 08:30 +0200 |
| Message-ID | <tN5iN-1a0-3@gated-at.bofh.it> |
| In reply to | #1653480 |
On Tue 30-05-17 16:55:01, Heiko Carstens wrote:
> On Tue, May 30, 2017 at 04:32:47PM +0200, Michal Hocko wrote:
> > On Tue 30-05-17 14:37:24, Heiko Carstens wrote:
> > > On Tue, May 30, 2017 at 02:18:06PM +0200, Michal Hocko wrote:
> > > > > So ZONE_DMA ends within ZONE_NORMAL. This shouldn't be possible, unless
> > > > > this restriction is gone?
> > > >
> > > > The patch below should help.
> > >
> > > It does fix this specific problem, but introduces a new one:
> > >
> > > # echo online_movable > /sys/devices/system/memory/memory16/state
> > > # cat /sys/devices/system/memory/memory16/valid_zones
> > > Movable
> > > # echo offline > /sys/devices/system/memory/memory16/state
> > > # cat /sys/devices/system/memory/memory16/valid_zones
> > > <--- no output
> > >
> > > Memory block 16 is the only one I onlined and offlineto ZONE_MOVABLE.
> >
> > Could you test the this on top please?
> > ---
> > diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
> > index 792c098e0e5f..a26f9f8e6365 100644
> > --- a/mm/memory_hotplug.c
> > +++ b/mm/memory_hotplug.c
> > @@ -937,13 +937,18 @@ void __ref move_pfn_range_to_zone(struct zone *zone,
> > set_zone_contiguous(zone);
> > }
> >
> > +/*
> > + * Returns a default kernel memory zone for the given pfn range.
> > + * If no kernel zone covers this pfn range it will automatically go
> > + * to the ZONE_NORMAL.
> > + */
> > struct zone *default_zone_for_pfn(int nid, unsigned long start_pfn,
> > unsigned long nr_pages)
> > {
> > struct pglist_data *pgdat = NODE_DATA(nid);
> > int zid;
> >
> > - for (zid = 0; zid < MAX_NR_ZONES; zid++) {
> > + for (zid = 0; zid <= ZONE_NORMAL; zid++) {
> > struct zone *zone = &pgdat->node_zones[zid];
> >
> > if (zone_intersects(zone, start_pfn, nr_pages))
>
> Still broken, but in different way(s):
>
> # cat /sys/devices/system/memory/memory16/valid_zones
> Normal Movable
> # echo online_movable > /sys/devices/system/memory/memory16/state
> # cat /sys/devices/system/memory/memory16/valid_zones
> Movable
> # cat /sys/devices/system/memory/memory18/valid_zones
> Movable
> # echo online > /sys/devices/system/memory/memory18/state
> # cat /sys/devices/system/memory/memory18/valid_zones
> Normal <--- should be Movable
> # cat /sys/devices/system/memory/memory17/valid_zones
> <--- no output
OK, so this is an independent problem and an unrelated one to the
patch I've posted. We need two patches actually. Damn, I hate
MMOP_ONLINE_KEEP. I will send 2 patches as a reply to this email.
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Heiko Carstens <heiko.carstens@de.ibm.com> |
|---|---|
| Date | 2017-06-01 09:00 +0200 |
| Message-ID | <tNsfo-7wq-11@gated-at.bofh.it> |
| In reply to | #1653871 |
On Wed, May 31, 2017 at 08:24:40AM +0200, Michal Hocko wrote:
> > # cat /sys/devices/system/memory/memory16/valid_zones
> > Normal Movable
> > # echo online_movable > /sys/devices/system/memory/memory16/state
> > # cat /sys/devices/system/memory/memory16/valid_zones
> > Movable
> > # cat /sys/devices/system/memory/memory18/valid_zones
> > Movable
> > # echo online > /sys/devices/system/memory/memory18/state
> > # cat /sys/devices/system/memory/memory18/valid_zones
> > Normal <--- should be Movable
> > # cat /sys/devices/system/memory/memory17/valid_zones
> > <--- no output
>
> OK, so this is an independent problem and an unrelated one to the
> patch I've posted. We need two patches actually. Damn, I hate
> MMOP_ONLINE_KEEP. I will send 2 patches as a reply to this email.
Tested with your patches on top of linux-next as of yesterday, however
starting at commit fa812e869a6fe7495a17150bb2639075081ef709 ("mm/zswap.c:
delete an error message for a failed memory allocation in
zswap_dstmem_prepare()"), since the "mm: per-lruvec slab stats" patch
series breaks everything ;)
Tested-by: Heiko Carstens <heiko.carstens@de.ibm.com>
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-06-01 09:20 +0200 |
| Message-ID | <tNsyJ-7UU-9@gated-at.bofh.it> |
| In reply to | #1654860 |
On Thu 01-06-17 08:49:54, Heiko Carstens wrote:
> On Wed, May 31, 2017 at 08:24:40AM +0200, Michal Hocko wrote:
> > > # cat /sys/devices/system/memory/memory16/valid_zones
> > > Normal Movable
> > > # echo online_movable > /sys/devices/system/memory/memory16/state
> > > # cat /sys/devices/system/memory/memory16/valid_zones
> > > Movable
> > > # cat /sys/devices/system/memory/memory18/valid_zones
> > > Movable
> > > # echo online > /sys/devices/system/memory/memory18/state
> > > # cat /sys/devices/system/memory/memory18/valid_zones
> > > Normal <--- should be Movable
> > > # cat /sys/devices/system/memory/memory17/valid_zones
> > > <--- no output
> >
> > OK, so this is an independent problem and an unrelated one to the
> > patch I've posted. We need two patches actually. Damn, I hate
> > MMOP_ONLINE_KEEP. I will send 2 patches as a reply to this email.
>
> Tested with your patches on top of linux-next as of yesterday, however
> starting at commit fa812e869a6fe7495a17150bb2639075081ef709 ("mm/zswap.c:
> delete an error message for a failed memory allocation in
> zswap_dstmem_prepare()"), since the "mm: per-lruvec slab stats" patch
> series breaks everything ;)
>
> Tested-by: Heiko Carstens <heiko.carstens@de.ibm.com>
Thanks a lot for testing! I will post those patches for wider review
later today.
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web