Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1233177 > unrolled thread
| Started by | Tang Chen <tangchen@cn.fujitsu.com> |
|---|---|
| First post | 2015-09-26 11:40 +0200 |
| Last post | 2015-09-28 04:00 +0200 |
| Articles | 3 — 2 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation. Tang Chen <tangchen@cn.fujitsu.com> - 2015-09-26 11:40 +0200
Re: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation. Tejun Heo <tj@kernel.org> - 2015-09-26 20:00 +0200
Re: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation. Tang Chen <tangchen@cn.fujitsu.com> - 2015-09-28 04:00 +0200
| From | Tang Chen <tangchen@cn.fujitsu.com> |
|---|---|
| Date | 2015-09-26 11:40 +0200 |
| Subject | Re: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation. |
| Message-ID | <qcU7w-4qo-19@gated-at.bofh.it> |
Hi, tj On 09/11/2015 03:29 AM, Tejun Heo wrote: > Hello, > > On Thu, Sep 10, 2015 at 12:27:45PM +0800, Tang Chen wrote: >> diff --git a/include/linux/gfp.h b/include/linux/gfp.h >> index ad35f30..1a1324f 100644 >> --- a/include/linux/gfp.h >> +++ b/include/linux/gfp.h >> @@ -307,13 +307,19 @@ static inline struct page *alloc_pages_node(int nid, gfp_t gfp_mask, >> if (nid < 0) >> nid = numa_node_id(); >> >> + if (!node_online(nid)) >> + nid = get_near_online_node(nid); >> + >> return __alloc_pages(gfp_mask, order, node_zonelist(nid, gfp_mask)); >> } > Why not just update node_data[]->node_zonelist in the first place? zonelist will be rebuilt in __offline_pages() when the zone is not populated any more. Here, getting the best near online node is for those cpus on memory-less nodes. In the original code, if nid is NUMA_NO_NODE, the node the current cpu resides in will be chosen. And if the node is memory-less node, the cpu will be mapped to its best near online node. But this patch-set will map the cpu to its original node, so numa_node_id() may return a memory-less node to allocator. And then memory allocation may fail. > Also, what's the synchronization rule here? How are allocators > synchronized against node hot [un]plugs? The rule is: node_to_near_node_map[] array will be updated each time node [un]hotplug happens. Now it is not protected by a lock. But I think acquiring a lock may cause performance regression to memory allocator. When rebuilding zonelist, stop_machine is used. So I think maybe updating the node_to_near_node_map[] array at the same time when zonelist is rebuilt could be a good idea. Thanks. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-26 20:00 +0200 |
| Subject | Re: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation. |
| Message-ID | <qd1Vo-77B-19@gated-at.bofh.it> |
| In reply to | #1233177 |
Hello, Tang. On Sat, Sep 26, 2015 at 05:31:07PM +0800, Tang Chen wrote: > >>@@ -307,13 +307,19 @@ static inline struct page *alloc_pages_node(int nid, gfp_t gfp_mask, > >> if (nid < 0) > >> nid = numa_node_id(); > >>+ if (!node_online(nid)) > >>+ nid = get_near_online_node(nid); > >>+ > >> return __alloc_pages(gfp_mask, order, node_zonelist(nid, gfp_mask)); > >> } > >Why not just update node_data[]->node_zonelist in the first place? > > zonelist will be rebuilt in __offline_pages() when the zone is not populated > any more. > > Here, getting the best near online node is for those cpus on memory-less > nodes. > > In the original code, if nid is NUMA_NO_NODE, the node the current cpu > resides in > will be chosen. And if the node is memory-less node, the cpu will be mapped > to its > best near online node. > > But this patch-set will map the cpu to its original node, so numa_node_id() > may return > a memory-less node to allocator. And then memory allocation may fail. Correct me if I'm wrong but the zonelist dictates which memory areas the page allocator is gonna try to from, right? What I'm wondering is why we aren't handling memory-less nodes by simply updating their zonelists. I mean, if, say, node 2 is memory-less, its zonelist can simply point to zones from other nodes, right? What am I missing here? Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tang Chen <tangchen@cn.fujitsu.com> |
|---|---|
| Date | 2015-09-28 04:00 +0200 |
| Message-ID | <qdvTr-8b5-7@gated-at.bofh.it> |
| In reply to | #1233251 |
Hi, tj, On 09/27/2015 01:53 AM, Tejun Heo wrote: > Hello, Tang. > > On Sat, Sep 26, 2015 at 05:31:07PM +0800, Tang Chen wrote: >>>> @@ -307,13 +307,19 @@ static inline struct page *alloc_pages_node(int nid, gfp_t gfp_mask, >>>> if (nid < 0) >>>> nid = numa_node_id(); >>>> + if (!node_online(nid)) >>>> + nid = get_near_online_node(nid); >>>> + >>>> return __alloc_pages(gfp_mask, order, node_zonelist(nid, gfp_mask)); >>>> } >>> Why not just update node_data[]->node_zonelist in the first place? >> zonelist will be rebuilt in __offline_pages() when the zone is not populated >> any more. >> >> Here, getting the best near online node is for those cpus on memory-less >> nodes. >> >> In the original code, if nid is NUMA_NO_NODE, the node the current cpu >> resides in >> will be chosen. And if the node is memory-less node, the cpu will be mapped >> to its >> best near online node. >> >> But this patch-set will map the cpu to its original node, so numa_node_id() >> may return >> a memory-less node to allocator. And then memory allocation may fail. > Correct me if I'm wrong but the zonelist dictates which memory areas > the page allocator is gonna try to from, right? What I'm wondering is > why we aren't handling memory-less nodes by simply updating their > zonelists. I mean, if, say, node 2 is memory-less, its zonelist can > simply point to zones from other nodes, right? What am I missing > here? Oh, yes, you are right. But I remember some time ago, Liu, Jiang has or was going to handle memory less node like this in his patch: https://lkml.org/lkml/2015/8/16/130 BTW, to Liu Jiang, how is your patches going on ? Thanks. > > Thanks. > -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web