Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1233177 > unrolled thread

Re: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation.

Started byTang Chen <tangchen@cn.fujitsu.com>
First post2015-09-26 11:40 +0200
Last post2015-09-28 04:00 +0200
Articles 3 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation. Tang Chen <tangchen@cn.fujitsu.com> - 2015-09-26 11:40 +0200
    Re: [PATCH v2 3/7] x86, gfp: Cache best near node for memory  allocation. Tejun Heo <tj@kernel.org> - 2015-09-26 20:00 +0200
      Re: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation. Tang Chen <tangchen@cn.fujitsu.com> - 2015-09-28 04:00 +0200

#1233177 — Re: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation.

FromTang Chen <tangchen@cn.fujitsu.com>
Date2015-09-26 11:40 +0200
SubjectRe: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation.
Message-ID<qcU7w-4qo-19@gated-at.bofh.it>
Hi, tj

On 09/11/2015 03:29 AM, Tejun Heo wrote:
> Hello,
>
> On Thu, Sep 10, 2015 at 12:27:45PM +0800, Tang Chen wrote:
>> diff --git a/include/linux/gfp.h b/include/linux/gfp.h
>> index ad35f30..1a1324f 100644
>> --- a/include/linux/gfp.h
>> +++ b/include/linux/gfp.h
>> @@ -307,13 +307,19 @@ static inline struct page *alloc_pages_node(int nid, gfp_t gfp_mask,
>>   	if (nid < 0)
>>   		nid = numa_node_id();
>>   
>> +	if (!node_online(nid))
>> +		nid = get_near_online_node(nid);
>> +
>>   	return __alloc_pages(gfp_mask, order, node_zonelist(nid, gfp_mask));
>>   }
> Why not just update node_data[]->node_zonelist in the first place?

zonelist will be rebuilt in __offline_pages() when the zone is not 
populated any more.

Here, getting the best near online node is for those cpus on memory-less 
nodes.

In the original code, if nid is NUMA_NO_NODE, the node the current cpu 
resides in
will be chosen. And if the node is memory-less node, the cpu will be 
mapped to its
best near online node.

But this patch-set will map the cpu to its original node, so 
numa_node_id() may return
a memory-less node to allocator. And then memory allocation may fail.

> Also, what's the synchronization rule here?  How are allocators
> synchronized against node hot [un]plugs?

The rule is: node_to_near_node_map[] array will be updated each time 
node [un]hotplug happens.

Now it is not protected by a lock. But I think acquiring a lock may 
cause performance regression
to memory allocator.

When rebuilding zonelist, stop_machine is used. So I think maybe 
updating the
node_to_near_node_map[] array at the same time when zonelist is rebuilt 
could be a good idea.

Thanks.



--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1233251 — Re: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation.

FromTejun Heo <tj@kernel.org>
Date2015-09-26 20:00 +0200
SubjectRe: [PATCH v2 3/7] x86, gfp: Cache best near node for memory allocation.
Message-ID<qd1Vo-77B-19@gated-at.bofh.it>
In reply to#1233177
Hello, Tang.

On Sat, Sep 26, 2015 at 05:31:07PM +0800, Tang Chen wrote:
> >>@@ -307,13 +307,19 @@ static inline struct page *alloc_pages_node(int nid, gfp_t gfp_mask,
> >>  	if (nid < 0)
> >>  		nid = numa_node_id();
> >>+	if (!node_online(nid))
> >>+		nid = get_near_online_node(nid);
> >>+
> >>  	return __alloc_pages(gfp_mask, order, node_zonelist(nid, gfp_mask));
> >>  }
> >Why not just update node_data[]->node_zonelist in the first place?
> 
> zonelist will be rebuilt in __offline_pages() when the zone is not populated
> any more.
> 
> Here, getting the best near online node is for those cpus on memory-less
> nodes.
> 
> In the original code, if nid is NUMA_NO_NODE, the node the current cpu
> resides in
> will be chosen. And if the node is memory-less node, the cpu will be mapped
> to its
> best near online node.
> 
> But this patch-set will map the cpu to its original node, so numa_node_id()
> may return
> a memory-less node to allocator. And then memory allocation may fail.

Correct me if I'm wrong but the zonelist dictates which memory areas
the page allocator is gonna try to from, right?  What I'm wondering is
why we aren't handling memory-less nodes by simply updating their
zonelists.  I mean, if, say, node 2 is memory-less, its zonelist can
simply point to zones from other nodes, right?  What am I missing
here?

Thanks.

-- 
tejun
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1233813

FromTang Chen <tangchen@cn.fujitsu.com>
Date2015-09-28 04:00 +0200
Message-ID<qdvTr-8b5-7@gated-at.bofh.it>
In reply to#1233251
Hi, tj,

On 09/27/2015 01:53 AM, Tejun Heo wrote:
> Hello, Tang.
>
> On Sat, Sep 26, 2015 at 05:31:07PM +0800, Tang Chen wrote:
>>>> @@ -307,13 +307,19 @@ static inline struct page *alloc_pages_node(int nid, gfp_t gfp_mask,
>>>>   	if (nid < 0)
>>>>   		nid = numa_node_id();
>>>> +	if (!node_online(nid))
>>>> +		nid = get_near_online_node(nid);
>>>> +
>>>>   	return __alloc_pages(gfp_mask, order, node_zonelist(nid, gfp_mask));
>>>>   }
>>> Why not just update node_data[]->node_zonelist in the first place?
>> zonelist will be rebuilt in __offline_pages() when the zone is not populated
>> any more.
>>
>> Here, getting the best near online node is for those cpus on memory-less
>> nodes.
>>
>> In the original code, if nid is NUMA_NO_NODE, the node the current cpu
>> resides in
>> will be chosen. And if the node is memory-less node, the cpu will be mapped
>> to its
>> best near online node.
>>
>> But this patch-set will map the cpu to its original node, so numa_node_id()
>> may return
>> a memory-less node to allocator. And then memory allocation may fail.
> Correct me if I'm wrong but the zonelist dictates which memory areas
> the page allocator is gonna try to from, right?  What I'm wondering is
> why we aren't handling memory-less nodes by simply updating their
> zonelists.  I mean, if, say, node 2 is memory-less, its zonelist can
> simply point to zones from other nodes, right?  What am I missing
> here?

Oh, yes, you are right. But I remember some time ago, Liu, Jiang has or was
going to handle memory less node like this in his patch:

https://lkml.org/lkml/2015/8/16/130

BTW, to Liu Jiang, how is your patches going on ?

Thanks.

>
> Thanks.
>

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web