Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1506875 > unrolled thread
| Started by | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| First post | 2016-10-24 06:40 +0200 |
| Last post | 2016-10-24 21:40 +0200 |
| Articles | 6 on this page of 46 — 6 participants |
Back to article view | Back to linux.kernel
[RFC 0/8] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:40 +0200
[RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:40 +0200
Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths Dave Hansen <dave.hansen@intel.com> - 2016-10-24 19:20 +0200
Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-10-25 06:20 +0200
Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths Balbir Singh <bsingharora@gmail.com> - 2016-10-25 09:20 +0200
Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths Balbir Singh <bsingharora@gmail.com> - 2016-10-25 09:30 +0200
[RFC 7/8] mm: Add a new migration function migrate_virtual_range() Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:40 +0200
[DEBUG 06/10] mm: Export definition of 'zone_names' array through mmzone.h Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
[DEBUG 00/10] Test and debug patches for coherent device memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
[DEBUG 02/10] powerpc/mm: Create numa nodes for hotplug memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
[DEBUG 03/10] powerpc/mm: Allow memory hotplug into a memory less node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
[DEBUG 08/10] powerpc: Enable CONFIG_MOVABLE_NODE for PPC64 platform Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
[DEBUG 01/10] dt-bindings: Add doc for ibm,hotplug-aperture Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
[DEBUG 05/10] powerpc/mm: Identify isolation seeking coherent memory nodes during boot Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
[DEBUG 10/10] test: Add a script to perform random VMA migrations across nodes Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
[DEBUG 07/10] mm: Add debugfs interface to dump each node's zonelist information Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
[DEBUG 04/10] mm: Enable CONFIG_MOVABLE_NODE on powerpc Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
[DEBUG 09/10] drivers: Add two drivers for coherent device memory tests Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-24 19:10 +0200
Re: [RFC 0/8] Define coherent device memory node "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-10-25 06:30 +0200
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-25 17:20 +0200
Re: [RFC 0/8] Define coherent device memory node "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-10-26 13:10 +0200
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-26 18:10 +0200
Re: [RFC 0/8] Define coherent device memory node "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-10-28 09:40 +0200
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-28 18:20 +0200
Re: [RFC 0/8] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-05 06:30 +0100
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-11-05 19:10 +0100
Re: [RFC 0/8] Define coherent device memory node "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-10-25 07:10 +0200
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-25 17:40 +0200
Re: [RFC 0/8] Define coherent device memory node "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-10-25 19:40 +0200
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-25 21:00 +0200
Re: [RFC 0/8] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-26 13:20 +0200
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-26 18:10 +0200
Re: [RFC 0/8] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-27 06:40 +0200
Re: [RFC 0/8] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-27 09:10 +0200
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-27 17:10 +0200
Re: [RFC 0/8] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-28 07:50 +0200
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-28 18:10 +0200
Re: [RFC 0/8] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-26 15:00 +0200
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-26 18:30 +0200
Re: [RFC 0/8] Define coherent device memory node Balbir Singh <bsingharora@gmail.com> - 2016-10-27 17:10 +0200
Re: [RFC 0/8] Define coherent device memory node Balbir Singh <bsingharora@gmail.com> - 2016-10-25 14:10 +0200
Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-25 17:30 +0200
Re: [RFC 0/8] Define coherent device memory node Dave Hansen <dave.hansen@intel.com> - 2016-10-24 20:10 +0200
Re: [RFC 0/8] Define coherent device memory node David Nellans <dnellans@nvidia.com> - 2016-10-24 20:40 +0200
Re: [RFC 0/8] Define coherent device memory node Dave Hansen <dave.hansen@intel.com> - 2016-10-24 21:40 +0200
Page 3 of 3 — ← Prev page 1 2 [3]
| From | Balbir Singh <bsingharora@gmail.com> |
|---|---|
| Date | 2016-10-27 17:10 +0200 |
| Message-ID | <swUtA-3yF-21@gated-at.bofh.it> |
| In reply to | #1509581 |
On 27/10/16 03:28, Jerome Glisse wrote: > On Wed, Oct 26, 2016 at 06:26:02PM +0530, Anshuman Khandual wrote: >> On 10/26/2016 12:22 AM, Jerome Glisse wrote: >>> On Tue, Oct 25, 2016 at 11:01:08PM +0530, Aneesh Kumar K.V wrote: >>>> Jerome Glisse <j.glisse@gmail.com> writes: >>>> >>>>> On Tue, Oct 25, 2016 at 10:29:38AM +0530, Aneesh Kumar K.V wrote: >>>>>> Jerome Glisse <j.glisse@gmail.com> writes: >>>>>>> On Mon, Oct 24, 2016 at 10:01:49AM +0530, Anshuman Khandual wrote: >>>>> >>>>> [...] >>>>> >>>>>>> You can take a look at hmm-v13 if you want to see how i do non LRU page >>>>>>> migration. While i put most of the migration code inside hmm_migrate.c it >>>>>>> could easily be move to migrate.c without hmm_ prefix. >>>>>>> >>>>>>> There is 2 missing piece with existing migrate code. First is to put memory >>>>>>> allocation for destination under control of who call the migrate code. Second >>>>>>> is to allow offloading the copy operation to device (ie not use the CPU to >>>>>>> copy data). >>>>>>> >>>>>>> I believe same requirement also make sense for platform you are targeting. >>>>>>> Thus same code can be use. >>>>>>> >>>>>>> hmm-v13 https://cgit.freedesktop.org/~glisse/linux/log/?h=hmm-v13 >>>>>>> >>>>>>> I haven't posted this patchset yet because we are doing some modifications >>>>>>> to the device driver API to accomodate some new features. But the ZONE_DEVICE >>>>>>> changes and the overall migration code will stay the same more or less (i have >>>>>>> patches that move it to migrate.c and share more code with existing migrate >>>>>>> code). >>>>>>> >>>>>>> If you think i missed anything about lru and page cache please point it to >>>>>>> me. Because when i audited code for that i didn't see any road block with >>>>>>> the few fs i was looking at (ext4, xfs and core page cache code). >>>>>>> >>>>>> >>>>>> The other restriction around ZONE_DEVICE is, it is not a managed zone. >>>>>> That prevents any direct allocation from coherent device by application. >>>>>> ie, we would like to force allocation from coherent device using >>>>>> interface like mbind(MPOL_BIND..) . Is that possible with ZONE_DEVICE ? >>>>> >>>>> To achieve this we rely on device fault code path ie when device take a page fault >>>>> with help of HMM it will use existing memory if any for fault address but if CPU >>>>> page table is empty (and it is not file back vma because of readback) then device >>>>> can directly allocate device memory and HMM will update CPU page table to point to >>>>> newly allocated device memory. >>>>> >>>> >>>> That is ok if the device touch the page first. What if we want the >>>> allocation touched first by cpu to come from GPU ?. Should we always >>>> depend on GPU driver to migrate such pages later from system RAM to GPU >>>> memory ? >>>> >>> >>> I am not sure what kind of workload would rather have every first CPU access for >>> a range to use device memory. So no my code does not handle that and it is pointless >>> for it as CPU can not access device memory for me. >> >> If the user space application can explicitly allocate device memory directly, we >> can save one round of migration when the device start accessing it. But then one >> can argue what problem statement the device would work on on a freshly allocated >> memory which has not been accessed by CPU for loading the data yet. Will look into >> this scenario in more detail. >> >>> >>> That said nothing forbid to add support for ZONE_DEVICE with mbind() like syscall. >>> Thought my personnal preference would still be to avoid use of such generic syscall >>> but have device driver set allocation policy through its own userspace API (device >>> driver could reuse internal of mbind() to achieve the end result). >> >> Okay, the basic premise of CDM node is to have a LRU based design where we can >> avoid use of driver specific user space memory management code altogether. > > And i think it is not a good fit, at least not for GPU. GPU device driver have a > big chunk of code dedicated to memory management. You can look at drm/ttm and at > userspace (most is in userspace). It is not because we want to reinvent the wheel > it is because they are some unique constraint. > Could you elaborate on the unique constraints a bit more? I looked at ttm briefly (specifically ttm_memory.c), I can see zones being replicated, it feels like a mini-mm is embedded in there. > >>> >>> I am not saying that eveything you want to do is doable now with HMM but, nothing >>> preclude achieving what you want to achieve using ZONE_DEVICE. I really don't think >>> any of the existing mm mechanism (kswapd, lru, numa, ...) are nice fit and can be reuse >>> with device memory. >> >> With CDM node based design, the expectation is to get all/maximum core VM mechanism >> working so that, driver has to do less device specific optimization. > > I think this is a bad idea, today, for GPU but i might be wrong. Why do you think so? What aspects do you think are wrong? I am guessing you mean that the GPU driver via the GEM/DRM/TTM layers should interact with the mm and manage their own memory and use some form of TTM mm abstraction? I'll study those systems if possible as well. > >>> >>> Each device is so different from the other that i don't believe in a one API fit all. >> >> Right, so as I had mentioned in the cover letter, pglist_data->coherent_device actually >> can become a bit mask indicating the type of coherent device the node is and that can >> be used to implement multiple types of requirement in core mm for various kinds of >> devices in the future. > > I really don't want to move GPU memory management into core mm, if you only concider GPGPU > then it _might_ make sense but for graphic side i definitly don't think so. There are way > to much device specific consideration to have in respect of memory management for GPU > (not only in between different vendor but difference between different generation). > Yes, GPGPU is of interest. We don't look at it as GPU memory management. The memory on the device is coherent, it is a part of the system. It comes online later and we would like to hotplug it out if required. Since it's sitting on a bus, we do need optimizations and the ability to migrate to and from it. I don't think it makes sense to replicate a lot of the mm core logic to manage this memory, IMHO. I think I'd like to point out is that it is wrong to assume only a GPU having coherent memory, the RFC clarifies. > >>> The drm GPU subsystem of the kernel is a testimony of how little can be share when it >>> comes to GPU. The only common code is modesetting. Everything that deals with how to >>> use GPU to compute stuff is per device and most of the logic is in userspace. So i do >> >> Whats the basic reason which prevents such code/functionality sharing ? > > While the higher level API (OpenGL, OpenCL, Vulkan, Cuda, ...) offer an abstraction model, > they are all different abstractions. They are just no way to have kernel expose a common > API that would allow all of the above to be implemented. > > Each GPU have complex memory management and requirement (not only differ between vendor > but also between generation of same vendor). They have different isa for each generation. > They have different way to schedule job for each generation. They offer different sync > mechanism. They have different page table format, mmu, ... > Agreed > Basicly each GPU generation is a platform on it is own, like arm, ppc, x86, ... so i do > not see a way to expose a common API and i don't think anyone who as work on any number > of GPU see one either. I wish but it is just not the case. > We are trying to leverage the ability to see coherent memory (across a set of devices plus system RAM) to keep memory management as simple as possible > >>> not see any commonality that could be abstracted at syscall level. I would rather let >>> device driver stack (kernel and userspace) take such decision and have the higher level >>> API (OpenCL, Cuda, C++17, ...) expose something that make sense for each of them. >>> Programmer target those high level API and they intend to use the mechanism each offer >>> to manage memory and memory placement. I would say forcing them to use a second linux >>> specific API to achieve the latter is wrong, at lest for now. >> >> But going forward dont we want a more closely integrated coherent device solution >> which does not depend too much on a device driver stack ? and can be used from a >> basic user space program ? > > That is something i want, but i strongly believe we are not there yet, we have no real > world experience. All we have in the open source community is the graphic stack (drm) > and the graphic stack clearly shows that today there is no common denominator between > GPU outside of modesetting. > :) > So while i share the same aim, i think for now we need to have real experience. Once we > have something like OpenCL >= 2.0, C++17 and couple other userspace API being actively > use on linux with different coherent devices then we can start looking at finding a > common denominator that make sense for enough devices. > > I am sure device driver would like to get rid of their custom memory management but i > don't think this is applicable now. I fear existing mm code would always make the worst > decision when it comes to memory placement, migration and reclaim. > Agreed, we don't want to make either placement/migration or reclaim slow. As I said earlier we should not restrict our thinking to just GPU devices. Balbir Singh.
[toc] | [prev] | [next] | [standalone]
| From | Balbir Singh <bsingharora@gmail.com> |
|---|---|
| Date | 2016-10-25 14:10 +0200 |
| Message-ID | <sw8Ih-5rG-7@gated-at.bofh.it> |
| In reply to | #1507467 |
On 25/10/16 04:09, Jerome Glisse wrote: > On Mon, Oct 24, 2016 at 10:01:49AM +0530, Anshuman Khandual wrote: > >> [...] > >> Core kernel memory features like reclamation, evictions etc. might >> need to be restricted or modified on the coherent device memory node as >> they can be performance limiting. The RFC does not propose anything on this >> yet but it can be looked into later on. For now it just disables Auto NUMA >> for any VMA which has coherent device memory. >> >> Seamless integration of coherent device memory with system memory >> will enable various other features, some of which can be listed as follows. >> >> a. Seamless migrations between system RAM and the coherent memory >> b. Will have asynchronous and high throughput migrations >> c. Be able to allocate huge order pages from these memory regions >> d. Restrict allocations to a large extent to the tasks using the >> device for workload acceleration >> >> Before concluding, will look into the reasons why the existing >> solutions don't work. There are two basic requirements which have to be >> satisfies before the coherent device memory can be integrated with core >> kernel seamlessly. >> >> a. PFN must have struct page >> b. Struct page must able to be inside standard LRU lists >> >> The above two basic requirements discard the existing method of >> device memory representation approaches like these which then requires the >> need of creating a new framework. > > I do not believe the LRU list is a hard requirement, yes when faulting in > a page inside the page cache it assumes it needs to be added to lru list. > But i think this can easily be work around. > > In HMM i am using ZONE_DEVICE and because memory is not accessible from CPU > (not everyone is bless with decent system bus like CAPI, CCIX, Gen-Z, ...) > so in my case a file back page must always be spawn first from a regular > page and once read from disk then i can migrate to GPU page. > I've not seen the HMM patchset, but read from disk will go to ZONE_DEVICE? Then get migrated? > So if you accept this intermediary step you can easily use ZONE_DEVICE for > device memory. This way no lru, no complex dance to make the memory out of > reach from regular memory allocator. > > I think we would have much to gain if we pool our effort on a single common > solution for device memory. In my case the device memory is not accessible > by the CPU (because PCIE restrictions), in your case it is. Thus the only > difference is that in my case it can not be map inside the CPU page table > while in yours it can. > I think thats a good idea to pool our efforts at the same time making progress >> >> (1) Traditional ioremap >> >> a. Memory is mapped into kernel (linear and virtual) and user space >> b. These PFNs do not have struct pages associated with it >> c. These special PFNs are marked with special flags inside the PTE >> d. Cannot participate in core VM functions much because of this >> e. Cannot do easy user space migrations >> >> (2) Zone ZONE_DEVICE >> >> a. Memory is mapped into kernel and user space >> b. PFNs do have struct pages associated with it >> c. These struct pages are allocated inside it's own memory range >> d. Unfortunately the struct page's union containing LRU has been >> used for struct dev_pagemap pointer >> e. Hence it cannot be part of any LRU (like Page cache) >> f. Hence file cached mapping cannot reside on these PFNs >> g. Cannot do easy migrations >> >> I had also explored non LRU representation of this coherent device >> memory where the integration with system RAM in the core VM is limited only >> to the following functions. Not being inside LRU is definitely going to >> reduce the scope of tight integration with system RAM. >> >> (1) Migration support between system RAM and coherent memory >> (2) Migration support between various coherent memory nodes >> (3) Isolation of the coherent memory >> (4) Mapping the coherent memory into user space through driver's >> struct vm_operations >> (5) HW poisoning of the coherent memory >> >> Allocating the entire memory of the coherent device node right >> after hot plug into ZONE_MOVABLE (where the memory is already inside the >> buddy system) will still expose a time window where other user space >> allocations can come into the coherent device memory node and prevent the >> intended isolation. So traditional hot plug is not the solution. Hence >> started looking into CMA based non LRU solution but then hit the following >> roadblocks. >> >> (1) CMA does not support hot plugging of new memory node >> a. CMA area needs to be marked during boot before buddy is >> initialized >> b. cma_alloc()/cma_release() can happen on the marked area >> c. Should be able to mark the CMA areas just after memory hot plug >> d. cma_alloc()/cma_release() can happen later after the hot plug >> e. This is not currently supported right now >> >> (2) Mapped non LRU migration of pages >> a. Recent work from Michan Kim makes non LRU page migratable >> b. But it still does not support migration of mapped non LRU pages >> c. With non LRU CMA reserved, again there are some additional >> challenges >> >> With hot pluggable CMA and non LRU mapped migration support there >> may be an alternate approach to represent coherent device memory. Please >> do review this RFC proposal and let me know your comments or suggestions. >> Thank you. > > You can take a look at hmm-v13 if you want to see how i do non LRU page > migration. While i put most of the migration code inside hmm_migrate.c it > could easily be move to migrate.c without hmm_ prefix. > > There is 2 missing piece with existing migrate code. First is to put memory > allocation for destination under control of who call the migrate code. Second > is to allow offloading the copy operation to device (ie not use the CPU to > copy data). > > I believe same requirement also make sense for platform you are targeting. > Thus same code can be use. > > hmm-v13 https://cgit.freedesktop.org/~glisse/linux/log/?h=hmm-v13 > Thanks for the link > I haven't posted this patchset yet because we are doing some modifications > to the device driver API to accomodate some new features. But the ZONE_DEVICE > changes and the overall migration code will stay the same more or less (i have > patches that move it to migrate.c and share more code with existing migrate > code). > > If you think i missed anything about lru and page cache please point it to > me. Because when i audited code for that i didn't see any road block with > the few fs i was looking at (ext4, xfs and core page cache code). > >> [...] > > Cheers, > Jérôme > Cheers, Balbir Singh.
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <j.glisse@gmail.com> |
|---|---|
| Date | 2016-10-25 17:30 +0200 |
| Message-ID | <swbPP-7oR-13@gated-at.bofh.it> |
| In reply to | #1508235 |
On Tue, Oct 25, 2016 at 11:07:39PM +1100, Balbir Singh wrote: > On 25/10/16 04:09, Jerome Glisse wrote: > > On Mon, Oct 24, 2016 at 10:01:49AM +0530, Anshuman Khandual wrote: > > > >> [...] > > > >> Core kernel memory features like reclamation, evictions etc. might > >> need to be restricted or modified on the coherent device memory node as > >> they can be performance limiting. The RFC does not propose anything on this > >> yet but it can be looked into later on. For now it just disables Auto NUMA > >> for any VMA which has coherent device memory. > >> > >> Seamless integration of coherent device memory with system memory > >> will enable various other features, some of which can be listed as follows. > >> > >> a. Seamless migrations between system RAM and the coherent memory > >> b. Will have asynchronous and high throughput migrations > >> c. Be able to allocate huge order pages from these memory regions > >> d. Restrict allocations to a large extent to the tasks using the > >> device for workload acceleration > >> > >> Before concluding, will look into the reasons why the existing > >> solutions don't work. There are two basic requirements which have to be > >> satisfies before the coherent device memory can be integrated with core > >> kernel seamlessly. > >> > >> a. PFN must have struct page > >> b. Struct page must able to be inside standard LRU lists > >> > >> The above two basic requirements discard the existing method of > >> device memory representation approaches like these which then requires the > >> need of creating a new framework. > > > > I do not believe the LRU list is a hard requirement, yes when faulting in > > a page inside the page cache it assumes it needs to be added to lru list. > > But i think this can easily be work around. > > > > In HMM i am using ZONE_DEVICE and because memory is not accessible from CPU > > (not everyone is bless with decent system bus like CAPI, CCIX, Gen-Z, ...) > > so in my case a file back page must always be spawn first from a regular > > page and once read from disk then i can migrate to GPU page. > > > > I've not seen the HMM patchset, but read from disk will go to ZONE_DEVICE? > Then get migrated? Because in my case device memory is not accessible by anything except the device (not entirely true but for sake of design it is) any page read from disk will be first read into regular page (from regular system memory). It is only once it is uptodate and in page cache that it can be migrated to a ZONE_DEVICE page. So read from disk use an intermediary page. Write back is kind of the same i plan on using a bounce page by leveraging existing bio bounce infrastructure. Cheers, Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Dave Hansen <dave.hansen@intel.com> |
|---|---|
| Date | 2016-10-24 20:10 +0200 |
| Message-ID | <svRR9-2Kt-73@gated-at.bofh.it> |
| In reply to | #1506875 |
On 10/23/2016 09:31 PM, Anshuman Khandual wrote: > To achieve seamless integration between system RAM and coherent > device memory it must be able to utilize core memory kernel features like > anon mapping, file mapping, page cache, driver managed pages, HW poisoning, > migrations, reclaim, compaction, etc. So, you need to support all these things, but not autonuma or hugetlbfs? What's the reasoning behind that? If you *really* don't want a "cdm" page to be migrated, then why isn't that policy set on the VMA in the first place? That would keep "cdm" pages from being made non-cdm. And, why would autonuma ever make a non-cdm page and migrate it in to cdm? There will be no NUMA access faults caused by the devices that are fed to autonuma. I'm confused.
[toc] | [prev] | [next] | [standalone]
| From | David Nellans <dnellans@nvidia.com> |
|---|---|
| Date | 2016-10-24 20:40 +0200 |
| Message-ID | <svSka-2XX-27@gated-at.bofh.it> |
| In reply to | #1507526 |
On 10/24/2016 01:04 PM, Dave Hansen wrote: > On 10/23/2016 09:31 PM, Anshuman Khandual wrote: >> To achieve seamless integration between system RAM and coherent >> device memory it must be able to utilize core memory kernel features like >> anon mapping, file mapping, page cache, driver managed pages, HW poisoning, >> migrations, reclaim, compaction, etc. > So, you need to support all these things, but not autonuma or hugetlbfs? > What's the reasoning behind that? > > If you *really* don't want a "cdm" page to be migrated, then why isn't > that policy set on the VMA in the first place? That would keep "cdm" > pages from being made non-cdm. And, why would autonuma ever make a > non-cdm page and migrate it in to cdm? There will be no NUMA access > faults caused by the devices that are fed to autonuma. > Pages are desired to be migrateable, both into (starting cpu zone movable->cdm) and out of (starting cdm->cpu zone movable) but only through explicit migration, not via autonuma. other pages in the same VMA should still be migrateable between CPU nodes via autonuma however. Its expected a lot of these allocations are going to end up in THPs. I'm not sure we need to explicitly disallow hugetlbfs support but the identified use case is definitely via THPs not tlbfs.
[toc] | [prev] | [next] | [standalone]
| From | Dave Hansen <dave.hansen@intel.com> |
|---|---|
| Date | 2016-10-24 21:40 +0200 |
| Message-ID | <svTge-3ya-19@gated-at.bofh.it> |
| In reply to | #1507570 |
On 10/24/2016 11:32 AM, David Nellans wrote: > On 10/24/2016 01:04 PM, Dave Hansen wrote: >> If you *really* don't want a "cdm" page to be migrated, then why isn't >> that policy set on the VMA in the first place? That would keep "cdm" >> pages from being made non-cdm. And, why would autonuma ever make a >> non-cdm page and migrate it in to cdm? There will be no NUMA access >> faults caused by the devices that are fed to autonuma. >> > Pages are desired to be migrateable, both into (starting cpu zone > movable->cdm) and out of (starting cdm->cpu zone movable) but only > through explicit migration, not via autonuma. OK, and is there a reason that the existing mbind code plus NUMA policies fails to give you this behavior? Does autonuma somehow override strict NUMA binding? > other pages in the same > VMA should still be migrateable between CPU nodes via autonuma however. That's not the way the implementation here works, as I understand it. See the VM_CDM patch and my responses to it. > Its expected a lot of these allocations are going to end up in THPs. > I'm not sure we need to explicitly disallow hugetlbfs support but the > identified use case is definitely via THPs not tlbfs. I think THP and hugetlbfs are implementations, not use cases. :) Is it too hard to support hugetlbfs that we should complicate its code to exclude it from this type of memory? Why?
[toc] | [prev] | [standalone]
Page 3 of 3 — ← Prev page 1 2 [3]
Back to top | Article view | linux.kernel
csiph-web