Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1631814 > unrolled thread
| Started by | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| First post | 2017-04-27 03:10 +0200 |
| Last post | 2017-04-27 18:20 +0200 |
| Articles | 15 on this page of 35 — 7 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
[PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-04-27 03:10 +0200
Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-04-27 10:40 +0200
Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Ingo Molnar <mingo@kernel.org> - 2017-04-28 08:40 +0200
[PATCH] mm, zone_device: Replace {get, put}_zone_device_page() with a single reference "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2017-04-28 10:20 +0200
[PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-04-28 19:30 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Jerome Glisse <jglisse@redhat.com> - 2017-04-28 19:40 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-04-28 19:50 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Jerome Glisse <jglisse@redhat.com> - 2017-04-28 20:10 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-04-28 21:10 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Jerome Glisse <jglisse@redhat.com> - 2017-04-28 21:20 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-04-28 21:30 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Jerome Glisse <jglisse@redhat.com> - 2017-04-28 21:40 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-04-29 12:20 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Jerome Glisse <jglisse@redhat.com> - 2017-05-01 01:20 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-05-01 03:50 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Jerome Glisse <jglisse@redhat.com> - 2017-05-01 04:00 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-05-01 04:50 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Logan Gunthorpe <logang@deltatee.com> - 2017-05-01 05:50 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-05-01 12:50 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Jerome Glisse <jglisse@redhat.com> - 2017-05-01 16:00 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-05-01 22:30 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Jerome Glisse <jglisse@redhat.com> - 2017-05-01 22:40 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-05-02 13:50 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Jerome Glisse <jglisse@redhat.com> - 2017-05-02 15:30 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Ingo Molnar <mingo@kernel.org> - 2017-04-29 16:20 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-05-01 04:50 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Ingo Molnar <mingo@kernel.org> - 2017-05-01 09:20 +0200
Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-05-01 11:40 +0200
[tip:x86/mm] mm, zone_device: Replace {get, put}_zone_device_page() with a single reference to fix pmem crash tip-bot for Dan Williams <tipbot@zytor.com> - 2017-05-01 10:40 +0200
Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-04-27 18:20 +0200
Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Logan Gunthorpe <logang@deltatee.com> - 2017-04-27 18:40 +0200
Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-04-27 18:40 +0200
Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Dan Williams <dan.j.williams@intel.com> - 2017-04-27 18:50 +0200
Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Logan Gunthorpe <logang@deltatee.com> - 2017-04-27 18:50 +0200
Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference Logan Gunthorpe <logang@deltatee.com> - 2017-04-27 18:20 +0200
Page 2 of 2 — ← Prev page 1 [2]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-05-01 22:30 +0200 |
| Subject | Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tCq7f-5fl-7@gated-at.bofh.it> |
| In reply to | #1633676 |
On Mon, May 1, 2017 at 6:55 AM, Jerome Glisse <jglisse@redhat.com> wrote: > On Mon, May 01, 2017 at 01:23:59PM +0300, Kirill A. Shutemov wrote: >> On Sun, Apr 30, 2017 at 07:14:24PM -0400, Jerome Glisse wrote: >> > On Sat, Apr 29, 2017 at 01:17:26PM +0300, Kirill A. Shutemov wrote: >> > > On Fri, Apr 28, 2017 at 03:33:07PM -0400, Jerome Glisse wrote: >> > > > On Fri, Apr 28, 2017 at 12:22:24PM -0700, Dan Williams wrote: >> > > > > Are you sure about needing to hook the 2 -> 1 transition? Could we >> > > > > change ZONE_DEVICE pages to not have an elevated reference count when >> > > > > they are created so you can keep the HMM references out of the mm hot >> > > > > path? >> > > > >> > > > 100% sure on that :) I need to callback into driver for 2->1 transition >> > > > no way around that. If we change ZONE_DEVICE to not have an elevated >> > > > reference count that you need to make a lot more change to mm so that >> > > > ZONE_DEVICE is never use as fallback for memory allocation. Also need >> > > > to make change to be sure that ZONE_DEVICE page never endup in one of >> > > > the path that try to put them back on lru. There is a lot of place that >> > > > would need to be updated and it would be highly intrusive and add a >> > > > lot of special cases to other hot code path. >> > > >> > > Could you explain more on where the requirement comes from or point me to >> > > where I can read about this. >> > > >> > >> > HMM ZONE_DEVICE pages are use like other pages (anonymous or file back page) >> > in _any_ vma. So i need to know when a page is freed ie either as result of >> > unmap, exit or migration or anything that would free the memory. For zone >> > device a page is free once its refcount reach 1 so i need to catch refcount >> > transition from 2->1 >> >> What if we would rework zone device to have pages with refcount 0 at >> start? > > That is a _lot_ of work from top of my head because it would need changes > to a lot of places and likely more hot code path that simply adding some- > thing to put_page() note that i only need something in put_page() i do not > need anything in the get page path. Is adding a conditional branch for > HMM pages in put_page() that much of a problem ? > > >> > This is the only way i can inform the device that the page is now free. See >> > >> > https://cgit.freedesktop.org/~glisse/linux/commit/?h=hmm-v21&id=52da8fe1a088b87b5321319add79e43b8372ed7d >> > >> > There is _no_ way around that. >> >> I'm still not convinced that it's impossible. >> >> Could you describe lifecycle for pages in case of HMM? > > Process malloc something, end it over to some function in the program > that use the GPU that function call GPU API (OpenCL, CUDA, ...) that > trigger a migration to device memory. > > So in the kernel you get a migration like any existing migration, > original page is unmap, if refcount is all ok (no pin) then a device > page is allocated and thing are migrated to device memory. > > What happen after is unknown. Either userspace/kernel driver decide > to migrate back to system memory, either there is an munmap, either > there is a CPU page fault, ... So from that point on the device page > as the exact same life as a regular page. > > Above i describe the migrate case, but you can also have new memory > allocation that directly allocate device memory. For instance if the > GPU do a page fault on an address that isn't back by anything then > we can directly allocate a device page. No migration involve in that > case. > > HMM pages are like any other pages in most respect. Exception are: > - no GUP > - no KSM > - no lru reclaim > - no NUMA balancing > - no regular migration (existing migrate_page) > > The fact that minimum refcount for ZONE_DEVICE is 1 already gives > us for free most of the above exception. To convert the refcount to > be like other pages would mean that all of the above would need to > be audited and probably modify to ignore ZONE_DEVICE pages (i am > pretty sure Dan do not want any of the above either). Right, adding HMM references to get_page() and put_page() seems less intrusive. Given how uncommon HMM hardware is (insert grumble about no visible upstream user of this functionality) I think the 'static branch' approach helps mitigate the impact for everything else. Looking back, I should have used that mechanism for the pmem use case, but it's moot now.
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2017-05-01 22:40 +0200 |
| Subject | Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tCqgV-5ie-9@gated-at.bofh.it> |
| In reply to | #1633854 |
On Mon, May 01, 2017 at 01:19:24PM -0700, Dan Williams wrote: > On Mon, May 1, 2017 at 6:55 AM, Jerome Glisse <jglisse@redhat.com> wrote: > > On Mon, May 01, 2017 at 01:23:59PM +0300, Kirill A. Shutemov wrote: > >> On Sun, Apr 30, 2017 at 07:14:24PM -0400, Jerome Glisse wrote: > >> > On Sat, Apr 29, 2017 at 01:17:26PM +0300, Kirill A. Shutemov wrote: > >> > > On Fri, Apr 28, 2017 at 03:33:07PM -0400, Jerome Glisse wrote: > >> > > > On Fri, Apr 28, 2017 at 12:22:24PM -0700, Dan Williams wrote: > >> > > > > Are you sure about needing to hook the 2 -> 1 transition? Could we > >> > > > > change ZONE_DEVICE pages to not have an elevated reference count when > >> > > > > they are created so you can keep the HMM references out of the mm hot > >> > > > > path? > >> > > > > >> > > > 100% sure on that :) I need to callback into driver for 2->1 transition > >> > > > no way around that. If we change ZONE_DEVICE to not have an elevated > >> > > > reference count that you need to make a lot more change to mm so that > >> > > > ZONE_DEVICE is never use as fallback for memory allocation. Also need > >> > > > to make change to be sure that ZONE_DEVICE page never endup in one of > >> > > > the path that try to put them back on lru. There is a lot of place that > >> > > > would need to be updated and it would be highly intrusive and add a > >> > > > lot of special cases to other hot code path. > >> > > > >> > > Could you explain more on where the requirement comes from or point me to > >> > > where I can read about this. > >> > > > >> > > >> > HMM ZONE_DEVICE pages are use like other pages (anonymous or file back page) > >> > in _any_ vma. So i need to know when a page is freed ie either as result of > >> > unmap, exit or migration or anything that would free the memory. For zone > >> > device a page is free once its refcount reach 1 so i need to catch refcount > >> > transition from 2->1 > >> > >> What if we would rework zone device to have pages with refcount 0 at > >> start? > > > > That is a _lot_ of work from top of my head because it would need changes > > to a lot of places and likely more hot code path that simply adding some- > > thing to put_page() note that i only need something in put_page() i do not > > need anything in the get page path. Is adding a conditional branch for > > HMM pages in put_page() that much of a problem ? > > > > > >> > This is the only way i can inform the device that the page is now free. See > >> > > >> > https://cgit.freedesktop.org/~glisse/linux/commit/?h=hmm-v21&id=52da8fe1a088b87b5321319add79e43b8372ed7d > >> > > >> > There is _no_ way around that. > >> > >> I'm still not convinced that it's impossible. > >> > >> Could you describe lifecycle for pages in case of HMM? > > > > Process malloc something, end it over to some function in the program > > that use the GPU that function call GPU API (OpenCL, CUDA, ...) that > > trigger a migration to device memory. > > > > So in the kernel you get a migration like any existing migration, > > original page is unmap, if refcount is all ok (no pin) then a device > > page is allocated and thing are migrated to device memory. > > > > What happen after is unknown. Either userspace/kernel driver decide > > to migrate back to system memory, either there is an munmap, either > > there is a CPU page fault, ... So from that point on the device page > > as the exact same life as a regular page. > > > > Above i describe the migrate case, but you can also have new memory > > allocation that directly allocate device memory. For instance if the > > GPU do a page fault on an address that isn't back by anything then > > we can directly allocate a device page. No migration involve in that > > case. > > > > HMM pages are like any other pages in most respect. Exception are: > > - no GUP > > - no KSM > > - no lru reclaim > > - no NUMA balancing > > - no regular migration (existing migrate_page) > > > > The fact that minimum refcount for ZONE_DEVICE is 1 already gives > > us for free most of the above exception. To convert the refcount to > > be like other pages would mean that all of the above would need to > > be audited and probably modify to ignore ZONE_DEVICE pages (i am > > pretty sure Dan do not want any of the above either). > > Right, adding HMM references to get_page() and put_page() seems less > intrusive. Given how uncommon HMM hardware is (insert grumble about no > visible upstream user of this functionality) I think the 'static > branch' approach helps mitigate the impact for everything else. > Looking back, I should have used that mechanism for the pmem use case, > but it's moot now. I do not need anything in get_page() all i need is something in put_page() to catch the 2 -> 1 refcount transition to know when a page is freed. Cheers, Jérôme
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2017-05-02 13:50 +0200 |
| Subject | Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tCEtz-6aQ-1@gated-at.bofh.it> |
| In reply to | #1633676 |
On Mon, May 01, 2017 at 09:55:48AM -0400, Jerome Glisse wrote: > On Mon, May 01, 2017 at 01:23:59PM +0300, Kirill A. Shutemov wrote: > > On Sun, Apr 30, 2017 at 07:14:24PM -0400, Jerome Glisse wrote: > > > On Sat, Apr 29, 2017 at 01:17:26PM +0300, Kirill A. Shutemov wrote: > > > > On Fri, Apr 28, 2017 at 03:33:07PM -0400, Jerome Glisse wrote: > > > > > On Fri, Apr 28, 2017 at 12:22:24PM -0700, Dan Williams wrote: > > > > > > Are you sure about needing to hook the 2 -> 1 transition? Could we > > > > > > change ZONE_DEVICE pages to not have an elevated reference count when > > > > > > they are created so you can keep the HMM references out of the mm hot > > > > > > path? > > > > > > > > > > 100% sure on that :) I need to callback into driver for 2->1 transition > > > > > no way around that. If we change ZONE_DEVICE to not have an elevated > > > > > reference count that you need to make a lot more change to mm so that > > > > > ZONE_DEVICE is never use as fallback for memory allocation. Also need > > > > > to make change to be sure that ZONE_DEVICE page never endup in one of > > > > > the path that try to put them back on lru. There is a lot of place that > > > > > would need to be updated and it would be highly intrusive and add a > > > > > lot of special cases to other hot code path. > > > > > > > > Could you explain more on where the requirement comes from or point me to > > > > where I can read about this. > > > > > > > > > > HMM ZONE_DEVICE pages are use like other pages (anonymous or file back page) > > > in _any_ vma. So i need to know when a page is freed ie either as result of > > > unmap, exit or migration or anything that would free the memory. For zone > > > device a page is free once its refcount reach 1 so i need to catch refcount > > > transition from 2->1 > > > > What if we would rework zone device to have pages with refcount 0 at > > start? > > That is a _lot_ of work from top of my head because it would need changes > to a lot of places and likely more hot code path that simply adding some- > thing to put_page() note that i only need something in put_page() i do not > need anything in the get page path. Is adding a conditional branch for > HMM pages in put_page() that much of a problem ? Well, it gets inlined everywhere. Removing zone_device code from get_page() and put_page() saved non-trivial ~140k in vmlinux for allyesconfig. Re-introducing part this bloat would be unfortunate. > > > This is the only way i can inform the device that the page is now free. See > > > > > > https://cgit.freedesktop.org/~glisse/linux/commit/?h=hmm-v21&id=52da8fe1a088b87b5321319add79e43b8372ed7d > > > > > > There is _no_ way around that. > > > > I'm still not convinced that it's impossible. > > > > Could you describe lifecycle for pages in case of HMM? > > Process malloc something, end it over to some function in the program > that use the GPU that function call GPU API (OpenCL, CUDA, ...) that > trigger a migration to device memory. > > So in the kernel you get a migration like any existing migration, > original page is unmap, if refcount is all ok (no pin) then a device > page is allocated and thing are migrated to device memory. > > What happen after is unknown. Either userspace/kernel driver decide > to migrate back to system memory, either there is an munmap, either > there is a CPU page fault, ... So from that point on the device page > as the exact same life as a regular page. > > Above i describe the migrate case, but you can also have new memory > allocation that directly allocate device memory. For instance if the > GPU do a page fault on an address that isn't back by anything then > we can directly allocate a device page. No migration involve in that > case. > > HMM pages are like any other pages in most respect. Exception are: > - no GUP Hm. How do you exclude GUP? And why is it required? -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2017-05-02 15:30 +0200 |
| Subject | Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tCG2m-7dm-19@gated-at.bofh.it> |
| In reply to | #1634358 |
On Tue, May 02, 2017 at 02:37:46PM +0300, Kirill A. Shutemov wrote: > On Mon, May 01, 2017 at 09:55:48AM -0400, Jerome Glisse wrote: > > On Mon, May 01, 2017 at 01:23:59PM +0300, Kirill A. Shutemov wrote: > > > On Sun, Apr 30, 2017 at 07:14:24PM -0400, Jerome Glisse wrote: > > > > On Sat, Apr 29, 2017 at 01:17:26PM +0300, Kirill A. Shutemov wrote: > > > > > On Fri, Apr 28, 2017 at 03:33:07PM -0400, Jerome Glisse wrote: > > > > > > On Fri, Apr 28, 2017 at 12:22:24PM -0700, Dan Williams wrote: > > > > > > > Are you sure about needing to hook the 2 -> 1 transition? Could we > > > > > > > change ZONE_DEVICE pages to not have an elevated reference count when > > > > > > > they are created so you can keep the HMM references out of the mm hot > > > > > > > path? > > > > > > > > > > > > 100% sure on that :) I need to callback into driver for 2->1 transition > > > > > > no way around that. If we change ZONE_DEVICE to not have an elevated > > > > > > reference count that you need to make a lot more change to mm so that > > > > > > ZONE_DEVICE is never use as fallback for memory allocation. Also need > > > > > > to make change to be sure that ZONE_DEVICE page never endup in one of > > > > > > the path that try to put them back on lru. There is a lot of place that > > > > > > would need to be updated and it would be highly intrusive and add a > > > > > > lot of special cases to other hot code path. > > > > > > > > > > Could you explain more on where the requirement comes from or point me to > > > > > where I can read about this. > > > > > > > > > > > > > HMM ZONE_DEVICE pages are use like other pages (anonymous or file back page) > > > > in _any_ vma. So i need to know when a page is freed ie either as result of > > > > unmap, exit or migration or anything that would free the memory. For zone > > > > device a page is free once its refcount reach 1 so i need to catch refcount > > > > transition from 2->1 > > > > > > What if we would rework zone device to have pages with refcount 0 at > > > start? > > > > That is a _lot_ of work from top of my head because it would need changes > > to a lot of places and likely more hot code path that simply adding some- > > thing to put_page() note that i only need something in put_page() i do not > > need anything in the get page path. Is adding a conditional branch for > > HMM pages in put_page() that much of a problem ? > > Well, it gets inlined everywhere. Removing zone_device code from > get_page() and put_page() saved non-trivial ~140k in vmlinux for > allyesconfig. > > Re-introducing part this bloat would be unfortunate. > > > > > This is the only way i can inform the device that the page is now free. See > > > > > > > > https://cgit.freedesktop.org/~glisse/linux/commit/?h=hmm-v21&id=52da8fe1a088b87b5321319add79e43b8372ed7d > > > > > > > > There is _no_ way around that. > > > > > > I'm still not convinced that it's impossible. > > > > > > Could you describe lifecycle for pages in case of HMM? > > > > Process malloc something, end it over to some function in the program > > that use the GPU that function call GPU API (OpenCL, CUDA, ...) that > > trigger a migration to device memory. > > > > So in the kernel you get a migration like any existing migration, > > original page is unmap, if refcount is all ok (no pin) then a device > > page is allocated and thing are migrated to device memory. > > > > What happen after is unknown. Either userspace/kernel driver decide > > to migrate back to system memory, either there is an munmap, either > > there is a CPU page fault, ... So from that point on the device page > > as the exact same life as a regular page. > > > > Above i describe the migrate case, but you can also have new memory > > allocation that directly allocate device memory. For instance if the > > GPU do a page fault on an address that isn't back by anything then > > we can directly allocate a device page. No migration involve in that > > case. > > > > HMM pages are like any other pages in most respect. Exception are: > > - no GUP > > Hm. How do you exclude GUP? And why is it required? Well it is not forbiden it just can not happen simply because as device memory is not accessible by CPU then the corresponding CPU page table entry is a special entry and thus GUP trigger a page fault that migrate thing back to regular memory. The why is simply because we need to always be able to migrate back to regular memory as device memory is not accessible by the CPU. So we can not allow anyone to pin it. Cheers, Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Ingo Molnar <mingo@kernel.org> |
|---|---|
| Date | 2017-04-29 16:20 +0200 |
| Subject | Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tBBo5-7iD-1@gated-at.bofh.it> |
| In reply to | #1633041 |
* Dan Williams <dan.j.williams@intel.com> wrote:
> Kirill points out that the calls to {get,put}_dev_pagemap() can be
> removed from the mm fast path if we take a single get_dev_pagemap()
> reference to signify that the page is alive and use the final put of the
> page to drop that reference.
>
> This does require some care to make sure that any waits for the
> percpu_ref to drop to zero occur *after* devm_memremap_page_release(),
> since it now maintains its own elevated reference.
>
> Cc: Ingo Molnar <mingo@redhat.com>
> Cc: Jérôme Glisse <jglisse@redhat.com>
> Cc: Andrew Morton <akpm@linux-foundation.org>
> Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> Suggested-by: Kirill Shutemov <kirill.shutemov@linux.intel.com>
> Tested-by: Kirill Shutemov <kirill.shutemov@linux.intel.com>
> Signed-off-by: Dan Williams <dan.j.williams@intel.com>
This changelog is lacking an explanation about how this solves the crashes you
were seeing.
Thanks,
Ingo
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-05-01 04:50 +0200 |
| Subject | Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tC9zs-3ad-7@gated-at.bofh.it> |
| In reply to | #1633299 |
On Sat, Apr 29, 2017 at 7:18 AM, Ingo Molnar <mingo@kernel.org> wrote:
>
> * Dan Williams <dan.j.williams@intel.com> wrote:
>
>> Kirill points out that the calls to {get,put}_dev_pagemap() can be
>> removed from the mm fast path if we take a single get_dev_pagemap()
>> reference to signify that the page is alive and use the final put of the
>> page to drop that reference.
>>
>> This does require some care to make sure that any waits for the
>> percpu_ref to drop to zero occur *after* devm_memremap_page_release(),
>> since it now maintains its own elevated reference.
>>
>> Cc: Ingo Molnar <mingo@redhat.com>
>> Cc: Jérôme Glisse <jglisse@redhat.com>
>> Cc: Andrew Morton <akpm@linux-foundation.org>
>> Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
>> Suggested-by: Kirill Shutemov <kirill.shutemov@linux.intel.com>
>> Tested-by: Kirill Shutemov <kirill.shutemov@linux.intel.com>
>> Signed-off-by: Dan Williams <dan.j.williams@intel.com>
>
> This changelog is lacking an explanation about how this solves the crashes you
> were seeing.
>
Kirill? It wasn't clear to me why the conversion to generic
get_user_pages_fast() caused the reference counts to be off.
[toc] | [prev] | [next] | [standalone]
| From | Ingo Molnar <mingo@kernel.org> |
|---|---|
| Date | 2017-05-01 09:20 +0200 |
| Subject | Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tCdMK-626-9@gated-at.bofh.it> |
| In reply to | #1633528 |
* Dan Williams <dan.j.williams@intel.com> wrote:
> On Sat, Apr 29, 2017 at 7:18 AM, Ingo Molnar <mingo@kernel.org> wrote:
> >
> > * Dan Williams <dan.j.williams@intel.com> wrote:
> >
> >> Kirill points out that the calls to {get,put}_dev_pagemap() can be
> >> removed from the mm fast path if we take a single get_dev_pagemap()
> >> reference to signify that the page is alive and use the final put of the
> >> page to drop that reference.
> >>
> >> This does require some care to make sure that any waits for the
> >> percpu_ref to drop to zero occur *after* devm_memremap_page_release(),
> >> since it now maintains its own elevated reference.
> >>
> >> Cc: Ingo Molnar <mingo@redhat.com>
> >> Cc: Jérôme Glisse <jglisse@redhat.com>
> >> Cc: Andrew Morton <akpm@linux-foundation.org>
> >> Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
> >> Suggested-by: Kirill Shutemov <kirill.shutemov@linux.intel.com>
> >> Tested-by: Kirill Shutemov <kirill.shutemov@linux.intel.com>
> >> Signed-off-by: Dan Williams <dan.j.williams@intel.com>
> >
> > This changelog is lacking an explanation about how this solves the crashes you
> > were seeing.
>
> Kirill? It wasn't clear to me why the conversion to generic
> get_user_pages_fast() caused the reference counts to be off.
Ok, the merge window is open and we really need this fix for x86/mm, so this is
what I've decoded:
The x86 conversion to the generic GUP code included a small change which causes
crashes and data corruption in the pmem code - not good.
The root cause is that the /dev/pmem driver code implicitly relies on the x86
get_user_pages() implementation doing a get_page() on the page refcount, because
get_page() does a get_zone_device_page() which properly refcounts pmem's separate
page struct arrays that are not present in the regular page struct structures.
(The pmem driver does this because it can cover huge memory areas.)
But the x86 conversion to the generic GUP code changed the get_page() to
page_cache_get_speculative() which is faster but doesn't do the
get_zone_device_page() call the pmem code relies on.
One way to solve the regression would be to change the generic GUP code to use
get_page(), but that would slow things down a bit and punish other generic-GUP
using architectures for an x86-ism they did not care about. (Arguably the pmem
driver was probably not working reliably for them: but nvdimm is an Intel
feature, so non-x86 exposure is probably still limited.)
So restructure the pmem code's interface with the MM instead: get rid of the
get/put_zone_device_page() distinction, integrate put_zone_device_page() into
__put_page() and and restructure the pmem completion-wait and teardown machinery.
This speeds up things while also making the pmem refcounting more robust going
forward.
... is this extension to the changelog correct?
I'll apply this for the time being - but can still amend the text before sending
it to Linus later today.
Thanks,
Ingo
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2017-05-01 11:40 +0200 |
| Subject | Re: [PATCH v2] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tCfYe-7mq-5@gated-at.bofh.it> |
| In reply to | #1633566 |
On Mon, May 01, 2017 at 09:12:59AM +0200, Ingo Molnar wrote: > ... is this extension to the changelog correct? Looks good to me. -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | tip-bot for Dan Williams <tipbot@zytor.com> |
|---|---|
| Date | 2017-05-01 10:40 +0200 |
| Subject | [tip:x86/mm] mm, zone_device: Replace {get, put}_zone_device_page() with a single reference to fix pmem crash |
| Message-ID | <tCf2a-6MZ-11@gated-at.bofh.it> |
| In reply to | #1633041 |
Commit-ID: 71389703839ebe9cb426c72d5f0bd549592e583c
Gitweb: http://git.kernel.org/tip/71389703839ebe9cb426c72d5f0bd549592e583c
Author: Dan Williams <dan.j.williams@intel.com>
AuthorDate: Fri, 28 Apr 2017 10:23:37 -0700
Committer: Ingo Molnar <mingo@kernel.org>
CommitDate: Mon, 1 May 2017 09:15:53 +0200
mm, zone_device: Replace {get, put}_zone_device_page() with a single reference to fix pmem crash
The x86 conversion to the generic GUP code included a small change which causes
crashes and data corruption in the pmem code - not good.
The root cause is that the /dev/pmem driver code implicitly relies on the x86
get_user_pages() implementation doing a get_page() on the page refcount, because
get_page() does a get_zone_device_page() which properly refcounts pmem's separate
page struct arrays that are not present in the regular page struct structures.
(The pmem driver does this because it can cover huge memory areas.)
But the x86 conversion to the generic GUP code changed the get_page() to
page_cache_get_speculative() which is faster but doesn't do the
get_zone_device_page() call the pmem code relies on.
One way to solve the regression would be to change the generic GUP code to use
get_page(), but that would slow things down a bit and punish other generic-GUP
using architectures for an x86-ism they did not care about. (Arguably the pmem
driver was probably not working reliably for them: but nvdimm is an Intel
feature, so non-x86 exposure is probably still limited.)
So restructure the pmem code's interface with the MM instead: get rid of the
get/put_zone_device_page() distinction, integrate put_zone_device_page() into
__put_page() and and restructure the pmem completion-wait and teardown machinery:
Kirill points out that the calls to {get,put}_dev_pagemap() can be
removed from the mm fast path if we take a single get_dev_pagemap()
reference to signify that the page is alive and use the final put of the
page to drop that reference.
This does require some care to make sure that any waits for the
percpu_ref to drop to zero occur *after* devm_memremap_page_release(),
since it now maintains its own elevated reference.
This speeds up things while also making the pmem refcounting more robust going
forward.
Suggested-by: Kirill Shutemov <kirill.shutemov@linux.intel.com>
Tested-by: Kirill Shutemov <kirill.shutemov@linux.intel.com>
Signed-off-by: Dan Williams <dan.j.williams@intel.com>
Reviewed-by: Logan Gunthorpe <logang@deltatee.com>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Andy Lutomirski <luto@kernel.org>
Cc: Borislav Petkov <bp@alien8.de>
Cc: Brian Gerst <brgerst@gmail.com>
Cc: Denys Vlasenko <dvlasenk@redhat.com>
Cc: H. Peter Anvin <hpa@zytor.com>
Cc: Josh Poimboeuf <jpoimboe@redhat.com>
Cc: Jérôme Glisse <jglisse@redhat.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: linux-mm@kvack.org
Link: http://lkml.kernel.org/r/149339998297.24933.1129582806028305912.stgit@dwillia2-desk3.amr.corp.intel.com
Signed-off-by: Ingo Molnar <mingo@kernel.org>
---
drivers/dax/pmem.c | 2 +-
drivers/nvdimm/pmem.c | 13 +++++++++++--
include/linux/mm.h | 14 --------------
kernel/memremap.c | 22 +++++++++-------------
mm/swap.c | 10 ++++++++++
5 files changed, 31 insertions(+), 30 deletions(-)
diff --git a/drivers/dax/pmem.c b/drivers/dax/pmem.c
index 033f49b3..cb0d742 100644
--- a/drivers/dax/pmem.c
+++ b/drivers/dax/pmem.c
@@ -43,6 +43,7 @@ static void dax_pmem_percpu_exit(void *data)
struct dax_pmem *dax_pmem = to_dax_pmem(ref);
dev_dbg(dax_pmem->dev, "%s\n", __func__);
+ wait_for_completion(&dax_pmem->cmp);
percpu_ref_exit(ref);
}
@@ -53,7 +54,6 @@ static void dax_pmem_percpu_kill(void *data)
dev_dbg(dax_pmem->dev, "%s\n", __func__);
percpu_ref_kill(ref);
- wait_for_completion(&dax_pmem->cmp);
}
static int dax_pmem_probe(struct device *dev)
diff --git a/drivers/nvdimm/pmem.c b/drivers/nvdimm/pmem.c
index 5b536be..fb7bbc7 100644
--- a/drivers/nvdimm/pmem.c
+++ b/drivers/nvdimm/pmem.c
@@ -25,6 +25,7 @@
#include <linux/badblocks.h>
#include <linux/memremap.h>
#include <linux/vmalloc.h>
+#include <linux/blk-mq.h>
#include <linux/pfn_t.h>
#include <linux/slab.h>
#include <linux/pmem.h>
@@ -231,6 +232,11 @@ static void pmem_release_queue(void *q)
blk_cleanup_queue(q);
}
+static void pmem_freeze_queue(void *q)
+{
+ blk_mq_freeze_queue_start(q);
+}
+
static void pmem_release_disk(void *disk)
{
del_gendisk(disk);
@@ -284,6 +290,9 @@ static int pmem_attach_disk(struct device *dev,
if (!q)
return -ENOMEM;
+ if (devm_add_action_or_reset(dev, pmem_release_queue, q))
+ return -ENOMEM;
+
pmem->pfn_flags = PFN_DEV;
if (is_nd_pfn(dev)) {
addr = devm_memremap_pages(dev, &pfn_res, &q->q_usage_counter,
@@ -303,10 +312,10 @@ static int pmem_attach_disk(struct device *dev,
pmem->size, ARCH_MEMREMAP_PMEM);
/*
- * At release time the queue must be dead before
+ * At release time the queue must be frozen before
* devm_memremap_pages is unwound
*/
- if (devm_add_action_or_reset(dev, pmem_release_queue, q))
+ if (devm_add_action_or_reset(dev, pmem_freeze_queue, q))
return -ENOMEM;
if (IS_ERR(addr))
diff --git a/include/linux/mm.h b/include/linux/mm.h
index a835edd..695da2a 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -762,19 +762,11 @@ static inline enum zone_type page_zonenum(const struct page *page)
}
#ifdef CONFIG_ZONE_DEVICE
-void get_zone_device_page(struct page *page);
-void put_zone_device_page(struct page *page);
static inline bool is_zone_device_page(const struct page *page)
{
return page_zonenum(page) == ZONE_DEVICE;
}
#else
-static inline void get_zone_device_page(struct page *page)
-{
-}
-static inline void put_zone_device_page(struct page *page)
-{
-}
static inline bool is_zone_device_page(const struct page *page)
{
return false;
@@ -790,9 +782,6 @@ static inline void get_page(struct page *page)
*/
VM_BUG_ON_PAGE(page_ref_count(page) <= 0, page);
page_ref_inc(page);
-
- if (unlikely(is_zone_device_page(page)))
- get_zone_device_page(page);
}
static inline void put_page(struct page *page)
@@ -801,9 +790,6 @@ static inline void put_page(struct page *page)
if (put_page_testzero(page))
__put_page(page);
-
- if (unlikely(is_zone_device_page(page)))
- put_zone_device_page(page);
}
#if defined(CONFIG_SPARSEMEM) && !defined(CONFIG_SPARSEMEM_VMEMMAP)
diff --git a/kernel/memremap.c b/kernel/memremap.c
index 07e85e5..23a6483 100644
--- a/kernel/memremap.c
+++ b/kernel/memremap.c
@@ -182,18 +182,6 @@ struct page_map {
struct vmem_altmap altmap;
};
-void get_zone_device_page(struct page *page)
-{
- percpu_ref_get(page->pgmap->ref);
-}
-EXPORT_SYMBOL(get_zone_device_page);
-
-void put_zone_device_page(struct page *page)
-{
- put_dev_pagemap(page->pgmap);
-}
-EXPORT_SYMBOL(put_zone_device_page);
-
static void pgmap_radix_release(struct resource *res)
{
resource_size_t key, align_start, align_size, align_end;
@@ -237,6 +225,10 @@ static void devm_memremap_pages_release(struct device *dev, void *data)
struct resource *res = &page_map->res;
resource_size_t align_start, align_size;
struct dev_pagemap *pgmap = &page_map->pgmap;
+ unsigned long pfn;
+
+ for_each_device_pfn(pfn, page_map)
+ put_page(pfn_to_page(pfn));
if (percpu_ref_tryget_live(pgmap->ref)) {
dev_WARN(dev, "%s: page mapping is still live!\n", __func__);
@@ -277,7 +269,10 @@ struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
*
* Notes:
* 1/ @ref must be 'live' on entry and 'dead' before devm_memunmap_pages() time
- * (or devm release event).
+ * (or devm release event). The expected order of events is that @ref has
+ * been through percpu_ref_kill() before devm_memremap_pages_release(). The
+ * wait for the completion of all references being dropped and
+ * percpu_ref_exit() must occur after devm_memremap_pages_release().
*
* 2/ @res is expected to be a host memory range that could feasibly be
* treated as a "System RAM" range, i.e. not a device mmio range, but
@@ -379,6 +374,7 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
*/
list_del(&page->lru);
page->pgmap = pgmap;
+ percpu_ref_get(ref);
}
devres_add(dev, page_map);
return __va(res->start);
diff --git a/mm/swap.c b/mm/swap.c
index c4910f1..a4e6113 100644
--- a/mm/swap.c
+++ b/mm/swap.c
@@ -97,6 +97,16 @@ static void __put_compound_page(struct page *page)
void __put_page(struct page *page)
{
+ if (is_zone_device_page(page)) {
+ put_dev_pagemap(page->pgmap);
+
+ /*
+ * The page belongs to the device that created pgmap. Do
+ * not return it to page allocator.
+ */
+ return;
+ }
+
if (unlikely(PageCompound(page)))
__put_compound_page(page);
else
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-27 18:20 +0200 |
| Subject | Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tAUj7-338-11@gated-at.bofh.it> |
| In reply to | #1631814 |
On Thu, Apr 27, 2017 at 9:11 AM, Logan Gunthorpe <logang@deltatee.com> wrote:
>
>
> On 26/04/17 06:55 PM, Dan Williams wrote:
>> @@ -277,7 +269,10 @@ struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
>> *
>> * Notes:
>> * 1/ @ref must be 'live' on entry and 'dead' before devm_memunmap_pages() time
>> - * (or devm release event).
>> + * (or devm release event). The expected order of events is that @ref has
>> + * been through percpu_ref_kill() before devm_memremap_pages_release(). The
>> + * wait for the completion of kill and percpu_ref_exit() must occur after
>> + * devm_memremap_pages_release().
>> *
>> * 2/ @res is expected to be a host memory range that could feasibly be
>> * treated as a "System RAM" range, i.e. not a device mmio range, but
>> @@ -379,6 +374,7 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
>> */
>> list_del(&page->lru);
>> page->pgmap = pgmap;
>> + percpu_ref_get(ref);
>> }
>> devres_add(dev, page_map);
>> return __va(res->start);
>> diff --git a/mm/swap.c b/mm/swap.c
>> index 5dabf444d724..01267dda6668 100644
>> --- a/mm/swap.c
>> +++ b/mm/swap.c
>> @@ -97,6 +97,16 @@ static void __put_compound_page(struct page *page)
>>
>> void __put_page(struct page *page)
>> {
>> + if (is_zone_device_page(page)) {
>> + put_dev_pagemap(page->pgmap);
>> +
>> + /*
>> + * The page belong to device, do not return it to
>> + * page allocator.
>> + */
>> + return;
>> + }
>> +
>> if (unlikely(PageCompound(page)))
>> __put_compound_page(page);
>> else
>>
>
> Forgive me if I'm missing something but this doesn't make sense to me.
> We are taking a reference once when the region is initialized and
> releasing it every time a page within the region's reference count drops
> to zero. That does not seem to be symmetric and I don't see how it
> tracks that pages are in use. Shouldn't get_dev_pagemap be called when
> any page is allocated or something like that (ie. the inverse of
> __put_page)?
You're overlooking that the page reference count 1 after
arch_add_memory(). So at the end of time we're just dropping the
arch_add_memory() reference to release the page and related
dev_pagemap.
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2017-04-27 18:40 +0200 |
| Subject | Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tAUCu-3ai-13@gated-at.bofh.it> |
| In reply to | #1632229 |
On 27/04/17 10:14 AM, Dan Williams wrote: > You're overlooking that the page reference count 1 after > arch_add_memory(). So at the end of time we're just dropping the > arch_add_memory() reference to release the page and related > dev_pagemap. Thanks, that does actually make a lot more sense to me now. However, there still appears to be an asymmetry in that the pgmap->ref is incremented once and decremented once per page... Logan
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-27 18:40 +0200 |
| Subject | Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tAUCv-3ai-33@gated-at.bofh.it> |
| In reply to | #1632247 |
On Thu, Apr 27, 2017 at 9:33 AM, Logan Gunthorpe <logang@deltatee.com> wrote:
>
>
> On 27/04/17 10:14 AM, Dan Williams wrote:
>> You're overlooking that the page reference count 1 after
>> arch_add_memory(). So at the end of time we're just dropping the
>> arch_add_memory() reference to release the page and related
>> dev_pagemap.
>
> Thanks, that does actually make a lot more sense to me now. However,
> there still appears to be an asymmetry in that the pgmap->ref is
> incremented once and decremented once per page...
>
No, this hunk...
@@ -379,6 +374,7 @@ void *devm_memremap_pages(struct device *dev,
struct resource *res,
*/
list_del(&page->lru);
page->pgmap = pgmap;
+ percpu_ref_get(ref);
}
...is inside a for_each_device_pfn() loop.
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-27 18:50 +0200 |
| Subject | Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tAUMa-3dZ-5@gated-at.bofh.it> |
| In reply to | #1632254 |
On Thu, Apr 27, 2017 at 9:45 AM, Logan Gunthorpe <logang@deltatee.com> wrote: > > > On 27/04/17 10:38 AM, Dan Williams wrote: >> ...is inside a for_each_device_pfn() loop. >> > > Ah, oops. Then that makes perfect sense. Thanks. > > You may have my review tag if you'd like: > > Reviewed-by: Logan Gunthorpe <logang@deltatee.com> Thanks!
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2017-04-27 18:50 +0200 |
| Subject | Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tAUMa-3dZ-7@gated-at.bofh.it> |
| In reply to | #1632254 |
On 27/04/17 10:38 AM, Dan Williams wrote: > ...is inside a for_each_device_pfn() loop. > Ah, oops. Then that makes perfect sense. Thanks. You may have my review tag if you'd like: Reviewed-by: Logan Gunthorpe <logang@deltatee.com> Logan
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2017-04-27 18:20 +0200 |
| Subject | Re: [PATCH] mm, zone_device: replace {get, put}_zone_device_page() with a single reference |
| Message-ID | <tAUj7-338-13@gated-at.bofh.it> |
| In reply to | #1631814 |
On 26/04/17 06:55 PM, Dan Williams wrote:
> @@ -277,7 +269,10 @@ struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
> *
> * Notes:
> * 1/ @ref must be 'live' on entry and 'dead' before devm_memunmap_pages() time
> - * (or devm release event).
> + * (or devm release event). The expected order of events is that @ref has
> + * been through percpu_ref_kill() before devm_memremap_pages_release(). The
> + * wait for the completion of kill and percpu_ref_exit() must occur after
> + * devm_memremap_pages_release().
> *
> * 2/ @res is expected to be a host memory range that could feasibly be
> * treated as a "System RAM" range, i.e. not a device mmio range, but
> @@ -379,6 +374,7 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
> */
> list_del(&page->lru);
> page->pgmap = pgmap;
> + percpu_ref_get(ref);
> }
> devres_add(dev, page_map);
> return __va(res->start);
> diff --git a/mm/swap.c b/mm/swap.c
> index 5dabf444d724..01267dda6668 100644
> --- a/mm/swap.c
> +++ b/mm/swap.c
> @@ -97,6 +97,16 @@ static void __put_compound_page(struct page *page)
>
> void __put_page(struct page *page)
> {
> + if (is_zone_device_page(page)) {
> + put_dev_pagemap(page->pgmap);
> +
> + /*
> + * The page belong to device, do not return it to
> + * page allocator.
> + */
> + return;
> + }
> +
> if (unlikely(PageCompound(page)))
> __put_compound_page(page);
> else
>
Forgive me if I'm missing something but this doesn't make sense to me.
We are taking a reference once when the region is initialized and
releasing it every time a page within the region's reference count drops
to zero. That does not seem to be symmetric and I don't see how it
tracks that pages are in use. Shouldn't get_dev_pagemap be called when
any page is allocated or something like that (ie. the inverse of
__put_page)?
Thanks,
Logan
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web