Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1680575 > unrolled thread

[PATCH 0/5] Cache coherent device memory (CDM) with HMM v3

Started byJérôme Glisse <jglisse@redhat.com>
First post2017-07-03 23:20 +0200
Last post2017-07-03 23:20 +0200
Articles 10 on this page of 30 — 5 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH 0/5] Cache coherent device memory (CDM) with HMM v3 Jérôme Glisse <jglisse@redhat.com> - 2017-07-03 23:20 +0200
    [PATCH 5/5] mm/memcontrol: support MEMORY_DEVICE_PRIVATE and MEMORY_DEVICE_PUBLIC Jérôme Glisse <jglisse@redhat.com> - 2017-07-03 23:20 +0200
    [PATCH 2/5] mm/device-public-memory: device memory cache coherent with CPU v2 Jérôme Glisse <jglisse@redhat.com> - 2017-07-03 23:20 +0200
      Re: [PATCH 2/5] mm/device-public-memory: device memory cache  coherent with CPU v2 Balbir Singh <bsingharora@gmail.com> - 2017-07-11 06:20 +0200
        Re: [PATCH 2/5] mm/device-public-memory: device memory cache  coherent with CPU v2 Jerome Glisse <jglisse@redhat.com> - 2017-07-11 17:00 +0200
          Re: [PATCH 2/5] mm/device-public-memory: device memory cache  coherent with CPU v2 Balbir Singh <bsingharora@gmail.com> - 2017-07-12 08:00 +0200
    [PATCH 1/5] mm/persistent-memory: match IORES_DESC name and enum memory_type one Jérôme Glisse <jglisse@redhat.com> - 2017-07-03 23:20 +0200
      Re: [PATCH 1/5] mm/persistent-memory: match IORES_DESC name and enum  memory_type one Dan Williams <dan.j.williams@intel.com> - 2017-07-04 01:50 +0200
        Re: [PATCH 1/5] mm/persistent-memory: match IORES_DESC name and enum  memory_type one Jerome Glisse <jglisse@redhat.com> - 2017-07-05 16:30 +0200
          Re: [PATCH 1/5] mm/persistent-memory: match IORES_DESC name and enum  memory_type one Dan Williams <dan.j.williams@intel.com> - 2017-07-05 18:20 +0200
            Re: [PATCH 1/5] mm/persistent-memory: match IORES_DESC name and enum  memory_type one Jerome Glisse <jglisse@redhat.com> - 2017-07-05 20:50 +0200
              Re: [PATCH 1/5] mm/persistent-memory: match IORES_DESC name and  enum memory_type one Balbir Singh <bsingharora@gmail.com> - 2017-07-11 05:50 +0200
              Re: [PATCH 1/5] mm/persistent-memory: match IORES_DESC name and enum  memory_type one Dan Williams <dan.j.williams@intel.com> - 2017-07-11 09:40 +0200
                Re: [PATCH 1/5] mm/persistent-memory: match IORES_DESC name and enum  memory_type one Jerome Glisse <jglisse@redhat.com> - 2017-07-11 17:10 +0200
                  Re: [PATCH 1/5] mm/persistent-memory: match IORES_DESC name and enum  memory_type one Dan Williams <dan.j.williams@intel.com> - 2017-07-11 19:00 +0200
    [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field Jérôme Glisse <jglisse@redhat.com> - 2017-07-03 23:20 +0200
      Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Michal Hocko <mhocko@kernel.org> - 2017-07-04 15:00 +0200
        Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Balbir Singh <bsingharora@gmail.com> - 2017-07-05 05:20 +0200
          Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Michal Hocko <mhocko@kernel.org> - 2017-07-05 08:40 +0200
            Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Balbir Singh <bsingharora@gmail.com> - 2017-07-05 12:30 +0200
        Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Jerome Glisse <jglisse@redhat.com> - 2017-07-05 16:40 +0200
          Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Michal Hocko <mhocko@kernel.org> - 2017-07-10 10:30 +0200
            Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Jerome Glisse <jglisse@redhat.com> - 2017-07-10 17:40 +0200
              Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Michal Hocko <mhocko@kernel.org> - 2017-07-10 18:10 +0200
                Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Jerome Glisse <jglisse@redhat.com> - 2017-07-10 18:30 +0200
                  Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Michal Hocko <mhocko@kernel.org> - 2017-07-10 18:40 +0200
                    Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Jerome Glisse <jglisse@redhat.com> - 2017-07-10 19:00 +0200
                      Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Michal Hocko <mhocko@kernel.org> - 2017-07-10 19:50 +0200
                        Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using  page->lru field Jerome Glisse <jglisse@redhat.com> - 2017-07-10 20:20 +0200
    [PATCH 3/5] mm/hmm: add new helper to hotplug CDM memory region Jérôme Glisse <jglisse@redhat.com> - 2017-07-03 23:20 +0200

Page 2 of 2 — ← Prev page 1 [2]


#1681563 — Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field

FromJerome Glisse <jglisse@redhat.com>
Date2017-07-05 16:40 +0200
SubjectRe: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field
Message-ID<tZTDb-7v9-11@gated-at.bofh.it>
In reply to#1680932
On Tue, Jul 04, 2017 at 02:51:13PM +0200, Michal Hocko wrote:
> On Mon 03-07-17 17:14:14, Jérôme Glisse wrote:
> > HMM pages (private or public device pages) are ZONE_DEVICE page and
> > thus you can not use page->lru fields of those pages. This patch
> > re-arrange the uncharge to allow single page to be uncharge without
> > modifying the lru field of the struct page.
> > 
> > There is no change to memcontrol logic, it is the same as it was
> > before this patch.
> 
> What is the memcg semantic of the memory? Why is it even charged? AFAIR
> this is not a reclaimable memory. If yes how are we going to deal with
> memory limits? What should happen if go OOM? Does killing an process
> actually help to release that memory? Isn't it pinned by a device?
> 
> For the patch itself. It is quite ugly but I haven't spotted anything
> obviously wrong with it. It is the memcg semantic with this class of
> memory which makes me worried.

So i am facing 3 choices. First one not account device memory at all.
Second one is account device memory like any other memory inside a
process. Third one is account device memory as something entirely new.

I pick the second one for two reasons. First because when migrating
back from device memory it means that migration can not fail because
of memory cgroup limit, this simplify an already complex migration
code. Second because i assume that device memory usage is a transient
state ie once device is done with its computation the most likely
outcome is memory is migrated back. From this assumption it means
that you do not want to allow a process to overuse regular memory
while it is using un-accounted device memory. It sounds safer to
account device memory and to keep the process within its memcg
boundary.

Admittedly here i am making an assumption and i can be wrong. Thing
is we do not have enough real data of how this will be use and how
much of an impact device memory will have. That is why for now i
would rather restrict myself to either not account it or account it
as usual.

If you prefer not accounting it until we have more experience on how
it is use and how it impacts memory resource management i am fine with
that too. It will make the migration code slightly more complex.

Cheers,
Jérôme

[toc] | [prev] | [next] | [standalone]


#1684048 — Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field

FromMichal Hocko <mhocko@kernel.org>
Date2017-07-10 10:30 +0200
SubjectRe: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field
Message-ID<u1CeR-2AZ-9@gated-at.bofh.it>
In reply to#1681563
On Wed 05-07-17 10:35:29, Jerome Glisse wrote:
> On Tue, Jul 04, 2017 at 02:51:13PM +0200, Michal Hocko wrote:
> > On Mon 03-07-17 17:14:14, Jérôme Glisse wrote:
> > > HMM pages (private or public device pages) are ZONE_DEVICE page and
> > > thus you can not use page->lru fields of those pages. This patch
> > > re-arrange the uncharge to allow single page to be uncharge without
> > > modifying the lru field of the struct page.
> > > 
> > > There is no change to memcontrol logic, it is the same as it was
> > > before this patch.
> > 
> > What is the memcg semantic of the memory? Why is it even charged? AFAIR
> > this is not a reclaimable memory. If yes how are we going to deal with
> > memory limits? What should happen if go OOM? Does killing an process
> > actually help to release that memory? Isn't it pinned by a device?
> > 
> > For the patch itself. It is quite ugly but I haven't spotted anything
> > obviously wrong with it. It is the memcg semantic with this class of
> > memory which makes me worried.
> 
> So i am facing 3 choices. First one not account device memory at all.
> Second one is account device memory like any other memory inside a
> process. Third one is account device memory as something entirely new.
> 
> I pick the second one for two reasons. First because when migrating
> back from device memory it means that migration can not fail because
> of memory cgroup limit, this simplify an already complex migration
> code. Second because i assume that device memory usage is a transient
> state ie once device is done with its computation the most likely
> outcome is memory is migrated back. From this assumption it means
> that you do not want to allow a process to overuse regular memory
> while it is using un-accounted device memory. It sounds safer to
> account device memory and to keep the process within its memcg
> boundary.
> 
> Admittedly here i am making an assumption and i can be wrong. Thing
> is we do not have enough real data of how this will be use and how
> much of an impact device memory will have. That is why for now i
> would rather restrict myself to either not account it or account it
> as usual.
> 
> If you prefer not accounting it until we have more experience on how
> it is use and how it impacts memory resource management i am fine with
> that too. It will make the migration code slightly more complex.

I can see why you want to do this but the semantic _has_ to be clear.
And as such make sure that the exiting task will simply unpin and
invalidate all the device memory (assuming this memory is not shared
which I am not sure is even possible).
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1684360 — Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field

FromJerome Glisse <jglisse@redhat.com>
Date2017-07-10 17:40 +0200
SubjectRe: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field
Message-ID<u1IX0-6Op-19@gated-at.bofh.it>
In reply to#1684048
On Mon, Jul 10, 2017 at 10:28:06AM +0200, Michal Hocko wrote:
> On Wed 05-07-17 10:35:29, Jerome Glisse wrote:
> > On Tue, Jul 04, 2017 at 02:51:13PM +0200, Michal Hocko wrote:
> > > On Mon 03-07-17 17:14:14, Jérôme Glisse wrote:
> > > > HMM pages (private or public device pages) are ZONE_DEVICE page and
> > > > thus you can not use page->lru fields of those pages. This patch
> > > > re-arrange the uncharge to allow single page to be uncharge without
> > > > modifying the lru field of the struct page.
> > > > 
> > > > There is no change to memcontrol logic, it is the same as it was
> > > > before this patch.
> > > 
> > > What is the memcg semantic of the memory? Why is it even charged? AFAIR
> > > this is not a reclaimable memory. If yes how are we going to deal with
> > > memory limits? What should happen if go OOM? Does killing an process
> > > actually help to release that memory? Isn't it pinned by a device?
> > > 
> > > For the patch itself. It is quite ugly but I haven't spotted anything
> > > obviously wrong with it. It is the memcg semantic with this class of
> > > memory which makes me worried.
> > 
> > So i am facing 3 choices. First one not account device memory at all.
> > Second one is account device memory like any other memory inside a
> > process. Third one is account device memory as something entirely new.
> > 
> > I pick the second one for two reasons. First because when migrating
> > back from device memory it means that migration can not fail because
> > of memory cgroup limit, this simplify an already complex migration
> > code. Second because i assume that device memory usage is a transient
> > state ie once device is done with its computation the most likely
> > outcome is memory is migrated back. From this assumption it means
> > that you do not want to allow a process to overuse regular memory
> > while it is using un-accounted device memory. It sounds safer to
> > account device memory and to keep the process within its memcg
> > boundary.
> > 
> > Admittedly here i am making an assumption and i can be wrong. Thing
> > is we do not have enough real data of how this will be use and how
> > much of an impact device memory will have. That is why for now i
> > would rather restrict myself to either not account it or account it
> > as usual.
> > 
> > If you prefer not accounting it until we have more experience on how
> > it is use and how it impacts memory resource management i am fine with
> > that too. It will make the migration code slightly more complex.
> 
> I can see why you want to do this but the semantic _has_ to be clear.
> And as such make sure that the exiting task will simply unpin and
> invalidate all the device memory (assuming this memory is not shared
> which I am not sure is even possible).

So there is 2 differents path out of device memory:
  - munmap/process exiting: memory will get uncharge from its memory
    cgroup just like regular memory
  - migration to non device memory, the memory cgroup charge get
    transfer to the new page just like for any other page

Do you want me to document all this in any specific place ? I will
add a comment in memory_control.c and in HMM documentations for this
but should i add it anywhere else ?

Note that the device memory is not pin. The whole point of HMM is to
do away with any pining. Thought as device page are not on lru they
are not reclaim like any other page. However we expect that device
driver might implement something akin to device memory reclaim to
make room for more important data base on statistic collected by the
device driver. If there is enough commonality accross devices then
we might implement a more generic mechanisms but at this point i
rather grow as we learn.

Cheers,
Jérôme

[toc] | [prev] | [next] | [standalone]


#1684391 — Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field

FromMichal Hocko <mhocko@kernel.org>
Date2017-07-10 18:10 +0200
SubjectRe: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field
Message-ID<u1Jq1-7dI-13@gated-at.bofh.it>
In reply to#1684360
On Mon 10-07-17 11:32:23, Jerome Glisse wrote:
> On Mon, Jul 10, 2017 at 10:28:06AM +0200, Michal Hocko wrote:
> > On Wed 05-07-17 10:35:29, Jerome Glisse wrote:
> > > On Tue, Jul 04, 2017 at 02:51:13PM +0200, Michal Hocko wrote:
> > > > On Mon 03-07-17 17:14:14, Jérôme Glisse wrote:
> > > > > HMM pages (private or public device pages) are ZONE_DEVICE page and
> > > > > thus you can not use page->lru fields of those pages. This patch
> > > > > re-arrange the uncharge to allow single page to be uncharge without
> > > > > modifying the lru field of the struct page.
> > > > > 
> > > > > There is no change to memcontrol logic, it is the same as it was
> > > > > before this patch.
> > > > 
> > > > What is the memcg semantic of the memory? Why is it even charged? AFAIR
> > > > this is not a reclaimable memory. If yes how are we going to deal with
> > > > memory limits? What should happen if go OOM? Does killing an process
> > > > actually help to release that memory? Isn't it pinned by a device?
> > > > 
> > > > For the patch itself. It is quite ugly but I haven't spotted anything
> > > > obviously wrong with it. It is the memcg semantic with this class of
> > > > memory which makes me worried.
> > > 
> > > So i am facing 3 choices. First one not account device memory at all.
> > > Second one is account device memory like any other memory inside a
> > > process. Third one is account device memory as something entirely new.
> > > 
> > > I pick the second one for two reasons. First because when migrating
> > > back from device memory it means that migration can not fail because
> > > of memory cgroup limit, this simplify an already complex migration
> > > code. Second because i assume that device memory usage is a transient
> > > state ie once device is done with its computation the most likely
> > > outcome is memory is migrated back. From this assumption it means
> > > that you do not want to allow a process to overuse regular memory
> > > while it is using un-accounted device memory. It sounds safer to
> > > account device memory and to keep the process within its memcg
> > > boundary.
> > > 
> > > Admittedly here i am making an assumption and i can be wrong. Thing
> > > is we do not have enough real data of how this will be use and how
> > > much of an impact device memory will have. That is why for now i
> > > would rather restrict myself to either not account it or account it
> > > as usual.
> > > 
> > > If you prefer not accounting it until we have more experience on how
> > > it is use and how it impacts memory resource management i am fine with
> > > that too. It will make the migration code slightly more complex.
> > 
> > I can see why you want to do this but the semantic _has_ to be clear.
> > And as such make sure that the exiting task will simply unpin and
> > invalidate all the device memory (assuming this memory is not shared
> > which I am not sure is even possible).
> 
> So there is 2 differents path out of device memory:
>   - munmap/process exiting: memory will get uncharge from its memory
>     cgroup just like regular memory

I might have missed that in your patch, I admit I only glanced through
that, but the memcg uncharged when the last reference to the page is
released. So if the device pins the page for some reason then the charge
will be there even when the oom victim unmaps the memory.

>   - migration to non device memory, the memory cgroup charge get
>     transfer to the new page just like for any other page
> 
> Do you want me to document all this in any specific place ? I will
> add a comment in memory_control.c and in HMM documentations for this
> but should i add it anywhere else ?

hmm documentation is sufficient and the uncharge path if it needs any
special handling.

> Note that the device memory is not pin. The whole point of HMM is to
> do away with any pining. Thought as device page are not on lru they
> are not reclaim like any other page. However we expect that device
> driver might implement something akin to device memory reclaim to
> make room for more important data base on statistic collected by the
> device driver. If there is enough commonality accross devices then
> we might implement a more generic mechanisms but at this point i
> rather grow as we learn.

Do we have any guarantee that devices will _never_ pin those pages? If
no then we have to make sure we can forcefully tear them down.

-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1684406 — Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field

FromJerome Glisse <jglisse@redhat.com>
Date2017-07-10 18:30 +0200
SubjectRe: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field
Message-ID<u1JJo-7kl-29@gated-at.bofh.it>
In reply to#1684391
On Mon, Jul 10, 2017 at 06:04:46PM +0200, Michal Hocko wrote:
> On Mon 10-07-17 11:32:23, Jerome Glisse wrote:
> > On Mon, Jul 10, 2017 at 10:28:06AM +0200, Michal Hocko wrote:
> > > On Wed 05-07-17 10:35:29, Jerome Glisse wrote:
> > > > On Tue, Jul 04, 2017 at 02:51:13PM +0200, Michal Hocko wrote:
> > > > > On Mon 03-07-17 17:14:14, Jérôme Glisse wrote:
> > > > > > HMM pages (private or public device pages) are ZONE_DEVICE page and
> > > > > > thus you can not use page->lru fields of those pages. This patch
> > > > > > re-arrange the uncharge to allow single page to be uncharge without
> > > > > > modifying the lru field of the struct page.
> > > > > > 
> > > > > > There is no change to memcontrol logic, it is the same as it was
> > > > > > before this patch.
> > > > > 
> > > > > What is the memcg semantic of the memory? Why is it even charged? AFAIR
> > > > > this is not a reclaimable memory. If yes how are we going to deal with
> > > > > memory limits? What should happen if go OOM? Does killing an process
> > > > > actually help to release that memory? Isn't it pinned by a device?
> > > > > 
> > > > > For the patch itself. It is quite ugly but I haven't spotted anything
> > > > > obviously wrong with it. It is the memcg semantic with this class of
> > > > > memory which makes me worried.
> > > > 
> > > > So i am facing 3 choices. First one not account device memory at all.
> > > > Second one is account device memory like any other memory inside a
> > > > process. Third one is account device memory as something entirely new.
> > > > 
> > > > I pick the second one for two reasons. First because when migrating
> > > > back from device memory it means that migration can not fail because
> > > > of memory cgroup limit, this simplify an already complex migration
> > > > code. Second because i assume that device memory usage is a transient
> > > > state ie once device is done with its computation the most likely
> > > > outcome is memory is migrated back. From this assumption it means
> > > > that you do not want to allow a process to overuse regular memory
> > > > while it is using un-accounted device memory. It sounds safer to
> > > > account device memory and to keep the process within its memcg
> > > > boundary.
> > > > 
> > > > Admittedly here i am making an assumption and i can be wrong. Thing
> > > > is we do not have enough real data of how this will be use and how
> > > > much of an impact device memory will have. That is why for now i
> > > > would rather restrict myself to either not account it or account it
> > > > as usual.
> > > > 
> > > > If you prefer not accounting it until we have more experience on how
> > > > it is use and how it impacts memory resource management i am fine with
> > > > that too. It will make the migration code slightly more complex.
> > > 
> > > I can see why you want to do this but the semantic _has_ to be clear.
> > > And as such make sure that the exiting task will simply unpin and
> > > invalidate all the device memory (assuming this memory is not shared
> > > which I am not sure is even possible).
> > 
> > So there is 2 differents path out of device memory:
> >   - munmap/process exiting: memory will get uncharge from its memory
> >     cgroup just like regular memory
> 
> I might have missed that in your patch, I admit I only glanced through
> that, but the memcg uncharged when the last reference to the page is
> released. So if the device pins the page for some reason then the charge
> will be there even when the oom victim unmaps the memory.

Device can not pin memory it is part of the "contract" when using HMM.
Device memory can never be pin. Nor by device driver nor by any other
means ie we want GUP to trigger a migration back to regular memory. We
will relax the GUP requirement a one point (especialy for direct I/O
and other short time GUP).


> >   - migration to non device memory, the memory cgroup charge get
> >     transfer to the new page just like for any other page
> > 
> > Do you want me to document all this in any specific place ? I will
> > add a comment in memory_control.c and in HMM documentations for this
> > but should i add it anywhere else ?
> 
> hmm documentation is sufficient and the uncharge path if it needs any
> special handling.

Uncharge happens in the ZONE_DEVICE special handling of page refcount
ie a ZONE_DEVICE is free when its refcount reach 1 not 0.

> 
> > Note that the device memory is not pin. The whole point of HMM is to
> > do away with any pining. Thought as device page are not on lru they
> > are not reclaim like any other page. However we expect that device
> > driver might implement something akin to device memory reclaim to
> > make room for more important data base on statistic collected by the
> > device driver. If there is enough commonality accross devices then
> > we might implement a more generic mechanisms but at this point i
> > rather grow as we learn.
> 
> Do we have any guarantee that devices will _never_ pin those pages? If
> no then we have to make sure we can forcefully tear them down.

Well yes we do, as long as i monitor how driver use thing :) Device we
are targetting are like CPU from MMU point of view ie you can tear down
a device page table entry without having the device to freak about it.
So there is no need for device to pin anything, if we update its page
table to non present entry any further access to the virtual address
will trigger a fault that is then handled by the device driver.

If the process is being kill than the GPU threads can be kill by the
device driver too. Otherwise the page fault is handled with the help
of HMM like any reguler CPU page fault. If for some reasons we can not
service the fault than the device driver is responsible to decide how
to handle various VM_FAULT_ERROR. Expectation is that it kills the
device threads and inform userspace through device specific API. I
think at one point down the road we will want to standardize way to
communicate fatal error condition that affect device threads.


I will review HMM documentation again to make sure this is all in
black and white. I am pretty sure that some of it is already there.

Bottom line is that we can always free and uncharge device memory
page just like any regular page.

Cheers,
Jérôme

[toc] | [prev] | [next] | [standalone]


#1684408 — Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field

FromMichal Hocko <mhocko@kernel.org>
Date2017-07-10 18:40 +0200
SubjectRe: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field
Message-ID<u1JT4-7ny-11@gated-at.bofh.it>
In reply to#1684406
On Mon 10-07-17 12:25:42, Jerome Glisse wrote:
[...]
> Bottom line is that we can always free and uncharge device memory
> page just like any regular page.

OK, this answers my earlier question. Then it should be feasible to
charge this memory. There are still some things to handle. E.g. how do
we consider this memory during oom victim selection (this is not
accounted as an anonymous memory in get_mm_counter, right?), maybe others.
But the primary point is that nobody pins the memory outside of the
mapping.
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1684418 — Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field

FromJerome Glisse <jglisse@redhat.com>
Date2017-07-10 19:00 +0200
SubjectRe: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field
Message-ID<u1Kcq-7u8-11@gated-at.bofh.it>
In reply to#1684408
On Mon, Jul 10, 2017 at 06:36:52PM +0200, Michal Hocko wrote:
> On Mon 10-07-17 12:25:42, Jerome Glisse wrote:
> [...]
> > Bottom line is that we can always free and uncharge device memory
> > page just like any regular page.
> 
> OK, this answers my earlier question. Then it should be feasible to
> charge this memory. There are still some things to handle. E.g. how do
> we consider this memory during oom victim selection (this is not
> accounted as an anonymous memory in get_mm_counter, right?), maybe others.
> But the primary point is that nobody pins the memory outside of the
> mapping.

At this point it is accounted as a regular page would be (anonymous, file
or share memory). I wanted mm_counters to reflect memcg but i can untie
that. Like i said at this point we are unsure how usage of such memory
will impact thing so i wanted to keep all thing as if it was regular
memory to avoid anuything to behave too much differently.

Jérôme

[toc] | [prev] | [next] | [standalone]


#1684539 — Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field

FromMichal Hocko <mhocko@kernel.org>
Date2017-07-10 19:50 +0200
SubjectRe: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field
Message-ID<u1KYN-81W-1@gated-at.bofh.it>
In reply to#1684418
On Mon 10-07-17 12:54:21, Jerome Glisse wrote:
> On Mon, Jul 10, 2017 at 06:36:52PM +0200, Michal Hocko wrote:
> > On Mon 10-07-17 12:25:42, Jerome Glisse wrote:
> > [...]
> > > Bottom line is that we can always free and uncharge device memory
> > > page just like any regular page.
> > 
> > OK, this answers my earlier question. Then it should be feasible to
> > charge this memory. There are still some things to handle. E.g. how do
> > we consider this memory during oom victim selection (this is not
> > accounted as an anonymous memory in get_mm_counter, right?), maybe others.
> > But the primary point is that nobody pins the memory outside of the
> > mapping.
> 
> At this point it is accounted as a regular page would be (anonymous, file
> or share memory). I wanted mm_counters to reflect memcg but i can untie
> that.

I am not sure I understand. If the device memory is accounted to the
same mm counter as the original page then it is correct. I will try to
double check the implementation (hopefully soon).

-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1684558 — Re: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field

FromJerome Glisse <jglisse@redhat.com>
Date2017-07-10 20:20 +0200
SubjectRe: [PATCH 4/5] mm/memcontrol: allow to uncharge page without using page->lru field
Message-ID<u1LrQ-8ub-23@gated-at.bofh.it>
In reply to#1684539
On Mon, Jul 10, 2017 at 07:48:58PM +0200, Michal Hocko wrote:
> On Mon 10-07-17 12:54:21, Jerome Glisse wrote:
> > On Mon, Jul 10, 2017 at 06:36:52PM +0200, Michal Hocko wrote:
> > > On Mon 10-07-17 12:25:42, Jerome Glisse wrote:
> > > [...]
> > > > Bottom line is that we can always free and uncharge device memory
> > > > page just like any regular page.
> > > 
> > > OK, this answers my earlier question. Then it should be feasible to
> > > charge this memory. There are still some things to handle. E.g. how do
> > > we consider this memory during oom victim selection (this is not
> > > accounted as an anonymous memory in get_mm_counter, right?), maybe others.
> > > But the primary point is that nobody pins the memory outside of the
> > > mapping.
> > 
> > At this point it is accounted as a regular page would be (anonymous, file
> > or share memory). I wanted mm_counters to reflect memcg but i can untie
> > that.
> 
> I am not sure I understand. If the device memory is accounted to the
> same mm counter as the original page then it is correct. I will try to
> double check the implementation (hopefully soon).

It is accounted like the original page. By same as memcg i mean i made
the same kind of choice for mm counter than i made for memcg. It is
all in the migrate code (migrate.c) ie i don't touch any of the mm
counter when migrating page.

Jérôme

[toc] | [prev] | [next] | [standalone]


#1680594 — [PATCH 3/5] mm/hmm: add new helper to hotplug CDM memory region

FromJérôme Glisse <jglisse@redhat.com>
Date2017-07-03 23:20 +0200
Subject[PATCH 3/5] mm/hmm: add new helper to hotplug CDM memory region
Message-ID<tZgVd-7a5-49@gated-at.bofh.it>
In reply to#1680575
Unlike unaddressable memory, coherent device memory has a real
resource associated with it on the system (as CPU can address
it). Add a new helper to hotplug such memory within the HMM
framework.

Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Reviewed-by: Balbir Singh <bsingharora@gmail.com>
---
 include/linux/hmm.h |  3 ++
 mm/hmm.c            | 85 +++++++++++++++++++++++++++++++++++++++++++++++++----
 2 files changed, 83 insertions(+), 5 deletions(-)

diff --git a/include/linux/hmm.h b/include/linux/hmm.h
index a40288309fd2..e44cb8edb137 100644
--- a/include/linux/hmm.h
+++ b/include/linux/hmm.h
@@ -392,6 +392,9 @@ struct hmm_devmem {
 struct hmm_devmem *hmm_devmem_add(const struct hmm_devmem_ops *ops,
 				  struct device *device,
 				  unsigned long size);
+struct hmm_devmem *hmm_devmem_add_resource(const struct hmm_devmem_ops *ops,
+					   struct device *device,
+					   struct resource *res);
 void hmm_devmem_remove(struct hmm_devmem *devmem);
 
 /*
diff --git a/mm/hmm.c b/mm/hmm.c
index eadf70829c34..28e54e3b4e1d 100644
--- a/mm/hmm.c
+++ b/mm/hmm.c
@@ -849,7 +849,11 @@ static void hmm_devmem_release(struct device *dev, void *data)
 	zone = page_zone(page);
 
 	mem_hotplug_begin();
-	__remove_pages(zone, start_pfn, npages);
+	if (resource->desc == IORES_DESC_DEVICE_PRIVATE_MEMORY)
+		__remove_pages(zone, start_pfn, npages);
+	else
+		arch_remove_memory(start_pfn << PAGE_SHIFT,
+				   npages << PAGE_SHIFT);
 	mem_hotplug_done();
 
 	hmm_devmem_radix_release(resource);
@@ -885,7 +889,11 @@ static int hmm_devmem_pages_create(struct hmm_devmem *devmem)
 	if (is_ram == REGION_INTERSECTS)
 		return -ENXIO;
 
-	devmem->pagemap.type = MEMORY_DEVICE_PRIVATE;
+	if (devmem->resource->desc == IORES_DESC_DEVICE_PUBLIC_MEMORY)
+		devmem->pagemap.type = MEMORY_DEVICE_PUBLIC;
+	else
+		devmem->pagemap.type = MEMORY_DEVICE_PRIVATE;
+
 	devmem->pagemap.res = devmem->resource;
 	devmem->pagemap.page_fault = hmm_devmem_fault;
 	devmem->pagemap.page_free = hmm_devmem_free;
@@ -924,8 +932,11 @@ static int hmm_devmem_pages_create(struct hmm_devmem *devmem)
 		nid = numa_mem_id();
 
 	mem_hotplug_begin();
-	ret = add_pages(nid, align_start >> PAGE_SHIFT,
-			align_size >> PAGE_SHIFT, false);
+	if (devmem->pagemap.type == MEMORY_DEVICE_PUBLIC)
+		ret = arch_add_memory(nid, align_start, align_size, false);
+	else
+		ret = add_pages(nid, align_start >> PAGE_SHIFT,
+				align_size >> PAGE_SHIFT, false);
 	if (ret) {
 		mem_hotplug_done();
 		goto error_add_memory;
@@ -1075,6 +1086,67 @@ struct hmm_devmem *hmm_devmem_add(const struct hmm_devmem_ops *ops,
 }
 EXPORT_SYMBOL(hmm_devmem_add);
 
+struct hmm_devmem *hmm_devmem_add_resource(const struct hmm_devmem_ops *ops,
+					   struct device *device,
+					   struct resource *res)
+{
+	struct hmm_devmem *devmem;
+	int ret;
+
+	if (res->desc != IORES_DESC_DEVICE_PUBLIC_MEMORY)
+		return ERR_PTR(-EINVAL);
+
+	static_branch_enable(&device_private_key);
+
+	devmem = devres_alloc_node(&hmm_devmem_release, sizeof(*devmem),
+				   GFP_KERNEL, dev_to_node(device));
+	if (!devmem)
+		return ERR_PTR(-ENOMEM);
+
+	init_completion(&devmem->completion);
+	devmem->pfn_first = -1UL;
+	devmem->pfn_last = -1UL;
+	devmem->resource = res;
+	devmem->device = device;
+	devmem->ops = ops;
+
+	ret = percpu_ref_init(&devmem->ref, &hmm_devmem_ref_release,
+			      0, GFP_KERNEL);
+	if (ret)
+		goto error_percpu_ref;
+
+	ret = devm_add_action(device, hmm_devmem_ref_exit, &devmem->ref);
+	if (ret)
+		goto error_devm_add_action;
+
+
+	devmem->pfn_first = devmem->resource->start >> PAGE_SHIFT;
+	devmem->pfn_last = devmem->pfn_first +
+			   (resource_size(devmem->resource) >> PAGE_SHIFT);
+
+	ret = hmm_devmem_pages_create(devmem);
+	if (ret)
+		goto error_devm_add_action;
+
+	devres_add(device, devmem);
+
+	ret = devm_add_action(device, hmm_devmem_ref_kill, &devmem->ref);
+	if (ret) {
+		hmm_devmem_remove(devmem);
+		return ERR_PTR(ret);
+	}
+
+	return devmem;
+
+error_devm_add_action:
+	hmm_devmem_ref_kill(&devmem->ref);
+	hmm_devmem_ref_exit(&devmem->ref);
+error_percpu_ref:
+	devres_free(devmem);
+	return ERR_PTR(ret);
+}
+EXPORT_SYMBOL(hmm_devmem_add_resource);
+
 /*
  * hmm_devmem_remove() - remove device memory (kill and free ZONE_DEVICE)
  *
@@ -1088,6 +1160,7 @@ void hmm_devmem_remove(struct hmm_devmem *devmem)
 {
 	resource_size_t start, size;
 	struct device *device;
+	bool cdm = false;
 
 	if (!devmem)
 		return;
@@ -1096,11 +1169,13 @@ void hmm_devmem_remove(struct hmm_devmem *devmem)
 	start = devmem->resource->start;
 	size = resource_size(devmem->resource);
 
+	cdm = devmem->resource->desc == IORES_DESC_DEVICE_PUBLIC_MEMORY;
 	hmm_devmem_ref_kill(&devmem->ref);
 	hmm_devmem_ref_exit(&devmem->ref);
 	hmm_devmem_pages_remove(devmem);
 
-	devm_release_mem_region(device, start, size);
+	if (!cdm)
+		devm_release_mem_region(device, start, size);
 }
 EXPORT_SYMBOL(hmm_devmem_remove);
 
-- 
2.13.0

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web