Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1530503 > unrolled thread
| Started by | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| First post | 2016-11-25 20:40 +0100 |
| Last post | 2016-11-28 19:40 +0100 |
| Articles | 20 on this page of 33 — 8 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-25 20:40 +0100
Re: Enabling peer to peer device transactions for PCIe devices Christian König <deathsimple@vodafone.de> - 2016-11-25 21:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-25 22:20 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-28 18:00 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-11-28 19:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-11-28 22:40 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-28 23:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-28 20:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-30 17:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-11-30 19:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices "Stephen Bates" <sbates@raithlin.com> - 2016-12-04 14:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices "Stephen Bates" <sbates@raithlin.com> - 2016-12-04 14:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-05 18:20 +0100
Re: Enabling peer to peer device transactions for PCIe devices Dan Williams <dan.j.williams@intel.com> - 2016-12-05 18:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Dan Williams <dan.j.williams@intel.com> - 2016-12-05 19:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-12-05 19:40 +0100
Re: Enabling peer to peer device transactions for PCIe devices Dan Williams <dan.j.williams@intel.com> - 2016-12-05 19:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-05 20:20 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-12-05 20:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-05 20:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-12-05 21:00 +0100
Re: Enabling peer to peer device transactions for PCIe devices Christoph Hellwig <hch@infradead.org> - 2016-12-05 21:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices "Stephen Bates" <sbates@raithlin.com> - 2016-12-06 09:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-06 17:40 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-12-06 18:00 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-06 18:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-12-06 22:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Dan Williams <dan.j.williams@intel.com> - 2016-12-06 23:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices Christoph Hellwig <hch@infradead.org> - 2016-12-06 18:20 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-05 19:10 +0100
RE: Enabling peer to peer device transactions for PCIe devices "Deucher, Alexander" <Alexander.Deucher@amd.com> - 2016-11-30 18:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Haggai Eran <haggaie@mellanox.com> - 2016-11-28 21:00 +0100
Re: Enabling peer to peer device transactions for PCIe devices Haggai Eran <haggaie@mellanox.com> - 2016-11-28 19:40 +0100
Page 1 of 2 [1] 2 Next page →
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-11-25 20:40 +0100 |
| Subject | Re: Enabling peer to peer device transactions for PCIe devices |
| Message-ID | <sHuvM-4HW-27@gated-at.bofh.it> |
On Fri, Nov 25, 2016 at 02:22:17PM +0100, Christian König wrote: > >Like you say below we have to handle short lived in the usual way, and > >that covers basically every device except IB MRs, including the > >command queue on a NVMe drive. > > Well a problem which wasn't mentioned so far is that while GPUs do have a > page table to mirror the CPU page table, they usually can't recover from > page faults. > So what we do is making sure that all memory accessed by the GPU Jobs stays > in place while those jobs run (pretty much the same pinning you do for the > DMA). Yes, it is DMA, so this is a valid approach. But, you don't need page faults from the GPU to do proper coherent page table mirroring. Basically when the driver submits the work to the GPU it 'faults' the pages into the CPU and mirror translation table (instead of pinning). Like in ODP, MMU notifiers/HMM are used to monitor for translation changes. If a change comes in the GPU driver checks if an executing command is touching those pages and blocks the MMU notifier until the command flushes, then unfaults the page (blocking future commands) and unblocks the mmu notifier. The code moving the page will move it and the next GPU command that needs it will refault it in the usual way, just like the CPU would. This might be much more efficient since it optimizes for the common case of unchanging translation tables. This assumes the commands are fairly short lived of course, the expectation of the mmu notifiers is that a flush is reasonably prompt .. > >Serguei, what is your plan in GPU land for migration? Ie if I have a > >CPU mapped page and the GPU moves it to VRAM, it becomes non-cachable > >- do you still allow the CPU to access it? Or do you swap it back to > >cachable memory if the CPU touches it? > > Depends on the policy in command, but currently it's the other way around > most of the time. > > E.g. we allocate memory in VRAM, the CPU writes to it WC and avoids reading > because that is slow, the GPU in turn can access it with full speed. > > When we run out of VRAM we move those allocations to system memory and > update both the CPU as well as the GPU page tables. > > So that move is transparent for both userspace as well as shaders running on > the GPU. That makes sense to me, but the objection that came back for non-cachable CPU mappings is that it basically breaks too much stuff subtly, eg atomics, unaligned accesses, the CPU threading memory model, all change on various architectures and break when caching is disabled. IMHO that is OK for specialty things like the GPU where the mmap comes in via drm or something and apps know to handle that buffer specially. But it is certainly not OK for DAX where the application is coded for normal file open()/mmap() is not prepared for a mmap where (eg) unaligned read accesses or atomics don't work depending on how the filesystem is setup. Which is why I think iopmem is still problematic.. At the very least I think a mmap flag or open flag should be needed to opt into this behavior and by default non-cachebale DAX mmaps should be paged into system ram when the CPU accesses them. I'm hearing most people say ZONE_DEVICE is the way to handle this, which means the missing remaing piece for RDMA is some kind of DMA core support for p2p address translation.. Jason
[toc] | [next] | [standalone]
| From | Christian König <deathsimple@vodafone.de> |
|---|---|
| Date | 2016-11-25 21:50 +0100 |
| Message-ID | <sHvBv-5oB-1@gated-at.bofh.it> |
| In reply to | #1530503 |
Am 25.11.2016 um 20:32 schrieb Jason Gunthorpe: > On Fri, Nov 25, 2016 at 02:22:17PM +0100, Christian König wrote: > >>> Like you say below we have to handle short lived in the usual way, and >>> that covers basically every device except IB MRs, including the >>> command queue on a NVMe drive. >> Well a problem which wasn't mentioned so far is that while GPUs do have a >> page table to mirror the CPU page table, they usually can't recover from >> page faults. >> So what we do is making sure that all memory accessed by the GPU Jobs stays >> in place while those jobs run (pretty much the same pinning you do for the >> DMA). > Yes, it is DMA, so this is a valid approach. > > But, you don't need page faults from the GPU to do proper coherent > page table mirroring. Basically when the driver submits the work to > the GPU it 'faults' the pages into the CPU and mirror translation > table (instead of pinning). > > Like in ODP, MMU notifiers/HMM are used to monitor for translation > changes. If a change comes in the GPU driver checks if an executing > command is touching those pages and blocks the MMU notifier until the > command flushes, then unfaults the page (blocking future commands) and > unblocks the mmu notifier. Yeah, we have a function to "import" anonymous pages from a CPU pointer which works exactly that way as well. We call this "userptr" and it's just a combination of get_user_pages() on command submission and making sure the returned list of pages stays valid using a MMU notifier. The "big" problem with this approach is that it is horrible slow. I mean seriously horrible slow so that we actually can't use it for some of the purposes we wanted to use it. > The code moving the page will move it and the next GPU command that > needs it will refault it in the usual way, just like the CPU would. And here comes the problem. CPU do this on a page by page basis, so they fault only what needed and everything else gets filled in on demand. This results that faulting a page is relatively light weight operation. But for GPU command submission we don't know which pages might be accessed beforehand, so what we do is walking all possible pages and make sure all of them are present. Now as far as I understand it the I/O subsystem for example assumes that it can easily change the CPU page tables without much overhead. So for example when a page can't modified it is temporary marked as readonly AFAIK (you are probably way deeper into this than me, so please confirm). That absolutely kills any performance for GPU command submissions. We have use cases where we practically ended up playing ping/pong between the GPU driver trying to grab the page with get_user_pages() and sombody else in the kernel marking it readonly. > This might be much more efficient since it optimizes for the common > case of unchanging translation tables. Yeah, completely agree. It works perfectly fine as long as you don't have two drivers trying to mess with the same page. > This assumes the commands are fairly short lived of course, the > expectation of the mmu notifiers is that a flush is reasonably prompt Correct, this is another problem. GFX command submissions usually don't take longer than a few milliseconds, but compute command submission can easily take multiple hours. I can easily imagine what would happen when kswapd is blocked by a GPU command submission for an hour or so while the system is under memory pressure :) I'm thinking on this problem for about a year now and going in circles for quite a while. So if you have ideas on this even if they sound totally crazy, feel free to come up. Cheers, Christian.
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-11-25 22:20 +0100 |
| Message-ID | <sHw4y-5Op-19@gated-at.bofh.it> |
| In reply to | #1530522 |
On Fri, Nov 25, 2016 at 09:40:10PM +0100, Christian König wrote: > We call this "userptr" and it's just a combination of get_user_pages() on > command submission and making sure the returned list of pages stays valid > using a MMU notifier. Doesn't that still pin the page? > The "big" problem with this approach is that it is horrible slow. I mean > seriously horrible slow so that we actually can't use it for some of the > purposes we wanted to use it. > > >The code moving the page will move it and the next GPU command that > >needs it will refault it in the usual way, just like the CPU would. > > And here comes the problem. CPU do this on a page by page basis, so they > fault only what needed and everything else gets filled in on demand. This > results that faulting a page is relatively light weight operation. > > But for GPU command submission we don't know which pages might be accessed > beforehand, so what we do is walking all possible pages and make sure all of > them are present. Little confused why this is slow? So you fault the entire user MM into your page tables at start of day and keep track of it with mmu notifiers? > >This might be much more efficient since it optimizes for the common > >case of unchanging translation tables. > > Yeah, completely agree. It works perfectly fine as long as you don't have > two drivers trying to mess with the same page. Well, the idea would be to not have the GPU block the other driver beyond hinting that the page shouldn't be swapped out. > >This assumes the commands are fairly short lived of course, the > >expectation of the mmu notifiers is that a flush is reasonably prompt > > Correct, this is another problem. GFX command submissions usually don't take > longer than a few milliseconds, but compute command submission can easily > take multiple hours. So, that won't work - you have the same issue as RDMA with work loads like that. If you can't somehow fence the hardware then pinning is the only solution. Felix has the right kind of suggestion for what is needed - globally stop the GPU, fence the DMA, fix the page tables, and start it up again. :\ > I can easily imagine what would happen when kswapd is blocked by a GPU > command submission for an hour or so while the system is under memory > pressure :) Right. The advantage of pinning is it tells the other stuff not to touch the page and doesn't block it, MMU notifiers have to be able to block&fence quickly. > I'm thinking on this problem for about a year now and going in circles for > quite a while. So if you have ideas on this even if they sound totally > crazy, feel free to come up. Well, it isn't a software problem. From what I've seen in this thread the GPU application requires coherent page table mirroring, so the only full & complete solution is going to be to actually implement that somehow in GPU hardware. Everything else is going to be deeply flawed somehow. Linux just doesn't have the support for this kind of stuff - and I'm honestly not sure something better is even possible considering the hardware constraints.... This doesn't have to be faulting, but really anything that lets you pause the GPU DMA and reload the page tables. You might look at trying to use the IOMMU and/or PCI ATS in very new hardware. IIRC the physical IOMMU hardware can do the fault and fence and block stuff, but I'm not sure about software support for using the IOMMU to create coherent user page table mirrors - that is something Linux doesn't do today. But there is demand for this kind of capability.. Jason
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-11-28 18:00 +0100 |
| Message-ID | <sIxrA-4y6-23@gated-at.bofh.it> |
| In reply to | #1530503 |
On Sun, Nov 27, 2016 at 04:02:16PM +0200, Haggai Eran wrote: > > Like in ODP, MMU notifiers/HMM are used to monitor for translation > > changes. If a change comes in the GPU driver checks if an executing > > command is touching those pages and blocks the MMU notifier until the > > command flushes, then unfaults the page (blocking future commands) and > > unblocks the mmu notifier. > I think blocking mmu notifiers against something that is basically > controlled by user-space can be problematic. This can block things like > memory reclaim. If you have user-space access to the device's queues, > user-space can block the mmu notifier forever. Right, I mentioned that.. > On PeerDirect, we have some kind of a middle-ground solution for pinning > GPU memory. We create a non-ODP MR pointing to VRAM but rely on > user-space and the GPU not to migrate it. If they do, the MR gets > destroyed immediately. That sounds horrible. How can that possibly work? What if the MR is being used when the GPU decides to migrate? I would not support that upstream without a lot more explanation.. I know people don't like requiring new hardware, but in this case we really do need ODP hardware to get all the semantics people want.. > Another thing I think is that while HMM is good for user-space > applications, for kernel p2p use there is no need for that. Using From what I understand we are not really talking about kernel p2p, everything proposed so far is being mediated by a userspace VMA, so I'd focus on making that work. Jason
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2016-11-28 19:30 +0100 |
| Message-ID | <sIyQG-5J3-37@gated-at.bofh.it> |
| In reply to | #1531445 |
On 28/11/16 09:57 AM, Jason Gunthorpe wrote: >> On PeerDirect, we have some kind of a middle-ground solution for pinning >> GPU memory. We create a non-ODP MR pointing to VRAM but rely on >> user-space and the GPU not to migrate it. If they do, the MR gets >> destroyed immediately. > > That sounds horrible. How can that possibly work? What if the MR is > being used when the GPU decides to migrate? I would not support that > upstream without a lot more explanation.. Yup, this was our experience when playing around with PeerDirect. There was nothing we could do if the GPU decided to invalidate the P2P mapping. It just meant the application would fail or need complicated logic to detect this and redo just about everything. And given that it was a reasonably rare occurrence during development it probably means not a lot of applications will be developed to handle it and most would end up being randomly broken in environments with memory pressure. Logan
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2016-11-28 22:40 +0100 |
| Message-ID | <sIBOx-7vt-15@gated-at.bofh.it> |
| In reply to | #1531513 |
On 28/11/16 12:35 PM, Serguei Sagalovitch wrote: > As soon as PeerDirect mapping is called then GPU must not "move" the > such memory. It is by PeerDirect design. It is similar how it is works > with system memory and RDMA MR: when "get_user_pages" is called then the > memory is pinned. We haven't touch this in a long time and perhaps it changed, but there definitely was a call back in the PeerDirect API to allow the GPU to invalidate the mapping. That's what we don't want. Logan
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-11-28 23:30 +0100 |
| Message-ID | <sICAW-85L-5@gated-at.bofh.it> |
| In reply to | #1531685 |
On Mon, Nov 28, 2016 at 04:55:23PM -0500, Serguei Sagalovitch wrote: > >We haven't touch this in a long time and perhaps it changed, but there > >definitely was a call back in the PeerDirect API to allow the GPU to > >invalidate the mapping. That's what we don't want. > I assume that you are talking about "invalidate_peer_memory()' callback? > I was told that it is the "last resort" because HCA (and driver) is not > able to handle it in the safe manner so it is basically "abort" everything. If it is a last resort to save system stability then kill the impacted process, that will release the MRs. Jason
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-11-28 20:10 +0100 |
| Message-ID | <sIzto-6c0-5@gated-at.bofh.it> |
| In reply to | #1531445 |
On Mon, Nov 28, 2016 at 06:19:40PM +0000, Haggai Eran wrote: > > > GPU memory. We create a non-ODP MR pointing to VRAM but rely on > > > user-space and the GPU not to migrate it. If they do, the MR gets > > > destroyed immediately. > > That sounds horrible. How can that possibly work? What if the MR is > > being used when the GPU decides to migrate? > Naturally this doesn't support migration. The GPU is expected to pin > these pages as long as the MR lives. The MR invalidation is done only as > a last resort to keep system correctness. That just forces applications to handle horrible unexpected failures. If this sort of thing is needed for correctness then OOM kill the offending process, don't corrupt its operation. > I think it is similar to how non-ODP MRs rely on user-space today to > keep them correct. If you do something like madvise(MADV_DONTNEED) on a > non-ODP MR's pages, you can still get yourself into a data corruption > situation (HCA sees one page and the process sees another for the same > virtual address). The pinning that we use only guarentees the HCA's page > won't be reused. That is not really data corruption - the data still goes where it was originally destined. That is an application violating the requirements of a MR. An application cannot munmap/mremap a VMA while a non ODP MR points to it and then keep using the MR. That is totally different from a GPU driver wanthing to mess with translation to physical pages. > > From what I understand we are not really talking about kernel p2p, > > everything proposed so far is being mediated by a userspace VMA, so > > I'd focus on making that work. > Fair enough, although we will need both eventually, and I hope the > infrastructure can be shared to some degree. What use case do you see for in kernel? Presumably in-kernel could use a vmap or something and the same basic flow? Jason
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-11-30 17:30 +0100 |
| Message-ID | <sJfVD-8ry-9@gated-at.bofh.it> |
| In reply to | #1531541 |
On Wed, Nov 30, 2016 at 12:45:58PM +0200, Haggai Eran wrote: > > That just forces applications to handle horrible unexpected > > failures. If this sort of thing is needed for correctness then OOM > > kill the offending process, don't corrupt its operation. > Yes, that sounds fine. Can we simply kill the process from the GPU driver? > Or do we need to extend the OOM killer to manage GPU pages? I don't know.. > >>> From what I understand we are not really talking about kernel p2p, > >>> everything proposed so far is being mediated by a userspace VMA, so > >>> I'd focus on making that work. > > > >> Fair enough, although we will need both eventually, and I hope the > >> infrastructure can be shared to some degree. > > > > What use case do you see for in kernel? > Two cases I can think of are RDMA access to an NVMe device's controller > memory buffer, I'm not sure on the use model there.. > and O_DIRECT operations that access GPU memory. This goes through user space so there is still a VMA.. > Also, HMM's migration between two GPUs could use peer to peer in the > kernel, although that is intended to be handled by the GPU driver if > I understand correctly. Hum, presumably these migrations are VMA backed as well... > > Presumably in-kernel could use a vmap or something and the same basic > > flow? > I think we can achieve the kernel's needs with ZONE_DEVICE and DMA-API support > for peer to peer. I'm not sure we need vmap. We need a way to have a scatterlist > of MMIO pfns, and ZONE_DEVICE allows that. Well, if there is no virtual map then we are back to how do you do migrations and other things people seem to want to do on these pages. Maybe the loose 'struct page' flow is not for those users. But I think if you want kGPU or similar then you probably need vmaps or something similar to represent the GPU pages in kernel memory. Jason
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2016-11-30 19:10 +0100 |
| Message-ID | <sJhup-10J-11@gated-at.bofh.it> |
| In reply to | #1533420 |
On 30/11/16 09:23 AM, Jason Gunthorpe wrote: >> Two cases I can think of are RDMA access to an NVMe device's controller >> memory buffer, > > I'm not sure on the use model there.. The NVMe fabrics stuff could probably make use of this. It's an in-kernel system to allow remote access to an NVMe device over RDMA. So they ought to be able to optimize their transfers by DMAing directly to the NVMe's CMB -- no userspace interface would be required but there would need some kernel infrastructure. Logan
[toc] | [prev] | [next] | [standalone]
| From | "Stephen Bates" <sbates@raithlin.com> |
|---|---|
| Date | 2016-12-04 14:10 +0100 |
| Message-ID | <sKEIh-7zM-7@gated-at.bofh.it> |
| In reply to | #1533479 |
>> >> The NVMe fabrics stuff could probably make use of this. It's an >> in-kernel system to allow remote access to an NVMe device over RDMA. So >> they ought to be able to optimize their transfers by DMAing directly to >> the NVMe's CMB -- no userspace interface would be required but there >> would need some kernel infrastructure. > > Yes, that's what I was thinking. The NVMe/f driver needs to map the CMB > for RDMA. I guess if it used ZONE_DEVICE like in the iopmem patches it > would be relatively easy to do. > Haggai, yes that was one of the use cases we considered when we put together the patchset.
[toc] | [prev] | [next] | [standalone]
| From | "Stephen Bates" <sbates@raithlin.com> |
|---|---|
| Date | 2016-12-04 14:50 +0100 |
| Message-ID | <sKFkZ-7QW-7@gated-at.bofh.it> |
| In reply to | #1533479 |
Hi All This has been a great thread (thanks to Alex for kicking it off) and I wanted to jump in and maybe try and put some summary around the discussion. I also wanted to propose we include this as a topic for LFS/MM because I think we need more discussion on the best way to add this functionality to the kernel. As far as I can tell the people looking for P2P support in the kernel fall into two main camps: 1. Those who simply want to expose static BARs on PCIe devices that can be used as the source/destination for DMAs from another PCIe device. This group has no need for memory invalidation and are happy to use physical/bus addresses and not virtual addresses. 2. Those who want to support devices that suffer from occasional memory pressure and need to invalidate memory regions from time to time. This camp also would like to use virtual addresses rather than physical ones to allow for things like migration. I am wondering if people agree with this assessment? I think something like the iopmem patches Logan and I submitted recently come close to addressing use case 1. There are some issues around routability but based on feedback to date that does not seem to be a show-stopper for an initial inclusion. For use-case 2 it looks like there are several options and some of them (like HMM) have been around for quite some time without gaining acceptance. I think there needs to be more discussion on this usecase and it could be some time before we get something upstreamable. I for one, would really like to see use case 1 get addressed soon because we have consumers for it coming soon in the form of CMBs for NVMe devices. Long term I think Jason summed it up really well. CPU vendors will put high-speed, open, switchable, coherent buses on their processors and all these problems will vanish. But I ain't holding my breathe for that to happen ;-). Cheers Stephen
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-12-05 18:20 +0100 |
| Message-ID | <sL55L-7o6-29@gated-at.bofh.it> |
| In reply to | #1535684 |
On Sun, Dec 04, 2016 at 07:23:00AM -0600, Stephen Bates wrote: > Hi All > > This has been a great thread (thanks to Alex for kicking it off) and I > wanted to jump in and maybe try and put some summary around the > discussion. I also wanted to propose we include this as a topic for LFS/MM > because I think we need more discussion on the best way to add this > functionality to the kernel. > > As far as I can tell the people looking for P2P support in the kernel fall > into two main camps: > > 1. Those who simply want to expose static BARs on PCIe devices that can be > used as the source/destination for DMAs from another PCIe device. This > group has no need for memory invalidation and are happy to use > physical/bus addresses and not virtual addresses. I didn't think there was much on this topic except for the CMB thing.. Even that is really a mapped kernel address.. > I think something like the iopmem patches Logan and I submitted recently > come close to addressing use case 1. There are some issues around > routability but based on feedback to date that does not seem to be a > show-stopper for an initial inclusion. If it is kernel only with physical addresess we don't need a uAPI for it, so I'm not sure #1 is at all related to iopmem. Most people who want #1 probably can just mmap /sys/../pci/../resourceX to get a user handle to it, or pass around __iomem pointers in the kernel. This has been asked for before with RDMA. I'm still not really clear what iopmem is for, or why DAX should ever be involved in this.. > For use-case 2 it looks like there are several options and some of them > (like HMM) have been around for quite some time without gaining > acceptance. I think there needs to be more discussion on this usecase and > it could be some time before we get something upstreamable. AFAIK, hmm makes parts easier, but isn't directly addressing this need.. I think you need to get ZONE_DEVICE accepted for non-cachable PCI BARs as the first step. From there is pretty clear we the DMA API needs to be updated to support that use and work can be done to solve the various problems there on the basis of using ZONE_DEVICE pages to figure out to the PCI-E end points Jason
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-12-05 18:50 +0100 |
| Message-ID | <sL5yN-7xB-5@gated-at.bofh.it> |
| In reply to | #1536248 |
On Mon, Dec 5, 2016 at 9:18 AM, Jason Gunthorpe <jgunthorpe@obsidianresearch.com> wrote: > On Sun, Dec 04, 2016 at 07:23:00AM -0600, Stephen Bates wrote: >> Hi All >> >> This has been a great thread (thanks to Alex for kicking it off) and I >> wanted to jump in and maybe try and put some summary around the >> discussion. I also wanted to propose we include this as a topic for LFS/MM >> because I think we need more discussion on the best way to add this >> functionality to the kernel. >> >> As far as I can tell the people looking for P2P support in the kernel fall >> into two main camps: >> >> 1. Those who simply want to expose static BARs on PCIe devices that can be >> used as the source/destination for DMAs from another PCIe device. This >> group has no need for memory invalidation and are happy to use >> physical/bus addresses and not virtual addresses. > > I didn't think there was much on this topic except for the CMB > thing.. Even that is really a mapped kernel address.. > >> I think something like the iopmem patches Logan and I submitted recently >> come close to addressing use case 1. There are some issues around >> routability but based on feedback to date that does not seem to be a >> show-stopper for an initial inclusion. > > If it is kernel only with physical addresess we don't need a uAPI for > it, so I'm not sure #1 is at all related to iopmem. > > Most people who want #1 probably can just mmap > /sys/../pci/../resourceX to get a user handle to it, or pass around > __iomem pointers in the kernel. This has been asked for before with > RDMA. > > I'm still not really clear what iopmem is for, or why DAX should ever > be involved in this.. Right, by default remap_pfn_range() does not establish DMA capable mappings. You can think of iopmem as remap_pfn_range() converted to use devm_memremap_pages(). Given the extra constraints of devm_memremap_pages() it seems reasonable to have those DMA capable mappings be optionally established via a separate driver.
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-12-05 19:10 +0100 |
| Message-ID | <sL5Sa-7U6-5@gated-at.bofh.it> |
| In reply to | #1536273 |
On Mon, Dec 5, 2016 at 10:02 AM, Jason Gunthorpe <jgunthorpe@obsidianresearch.com> wrote: > On Mon, Dec 05, 2016 at 09:40:38AM -0800, Dan Williams wrote: > >> > If it is kernel only with physical addresess we don't need a uAPI for >> > it, so I'm not sure #1 is at all related to iopmem. >> > >> > Most people who want #1 probably can just mmap >> > /sys/../pci/../resourceX to get a user handle to it, or pass around >> > __iomem pointers in the kernel. This has been asked for before with >> > RDMA. >> > >> > I'm still not really clear what iopmem is for, or why DAX should ever >> > be involved in this.. >> >> Right, by default remap_pfn_range() does not establish DMA capable >> mappings. You can think of iopmem as remap_pfn_range() converted to >> use devm_memremap_pages(). Given the extra constraints of >> devm_memremap_pages() it seems reasonable to have those DMA capable >> mappings be optionally established via a separate driver. > > Except the iopmem driver claims the PCI ID, and presents a block > interface which is really *NOT* what people who have asked for this in > the past have wanted. IIRC it was embedded stuff eg RDMA video > directly out of a capture card or a similar kind of thinking. > > It is a good point about devm_memremap_pages limitations, but maybe > that just says to create a /sys/.../resource_dmableX ? > > Or is there some reason why people want a filesystem on top of BAR > memory? That does not seem to have been covered yet.. > I've already recommended that iopmem not be a block device and instead be a device-dax instance. I also don't think it should claim the PCI ID, rather the driver that wants to map one of its bars this way can register the memory region with the device-dax core. I'm not sure there are enough device drivers that want to do this to have it be a generic /sys/.../resource_dmableX capability. It still seems to be an exotic one-off type of configuration.
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2016-12-05 19:40 +0100 |
| Message-ID | <sL6lc-87B-29@gated-at.bofh.it> |
| In reply to | #1536286 |
On 05/12/16 11:08 AM, Dan Williams wrote: > I've already recommended that iopmem not be a block device and instead > be a device-dax instance. I also don't think it should claim the PCI > ID, rather the driver that wants to map one of its bars this way can > register the memory region with the device-dax core. > > I'm not sure there are enough device drivers that want to do this to > have it be a generic /sys/.../resource_dmableX capability. It still > seems to be an exotic one-off type of configuration. Yes, this is essentially my thinking. Except I think the userspace interface should really depend on the device itself. Device dax is a good choice for many and I agree the block device approach wouldn't be ideal. Specifically for NVME CMB: I think it would make a lot of sense to just hand out these mappings with an mmap call on /dev/nvmeX. I expect CMB buffers would be volatile and thus you wouldn't need to keep track of where in the BAR the region came from. Thus, the mmap call would just be an allocator from BAR memory. If device-dax were used, userspace would need to lookup which device-dax instance corresponds to which nvme drive. Logan
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-12-05 19:50 +0100 |
| Message-ID | <sL6uR-8aT-1@gated-at.bofh.it> |
| In reply to | #1536331 |
On Mon, Dec 5, 2016 at 10:39 AM, Logan Gunthorpe <logang@deltatee.com> wrote: > On 05/12/16 11:08 AM, Dan Williams wrote: >> >> I've already recommended that iopmem not be a block device and instead >> be a device-dax instance. I also don't think it should claim the PCI >> ID, rather the driver that wants to map one of its bars this way can >> register the memory region with the device-dax core. >> >> I'm not sure there are enough device drivers that want to do this to >> have it be a generic /sys/.../resource_dmableX capability. It still >> seems to be an exotic one-off type of configuration. > > > Yes, this is essentially my thinking. Except I think the userspace interface > should really depend on the device itself. Device dax is a good choice for > many and I agree the block device approach wouldn't be ideal. > > Specifically for NVME CMB: I think it would make a lot of sense to just hand > out these mappings with an mmap call on /dev/nvmeX. I expect CMB buffers > would be volatile and thus you wouldn't need to keep track of where in the > BAR the region came from. Thus, the mmap call would just be an allocator > from BAR memory. If device-dax were used, userspace would need to lookup > which device-dax instance corresponds to which nvme drive. > I'm not opposed to mapping /dev/nvmeX. However, the lookup is trivial to accomplish in sysfs through /sys/dev/char to find the sysfs path of the device-dax instance under the nvme device, or if you already have the nvme sysfs path the dax instance(s) will appear under the "dax" sub-directory.
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-12-05 20:20 +0100 |
| Message-ID | <sL6XT-81-15@gated-at.bofh.it> |
| In reply to | #1536333 |
On Mon, Dec 05, 2016 at 10:48:58AM -0800, Dan Williams wrote: > On Mon, Dec 5, 2016 at 10:39 AM, Logan Gunthorpe <logang@deltatee.com> wrote: > > On 05/12/16 11:08 AM, Dan Williams wrote: > >> > >> I've already recommended that iopmem not be a block device and instead > >> be a device-dax instance. I also don't think it should claim the PCI > >> ID, rather the driver that wants to map one of its bars this way can > >> register the memory region with the device-dax core. > >> > >> I'm not sure there are enough device drivers that want to do this to > >> have it be a generic /sys/.../resource_dmableX capability. It still > >> seems to be an exotic one-off type of configuration. > > > > > > Yes, this is essentially my thinking. Except I think the userspace interface > > should really depend on the device itself. Device dax is a good choice for > > many and I agree the block device approach wouldn't be ideal. > > > > Specifically for NVME CMB: I think it would make a lot of sense to just hand > > out these mappings with an mmap call on /dev/nvmeX. I expect CMB buffers > > would be volatile and thus you wouldn't need to keep track of where in the > > BAR the region came from. Thus, the mmap call would just be an allocator > > from BAR memory. If device-dax were used, userspace would need to lookup > > which device-dax instance corresponds to which nvme drive. > > I'm not opposed to mapping /dev/nvmeX. However, the lookup is trivial > to accomplish in sysfs through /sys/dev/char to find the sysfs path > of But CMB sounds much more like the GPU case where there is a specialized allocator handing out the BAR to consumers, so I'm not sure a general purpose chardev makes a lot of sense? Jason
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2016-12-05 20:30 +0100 |
| Message-ID | <sL77A-bb-7@gated-at.bofh.it> |
| In reply to | #1536350 |
On 05/12/16 12:14 PM, Jason Gunthorpe wrote: > But CMB sounds much more like the GPU case where there is a > specialized allocator handing out the BAR to consumers, so I'm not > sure a general purpose chardev makes a lot of sense? I don't think it will ever need to be as complicated as the GPU case. There will probably only ever be a relatively small amount of memory behind the CMB and really the only users are those doing P2P work. Thus the specialized allocator could be pretty simple and I expect it would be fine to just return -ENOMEM if there is not enough memory. Also, if it was implemented this way, if there was a need to make the allocator more complicated it could easily be added later as the userspace interface is just mmap to obtain a buffer. Logan
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-12-05 20:50 +0100 |
| Message-ID | <sL7qV-hs-15@gated-at.bofh.it> |
| In reply to | #1536353 |
On Mon, Dec 05, 2016 at 12:27:20PM -0700, Logan Gunthorpe wrote: > > > On 05/12/16 12:14 PM, Jason Gunthorpe wrote: > >But CMB sounds much more like the GPU case where there is a > >specialized allocator handing out the BAR to consumers, so I'm not > >sure a general purpose chardev makes a lot of sense? > > I don't think it will ever need to be as complicated as the GPU case. There > will probably only ever be a relatively small amount of memory behind the > CMB and really the only users are those doing P2P work. Thus the specialized > allocator could be pretty simple and I expect it would be fine to just > return -ENOMEM if there is not enough memory. NVMe might have to deal with pci-e hot-unplug, which is a similar problem-class to the GPU case.. In any event the allocator still needs to track which regions are in use and be able to hook 'free' from userspace. That does suggest it should be integrated into the nvme driver and not a bolt on driver.. Jason
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web