Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1530503 > unrolled thread
| Started by | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| First post | 2016-11-25 20:40 +0100 |
| Last post | 2016-11-28 19:40 +0100 |
| Articles | 13 on this page of 33 — 8 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-25 20:40 +0100
Re: Enabling peer to peer device transactions for PCIe devices Christian König <deathsimple@vodafone.de> - 2016-11-25 21:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-25 22:20 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-28 18:00 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-11-28 19:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-11-28 22:40 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-28 23:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-28 20:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-11-30 17:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-11-30 19:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices "Stephen Bates" <sbates@raithlin.com> - 2016-12-04 14:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices "Stephen Bates" <sbates@raithlin.com> - 2016-12-04 14:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-05 18:20 +0100
Re: Enabling peer to peer device transactions for PCIe devices Dan Williams <dan.j.williams@intel.com> - 2016-12-05 18:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Dan Williams <dan.j.williams@intel.com> - 2016-12-05 19:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-12-05 19:40 +0100
Re: Enabling peer to peer device transactions for PCIe devices Dan Williams <dan.j.williams@intel.com> - 2016-12-05 19:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-05 20:20 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-12-05 20:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-05 20:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-12-05 21:00 +0100
Re: Enabling peer to peer device transactions for PCIe devices Christoph Hellwig <hch@infradead.org> - 2016-12-05 21:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices "Stephen Bates" <sbates@raithlin.com> - 2016-12-06 09:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-06 17:40 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-12-06 18:00 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-06 18:30 +0100
Re: Enabling peer to peer device transactions for PCIe devices Logan Gunthorpe <logang@deltatee.com> - 2016-12-06 22:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Dan Williams <dan.j.williams@intel.com> - 2016-12-06 23:10 +0100
Re: Enabling peer to peer device transactions for PCIe devices Christoph Hellwig <hch@infradead.org> - 2016-12-06 18:20 +0100
Re: Enabling peer to peer device transactions for PCIe devices Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2016-12-05 19:10 +0100
RE: Enabling peer to peer device transactions for PCIe devices "Deucher, Alexander" <Alexander.Deucher@amd.com> - 2016-11-30 18:50 +0100
Re: Enabling peer to peer device transactions for PCIe devices Haggai Eran <haggaie@mellanox.com> - 2016-11-28 21:00 +0100
Re: Enabling peer to peer device transactions for PCIe devices Haggai Eran <haggaie@mellanox.com> - 2016-11-28 19:40 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2016-12-05 21:00 +0100 |
| Message-ID | <sL7AC-kR-29@gated-at.bofh.it> |
| In reply to | #1536365 |
On 05/12/16 12:46 PM, Jason Gunthorpe wrote: > NVMe might have to deal with pci-e hot-unplug, which is a similar > problem-class to the GPU case.. Sure, but if the NVMe device gets hot-unplugged it means that all the CMB mappings are useless and need to be torn down. This probably means killing any process that has mappings open. > In any event the allocator still needs to track which regions are in > use and be able to hook 'free' from userspace. That does suggest it > should be integrated into the nvme driver and not a bolt on driver.. Yup, that's correct. And yes, I've never suggested this to be a bolt on driver -- I always expected for it to get integrated into the nvme driver. (iopmem was not meant for this.) Logan
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-12-05 21:10 +0100 |
| Message-ID | <sL7Ki-Dr-29@gated-at.bofh.it> |
| In reply to | #1536365 |
On Mon, Dec 05, 2016 at 12:46:14PM -0700, Jason Gunthorpe wrote: > In any event the allocator still needs to track which regions are in > use and be able to hook 'free' from userspace. That does suggest it > should be integrated into the nvme driver and not a bolt on driver.. Two totally different use cases: - a card that exposes directly byte addressable storage as a PCI-e bar. Thin of it as a nvdimm on a PCI-e card. That's the iopmem case. - the NVMe CMB which exposes a byte addressable indirection buffer for I/O, but does not actually provide byte addressable persistent storage. This is something that needs to be added to the NVMe driver (and the block layer for the abstraction probably).
[toc] | [prev] | [next] | [standalone]
| From | "Stephen Bates" <sbates@raithlin.com> |
|---|---|
| Date | 2016-12-06 09:30 +0100 |
| Message-ID | <sLjip-84n-17@gated-at.bofh.it> |
| In reply to | #1536333 |
>>> I've already recommended that iopmem not be a block device and >>> instead be a device-dax instance. I also don't think it should claim >>> the PCI ID, rather the driver that wants to map one of its bars this >>> way can register the memory region with the device-dax core. >>> >>> I'm not sure there are enough device drivers that want to do this to >>> have it be a generic /sys/.../resource_dmableX capability. It still >>> seems to be an exotic one-off type of configuration. >> >> >> Yes, this is essentially my thinking. Except I think the userspace >> interface should really depend on the device itself. Device dax is a >> good choice for many and I agree the block device approach wouldn't be >> ideal. I tend to agree here. The block device interface has seen quite a bit of resistance and /dev/dax looks like a better approach for most. We can look at doing it that way in v2. >> >> Specifically for NVME CMB: I think it would make a lot of sense to just >> hand out these mappings with an mmap call on /dev/nvmeX. I expect CMB >> buffers would be volatile and thus you wouldn't need to keep track of >> where in the BAR the region came from. Thus, the mmap call would just be >> an allocator from BAR memory. If device-dax were used, userspace would >> need to lookup which device-dax instance corresponds to which nvme >> drive. >> > > I'm not opposed to mapping /dev/nvmeX. However, the lookup is trivial > to accomplish in sysfs through /sys/dev/char to find the sysfs path of the > device-dax instance under the nvme device, or if you already have the nvme > sysfs path the dax instance(s) will appear under the "dax" sub-directory. > Personally I think mapping the dax resource in the sysfs tree is a nice way to do this and a bit more intuitive than mapping a /dev/nvmeX.
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-12-06 17:40 +0100 |
| Message-ID | <sLqWB-4tQ-21@gated-at.bofh.it> |
| In reply to | #1536753 |
> > I'm not opposed to mapping /dev/nvmeX. However, the lookup is trivial > > to accomplish in sysfs through /sys/dev/char to find the sysfs path of the > > device-dax instance under the nvme device, or if you already have the nvme > > sysfs path the dax instance(s) will appear under the "dax" sub-directory. > > Personally I think mapping the dax resource in the sysfs tree is a nice > way to do this and a bit more intuitive than mapping a /dev/nvmeX. It is still not at all clear to me what userpsace is supposed to do with this on nvme.. How is the CMB usable from userspace? Jason
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2016-12-06 18:00 +0100 |
| Message-ID | <sLrfX-4AN-21@gated-at.bofh.it> |
| In reply to | #1537096 |
Hey, On 06/12/16 09:38 AM, Jason Gunthorpe wrote: >>> I'm not opposed to mapping /dev/nvmeX. However, the lookup is trivial >>> to accomplish in sysfs through /sys/dev/char to find the sysfs path of the >>> device-dax instance under the nvme device, or if you already have the nvme >>> sysfs path the dax instance(s) will appear under the "dax" sub-directory. >> >> Personally I think mapping the dax resource in the sysfs tree is a nice >> way to do this and a bit more intuitive than mapping a /dev/nvmeX. > > It is still not at all clear to me what userpsace is supposed to do > with this on nvme.. How is the CMB usable from userspace? The flow is pretty simple. For example to write to NVMe from an RDMA device: 1) Obtain a chunk of the CMB to use as a buffer(either by mmaping /dev/nvmx, the device dax char device or through a block layer interface (which sounds like a good suggestion from Christoph, but I'm not really sure how it would look). 2) Create an MR with the buffer and use an RDMA function to fill it with data from a remote host. This will cause the RDMA hardware to write directly to the memory in the NVMe card. 3) Using O_DIRECT, write the buffer to a file on the NVMe filesystem. When the address reaches hardware the NVMe will recognize it as local memory and copy it directly there. Thus we are able to transfer data to any file on an NVMe device without going through system memory. This has benefits on systems with lots of activity in system memory but step 3 is likely to be slowish due to the need to pin/unpin the memory for every transaction. Logan
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-12-06 18:30 +0100 |
| Message-ID | <sLrJ0-4ZP-31@gated-at.bofh.it> |
| In reply to | #1537118 |
On Tue, Dec 06, 2016 at 09:51:15AM -0700, Logan Gunthorpe wrote: > Hey, > > On 06/12/16 09:38 AM, Jason Gunthorpe wrote: > >>> I'm not opposed to mapping /dev/nvmeX. However, the lookup is trivial > >>> to accomplish in sysfs through /sys/dev/char to find the sysfs path of the > >>> device-dax instance under the nvme device, or if you already have the nvme > >>> sysfs path the dax instance(s) will appear under the "dax" sub-directory. > >> > >> Personally I think mapping the dax resource in the sysfs tree is a nice > >> way to do this and a bit more intuitive than mapping a /dev/nvmeX. > > > > It is still not at all clear to me what userpsace is supposed to do > > with this on nvme.. How is the CMB usable from userspace? > > The flow is pretty simple. For example to write to NVMe from an RDMA device: > > 1) Obtain a chunk of the CMB to use as a buffer(either by mmaping > /dev/nvmx, the device dax char device or through a block layer interface > (which sounds like a good suggestion from Christoph, but I'm not really > sure how it would look). Okay, so clearly this needs a kernel side NVMe specific allocator and locking so users don't step on each other.. Or as Christoph says some kind of general mechanism to get these bounce buffers.. > 2) Create an MR with the buffer and use an RDMA function to fill it with > data from a remote host. This will cause the RDMA hardware to write > directly to the memory in the NVMe card. > > 3) Using O_DIRECT, write the buffer to a file on the NVMe filesystem. > When the address reaches hardware the NVMe will recognize it as local > memory and copy it directly there. Ah, I see. As a first draft I'd stick with some kind of API built into the /dev/nvmeX that backs the filesystem. The user app would fstat the target file, open /dev/block/MAJOR(st_dev):MINOR(st_dev), do some ioctl to get a CMB mmap, and then proceed from there.. When that is all working kernel-side, it would make sense to look at a more general mechanism that could be used unprivileged?? > Thus we are able to transfer data to any file on an NVMe device without > going through system memory. This has benefits on systems with lots of > activity in system memory but step 3 is likely to be slowish due to the > need to pin/unpin the memory for every transaction. This is similar to the GPU issues too.. On NVMe you don't need to pin the pages, you just need to lock that VMA so it doesn't get freed from the NVMe CMB allocator while the IO is running... Probably in the long run the get_user_pages is going to have to be pushed down into drivers.. Future MMU coherent IO hardware also does not need the pinning or other overheads. Jason
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2016-12-06 22:50 +0100 |
| Message-ID | <sLvMC-7vq-15@gated-at.bofh.it> |
| In reply to | #1537153 |
Hey, > Okay, so clearly this needs a kernel side NVMe specific allocator > and locking so users don't step on each other.. Yup, ideally. That's why device dax isn't ideal for this application: it doesn't provide any way to prevent users from stepping on each other. > Or as Christoph says some kind of general mechanism to get these > bounce buffers.. Yeah, I imagine a general allocate from BAR/region system would be very useful. > Ah, I see. > > As a first draft I'd stick with some kind of API built into the > /dev/nvmeX that backs the filesystem. The user app would fstat the > target file, open /dev/block/MAJOR(st_dev):MINOR(st_dev), do some > ioctl to get a CMB mmap, and then proceed from there.. > > When that is all working kernel-side, it would make sense to look at a > more general mechanism that could be used unprivileged?? That makes a lot of sense to me. I suggested mmapping the char device because it's really easy, but I can see that an ioctl on the block device does seem more general and device agnostic. > This is similar to the GPU issues too.. On NVMe you don't need to pin > the pages, you just need to lock that VMA so it doesn't get freed from > the NVMe CMB allocator while the IO is running... > Probably in the long run the get_user_pages is going to have to be > pushed down into drivers.. Future MMU coherent IO hardware also does > not need the pinning or other overheads. Yup. Yup. Logan
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-12-06 23:10 +0100 |
| Message-ID | <sLw5X-7QZ-17@gated-at.bofh.it> |
| In reply to | #1537286 |
On Tue, Dec 6, 2016 at 1:47 PM, Logan Gunthorpe <logang@deltatee.com> wrote: > Hey, > >> Okay, so clearly this needs a kernel side NVMe specific allocator >> and locking so users don't step on each other.. > > Yup, ideally. That's why device dax isn't ideal for this application: it > doesn't provide any way to prevent users from stepping on each other. On this particular point I'm in the process of posting patches that allow device-dax sub-division, so you could carve up a bar into multiple devices of various sizes.
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-12-06 18:20 +0100 |
| Message-ID | <sLrzj-4WB-21@gated-at.bofh.it> |
| In reply to | #1537096 |
On Tue, Dec 06, 2016 at 09:38:50AM -0700, Jason Gunthorpe wrote: > > > I'm not opposed to mapping /dev/nvmeX. However, the lookup is trivial > > > to accomplish in sysfs through /sys/dev/char to find the sysfs path of the > > > device-dax instance under the nvme device, or if you already have the nvme > > > sysfs path the dax instance(s) will appear under the "dax" sub-directory. > > > > Personally I think mapping the dax resource in the sysfs tree is a nice > > way to do this and a bit more intuitive than mapping a /dev/nvmeX. > > It is still not at all clear to me what userpsace is supposed to do > with this on nvme.. How is the CMB usable from userspace? I don't think trying to expose it to userspace makes any sense. Exposing it to in-kernel storage targets on the other hand makes a lot of sense.
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2016-12-05 19:10 +0100 |
| Message-ID | <sL5Sa-7U6-7@gated-at.bofh.it> |
| In reply to | #1536273 |
On Mon, Dec 05, 2016 at 09:40:38AM -0800, Dan Williams wrote: > > If it is kernel only with physical addresess we don't need a uAPI for > > it, so I'm not sure #1 is at all related to iopmem. > > > > Most people who want #1 probably can just mmap > > /sys/../pci/../resourceX to get a user handle to it, or pass around > > __iomem pointers in the kernel. This has been asked for before with > > RDMA. > > > > I'm still not really clear what iopmem is for, or why DAX should ever > > be involved in this.. > > Right, by default remap_pfn_range() does not establish DMA capable > mappings. You can think of iopmem as remap_pfn_range() converted to > use devm_memremap_pages(). Given the extra constraints of > devm_memremap_pages() it seems reasonable to have those DMA capable > mappings be optionally established via a separate driver. Except the iopmem driver claims the PCI ID, and presents a block interface which is really *NOT* what people who have asked for this in the past have wanted. IIRC it was embedded stuff eg RDMA video directly out of a capture card or a similar kind of thinking. It is a good point about devm_memremap_pages limitations, but maybe that just says to create a /sys/.../resource_dmableX ? Or is there some reason why people want a filesystem on top of BAR memory? That does not seem to have been covered yet.. Jason
[toc] | [prev] | [next] | [standalone]
| From | "Deucher, Alexander" <Alexander.Deucher@amd.com> |
|---|---|
| Date | 2016-11-30 18:50 +0100 |
| Message-ID | <sJhb3-F2-1@gated-at.bofh.it> |
| In reply to | #1531541 |
> -----Original Message----- > From: Haggai Eran [mailto:haggaie@mellanox.com] > Sent: Wednesday, November 30, 2016 5:46 AM > To: Jason Gunthorpe > Cc: linux-kernel@vger.kernel.org; linux-rdma@vger.kernel.org; linux- > nvdimm@ml01.01.org; Koenig, Christian; Suthikulpanit, Suravee; Bridgman, > John; Deucher, Alexander; Linux-media@vger.kernel.org; > dan.j.williams@intel.com; logang@deltatee.com; dri- > devel@lists.freedesktop.org; Max Gurtovoy; linux-pci@vger.kernel.org; > Sagalovitch, Serguei; Blinzer, Paul; Kuehling, Felix; Sander, Ben > Subject: Re: Enabling peer to peer device transactions for PCIe devices > > On 11/28/2016 9:02 PM, Jason Gunthorpe wrote: > > On Mon, Nov 28, 2016 at 06:19:40PM +0000, Haggai Eran wrote: > >>>> GPU memory. We create a non-ODP MR pointing to VRAM but rely on > >>>> user-space and the GPU not to migrate it. If they do, the MR gets > >>>> destroyed immediately. > >>> That sounds horrible. How can that possibly work? What if the MR is > >>> being used when the GPU decides to migrate? > >> Naturally this doesn't support migration. The GPU is expected to pin > >> these pages as long as the MR lives. The MR invalidation is done only as > >> a last resort to keep system correctness. > > > > That just forces applications to handle horrible unexpected > > failures. If this sort of thing is needed for correctness then OOM > > kill the offending process, don't corrupt its operation. > Yes, that sounds fine. Can we simply kill the process from the GPU driver? > Or do we need to extend the OOM killer to manage GPU pages? Christian sent out an RFC patch a while back that extended the OOM to cover memory allocated for the GPU: https://lists.freedesktop.org/archives/dri-devel/2015-September/089778.html Alex > > > > >> I think it is similar to how non-ODP MRs rely on user-space today to > >> keep them correct. If you do something like madvise(MADV_DONTNEED) > on a > >> non-ODP MR's pages, you can still get yourself into a data corruption > >> situation (HCA sees one page and the process sees another for the same > >> virtual address). The pinning that we use only guarentees the HCA's page > >> won't be reused. > > > > That is not really data corruption - the data still goes where it was > > originally destined. That is an application violating the > > requirements of a MR. > I guess it is a matter of terminology. If you compare it to the ODP case > or the CPU case then you usually expect a single virtual address to map to > a single physical page. Violating this cause some of your writes to be dropped > which is a data corruption in my book, even if the application caused it. > > > An application cannot munmap/mremap a VMA > > while a non ODP MR points to it and then keep using the MR. > Right. And it is perfectly fine to have some similar requirements from the > application > when doing peer to peer with a non-ODP MR. > > > That is totally different from a GPU driver wanthing to mess with > > translation to physical pages. > > > >>> From what I understand we are not really talking about kernel p2p, > >>> everything proposed so far is being mediated by a userspace VMA, so > >>> I'd focus on making that work. > > > >> Fair enough, although we will need both eventually, and I hope the > >> infrastructure can be shared to some degree. > > > > What use case do you see for in kernel? > Two cases I can think of are RDMA access to an NVMe device's controller > memory buffer, and O_DIRECT operations that access GPU memory. > Also, HMM's migration between two GPUs could use peer to peer in the > kernel, > although that is intended to be handled by the GPU driver if I understand > correctly. > > > Presumably in-kernel could use a vmap or something and the same basic > > flow? > I think we can achieve the kernel's needs with ZONE_DEVICE and DMA-API > support > for peer to peer. I'm not sure we need vmap. We need a way to have a > scatterlist > of MMIO pfns, and ZONE_DEVICE allows that. > > Haggai
[toc] | [prev] | [next] | [standalone]
| From | Haggai Eran <haggaie@mellanox.com> |
|---|---|
| Date | 2016-11-28 21:00 +0100 |
| Message-ID | <sIzto-6c0-7@gated-at.bofh.it> |
| In reply to | #1531445 |
On Mon, 2016-11-28 at 09:57 -0700, Jason Gunthorpe wrote: +AD4- On Sun, Nov 27, 2016 at 04:02:16PM +-0200, Haggai Eran wrote: +AD4- +AD4- I think blocking mmu notifiers against something that is basica= lly +AD4- +AD4- controlled by user-space can be problematic. This can block thi= ngs +AD4- +AD4- like +AD4- +AD4- memory reclaim. If you have user-space access to the device's +AD4- +AD4- queues, +AD4- +AD4- user-space can block the mmu notifier forever. +AD4- Right, I mentioned that.. Sorry, I must have missed it. +AD4- +AD4- On PeerDirect, we have some kind of a middle-ground solution fo= r +AD4- +AD4- pinning +AD4- +AD4- GPU memory. We create a non-ODP MR pointing to VRAM but rely on +AD4- +AD4- user-space and the GPU not to migrate it. If they do, the MR ge= ts +AD4- +AD4- destroyed immediately. +AD4- That sounds horrible. How can that possibly work? What if the MR is +AD4- being used when the GPU decides to migrate?=20 Naturally this doesn't support migration. The GPU is expected to pin these pages as long as the MR lives. The MR invalidation is done only as a last resort to keep system correctness. I think it is similar to how non-ODP MRs rely on user-space today to keep them correct. If you do something like madvise(MADV+AF8-DONTNEED) on a non-ODP MR's pages, you can still get yourself into a data corruption situation (HCA sees one page and the process sees another for the same virtual address). The pinning that we use only guarentees the HCA's page won't be reused. +AD4- I would not support that +AD4- upstream without a lot more explanation.. +AD4-=20 +AD4- I know people don't like requiring new hardware, but in this case we +AD4- really do need ODP hardware to get all the semantics people want.. +AD4-=20 +AD4- +AD4-=20 +AD4- +AD4- Another thing I think is that while HMM is good for user-space +AD4- +AD4- applications, for kernel p2p use there is no need for that. Usi= ng +AD4- From what I understand we are not really talking about kernel p2p, +AD4- everything proposed so far is being mediated by a userspace VMA, so +AD4- I'd focus on making that work. Fair enough, although we will need both eventually, and I hope the infrastructure can be shared to some degree.=
[toc] | [prev] | [next] | [standalone]
| From | Haggai Eran <haggaie@mellanox.com> |
|---|---|
| Date | 2016-11-28 19:40 +0100 |
| Message-ID | <sIz0m-5MH-29@gated-at.bofh.it> |
| In reply to | #1530503 |
On Mon, 2016-11-28 at 09:48 -0500, Serguei Sagalovitch wrote: +AD4- On 2016-11-27 09:02 AM, Haggai Eran wrote +AD4- +AD4- +AD4- +AD4- On PeerDirect, we have some kind of a middle-ground solution for +AD4- +AD4- pinning +AD4- +AD4- GPU memory. We create a non-ODP MR pointing to VRAM but rely on +AD4- +AD4- user-space and the GPU not to migrate it. If they do, the MR gets +AD4- +AD4- destroyed immediately. This should work on legacy devices without +AD4- +AD4- ODP +AD4- +AD4- support, and allows the system to safely terminate a process that +AD4- +AD4- misbehaves. The downside of course is that it cannot transparently +AD4- +AD4- migrate memory but I think for user-space RDMA doing that +AD4- +AD4- transparently +AD4- +AD4- requires hardware support for paging, via something like HMM. +AD4- +AD4- +AD4- +AD4- ... +AD4- May be I am wrong but my understanding is that PeerDirect logic +AD4- basically +AD4- follow+AKAAoAAi-RDMA register MR+ACI- logic Yes. The only difference from regular MRs is the invalidation process I mentioned, and the fact that we get the addresses not from get+AF8-user+AF8-pages but from a peer driver. +AD4- so basically nothing prevent to +ACI-terminate+ACI- +AD4- process for +ACI-MMU notifier+ACI- case when we are very low on memory +AD4- not making it similar (not worse) then PeerDirect case. I'm not sure I understand. I don't think any solution prevents terminating an application. The paragraph above is just trying to explain how a non-ODP device/MR can handle an invalidation. +AD4- +AD4- +AD4- I'm hearing most people say ZONE+AF8-DEVICE is the way to handle this, +AD4- +AD4- +AD4- which means the missing remaing piece for RDMA is some kind of DMA +AD4- +AD4- +AD4- core support for p2p address translation.. +AD4- +AD4- Yes, this is definitely something we need. I think Will Davis's +AD4- +AD4- patches +AD4- +AD4- are a good start. +AD4- +AD4- +AD4- +AD4- Another thing I think is that while HMM is good for user-space +AD4- +AD4- applications, for kernel p2p use there is no need for that. +AD4- About HMM: I do not think that in the current form HMM would+AKAAoA-fit in +AD4- requirement for generic P2P transfer case. My understanding is that at +AD4- the current stage HMM is good for +ACI-caching+ACI- system memory +AD4- in device memory for fast GPU access but in RDMA MR non-ODP case +AD4- it will not work because+AKAAoA-the location of memory should not be +AD4- changed so memory should be allocated directly in PCIe memory. The way I see it there are two ways to handle non-ODP MRs. Either you prevent the GPU from migrating / reusing the MR's VRAM pages for as long as the MR is alive (if I understand correctly you didn't like this solution), or you allow the GPU to somehow notify the HCA to invalidate the MR. If you do that, you can use mmu notifiers or HMM or something else, but HMM provides a nice framework to facilitate that notification. +AD4- +AD4- +AD4- +AD4- Using ZONE+AF8-DEVICE with or without something like DMA-BUF to pin and +AD4- +AD4- unpin +AD4- +AD4- pages for the short duration as you wrote above could work fine for +AD4- +AD4- kernel uses in which we can guarantee they are short. +AD4- Potentially there is another issue related to pin/unpin. If memory +AD4- could +AD4- be used a lot of time then there is no sense to rebuild and program +AD4- s/g tables each time if location of memory was not changed. Is this about the kernel use or user-space? In user-space I think the MR concept captures a long-lived s/g table so you don't need to rebuild it (unless the mapping changes). Haggai=
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web