Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1624403 > unrolled thread
| Started by | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| First post | 2017-04-16 17:50 +0200 |
| Last post | 2017-04-20 02:10 +0200 |
| Articles | 20 on this page of 74 — 7 participants |
Back to article view | Back to linux.kernel
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-16 17:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-16 18:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-17 00:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-17 07:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-17 09:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-17 19:00 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-17 19:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-18 07:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jerome Glisse <jglisse@redhat.com> - 2017-04-17 20:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-18 08:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-17 23:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-18 07:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-18 08:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-17 00:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-18 18:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-18 19:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-18 20:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-18 20:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-19 03:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-19 01:00 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-19 02:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-18 20:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-18 21:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-18 21:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-18 21:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-18 22:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-18 21:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jerome Glisse <jglisse@redhat.com> - 2017-04-18 22:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-18 22:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-18 22:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-19 03:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-18 23:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-18 23:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-18 23:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-18 23:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-19 00:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-19 00:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-19 00:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-19 01:00 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-19 01:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-19 03:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-19 00:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-19 01:00 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-19 01:00 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-19 01:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-19 03:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-18 23:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-19 00:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-19 01:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-19 04:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-19 03:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-19 18:00 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-19 18:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-19 19:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jerome Glisse <jglisse@redhat.com> - 2017-04-19 19:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-19 19:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-19 20:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-19 20:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-19 20:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-19 20:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-19 20:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-20 22:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory "Stephen Bates" <sbates@raithlin.com> - 2017-04-21 01:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-21 07:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory "Stephen Bates" <sbates@raithlin.com> - 2017-04-20 22:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-19 19:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-19 20:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-19 20:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-19 21:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-19 21:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-19 21:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-19 22:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-20 01:00 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@gmail.com> - 2017-04-20 02:10 +0200
Page 2 of 4 — ← Prev page 1 [2] 3 4 Next page →
| From | Benjamin Herrenschmidt <benh@kernel.crashing.org> |
|---|---|
| Date | 2017-04-19 02:10 +0200 |
| Message-ID | <txKgj-3qF-21@gated-at.bofh.it> |
| In reply to | #1625488 |
On Tue, 2017-04-18 at 10:27 -0700, Dan Williams wrote: > > FWIW, RDMA probably wouldn't want to use a p2mem device either, we > > already have APIs that map BAR memory to user space, and would like to > > keep using them. A 'enable P2P for bar' helper function sounds better > > to me. > > ...and I think it's not a helper function as much as asking the bus > provider "can these two device dma to each other". The "helper" is the > dma api redirecting through a software-iommu that handles bus address > translation differently than it would handle host memory dma mapping. Do we even need tat function ? The dma_ops have a dma_supported() call... If we have those override ops built into the "dma_target" object, then these things can make that decision knowing both the source and target device. Cheers, Ben.
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2017-04-18 20:40 +0200 |
| Message-ID | <txGcF-140-3@gated-at.bofh.it> |
| In reply to | #1625454 |
On 18/04/17 10:45 AM, Jason Gunthorpe wrote: > From Ben's comments, I would think that the 'first class' support that > is needed here is simply a function to return the 'struct device' > backing a CPU address range. Yes, and Dan's get_dev_pagemap suggestion gets us 90% of the way there. It's just a disagreement as to what struct device is inside the pagemap. Care needs to be taken to ensure that struct device doesn't conflict with hmm and doesn't limit other potential future users of ZONE_DEVICE. > If there is going to be more core support for this stuff I think it > will be under the topic of more robustly describing the fabric to the > core and core helpers to extract data from the description: eg compute > the path, check if the path crosses translation, etc Agreed, those helpers would be useful to everyone. > I think the key agreement to get out of Logan's series is that P2P DMA > means: > - The BAR will be backed by struct pages > - Passing the CPU __iomem address of the BAR to the DMA API is > valid and, long term, dma ops providers are expected to fail > or return the right DMA address Well, yes but we have a _lot_ of work to do to make it safe to pass around struct pages backed with __iomem. That's where our next focus will be. I've already taken very initial steps toward this with my scatterlist map patchset. > - Mapping BAR memory into userspace and back to the kernel via > get_user_pages works transparently, and with the DMA API above Again, we've had a lot of push back for the memory to go to userspace at all. It does work, but people expect userspace to screw it up in a lot of ways. Among the people pushing back on that: Christoph Hellwig has specifically said he wants to see this stay with in-kernel users only until the apis can be worked out. This is one of the reasons we decided to go with enabling nvme-fabrics as everything remains in the kernel. And with that decision we needed a common in-kernel allocation infrastructure: this is what p2pmem really is at this point. > - The dma ops provider must be able to tell if source memory is bar > mapped and recover the pci device backing the mapping. Do you mean to say that every dma-ops provider needs to be taught about p2p backed pages? I was hoping we could have dma_map_* just use special p2p dma-ops if it was passed p2p pages (though there are some complications to this too). > At least this is what we'd like in RDMA :) > > FWIW, RDMA probably wouldn't want to use a p2mem device either, we > already have APIs that map BAR memory to user space, and would like to > keep using them. A 'enable P2P for bar' helper function sounds better > to me. Well, in the end that will likely come down to just devm_memremap_pages with some (presently undecided) struct device that can be used to get special p2p dma-ops for the bus. Logan
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2017-04-18 21:10 +0200 |
| Message-ID | <txGFI-1tx-11@gated-at.bofh.it> |
| In reply to | #1625525 |
On Tue, Apr 18, 2017 at 12:30:59PM -0600, Logan Gunthorpe wrote:
> > - The dma ops provider must be able to tell if source memory is bar
> > mapped and recover the pci device backing the mapping.
>
> Do you mean to say that every dma-ops provider needs to be taught about
> p2p backed pages? I was hoping we could have dma_map_* just use special
> p2p dma-ops if it was passed p2p pages (though there are some
> complications to this too).
I think that is how it will end up working out if this is the path..
Ultimately every dma_ops will need special code to support P2P with
the special hardware that ops is controlling, so it makes some sense
to start by pushing the check down there in the first place. This
advice is partially motivated by how dma_map_sg is just a small
wrapper around the function pointer call...
Something like:
foo_dma_map_sg(...)
{
for (every page in sg)
if (page is p2p)
dma_addr[I] = p2p_same_segment_map_page(...);
}
Where p2p_same_segment_map_page checks if the two devices are on the
'same switch' and if so returns the address translated to match the
bus address programmed into the BAR or fails. We knows this case is
required to work by the PCI spec, so it makes sense to use it as the
first canned helper.
This also proves out the basic idea that the dma ops can recover the
pci device and perform an inspection of the traversed fabric path.
From there every arch would have to expand the implementation to
support a wider range of things. Eg x86 with no iommu and no offset
could allow every address to be used based on a host bridge white
list.
Jason
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2017-04-18 21:40 +0200 |
| Message-ID | <txH8K-1CQ-1@gated-at.bofh.it> |
| In reply to | #1625554 |
On 18/04/17 01:01 PM, Jason Gunthorpe wrote: > Ultimately every dma_ops will need special code to support P2P with > the special hardware that ops is controlling, so it makes some sense > to start by pushing the check down there in the first place. This > advice is partially motivated by how dma_map_sg is just a small > wrapper around the function pointer call... Yes, I noticed this problem too and that makes sense. It just means every dma_ops will probably need to be modified to either support p2p pages or fail on them. Though, the only real difficulty there is that it will be a lot of work. > Where p2p_same_segment_map_page checks if the two devices are on the > 'same switch' and if so returns the address translated to match the > bus address programmed into the BAR or fails. We knows this case is > required to work by the PCI spec, so it makes sense to use it as the > first canned helper. I've also suggested that this check should probably be done (or perhaps duplicated) before we even get to the map stage. In the case of nvme-fabrics we'd probably want to let the user know when they try to configure it or at least fall back to allocating regular memory instead. It would be a difficult situation to have already copied a block of data from a NIC to p2p memory only to have it be deemed unmappable on the NVMe device it's destined for. (Or vice-versa.) This was another issue p2pmem was attempting to solve. Logan
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2017-04-18 21:50 +0200 |
| Message-ID | <txHiq-1Gh-11@gated-at.bofh.it> |
| In reply to | #1625569 |
On Tue, Apr 18, 2017 at 01:35:32PM -0600, Logan Gunthorpe wrote: > > Ultimately every dma_ops will need special code to support P2P with > > the special hardware that ops is controlling, so it makes some sense > > to start by pushing the check down there in the first place. This > > advice is partially motivated by how dma_map_sg is just a small > > wrapper around the function pointer call... > > Yes, I noticed this problem too and that makes sense. It just means > every dma_ops will probably need to be modified to either support p2p > pages or fail on them. Though, the only real difficulty there is that it > will be a lot of work. I think this is why progress on this keeps getting stuck - every solution is a lot of work. > > Where p2p_same_segment_map_page checks if the two devices are on the > > 'same switch' and if so returns the address translated to match the > > bus address programmed into the BAR or fails. We knows this case is > > required to work by the PCI spec, so it makes sense to use it as the > > first canned helper. > > I've also suggested that this check should probably be done (or perhaps > duplicated) before we even get to the map stage. Since the mechanics of the check is essentially unique to every dma-ops I would not hoist it out of the map function without a really good reason. > In the case of nvme-fabrics we'd probably want to let the user know > when they try to configure it or at least fall back to allocating > regular memory instead. You could try to do a dummy mapping / create a MR early on to detect this. FWIW, I wonder if from a RDMA perspective we have another problem.. Should we allow P2P memory to be used with the local DMA lkey? There are potential designs around virtualization that would not allow that. Should we mandate that P2P memory be in its own MR? Jason
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2017-04-18 22:10 +0200 |
| Message-ID | <txHBM-22s-9@gated-at.bofh.it> |
| In reply to | #1625576 |
On 18/04/17 01:48 PM, Jason Gunthorpe wrote: > I think this is why progress on this keeps getting stuck - every > solution is a lot of work. Yup! There's also a ton of work just to get the iomem safety issues addressed. Let alone the dma mapping issues. > You could try to do a dummy mapping / create a MR early on to detect > this. Ok, that could be a workable solution. > FWIW, I wonder if from a RDMA perspective we have another > problem.. Should we allow P2P memory to be used with the local DMA > lkey? There are potential designs around virtualization that would not > allow that. Should we mandate that P2P memory be in its own MR? I can't say I understand these issues... Logan
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-18 21:50 +0200 |
| Message-ID | <txHiq-1Gh-21@gated-at.bofh.it> |
| In reply to | #1625569 |
On Tue, Apr 18, 2017 at 12:35 PM, Logan Gunthorpe <logang@deltatee.com> wrote: > > > On 18/04/17 01:01 PM, Jason Gunthorpe wrote: >> Ultimately every dma_ops will need special code to support P2P with >> the special hardware that ops is controlling, so it makes some sense >> to start by pushing the check down there in the first place. This >> advice is partially motivated by how dma_map_sg is just a small >> wrapper around the function pointer call... > > Yes, I noticed this problem too and that makes sense. It just means > every dma_ops will probably need to be modified to either support p2p > pages or fail on them. Though, the only real difficulty there is that it > will be a lot of work. I don't think you need to go touch all dma_ops, I think you can just arrange for devices that are going to do dma to get redirected to a p2p aware provider of operations that overrides the system default dma_ops. I.e. just touch get_dma_ops().
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2017-04-18 22:30 +0200 |
| Message-ID | <txHV8-29d-17@gated-at.bofh.it> |
| In reply to | #1625579 |
> On Tue, Apr 18, 2017 at 12:35 PM, Logan Gunthorpe <logang@deltatee.com> > wrote: > > > > > > On 18/04/17 01:01 PM, Jason Gunthorpe wrote: > >> Ultimately every dma_ops will need special code to support P2P with > >> the special hardware that ops is controlling, so it makes some sense > >> to start by pushing the check down there in the first place. This > >> advice is partially motivated by how dma_map_sg is just a small > >> wrapper around the function pointer call... > > > > Yes, I noticed this problem too and that makes sense. It just means > > every dma_ops will probably need to be modified to either support p2p > > pages or fail on them. Though, the only real difficulty there is that it > > will be a lot of work. > > I don't think you need to go touch all dma_ops, I think you can just > arrange for devices that are going to do dma to get redirected to a > p2p aware provider of operations that overrides the system default > dma_ops. I.e. just touch get_dma_ops(). This would not work well for everyone, for instance on GPU we usualy have buffer object with a mix of device memory and regular system memory but call dma sg map once for the list. Cheers, Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-18 22:40 +0200 |
| Message-ID | <txI4O-2c8-7@gated-at.bofh.it> |
| In reply to | #1625635 |
On Tue, Apr 18, 2017 at 1:29 PM, Jerome Glisse <jglisse@redhat.com> wrote: >> On Tue, Apr 18, 2017 at 12:35 PM, Logan Gunthorpe <logang@deltatee.com> >> wrote: >> > >> > >> > On 18/04/17 01:01 PM, Jason Gunthorpe wrote: >> >> Ultimately every dma_ops will need special code to support P2P with >> >> the special hardware that ops is controlling, so it makes some sense >> >> to start by pushing the check down there in the first place. This >> >> advice is partially motivated by how dma_map_sg is just a small >> >> wrapper around the function pointer call... >> > >> > Yes, I noticed this problem too and that makes sense. It just means >> > every dma_ops will probably need to be modified to either support p2p >> > pages or fail on them. Though, the only real difficulty there is that it >> > will be a lot of work. >> >> I don't think you need to go touch all dma_ops, I think you can just >> arrange for devices that are going to do dma to get redirected to a >> p2p aware provider of operations that overrides the system default >> dma_ops. I.e. just touch get_dma_ops(). > > This would not work well for everyone, for instance on GPU we usualy > have buffer object with a mix of device memory and regular system > memory but call dma sg map once for the list. > ...and that dma_map goes through get_dma_ops(), so I don't see the conflict?
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2017-04-18 22:50 +0200 |
| Message-ID | <txIet-2fd-11@gated-at.bofh.it> |
| In reply to | #1625643 |
On 18/04/17 02:31 PM, Dan Williams wrote: > On Tue, Apr 18, 2017 at 1:29 PM, Jerome Glisse <jglisse@redhat.com> wrote: >>> On Tue, Apr 18, 2017 at 12:35 PM, Logan Gunthorpe <logang@deltatee.com> >>> wrote: >>>> >>>> >>>> On 18/04/17 01:01 PM, Jason Gunthorpe wrote: >>>>> Ultimately every dma_ops will need special code to support P2P with >>>>> the special hardware that ops is controlling, so it makes some sense >>>>> to start by pushing the check down there in the first place. This >>>>> advice is partially motivated by how dma_map_sg is just a small >>>>> wrapper around the function pointer call... >>>> >>>> Yes, I noticed this problem too and that makes sense. It just means >>>> every dma_ops will probably need to be modified to either support p2p >>>> pages or fail on them. Though, the only real difficulty there is that it >>>> will be a lot of work. >>> >>> I don't think you need to go touch all dma_ops, I think you can just >>> arrange for devices that are going to do dma to get redirected to a >>> p2p aware provider of operations that overrides the system default >>> dma_ops. I.e. just touch get_dma_ops(). >> >> This would not work well for everyone, for instance on GPU we usualy >> have buffer object with a mix of device memory and regular system >> memory but call dma sg map once for the list. >> > > ...and that dma_map goes through get_dma_ops(), so I don't see the conflict? The main conflict is in dma_map_sg which only does get_dma_ops once but the sg may contain memory of different types. Logan
[toc] | [prev] | [next] | [standalone]
| From | Benjamin Herrenschmidt <benh@kernel.crashing.org> |
|---|---|
| Date | 2017-04-19 03:20 +0200 |
| Message-ID | <txMrL-5bR-7@gated-at.bofh.it> |
| In reply to | #1625645 |
On Tue, 2017-04-18 at 14:48 -0600, Logan Gunthorpe wrote: > > ...and that dma_map goes through get_dma_ops(), so I don't see the conflict? > > The main conflict is in dma_map_sg which only does get_dma_ops once but > the sg may contain memory of different types. We can handle that in our "overriden" dma ops. It's a bit tricky but it *could* break it down into segments and forward portions back to the original dma ops. Ben.
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2017-04-18 23:10 +0200 |
| Message-ID | <txIxP-2AU-17@gated-at.bofh.it> |
| In reply to | #1625579 |
On Tue, Apr 18, 2017 at 12:48:35PM -0700, Dan Williams wrote:
> > Yes, I noticed this problem too and that makes sense. It just means
> > every dma_ops will probably need to be modified to either support p2p
> > pages or fail on them. Though, the only real difficulty there is that it
> > will be a lot of work.
>
> I don't think you need to go touch all dma_ops, I think you can just
> arrange for devices that are going to do dma to get redirected to a
> p2p aware provider of operations that overrides the system default
> dma_ops. I.e. just touch get_dma_ops().
I don't follow, when does get_dma_ops() return a p2p aware provider?
It has no way to know if the DMA is going to involve p2p, get_dma_ops
is called with the device initiating the DMA.
So you'd always return the P2P shim on a system that has registered
P2P memory?
Even so, how does this shim work? dma_ops are not really intended to
be stacked. How would we make unmap work, for instance? What happens
when the underlying iommu dma ops actually natively understands p2p
and doesn't want the shim?
I think this opens an even bigger can of worms..
Lets find a strategy to safely push this into dma_ops.
What about something more incremental like this instead:
- dma_ops will set map_sg_p2p == map_sg when they are updated to
support p2p, otherwise DMA on P2P pages will fail for those ops.
- When all ops support p2p we remove the if and ops->map_sg then
just call map_sg_p2p
- For now the scatterlist maintains a bit when pages are added indicating if
p2p memory might be present in the list.
- Unmap for p2p and non-p2p is the same, the underlying ops driver has
to make it work.
diff --git a/include/linux/dma-mapping.h b/include/linux/dma-mapping.h
index 0977317c6835c2..505ed7d502053d 100644
--- a/include/linux/dma-mapping.h
+++ b/include/linux/dma-mapping.h
@@ -103,6 +103,9 @@ struct dma_map_ops {
int (*map_sg)(struct device *dev, struct scatterlist *sg,
int nents, enum dma_data_direction dir,
unsigned long attrs);
+ int (*map_sg_p2p)(struct device *dev, struct scatterlist *sg,
+ int nents, enum dma_data_direction dir,
+ unsigned long attrs);
void (*unmap_sg)(struct device *dev,
struct scatterlist *sg, int nents,
enum dma_data_direction dir,
@@ -244,7 +247,15 @@ static inline int dma_map_sg_attrs(struct device *dev, struct scatterlist *sg,
for_each_sg(sg, s, nents, i)
kmemcheck_mark_initialized(sg_virt(s), s->length);
BUG_ON(!valid_dma_direction(dir));
- ents = ops->map_sg(dev, sg, nents, dir, attrs);
+
+ if (sg_has_p2p(sg)) {
+ if (ops->map_sg_p2p)
+ ents = ops->map_sg_p2p(dev, sg, nents, dir, attrs);
+ else
+ return 0;
+ } else
+ ents = ops->map_sg(dev, sg, nents, dir, attrs);
+
BUG_ON(ents < 0);
debug_dma_map_sg(dev, sg, nents, ents, dir);
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-18 23:20 +0200 |
| Message-ID | <txIHv-2Es-1@gated-at.bofh.it> |
| In reply to | #1625654 |
On Tue, Apr 18, 2017 at 2:03 PM, Jason Gunthorpe <jgunthorpe@obsidianresearch.com> wrote: > On Tue, Apr 18, 2017 at 12:48:35PM -0700, Dan Williams wrote: > >> > Yes, I noticed this problem too and that makes sense. It just means >> > every dma_ops will probably need to be modified to either support p2p >> > pages or fail on them. Though, the only real difficulty there is that it >> > will be a lot of work. >> >> I don't think you need to go touch all dma_ops, I think you can just >> arrange for devices that are going to do dma to get redirected to a >> p2p aware provider of operations that overrides the system default >> dma_ops. I.e. just touch get_dma_ops(). > > I don't follow, when does get_dma_ops() return a p2p aware provider? > It has no way to know if the DMA is going to involve p2p, get_dma_ops > is called with the device initiating the DMA. > > So you'd always return the P2P shim on a system that has registered > P2P memory? > > Even so, how does this shim work? dma_ops are not really intended to > be stacked. How would we make unmap work, for instance? What happens > when the underlying iommu dma ops actually natively understands p2p > and doesn't want the shim? > > I think this opens an even bigger can of worms.. No, I don't think it does. You'd only shim when the target page is backed by a device, not host memory, and you can figure this out by a is_zone_device_page()-style lookup.
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2017-04-18 23:30 +0200 |
| Message-ID | <txIRb-2HH-5@gated-at.bofh.it> |
| In reply to | #1625657 |
On Tue, Apr 18, 2017 at 02:11:33PM -0700, Dan Williams wrote: > > I think this opens an even bigger can of worms.. > > No, I don't think it does. You'd only shim when the target page is > backed by a device, not host memory, and you can figure this out by a > is_zone_device_page()-style lookup. The bigger can of worms is how do you meaningfully stack dma_ops. What does the p2p provider do when it detects a p2p page? Jason
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-18 23:40 +0200 |
| Message-ID | <txJ0S-2KJ-13@gated-at.bofh.it> |
| In reply to | #1625659 |
On Tue, Apr 18, 2017 at 2:22 PM, Jason Gunthorpe <jgunthorpe@obsidianresearch.com> wrote: > On Tue, Apr 18, 2017 at 02:11:33PM -0700, Dan Williams wrote: >> > I think this opens an even bigger can of worms.. >> >> No, I don't think it does. You'd only shim when the target page is >> backed by a device, not host memory, and you can figure this out by a >> is_zone_device_page()-style lookup. > > The bigger can of worms is how do you meaningfully stack dma_ops. This goes back to my original comment to make this capability a function of the pci bridge itself. The kernel has an implementation of a dynamically created bridge device that injects its own dma_ops for the devices behind the bridge. See vmd_setup_dma_ops() in drivers/pci/host/vmd.c. > What does the p2p provider do when it detects a p2p page? Check to see if the arch requires this offset translation that Ben brought up and if not provide the physical address as the patches are doing now.
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2017-04-19 00:20 +0200 |
| Message-ID | <txJDz-3bW-1@gated-at.bofh.it> |
| In reply to | #1625663 |
On 18/04/17 03:36 PM, Dan Williams wrote: > On Tue, Apr 18, 2017 at 2:22 PM, Jason Gunthorpe > <jgunthorpe@obsidianresearch.com> wrote: >> On Tue, Apr 18, 2017 at 02:11:33PM -0700, Dan Williams wrote: >>>> I think this opens an even bigger can of worms.. >>> >>> No, I don't think it does. You'd only shim when the target page is >>> backed by a device, not host memory, and you can figure this out by a >>> is_zone_device_page()-style lookup. >> >> The bigger can of worms is how do you meaningfully stack dma_ops. > > This goes back to my original comment to make this capability a > function of the pci bridge itself. The kernel has an implementation of > a dynamically created bridge device that injects its own dma_ops for > the devices behind the bridge. See vmd_setup_dma_ops() in > drivers/pci/host/vmd.c. Well the issue I think Jason is pointing out is that the ops don't stack. The map_* function in the injected dma_ops needs to be able to call the original map_* for any page that is not p2p memory. This is especially annoying in the map_sg function which may need to call a different op based on the contents of the sgl. (And please correct me if I'm not seeing how this can be done in the vmd example.) Also, what happens if p2p pages end up getting passed to a device that doesn't have the injected dma_ops? However, the concept of replacing the dma_ops for all devices behind a supporting bridge is interesting and may be a good piece of the final solution. Logan
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-19 00:30 +0200 |
| Message-ID | <txJNf-3f2-9@gated-at.bofh.it> |
| In reply to | #1625670 |
On Tue, Apr 18, 2017 at 3:15 PM, Logan Gunthorpe <logang@deltatee.com> wrote: > > > On 18/04/17 03:36 PM, Dan Williams wrote: >> On Tue, Apr 18, 2017 at 2:22 PM, Jason Gunthorpe >> <jgunthorpe@obsidianresearch.com> wrote: >>> On Tue, Apr 18, 2017 at 02:11:33PM -0700, Dan Williams wrote: >>>>> I think this opens an even bigger can of worms.. >>>> >>>> No, I don't think it does. You'd only shim when the target page is >>>> backed by a device, not host memory, and you can figure this out by a >>>> is_zone_device_page()-style lookup. >>> >>> The bigger can of worms is how do you meaningfully stack dma_ops. >> >> This goes back to my original comment to make this capability a >> function of the pci bridge itself. The kernel has an implementation of >> a dynamically created bridge device that injects its own dma_ops for >> the devices behind the bridge. See vmd_setup_dma_ops() in >> drivers/pci/host/vmd.c. > > Well the issue I think Jason is pointing out is that the ops don't > stack. The map_* function in the injected dma_ops needs to be able to > call the original map_* for any page that is not p2p memory. This is > especially annoying in the map_sg function which may need to call a > different op based on the contents of the sgl. (And please correct me if > I'm not seeing how this can be done in the vmd example.) Unlike the pci bus address offset case which I think is fundamental to support since shipping archs do this today, I think it is ok to say p2p is restricted to a single sgl that gets to talk to host memory or a single device. That said, what's wrong with a p2p aware map_sg implementation calling up to the host memory map_sg implementation on a per sgl basis? > Also, what happens if p2p pages end up getting passed to a device that > doesn't have the injected dma_ops? This goes back to limiting p2p to a single pci host bridge. If the p2p capability is coordinated with the bridge rather than between the individual devices then we have a central point to catch this case. ...of course this is all hand wavy until someone writes the code and proves otherwise. > However, the concept of replacing the dma_ops for all devices behind a > supporting bridge is interesting and may be a good piece of the final > solution. It's at least a proof point for injecting special behavior for devices behind a (virtual) pci bridge without needing to go touch a bunch of drivers.
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2017-04-19 00:50 +0200 |
| Message-ID | <txK6B-3m1-1@gated-at.bofh.it> |
| In reply to | #1625672 |
On Tue, Apr 18, 2017 at 03:28:17PM -0700, Dan Williams wrote: > Unlike the pci bus address offset case which I think is fundamental to > support since shipping archs do this toda But we can support this by modifying those arch's unique dma_ops directly. Eg as I explained, my p2p_same_segment_map_page() helper concept would do the offset adjustment for same-segement DMA. If PPC calls that in their IOMMU drivers then they will have proper support for this basic p2p, and the right framework to move on to more advanced cases of p2p. This really seems like much less trouble than trying to wrapper all the arch's dma ops, and doesn't have the wonky restrictions. > I think it is ok to say p2p is restricted to a single sgl that gets > to talk to host memory or a single device. RDMA and GPU would be sad with this restriction... > That said, what's wrong with a p2p aware map_sg implementation > calling up to the host memory map_sg implementation on a per sgl > basis? Setting up the iommu is fairly expensive, so getting rid of the batching would kill performance.. Jason
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-19 01:00 +0200 |
| Message-ID | <txKgi-3qF-15@gated-at.bofh.it> |
| In reply to | #1625679 |
On Tue, Apr 18, 2017 at 3:42 PM, Jason Gunthorpe <jgunthorpe@obsidianresearch.com> wrote: > On Tue, Apr 18, 2017 at 03:28:17PM -0700, Dan Williams wrote: > >> Unlike the pci bus address offset case which I think is fundamental to >> support since shipping archs do this toda > > But we can support this by modifying those arch's unique dma_ops > directly. > > Eg as I explained, my p2p_same_segment_map_page() helper concept would > do the offset adjustment for same-segement DMA. > > If PPC calls that in their IOMMU drivers then they will have proper > support for this basic p2p, and the right framework to move on to more > advanced cases of p2p. > > This really seems like much less trouble than trying to wrapper all > the arch's dma ops, and doesn't have the wonky restrictions. I don't think the root bus iommu drivers have any business knowing or caring about dma happening between devices lower in the hierarchy. >> I think it is ok to say p2p is restricted to a single sgl that gets >> to talk to host memory or a single device. > > RDMA and GPU would be sad with this restriction... > >> That said, what's wrong with a p2p aware map_sg implementation >> calling up to the host memory map_sg implementation on a per sgl >> basis? > > Setting up the iommu is fairly expensive, so getting rid of the > batching would kill performance.. When we're crossing device and host memory boundaries how much batching is possible? As far as I can see you'll always be splitting the sgl on these dma mapping boundaries.
[toc] | [prev] | [next] | [standalone]
| From | Jason Gunthorpe <jgunthorpe@obsidianresearch.com> |
|---|---|
| Date | 2017-04-19 01:30 +0200 |
| Message-ID | <txKJk-3Uf-11@gated-at.bofh.it> |
| In reply to | #1625689 |
On Tue, Apr 18, 2017 at 03:51:27PM -0700, Dan Williams wrote: > > This really seems like much less trouble than trying to wrapper all > > the arch's dma ops, and doesn't have the wonky restrictions. > > I don't think the root bus iommu drivers have any business knowing or > caring about dma happening between devices lower in the hierarchy. Maybe not, but performance requires some odd choices in this code.. :( > > Setting up the iommu is fairly expensive, so getting rid of the > > batching would kill performance.. > > When we're crossing device and host memory boundaries how much > batching is possible? As far as I can see you'll always be splitting > the sgl on these dma mapping boundaries. Splitting the sgl is different from iommu batching. As an example, an O_DIRECT write of 1 MB with a single 4K P2P page in the middle. The optimum behavior is to allocate a 1MB-4K iommu range and fill it with the CPU memory. Then return a SGL with three entires, two pointing into the range and one to the p2p. It is creating each range which tends to be expensive, so creating two ranges (or worse, if every SGL created a range it would be 255) is very undesired. Jason
[toc] | [prev] | [next] | [standalone]
Page 2 of 4 — ← Prev page 1 [2] 3 4 Next page →
Back to top | Article view | linux.kernel
csiph-web