Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1621832 > unrolled thread
| Started by | Benjamin Herrenschmidt <benh@kernel.crashing.org> |
|---|---|
| First post | 2017-04-12 08:30 +0200 |
| Last post | 2017-04-16 07:40 +0200 |
| Articles | 6 on this page of 26 — 7 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-12 08:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-12 19:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-13 00:00 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-13 23:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-14 00:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Bjorn Helgaas <helgaas@kernel.org> - 2017-04-14 01:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Jason Gunthorpe <jgunthorpe@obsidianresearch.com> - 2017-04-14 06:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-14 06:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@au1.ibm.com> - 2017-04-14 13:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-14 15:00 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-14 13:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-14 19:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Bjorn Helgaas <helgaas@kernel.org> - 2017-04-14 21:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-15 00:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-15 19:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-16 00:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-16 05:10 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-16 06:50 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Dan Williams <dan.j.williams@intel.com> - 2017-04-16 18:00 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-16 18:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-17 01:00 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Knut Omang <knut.omang@oracle.com> - 2017-04-24 09:40 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-24 18:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-17 00:30 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Benjamin Herrenschmidt <benh@kernel.crashing.org> - 2017-04-16 00:20 +0200
Re: [RFC 0/8] Copy Offload with Peer-to-Peer PCI Memory Logan Gunthorpe <logang@deltatee.com> - 2017-04-16 07:40 +0200
Page 2 of 2 — ← Prev page 1 [2]
| From | Benjamin Herrenschmidt <benh@kernel.crashing.org> |
|---|---|
| Date | 2017-04-17 01:00 +0200 |
| Message-ID | <tx1jb-12F-3@gated-at.bofh.it> |
| In reply to | #1624409 |
On Sun, 2017-04-16 at 10:34 -0600, Logan Gunthorpe wrote: > > On 16/04/17 09:53 AM, Dan Williams wrote: > > ZONE_DEVICE allows you to redirect via get_dev_pagemap() to retrieve > > context about the physical address in question. I'm thinking you can > > hang bus address translation data off of that structure. This seems > > vaguely similar to what HMM is doing. > > Thanks! I didn't realize you had the infrastructure to look up a device > from a pfn/page. That would really come in handy for us. It does indeed. I won't be able to play with that much for a few weeks (see my other email) so if you're going to tackle this while I'm away, can you work with Jerome to make sure you don't conflict with HMM ? I really want a way for HMM to be able to layout struct pages over the GPU BARs rather than in "allocated free space" for the case where the BAR is big enough to cover all of the GPU memory. In general, I'd like a simple & generic way for any driver to ask the core to layout DMA'ble struct pages over BAR space. I an not convinced this requires a "p2mem device" to be created on top of this though but that's a different discussion. Of course the actual ability to perform the DMA mapping will be subject to various restrictions that will have to be implemented in the actual "dma_ops override" backend. We can have generic code to handle the case where devices reside on the same domain, which can deal with switch configuration etc... we will need to have iommu specific code to handle the case going through the fabric. Virtualization is a separate can of worms due to how qemu completely fakes the MMIO space, we can look into that later. Cheers, Ben.
[toc] | [prev] | [next] | [standalone]
| From | Knut Omang <knut.omang@oracle.com> |
|---|---|
| Date | 2017-04-24 09:40 +0200 |
| Message-ID | <tzGLg-45f-9@gated-at.bofh.it> |
| In reply to | #1624483 |
On Mon, 2017-04-17 at 08:31 +1000, Benjamin Herrenschmidt wrote: > On Sun, 2017-04-16 at 10:34 -0600, Logan Gunthorpe wrote: > > > > On 16/04/17 09:53 AM, Dan Williams wrote: > > > ZONE_DEVICE allows you to redirect via get_dev_pagemap() to retrieve > > > context about the physical address in question. I'm thinking you can > > > hang bus address translation data off of that structure. This seems > > > vaguely similar to what HMM is doing. > > > > Thanks! I didn't realize you had the infrastructure to look up a device > > from a pfn/page. That would really come in handy for us. > > It does indeed. I won't be able to play with that much for a few weeks > (see my other email) so if you're going to tackle this while I'm away, > can you work with Jerome to make sure you don't conflict with HMM ? > > I really want a way for HMM to be able to layout struct pages over the > GPU BARs rather than in "allocated free space" for the case where the > BAR is big enough to cover all of the GPU memory. > > In general, I'd like a simple & generic way for any driver to ask the > core to layout DMA'ble struct pages over BAR space. I an not convinced > this requires a "p2mem device" to be created on top of this though but > that's a different discussion. > > Of course the actual ability to perform the DMA mapping will be subject > to various restrictions that will have to be implemented in the actual > "dma_ops override" backend. We can have generic code to handle the case > where devices reside on the same domain, which can deal with switch > configuration etc... we will need to have iommu specific code to handle > the case going through the fabric. > > Virtualization is a separate can of worms due to how qemu completely > fakes the MMIO space, we can look into that later. My first reflex when reading this thread was to think that this whole domain lends it self excellently to testing via Qemu. Could it be that doing this in the opposite direction might be a safer approach in the long run even though (significant) more work up-front? Eg. start by fixing/providing/documenting suitable model(s) for testing this in Qemu, then implement the patch set based on those models? Thanks, Knut > > Cheers, > Ben. > > -- > To unsubscribe from this list: send the line "unsubscribe linux-rdma" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2017-04-24 18:20 +0200 |
| Message-ID | <tzOSu-O4-3@gated-at.bofh.it> |
| In reply to | #1629270 |
On 24/04/17 01:36 AM, Knut Omang wrote: > My first reflex when reading this thread was to think that this whole domain > lends it self excellently to testing via Qemu. Could it be that doing this in > the opposite direction might be a safer approach in the long run even though > (significant) more work up-front? That's an interesting idea. We did do some very limited testing on qemu with one iteration of our work. However, it's difficult because there is no support for any RDMA devices which are a part of our primary use case. I also imagine it would be quite difficult to develop those models given the array of hardware that needs to be supported and the deep functional knowledge required to figure out appropriate restrictions. Logan
[toc] | [prev] | [next] | [standalone]
| From | Benjamin Herrenschmidt <benh@kernel.crashing.org> |
|---|---|
| Date | 2017-04-17 00:30 +0200 |
| Message-ID | <tx0Qa-RT-11@gated-at.bofh.it> |
| In reply to | #1624404 |
On Sun, 2017-04-16 at 08:53 -0700, Dan Williams wrote: > > Just thinking out loud ... I don't have a firm idea or a design. But > > peer to peer is definitely a problem we need to tackle generically, the > > demand for it keeps coming up. > > ZONE_DEVICE allows you to redirect via get_dev_pagemap() to retrieve > context about the physical address in question. I'm thinking you can > hang bus address translation data off of that structure. This seems > vaguely similar to what HMM is doing. Ok, that's interesting. That would be a way to handle the "lookup" I was mentioning in the email I sent a few minutes ago. We would probably need to put some "structure" to that context. I'm very short on time to look into the details of this for at least a month (I'm taking about 3 weeks off for personal reasons next week), but I'm happy to dive more into this when I'm back and sort out with Jerome how to make it all co-habitate nicely with HMM. Cheers, Ben.
[toc] | [prev] | [next] | [standalone]
| From | Benjamin Herrenschmidt <benh@kernel.crashing.org> |
|---|---|
| Date | 2017-04-16 00:20 +0200 |
| Message-ID | <twEcV-3Xy-1@gated-at.bofh.it> |
| In reply to | #1624122 |
On Sat, 2017-04-15 at 11:41 -0600, Logan Gunthorpe wrote: > Thanks, Benjamin, for the summary of some of the issues. > > On 14/04/17 04:07 PM, Benjamin Herrenschmidt wrote > > So I assume the p2p code provides a way to address that too via special > > dma_ops ? Or wrappers ? > > Not at this time. We will probably need a way to ensure the iommus do > not attempt to remap these addresses. You can't. If the iommu is on, everything is remapped. Or do you mean to have dma_map_* not do a remapping ? That's the problem again, same as before, for that to work, the dma_map_* ops would have to do something special that depends on *both* the source and target device. The current DMA infrastructure doesn't have anything like that. It's a rather fundamental issue to your design that you need to address. The dma_ops today are architecture specific and have no way to differenciate between normal and those special P2P DMA pages. > Though if it does, I'd expect > everything would still work you just wouldn't get the performance or > traffic flow you are looking for. We've been testing with the software > iommu which doesn't have this problem. So first, no, it's more than "you wouldn't get the performance". On some systems it may also just not work. Also what do you mean by "the SW iommu doesn't have this problem" ? It catches the fact that addresses don't point to RAM and maps differently ? > > The problem is that the latter while seemingly easier, is also slower > > and not supported by all platforms and architectures (for example, > > POWER currently won't allow it, or rather only allows a store-only > > subset of it under special circumstances). > > Yes, I think situations where we have to cross host bridges will remain > unsupported by this work for a long time. There are two many cases where > it just doesn't work or it performs too poorly to be useful. And the situation where you don't cross bridges is the one where you need to also take into account the offsets. *both* cases mean that you need somewhat to intervene at the dma_ops level to handle this. Which means having a way to identify your special struct pages or PFNs to allow the arch to add a special case to the dma_ops. > > I don't fully understand how p2pmem "solves" that by creating struct > > pages. The offset problem is one issue. But there's the iommu issue as > > well, the driver cannot just use the normal dma_map ops. > > We are not using a proper iommu and we are dealing with systems that > have zero offset. This case is also easily supported. I expect fixing > the iommus to not map these addresses would also be reasonably achievable. So you are designing something that is built from scratch to only work on a specific limited category of systems and is also incompatible with virtualization. This is an interesting experiement to look at I suppose, but if you ever want this upstream I would like at least for you to develop a strategy to support the wider case, if not an actual implementation. Cheers, Ben.
[toc] | [prev] | [next] | [standalone]
| From | Logan Gunthorpe <logang@deltatee.com> |
|---|---|
| Date | 2017-04-16 07:40 +0200 |
| Message-ID | <twL4K-8hc-3@gated-at.bofh.it> |
| In reply to | #1624139 |
On 15/04/17 04:17 PM, Benjamin Herrenschmidt wrote: > You can't. If the iommu is on, everything is remapped. Or do you mean > to have dma_map_* not do a remapping ? Well, yes, you'd have to change the code so that iomem pages do not get remapped and the raw BAR address is passed to the DMA engine. I said specifically we haven't done this at this time but it really doesn't seem like an unsolvable problem. It is something we will need to address before a proper patch set is posted though. > That's the problem again, same as before, for that to work, the > dma_map_* ops would have to do something special that depends on *both* > the source and target device. No, I don't think you have to do things different based on the source. Have the p2pmem device layer restrict allocating p2pmem based on the devices in use (similar to how the RFC code works now) and when the dma mapping code sees iomem pages it just needs to leave the address alone so it's used directly by the dma in question. It's much better to make the decision on which memory to use when you allocate it. If you wait until you map it, it would be a pain to fall back to system memory if it doesn't look like it will work. So, if when you allocate it, you know everything will work you just need the dma mapping layer to stay out of the way. > The dma_ops today are architecture specific and have no way to > differenciate between normal and those special P2P DMA pages. Correct, unless Dan's idea works (which will need some investigation), we'd need a flag in struct page or some other similar method to determine that these are special iomem pages. >> Though if it does, I'd expect >> everything would still work you just wouldn't get the performance or >> traffic flow you are looking for. We've been testing with the software >> iommu which doesn't have this problem. > > So first, no, it's more than "you wouldn't get the performance". On > some systems it may also just not work. Also what do you mean by "the > SW iommu doesn't have this problem" ? It catches the fact that > addresses don't point to RAM and maps differently ? I haven't tested it but I can't imagine why an iommu would not correctly map the memory in the bar. But that's _way_ beside the point. We _really_ want to avoid that situation anyway. If the iommu maps the memory it defeats what we are trying to accomplish. I believe the sotfware iommu only uses bounce buffers if the DMA engine in use cannot address the memory. So in most cases, with modern hardware, it just passes the BAR's address to the DMA engine and everything works. The code posted in the RFC does in fact work without needing to do any of this fussing. >>> The problem is that the latter while seemingly easier, is also slower >>> and not supported by all platforms and architectures (for example, >>> POWER currently won't allow it, or rather only allows a store-only >>> subset of it under special circumstances). >> >> Yes, I think situations where we have to cross host bridges will remain >> unsupported by this work for a long time. There are two many cases where >> it just doesn't work or it performs too poorly to be useful. > > And the situation where you don't cross bridges is the one where you > need to also take into account the offsets. I think for the first incarnation we will just not support systems that have offsets. This makes things much easier and still supports all the use cases we are interested in. > So you are designing something that is built from scratch to only work > on a specific limited category of systems and is also incompatible with > virtualization. Yes, we are starting with support for specific use cases. Almost all technology starts that way. Dax has been in the kernel for years and only recently has someone submitted patches for it to support pmem on powerpc. This is not unusual. If you had forced the pmem developers to support all architectures in existence before allowing them upstream they couldn't possibly be as far as they are today. Virtualization specifically would be a _lot_ more difficult than simply supporting offsets. The actual topology of the bus will probably be lost on the guest OS and it would therefor have a difficult time figuring out when it's acceptable to use p2pmem. I also have a difficult time seeing a use case for it and thus I have a hard time with the argument that we can't support use cases that do want it because use cases that don't want it (perhaps yet) won't work. > This is an interesting experiement to look at I suppose, but if you > ever want this upstream I would like at least for you to develop a > strategy to support the wider case, if not an actual implementation. I think there are plenty of avenues forward to support offsets, etc. It's just work. Nothing we'd be proposing would be incompatible with it. We just don't want to have to do it all upfront especially when no one really knows how well various architecture's hardware supports this or if anyone even wants to run it on systems such as those. (Keep in mind this is a pretty specific optimization that mostly helps systems designed in specific ways -- not a general "everybody gets faster" type situation.) Get the cases working we know will work, can easily support and people actually want. Then expand it to support others as people come around with hardware to test and use cases for it. Logan
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web