Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1652907 > unrolled thread
| Started by | Michal Hocko <mhocko@kernel.org> |
|---|---|
| First post | 2017-05-30 09:50 +0200 |
| Last post | 2017-05-31 14:30 +0200 |
| Articles | 10 on this page of 30 — 5 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-05-30 09:50 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Mike Rapoport <rppt@linux.vnet.ibm.com> - 2017-05-30 12:20 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-05-30 12:50 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Andrea Arcangeli <aarcange@redhat.com> - 2017-05-30 16:10 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-05-30 16:50 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-05-30 17:00 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Andrea Arcangeli <aarcange@redhat.com> - 2017-05-30 18:10 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Vlastimil Babka <vbabka@suse.cz> - 2017-05-31 08:40 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-05-31 10:30 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Mike Rapoport <rppt@linux.vnet.ibm.com> - 2017-05-31 11:30 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-05-31 12:30 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-05-31 12:30 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Mike Rapoport <rppt@linux.vnet.ibm.com> - 2017-06-01 13:10 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-06-01 14:30 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Andrea Arcangeli <aarcange@redhat.com> - 2017-05-30 17:50 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-05-31 14:10 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Mike Rapoprt <rppt@linux.vnet.ibm.com> - 2017-05-31 14:40 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Andrea Arcangeli <aarcange@redhat.com> - 2017-05-31 16:20 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-05-31 16:40 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Andrea Arcangeli <aarcange@redhat.com> - 2017-05-31 17:50 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Mike Rapoport <rppt@linux.vnet.ibm.com> - 2017-06-01 09:00 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-05-31 16:20 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Mike Rapoport <rppt@linux.vnet.ibm.com> - 2017-06-01 09:00 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-06-01 10:10 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Mike Rapoport <rppt@linux.vnet.ibm.com> - 2017-06-01 10:40 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Andrea Arcangeli <aarcange@redhat.com> - 2017-06-01 15:50 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Mike Rapoport <rppt@linux.vnet.ibm.com> - 2017-06-02 11:20 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Mike Rapoport <rppt@linux.vnet.ibm.com> - 2017-05-31 11:10 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Michal Hocko <mhocko@kernel.org> - 2017-05-31 14:10 +0200
Re: [PATCH] mm: introduce MADV_CLR_HUGEPAGE Mike Rapoprt <rppt@linux.vnet.ibm.com> - 2017-05-31 14:30 +0200
Page 2 of 2 — ← Prev page 1 [2]
| From | Mike Rapoport <rppt@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-06-01 09:00 +0200 |
| Message-ID | <tNsfo-7wq-7@gated-at.bofh.it> |
| In reply to | #1654298 |
On Wed, May 31, 2017 at 04:18:09PM +0200, Andrea Arcangeli wrote: > On Wed, May 31, 2017 at 03:39:22PM +0300, Mike Rapoport wrote: > > For the CRIU usecase, disabling THP for a while and re-enabling it > > back will do the trick, provided VMAs flags are not affected, like > > in the patch you've sent. Moreover, we may even get away with > > Are you going to check uname -r to know when the kABI changed in your > favor (so CRIU cannot ever work with enterprise backports unless you > expand the uname -r coverage), or how do you know the patch is > applied? CRIU does not rely on uname -r. We have code that checks what kernel features we can actually use. For instance, we use UFFDIO_API to see if we can do post-copy at all. > Optimistically assuming people is going to run new CRIU code only on > new kernels looks very risky, it would leads to silent random memory > corruption, so I doubt you can get away without a uname -r check. > > This is fairly simple change too, its main cons is that it adds a > branch to the page fault fast path, the old behavior of the prctl and > the new madvise were both zero cost. > > Still if the prctl is preferred despite the added branch, to avoid > uname -r clashes, to me it sounds better to add a new prctl ID and > keep the old one too. The old one could be implemented the same way as > the new one if you want to save a few bytes of .text. But the old one > should probably do a printk_once to print a deprecation warning so the > old ID with weaker (zero runtime cost) semantics can be removed later. >
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-05-31 16:20 +0200 |
| Message-ID | <tNcDE-5Wy-27@gated-at.bofh.it> |
| In reply to | #1654183 |
On Wed 31-05-17 15:39:22, Mike Rapoprt wrote: > > > On May 31, 2017 3:08:22 PM GMT+03:00, Michal Hocko <mhocko@kernel.org> wrote: [...] > > From what Mike said a global disable THP for the whole process > >while the post-copy is in progress is a better solution anyway. > > For the CRIU usecase, disabling THP for a while and re-enabling > it back will do the trick, provided VMAs flags are not affected, > like in the patch you've sent. Moreover, we may even get away with > ioctl(UFFDIO_COPY) if it's overhead shows to be negligible. Still, > I believe that MADV_RESET_HUGEPAGE (or some better named) command has > the value on its own. I would prefer if we could go the prctl if possible and add a new MADV_RESET_HUGEPAGE if there is really a usecase for it. -- Michal Hocko SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Mike Rapoport <rppt@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-06-01 09:00 +0200 |
| Message-ID | <tNsfn-7wq-5@gated-at.bofh.it> |
| In reply to | #1653289 |
On Tue, May 30, 2017 at 04:39:41PM +0200, Michal Hocko wrote:
> On Tue 30-05-17 16:04:56, Andrea Arcangeli wrote:
> >
> > UFFDIO_COPY while not being a major slowdown for sure, it's likely
> > measurable at the microbenchmark level because it would add a
> > enter/exit kernel to every 4k memcpy. It's not hard to imagine that as
> > measurable. How that impacts the total precopy time I don't know, it
> > would need to be benchmarked to be sure.
>
> Yes, please!
I've run a simple test (below) that fills 1G of memory either with memcpy
of ioctl(UFFDIO_COPY) in 4K chunks.
The machine I used has two "Intel(R) Xeon(R) CPU E5-2680 0 @ 2.70GHz" and
128G of RAM.
I've averaged elapsed time reported by /usr/bin/time over 100 runs and here
what I've got:
memcpy with THP on: 0.3278 sec
memcpy with THP off: 0.5295 sec
UFFDIO_COPY: 0.44 sec
That said, for the CRIU usecase UFFDIO_COPY seems faster that disabling THP
and then doing memcpy.
--
Sincerely yours,
Mike.
----------------------------------------------------------
{
...
src = mmap(NULL, page_size, PROT_READ | PROT_WRITE,
MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (src == MAP_FAILED)
fprintf(stderr, "map src failed\n"), exit(1);
*((unsigned long *)src) = 1;
if (disable_huge && prctl(PR_SET_THP_DISABLE, 1, 0, 0, 0))
fprintf(stderr, "ptctl failed\n"), exit(1);
dst = mmap(NULL, page_size * nr_pages, PROT_READ | PROT_WRITE,
MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (dst == MAP_FAILED)
fprintf(stderr, "map dst failed\n"), exit(1);
if (use_uffd && userfaultfd_register(dst))
fprintf(stderr, "userfault_register failed\n"), exit(1);
for (i = 0; i < nr_pages; i++) {
char *address = dst + i * page_size;
if (use_uffd) {
struct uffdio_copy uffdio_copy;
uffdio_copy.dst = (unsigned long)address;
uffdio_copy.src = (unsigned long)src;
uffdio_copy.len = page_size;
uffdio_copy.mode = 0;
uffdio_copy.copy = 0;
ret = ioctl(uffd, UFFDIO_COPY, &uffdio_copy);
if (ret)
fprintf(stderr, "copy: %d, %d\n", ret, errno),
exit(1);
} else {
memcpy(address, src, page_size);
}
}
return 0;
}
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-06-01 10:10 +0200 |
| Message-ID | <tNtl8-8tU-15@gated-at.bofh.it> |
| In reply to | #1654857 |
On Thu 01-06-17 09:53:02, Mike Rapoport wrote: > On Tue, May 30, 2017 at 04:39:41PM +0200, Michal Hocko wrote: > > On Tue 30-05-17 16:04:56, Andrea Arcangeli wrote: > > > > > > UFFDIO_COPY while not being a major slowdown for sure, it's likely > > > measurable at the microbenchmark level because it would add a > > > enter/exit kernel to every 4k memcpy. It's not hard to imagine that as > > > measurable. How that impacts the total precopy time I don't know, it > > > would need to be benchmarked to be sure. > > > > Yes, please! > > I've run a simple test (below) that fills 1G of memory either with memcpy > of ioctl(UFFDIO_COPY) in 4K chunks. > The machine I used has two "Intel(R) Xeon(R) CPU E5-2680 0 @ 2.70GHz" and > 128G of RAM. > I've averaged elapsed time reported by /usr/bin/time over 100 runs and here > what I've got: > > memcpy with THP on: 0.3278 sec > memcpy with THP off: 0.5295 sec > UFFDIO_COPY: 0.44 sec I assume that the standard deviation is small? > That said, for the CRIU usecase UFFDIO_COPY seems faster that disabling THP > and then doing memcpy. That is a bit surprising. I didn't think that the userfault syscall (ioctl) can be faster than a regular #PF but considering that __mcopy_atomic bypasses the page fault path and it can be optimized for the anon case suggests that we can save some cycles for each page and so the cumulative savings can be visible. -- Michal Hocko SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Mike Rapoport <rppt@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-06-01 10:40 +0200 |
| Message-ID | <tNtOa-cj-15@gated-at.bofh.it> |
| In reply to | #1654900 |
On Thu, Jun 01, 2017 at 10:09:09AM +0200, Michal Hocko wrote: > On Thu 01-06-17 09:53:02, Mike Rapoport wrote: > > On Tue, May 30, 2017 at 04:39:41PM +0200, Michal Hocko wrote: > > > On Tue 30-05-17 16:04:56, Andrea Arcangeli wrote: > > > > > > > > UFFDIO_COPY while not being a major slowdown for sure, it's likely > > > > measurable at the microbenchmark level because it would add a > > > > enter/exit kernel to every 4k memcpy. It's not hard to imagine that as > > > > measurable. How that impacts the total precopy time I don't know, it > > > > would need to be benchmarked to be sure. > > > > > > Yes, please! > > > > I've run a simple test (below) that fills 1G of memory either with memcpy > > of ioctl(UFFDIO_COPY) in 4K chunks. > > The machine I used has two "Intel(R) Xeon(R) CPU E5-2680 0 @ 2.70GHz" and > > 128G of RAM. > > I've averaged elapsed time reported by /usr/bin/time over 100 runs and here > > what I've got: > > > > memcpy with THP on: 0.3278 sec > > memcpy with THP off: 0.5295 sec > > UFFDIO_COPY: 0.44 sec > > I assume that the standard deviation is small? Yes. > > That said, for the CRIU usecase UFFDIO_COPY seems faster that disabling THP > > and then doing memcpy. > > That is a bit surprising. I didn't think that the userfault syscall > (ioctl) can be faster than a regular #PF but considering that > __mcopy_atomic bypasses the page fault path and it can be optimized for > the anon case suggests that we can save some cycles for each page and so > the cumulative savings can be visible. > > -- > Michal Hocko > SUSE Labs >
[toc] | [prev] | [next] | [standalone]
| From | Andrea Arcangeli <aarcange@redhat.com> |
|---|---|
| Date | 2017-06-01 15:50 +0200 |
| Message-ID | <tNyEa-3cj-11@gated-at.bofh.it> |
| In reply to | #1654900 |
On Thu, Jun 01, 2017 at 10:09:09AM +0200, Michal Hocko wrote: > That is a bit surprising. I didn't think that the userfault syscall > (ioctl) can be faster than a regular #PF but considering that > __mcopy_atomic bypasses the page fault path and it can be optimized for > the anon case suggests that we can save some cycles for each page and so > the cumulative savings can be visible. __mcopy_atomic works not just for anonymous memory, hugetlbfs/shmem are covered too and there are branches to handle those. If you were to run more than one precopy pass UFFDIO_COPY shall become slower than the userland access starting from the second pass. At the light of this if CRIU can only do one single pass of precopy, CRIU is probably better off using UFFDIO_COPY than using prctl or madvise to temporarily turn off THP. With QEMU as opposed we set MADV_HUGEPAGE during precopy on destination to maximize the THP utilization for all those 2M naturally aligned guest regions that aren't re-dirtied in the source, so we're better off without using UFFDIO_COPY in precopy even during the first pass to avoid the enter/kernel for subpages that are written to destination in a already instantiated THP. At least until we teach QEMU to map 2M at once if possible (UFFDIO_COPY would then also require an enhancement, because currently it won't map THP on the fly). Thanks, Andrea
[toc] | [prev] | [next] | [standalone]
| From | Mike Rapoport <rppt@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-06-02 11:20 +0200 |
| Message-ID | <tNQUp-7vd-9@gated-at.bofh.it> |
| In reply to | #1655154 |
On Thu, Jun 01, 2017 at 03:45:22PM +0200, Andrea Arcangeli wrote: > On Thu, Jun 01, 2017 at 10:09:09AM +0200, Michal Hocko wrote: > > That is a bit surprising. I didn't think that the userfault syscall > > (ioctl) can be faster than a regular #PF but considering that > > __mcopy_atomic bypasses the page fault path and it can be optimized for > > the anon case suggests that we can save some cycles for each page and so > > the cumulative savings can be visible. > > __mcopy_atomic works not just for anonymous memory, hugetlbfs/shmem > are covered too and there are branches to handle those. > > If you were to run more than one precopy pass UFFDIO_COPY shall become > slower than the userland access starting from the second pass. > > At the light of this if CRIU can only do one single pass of precopy, > CRIU is probably better off using UFFDIO_COPY than using prctl or > madvise to temporarily turn off THP. CRIU does memory tracking differently from QEMU. Every round of pre-copy in CRIU means we dump the dirty pages into an image file. The restore then chooses what image file to use. Anyway, we fill the memory only once at restore time, hence UFFDIO_COPY would be better than disabling THP. > With QEMU as opposed we set MADV_HUGEPAGE during precopy on > destination to maximize the THP utilization for all those 2M naturally > aligned guest regions that aren't re-dirtied in the source, so we're > better off without using UFFDIO_COPY in precopy even during the first > pass to avoid the enter/kernel for subpages that are written to > destination in a already instantiated THP. At least until we teach > QEMU to map 2M at once if possible (UFFDIO_COPY would then also > require an enhancement, because currently it won't map THP on the > fly). > > Thanks, > Andrea > -- Sincerely yours, Mike.
[toc] | [prev] | [next] | [standalone]
| From | Mike Rapoport <rppt@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-05-31 11:10 +0200 |
| Message-ID | <tN7ND-2Od-13@gated-at.bofh.it> |
| In reply to | #1653105 |
On Tue, May 30, 2017 at 12:39:30PM +0200, Michal Hocko wrote: > On Tue 30-05-17 13:19:22, Mike Rapoport wrote: > > > > But then we'll have to populate these regions with > > > > UFFDIO_COPY which adds quite an overhead. > > > > > > How big is the performance impact? > > > > I don't have the numbers handy, but for each post-copy range it means that > > instead of memcpy() we will use ioctl(UFFDIO_COPY). > > It would be good to measure that though. I will, but I won't expect huge difference here. Anyway, memcpy() will touch still unpopulated pages, so we'll anyway enter/exit kernel. > You are proposing a new user > API and the THP api is quite convoluted already so there better be a > very good reason to add a new API. So far I can only see that it would > be more convinient to add another madvise command and that is rather > insufficient justification IMHO. Well, the most convenient for my use case would be simply disable THP before restore and re-enable it afterwards. And the need to use ioctl(UFFDIO_COPY) is not that less convenient that the proposed madvise command. I've proposed the new madvise command because I firmly believe it is missing. All madvise() commands that set some flag in vma->vm_flags have the counter-command that resets that flag. Except for THP. The THP-related flags can define three states for a VMA, pretty much like VM_SEQ_READ and VM_RAND_READ. And it requires three madvise commands to allow setting any of the desired states, just like with MADV_RANDOM, MADV_SEQUENTIAL and MADV_NORMAL. > Also do you expect somebody else would use new madvise? What would be the > usecase? I can think of an application that wants to keep 4K pages to save physical memory for certain phase, e.g. until these pages are populated with very few data. After the memory usage increases, the application may wish to stop preventing khugepged from merging these pages, but it does not have strong inclination to force use of huge pages. > -- > Michal Hocko > SUSE Labs -- Sincerely yours, Mike.
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-05-31 14:10 +0200 |
| Message-ID | <tNaBR-4zj-33@gated-at.bofh.it> |
| In reply to | #1654025 |
On Wed 31-05-17 12:08:45, Mike Rapoport wrote: > On Tue, May 30, 2017 at 12:39:30PM +0200, Michal Hocko wrote: [...] > > Also do you expect somebody else would use new madvise? What would be the > > usecase? > > I can think of an application that wants to keep 4K pages to save physical > memory for certain phase, e.g. until these pages are populated with very > few data. After the memory usage increases, the application may wish to > stop preventing khugepged from merging these pages, but it does not have > strong inclination to force use of huge pages. Well, is actually anybody going to do that? -- Michal Hocko SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Mike Rapoprt <rppt@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-05-31 14:30 +0200 |
| Message-ID | <tNaVc-4Fy-5@gated-at.bofh.it> |
| In reply to | #1654169 |
On May 31, 2017 3:05:56 PM GMT+03:00, Michal Hocko <mhocko@kernel.org> wrote: >On Wed 31-05-17 12:08:45, Mike Rapoport wrote: >> On Tue, May 30, 2017 at 12:39:30PM +0200, Michal Hocko wrote: >[...] >> > Also do you expect somebody else would use new madvise? What would >be the >> > usecase? >> >> I can think of an application that wants to keep 4K pages to save >physical >> memory for certain phase, e.g. until these pages are populated with >very >> few data. After the memory usage increases, the application may wish >to >> stop preventing khugepged from merging these pages, but it does not >have >> strong inclination to force use of huge pages. > >Well, is actually anybody going to do that? Well, I don't​ know, it's pretty much future telling :) For sure, without the new madvise nobody will be even able to do that.
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web