Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1348973 > unrolled thread
| Started by | Liang Li <liang.z.li@intel.com> |
|---|---|
| First post | 2016-03-03 12:00 +0100 |
| Last post | 2016-03-16 02:30 +0100 |
| Articles | 20 on this page of 59 — 8 participants |
Back to article view | Back to linux.kernel
[RFC qemu 0/4] A PV solution for live migration optimization Liang Li <liang.z.li@intel.com> - 2016-03-03 12:00 +0100
[RFC qemu 3/4] migration: not set migration bitmap in setup stage Liang Li <liang.z.li@intel.com> - 2016-03-03 12:00 +0100
[RFC qemu 1/4] pc: Add code to get the lowmem form PCMachineState Liang Li <liang.z.li@intel.com> - 2016-03-03 12:00 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Roman Kagan <rkagan@virtuozzo.com> - 2016-03-03 15:10 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-04 02:40 +0100
Re: [RFC qemu 0/4] A PV solution for live migration optimization "Dr. David Alan Gilbert" <dgilbert@redhat.com> - 2016-03-03 18:50 +0100
RE: [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-04 03:00 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Roman Kagan <rkagan@virtuozzo.com> - 2016-03-04 09:20 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-04 10:10 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Roman Kagan <rkagan@virtuozzo.com> - 2016-03-04 11:30 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-04 15:30 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Michael S. Tsirkin" <mst@redhat.com> - 2016-03-04 15:50 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-04 16:50 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Michael S. Tsirkin" <mst@redhat.com> - 2016-03-05 21:00 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-07 08:00 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Michael S. Tsirkin" <mst@redhat.com> - 2016-03-07 12:50 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-07 16:10 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Roman Kagan <rkagan@virtuozzo.com> - 2016-03-09 15:30 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-09 16:30 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Michael S. Tsirkin" <mst@redhat.com> - 2016-03-09 16:40 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-10 02:50 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Michael S. Tsirkin" <mst@redhat.com> - 2016-03-10 13:30 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Michael S. Tsirkin" <mst@redhat.com> - 2016-03-09 16:50 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Roman Kagan <rkagan@virtuozzo.com> - 2016-03-09 18:10 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Michael S. Tsirkin" <mst@redhat.com> - 2016-03-09 18:40 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Roman Kagan <rkagan@virtuozzo.com> - 2016-03-10 11:30 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Rik van Riel <riel@redhat.com> - 2016-03-09 20:40 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Roman Kagan <rkagan@virtuozzo.com> - 2016-03-10 10:40 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Paolo Bonzini <pbonzini@redhat.com> - 2016-03-04 17:30 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Dr. David Alan Gilbert" <dgilbert@redhat.com> - 2016-03-04 20:00 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-07 06:40 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Roman Kagan <rkagan@virtuozzo.com> - 2016-03-09 14:30 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-09 15:20 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-09 07:20 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Roman Kagan <rkagan@virtuozzo.com> - 2016-03-04 09:00 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-04 09:30 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Roman Kagan <rkagan@virtuozzo.com> - 2016-03-04 09:40 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Dr. David Alan Gilbert" <dgilbert@redhat.com> - 2016-03-04 10:10 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-04 10:20 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Michael S. Tsirkin" <mst@redhat.com> - 2016-03-04 10:50 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-04 11:20 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Michael S. Tsirkin" <mst@redhat.com> - 2016-03-04 11:40 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-04 16:20 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Michael S. Tsirkin" <mst@redhat.com> - 2016-03-08 15:10 +0100
RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-08 15:20 +0100
Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization Roman Kagan <rkagan@virtuozzo.com> - 2016-03-04 10:40 +0100
Re: [RFC qemu 0/4] A PV solution for live migration optimization Amit Shah <amit.shah@redhat.com> - 2016-03-08 12:20 +0100
RE: [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-08 14:20 +0100
RE: [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-10 08:50 +0100
Re: [RFC qemu 0/4] A PV solution for live migration optimization Amit Shah <amit.shah@redhat.com> - 2016-03-10 09:00 +0100
RE: [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-10 09:40 +0100
Re: [RFC qemu 0/4] A PV solution for live migration optimization "Dr. David Alan Gilbert" <dgilbert@redhat.com> - 2016-03-10 12:20 +0100
RE: [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-11 03:40 +0100
Re: [RFC qemu 0/4] A PV solution for live migration optimization "Dr. David Alan Gilbert" <dgilbert@redhat.com> - 2016-03-14 18:10 +0100
RE: [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-15 04:40 +0100
Re: [RFC qemu 0/4] A PV solution for live migration optimization "Michael S. Tsirkin" <mst@redhat.com> - 2016-03-15 11:40 +0100
RE: [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-15 12:20 +0100
Re: [RFC qemu 0/4] A PV solution for live migration optimization "Dr. David Alan Gilbert" <dgilbert@redhat.com> - 2016-03-15 21:00 +0100
RE: [RFC qemu 0/4] A PV solution for live migration optimization "Li, Liang Z" <liang.z.li@intel.com> - 2016-03-16 02:30 +0100
Page 1 of 3 [1] 2 3 Next page →
| From | Liang Li <liang.z.li@intel.com> |
|---|---|
| Date | 2016-03-03 12:00 +0100 |
| Subject | [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <r8z97-5Rb-3@gated-at.bofh.it> |
The current QEMU live migration implementation mark the all the
guest's RAM pages as dirtied in the ram bulk stage, all these pages
will be processed and that takes quit a lot of CPU cycles.
From guest's point of view, it doesn't care about the content in free
pages. We can make use of this fact and skip processing the free
pages in the ram bulk stage, it can save a lot CPU cycles and reduce
the network traffic significantly while speed up the live migration
process obviously.
This patch set is the QEMU side implementation.
The virtio-balloon is extended so that QEMU can get the free pages
information from the guest through virtio.
After getting the free pages information (a bitmap), QEMU can use it
to filter out the guest's free pages in the ram bulk stage. This make
the live migration process much more efficient.
This RFC version doesn't take the post-copy and RDMA into
consideration, maybe both of them can benefit from this PV solution
by with some extra modifications.
Performance data
================
Test environment:
CPU: Intel (R) Xeon(R) CPU ES-2699 v3 @ 2.30GHz
Host RAM: 64GB
Host Linux Kernel: 4.2.0 Host OS: CentOS 7.1
Guest Linux Kernel: 4.5.rc6 Guest OS: CentOS 6.6
Network: X540-AT2 with 10 Gigabit connection
Guest RAM: 8GB
Case 1: Idle guest just boots:
============================================
| original | pv
-------------------------------------------
total time(ms) | 1894 | 421
--------------------------------------------
transferred ram(KB) | 398017 | 353242
============================================
Case 2: The guest has ever run some memory consuming workload, the
workload is terminated just before live migration.
============================================
| original | pv
-------------------------------------------
total time(ms) | 7436 | 552
--------------------------------------------
transferred ram(KB) | 8146291 | 361375
============================================
Liang Li (4):
pc: Add code to get the lowmem form PCMachineState
virtio-balloon: Add a new feature to balloon device
migration: not set migration bitmap in setup stage
migration: filter out guest's free pages in ram bulk stage
balloon.c | 30 ++++++++-
hw/i386/pc.c | 5 ++
hw/i386/pc_piix.c | 1 +
hw/i386/pc_q35.c | 1 +
hw/virtio/virtio-balloon.c | 81 ++++++++++++++++++++++++-
include/hw/i386/pc.h | 3 +-
include/hw/virtio/virtio-balloon.h | 17 +++++-
include/standard-headers/linux/virtio_balloon.h | 1 +
include/sysemu/balloon.h | 10 ++-
migration/ram.c | 64 +++++++++++++++----
10 files changed, 195 insertions(+), 18 deletions(-)
--
1.8.3.1
[toc] | [next] | [standalone]
| From | Liang Li <liang.z.li@intel.com> |
|---|---|
| Date | 2016-03-03 12:00 +0100 |
| Subject | [RFC qemu 3/4] migration: not set migration bitmap in setup stage |
| Message-ID | <r8z98-5Rb-27@gated-at.bofh.it> |
| In reply to | #1348973 |
Set ram_list.dirty_memory instead of migration bitmap, the migration
bitmap will be update when doing migration_bitmap_sync().
Set migration_dirty_pages to 0 and it will be updated by
migration_dirty_pages() too.
The following patch is based on this change.
Signed-off-by: Liang Li <liang.z.li@intel.com>
---
migration/ram.c | 12 ++++++------
1 file changed, 6 insertions(+), 6 deletions(-)
diff --git a/migration/ram.c b/migration/ram.c
index 704f6a9..ee2547d 100644
--- a/migration/ram.c
+++ b/migration/ram.c
@@ -1931,19 +1931,19 @@ static int ram_save_setup(QEMUFile *f, void *opaque)
ram_bitmap_pages = last_ram_offset() >> TARGET_PAGE_BITS;
migration_bitmap_rcu = g_new0(struct BitmapRcu, 1);
migration_bitmap_rcu->bmap = bitmap_new(ram_bitmap_pages);
- bitmap_set(migration_bitmap_rcu->bmap, 0, ram_bitmap_pages);
if (migrate_postcopy_ram()) {
migration_bitmap_rcu->unsentmap = bitmap_new(ram_bitmap_pages);
bitmap_set(migration_bitmap_rcu->unsentmap, 0, ram_bitmap_pages);
}
- /*
- * Count the total number of pages used by ram blocks not including any
- * gaps due to alignment or unplugs.
- */
- migration_dirty_pages = ram_bytes_total() >> TARGET_PAGE_BITS;
+ migration_dirty_pages = 0;
+ QLIST_FOREACH_RCU(block, &ram_list.blocks, next) {
+ cpu_physical_memory_set_dirty_range(block->offset,
+ block->used_length,
+ DIRTY_MEMORY_MIGRATION);
+ }
memory_global_dirty_log_start();
migration_bitmap_sync();
qemu_mutex_unlock_ramlist();
--
1.8.3.1
[toc] | [prev] | [next] | [standalone]
| From | Liang Li <liang.z.li@intel.com> |
|---|---|
| Date | 2016-03-03 12:00 +0100 |
| Subject | [RFC qemu 1/4] pc: Add code to get the lowmem form PCMachineState |
| Message-ID | <r8z98-5Rb-23@gated-at.bofh.it> |
| In reply to | #1348973 |
The lowmem will be used by the following patch to get
a correct free pages bitmap.
Signed-off-by: Liang Li <liang.z.li@intel.com>
---
hw/i386/pc.c | 5 +++++
hw/i386/pc_piix.c | 1 +
hw/i386/pc_q35.c | 1 +
include/hw/i386/pc.h | 3 ++-
4 files changed, 9 insertions(+), 1 deletion(-)
diff --git a/hw/i386/pc.c b/hw/i386/pc.c
index 0aeefd2..f794a84 100644
--- a/hw/i386/pc.c
+++ b/hw/i386/pc.c
@@ -1115,6 +1115,11 @@ void pc_hot_add_cpu(const int64_t id, Error **errp)
object_unref(OBJECT(cpu));
}
+ram_addr_t pc_get_lowmem(PCMachineState *pcms)
+{
+ return pcms->lowmem;
+}
+
void pc_cpus_init(PCMachineState *pcms)
{
int i;
diff --git a/hw/i386/pc_piix.c b/hw/i386/pc_piix.c
index 6f8c2cd..268a08c 100644
--- a/hw/i386/pc_piix.c
+++ b/hw/i386/pc_piix.c
@@ -113,6 +113,7 @@ static void pc_init1(MachineState *machine,
}
}
+ pcms->lowmem = lowmem;
if (machine->ram_size >= lowmem) {
pcms->above_4g_mem_size = machine->ram_size - lowmem;
pcms->below_4g_mem_size = lowmem;
diff --git a/hw/i386/pc_q35.c b/hw/i386/pc_q35.c
index 46522c9..8d9bd39 100644
--- a/hw/i386/pc_q35.c
+++ b/hw/i386/pc_q35.c
@@ -101,6 +101,7 @@ static void pc_q35_init(MachineState *machine)
}
}
+ pcms->lowmem = lowmem;
if (machine->ram_size >= lowmem) {
pcms->above_4g_mem_size = machine->ram_size - lowmem;
pcms->below_4g_mem_size = lowmem;
diff --git a/include/hw/i386/pc.h b/include/hw/i386/pc.h
index 8b3546e..3694c91 100644
--- a/include/hw/i386/pc.h
+++ b/include/hw/i386/pc.h
@@ -60,7 +60,7 @@ struct PCMachineState {
bool nvdimm;
/* RAM information (sizes, addresses, configuration): */
- ram_addr_t below_4g_mem_size, above_4g_mem_size;
+ ram_addr_t below_4g_mem_size, above_4g_mem_size, lowmem;
/* CPU and apic information: */
bool apic_xrupt_override;
@@ -229,6 +229,7 @@ void pc_hot_add_cpu(const int64_t id, Error **errp);
void pc_acpi_init(const char *default_dsdt);
void pc_guest_info_init(PCMachineState *pcms);
+ram_addr_t pc_get_lowmem(PCMachineState *pcms);
#define PCI_HOST_PROP_PCI_HOLE_START "pci-hole-start"
#define PCI_HOST_PROP_PCI_HOLE_END "pci-hole-end"
--
1.8.3.1
[toc] | [prev] | [next] | [standalone]
| From | Roman Kagan <rkagan@virtuozzo.com> |
|---|---|
| Date | 2016-03-03 15:10 +0100 |
| Subject | Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <r8C6Z-88W-3@gated-at.bofh.it> |
| In reply to | #1348973 |
On Thu, Mar 03, 2016 at 06:44:24PM +0800, Liang Li wrote: > The current QEMU live migration implementation mark the all the > guest's RAM pages as dirtied in the ram bulk stage, all these pages > will be processed and that takes quit a lot of CPU cycles. > > From guest's point of view, it doesn't care about the content in free > pages. We can make use of this fact and skip processing the free > pages in the ram bulk stage, it can save a lot CPU cycles and reduce > the network traffic significantly while speed up the live migration > process obviously. > > This patch set is the QEMU side implementation. > > The virtio-balloon is extended so that QEMU can get the free pages > information from the guest through virtio. > > After getting the free pages information (a bitmap), QEMU can use it > to filter out the guest's free pages in the ram bulk stage. This make > the live migration process much more efficient. > > This RFC version doesn't take the post-copy and RDMA into > consideration, maybe both of them can benefit from this PV solution > by with some extra modifications. > > Performance data > ================ > > Test environment: > > CPU: Intel (R) Xeon(R) CPU ES-2699 v3 @ 2.30GHz > Host RAM: 64GB > Host Linux Kernel: 4.2.0 Host OS: CentOS 7.1 > Guest Linux Kernel: 4.5.rc6 Guest OS: CentOS 6.6 > Network: X540-AT2 with 10 Gigabit connection > Guest RAM: 8GB > > Case 1: Idle guest just boots: > ============================================ > | original | pv > ------------------------------------------- > total time(ms) | 1894 | 421 > -------------------------------------------- > transferred ram(KB) | 398017 | 353242 > ============================================ > > > Case 2: The guest has ever run some memory consuming workload, the > workload is terminated just before live migration. > ============================================ > | original | pv > ------------------------------------------- > total time(ms) | 7436 | 552 > -------------------------------------------- > transferred ram(KB) | 8146291 | 361375 > ============================================ Both cases look very artificial to me. Normally you migrate VMs which have started long ago and which can't have their services terminated before the migration, so I wouldn't expect any useful amount of free pages obtained this way. OTOH I don't see why you can't just inflate the balloon before the migration, and really optimize the amount of transferred data this way? With the recently proposed VIRTIO_BALLOON_S_AVAIL you can have a fairly good estimate of the optimal balloon size, and with the recently merged balloon deflation on OOM it's a safe thing to do without exposing the guest workloads to OOM risks. Roman.
[toc] | [prev] | [next] | [standalone]
| From | "Li, Liang Z" <liang.z.li@intel.com> |
|---|---|
| Date | 2016-03-04 02:40 +0100 |
| Subject | RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <r8MSK-7vn-9@gated-at.bofh.it> |
| In reply to | #1349181 |
> On Thu, Mar 03, 2016 at 06:44:24PM +0800, Liang Li wrote: > > The current QEMU live migration implementation mark the all the > > guest's RAM pages as dirtied in the ram bulk stage, all these pages > > will be processed and that takes quit a lot of CPU cycles. > > > > From guest's point of view, it doesn't care about the content in free > > pages. We can make use of this fact and skip processing the free pages > > in the ram bulk stage, it can save a lot CPU cycles and reduce the > > network traffic significantly while speed up the live migration > > process obviously. > > > > This patch set is the QEMU side implementation. > > > > The virtio-balloon is extended so that QEMU can get the free pages > > information from the guest through virtio. > > > > After getting the free pages information (a bitmap), QEMU can use it > > to filter out the guest's free pages in the ram bulk stage. This make > > the live migration process much more efficient. > > > > This RFC version doesn't take the post-copy and RDMA into > > consideration, maybe both of them can benefit from this PV solution by > > with some extra modifications. > > > > Performance data > > ================ > > > > Test environment: > > > > CPU: Intel (R) Xeon(R) CPU ES-2699 v3 @ 2.30GHz Host RAM: 64GB > > Host Linux Kernel: 4.2.0 Host OS: CentOS 7.1 > > Guest Linux Kernel: 4.5.rc6 Guest OS: CentOS 6.6 > > Network: X540-AT2 with 10 Gigabit connection Guest RAM: 8GB > > > > Case 1: Idle guest just boots: > > ============================================ > > | original | pv > > ------------------------------------------- > > total time(ms) | 1894 | 421 > > -------------------------------------------- > > transferred ram(KB) | 398017 | 353242 > > ============================================ > > > > > > Case 2: The guest has ever run some memory consuming workload, the > > workload is terminated just before live migration. > > ============================================ > > | original | pv > > ------------------------------------------- > > total time(ms) | 7436 | 552 > > -------------------------------------------- > > transferred ram(KB) | 8146291 | 361375 > > ============================================ > > Both cases look very artificial to me. Normally you migrate VMs which have > started long ago and which can't have their services terminated before the > migration, so I wouldn't expect any useful amount of free pages obtained > this way. > Yes, it's somewhat artificial, just to emphasize the effect. And I think these two cases are very easy to reproduce. Using the real workload and do the test in production environment will be more convince. We can predict that as long as the guest doesn't use out of its memory, this solution may still take affect and shorten the total live migration time. (Off cause, we should consider the time cost of the virtio communication.) > OTOH I don't see why you can't just inflate the balloon before the migration, > and really optimize the amount of transferred data this way? > With the recently proposed VIRTIO_BALLOON_S_AVAIL you can have a fairly > good estimate of the optimal balloon size, and with the recently merged > balloon deflation on OOM it's a safe thing to do without exposing the guest > workloads to OOM risks. > > Roman. Thanks for your information. The size of the free page bitmap is not very large, for a guest with 8GB RAM, only 256KB extra memory is required. Comparing to this solution, inflate the balloon is more expensive. If the balloon size is not so optimal and guest request more memory during live migration, the guest's performance will be impacted. Liang
[toc] | [prev] | [next] | [standalone]
| From | "Dr. David Alan Gilbert" <dgilbert@redhat.com> |
|---|---|
| Date | 2016-03-03 18:50 +0100 |
| Message-ID | <r8FxU-29v-5@gated-at.bofh.it> |
| In reply to | #1348973 |
* Liang Li (liang.z.li@intel.com) wrote: > The current QEMU live migration implementation mark the all the > guest's RAM pages as dirtied in the ram bulk stage, all these pages > will be processed and that takes quit a lot of CPU cycles. > > From guest's point of view, it doesn't care about the content in free > pages. We can make use of this fact and skip processing the free > pages in the ram bulk stage, it can save a lot CPU cycles and reduce > the network traffic significantly while speed up the live migration > process obviously. > > This patch set is the QEMU side implementation. > > The virtio-balloon is extended so that QEMU can get the free pages > information from the guest through virtio. > > After getting the free pages information (a bitmap), QEMU can use it > to filter out the guest's free pages in the ram bulk stage. This make > the live migration process much more efficient. Hi, An interesting solution; I know a few different people have been looking at how to speed up ballooned VM migration. I wonder if it would be possible to avoid the kernel changes by parsing /proc/self/pagemap - if that can be used to detect unmapped/zero mapped pages in the guest ram, would it achieve the same result? > This RFC version doesn't take the post-copy and RDMA into > consideration, maybe both of them can benefit from this PV solution > by with some extra modifications. For postcopy to be safe, you would still need to send a message to the destination telling it that there were zero pages, otherwise the destination can't tell if it's supposed to request the page from the source or treat the page as zero. Dave > > Performance data > ================ > > Test environment: > > CPU: Intel (R) Xeon(R) CPU ES-2699 v3 @ 2.30GHz > Host RAM: 64GB > Host Linux Kernel: 4.2.0 Host OS: CentOS 7.1 > Guest Linux Kernel: 4.5.rc6 Guest OS: CentOS 6.6 > Network: X540-AT2 with 10 Gigabit connection > Guest RAM: 8GB > > Case 1: Idle guest just boots: > ============================================ > | original | pv > ------------------------------------------- > total time(ms) | 1894 | 421 > -------------------------------------------- > transferred ram(KB) | 398017 | 353242 > ============================================ > > > Case 2: The guest has ever run some memory consuming workload, the > workload is terminated just before live migration. > ============================================ > | original | pv > ------------------------------------------- > total time(ms) | 7436 | 552 > -------------------------------------------- > transferred ram(KB) | 8146291 | 361375 > ============================================ > > Liang Li (4): > pc: Add code to get the lowmem form PCMachineState > virtio-balloon: Add a new feature to balloon device > migration: not set migration bitmap in setup stage > migration: filter out guest's free pages in ram bulk stage > > balloon.c | 30 ++++++++- > hw/i386/pc.c | 5 ++ > hw/i386/pc_piix.c | 1 + > hw/i386/pc_q35.c | 1 + > hw/virtio/virtio-balloon.c | 81 ++++++++++++++++++++++++- > include/hw/i386/pc.h | 3 +- > include/hw/virtio/virtio-balloon.h | 17 +++++- > include/standard-headers/linux/virtio_balloon.h | 1 + > include/sysemu/balloon.h | 10 ++- > migration/ram.c | 64 +++++++++++++++---- > 10 files changed, 195 insertions(+), 18 deletions(-) > > -- > 1.8.3.1 > -- Dr. David Alan Gilbert / dgilbert@redhat.com / Manchester, UK
[toc] | [prev] | [next] | [standalone]
| From | "Li, Liang Z" <liang.z.li@intel.com> |
|---|---|
| Date | 2016-03-04 03:00 +0100 |
| Message-ID | <r8Nc6-7Fq-19@gated-at.bofh.it> |
| In reply to | #1349447 |
> Subject: Re: [RFC qemu 0/4] A PV solution for live migration optimization > > * Liang Li (liang.z.li@intel.com) wrote: > > The current QEMU live migration implementation mark the all the > > guest's RAM pages as dirtied in the ram bulk stage, all these pages > > will be processed and that takes quit a lot of CPU cycles. > > > > From guest's point of view, it doesn't care about the content in free > > pages. We can make use of this fact and skip processing the free pages > > in the ram bulk stage, it can save a lot CPU cycles and reduce the > > network traffic significantly while speed up the live migration > > process obviously. > > > > This patch set is the QEMU side implementation. > > > > The virtio-balloon is extended so that QEMU can get the free pages > > information from the guest through virtio. > > > > After getting the free pages information (a bitmap), QEMU can use it > > to filter out the guest's free pages in the ram bulk stage. This make > > the live migration process much more efficient. > > Hi, > An interesting solution; I know a few different people have been looking at > how to speed up ballooned VM migration. > Ooh, different solutions for the same purpose, and both based on the balloon. > I wonder if it would be possible to avoid the kernel changes by parsing > /proc/self/pagemap - if that can be used to detect unmapped/zero mapped > pages in the guest ram, would it achieve the same result? > Only detect the unmapped/zero mapped pages is not enough. Consider the situation like case 2, it can't achieve the same result. > > This RFC version doesn't take the post-copy and RDMA into > > consideration, maybe both of them can benefit from this PV solution by > > with some extra modifications. > > For postcopy to be safe, you would still need to send a message to the > destination telling it that there were zero pages, otherwise the destination > can't tell if it's supposed to request the page from the source or treat the > page as zero. > > Dave I will consider this later, thanks, Dave. Liang > > > > > Performance data > > ================ > > > > Test environment: > > > > CPU: Intel (R) Xeon(R) CPU ES-2699 v3 @ 2.30GHz Host RAM: 64GB > > Host Linux Kernel: 4.2.0 Host OS: CentOS 7.1 > > Guest Linux Kernel: 4.5.rc6 Guest OS: CentOS 6.6 > > Network: X540-AT2 with 10 Gigabit connection Guest RAM: 8GB > > > > Case 1: Idle guest just boots: > > ============================================ > > | original | pv > > ------------------------------------------- > > total time(ms) | 1894 | 421 > > -------------------------------------------- > > transferred ram(KB) | 398017 | 353242 > > ============================================ > > > > > > Case 2: The guest has ever run some memory consuming workload, the > > workload is terminated just before live migration. > > ============================================ > > | original | pv > > ------------------------------------------- > > total time(ms) | 7436 | 552 > > -------------------------------------------- > > transferred ram(KB) | 8146291 | 361375 > > ============================================ > >
[toc] | [prev] | [next] | [standalone]
| From | Roman Kagan <rkagan@virtuozzo.com> |
|---|---|
| Date | 2016-03-04 09:20 +0100 |
| Subject | Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <r8T7P-3PF-3@gated-at.bofh.it> |
| In reply to | #1349769 |
On Fri, Mar 04, 2016 at 01:52:53AM +0000, Li, Liang Z wrote: > > I wonder if it would be possible to avoid the kernel changes by parsing > > /proc/self/pagemap - if that can be used to detect unmapped/zero mapped > > pages in the guest ram, would it achieve the same result? > > Only detect the unmapped/zero mapped pages is not enough. Consider the > situation like case 2, it can't achieve the same result. Your case 2 doesn't exist in the real world. If people could stop their main memory consumer in the guest prior to migration they wouldn't need live migration at all. I tend to think you can safely assume there's no free memory in the guest, so there's little point optimizing for it. OTOH it makes perfect sense optimizing for the unmapped memory that's made up, in particular, by the ballon, and consider inflating the balloon right before migration unless you already maintain it at the optimal size for other reasons (like e.g. a global resource manager optimizing the VM density). Roman.
[toc] | [prev] | [next] | [standalone]
| From | "Li, Liang Z" <liang.z.li@intel.com> |
|---|---|
| Date | 2016-03-04 10:10 +0100 |
| Subject | RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <r8TUf-4rg-35@gated-at.bofh.it> |
| In reply to | #1349935 |
> On Fri, Mar 04, 2016 at 01:52:53AM +0000, Li, Liang Z wrote: > > > I wonder if it would be possible to avoid the kernel changes by > > > parsing /proc/self/pagemap - if that can be used to detect > > > unmapped/zero mapped pages in the guest ram, would it achieve the > same result? > > > > Only detect the unmapped/zero mapped pages is not enough. Consider > the > > situation like case 2, it can't achieve the same result. > > Your case 2 doesn't exist in the real world. If people could stop their main > memory consumer in the guest prior to migration they wouldn't need live > migration at all. The case 2 is just a simplified scenario, not a real case. As long as the guest's memory usage does not keep increasing, or not always run out, it can be covered by the case 2. > I tend to think you can safely assume there's no free memory in the guest, so > there's little point optimizing for it. If this is true, we should not inflate the balloon either. > OTOH it makes perfect sense optimizing for the unmapped memory that's > made up, in particular, by the ballon, and consider inflating the balloon right > before migration unless you already maintain it at the optimal size for other > reasons (like e.g. a global resource manager optimizing the VM density). > Yes, I believe the current balloon works and it's simple. Do you take the performance impact for consideration? For and 8G guest, it takes about 5s to inflating the balloon. But it only takes 20ms to traverse the free_list and construct the free pages bitmap. In this period, the guest are very busy. By inflating the balloon, all the guest's pages are still be processed (zero page checking). The only advantage of ' inflating the balloon before live migration' is simple, nothing more. Liang > Roman.
[toc] | [prev] | [next] | [standalone]
| From | Roman Kagan <rkagan@virtuozzo.com> |
|---|---|
| Date | 2016-03-04 11:30 +0100 |
| Subject | Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <r8V9E-5bS-11@gated-at.bofh.it> |
| In reply to | #1349976 |
On Fri, Mar 04, 2016 at 09:08:44AM +0000, Li, Liang Z wrote: > > On Fri, Mar 04, 2016 at 01:52:53AM +0000, Li, Liang Z wrote: > > > > I wonder if it would be possible to avoid the kernel changes by > > > > parsing /proc/self/pagemap - if that can be used to detect > > > > unmapped/zero mapped pages in the guest ram, would it achieve the > > same result? > > > > > > Only detect the unmapped/zero mapped pages is not enough. Consider > > the > > > situation like case 2, it can't achieve the same result. > > > > Your case 2 doesn't exist in the real world. If people could stop their main > > memory consumer in the guest prior to migration they wouldn't need live > > migration at all. > > The case 2 is just a simplified scenario, not a real case. > As long as the guest's memory usage does not keep increasing, or not always run out, > it can be covered by the case 2. The memory usage will keep increasing due to ever growing caches, etc, so you'll be left with very little free memory fairly soon. > > I tend to think you can safely assume there's no free memory in the guest, so > > there's little point optimizing for it. > > If this is true, we should not inflate the balloon either. We certainly should if there's "available" memory, i.e. not free but cheap to reclaim. > > OTOH it makes perfect sense optimizing for the unmapped memory that's > > made up, in particular, by the ballon, and consider inflating the balloon right > > before migration unless you already maintain it at the optimal size for other > > reasons (like e.g. a global resource manager optimizing the VM density). > > > > Yes, I believe the current balloon works and it's simple. Do you take the performance impact for consideration? > For and 8G guest, it takes about 5s to inflating the balloon. But it only takes 20ms to traverse the free_list and > construct the free pages bitmap. I don't have any feeling of how important the difference is. And if the limiting factor for balloon inflation speed is the granularity of communication it may be worth optimizing that, because quick balloon reaction may be important in certain resource management scenarios. > By inflating the balloon, all the guest's pages are still be processed (zero page checking). Not sure what you mean. If you describe the current state of affairs that's exactly the suggested optimization point: skip unmapped pages. > The only advantage of ' inflating the balloon before live migration' is simple, nothing more. That's a big advantage. Another one is that it does something useful in real-world scenarios. Roman.
[toc] | [prev] | [next] | [standalone]
| From | "Li, Liang Z" <liang.z.li@intel.com> |
|---|---|
| Date | 2016-03-04 15:30 +0100 |
| Subject | RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <r8YTT-810-3@gated-at.bofh.it> |
| In reply to | #1350126 |
> Subject: Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration > optimization > > On Fri, Mar 04, 2016 at 09:08:44AM +0000, Li, Liang Z wrote: > > > On Fri, Mar 04, 2016 at 01:52:53AM +0000, Li, Liang Z wrote: > > > > > I wonder if it would be possible to avoid the kernel changes > > > > > by parsing /proc/self/pagemap - if that can be used to detect > > > > > unmapped/zero mapped pages in the guest ram, would it achieve > > > > > the > > > same result? > > > > > > > > Only detect the unmapped/zero mapped pages is not enough. > Consider > > > the > > > > situation like case 2, it can't achieve the same result. > > > > > > Your case 2 doesn't exist in the real world. If people could stop > > > their main memory consumer in the guest prior to migration they > > > wouldn't need live migration at all. > > > > The case 2 is just a simplified scenario, not a real case. > > As long as the guest's memory usage does not keep increasing, or not > > always run out, it can be covered by the case 2. > > The memory usage will keep increasing due to ever growing caches, etc, so > you'll be left with very little free memory fairly soon. > I don't think so. > > > I tend to think you can safely assume there's no free memory in the > > > guest, so there's little point optimizing for it. > > > > If this is true, we should not inflate the balloon either. > > We certainly should if there's "available" memory, i.e. not free but cheap to > reclaim. > What's your mean by "available" memory? if they are not free, I don't think it's cheap. > > > OTOH it makes perfect sense optimizing for the unmapped memory > > > that's made up, in particular, by the ballon, and consider inflating > > > the balloon right before migration unless you already maintain it at > > > the optimal size for other reasons (like e.g. a global resource manager > optimizing the VM density). > > > > > > > Yes, I believe the current balloon works and it's simple. Do you take the > performance impact for consideration? > > For and 8G guest, it takes about 5s to inflating the balloon. But it > > only takes 20ms to traverse the free_list and construct the free pages > bitmap. > > I don't have any feeling of how important the difference is. And if the > limiting factor for balloon inflation speed is the granularity of communication > it may be worth optimizing that, because quick balloon reaction may be > important in certain resource management scenarios. > > > By inflating the balloon, all the guest's pages are still be processed (zero > page checking). > > Not sure what you mean. If you describe the current state of affairs that's > exactly the suggested optimization point: skip unmapped pages. > You'd better check the live migration code. > > The only advantage of ' inflating the balloon before live migration' is simple, > nothing more. > > That's a big advantage. Another one is that it does something useful in real- > world scenarios. > I don't think the heave performance impaction is something useful in real world scenarios. Liang > Roman.
[toc] | [prev] | [next] | [standalone]
| From | "Michael S. Tsirkin" <mst@redhat.com> |
|---|---|
| Date | 2016-03-04 15:50 +0100 |
| Subject | Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <r8Zdf-89n-3@gated-at.bofh.it> |
| In reply to | #1350256 |
On Fri, Mar 04, 2016 at 02:26:49PM +0000, Li, Liang Z wrote:
> > Subject: Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration
> > optimization
> >
> > On Fri, Mar 04, 2016 at 09:08:44AM +0000, Li, Liang Z wrote:
> > > > On Fri, Mar 04, 2016 at 01:52:53AM +0000, Li, Liang Z wrote:
> > > > > > I wonder if it would be possible to avoid the kernel changes
> > > > > > by parsing /proc/self/pagemap - if that can be used to detect
> > > > > > unmapped/zero mapped pages in the guest ram, would it achieve
> > > > > > the
> > > > same result?
> > > > >
> > > > > Only detect the unmapped/zero mapped pages is not enough.
> > Consider
> > > > the
> > > > > situation like case 2, it can't achieve the same result.
> > > >
> > > > Your case 2 doesn't exist in the real world. If people could stop
> > > > their main memory consumer in the guest prior to migration they
> > > > wouldn't need live migration at all.
> > >
> > > The case 2 is just a simplified scenario, not a real case.
> > > As long as the guest's memory usage does not keep increasing, or not
> > > always run out, it can be covered by the case 2.
> >
> > The memory usage will keep increasing due to ever growing caches, etc, so
> > you'll be left with very little free memory fairly soon.
> >
>
> I don't think so.
Here's my laptop:
KiB Mem : 16048560 total, 8574956 free, 3360532 used, 4113072 buff/cache
But here's a server:
KiB Mem: 32892768 total, 20092812 used, 12799956 free, 368704 buffers
What is the difference? A ton of tiny daemons not doing anything,
staying resident in memory.
> > > > I tend to think you can safely assume there's no free memory in the
> > > > guest, so there's little point optimizing for it.
> > >
> > > If this is true, we should not inflate the balloon either.
> >
> > We certainly should if there's "available" memory, i.e. not free but cheap to
> > reclaim.
> >
>
> What's your mean by "available" memory? if they are not free, I don't think it's cheap.
clean pages are cheap to drop as they don't have to be written.
whether they will be ever be used is another matter.
> > > > OTOH it makes perfect sense optimizing for the unmapped memory
> > > > that's made up, in particular, by the ballon, and consider inflating
> > > > the balloon right before migration unless you already maintain it at
> > > > the optimal size for other reasons (like e.g. a global resource manager
> > optimizing the VM density).
> > > >
> > >
> > > Yes, I believe the current balloon works and it's simple. Do you take the
> > performance impact for consideration?
> > > For and 8G guest, it takes about 5s to inflating the balloon. But it
> > > only takes 20ms to traverse the free_list and construct the free pages
> > bitmap.
> >
> > I don't have any feeling of how important the difference is. And if the
> > limiting factor for balloon inflation speed is the granularity of communication
> > it may be worth optimizing that, because quick balloon reaction may be
> > important in certain resource management scenarios.
> >
> > > By inflating the balloon, all the guest's pages are still be processed (zero
> > page checking).
> >
> > Not sure what you mean. If you describe the current state of affairs that's
> > exactly the suggested optimization point: skip unmapped pages.
> >
>
> You'd better check the live migration code.
What's there to check in migration code?
Here's the extent of what balloon does on output:
while (iov_to_buf(elem->out_sg, elem->out_num, offset, &pfn, 4) == 4) {
ram_addr_t pa;
ram_addr_t addr;
int p = virtio_ldl_p(vdev, &pfn);
pa = (ram_addr_t) p << VIRTIO_BALLOON_PFN_SHIFT;
offset += 4;
/* FIXME: remove get_system_memory(), but how? */
section = memory_region_find(get_system_memory(), pa, 1);
if (!int128_nz(section.size) || !memory_region_is_ram(section.mr))
continue;
trace_virtio_balloon_handle_output(memory_region_name(section.mr),
pa);
/* Using memory_region_get_ram_ptr is bending the rules a bit, but
should be OK because we only want a single page. */
addr = section.offset_within_region;
balloon_page(memory_region_get_ram_ptr(section.mr) + addr,
!!(vq == s->dvq));
memory_region_unref(section.mr);
}
so all that happens when we get a page is balloon_page.
and
static void balloon_page(void *addr, int deflate)
{
#if defined(__linux__)
if (!qemu_balloon_is_inhibited() && (!kvm_enabled() ||
kvm_has_sync_mmu())) {
qemu_madvise(addr, TARGET_PAGE_SIZE,
deflate ? QEMU_MADV_WILLNEED : QEMU_MADV_DONTNEED);
}
#endif
}
Do you see anything that tracks pages to help migration skip
the ballooned memory? I don't.
> > > The only advantage of ' inflating the balloon before live migration' is simple,
> > nothing more.
> >
> > That's a big advantage. Another one is that it does something useful in real-
> > world scenarios.
> >
>
> I don't think the heave performance impaction is something useful in real world scenarios.
>
> Liang
> > Roman.
So fix the performance then. You will have to try harder if you want to
convince people that the performance is due to bad host/guest interface,
and so we have to change *that*.
--
MST
[toc] | [prev] | [next] | [standalone]
| From | "Li, Liang Z" <liang.z.li@intel.com> |
|---|---|
| Date | 2016-03-04 16:50 +0100 |
| Subject | RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <r909j-mc-9@gated-at.bofh.it> |
| In reply to | #1350259 |
> > > > > > Only detect the unmapped/zero mapped pages is not enough.
> > > Consider
> > > > > the
> > > > > > situation like case 2, it can't achieve the same result.
> > > > >
> > > > > Your case 2 doesn't exist in the real world. If people could
> > > > > stop their main memory consumer in the guest prior to migration
> > > > > they wouldn't need live migration at all.
> > > >
> > > > The case 2 is just a simplified scenario, not a real case.
> > > > As long as the guest's memory usage does not keep increasing, or
> > > > not always run out, it can be covered by the case 2.
> > >
> > > The memory usage will keep increasing due to ever growing caches,
> > > etc, so you'll be left with very little free memory fairly soon.
> > >
> >
> > I don't think so.
>
> Here's my laptop:
> KiB Mem : 16048560 total, 8574956 free, 3360532 used, 4113072 buff/cache
>
> But here's a server:
> KiB Mem: 32892768 total, 20092812 used, 12799956 free, 368704 buffers
>
> What is the difference? A ton of tiny daemons not doing anything, staying
> resident in memory.
>
> > > > > I tend to think you can safely assume there's no free memory in
> > > > > the guest, so there's little point optimizing for it.
> > > >
> > > > If this is true, we should not inflate the balloon either.
> > >
> > > We certainly should if there's "available" memory, i.e. not free but
> > > cheap to reclaim.
> > >
> >
> > What's your mean by "available" memory? if they are not free, I don't think
> it's cheap.
>
> clean pages are cheap to drop as they don't have to be written.
> whether they will be ever be used is another matter.
>
> > > > > OTOH it makes perfect sense optimizing for the unmapped memory
> > > > > that's made up, in particular, by the ballon, and consider
> > > > > inflating the balloon right before migration unless you already
> > > > > maintain it at the optimal size for other reasons (like e.g. a
> > > > > global resource manager
> > > optimizing the VM density).
> > > > >
> > > >
> > > > Yes, I believe the current balloon works and it's simple. Do you
> > > > take the
> > > performance impact for consideration?
> > > > For and 8G guest, it takes about 5s to inflating the balloon. But
> > > > it only takes 20ms to traverse the free_list and construct the
> > > > free pages
> > > bitmap.
> > >
> > > I don't have any feeling of how important the difference is. And if
> > > the limiting factor for balloon inflation speed is the granularity
> > > of communication it may be worth optimizing that, because quick
> > > balloon reaction may be important in certain resource management
> scenarios.
> > >
> > > > By inflating the balloon, all the guest's pages are still be
> > > > processed (zero
> > > page checking).
> > >
> > > Not sure what you mean. If you describe the current state of
> > > affairs that's exactly the suggested optimization point: skip unmapped
> pages.
> > >
> >
> > You'd better check the live migration code.
>
> What's there to check in migration code?
> Here's the extent of what balloon does on output:
>
>
> while (iov_to_buf(elem->out_sg, elem->out_num, offset, &pfn, 4) == 4)
> {
> ram_addr_t pa;
> ram_addr_t addr;
> int p = virtio_ldl_p(vdev, &pfn);
>
> pa = (ram_addr_t) p << VIRTIO_BALLOON_PFN_SHIFT;
> offset += 4;
>
> /* FIXME: remove get_system_memory(), but how? */
> section = memory_region_find(get_system_memory(), pa, 1);
> if (!int128_nz(section.size) || !memory_region_is_ram(section.mr))
> continue;
>
>
> trace_virtio_balloon_handle_output(memory_region_name(section.mr),
> pa);
> /* Using memory_region_get_ram_ptr is bending the rules a bit, but
> should be OK because we only want a single page. */
> addr = section.offset_within_region;
> balloon_page(memory_region_get_ram_ptr(section.mr) + addr,
> !!(vq == s->dvq));
> memory_region_unref(section.mr);
> }
>
> so all that happens when we get a page is balloon_page.
> and
>
> static void balloon_page(void *addr, int deflate) { #if defined(__linux__)
> if (!qemu_balloon_is_inhibited() && (!kvm_enabled() ||
> kvm_has_sync_mmu())) {
> qemu_madvise(addr, TARGET_PAGE_SIZE,
> deflate ? QEMU_MADV_WILLNEED : QEMU_MADV_DONTNEED);
> }
> #endif
> }
>
>
> Do you see anything that tracks pages to help migration skip the ballooned
> memory? I don't.
>
No. And it's exactly what I mean. The ballooned memory is still processed during
live migration without skipping. The live migration code is in migration/ram.c.
>
> > > > The only advantage of ' inflating the balloon before live
> > > > migration' is simple,
> > > nothing more.
> > >
> > > That's a big advantage. Another one is that it does something
> > > useful in real- world scenarios.
> > >
> >
> > I don't think the heave performance impaction is something useful in real
> world scenarios.
> >
> > Liang
> > > Roman.
>
> So fix the performance then. You will have to try harder if you want to
> convince people that the performance is due to bad host/guest interface,
> and so we have to change *that*.
>
Actually, the PV solution is irrelevant with the balloon mechanism, I just use it
to transfer information between host and guest.
I am not sure if I should implement a new virtio device, and I want to get the answer from
the community.
In this RFC patch, to make things simple, I choose to extend the virtio-balloon and use the
extended interface to transfer the request and free_page_bimap content.
I am not intend to change the current virtio-balloon implementation.
Liang
> --
> MST
[toc] | [prev] | [next] | [standalone]
| From | "Michael S. Tsirkin" <mst@redhat.com> |
|---|---|
| Date | 2016-03-05 21:00 +0100 |
| Subject | Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <r9qwO-2v2-7@gated-at.bofh.it> |
| In reply to | #1350324 |
On Fri, Mar 04, 2016 at 03:49:37PM +0000, Li, Liang Z wrote:
> > > > > > > Only detect the unmapped/zero mapped pages is not enough.
> > > > Consider
> > > > > > the
> > > > > > > situation like case 2, it can't achieve the same result.
> > > > > >
> > > > > > Your case 2 doesn't exist in the real world. If people could
> > > > > > stop their main memory consumer in the guest prior to migration
> > > > > > they wouldn't need live migration at all.
> > > > >
> > > > > The case 2 is just a simplified scenario, not a real case.
> > > > > As long as the guest's memory usage does not keep increasing, or
> > > > > not always run out, it can be covered by the case 2.
> > > >
> > > > The memory usage will keep increasing due to ever growing caches,
> > > > etc, so you'll be left with very little free memory fairly soon.
> > > >
> > >
> > > I don't think so.
> >
> > Here's my laptop:
> > KiB Mem : 16048560 total, 8574956 free, 3360532 used, 4113072 buff/cache
> >
> > But here's a server:
> > KiB Mem: 32892768 total, 20092812 used, 12799956 free, 368704 buffers
> >
> > What is the difference? A ton of tiny daemons not doing anything, staying
> > resident in memory.
> >
> > > > > > I tend to think you can safely assume there's no free memory in
> > > > > > the guest, so there's little point optimizing for it.
> > > > >
> > > > > If this is true, we should not inflate the balloon either.
> > > >
> > > > We certainly should if there's "available" memory, i.e. not free but
> > > > cheap to reclaim.
> > > >
> > >
> > > What's your mean by "available" memory? if they are not free, I don't think
> > it's cheap.
> >
> > clean pages are cheap to drop as they don't have to be written.
> > whether they will be ever be used is another matter.
> >
> > > > > > OTOH it makes perfect sense optimizing for the unmapped memory
> > > > > > that's made up, in particular, by the ballon, and consider
> > > > > > inflating the balloon right before migration unless you already
> > > > > > maintain it at the optimal size for other reasons (like e.g. a
> > > > > > global resource manager
> > > > optimizing the VM density).
> > > > > >
> > > > >
> > > > > Yes, I believe the current balloon works and it's simple. Do you
> > > > > take the
> > > > performance impact for consideration?
> > > > > For and 8G guest, it takes about 5s to inflating the balloon. But
> > > > > it only takes 20ms to traverse the free_list and construct the
> > > > > free pages
> > > > bitmap.
> > > >
> > > > I don't have any feeling of how important the difference is. And if
> > > > the limiting factor for balloon inflation speed is the granularity
> > > > of communication it may be worth optimizing that, because quick
> > > > balloon reaction may be important in certain resource management
> > scenarios.
> > > >
> > > > > By inflating the balloon, all the guest's pages are still be
> > > > > processed (zero
> > > > page checking).
> > > >
> > > > Not sure what you mean. If you describe the current state of
> > > > affairs that's exactly the suggested optimization point: skip unmapped
> > pages.
> > > >
> > >
> > > You'd better check the live migration code.
> >
> > What's there to check in migration code?
> > Here's the extent of what balloon does on output:
> >
> >
> > while (iov_to_buf(elem->out_sg, elem->out_num, offset, &pfn, 4) == 4)
> > {
> > ram_addr_t pa;
> > ram_addr_t addr;
> > int p = virtio_ldl_p(vdev, &pfn);
> >
> > pa = (ram_addr_t) p << VIRTIO_BALLOON_PFN_SHIFT;
> > offset += 4;
> >
> > /* FIXME: remove get_system_memory(), but how? */
> > section = memory_region_find(get_system_memory(), pa, 1);
> > if (!int128_nz(section.size) || !memory_region_is_ram(section.mr))
> > continue;
> >
> >
> > trace_virtio_balloon_handle_output(memory_region_name(section.mr),
> > pa);
> > /* Using memory_region_get_ram_ptr is bending the rules a bit, but
> > should be OK because we only want a single page. */
> > addr = section.offset_within_region;
> > balloon_page(memory_region_get_ram_ptr(section.mr) + addr,
> > !!(vq == s->dvq));
> > memory_region_unref(section.mr);
> > }
> >
> > so all that happens when we get a page is balloon_page.
> > and
> >
> > static void balloon_page(void *addr, int deflate) { #if defined(__linux__)
> > if (!qemu_balloon_is_inhibited() && (!kvm_enabled() ||
> > kvm_has_sync_mmu())) {
> > qemu_madvise(addr, TARGET_PAGE_SIZE,
> > deflate ? QEMU_MADV_WILLNEED : QEMU_MADV_DONTNEED);
> > }
> > #endif
> > }
> >
> >
> > Do you see anything that tracks pages to help migration skip the ballooned
> > memory? I don't.
> >
>
> No. And it's exactly what I mean. The ballooned memory is still processed during
> live migration without skipping. The live migration code is in migration/ram.c.
So if guest acknowledged VIRTIO_BALLOON_F_MUST_TELL_HOST,
we can teach qemu to skip these pages.
Want to write a patch to do this?
> >
> > > > > The only advantage of ' inflating the balloon before live
> > > > > migration' is simple,
> > > > nothing more.
> > > >
> > > > That's a big advantage. Another one is that it does something
> > > > useful in real- world scenarios.
> > > >
> > >
> > > I don't think the heave performance impaction is something useful in real
> > world scenarios.
> > >
> > > Liang
> > > > Roman.
> >
> > So fix the performance then. You will have to try harder if you want to
> > convince people that the performance is due to bad host/guest interface,
> > and so we have to change *that*.
> >
>
> Actually, the PV solution is irrelevant with the balloon mechanism, I just use it
> to transfer information between host and guest.
> I am not sure if I should implement a new virtio device, and I want to get the answer from
> the community.
> In this RFC patch, to make things simple, I choose to extend the virtio-balloon and use the
> extended interface to transfer the request and free_page_bimap content.
>
> I am not intend to change the current virtio-balloon implementation.
>
> Liang
And the answer would depend on the answer to my question above.
Does balloon need an interface passing page bitmaps around?
Does this speed up any operations?
OTOH what if you use the regular balloon interface with your patches?
> > --
> > MST
[toc] | [prev] | [next] | [standalone]
| From | "Li, Liang Z" <liang.z.li@intel.com> |
|---|---|
| Date | 2016-03-07 08:00 +0100 |
| Subject | RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <r9Xj4-7Of-5@gated-at.bofh.it> |
| In reply to | #1350964 |
> > No. And it's exactly what I mean. The ballooned memory is still > > processed during live migration without skipping. The live migration code is > in migration/ram.c. > > So if guest acknowledged VIRTIO_BALLOON_F_MUST_TELL_HOST, we can > teach qemu to skip these pages. > Want to write a patch to do this? > Yes, we really can teach qemu to skip these pages and it's not hard. The problem is the poor performance, this PV solution is aimed to make it more efficient and reduce the performance impact on guest. > > > > > > > > > The only advantage of ' inflating the balloon before live > > > > > > migration' is simple, > > > > > nothing more. > > > > > > > > > > That's a big advantage. Another one is that it does something > > > > > useful in real- world scenarios. > > > > > > > > > > > > > I don't think the heave performance impaction is something useful > > > > in real > > > world scenarios. > > > > > > > > Liang > > > > > Roman. > > > > > > So fix the performance then. You will have to try harder if you want > > > to convince people that the performance is due to bad host/guest > > > interface, and so we have to change *that*. > > > > > > > Actually, the PV solution is irrelevant with the balloon mechanism, I > > just use it to transfer information between host and guest. > > I am not sure if I should implement a new virtio device, and I want to > > get the answer from the community. > > In this RFC patch, to make things simple, I choose to extend the > > virtio-balloon and use the extended interface to transfer the request and > free_page_bimap content. > > > > I am not intend to change the current virtio-balloon implementation. > > > > Liang > > And the answer would depend on the answer to my question above. > Does balloon need an interface passing page bitmaps around? Yes, I need a new interface. > Does this speed up any operations? No, a new interface will not speed up anything, but it is the easiest way to solve the compatibility issue. > OTOH what if you use the regular balloon interface with your patches? > The regular balloon interfaces have their specific function and I can't use them in my patches. If using these regular interface, I have to do a lot of changes to keep the compatibility. > > > > -- > > > MST
[toc] | [prev] | [next] | [standalone]
| From | "Michael S. Tsirkin" <mst@redhat.com> |
|---|---|
| Date | 2016-03-07 12:50 +0100 |
| Subject | Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <ra1PH-2jh-5@gated-at.bofh.it> |
| In reply to | #1351330 |
On Mon, Mar 07, 2016 at 06:49:19AM +0000, Li, Liang Z wrote: > > > No. And it's exactly what I mean. The ballooned memory is still > > > processed during live migration without skipping. The live migration code is > > in migration/ram.c. > > > > So if guest acknowledged VIRTIO_BALLOON_F_MUST_TELL_HOST, we can > > teach qemu to skip these pages. > > Want to write a patch to do this? > > > > Yes, we really can teach qemu to skip these pages and it's not hard. > The problem is the poor performance, this PV solution Balloon is always PV. And do not call patches solutions please. > is aimed to make it more > efficient and reduce the performance impact on guest. We need to get a bit beyond this. You are making multiple changes, it seems to make sense to split it all up, and analyse each change separately. If you don't this patchset will be stuck: as you have seen people aren't convinced it actually helps with real workloads. > > > > > > > > > > > The only advantage of ' inflating the balloon before live > > > > > > > migration' is simple, > > > > > > nothing more. > > > > > > > > > > > > That's a big advantage. Another one is that it does something > > > > > > useful in real- world scenarios. > > > > > > > > > > > > > > > > I don't think the heave performance impaction is something useful > > > > > in real > > > > world scenarios. > > > > > > > > > > Liang > > > > > > Roman. > > > > > > > > So fix the performance then. You will have to try harder if you want > > > > to convince people that the performance is due to bad host/guest > > > > interface, and so we have to change *that*. > > > > > > > > > > Actually, the PV solution is irrelevant with the balloon mechanism, I > > > just use it to transfer information between host and guest. > > > I am not sure if I should implement a new virtio device, and I want to > > > get the answer from the community. > > > In this RFC patch, to make things simple, I choose to extend the > > > virtio-balloon and use the extended interface to transfer the request and > > free_page_bimap content. > > > > > > I am not intend to change the current virtio-balloon implementation. > > > > > > Liang > > > > And the answer would depend on the answer to my question above. > > Does balloon need an interface passing page bitmaps around? > > Yes, I need a new interface. Possibly, but you will need to justify this at some level if you care about upstreaming your patches. > > Does this speed up any operations? > > No, a new interface will not speed up anything, but it is the easiest way to solve the compatibility issue. A bunch of new code is often easier to write than to figure out the old one, but if we keep piling it up we'll end up with an unmaintainable mess. So we are rather careful about adding new interfaces, and we try to make them generic sometimes even at cost of slight inefficiencies. > > OTOH what if you use the regular balloon interface with your patches? > > > > The regular balloon interfaces have their specific function and I can't use them in my patches. > If using these regular interface, I have to do a lot of changes to keep the compatibility. Why can't you? What exactly do we need to change? If we put things in terms of the balloon, that supports adding and removing pages. Using these terms, let's enumerate: - a new method (e.g. new virtqueue) that adds and immediately removes page in a balloon clearly, you can add then remove using the existing interfaces is a single command significantly faster than using existing two vqs? - a new kind of request that says "add (and immediately remove?) as many pages as you can" sounds rather benign - a new kind of message that adds multiple pages using a bitmap (instead of an address list) again, is this significantly faster? Does not look like compatibility is an issue, to me. At some level, your patches look like page hints. If we have more patches in mind that use page hints, then a new hint device might make sense. However, people experimented with page hints in the past, so far this always went nowhere. E.g. I CC Rick who saw some problems when page hints interact with huge pages. Rick, could you elaborate please? -- MST
[toc] | [prev] | [next] | [standalone]
| From | "Li, Liang Z" <liang.z.li@intel.com> |
|---|---|
| Date | 2016-03-07 16:10 +0100 |
| Subject | RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <ra4Xf-4ss-3@gated-at.bofh.it> |
| In reply to | #1351559 |
> Cc: Roman Kagan; Dr. David Alan Gilbert; ehabkost@redhat.com; > kvm@vger.kernel.org; quintela@redhat.com; linux-kernel@vger.kernel.org; > qemu-devel@nongnu.org; linux-mm@kvack.org; amit.shah@redhat.com; > pbonzini@redhat.com; akpm@linux-foundation.org; > virtualization@lists.linux-foundation.org; rth@twiddle.net; riel@redhat.com > Subject: Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration > optimization > > On Mon, Mar 07, 2016 at 06:49:19AM +0000, Li, Liang Z wrote: > > > > No. And it's exactly what I mean. The ballooned memory is still > > > > processed during live migration without skipping. The live > > > > migration code is > > > in migration/ram.c. > > > > > > So if guest acknowledged VIRTIO_BALLOON_F_MUST_TELL_HOST, we > can > > > teach qemu to skip these pages. > > > Want to write a patch to do this? > > > > > > > Yes, we really can teach qemu to skip these pages and it's not hard. > > The problem is the poor performance, this PV solution > > Balloon is always PV. And do not call patches solutions please. > OK. > > is aimed to make it more > > efficient and reduce the performance impact on guest. > > We need to get a bit beyond this. You are making multiple changes, it seems > to make sense to split it all up, and analyse each change separately. If you > don't this patchset will be stuck: as you have seen people aren't convinced it > actually helps with real workloads. > Really, changing the virtio spec must have good reasons. > > > > > > > > > > > > > The only advantage of ' inflating the balloon before live > > > > > > > > migration' is simple, > > > > > > > nothing more. > > > > > > > > > > > > > > That's a big advantage. Another one is that it does > > > > > > > something useful in real- world scenarios. > > > > > > > > > > > > > > > > > > > I don't think the heave performance impaction is something > > > > > > useful in real > > > > > world scenarios. > > > > > > > > > > > > Liang > > > > > > > Roman. > > > > > > > > > > So fix the performance then. You will have to try harder if you > > > > > want to convince people that the performance is due to bad > > > > > host/guest interface, and so we have to change *that*. > > > > > > > > > > > > > Actually, the PV solution is irrelevant with the balloon > > > > mechanism, I just use it to transfer information between host and > guest. > > > > I am not sure if I should implement a new virtio device, and I > > > > want to get the answer from the community. > > > > In this RFC patch, to make things simple, I choose to extend the > > > > virtio-balloon and use the extended interface to transfer the > > > > request and > > > free_page_bimap content. > > > > > > > > I am not intend to change the current virtio-balloon implementation. > > > > > > > > Liang > > > > > > And the answer would depend on the answer to my question above. > > > Does balloon need an interface passing page bitmaps around? > > > > Yes, I need a new interface. > > Possibly, but you will need to justify this at some level if you care about > upstreaming your patches. > > > > Does this speed up any operations? > > > > No, a new interface will not speed up anything, but it is the easiest way to > solve the compatibility issue. > > A bunch of new code is often easier to write than to figure out the old one, > but if we keep piling it up we'll end up with an unmaintainable mess. So we > are rather careful about adding new interfaces, and we try to make them > generic sometimes even at cost of slight inefficiencies. > > > > OTOH what if you use the regular balloon interface with your patches? > > > > > > > The regular balloon interfaces have their specific function and I can't use > them in my patches. > > If using these regular interface, I have to do a lot of changes to keep the > compatibility. > > Why can't you? > > What exactly do we need to change? > > If we put things in terms of the balloon, that supports adding and removing > pages. > > Using these terms, let's enumerate: > - a new method (e.g. new virtqueue) that adds and immediately removes > page in a balloon > clearly, you can add then remove using the existing interfaces > is a single command significantly faster than using existing two vqs? > - a new kind of request that says "add (and immediately remove?) as many > pages as you can" > sounds rather benign > - a new kind of message that adds multiple pages using a bitmap > (instead of an address list) > again, is this significantly faster? More of less faster because of less data traffic. I didn't measure this, I will do it and take a deep look at the way you suggest if we choose to make use of the virtio-balloon interface. > > Does not look like compatibility is an issue, to me. > > > At some level, your patches look like page hints. > If we have more patches in mind that use page hints, then a new hint device > might make sense. > Yes, I have ever considered to implement a new device, use the virtio-balloon to transfer the free pages information which is irrelevant with the balloon mechanism is some more or less confusing. > However, people experimented with page hints in the past, so far this always > went nowhere. E.g. I CC Rick who saw some problems when page hints > interact with huge pages. Rick, could you elaborate please? > Thanks a lot. Can't wait to know the problems. Liang > > -- > MST
[toc] | [prev] | [next] | [standalone]
| From | Roman Kagan <rkagan@virtuozzo.com> |
|---|---|
| Date | 2016-03-09 15:30 +0100 |
| Subject | Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <raNhE-DP-3@gated-at.bofh.it> |
| In reply to | #1351559 |
On Mon, Mar 07, 2016 at 01:40:06PM +0200, Michael S. Tsirkin wrote: > On Mon, Mar 07, 2016 at 06:49:19AM +0000, Li, Liang Z wrote: > > > > No. And it's exactly what I mean. The ballooned memory is still > > > > processed during live migration without skipping. The live migration code is > > > in migration/ram.c. > > > > > > So if guest acknowledged VIRTIO_BALLOON_F_MUST_TELL_HOST, we can > > > teach qemu to skip these pages. > > > Want to write a patch to do this? > > > > > > > Yes, we really can teach qemu to skip these pages and it's not hard. > > The problem is the poor performance, this PV solution > > Balloon is always PV. And do not call patches solutions please. > > > is aimed to make it more > > efficient and reduce the performance impact on guest. > > We need to get a bit beyond this. You are making multiple > changes, it seems to make sense to split it all up, and analyse each > change separately. Couldn't agree more. There are three stages in this optimization: 1) choosing which pages to skip 2) communicating them from guest to host 3) skip transferring uninteresting pages to the remote side on migration For (3) there seems to be a low-hanging fruit to amend migration/ram.c:iz_zero_range() to consult /proc/self/pagemap. This would work for guest RAM that hasn't been touched yet or which has been ballooned out. For (1) I've been trying to make a point that skipping clean pages is much more likely to result in noticable benefit than free pages only. As for (2), we do seem to have a problem with the existing balloon: according to your measurements it's very slow; besides, I guess it plays badly with transparent huge pages (as both the guest and the host work with one 4k page at a time). This is a problem for other use cases of balloon (e.g. as a facility for resource management); tackling that appears a more natural application for optimization efforts. Thanks, Roman.
[toc] | [prev] | [next] | [standalone]
| From | "Li, Liang Z" <liang.z.li@intel.com> |
|---|---|
| Date | 2016-03-09 16:30 +0100 |
| Subject | RE: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <raOdI-1eq-11@gated-at.bofh.it> |
| In reply to | #1354167 |
> On Mon, Mar 07, 2016 at 01:40:06PM +0200, Michael S. Tsirkin wrote: > > On Mon, Mar 07, 2016 at 06:49:19AM +0000, Li, Liang Z wrote: > > > > > No. And it's exactly what I mean. The ballooned memory is still > > > > > processed during live migration without skipping. The live > > > > > migration code is > > > > in migration/ram.c. > > > > > > > > So if guest acknowledged VIRTIO_BALLOON_F_MUST_TELL_HOST, we > can > > > > teach qemu to skip these pages. > > > > Want to write a patch to do this? > > > > > > > > > > Yes, we really can teach qemu to skip these pages and it's not hard. > > > The problem is the poor performance, this PV solution > > > > Balloon is always PV. And do not call patches solutions please. > > > > > is aimed to make it more > > > efficient and reduce the performance impact on guest. > > > > We need to get a bit beyond this. You are making multiple changes, it > > seems to make sense to split it all up, and analyse each change > > separately. > > Couldn't agree more. > > There are three stages in this optimization: > > 1) choosing which pages to skip > > 2) communicating them from guest to host > > 3) skip transferring uninteresting pages to the remote side on migration > > For (3) there seems to be a low-hanging fruit to amend > migration/ram.c:iz_zero_range() to consult /proc/self/pagemap. This would > work for guest RAM that hasn't been touched yet or which has been > ballooned out. > > For (1) I've been trying to make a point that skipping clean pages is much > more likely to result in noticable benefit than free pages only. > I am considering to drop the pagecache before getting the free pages. > As for (2), we do seem to have a problem with the existing balloon: > according to your measurements it's very slow; besides, I guess it plays badly I didn't say communicating is slow. Even this is very slow, my solution use bitmap instead of PFNs, there is fewer data traffic, so it's faster than the existing balloon which use PFNs. > with transparent huge pages (as both the guest and the host work with one > 4k page at a time). This is a problem for other use cases of balloon (e.g. as a > facility for resource management); tackling that appears a more natural > application for optimization efforts. > > Thanks, > Roman.
[toc] | [prev] | [next] | [standalone]
| From | "Michael S. Tsirkin" <mst@redhat.com> |
|---|---|
| Date | 2016-03-09 16:40 +0100 |
| Subject | Re: [Qemu-devel] [RFC qemu 0/4] A PV solution for live migration optimization |
| Message-ID | <raOno-1of-17@gated-at.bofh.it> |
| In reply to | #1354214 |
On Wed, Mar 09, 2016 at 03:27:54PM +0000, Li, Liang Z wrote: > > On Mon, Mar 07, 2016 at 01:40:06PM +0200, Michael S. Tsirkin wrote: > > > On Mon, Mar 07, 2016 at 06:49:19AM +0000, Li, Liang Z wrote: > > > > > > No. And it's exactly what I mean. The ballooned memory is still > > > > > > processed during live migration without skipping. The live > > > > > > migration code is > > > > > in migration/ram.c. > > > > > > > > > > So if guest acknowledged VIRTIO_BALLOON_F_MUST_TELL_HOST, we > > can > > > > > teach qemu to skip these pages. > > > > > Want to write a patch to do this? > > > > > > > > > > > > > Yes, we really can teach qemu to skip these pages and it's not hard. > > > > The problem is the poor performance, this PV solution > > > > > > Balloon is always PV. And do not call patches solutions please. > > > > > > > is aimed to make it more > > > > efficient and reduce the performance impact on guest. > > > > > > We need to get a bit beyond this. You are making multiple changes, it > > > seems to make sense to split it all up, and analyse each change > > > separately. > > > > Couldn't agree more. > > > > There are three stages in this optimization: > > > > 1) choosing which pages to skip > > > > 2) communicating them from guest to host > > > > 3) skip transferring uninteresting pages to the remote side on migration > > > > For (3) there seems to be a low-hanging fruit to amend > > migration/ram.c:iz_zero_range() to consult /proc/self/pagemap. This would > > work for guest RAM that hasn't been touched yet or which has been > > ballooned out. > > > > For (1) I've been trying to make a point that skipping clean pages is much > > more likely to result in noticable benefit than free pages only. > > > > I am considering to drop the pagecache before getting the free pages. > > > As for (2), we do seem to have a problem with the existing balloon: > > according to your measurements it's very slow; besides, I guess it plays badly > > I didn't say communicating is slow. Even this is very slow, my solution use bitmap instead of > PFNs, there is fewer data traffic, so it's faster than the existing balloon which use PFNs. By how much? > > with transparent huge pages (as both the guest and the host work with one > > 4k page at a time). This is a problem for other use cases of balloon (e.g. as a > > facility for resource management); tackling that appears a more natural > > application for optimization efforts. > > > > Thanks, > > Roman.
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | linux.kernel
csiph-web