[PULL 14/18] migration: Enlarge postcopy recovery to capture !-EIO too

qemu-devel

[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

[PULL 14/18] migration: Enlarge postcopy recovery to capture !-EIO too

From:	Dr. David Alan Gilbert (git)
Subject:	[PULL 14/18] migration: Enlarge postcopy recovery to capture !-EIO too
Date:	Wed, 2 Mar 2022 18:29:32 +0000

From: Peter Xu <peterx@redhat.com>

We used to have quite a few places making sure -EIO happened and that's the
only way to trigger postcopy recovery.  That's based on the assumption that
we'll only return -EIO for channel issues.

It'll work in 99.99% cases but logically that won't cover some corner cases.
One example is e.g. ram_block_from_stream() could fail with an interrupted
network, then -EINVAL will be returned instead of -EIO.

I remembered Dave Gilbert pointed that out before, but somehow this is
overlooked.  Neither did I encounter anything outside the -EIO error.

However we'd better touch that up before it triggers a rare VM data loss during
live migrating.

To cover as much those cases as possible, remove the -EIO restriction on
triggering the postcopy recovery, because even if it's not a channel failure,
we can't do anything better than halting QEMU anyway - the corpse of the
process may even be used by a good hand to dig out useful memory regions, or
the admin could simply kill the process later on.

Reviewed-by: Dr. David Alan Gilbert <dgilbert@redhat.com>
Signed-off-by: Peter Xu <peterx@redhat.com>
Message-Id: <20220301083925.33483-11-peterx@redhat.com>
Signed-off-by: Dr. David Alan Gilbert <dgilbert@redhat.com>
---
 migration/migration.c    | 4 ++--
 migration/postcopy-ram.c | 2 +-
 2 files changed, 3 insertions(+), 3 deletions(-)

diff --git a/migration/migration.c b/migration/migration.c
index bcc385b94b..306e2ac60e 100644
--- a/migration/migration.c
+++ b/migration/migration.c
@@ -2865,7 +2865,7 @@ retry:
 out:
     res = qemu_file_get_error(rp);
     if (res) {
-        if (res == -EIO && migration_in_postcopy()) {
+        if (res && migration_in_postcopy()) {
             /*
              * Maybe there is something we can do: it looks like a
              * network down issue, and we pause for a recovery.
@@ -3466,7 +3466,7 @@ static MigThrError migration_detect_error(MigrationState 
*s)
         error_free(local_error);
     }
 
-    if (state == MIGRATION_STATUS_POSTCOPY_ACTIVE && ret == -EIO) {
+    if (state == MIGRATION_STATUS_POSTCOPY_ACTIVE && ret) {
         /*
          * For postcopy, we allow the network to be down for a
          * while. After that, it can be continued by a
diff --git a/migration/postcopy-ram.c b/migration/postcopy-ram.c
index d08d396c63..b0d12d5053 100644
--- a/migration/postcopy-ram.c
+++ b/migration/postcopy-ram.c
@@ -1039,7 +1039,7 @@ retry:
                                         msg.arg.pagefault.address);
             if (ret) {
                 /* May be network failure, try to wait for recovery */
-                if (ret == -EIO && postcopy_pause_fault_thread(mis)) {
+                if (postcopy_pause_fault_thread(mis)) {
                     /* We got reconnected somehow, try to continue */
                     goto retry;
                 } else {
-- 
2.35.1

[Prev in Thread]

Current Thread

[Next in Thread]

[PULL 01/18] clock-vmstate: Add missing END_OF_LIST, (continued)
- [PULL 01/18] clock-vmstate: Add missing END_OF_LIST, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 03/18] monitor/hmp: add support for flag argument with value, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 04/18] qapi/monitor: refactor set/expire_password with enums, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 02/18] virtiofsd: Let meson check for statx.stx_mnt_id, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 08/18] migration: Finer grained tracepoints for POSTCOPY_LISTEN, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 07/18] migration: Dump sub-cmd name in loadvm_process_command tp, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 05/18] qapi/monitor: allow VNC display id in set/expire_password, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 06/18] migration/rdma: set the REUSEADDR option for destination, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 09/18] migration: Tracepoint change in postcopy-run bottom half, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 11/18] migration: Dump ramblock and offset too when non-same-page detected, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 14/18] migration: Enlarge postcopy recovery to capture !-EIO too, Dr. David Alan Gilbert (git) <=
- [PULL 10/18] migration: Introduce postcopy channels on dest node, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 13/18] migration: Move static var in ram_block_from_stream() into global, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 12/18] migration: Add postcopy_thread_create(), Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 15/18] migration: postcopy_pause_fault_thread() never fails, Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 16/18] migration: Add migration_incoming_transport_cleanup(), Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 17/18] tests: Pass in MigrateStart** into test_migrate_start(), Dr. David Alan Gilbert (git), 2022/03/02
- [PULL 18/18] migration: Remove load_state_old and minimum_version_id_old, Dr. David Alan Gilbert (git), 2022/03/02
- Re: [PULL 00/18] migration queue, Peter Maydell, 2022/03/03
  - Re: [PULL 00/18] migration queue, Philippe Mathieu-Daudé, 2022/03/08
    - Re: [PULL 00/18] migration queue, Dr. David Alan Gilbert, 2022/03/08

Prev by Date: [PULL 11/18] migration: Dump ramblock and offset too when non-same-page detected
Next by Date: [PULL 10/18] migration: Introduce postcopy channels on dest node
Previous by thread: [PULL 11/18] migration: Dump ramblock and offset too when non-same-page detected
Next by thread: [PULL 10/18] migration: Introduce postcopy channels on dest node
Index(es):
- Date
- Thread