[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]
[PATCH v2 10/25] migration: Enlarge postcopy recovery to capture !-EIO t
From: |
Peter Xu |
Subject: |
[PATCH v2 10/25] migration: Enlarge postcopy recovery to capture !-EIO too |
Date: |
Tue, 1 Mar 2022 16:39:10 +0800 |
We used to have quite a few places making sure -EIO happened and that's the
only way to trigger postcopy recovery. That's based on the assumption that
we'll only return -EIO for channel issues.
It'll work in 99.99% cases but logically that won't cover some corner cases.
One example is e.g. ram_block_from_stream() could fail with an interrupted
network, then -EINVAL will be returned instead of -EIO.
I remembered Dave Gilbert pointed that out before, but somehow this is
overlooked. Neither did I encounter anything outside the -EIO error.
However we'd better touch that up before it triggers a rare VM data loss during
live migrating.
To cover as much those cases as possible, remove the -EIO restriction on
triggering the postcopy recovery, because even if it's not a channel failure,
we can't do anything better than halting QEMU anyway - the corpse of the
process may even be used by a good hand to dig out useful memory regions, or
the admin could simply kill the process later on.
Reviewed-by: Dr. David Alan Gilbert <dgilbert@redhat.com>
Signed-off-by: Peter Xu <peterx@redhat.com>
---
migration/migration.c | 4 ++--
migration/postcopy-ram.c | 2 +-
2 files changed, 3 insertions(+), 3 deletions(-)
diff --git a/migration/migration.c b/migration/migration.c
index 6e4cc9cc87..67520d3105 100644
--- a/migration/migration.c
+++ b/migration/migration.c
@@ -2877,7 +2877,7 @@ retry:
out:
res = qemu_file_get_error(rp);
if (res) {
- if (res == -EIO && migration_in_postcopy()) {
+ if (res && migration_in_postcopy()) {
/*
* Maybe there is something we can do: it looks like a
* network down issue, and we pause for a recovery.
@@ -3478,7 +3478,7 @@ static MigThrError migration_detect_error(MigrationState
*s)
error_free(local_error);
}
- if (state == MIGRATION_STATUS_POSTCOPY_ACTIVE && ret == -EIO) {
+ if (state == MIGRATION_STATUS_POSTCOPY_ACTIVE && ret) {
/*
* For postcopy, we allow the network to be down for a
* while. After that, it can be continued by a
diff --git a/migration/postcopy-ram.c b/migration/postcopy-ram.c
index d08d396c63..b0d12d5053 100644
--- a/migration/postcopy-ram.c
+++ b/migration/postcopy-ram.c
@@ -1039,7 +1039,7 @@ retry:
msg.arg.pagefault.address);
if (ret) {
/* May be network failure, try to wait for recovery */
- if (ret == -EIO && postcopy_pause_fault_thread(mis)) {
+ if (postcopy_pause_fault_thread(mis)) {
/* We got reconnected somehow, try to continue */
goto retry;
} else {
--
2.32.0
- [PATCH v2 00/25] migration: Postcopy Preemption, Peter Xu, 2022/03/01
- [PATCH v2 01/25] migration: Dump sub-cmd name in loadvm_process_command tp, Peter Xu, 2022/03/01
- [PATCH v2 02/25] migration: Finer grained tracepoints for POSTCOPY_LISTEN, Peter Xu, 2022/03/01
- [PATCH v2 03/25] migration: Tracepoint change in postcopy-run bottom half, Peter Xu, 2022/03/01
- [PATCH v2 04/25] migration: Introduce postcopy channels on dest node, Peter Xu, 2022/03/01
- [PATCH v2 05/25] migration: Dump ramblock and offset too when non-same-page detected, Peter Xu, 2022/03/01
- [PATCH v2 06/25] migration: Add postcopy_thread_create(), Peter Xu, 2022/03/01
- [PATCH v2 07/25] migration: Move static var in ram_block_from_stream() into global, Peter Xu, 2022/03/01
- [PATCH v2 08/25] migration: Add pss.postcopy_requested status, Peter Xu, 2022/03/01
- [PATCH v2 09/25] migration: Move migrate_allow_multifd and helpers into migration.c, Peter Xu, 2022/03/01
- [PATCH v2 10/25] migration: Enlarge postcopy recovery to capture !-EIO too,
Peter Xu <=
- [PATCH v2 11/25] migration: postcopy_pause_fault_thread() never fails, Peter Xu, 2022/03/01
- [PATCH v2 14/25] migration: Add migration_incoming_transport_cleanup(), Peter Xu, 2022/03/01
- [PATCH v2 15/25] migration: Allow migrate-recover to run multiple times, Peter Xu, 2022/03/01
- [PATCH v2 13/25] migration: Move channel setup out of postcopy_try_recover(), Peter Xu, 2022/03/01
- [PATCH v2 16/25] migration: Add postcopy-preempt capability, Peter Xu, 2022/03/01
- [PATCH v2 17/25] migration: Postcopy preemption preparation on channel creation, Peter Xu, 2022/03/01
- [PATCH v2 18/25] migration: Postcopy preemption enablement, Peter Xu, 2022/03/01
- [PATCH v2 19/25] migration: Postcopy recover with preempt enabled, Peter Xu, 2022/03/01
- [PATCH v2 12/25] migration: Export ram_load_postcopy(), Peter Xu, 2022/03/01
- [PATCH v2 21/25] migration: Parameter x-postcopy-preempt-break-huge, Peter Xu, 2022/03/01