Commit 0bd5a1b3d6 for qemu.org
commit 0bd5a1b3d60cca595e19bb73eec83307259be671
Author: Kevin Wolf <kwolf@redhat.com>
Date: Fri Oct 2 15:21:34 2026 +0200
block: Fix corruption from broken RMW atomicity
When emulating 512 byte sector devices on top of 4k native host storage,
QEMU performs a read-modify-write cycle for misaligned write requests.
These requests are marked as serialising so that any other writes to the
same 4k block will have to wait and can't overwrite the block in the
middle, which would make the RMW write back stale data.
In principle, this mechanism works fine, but there is a race condition
in the implementation: bdrv_find_conflicting_request() assumes that
req->waiting_for == NULL means that the request is the active owner of
the data blocks it touches and everyone else has to wait for this
request. However, there is a window between tracked_request_begin() and
bdrv_wait_serialising_requests_locked() where any new request has
req->waiting_for == NULL even if it conflicts with an already running
serialising request. So at this point, we have seemingly two active
owners at the same time until the new request checks for conflicts.
For a third request coming in, this isn't a problem, it just might wait
for a little longer than necessary. However, if the original RMW request
calls bdrv_wait_serialising_requests_locked() again (which it does in
bdrv_aligned_pwritev() -> bdrv_co_write_req_prepare() before writing
back the RMW buffer), it will now wait on the other request that snuck
in, which is fatal and makes the RMW buffer stale. This scenario causes
data corruption.
Fix this by checking for conflicts already in tracked_request_begin()
while holding the bs->reqs_lock mutex, so that the race window is closed
and req->waiting_for == NULL really has the intended meaning.
For good measure, add an active_owner boolean that is set when a request
is marked serialising and has completed waiting for other requests, and
assert that it is not set when waiting for another request, so that in
the case of another bug, we'd crash instead of causing data corruption.
Cc: qemu-stable@nongnu.org
Buglink: https://redhat.atlassian.net/browse/RHEL-272836
Signed-off-by: Kevin Wolf <kwolf@redhat.com>
Message-ID: <20261002132135.93345-2-kwolf@redhat.com>
Signed-off-by: Kevin Wolf <kwolf@redhat.com>
diff --git a/block/io.c b/block/io.c
index a916b236c3..bd562d472f 100644
--- a/block/io.c
+++ b/block/io.c
@@ -50,6 +50,9 @@
static int coroutine_fn bdrv_co_do_pwrite_zeroes(BlockDriverState *bs,
int64_t offset, int64_t bytes, BdrvRequestFlags flags);
+static void coroutine_fn
+bdrv_wait_serialising_requests_locked(BdrvTrackedRequest *self);
+
static void GRAPH_RDLOCK
bdrv_parent_drained_begin(BlockDriverState *bs, BdrvChild *ignore)
{
@@ -616,6 +619,7 @@ static void coroutine_fn tracked_request_begin(BdrvTrackedRequest *req,
.type = type,
.co = qemu_coroutine_self(),
.serialising = false,
+ .active_owner = false,
.overlap_offset = offset,
.overlap_bytes = bytes,
};
@@ -624,6 +628,15 @@ static void coroutine_fn tracked_request_begin(BdrvTrackedRequest *req,
qemu_mutex_lock(&bs->reqs_lock);
QLIST_INSERT_HEAD(&bs->tracked_requests, req, list);
+
+ /*
+ * Set .waiting_for while we're holding the lock, otherwise we'd open a race
+ * window where another request in the middle of its RMW cycle starts
+ * waiting for us, leading to potential data corruption if this request
+ * overwrites a block the RMW request has already read.
+ */
+ bdrv_wait_serialising_requests_locked(req);
+
qemu_mutex_unlock(&bs->reqs_lock);
}
@@ -684,6 +697,15 @@ bdrv_wait_serialising_requests_locked(BdrvTrackedRequest *self)
BdrvTrackedRequest *req;
while ((req = bdrv_find_conflicting_request(self))) {
+ /*
+ * If this request is the active owner of the data blocks it accesses,
+ * no other conflicting request may exist; other requests have to wait
+ * for this one, not the other way around. If this condition is
+ * violated, RMW operations may be interrupted in the middle, operate on
+ * stale data and introduce data corruption.
+ */
+ assert(!self->active_owner);
+
self->waiting_for = req;
qemu_co_queue_wait(&req->wait_queue, &self->bs->reqs_lock);
self->waiting_for = NULL;
@@ -803,6 +825,7 @@ void coroutine_fn bdrv_make_request_serialising(BdrvTrackedRequest *req,
tracked_request_set_serialising(req, align);
bdrv_wait_serialising_requests_locked(req);
+ req->active_owner = true;
qemu_mutex_unlock(&req->bs->reqs_lock);
}
diff --git a/include/block/block_int-common.h b/include/block/block_int-common.h
index 147c08155f..46428b034d 100644
--- a/include/block/block_int-common.h
+++ b/include/block/block_int-common.h
@@ -80,6 +80,7 @@ typedef struct BdrvTrackedRequest {
enum BdrvTrackedRequestType type;
bool serialising;
+ bool active_owner;
int64_t overlap_offset;
int64_t overlap_bytes;