Commit 79c168e2aa95 for kernel

commit 79c168e2aa957ed0054c93db17caba653b95e377
Author: Lukas Wunner <lukas@wunner.de>
Date:   Thu Oct 8 14:26:00 2026 +0200

    PCI/AER: Skip error recovery on false alarms

    Alex is seeing a probe failure of the amdgpu driver after the Root Port
    above an AMD Navi10 GPU has been reset.  The reset was performed to recover
    from a Firmware First reported Fatal Error.

    However all status registers in the Root Port's AER Extended Capability are
    blank, so apparently the platform firmware raised a false alarm.

    The issue is only occurring since commit eddba19b8b5f ("PCI/AER: Support
    Advisory Non-Fatal Errors").  It looks like enabling Advisory Non-Fatal
    Errors causes code paths to be exercised in platform firmware which were
    never validated before.

    Skip error recovery on false alarms, i.e. if no unmasked errors were
    actually signaled.

    Note that this will also skip recovery if both the Status and Mask
    registers are "all ones", as would be the case for inaccessible devices.
    However that seems justified because it would imply either a hot-unplug
    event or a Surprise Down Error further up in the hierarchy.  Interfering
    with recovery from that seems uncalled for.

    Fixes: eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal Errors")
    Reported-by: Alex Deucher <alexander.deucher@amd.com>
    Closes: https://bugzilla.kernel.org/show_bug.cgi?id=222095
    Signed-off-by: Lukas Wunner <lukas@wunner.de>
    Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
    Tested-by: Alex Deucher <alexander.deucher@amd.com>
    Acked-by: Alex Deucher <alexander.deucher@amd.com>
    Link: https://patch.msgid.link/0552ed277e40a288e0157af799257ee6ec722534.1791460615.git.lukas@wunner.de

diff --git a/drivers/pci/pcie/aer.c b/drivers/pci/pcie/aer.c
index d8dcd238fda1..922a726a52a5 100644
--- a/drivers/pci/pcie/aer.c
+++ b/drivers/pci/pcie/aer.c
@@ -1360,8 +1360,10 @@ static DEFINE_KFIFO(aer_recover_ring, struct aer_recover_entry,

 static void aer_recover_work_func(struct work_struct *work)
 {
+	struct aer_capability_regs *regs;
 	struct aer_recover_entry entry;
 	struct pci_dev *pdev;
+	u32 err;

 	while (kfifo_get(&aer_recover_ring, &entry)) {
 		pdev = pci_get_domain_bus_and_slot(entry.domain, entry.bus,
@@ -1375,6 +1377,12 @@ static void aer_recover_work_func(struct work_struct *work)
 		}
 		pci_print_aer(pdev, entry.severity, entry.regs);

+		regs = entry.regs;
+		if (entry.severity == AER_CORRECTABLE)
+			err = regs->cor_status & ~regs->cor_mask;
+		else
+			err = regs->uncor_status & ~regs->uncor_mask;
+
 		/*
 		 * Memory for aer_capability_regs(entry.regs) is being
 		 * allocated from the ghes_estatus_pool to protect it from
@@ -1385,12 +1393,14 @@ static void aer_recover_work_func(struct work_struct *work)
 		ghes_estatus_pool_region_free((unsigned long)entry.regs,
 					    sizeof(struct aer_capability_regs));

-		if (entry.severity == AER_NONFATAL)
-			pcie_do_recovery(pdev, pci_channel_io_normal,
-					 aer_root_reset);
-		else if (entry.severity == AER_FATAL)
-			pcie_do_recovery(pdev, pci_channel_io_frozen,
-					 aer_root_reset);
+		if (err) {
+			if (entry.severity == AER_NONFATAL)
+				pcie_do_recovery(pdev, pci_channel_io_normal,
+						 aer_root_reset);
+			else if (entry.severity == AER_FATAL)
+				pcie_do_recovery(pdev, pci_channel_io_frozen,
+						 aer_root_reset);
+		}
 		pci_dev_put(pdev);
 	}
 }