Skip to content

QCLINUX: arm64: configs: Enable GPU virtualization - #1118

Open
Shivam Rawat (shivrawa) wants to merge 1 commit into
qualcomm-linux:qcom-6.18.yfrom
shivrawa:qcom-6.18.y
Open

Shivam Rawat (shivrawa) wants to merge 1 commit into
qualcomm-linux:qcom-6.18.yfrom
shivrawa:qcom-6.18.y

Conversation

@shivrawa

@shivrawa Shivam Rawat (shivrawa) commented Sep 15, 2026

Copy link
Copy Markdown

Enable DRM VirtIO GPU configs to support GPU virtualization.

Target milestone: qli-2.1 pull-request freeze

CRs-fixed: 4677030

Enable DRM VirtIO GPU configs to support GPU virtualization.

Signed-off-by: Shivam Rawat <shivrawa@qti.qualcomm.com>
@qswat-orbit-external

Copy link
Copy Markdown

Merge Check Failed: No CR Numbers Found

Error: No Change Request numbers were found.

Please add Change Request numbers to your pull request description in the format CRs-Fixed: 12345 or link GitHub issues that are associated with Change Requests.

@qswat-orbit-external

Copy link
Copy Markdown

Merge Check Failed: CR Not Eligible for Merge

CR 4677030 is not eligible for merge.

The parent software image for kernel.qli.2.0 is not development complete.

Entity: kernel.qli.2.0
CR: 4677030
Reason: CR_CANNOT_MERGE

Please ensure the CR passes both CCT (ComponentChangeTasks) and ICT (Integration Change Tasks) validations.

1 similar comment
@qswat-orbit-external

Copy link
Copy Markdown

Merge Check Failed: CR Not Eligible for Merge

CR 4677030 is not eligible for merge.

The parent software image for kernel.qli.2.0 is not development complete.

Entity: kernel.qli.2.0
CR: 4677030
Reason: CR_CANNOT_MERGE

Please ensure the CR passes both CCT (ComponentChangeTasks) and ICT (Integration Change Tasks) validations.

@shivrawa

Copy link
Copy Markdown
Author

Shivendra Pratap (@quicAspratap) could you please review

@qcomlnxci

Copy link
Copy Markdown

Test Matrix

Test Case hamoa-iot-evk-multimedia lemans-evk-multimedia monaco-evk-multimedia purwa-iot-evk-multimedia qcs615-ride-multimedia qcs6490-rb3gen2-multimedia qcs8300-ride-multimedia qcs9100-ride-r3-multimedia shikra-iqs-evk-multimedia
Audio_Card_Registration ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip ✅ Pass ⚠️ skip ⚠️ skip ⚠️ skip
BT_FW_KMD_Service ✅ Pass ✅ Pass ✅ Pass ✅ Pass ❌ Fail ✅ Pass ✅ Pass ✅ Pass ✅ Pass
BT_ON_OFF ✅ Pass ✅ Pass ✅ Pass ✅ Pass ❌ Fail ✅ Pass ✅ Pass ✅ Pass ✅ Pass
BT_SCAN ✅ Pass ✅ Pass ✅ Pass ✅ Pass ❌ Fail ✅ Pass ✅ Pass ✅ Pass ❌ Fail
CPUFreq_Validation ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
CPU_affinity ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
DSP_AudioPD ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
Ethernet_Basic_Validation ⚠️ skip ✅ Pass ✅ Pass ⚠️ skip ⚠️ skip ⚠️ skip ❌ Fail ❌ Fail ⚠️ skip
Freq_Scaling ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
GIC ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ❌ Fail
IPA ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
Interrupts ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
KVM_Driver ❌ Fail ✅ Pass ✅ Pass ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail
KVM_EL2_DTB ❌ Fail ✅ Pass ✅ Pass ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail
KVM_Infra ❌ Fail ✅ Pass ✅ Pass ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail
OpenCV ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
PCIe ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
Probe_Failure_Check ❌ Fail ❌ Fail ❌ Fail ❌ Fail ✅ Pass ❌ Fail ❌ Fail ❌ Fail ❌ Fail
RMNET ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
UFS_Validation ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
USBHost ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ❌ Fail ❌ Fail ❌ Fail ❌ Fail
WiFi_Firmware_Driver ✅ Pass ✅ Pass ❌ Fail ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
WiFi_OnOff ✅ Pass ✅ Pass ❌ Fail ✅ Pass ⚠️ skip ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
adsp_remoteproc ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
cdsp_remoteproc ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
gpdsp_remoteproc ⚠️ skip ✅ Pass ✅ Pass ⚠️ skip ⚠️ skip ⚠️ skip ✅ Pass ✅ Pass ⚠️ skip
hotplug ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
irq ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
kaslr ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
pinctrl ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
qcom_hwrng ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️
rngtest ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
shmbridge ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
smmu ❌ Fail ❌ Fail ✅ Pass ❌ Fail ❌ Fail ✅ Pass ✅ Pass ❌ Fail ✅ Pass
watchdog ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
wpss_remoteproc ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass

@shivrawa

Copy link
Copy Markdown
Author

Komal Bajaj (@Komal-Bajaj) could you please check and help why these checkers are failing. I don't see the proper error log related to the code changes. If not, can you please override them and merge

@qlijarvis

Copy link
Copy Markdown

LAVA Failed Case Triage Summary

PR: #1118

Job 227677 | SoC qcs615-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/227677

Failed test cases in LAVA job 227677 (SoC: qcs615-ride).

  Case 1: smmu
  1. Failed case: smmu
  2. Root cause: Test expects video-decoder and video-encoder child devices to appear as separate entries in /sys/kernel/iommu_groups/, but the Venus driver with "non legacy binding" on qcs615-ride only registers the parent device (aa00000.video-codec) to IOMMU group 7. The child video-decoder and video-encoder devices are logical V4L2 video nodes, not separate platform devices with independent IOMMU group attachments.
  3. Possible fix: Update the smmu test script to recognize that for Venus video codec with non-legacy binding, the parent device (aa00000.video-codec) IOMMU group attachment is sufficient and child video-decoder/encoder nodes do not require separate IOMMU group entries. Alternatively, verify this is the expected behavior for qcs615 and adjust test expectations accordingly.
  4. Detail analysis attachment: failed_case_job227677_1_detailed.md
  Case 2: BT_FW_KMD_Service
  1. Failed case: BT_FW_KMD_Service
  2. Root cause: WCN6855 Bluetooth firmware initialization failure on qcs615-ride — UART communication timeout (-ETIMEDOUT, errno 110) when attempting to read QCA version information during hci0 setup; the Bluetooth controller never responds to HCI commands (command 0xfc00 tx timeout), resulting in invalid BD address (00:00:00:00:00:00) and non-functional Bluetooth stack despite driver and service being loaded.
  3. Possible fix: This is a pre-existing hardware/firmware/platform issue unrelated to the PR (which only enables GPU virtualization configs). Verify: (1) WCN6855 firmware files are present and correct version for qcs615-ride, (2) UART pinctrl/clocks/regulators for the Bluetooth UART (serial@a8c000) are correctly configured in device tree, (3) WCN6855 power sequencing (BT_EN GPIO, regulators) is correct, (4) check for known qcs615-ride board-specific Bluetooth errata or firmware compatibility issues. If this is a known board issue, document as expected failure for qcs615-ride until resolved.
  4. Detail analysis attachment: failed_case_job227677_2_detailed.md
  Case 3: BT_ON_OFF
  1. Failed case: BT_ON_OFF
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: This is a pre-existing platform/infrastructure issue, not a PR regression. Recommended actions: (1) Verify WCN6855 Bluetooth firmware is present and correct version for qcs615-ride in the test image (/lib/firmware/qca/). (2) Check UART hardware path: verify Bluetooth UART device tree node is correct, UART controller is functional, and UART pins are properly configured. (3) Check for known WCN6855 Bluetooth errata on qcs615 platform. (4) If this is a known intermittent issue on qcs615-ride lab boards, consider marking BT tests as expected-fail for this platform until hardware/firmware is fixed. (5) Do NOT block PR QCLINUX: arm64: configs: Enable GPU virtualization #1118 merge based on this failure — it is unrelated to the GPU virtualization config changes in the PR.
  4. Detail analysis attachment: failed_case_job227677_3_detailed.md
  Case 4: BT_SCAN — Bluetooth Hardware Communication Timeout
  1. Failed case: BT_SCAN — Bluetooth Hardware Communication Timeout
  2. Root cause: Bluetooth driver (btqca) cannot communicate with WCN6855 hardware on qcs615-ride; all HCI commands time out with -ETIMEDOUT (-110). UART serial communication path between AP and BT controller is non-functional. Repeated pattern: "Bluetooth: hci0: command tx timeout" → "Bluetooth: hci0: Reading QCA version information failed (-110)". This is a pre-existing board/infrastructure issue, not a PR-introduced regression (PR only enables GPU virtualization configs with no Bluetooth dependency).
  3. Possible fix: Verify qcs615-ride board hardware: (1) Check UART GPIO/pinctrl configuration in device tree for BT UART pins; (2) Verify BT power regulators are enabled and stable; (3) Check UART clock configuration; (4) Verify BT firmware files are present and correct version; (5) Test with a known-good qcs615-ride board to isolate hardware vs software issue. If board is confirmed functional, bisect kernel to find when BT communication broke.
  4. Detail analysis attachment: failed_case_job227677_4_detailed.md
  Case 5: KVM_Driver — Platform Limitation (HYP Mode Not Available)
  1. Failed case: KVM_Driver — Platform Limitation (HYP Mode Not Available)
  2. Root cause: QCS615 platform does not support ARM EL2 (Hypervisor mode), causing KVM driver initialization to fail with "HYP mode not available" at boot, which prevents /dev/kvm device node creation required by the test.
  3. Possible fix: Mark KVM_Driver, KVM_EL2_DTB, and KVM_Infra tests as "skip" or "not applicable" for qcs615-ride platform in the LAVA test suite configuration, as this SoC does not have hardware virtualization support (no EL2/HYP mode).
  4. Detail analysis attachment: failed_case_job227677_5_detailed.md
  Case 6: KVM_EL2_DTB — KVM Hypervisor Mode Unavailable
  1. Failed case: KVM_EL2_DTB — KVM Hypervisor Mode Unavailable
  2. Root cause: The qcs615-ride platform is running under the Gunyah hypervisor (version gunyah-cdfb73831), which occupies EL2 (hypervisor mode). KVM requires exclusive access to EL2 to function, but cannot access it when another hypervisor is already running at that privilege level. The kernel message kvm [1]: HYP mode not available at boot time confirms KVM detected this condition and disabled itself, preventing /dev/kvm device node creation.
  3. Possible fix: This is not a PR-introduced regression (the PR only enables GPU virtualization configs). This is a platform configuration issue. To enable KVM on qcs615-ride: (1) boot without the Gunyah hypervisor, or (2) use nested virtualization if Gunyah supports it, or (3) run KVM tests on a platform that boots Linux directly at EL1 without a hypervisor. For CI: either skip KVM tests on Gunyah-enabled platforms or add a platform-specific test gate that checks for /dev/kvm availability before running KVM tests.
  4. Detail analysis attachment: failed_case_job227677_6_detailed.md
  Case 7: KVM_Infra — Platform Hardware/Firmware Limitation
  1. Failed case: KVM_Infra — Platform Hardware/Firmware Limitation
  2. Root cause: QCS615 Ride platform does not provide EL2 (ARM Hypervisor mode) access to the Linux kernel. During boot, KVM ARM driver detects "HYP mode not available" and aborts initialization, preventing /dev/kvm device creation. This is a platform/firmware configuration issue specific to QCS615, not a kernel bug or PR regression.
  3. Possible fix: Add platform-specific test skip rule in LAVA CI: if [ "$DEVICE_TYPE" = "qcs615-ride" ]; then skip_test KVM_Infra; fi. QCS615 Ride does not support KVM because firmware boots Linux at EL1 (likely with Gunyah hypervisor at EL2 for automotive use cases). If KVM is required, contact Qualcomm platform team to request firmware update to expose EL2 to Linux or enable VHE.
  4. Detail analysis attachment: failed_case_job227677_7_detailed.md
  Case 8: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM is not available on qcs615-ride because the platform does not support HYP (EL2) mode - kernel message "kvm [1]: HYP mode not available" indicates the hardware/firmware does not provide virtualization extensions or EL2 is disabled.
  3. Possible fix: This is a pre-existing platform limitation, not a PR-introduced regression. The PR enables GPU virtualization configs (CONFIG_DRM_VIRTIO_GPU) which are guest-side drivers for use inside VMs, but qcs615-ride does not support hosting VMs. Either: (1) skip KVM tests on qcs615-ride in CI, or (2) run KVM tests only on platforms with HYP mode support (e.g., SA8775P, SM8550).
  4. Detail analysis attachment: failed_case_job227677_8_detailed.md
Job 227678 | SoC qcs9100-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/227678

Failed test cases in LAVA job 227678 (SoC: qcs9100-ride).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Pre-existing platform issues unrelated to PR changes — four PMIC temp-alarm devices remain in deferred probe state (missing thermal zone dependency), regulatory.db firmware file absent from rootfs (benign cfg80211 warning), and Aquantia AQR115C Ethernet PHY probe fails due to missing firmware-name DT property (error -22 / -EINVAL). The PR only enables GPU virtualization config options and does not touch SPMI, thermal, networking, or firmware paths.
  3. Possible fix: These are known platform configuration gaps on qcs9100-ride, not regressions. For temp-alarm: ensure qcom-spmi-temp-alarm driver and thermal zone bindings are present in DT and driver is built-in or loaded early. For regulatory.db: install wireless-regdb package or suppress the warning (non-critical). For Aquantia PHY: add firmware-name property to stmmac-0:08 DT node or provide the required firmware file. No action required on this PR.
  4. Detail analysis attachment: failed_case_job227678_1_detailed.md
  Case 2: smmu
  1. Failed case: smmu
  2. Root cause: Video codec device aa00000.video-codec exists in device tree but driver did not probe or bind to IOMMU — device is missing IOMMU group attachment, causing SMMU test validation to fail on qcs9100-ride platform.
  3. Possible fix: Investigate why the video codec driver at aa00000.video-codec is not probing or binding to IOMMU on qcs9100-ride — check device tree iommus property, verify driver is built and loaded, check dmesg for probe failures; if this is expected platform behavior (driver not enabled for this SoC), update the SMMU test's critical master list to exclude aa00000.video-codec for qcs9100-ride.
  4. Detail analysis attachment: failed_case_job227678_2_detailed.md
  Case 3: USBHost (Test Infrastructure Issue — No External USB Devices Connected)
  1. Failed case: USBHost (Test Infrastructure Issue — No External USB Devices Connected)
  2. Root cause: Test expects external USB devices (keyboard, mouse, flash drive, etc.) to be physically connected to the board's USB host ports, but the qcs9100-ride board in the LAVA lab has no external devices connected. USB host controllers initialized successfully and root hubs enumerated correctly, confirming the USB subsystem is fully functional. This is a test infrastructure/hardware setup issue, not a kernel regression.
  3. Possible fix: Connect a USB flash drive or other USB device to one of the qcs9100-ride board's USB host ports in the LAVA lab to satisfy the test's hardware expectations. Alternatively, update the USBHost test to validate USB host controller functionality without requiring external devices, or skip the test on boards where external USB devices are not available.
  4. Detail analysis attachment: failed_case_job227678_3_detailed.md
  Case 4: ** Driver Probe Failure — Ethernet PHY
  1. Failed case: ** Driver Probe Failure — Ethernet PHY
  2. Root cause: ** The Aquantia AQR115C PHY driver probe failed with error -EINVAL because the device tree node for the PHY (stmmac-0:08) is missing the required firmware-name property, preventing the driver from loading the PHY firmware and causing the end0 Ethernet interface to fail when attempting to attach to the PHY.
  3. Possible fix: Add the firmware-name property to the Aquantia AQR115C PHY device tree node in the qcs9100-ride DTS file (e.g., firmware-name = "Rhe-05.06-Candidate7-AQR_Mediatek_23B_P5_ID45824_LCLVER1.cld"; or the appropriate firmware filename for this PHY). This is a pre-existing device tree issue unrelated to the PR under test (which only modifies kernel config for GPU virtualization).
  4. Detail analysis attachment: failed_case_job227678_4_detailed.md
  Case 5: KVM Driver Initialization Failure — HYP mode not available
  1. Failed case: KVM Driver Initialization Failure — HYP mode not available
  2. Root cause: KVM driver initialization failed because EL2 (HYP mode) is occupied by the Gunyah hypervisor running on qcs9100-ride. KVM requires exclusive access to EL2 to provide virtualization support, but Gunyah is already using EL2 as the platform's primary hypervisor. The kernel message kvm [1]: HYP mode not available at boot time (line 2516) indicates KVM detected it cannot access HYP mode, preventing /dev/kvm device node creation.
  3. Possible fix: This is a pre-existing platform limitation, not a PR-introduced regression. The PR only enables VirtIO GPU drivers (CONFIG_DRM_VIRTIO_GPU) for GPU virtualization under Gunyah, and does not affect KVM. To resolve: (1) Exclude KVM tests from qcs9100-ride CI runs, as this platform uses Gunyah hypervisor and cannot support KVM; or (2) Disable CONFIG_KVM in the kernel config for Gunyah-based platforms to prevent the driver from attempting initialization; or (3) Update the test suite to skip KVM tests when Gunyah hypervisor is detected at boot.
  4. Detail analysis attachment: failed_case_job227678_5_detailed.md
  Case 6: KVM_EL2_DTB — KVM/ARM Virtualization Not Available (Gunyah Hypervisor Conflict)
  1. Failed case: KVM_EL2_DTB — KVM/ARM Virtualization Not Available (Gunyah Hypervisor Conflict)
  2. Root cause: The qcs9100-ride platform boots with Gunyah hypervisor in EL2, which prevents KVM/ARM from initializing. KVM requires exclusive EL2 access, but Gunyah occupies EL2 and provides its own virtualization interface. The kernel logs show "kvm [1]: HYP mode not available" at boot, and bootloader logs confirm "Gunyah based bootup" with hypervisor version "gunyah-cdfb73831". This is a platform configuration issue, not a kernel regression introduced by the PR (which only enables DRM VirtIO GPU configs).
  3. Possible fix: This is expected behavior on Gunyah-enabled platforms. KVM/ARM and Gunyah are mutually exclusive hypervisors. To enable KVM testing: (1) flash a non-Gunyah firmware/bootloader configuration that boots Linux directly in EL1 without a hypervisor, OR (2) exclude KVM tests from the qcs9100-ride LAVA job definition, as this platform is configured for Gunyah virtualization, not KVM/ARM. The PR changes (GPU virtualization configs) are unrelated to this failure.
  4. Detail analysis attachment: failed_case_job227678_6_detailed.md
  Case 7: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Suppress KVM_Infra, KVM_Driver, and KVM_EL2_DTB test cases for qcs9100-ride and all other Qualcomm platforms configured with Gunyah hypervisor. KVM host functionality is architecturally incompatible with platforms running under a hypervisor. If KVM guest support is needed, use nested virtualization (if supported by Gunyah) or run KVM tests only on bare-metal configurations without Gunyah.
  4. Detail analysis attachment: failed_case_job227678_7_detailed.md
  Case 8: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM infrastructure test fails because /dev/kvm device node is not created; kernel reports "kvm [1]: HYP mode not available" at boot (line 3127 of log), indicating the qcs9100-ride platform does not support ARM virtualization extensions (EL2 hypervisor mode) in its current hardware/firmware configuration.
  3. Possible fix: This is a pre-existing platform limitation, not a PR-introduced regression (PR only enables GPU virtualization config options). Mark KVM tests as expected-fail or skip for qcs9100-ride platform in the LAVA test definition, or enable HYP mode support in the platform firmware/bootloader if virtualization is required.
  4. Detail analysis attachment: failed_case_job227678_8_detailed.md
Job 227679 | SoC hamoa-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/227679

Failed test cases in LAVA job 227679 (SoC: hamoa-evk).

  Case 1: Probe_Failure_Check — Pre-existing Platform Issues
  1. Failed case: Probe_Failure_Check — Pre-existing Platform Issues
  2. Root cause: Three unrelated driver probe failures detected by zero-tolerance test: (1) qcom_qseecom_uefisecapp failed with -EBUSY (TrustZone resource conflict), (2) qcom-spmi-lpg failed with -EINVAL (DT multi-LED "reg" property validation error), (3) regulatory.db firmware load failed with -ENOENT (missing file, benign with built-in fallback). All three are pre-existing hamoa-evk platform/configuration issues completely unrelated to the PR's GPU virtualization config changes (CONFIG_DRM_VIRTIO_GPU).
  3. Possible fix: Suppress these three known benign failures in the Probe_Failure_Check test for hamoa-evk by adding suppression patterns: qcom_qseecom_uefisecapp.*probe.*failed with error -16, qcom-spmi-lpg.*probe.*failed with error -22, and regulatory: Direct firmware load for regulatory.db failed with error -2. Long-term: (1) investigate TZ resource allocation for QSEE UEFI secure app, (2) fix DT "reg" property in hamoa-evk PMIC LED multi-led node per leds-qcom-lpg.yaml binding, (3) add wireless-regdb package to rootfs or accept built-in fallback.
  4. Detail analysis attachment: failed_case_job227679_1_detailed.md
  Case 2: smmu
  1. Failed case: smmu
  2. Root cause: Six USB PHY devices (a0f8800.usb, a2f8800.usb, a4f8800.usb, a6f8800.usb, a8f8800.usb) and one Video Codec device (aa00000.video-codec) on hamoa-evk lack IOMMU group attachments because their device tree nodes do not include iommus properties, causing the SMMU validation test to fail its critical master protection checks.
  3. Possible fix: Add iommus properties to the device tree nodes for the six USB PHY devices and the Video Codec device in arch/arm64/boot/dts/qcom/x1e80100.dtsi, referencing the appropriate SMMU instance and stream IDs, then verify the devices appear in /sys/kernel/iommu_groups after reboot.
  4. Detail analysis attachment: failed_case_job227679_2_detailed.md
  Case 3: KVM_Driver — /dev/kvm not available (driver initialization failure)
  1. Failed case: KVM_Driver — /dev/kvm not available (driver initialization failure)
  2. Root cause: KVM driver initialization failed because the Hamoa IoT EVK (x7181) platform does not support EL2/HYP mode, which is a prerequisite for KVM. Boot log shows kvm [1]: HYP mode not available at kernel initialization. This is a pre-existing platform limitation, not a PR-introduced regression.
  3. Possible fix: This is a pre-existing platform limitation, not a test failure caused by the PR. The PR only adds guest-side VirtIO GPU drivers (CONFIG_DRM_VIRTIO_GPU) and does not modify KVM host functionality. Recommended actions: (1) Mark KVM tests as "expected to fail" or "skip" for Hamoa IoT EVK platform in the LAVA job definition, as this platform does not support virtualization. (2) If KVM support is required, use a different platform that supports EL2/HYP mode (e.g., RB5, RB3Gen2, or other platforms with virtualization extensions enabled in firmware/bootloader).
  4. Detail analysis attachment: failed_case_job227679_3_detailed.md
  Case 4: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: KVM driver initialization failed because Linux is running as a guest VM under the Gunyah hypervisor on Hamoa EVK, executing at EL1 without access to EL2 (HYP mode). KVM requires EL2 access to provide virtualization services, which is unavailable in this nested virtualization configuration where Gunyah owns EL2.
  3. Possible fix: This is a platform configuration limitation, not a kernel bug. To enable KVM on Hamoa: (1) boot Linux directly at EL2 without Gunyah hypervisor, OR (2) if nested virtualization is required, configure Gunyah to expose virtual EL2 (VHE) to the guest Linux kernel. The PR change (enabling DRM_VIRTIO_GPU configs) is unrelated and does not cause this failure.
  4. Detail analysis attachment: failed_case_job227679_4_detailed.md
  Case 5: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM cannot initialize because HYP (EL2) mode is not available to Linux — the Gunyah hypervisor is running at EL2 and has not configured nested virtualization support for the Primary VM (PVM), preventing KVM from accessing EL2 resources required to create /dev/kvm.
  3. Possible fix: This is a platform/hypervisor configuration issue, not a kernel regression. To enable KVM on hamoa-evk: (1) verify the Gunyah hypervisor firmware version supports nested virtualization for the PVM, (2) ensure the hypervisor configuration grants EL2 access to the PVM, or (3) if nested virtualization is not supported on this platform, mark KVM tests as expected-fail for hamoa-evk in the CI test matrix.
  4. Detail analysis attachment: failed_case_job227679_5_detailed.md
  Case 6: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM initialization failed because the Hamoa IoT EVK platform is running under the Gunyah hypervisor, which does not expose EL2 (HYP mode) to the guest OS. The kernel message "kvm [1]: HYP mode not available" indicates that CONFIG_KVM is enabled but the CPU is not running at EL2, preventing /dev/kvm device creation.
  3. Possible fix: This is not a PR-introduced regression (PR only adds GPU virtualization configs). The test failure is expected on this platform. Either: (1) skip KVM tests on Gunyah-based platforms in the CI test suite, or (2) enable nested virtualization support in the Gunyah hypervisor configuration for this board if the hardware supports it.
  4. Detail analysis attachment: failed_case_job227679_6_detailed.md
Job 227680 | SoC purwa-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/227680

Failed test cases in LAVA job 227680 (SoC: purwa-evk).

  Case 1: ** Probe_Failure_Check
  1. Failed case: ** Probe_Failure_Check
  2. Root cause: ** Five driver probe failures detected during boot on purwa-evk: qcom_qseecom_uefisecapp (-EBUSY, device busy), qcom-spmi-lpg (-EINVAL, invalid DT configuration), two qcom-pcie instances (-ENODATA, missing PCIe endpoint/link training failure), and regulatory.db firmware (-ENOENT, missing firmware file in rootfs).
  3. Possible fix: These are pre-existing platform/configuration issues unrelated to PR QCLINUX: arm64: configs: Enable GPU virtualization #1118 (which only enables GPU virtualization configs). The PR does not modify any drivers, device trees, or firmware paths. Mark as NOT PR-INTRODUCED. To resolve: (1) qcom_qseecom_uefisecapp: verify TrustZone app availability; (2) qcom-spmi-lpg: fix PWM DT node configuration at c42d000.spmi:pmic@1:pwm; (3) qcom-pcie: verify PCIe endpoint presence and DT regulators for 1bd0000.pcie and 1bf8000.pci; (4) regulatory.db: add wireless-regdb package to rootfs or disable cfg80211 built-in regulatory DB requirement.
  4. Detail analysis attachment: failed_case_job227680_1_detailed.md
  Case 2: smmu
  1. Failed case: smmu
  2. Root cause: Test validation failure — six critical devices (five USB controllers at a0f8800, a2f8800, a4f8800, a6f8800, a8f8800 and one Video codec at aa00000) are missing IOMMU group attachments on purwa-evk (X1E80100 SoC), failing the test's strict requirement that all critical masters must be protected by SMMU.
  3. Possible fix: This is a pre-existing platform configuration issue unrelated to PR QCLINUX: arm64: configs: Enable GPU virtualization #1118 (which only enables GPU virtualization config options). The missing IOMMU attachments indicate incomplete device tree iommus properties for these USB and Video devices on the purwa-evk platform. Add iommus properties to the device tree nodes for USB controllers a0f8800.usb, a2f8800.usb, a4f8800.usb, a6f8800.usb, a8f8800.usb and video-codec aa00000.video-codec in arch/arm64/boot/dts/qcom/x1e80100.dtsi.
  4. Detail analysis attachment: failed_case_job227680_2_detailed.md
  Case 3: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: KVM driver initialization failed because the platform is running under the Gunyah hypervisor (detected at boot: "Hypervisor cold boot, version: gunyah-mobile-ad1fb25c6"), which prevents KVM from accessing EL2 (HYP mode). The kernel message "kvm [1]: HYP mode not available" at boot time indicates the CPU is not in a state that allows KVM to initialize, as the Gunyah hypervisor is already occupying EL2. This is a platform configuration issue specific to purwa-evk running with Gunyah, not a kernel regression introduced by the PR (which only enables GPU virtualization configs).
  3. Possible fix: This is not a PR-introduced regression. The KVM_Driver test is expected to fail on purwa-evk when running under the Gunyah hypervisor, as KVM and Gunyah cannot coexist (both require EL2). Either: (1) exclude KVM tests from the purwa-evk test suite when Gunyah is enabled, or (2) boot purwa-evk without the Gunyah hypervisor if KVM functionality is required for testing. The PR changes (enabling DRM_VIRTIO_GPU configs) are unrelated to this failure.
  4. Detail analysis attachment: failed_case_job227680_3_detailed.md
  Case 4: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: KVM cannot initialize on purwa-evk because the platform is running under the Gunyah hypervisor, which occupies EL2 and prevents nested virtualization. The kernel message "kvm [1]: HYP mode not available" confirms KVM detected it cannot access EL2 hypervisor mode, resulting in /dev/kvm device node not being created.
  3. Possible fix: This is a pre-existing platform limitation, not a PR-introduced regression (the PR only enables GPU virtualization configs). To enable KVM on this platform, either: (1) configure Gunyah to support nested virtualization and expose EL2 to the guest kernel, or (2) boot without Gunyah hypervisor if KVM testing is required. For CI purposes, suppress KVM tests on Gunyah-based platforms or run them only on bare-metal configurations.
  4. Detail analysis attachment: failed_case_job227680_4_detailed.md
  Case 5: KVM Infrastructure Limitation — Gunyah Hypervisor Platform
  1. Failed case: KVM Infrastructure Limitation — Gunyah Hypervisor Platform
  2. Root cause: KVM cannot initialize on purwa-evk because the Gunyah hypervisor is already running at EL2 (Hypervisor Exception Level). KVM requires exclusive access to EL2 to function, but Gunyah (a Type-1 hypervisor) occupies this privilege level. The kernel message kvm [1]: HYP mode not available is the expected and correct behavior on Gunyah-based platforms.
  3. Possible fix: This is not a bug — it is a platform architectural limitation. The KVM_Infra, KVM_Driver, and KVM_EL2_DTB tests should be skipped or marked as "not applicable" for purwa-evk and other Gunyah-based platforms. Update the LAVA test job definition to exclude KVM tests for device type purwa-evk, or add platform detection logic to the test scripts to skip KVM tests when Gunyah hypervisor is detected (check for /sys/firmware/devicetree/base/reserved-memory/gunyah-hyp@* or hypervisor boot messages).
  4. Detail analysis attachment: failed_case_job227680_5_detailed.md
  Case 6: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM HYP mode is not available on the purwa-evk platform — kernel message "kvm [1]: HYP mode not available" indicates the hardware/firmware does not support EL2 virtualization, preventing /dev/kvm device creation.
  3. Possible fix: This is a pre-existing platform limitation, not a PR-introduced regression. The PR (GPU virtualization config changes) is unrelated to KVM HYP mode availability. If KVM support is required on purwa-evk, verify that: (1) the bootloader/firmware enables EL2, (2) the device tree does not disable virtualization, and (3) the SoC/board physically supports ARM virtualization extensions. Otherwise, skip KVM tests on this platform.
  4. Detail analysis attachment: failed_case_job227680_6_detailed.md
Job 227681 | SoC monaco-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/227681

Failed test cases in LAVA job 227681 (SoC: monaco-evk).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Add the missing WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin to the rootfs /lib/firmware directory in the build recipe. Verify the firmware package (e.g., linux-firmware-ath11k or vendor-specific firmware package) is included in the Yocto image recipe for Monaco EVK builds.
  4. Detail analysis attachment: failed_case_job227681_1_detailed.md
  Case 2: WiFi Driver Probe Failure — ath11k_pci probe timeout
  1. Failed case: WiFi Driver Probe Failure — ath11k_pci probe timeout
  2. Root cause: ath11k_pci driver probe failed with -ETIMEDOUT (-110) on Monaco EVK because the required WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin is missing from the rootfs (-ENOENT). MHI bus power-up timed out waiting for firmware load, causing the WCN6855 WiFi chip initialization to fail. This is a pre-existing platform/image issue unrelated to PR QCLINUX: arm64: configs: Enable GPU virtualization #1118 (which only enables GPU virtualization kernel configs).
  3. Possible fix: Add the missing ath11k WCN6855 firmware files to the Monaco EVK rootfs image. The required firmware path is ath11k/WCN6855/hw2.1/nfa765/amss.bin and associated board files. Verify firmware package installation in the Yocto/build recipe for Monaco EVK. This is not a kernel regression introduced by the PR.
  4. Detail analysis attachment: failed_case_job227681_2_detailed.md
  Case 3: ** WiFi Driver Probe Failure — ath11k_pci firmware missing
  1. Failed case: ** WiFi Driver Probe Failure — ath11k_pci firmware missing
  2. Root cause: ** WiFi driver probe failed with -110 (ETIMEDOUT) because the required firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin is missing from the rootfs (-ENOENT). MHI cannot power up the WCN6855 WiFi device without firmware, causing probe to time out. This is a pre-existing firmware packaging issue on monaco-evk, not a regression introduced by PR QCLINUX: arm64: configs: Enable GPU virtualization #1118 (which only modifies GPU virtualization configs).
  3. Possible fix: Add the missing firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin to the rootfs firmware directory (/lib/firmware/). Verify the firmware package for WCN6855 hw2.1 includes the nfa765 variant for monaco-evk. If the firmware path is incorrect, update the device tree or driver board file to reference the correct firmware variant path.
  4. Detail analysis attachment: failed_case_job227681_3_detailed.md
  Case 4: 0_qcom-next-ci-premerge-tests
  1. Failed case: 0_qcom-next-ci-premerge-tests
  2. Root cause: LAVA test infrastructure issue — test runner completed and exited normally (<LAVA_TEST_RUNNER EXIT> signal sent), but LAVA dispatcher marked the test definition as failed with "Marking unfinished test run as failed", indicating the dispatcher expected additional output or a different completion signal that was not received.
  3. Possible fix: Re-trigger the CI job. If the issue recurs, review the LAVA job definition to ensure the test runner completion signal matches what the dispatcher expects, or increase the test action timeout if the runner is being terminated prematurely.
  4. Detail analysis attachment: failed_case_job227681_4_detailed.md
Job 227682 | SoC qcs8300-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/227682

Failed test cases in LAVA job 227682 (SoC: qcs8300-ride).

  Case 1: Probe_Failure_Check — Pre-existing Driver/Firmware Issues
  1. Failed case: Probe_Failure_Check — Pre-existing Driver/Firmware Issues
  2. Root cause: The Probe_Failure_Check test detected two pre-existing kernel probe/firmware errors: (1) Aquantia AQR115C Ethernet PHY driver probe failure with -EINVAL (-22) due to missing firmware-name DT property, and (2) cfg80211 regulatory.db firmware load failure with -ENOENT (-2) which is a known benign issue on systems without the regulatory database file. Neither failure is related to the PR changes (GPU virtualization config enablement) and both are platform-specific issues on qcs8300-ride.
  3. Possible fix: These are pre-existing platform issues unrelated to PR QCLINUX: arm64: configs: Enable GPU virtualization #1118. (1) For the Aquantia PHY probe failure: add the missing firmware-name property to the Ethernet PHY device tree node in qcs8300-ride DTS, or update the PHY driver to handle missing firmware-name gracefully. (2) For the regulatory.db firmware failure: this is benign and can be suppressed in the test — cfg80211 falls back to built-in regulatory data when the file is missing. The PR should be approved as the failures are not regressions introduced by the GPU virtualization config changes.
  4. Detail analysis attachment: failed_case_job227682_1_detailed.md
  Case 2: USBHost — Test Infrastructure Issue (No Physical USB Device Connected)
  1. Failed case: USBHost — Test Infrastructure Issue (No Physical USB Device Connected)
  2. Root cause: The USBHost test expects a functional USB device to be connected to the qcs8300-ride board's USB host port, but only the USB root hub is enumerated (Bus 001 Device 001: ID 1d6b:0002 Linux Foundation 2.0 root hub). The USB host controller (xhci-hcd) initialized successfully and detected 1 port, but no physical USB device is plugged into that port. This is a test setup issue, not a kernel regression. The PR changes (CONFIG_DRM_VIRTIO_GPU GPU virtualization configs) are unrelated to USB functionality.
  3. Possible fix: Connect a functional USB device (e.g., USB flash drive, USB keyboard, or USB hub with devices) to the USB host port on the qcs8300-ride board before running the USBHost test. If the board's USB host port is not accessible or not functional in the LAVA lab setup, mark this test as "skip" or "not applicable" for this board configuration, or update the test to verify only that the USB host controller driver loads successfully without requiring a physical device.
  4. Detail analysis attachment: failed_case_job227682_2_detailed.md
  Case 3: Ethernet_Basic_Validation
  1. Failed case: Ethernet_Basic_Validation
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Add the firmware-name property to the Aquantia AQR115C PHY device tree node in the qcs8300-ride DTS file (likely arch/arm64/boot/dts/qcom/qcs8300-ride.dts or overlay). The property should specify the correct firmware file path for the AQR115C PHY chip (e.g., firmware-name = "Rhe-05.06-Candidate9-AQR_Mediatek_23B_P5_ID45824_LCLVER1.cld"; or the appropriate firmware file available in /lib/firmware/). Verify the firmware file exists in the rootfs firmware directory.
  4. Detail analysis attachment: failed_case_job227682_3_detailed.md
  Case 4: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: KVM driver failed to initialize on qcs8300-ride (Monaco) platform despite CONFIG_KVM=y — /dev/kvm device node was not created because the KVM subsystem did not probe successfully. No KVM initialization messages appear in the boot log, indicating the ARM KVM driver silently failed to initialize, most likely because EL2 (hypervisor mode) is not available or is disabled by the platform firmware/bootloader on this SoC.
  3. Possible fix: Verify that the qcs8300-ride platform firmware/bootloader enables EL2 (ARM virtualization extensions). Check bootloader logs for EL2 availability and ensure the platform supports KVM. If EL2 is disabled in firmware, update the bootloader configuration to enable virtualization support. If the platform does not support KVM in hardware, mark the KVM tests as SKIP for this SoC in the CI test matrix.
  4. Detail analysis attachment: failed_case_job227682_4_detailed.md
  Case 5: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: KVM driver failed to initialize on qcs8300-ride platform because the hardware does not support virtualization extensions (VHE/nVHE) or the bootloader did not boot the kernel at EL2. CONFIG_KVM is enabled but the KVM subsystem silently skipped initialization when it detected the platform lacks the required hypervisor mode support, resulting in no /dev/kvm device node creation.
  3. Possible fix: This is a pre-existing platform limitation, not a regression introduced by PR QCLINUX: arm64: configs: Enable GPU virtualization #1118 (which only adds GPU virtualization configs). Mark the KVM test suite as "not applicable" for qcs8300-ride in the LAVA test configuration, or update the platform firmware/bootloader to enable EL2/hypervisor mode if the hardware supports it. To verify hardware capability, check CPU feature registers (ID_AA64PFR0_EL1) for EL2 support.
  4. Detail analysis attachment: failed_case_job227682_5_detailed.md
  Case 6: KVM_Infra — /dev/kvm device node unavailable
  1. Failed case: KVM_Infra — /dev/kvm device node unavailable
  2. Root cause: QCS8300 (Monaco) platform runs as a guest under Gunyah hypervisor (evidenced by gunyah-md-region reservation and arch_timer running in virt mode); KVM requires EL2 hypervisor mode but the kernel is running at EL1 as a guest VM, preventing KVM initialization and /dev/kvm creation despite CONFIG_KVM being enabled.
  3. Possible fix: This is a platform architecture limitation, not a PR-introduced regression. KVM tests should be skipped on QCS8300 targets running under Gunyah hypervisor. Add platform detection to the LAVA test definition to skip KVM tests when running as a hypervisor guest (check for /sys/hypervisor or virt timer mode).
  4. Detail analysis attachment: failed_case_job227682_6_detailed.md
  Case 7: LAVA Test Definition Meta-Failure (0_qcom-next-ci-premerge-tests)
  1. Failed case: LAVA Test Definition Meta-Failure (0_qcom-next-ci-premerge-tests)
  2. Root cause: LAVA marked the test definition as failed because 6 individual tests within it failed (KVM_Driver, KVM_EL2_DTB, KVM_Infra, Probe_Failure_Check, USBHost, Ethernet_Basic_Validation); this is standard LAVA behavior, not a kernel crash or infrastructure failure.
  3. Possible fix: This is not a failure that requires a "fix" - it is LAVA's expected behavior. To resolve the overall test definition failure, investigate and fix the 6 individual test failures. The KVM failures are likely due to missing KVM device node creation (CONFIG_KVM enabled but /dev/kvm not present), which may require hypervisor/firmware configuration or additional kernel config. The other failures (Probe_Failure_Check, USBHost, Ethernet) are unrelated to the PR patch and appear to be pre-existing platform issues on qcs8300-ride.
  4. Detail analysis attachment: failed_case_job227682_7_detailed.md
Job 227683 | SoC qcs6490-rb3gen2

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/227683

Failed test cases in LAVA job 227683 (SoC: qcs6490-rb3gen2).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Add the missing firmware file renesas_usb_fw.mem (from the linux-firmware package) to /lib/firmware/ in the LAVA root filesystem image. The firmware is available in the upstream linux-firmware repository at https://git.kernel.org/pub/scm/linux/kernel/git/firmware/linux-firmware.git/tree/renesas_usb_fw.mem. After adding the firmware, rebuild the root filesystem image and re-trigger the LAVA job to verify the xhci-pci-renesas driver probes successfully.
  4. Detail analysis attachment: failed_case_job227683_1_detailed.md
  Case 2: USBHost
  1. Failed case: USBHost
  2. Root cause: USB host controller drivers (dwc3-qcom/xhci-hcd) did not probe the on-SoC USB controllers (8c00000.usb, a600000.usb) despite device tree nodes being present and registered. No USB host controller initialization occurred, resulting in zero enumerated USB devices. This is a pre-existing kernel configuration or driver loading issue unrelated to the PR (which only enables GPU virtualization configs).
  3. Possible fix: Verify USB host controller driver configs (CONFIG_USB_DWC3, CONFIG_USB_DWC3_QCOM, CONFIG_USB_XHCI_HCD, CONFIG_USB_XHCI_PLATFORM) are enabled (=y or =m). If built as modules, ensure they are loaded at boot. Check for deferred probe issues by examining /sys/kernel/debug/devices_deferred on the target. This failure is not caused by the PR changes (GPU virtualization) and represents a baseline kernel configuration gap.
  4. Detail analysis attachment: failed_case_job227683_2_detailed.md
  Case 3: KVM_Driver — /dev/kvm not available
  1. Failed case: KVM_Driver — /dev/kvm not available
  2. Root cause: KVM initialization failed with "HYP mode not available" because the qcs6490-rb3gen2 platform is running under the Gunyah hypervisor at EL2, preventing the Linux kernel from accessing EL2 to initialize KVM; this is a platform configuration constraint, not a kernel regression.
  3. Possible fix: This is expected behavior on platforms running under a hypervisor. If KVM functionality is required, boot the kernel without the Gunyah hypervisor (bare-metal EL2 access), or use nested virtualization if supported by the hypervisor. The PR (GPU virtualization config changes) did not cause this failure.
  4. Detail analysis attachment: failed_case_job227683_3_detailed.md
  Case 4: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: This is expected platform behavior, not a bug. To enable KVM testing on this platform, either: (1) disable Gunyah hypervisor in the firmware/bootloader configuration to allow KVM to own EL2, or (2) exclude KVM tests from the CI test suite for platforms running Gunyah hypervisor, or (3) run KVM tests on a different platform variant without Gunyah enabled.
  4. Detail analysis attachment: failed_case_job227683_4_detailed.md
  Case 5: KVM Infrastructure Unavailable — Gunyah Hypervisor Conflict
  1. Failed case: KVM Infrastructure Unavailable — Gunyah Hypervisor Conflict
  2. Root cause: KVM cannot initialize because the Gunyah hypervisor (gunyah-cdfb73831) is already running at EL2 on qcs6490-rb3gen2, preventing KVM from obtaining required HYP mode access; kernel logs "kvm [1]: HYP mode not available" and /dev/kvm is never created.
  3. Possible fix: This is a platform configuration issue, not a PR regression (PR only adds GPU virtualization configs). To enable KVM testing: either (1) boot without Gunyah hypervisor to allow KVM native EL2 access, or (2) configure nested virtualization if Gunyah supports it, or (3) exclude KVM tests from qcs6490-rb3gen2 CI runs since this platform is configured for Gunyah-based virtualization, not KVM.
  4. Detail analysis attachment: failed_case_job227683_5_detailed.md
  Case 6: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM host support unavailable on qcs6490-rb3gen2 platform — kernel message "kvm [1]: HYP mode not available" indicates the SoC/firmware does not provide EL2 (hypervisor) mode access required for KVM operation, despite CONFIG_KVM being enabled in the kernel configuration.
  3. Possible fix: This is a pre-existing platform limitation, not a PR-introduced regression. The qcs6490-rb3gen2 board does not support KVM host mode because the bootloader/firmware does not enable EL2 access. Mark KVM tests as expected-fail for this platform, or enable EL2 in the board's bootloader/firmware configuration if virtualization host support is required.
  4. Detail analysis attachment: failed_case_job227683_6_detailed.md
Job 227684 | SoC shikra-iqs-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/227684

Failed test cases in LAVA job 227684 (SoC: shikra-iqs-evk).

  Case 1: GIC Test Script Bug — False Failure
  1. Failed case: GIC Test Script Bug — False Failure
  2. Root cause: The GIC test script at line 75 has a parsing bug that incorrectly attempts to extract interrupt counts for CPUs 4-7 when the shikra-iqs-evk platform has only 4 CPUs (0-3). The script parses /proc/interrupts output but treats the "GICv3", "Level", and "arch_timer" descriptor strings as additional CPU columns, causing bash integer comparison errors and false test failures for non-existent CPUs 4-7.
  3. Possible fix: Update the GIC test script to dynamically detect the actual CPU count from /sys/devices/system/cpu/online or /proc/cpuinfo before parsing /proc/interrupts, and only validate interrupt counts for CPUs that actually exist on the platform. The script should parse the /proc/interrupts format correctly by identifying the CPU column count from the header line rather than assuming a fixed 8-CPU layout.
  4. Detail analysis attachment: failed_case_job227684_1_detailed.md
  Case 2: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Pre-existing platform-specific probe failures on shikra-iqs-evk unrelated to PR changes (GPU virtualization config). Failures include: coresight-etm4x (-22/EINVAL, likely DT/hardware mismatch), cpufreq-dt (-17/EEXIST, benign duplicate registration), regulatory.db (-2/ENOENT, missing firmware file), and lt9611c display bridge (-5/EIO, hardware communication failure).
  3. Possible fix: These are known baseline issues on shikra-iqs-evk. The Probe_Failure_Check test should be updated to exclude known-benign failures (cpufreq-dt EEXIST, regulatory.db missing firmware) and platform-specific hardware issues (coresight ETM on shikra, lt9611c bridge). PR1118 is not the root cause and should not be blocked by this test failure.
  4. Detail analysis attachment: failed_case_job227684_2_detailed.md
  Case 3: ** USBHost
  1. Failed case: ** USBHost
  2. Root cause: ** Test infrastructure issue — the USBHost test failed because no USB device was physically connected to the shikra-iqs-evk board's USB host port during the test run. The USB subsystem initialized correctly (usbcore drivers loaded, USB controller added to IOMMU group 5), and the PR changes (GPU virtualization configs) are unrelated to USB functionality. This is not a kernel bug or regression.
  3. Possible fix: Connect a USB device (flash drive, keyboard, or mouse) to the shikra-iqs-evk board's USB host port in the LAVA lab infrastructure. Alternatively, update the USBHost test script to SKIP (not FAIL) when no USB devices are present, if the test environment does not guarantee a USB device is connected.
  4. Detail analysis attachment: failed_case_job227684_3_detailed.md
  Case 4: BT_SCAN — Test Environment Issue (No Scannable Devices)
  1. Failed case: BT_SCAN — Test Environment Issue (No Scannable Devices)
  2. Root cause: BT_SCAN test failed because no Bluetooth devices were available in the LAVA lab environment to be discovered during the scan window (3 attempts × 15 seconds each). The Bluetooth stack is fully functional (hci0 adapter operational, power on/off working, bluetoothctl commands executing correctly), but the test requires at least one nearby Bluetooth device to pass. This is a test infrastructure limitation, not a kernel regression. The PR changes only GPU virtualization config (CONFIG_DRM_VIRTIO_GPU) and has no relationship to Bluetooth functionality.
  3. Possible fix: This is not a kernel bug requiring a code fix. The test failure is environmental. To resolve: (1) Ensure at least one Bluetooth beacon/device is powered on and within range of the shikra-iqs-evk board in the LAVA lab during test execution, OR (2) Mark BT_SCAN as an optional/informational test that does not block PR merge when the test environment lacks scannable devices, OR (3) Add a test environment readiness check that skips BT_SCAN with a clear message when no reference Bluetooth device is detected in the lab.
  4. Detail analysis attachment: failed_case_job227684_4_detailed.md
  Case 5: ** KVM_Driver
  1. Failed case: ** KVM_Driver
  2. Root cause: ** Shikra IQS EVK platform does not provide EL2 (HYP mode) access to the kernel — KVM initialization detects "HYP mode not available" during boot and does not create /dev/kvm device node. This is a platform/firmware configuration limitation, not a kernel regression.
  3. Possible fix: Configure the Shikra IQS EVK bootloader/firmware to boot the kernel at EL2 or preserve EL2 access for KVM. This requires updating UEFI/ABL configuration to enable virtualization support. Alternatively, if EL2 support is not available on this platform, mark KVM tests as "not applicable" for Shikra IQS EVK in the CI test matrix.
  4. Detail analysis attachment: failed_case_job227684_5_detailed.md
  Case 6: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: KVM hypervisor mode (HYP/EL2) is not available on the Shikra IQS EVK platform, preventing KVM initialization and /dev/kvm device node creation; kernel log shows "kvm [1]: HYP mode not available" at boot time (line 3241, timestamp 3.298864s).
  3. Possible fix: This is a platform hardware/firmware limitation, not a kernel regression introduced by PR QCLINUX: arm64: configs: Enable GPU virtualization #1118 (which only adds DRM virtio-gpu guest driver configs). The Shikra IQS EVK does not support KVM virtualization because it lacks EL2 hypervisor mode support in its boot firmware or CPU configuration. To enable KVM on this platform: (1) verify the SoC supports virtualization extensions (ARMv8.0-A Virtualization Extensions), (2) ensure the bootloader (ABL/UEFI) boots the kernel at EL2 instead of EL1, (3) check if secure firmware (TZ/ATF) allows EL2 access, or (4) use a different platform that supports KVM for virtualization testing.
  4. Detail analysis attachment: failed_case_job227684_6_detailed.md
  Case 7: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: This is a platform/firmware limitation, not a kernel bug. The Shikra IQS EVK board firmware does not boot the kernel in EL2 (HYP mode), which is a prerequisite for KVM on ARM64. To enable KVM: (1) Update board firmware/bootloader to boot kernel at EL2, or (2) Use a hypervisor-enabled boot flow, or (3) Skip KVM tests on this platform as KVM is not supported without EL2 access. The PR patch (GPU virtualization config) is unrelated and did not cause this failure.
  4. Detail analysis attachment: failed_case_job227684_7_detailed.md
  Case 8: ** Kernel Crash — Synchronous External Abort in qcom_rng driver
  1. Failed case: ** Kernel Crash — Synchronous External Abort in qcom_rng driver
  2. Root cause: ** Hardware bus error (synchronous external abort) when qcom_rng driver attempted to read from RNG hardware MMIO register at address 0xffff800082c7d004 during qcom_hwrng test execution. The hardware did not respond to the read request, causing the CPU to raise a fatal exception. This is a pre-existing platform/hardware issue on the Shikra IQS EVK board, not introduced by PR1118 (which only adds GPU virtualization config options).
  3. Possible fix: This is not a PR-introduced regression. Recommended actions: (1) Verify RNG hardware power/clock configuration in device tree for Shikra IQS EVK. (2) Check if RNG hardware block is properly powered and clocked during test execution. (3) Inspect the specific LAVA lab board (serial E2030216110A0F24) for hardware defects. (4) Add runtime PM checks in qcom_rng driver to ensure hardware is accessible before MMIO reads. (5) Re-run the test on a different Shikra board to confirm if issue is board-specific.
  4. Detail analysis attachment: failed_case_job227684_8_detailed.md
  Case 9: Kernel Crash — synchronous external abort in qcom_rng driver
  1. Failed case: Kernel Crash — synchronous external abort in qcom_rng driver
  2. Root cause: Hardware RNG driver (qcom_rng) triggered a synchronous external abort at PC qcom_rng_read+0xc4 when the qcom_hwrng test attempted to read entropy from /dev/hwrng. The crash occurred at kernel timestamp 1012.954826s during a dd read operation. The board subsequently panicked and rebooted into EDL/ramdump mode, causing the LAVA test shell to time out after 2400 seconds waiting for test completion. This is a pre-existing hardware/driver issue unrelated to the PR (which only enables GPU virtualization configs).
  3. Possible fix: This is a pre-existing kernel crash in the qcom_rng driver on shikra-iqs-evk, not introduced by PR QCLINUX: arm64: configs: Enable GPU virtualization #1118. The PR changes are limited to enabling CONFIG_DRM_VIRTIO_GPU and CONFIG_DRM_VIRTIO_GPU_KMS in arch/arm64/configs/qcom.config, which do not affect the hardware RNG subsystem. The crash should be triaged separately as a baseport/hardware issue. To unblock PR validation: (1) skip or disable the qcom_hwrng test case in the LAVA job definition for shikra-iqs-evk until the RNG driver issue is resolved, or (2) re-run the job to confirm reproducibility — if the crash is intermittent hardware-related, a retry may pass.
  4. Detail analysis attachment: failed_case_job227684_9_detailed.md
  Case 10: lava-test-retry
  1. Failed case: lava-test-retry
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: This is a pre-existing platform/firmware issue, not introduced by PR QCLINUX: arm64: configs: Enable GPU virtualization #1118 (which only enables GPU virtualization config options). The qcom_rng driver is attempting to access hardware that is not properly initialized or powered on the shikra-iqs-evk platform. Recommended actions: (1) Verify the PRNG/TRNG hardware block is enabled in the device tree for shikra-iqs-evk; (2) Check if the qcom_rng driver probe succeeded and the device is in the correct power state before the test runs; (3) Add error handling in the qcom_rng driver to gracefully handle MMIO access failures instead of crashing; (4) If this is a known shikra-iqs-evk hardware limitation, disable the qcom_hwrng test for this platform or mark it as expected-fail in the LAVA job definition.
  4. Detail analysis attachment: failed_case_job227684_10_detailed.md
  Case 11: ** Kernel Crash — Synchronous External Abort in qcom_rng driver
  1. Failed case: ** Kernel Crash — Synchronous External Abort in qcom_rng driver
  2. Root cause: ** Synchronous external abort at qcom_rng_read+0xc4 during hardware register read from the Qualcomm RNG block; the CPU load instruction faulted, indicating the hardware register was inaccessible (likely due to power/clock gating, incorrect MMIO mapping, or bus path issue on shikra-iqs-evk platform).
  3. Possible fix: Investigate qcom_rng driver runtime PM and clock/power dependencies on shikra platform; verify MMIO mapping correctness; add defensive register access checks; consider enabling CONFIG_IOMMU_DEBUGFS and CONFIG_ARM_SMMU_TESTBUS_DUMP to capture bus state at fault time; re-run test with qcom_rng driver debug enabled.
  4. Detail analysis attachment: failed_case_job227684_11_detailed.md
Job 227685 | SoC lemans-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/227685

Failed test cases in LAVA job 227685 (SoC: lemans-evk).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Four SPMI PMIC temp-alarm devices (pmic@0/2/4/6:temp-alarm@a00) remain in deferred probe state, and three firmware load failures (regulatory.db, qca/wcnhpbtfw21.tlv, qca/hpbtfw21.tlv) are detected; these are pre-existing platform/firmware packaging issues unrelated to the PR (which only enables GPU virtualization configs).
  3. Possible fix: For temp-alarm deferred probes: investigate missing thermal zone dependencies in the lemans-evk device tree or kernel config. For firmware failures: these are known benign (regulatory.db is optional; Bluetooth firmware fallback paths exist) — suppress these patterns in the Probe_Failure_Check test or package the missing firmware files in the rootfs.
  4. Detail analysis attachment: failed_case_job227685_1_detailed.md
  Case 2: smmu
  1. Failed case: smmu
  2. Root cause: Test script expects legacy Venus video codec device aa00000.video-codec to be attached to an IOMMU group, but the device did not probe because lemans-evk uses the new Iris video codec driver instead; Iris devices (iris_non_pixel.0, iris_pixel.0) are correctly attached to IOMMU groups 25 and 26.
  3. Possible fix: Update the SMMU test script to recognize Iris video codec devices (iris_non_pixel.0, iris_pixel.0) as valid video codec IOMMU group attachments on platforms that use the Iris driver instead of legacy Venus.
  4. Detail analysis attachment: failed_case_job227685_2_detailed.md
  Case 3: smmu
  1. Failed case: smmu
  2. Root cause: The video codec device (aa00000.video-codec) is missing IOMMU group attachment on lemans-evk, causing the SMMU validation test to fail when checking that all critical masters are protected by IOMMU.
  3. Possible fix: This is a pre-existing platform/DT issue unrelated to PR QCLINUX: arm64: configs: Enable GPU virtualization #1118 (which only enables GPU virtualization kernel configs). The video codec device tree node for lemans-evk needs an iommus property to attach it to an IOMMU group. Verify the device tree for lemans-evk includes proper IOMMU bindings for the video codec at address aa00000, or mark this device as non-critical in the SMMU test if IOMMU protection is not required for this platform.
  4. Detail analysis attachment: failed_case_job227685_3_detailed.md

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants