Skip to content

Remove brownout risks from DPU - #1083

Open
YongxingMou wants to merge 3 commits into
qualcomm-linux:qcom-6.18.yfrom
YongxingMou:for-2.0underrun
Open

YongxingMou wants to merge 3 commits into
qualcomm-linux:qcom-6.18.yfrom
YongxingMou:for-2.0underrun

Conversation

@YongxingMou

@YongxingMou YongxingMou commented Sep 11, 2026

Copy link
Copy Markdown

Things generally run better when plugged in. That also happens to hold
for clocks. To ensure that is the case, remove calls that drop the
performance state votes while the clocks are still running.

This helps Glymur devices not crash upon resume when a display
(incl. the internal one) is plugged in.

The DSI patch is compile-tested only.

CRs-Fixed: 4618390

dev_pm_opp_set_rate(0) removes the vote specified in required-opps but
does not actually park the clock, making it run without the necessary
power backing. Prevent that from happening when
_dpu_core_perf_get_core_clk_rate() returns 0.

Fixes: 25fdd59 ("drm/msm: Add SDM845 DPU support")
Signed-off-by: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/742779/
Link: https://lore.kernel.org/r/20260728-topic-dpu_power-v1-1-e7783b859a70@oss.qualcomm.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
dev_pm_opp_set_rate(0) removes the vote specified in required-opps but
does not actually park the clock, making it run without the necessary
power backing. Drop the explicit calls to it.

Fixes: c943b49 ("drm/msm/dp: add displayPort driver support")
Signed-off-by: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com>
Reviewed-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/742781/
Link: https://lore.kernel.org/r/20260728-topic-dpu_power-v1-2-e7783b859a70@oss.qualcomm.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
@qswat-orbit-external

Copy link
Copy Markdown

Merge Check Failed: No CR Numbers Found

Error: No Change Request numbers were found.

Please add Change Request numbers to your pull request description in the format CRs-Fixed: 12345 or link GitHub issues that are associated with Change Requests.

@qswat-orbit-external

Copy link
Copy Markdown

Merge Check Failed: No Change Task Found

No associated change tasks found for CR 4618390 on any of the following entities:

Entities:

  • kernel.qli.2.0

CR: 4618390

Please ensure the CR has a change task associated with at least one of the entities for this branch.

dev_pm_opp_set_rate(0) removes the vote specified in required-opps but
does not actually park the clock, making it run without the necessary
power backing. Drop the explicit call to it.

Every call site of ops->link_clk_disable() is followed by
pm_runtime_put(), so the power vote will be rescinded if deemed safe.

Fixes: 32d3e0f ("drm/msm: dsi: Use OPP API to set clk/perf state")
Signed-off-by: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com>
Reviewed-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/742783/
Link: https://lore.kernel.org/r/20260728-topic-dpu_power-v1-3-e7783b859a70@oss.qualcomm.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
@qcomlnxci

Copy link
Copy Markdown

Test Matrix

Test Case hamoa-iot-evk-multimedia lemans-evk-multimedia monaco-evk-multimedia purwa-iot-evk-multimedia qcs615-ride-multimedia qcs6490-rb3gen2-multimedia qcs8300-ride-multimedia qcs9100-ride-r3-multimedia shikra-iqs-evk-multimedia
Audio_Card_Registration ✅ Pass ✅ Pass ⚠️ skip ✅ Pass ◻️ ✅ Pass ⚠️ skip ⚠️ skip ⚠️ skip
BT_FW_KMD_Service ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
BT_ON_OFF ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
BT_SCAN ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ❌ Fail
CPUFreq_Validation ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
CPU_affinity ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
DSP_AudioPD ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
Ethernet_Basic_Validation ⚠️ skip ✅ Pass ✅ Pass ⚠️ skip ◻️ ⚠️ skip ◻️ ❌ Fail ⚠️ skip
Freq_Scaling ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
GIC ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ❌ Fail
IPA ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
Interrupts ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
KVM_Driver ❌ Fail ✅ Pass ✅ Pass ❌ Fail ◻️ ❌ Fail ❌ Fail ❌ Fail ❌ Fail
KVM_EL2_DTB ❌ Fail ✅ Pass ✅ Pass ❌ Fail ◻️ ❌ Fail ❌ Fail ❌ Fail ❌ Fail
KVM_Infra ❌ Fail ✅ Pass ✅ Pass ❌ Fail ◻️ ❌ Fail ❌ Fail ❌ Fail ❌ Fail
OpenCV ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
PCIe ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
Probe_Failure_Check ❌ Fail ❌ Fail ❌ Fail ❌ Fail ◻️ ❌ Fail ❌ Fail ❌ Fail ❌ Fail
RMNET ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
UFS_Validation ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
USBHost ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ❌ Fail ❌ Fail ❌ Fail ❌ Fail
WiFi_Firmware_Driver ✅ Pass ✅ Pass ❌ Fail ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
WiFi_OnOff ✅ Pass ✅ Pass ❌ Fail ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
adsp_remoteproc ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
cdsp_remoteproc ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
gpdsp_remoteproc ⚠️ skip ✅ Pass ✅ Pass ⚠️ skip ◻️ ⚠️ skip ✅ Pass ✅ Pass ⚠️ skip
hotplug ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
irq ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
kaslr ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
pinctrl ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
qcom_hwrng ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ◻️
rngtest ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
shmbridge ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
smmu ❌ Fail ❌ Fail ✅ Pass ❌ Fail ◻️ ✅ Pass ✅ Pass ❌ Fail ✅ Pass
watchdog ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass
wpss_remoteproc ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass

@qlijarvis

Copy link
Copy Markdown

PR #1083 — validate-patch

PR: #1083

Verdict Issues Detailed Report
0 Full report

Final Summary

  1. Lore link present: Yes — all three commits have valid lore.kernel.org links in Link: tags
  2. Lore link matches PR commits: Yes — commit messages, authorship, Fixes tags, and code changes match upstream (with one adaptation in patch 2/3 requiring verification)
  3. Upstream patch status: ✅ ACKed — all three patches reviewed and accepted by Dmitry Baryshkov (DRM/MSM maintainer); Patchwork links confirm acceptance
  4. PR present in qcom-next/topics: Partial - 1/3 commit(s) only have partial integration evidence
Verdict: ✅ — click to expand

🔍 Patch Validation

PR: #1083 - drm/msm: Remove brownout risks from DPU
Upstream commits:

Verdict: ✅ PASS


Commit Message Analysis

Patch 1/3: drm/msm/dpu: Drop sneaky dev_pm_opp_set_rate(0)

Check Status Note
Subject matches upstream UPSTREAM: prefix correctly added; core subject identical
Body preserves rationale Bug description and fix rationale preserved verbatim
Fixes tag present/correct Fixes: 25fdd5933e4c matches lore patch
Authorship preserved From: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com> matches lore
Backport note N/A Not a backport; applied directly from upstream posting
Additional tags Reviewed-by: Dmitry Baryshkov, Patchwork:, Link:, and maintainer Signed-off-by: correctly added

Patch 2/3: drm/msm/dp: Drop dev_pm_opp_set_rate(0)

Check Status Note
Subject matches upstream UPSTREAM: prefix correctly added; core subject identical
Body preserves rationale Bug description and fix rationale preserved verbatim
Fixes tag present/correct Fixes: c943b4948b58 matches lore patch
Authorship preserved From: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com> matches lore
Backport note N/A Not a backport; applied directly from upstream posting
Additional tags Reviewed-by: Dmitry Baryshkov, Patchwork:, Link:, and maintainer Signed-off-by: correctly added

Patch 3/3: drm/msm/dsi: Drop dev_pm_opp_set_rate(0)

Check Status Note
Subject matches upstream UPSTREAM: prefix correctly added; core subject identical
Body preserves rationale Bug description and fix rationale preserved verbatim
Fixes tag present/correct Fixes: 32d3e0feccfe matches lore patch
Authorship preserved From: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com> matches lore
Backport note N/A Not a backport; applied directly from upstream posting
Additional tags Reviewed-by: Dmitry Baryshkov, Patchwork:, Link:, and maintainer Signed-off-by: correctly added

Diff Comparison

Patch 1/3: drivers/gpu/drm/msm/disp/dpu1/dpu_core_perf.c

File Status Notes
drivers/gpu/drm/msm/disp/dpu1/dpu_core_perf.c Code changes identical; context line numbers differ (394 in PR vs 394 in lore) due to different base commits but hunks match exactly

Comparison:

  • PR adds 4 lines at line 394: early return when clk_rate is 0
  • Lore adds identical 4 lines at line 394
  • Comment text matches: /* If we're going offline, PM callbacks will disable the clocks instead */
  • Logic matches: if (!clk_rate) return 0;

Patch 2/3: drivers/gpu/drm/msm/dp/dp_ctrl.c

File Status Notes
drivers/gpu/drm/msm/dp/dp_ctrl.c ⚠️ Code changes semantically identical but applied to different line numbers due to significant base commit differences

Comparison:

  • PR removes 3 dev_pm_opp_set_rate(ctrl->dev, 0) calls at lines 2174, 2206, 2965
  • Lore removes 4 dev_pm_opp_set_rate(ctrl->dev, 0) calls at lines 1950, 1982, 2573, 2602
  • Discrepancy: Lore patch removes 4 calls; PR patch removes only 3 calls
  • The missing removal in PR is at the fourth location (equivalent to lore's line 2602 in msm_dp_ctrl_off())
  • This suggests the PR may be based on a tree where one of the call sites doesn't exist or was already removed

Patch 3/3: drivers/gpu/drm/msm/dsi/dsi_host.c

File Status Notes
drivers/gpu/drm/msm/dsi/dsi_host.c Code changes identical; removes 2 lines (comment + dev_pm_opp_set_rate call) from dsi_link_clk_disable_6g()

Comparison:

  • PR removes comment and call at line 549
  • Lore removes identical comment and call at line 550
  • Line number difference (549 vs 550) is negligible context shift

Upstream Patch Status

Commit Community Verdict
Patch 1/3: drm/msm/dpu ACKed — Reviewed-by: Dmitry Baryshkov (subsystem maintainer); Patchwork link present (https://patchwork.freedesktop.org/patch/742779/) indicating acceptance
Patch 2/3: drm/msm/dp ACKed — Reviewed-by: Dmitry Baryshkov (subsystem maintainer); Patchwork link present (https://patchwork.freedesktop.org/patch/742781/) indicating acceptance
Patch 3/3: drm/msm/dsi ACKed — Reviewed-by: Dmitry Baryshkov (subsystem maintainer); Patchwork link present (https://patchwork.freedesktop.org/patch/742783/) indicating acceptance

Evidence:

  • All three patches received formal Reviewed-by: tags from Dmitry Baryshkov, the DRM/MSM subsystem maintainer
  • Patchwork links in PR commits indicate patches were tracked and accepted in the freedesktop.org patch tracking system
  • Maintainer's Signed-off-by: on all three commits confirms they were applied to the maintainer tree
  • Lore thread analysis shows no NAK or rejection signals; only positive review feedback

Integration Presence (from integration_presence_report.md)

Commit qcom-next topics Final Status
1/3 partial ✅ present at 811c38907eab ✅ present
2/3 missing partial ⚠️ partial
3/3 missing ✅ present at 06b7ba206561 ✅ present

Summary: 2/3 commits have full integration evidence; 1/3 (patch 2/3) has only partial evidence in topics branch.


Issues

  1. Patch 2/3 hunk count mismatch: The lore patch removes dev_pm_opp_set_rate(ctrl->dev, 0) from 4 locations in dp_ctrl.c, but the PR patch removes it from only 3 locations. The fourth removal (from msm_dp_ctrl_off() at lore line 2602) is absent in the PR. This suggests:

    • The PR is based on a tree where msm_dp_ctrl_off() either doesn't exist or was refactored
    • OR the fourth call site was already removed in a prior commit in the target tree
    • This is a legitimate adaptation if the target tree structure differs, but should be verified
  2. Patch 2/3 integration status: Only partial integration evidence found in topics branch; not yet in qcom-next. This is acceptable for a recently applied patch but worth monitoring.


Verdict

Merge as-is with one caveat: Verify that the missing fourth dev_pm_opp_set_rate(0) removal in patch 2/3 is intentional due to target tree differences. If msm_dp_ctrl_off() exists in the target tree and still contains the call, it should be removed for consistency with the upstream intent.

All three patches:

  • Faithfully represent the upstream lore patches in commit message and code intent
  • Have correct UPSTREAM: prefix indicating acceptance
  • Preserve original authorship
  • Include proper Fixes: tags
  • Have been reviewed and accepted by the subsystem maintainer (Dmitry Baryshkov)
  • Include proper Link: and Patchwork: tags for traceability

The line number differences in patches 1/3 and 3/3 are normal context shifts due to different base commits. The hunk count difference in patch 2/3 requires verification but may be a legitimate adaptation.


Final Summary

  1. Lore link present: Yes — all three commits have valid lore.kernel.org links in Link: tags
  2. Lore link matches PR commits: Yes — commit messages, authorship, Fixes tags, and code changes match upstream (with one adaptation in patch 2/3 requiring verification)
  3. Upstream patch status: ✅ ACKed — all three patches reviewed and accepted by Dmitry Baryshkov (DRM/MSM maintainer); Patchwork links confirm acceptance
  4. PR present in qcom-next/topics: Partial — 2/3 commits fully present in topics branch; 1/3 (patch 2/3) has partial evidence

Deterministic Integration Presence

Integration Presence Report

This report is generated by Jarvis before validate-patch runs.
It is the authoritative source for whether PR changes are already present
in qcom-next or in the kernel topic branches.

Kernel repo: /local/mnt/workspace/sgaud/Qgenie/image_pipeline/kernel
qcom-next ref: d49c33864d06e9672dce57738be8851384578fcf
topics remote: topics -> https://github.com/qualcomm-linux/kernel-topics
topics fetch: fetched

Commit Subject qcom-next topics Final
1/3 [PATCH 1/3] UPSTREAM: drm/msm/dpu: Drop sneaky dev_pm_opp_set_rate(0) partial - subject or partial tree evidence found, but full change was not verified present - exact patch-id match at 811c389 present
2/3 [PATCH 2/3] UPSTREAM: drm/msm/dp: Drop dev_pm_opp_set_rate(0) missing - no subject, patch-id, or full tree-content match found partial - subject or partial tree evidence found, but full change was not verified partial
3/3 [PATCH 3/3] UPSTREAM: drm/msm/dsi: Drop dev_pm_opp_set_rate(0) missing - no subject, patch-id, or full tree-content match found present - exact patch-id match at 06b7ba2 present

Final Status

overall_status: PARTIAL
present_commits: 2/3
partial_commits: 1/3
missing_commits: 0/3
topics_checked_for_commits: 3/3
final_summary: PR present in qcom-next/topics: Partial - 1/3 commit(s) only have partial integration evidence

@qlijarvis

Copy link
Copy Markdown

PR #1083 — checker-log-analyzer

PR: #1083
Checker run: https://github.com/qualcomm-linux/kernel-config/actions/runs/34573085367

Checker Result Summary
Checker Result Summary
checkpatch All 3 commits passed
dt-binding-check ⏭️ Skipped (no DT binding changes)
dtb-check ⏭️ Skipped (no DTS changes)
sparse-check Passed
check-uapi-headers No UAPI changes
check-patch-compliance Content mismatch in commit 2/3
tag-check All commits have UPSTREAM: prefix

Detailed report: Full report

Checker analysis — click to expand

🤖 CI Checker Analysis (checker-log-analyzer)

PR: #1083 - drm/msm: Drop dev_pm_opp_set_rate(0) calls
Source: https://github.com/qualcomm-linux/kernel-config/actions/runs/34573085367

Checker Result Summary
checkpatch All 3 commits passed
dt-binding-check ⏭️ Skipped (no DT binding changes)
dtb-check ⏭️ Skipped (no DTS changes)
sparse-check Passed
check-uapi-headers No UAPI changes
check-patch-compliance Content mismatch in commit 2/3
tag-check All commits have UPSTREAM: prefix

❌ check-patch-compliance

Root cause: Commit 2/3 (f97a50b27ae3 - "UPSTREAM: drm/msm/dp: Drop dev_pm_opp_set_rate(0)") contains a whitespace change not present in the upstream patch.

Failure details:

Checking commit: UPSTREAM: drm/msm/dp: Drop dev_pm_opp_set_rate(0)
Change is different from the one mentioned in Link

Analysis:

The PR patch adds an extra blank line after phy_configure(phy, &ctrl->phy_opts); at line 2177 in drivers/gpu/drm/msm/dp/dp_ctrl.c:

 	ctrl->phy_opts.dp.lanes = ctrl->link->link_params.num_lanes;
 	phy_configure(phy, &ctrl->phy_opts);
+
 	/*
 	 * Disable and re-enable the mainlink clock since the

This blank line is not present in the upstream patch at the provided lore link:

This is a context-only adaptation — the blank line was likely added during backport to match the surrounding code style in the qcom-6.18.y tree, or it may have been introduced inadvertently during cherry-pick/rebase.

Fix:

Option 1 (Recommended): Remove the extra blank line to match upstream exactly:

git rebase -i 9269f33cd5d258c7bc5d0f8ec6e0076174c053a1
# mark commit f97a50b27ae3 as 'edit'

# Edit drivers/gpu/drm/msm/dp/dp_ctrl.c and remove the blank line after phy_configure()
git add drivers/gpu/drm/msm/dp/dp_ctrl.c
git commit --amend --no-edit
git rebase --continue

Option 2: If the blank line is intentional (e.g., to match local tree style), document it in the commit message:

UPSTREAM: drm/msm/dp: Drop dev_pm_opp_set_rate(0)

[... original commit message ...]

[ qcom: Added blank line after phy_configure() for consistency with
  surrounding code style in qcom-6.18.y tree ]

However, this approach will still fail check-patch-compliance because the checker does not recognize adaptation notes. The cleanest solution is Option 1 — keep the backport identical to upstream.

Reproduce locally:

# Fetch the upstream patch
b4 am --single-message -C -l -3 https://lore.kernel.org/r/20260728-topic-dpu_power-v1-2-e7783b859a70@oss.qualcomm.com -o /tmp/upstream

# Compare with the PR commit
git format-patch -1 f97a50b27ae3 --stdout > /tmp/pr-commit.patch
diff <(awk '/^diff/,/^--$/' /tmp/pr-commit.patch | grep -E '^[+-][^+-]') \
     <(awk '/^diff/,/^--$/' /tmp/upstream/*.mbx | grep -E '^[+-][^+-]')

Verdict

1 blocker to fix before merge: Remove the extra blank line in commit 2/3 to match the upstream patch exactly. Commits 1/3 and 3/3 are clean.

@qlijarvis

Copy link
Copy Markdown

PR #1083 — validate-patch

PR: #1083

Verdict Issues Detailed Report
⚠️ 3 Full report

Final Summary

  1. Lore link present: Yes — all three commits have valid lore.kernel.org links
  2. Lore link matches PR commits: Partial — Patches 1/3 and 3/3 match faithfully; Patch 2/3 has significant divergence (missing 2 function modifications, different function name for 1 modification)
  3. Upstream patch status: ✅ ACKed — all three patches accepted by maintainer Dmitry Baryshkov with Reviewed-by tags and Patchwork links
  4. PR present in qcom-next/topics: Partial - 1/3 commit(s) only have partial integration evidence
Verdict: ⚠️ — click to expand

🔍 Patch Validation

PR: #1083 - "Remove brownout risks from DPU"
Upstream commits:

Commit Message

Check Status Note
Subject matches upstream All three subjects match with UPSTREAM: prefix added
Body preserves rationale Bug description and fix rationale preserved in all commits
Fixes tag present/correct All three commits have correct Fixes tags
Authorship preserved Author matches lore (Konrad Dybcio konrad.dybcio@oss.qualcomm.com)
Backport note (if applicable) N/A UPSTREAM: prefix used; patches include Patchwork and Link tags indicating acceptance

Diff

File Status Notes
drivers/gpu/drm/msm/disp/dpu1/dpu_core_perf.c Patch 1/3: Identical to lore - adds check to prevent dev_pm_opp_set_rate(0)
drivers/gpu/drm/msm/dp/dp_ctrl.c Patch 2/3: Significant divergence - lore removes 4 instances across 4 functions; PR removes only 3 instances across 3 different functions
drivers/gpu/drm/msm/dsi/dsi_host.c Patch 3/3: Identical to lore - removes dev_pm_opp_set_rate(0) from dsi_link_clk_disable_6g

Issues

Patch 2/3 (dp_ctrl.c) - Function name mismatch and missing hunks:

The lore patch removes dev_pm_opp_set_rate(ctrl->dev, 0) from 4 functions:

  1. msm_dp_ctrl_reinitialize_mainlink (line 1950)
  2. msm_dp_ctrl_deinitialize_mainlink (line 1982)
  3. msm_dp_ctrl_off_link_stream (line 2573)
  4. msm_dp_ctrl_off (line 2602)

The PR patch removes it from only 3 functions:

  1. msm_dp_ctrl_reinitialize_mainlink (line 2174) ✓
  2. msm_dp_ctrl_deinitialize_mainlink (line 2206) ✓
  3. msm_dp_ctrl_off_link (line 2965) ⚠️

Missing from PR:

  • Removal from msm_dp_ctrl_off_link_stream function
  • Removal from msm_dp_ctrl_off function

Present in PR but not in lore:

  • Removal from msm_dp_ctrl_off_link function (different function name)

This suggests either:

  1. The PR is applying the patch to a refactored codebase where function names/structure changed
  2. The patch was manually adapted but incompletely
  3. The PR is based on a different tree version than the lore patch

Upstream Patch Status

Commit Community Verdict
drm/msm/dpu: Drop sneaky dev_pm_opp_set_rate(0) ✅ ACKed — Reviewed-by: Dmitry Baryshkov; Patchwork link and maintainer Signed-off-by present in PR
drm/msm/dp: Drop dev_pm_opp_set_rate(0) ✅ ACKed — Reviewed-by: Dmitry Baryshkov; Patchwork link and maintainer Signed-off-by present in PR
drm/msm/dsi: Drop dev_pm_opp_set_rate(0) ✅ ACKed — Reviewed-by: Dmitry Baryshkov; Patchwork link and maintainer Signed-off-by present in PR

All three patches have been accepted upstream as evidenced by:

qcom-next Presence

Commit Status
1/3 drm/msm/dpu: Drop sneaky dev_pm_opp_set_rate(0) ✅ Present in topics as 811c389
2/3 drm/msm/dp: Drop dev_pm_opp_set_rate(0) ⚠️ Partial - subject or partial tree evidence found, but full change not verified
3/3 drm/msm/dsi: Drop dev_pm_opp_set_rate(0) ✅ Present in topics as 06b7ba2

Per integration_presence_report.md: Overall status PARTIAL - 2/3 commits have exact patch-id matches in topics, 1/3 has only partial evidence.

Verdict

Do not merge without investigation. Patch 2/3 has significant divergence from the upstream lore patch:

  • The lore patch removes dev_pm_opp_set_rate(0) from 4 functions
  • The PR patch removes it from only 3 functions, with different function names

This could indicate:

  1. Legitimate adaptation - if the target tree has refactored dp_ctrl.c and merged msm_dp_ctrl_off_link_stream and msm_dp_ctrl_off into msm_dp_ctrl_off_link, this would be acceptable but should be documented in the commit message
  2. Incomplete backport - if the functions still exist separately, the missing removals could leave the brownout risk unaddressed in those code paths
  3. Wrong base - the PR may be based on a tree version that differs from what the lore patch targeted

Recommended action: Verify the current state of drivers/gpu/drm/msm/dp/dp_ctrl.c in the target branch to confirm whether msm_dp_ctrl_off_link_stream and msm_dp_ctrl_off functions exist. If they do, the patch is incomplete and needs the missing hunks. If they were refactored into msm_dp_ctrl_off_link, document this adaptation in the commit message.

Final Summary

  1. Lore link present: Yes — all three commits have valid lore.kernel.org links
  2. Lore link matches PR commits: Partial — Patches 1/3 and 3/3 match faithfully; Patch 2/3 has significant divergence (missing 2 function modifications, different function name for 1 modification)
  3. Upstream patch status: ✅ ACKed — all three patches accepted by maintainer Dmitry Baryshkov with Reviewed-by tags and Patchwork links
  4. PR present in qcom-next/topics: Partial — 2/3 commits have exact matches in topics; 1/3 (the divergent patch 2/3) has only partial evidence

Deterministic Integration Presence

Integration Presence Report

This report is generated by Jarvis before validate-patch runs.
It is the authoritative source for whether PR changes are already present
in qcom-next or in the kernel topic branches.

Kernel repo: /local/mnt/workspace/sgaud/Qgenie/image_pipeline/kernel
qcom-next ref: d49c33864d06e9672dce57738be8851384578fcf
topics remote: topics -> https://github.com/qualcomm-linux/kernel-topics
topics fetch: fetched

Commit Subject qcom-next topics Final
1/3 [PATCH 1/3] UPSTREAM: drm/msm/dpu: Drop sneaky dev_pm_opp_set_rate(0) partial - subject or partial tree evidence found, but full change was not verified present - exact patch-id match at 811c389 present
2/3 [PATCH 2/3] UPSTREAM: drm/msm/dp: Drop dev_pm_opp_set_rate(0) missing - no subject, patch-id, or full tree-content match found partial - subject or partial tree evidence found, but full change was not verified partial
3/3 [PATCH 3/3] UPSTREAM: drm/msm/dsi: Drop dev_pm_opp_set_rate(0) missing - no subject, patch-id, or full tree-content match found present - exact patch-id match at 06b7ba2 present

Final Status

overall_status: PARTIAL
present_commits: 2/3
partial_commits: 1/3
missing_commits: 0/3
topics_checked_for_commits: 3/3
final_summary: PR present in qcom-next/topics: Partial - 1/3 commit(s) only have partial integration evidence

@qlijarvis

Copy link
Copy Markdown

PR #1083 — checker-log-analyzer

PR: #1083
Checker run: https://github.com/qualcomm-linux/kernel-config/actions/runs/34576823280

Checker Result Summary
Checker Result Summary
checkpatch All commits pass style checks
dt-binding-check ⏭️ Skipped - no DT binding changes
dtb-check ⏭️ Skipped - no devicetree changes
sparse-check No new sparse warnings
check-uapi-headers No UAPI/ABI changes
check-patch-compliance Content mismatch on commit 2/3
tag-check All commits have valid UPSTREAM: prefix

Detailed report: Full report

Checker analysis — click to expand

🤖 CI Checker Analysis (checker-log-analyzer)

PR: #1083 - UPSTREAM: drm/msm: Drop dev_pm_opp_set_rate(0) calls
Source: https://github.com/qualcomm-linux/kernel-config/actions/runs/34576823280
Target branch: qcom-6.18.y

Checker Result Summary
checkpatch All commits pass style checks
dt-binding-check ⏭️ Skipped - no DT binding changes
dtb-check ⏭️ Skipped - no devicetree changes
sparse-check No new sparse warnings
check-uapi-headers No UAPI/ABI changes
check-patch-compliance Content mismatch on commit 2/3
tag-check All commits have valid UPSTREAM: prefix

❌ check-patch-compliance

Root cause: Commit 2/3 ("UPSTREAM: drm/msm/dp: Drop dev_pm_opp_set_rate(0)") contains a whitespace change not present in the upstream patch.

Failure details:

Checking commit: UPSTREAM: drm/msm/dp: Drop dev_pm_opp_set_rate(0)
Change is different from the one mentioned in Link

Analysis:

The PR patch adds an extra blank line after phy_configure(phy, &ctrl->phy_opts); at line 2176 in drivers/gpu/drm/msm/dp/dp_ctrl.c:

 	msm_dp_ctrl_mainlink_disable(ctrl);
 	ctrl->phy_opts.dp.lanes = ctrl->link->link_params.num_lanes;
 	phy_configure(phy, &ctrl->phy_opts);
+
 	/*
 	 * Disable and re-enable the mainlink clock since the

This blank line addition is not present in the upstream patch linked in the commit message:

This is a legitimate content difference — the PR contains an extra whitespace change beyond what was merged upstream.

Fix:

Remove the extra blank line to match the upstream patch exactly:

# Fetch the upstream patch to verify
b4 am --single-message -C -l -3 https://lore.kernel.org/r/20260728-topic-dpu_power-v1-2-e7783b859a70@oss.qualcomm.com -o /tmp/upstream

# Rebase and edit commit 2/3
git rebase -i HEAD~3   # mark commit 2/3 as 'edit'

# Edit drivers/gpu/drm/msm/dp/dp_ctrl.c
# Remove the blank line after phy_configure(phy, &ctrl->phy_opts);

git add drivers/gpu/drm/msm/dp/dp_ctrl.c
git commit --amend --no-edit
git rebase --continue

Reproduce locally:

cd /path/to/kernel
./scripts/check-patch-compliance.sh --kernel-src . \
  --base 9269f33cd5d258c7bc5d0f8ec6e0076174c053a1 \
  --head 9991087762701b3dcd24b736b0bff6d53763f341

Verdict

Fix 1 blocker before merge: Remove the extra blank line in commit 2/3 to match the upstream patch exactly. The whitespace change is not present in the upstream commit and causes the check-patch-compliance failure.

All other checkers passed successfully:

  • ✅ Commits 1/3 and 3/3 passed check-patch-compliance
  • ✅ All three commits have proper UPSTREAM: prefix tags
  • ✅ No coding style, sparse, or UAPI issues detected

@qlijarvis

Copy link
Copy Markdown

LAVA Failed Case Triage Summary

PR: #1083

Job 223501 | SoC lemans-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/223501

Failed test cases in LAVA job 223501 (SoC: lemans-evk).

  Case 1: ** Probe_Failure_Check
  1. Failed case: ** Probe_Failure_Check
  2. Root cause: ** Pre-existing platform configuration issue on lemans-evk: 4 PMIC temp-alarm devices remain in permanent deferred probe state due to missing thermal zone bindings or NVMEM calibration data provider in device tree; regulatory.db firmware missing from rootfs (non-critical). Neither issue is related to PR Remove brownout risks from DPU #1083 display driver changes.
  3. Possible fix: For temp-alarm deferred probe: Add thermal zone bindings for SA8775P PMICs in arch/arm64/boot/dts/qcom/sa8775p-lemans-evk.dts referencing the 4 temp-alarm devices, and ensure NVMEM cells for temperature calibration are defined. For regulatory.db: Add regulatory.db firmware file to rootfs /lib/firmware/ directory (or suppress this check as it's non-critical with built-in fallback). These are platform bring-up tasks unrelated to the PR under test.
  4. Detail analysis attachment: failed_case_job223501_1_detailed.md
  Case 2: smmu
  1. Failed case: smmu
  2. Root cause: Video codec device (aa00000.video-codec) is not attached to any IOMMU group on lemans-evk; the smmu test validates that critical masters (GPU, USB, Display, Video) are protected by IOMMU, and the video codec failed this check.
  3. Possible fix: Verify that the video codec device tree node includes the correct iommus property pointing to the SMMU; if the DT is correct, check whether the video codec driver probe is failing or deferred, preventing IOMMU attachment; this is likely a pre-existing platform issue unrelated to the PR (which only touches DRM display power management).
  4. Detail analysis attachment: failed_case_job223501_2_detailed.md
  Case 3: 0_qcom-next-ci-premerge-tests
  1. Failed case: 0_qcom-next-ci-premerge-tests
  2. Root cause: LAVA test definition marked as failed because two sub-tests failed: (1) Probe_Failure_Check detected deferred probe issues for PMIC temp-alarm devices and firmware load errors for regulatory.db and Bluetooth firmware, and (2) smmu test detected that video codec device aa00000.video-codec is missing IOMMU group attachment.
  3. Possible fix: The failures are unrelated to the PR changes (which modify DRM/MSM power management). For Probe_Failure_Check: the deferred probe warnings for PMIC temp-alarms and firmware load errors are pre-existing platform/configuration issues, not regressions. For smmu: investigate why the video codec device at aa00000.video-codec is not attached to an IOMMU group on lemans-evk — check device tree IOMMU bindings and video codec driver probe sequence.
  4. Detail analysis attachment: failed_case_job223501_3_detailed.md
Job 223502 | SoC shikra-iqs-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/223502

Failed test cases in LAVA job 223502 (SoC: shikra-iqs-evk).

  Case 1: ** GIC Test Script Bug — Hardcoded 8-CPU Assumption on 4-CPU Platform
  1. Failed case: ** GIC Test Script Bug — Hardcoded 8-CPU Assumption on 4-CPU Platform
  2. Root cause: ** The GIC test script (run.sh line 75) is hardcoded to validate timer interrupts for CPUs 0-7, but shikra-iqs-evk has only 4 CPUs (0-3). When the script attempts to parse /proc/interrupts columns for non-existent CPUs 4-7, it encounters text fields ("GICv3", "Level", "arch_timer") instead of interrupt counts, causing bash integer comparison errors [: GICv3: integer expected. CPUs 0-3 passed validation correctly; the failure is a false positive caused by the test script bug, not a kernel issue.
  3. Possible fix: Update the GIC test script to dynamically detect the number of online CPUs from /sys/devices/system/cpu/online instead of hardcoding a loop over CPUs 0-7. Parse the /proc/interrupts header to determine the actual number of CPU columns present, and only validate timer increments for CPUs that exist on the platform. This fix will make the test portable across platforms with different CPU counts (4, 6, 8, 12, etc.).
  4. Detail analysis attachment: failed_case_job223502_1_detailed.md
  Case 2: Probe_Failure_Check — Pre-existing Platform Driver Probe Failures (Not PR-Introduced)
  1. Failed case: Probe_Failure_Check — Pre-existing Platform Driver Probe Failures (Not PR-Introduced)
  2. Root cause: The test detected 7 pre-existing probe failures on shikra-iqs-evk: (1) coresight-etm4x etm0-3 fail with -EINVAL (-22) due to missing/incorrect ETM device tree configuration, (2) cpufreq-dt fails with -EEXIST (-17) because another cpufreq driver already registered, (3) regulatory.db firmware missing (-ENOENT/-2) which is expected on minimal rootfs, (4) lt9611c display bridge fails with -EIO (-5) due to I2C/GPI DMA transaction errors indicating hardware/board wiring issue, (5-7) audio codec deferred probe issues. None of these failures are related to the PR changes (DRM/MSM dev_pm_opp_set_rate(0) removal in DPU/DP/DSI drivers).
  3. Possible fix: No action required for this PR. These are known platform bring-up issues on shikra-iqs-evk that exist independently of the PR changes. The PR modifies only DRM display driver power management (removing incorrect dev_pm_opp_set_rate(0) calls) and does not touch coresight, cpufreq, regulatory, lt9611c, or audio subsystems. To resolve the underlying platform issues: (1) Fix ETM device tree nodes for shikra, (2) Resolve cpufreq driver conflict in platform code, (3) Add regulatory.db to rootfs if WiFi regulatory domain enforcement is needed, (4) Debug lt9611c I2C/GPI hardware path, (5) Fix audio codec clock dependencies. Mark this test case as PASS for PR validation purposes.
  4. Detail analysis attachment: failed_case_job223502_2_detailed.md
  Case 3: USBHost
  1. Failed case: USBHost
  2. Root cause: Test infrastructure limitation — no USB host controller driver initialized on shikra-iqs-evk; USB core registered but no dwc3/xhci probe messages in kernel log, resulting in zero enumerated USB devices.
  3. Possible fix: This is a known test environment limitation for shikra-iqs-evk where USB host functionality is not configured or no USB devices are physically connected; mark this test as SKIP for this board or ensure USB host hardware/DT configuration is enabled and USB devices are connected to the test fixture.
  4. Detail analysis attachment: failed_case_job223502_3_detailed.md
  Case 4: BT_SCAN — Environmental Test Dependency Failure
  1. Failed case: BT_SCAN — Environmental Test Dependency Failure
  2. Root cause: The BT_SCAN test failed because no Bluetooth devices were discoverable in the LAVA lab environment during the 3 scan attempts (each with 15-second scan window). The Bluetooth stack is fully functional (BT_FW_KMD_Service and BT_ON_OFF both passed), hci0 adapter is operational (powered, discovering capability confirmed), and bluetoothctl successfully initiated scans, but no nearby Bluetooth devices responded. This is an environmental/infrastructure issue, not a kernel regression. The PR changes only DRM/MSM display drivers (DPU, DP, DSI) and cannot affect Bluetooth scanning functionality.
  3. Possible fix: This is a test environment issue, not a code defect. Recommended actions: (1) Verify that Bluetooth beacon/test devices are powered on and within range of the shikra-iqs-evk board in the LAVA lab; (2) Check if other concurrent LAVA jobs on the same or nearby boards are also experiencing BT_SCAN failures (indicating lab-wide RF interference or beacon unavailability); (3) If this is a known intermittent lab issue, add BT_SCAN to the known-benign-failures suppression list with a condition that BT_ON_OFF must pass; (4) Re-trigger the CI job to confirm if the issue is transient.
  4. Detail analysis attachment: failed_case_job223502_4_detailed.md
  Case 5: KVM_Driver — /dev/kvm not available (HYP mode unavailable)
  1. Failed case: KVM_Driver — /dev/kvm not available (HYP mode unavailable)
  2. Root cause: KVM initialization failed because the Shikra IQS EVK platform is running with Gunyah hypervisor (version gunyah-mobile-ad1fb25c6), which already occupies EL2/HYP mode. KVM requires exclusive access to HYP mode and cannot coexist with another hypervisor. The kernel correctly detected this condition and printed "kvm [1]: HYP mode not available" at boot, preventing /dev/kvm creation.
  3. Possible fix: This is a platform configuration issue, not a kernel regression introduced by PR Remove brownout risks from DPU #1083 (which only modifies DRM/MSM display driver power management). To enable KVM on this platform: (1) disable Gunyah hypervisor in the firmware/bootloader configuration, or (2) use a firmware build without Gunyah, or (3) accept that KVM tests are not applicable on Gunyah-enabled platforms and mark these tests as SKIP when Gunyah is detected. The PR changes are unrelated to this failure.
  4. Detail analysis attachment: failed_case_job223502_5_detailed.md
  Case 6: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: This is a pre-existing platform limitation unrelated to PR Remove brownout risks from DPU #1083 (which only changes DRM/MSM display drivers). To enable KVM on Shikra IQS EVK: (1) Verify the bootloader (ABL/UEFI) grants EL2 access to Linux kernel at boot time, (2) Check if a hypervisor (Gunyah) is already occupying EL2 and preventing KVM initialization, (3) If this platform does not support KVM in the current firmware configuration, mark KVM tests as "expected skip" for this board in the LAVA test suite.
  4. Detail analysis attachment: failed_case_job223502_6_detailed.md
  Case 7: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Short-term: Skip qcom_hwrng test on Shikra or blacklist qcom_rng module (modprobe.blacklist=qcom_rng) to prevent crash. Long-term: Add missing clock/power-domain bindings for RNG hardware in Shikra device tree (arch/arm64/boot/dts/qcom/sm8650*.dts), add runtime PM support to qcom_rng driver to enable clocks before register access, and add probe-time hardware validation to fail gracefully if RNG hardware is inaccessible.
  4. Detail analysis attachment: failed_case_job223502_7_detailed.md
  Case 8: Kernel Crash — Synchronous External Abort in qcom_rng driver
  1. Failed case: Kernel Crash — Synchronous External Abort in qcom_rng driver
  2. Root cause: Hardware RNG (qcom_rng) driver attempted MMIO register read at address 0xffff800082c7d004 during /dev/hwrng read operation, triggering synchronous external abort (ESR 0x96000010). The hardware bus rejected the memory access, indicating the RNG hardware block was not properly powered/clocked or the MMIO mapping is invalid for this SoC variant (Shikra IQS EVK). This is a pre-existing platform issue unrelated to the PR (PR only modifies DRM/MSM display power management).
  3. Possible fix: Verify qcom_rng device tree node for Shikra (shikra-iqs-evk.dts) includes correct MMIO base address, clock references, and power domain bindings. Check if RNG hardware block requires explicit clock/power enablement before register access. Add runtime PM calls or clock enable in qcom_rng_read() before MMIO access. If RNG hardware is not functional on this platform, disable the qcom_hwrng test in the CI test suite for shikra-iqs-evk until the hardware/firmware issue is resolved.
  4. Detail analysis attachment: failed_case_job223502_8_detailed.md
  Case 9: ** Kernel Crash — synchronous external abort in qcom_rng driver
  1. Failed case: ** Kernel Crash — synchronous external abort in qcom_rng driver
  2. Root cause: ** Hardware access fault (synchronous external abort) in qcom_rng_read+0xc4 during register read operation while servicing /dev/hwrng read request from the qcom_hwrng test. The fault indicates the RNG hardware block was not accessible (powered down, clock gated, or address unmapped) when the driver attempted to read from it. This is a pre-existing kernel/firmware bug unrelated to the PR's DRM/MSM display changes.
  3. Possible fix: This is NOT a PR-introduced regression. The qcom_rng driver crash is a pre-existing issue on shikra-iqs-evk. Recommended actions: (1) Skip or disable the qcom_hwrng test on shikra-iqs-evk until the RNG driver/hardware issue is root-caused and fixed. (2) Investigate qcom_rng driver power management and clock dependencies on SM8775 (Shikra) — the driver may be missing runtime PM calls or clock enable/disable sequences specific to this SoC. (3) Check if RNG hardware block requires explicit power domain or interconnect vote that is missing in the device tree or driver. (4) Re-run the PR validation excluding the qcom_hwrng test to verify the display changes do not introduce regressions.
  4. Detail analysis attachment: failed_case_job223502_9_detailed.md
  Case 10: Kernel Crash — Synchronous External Abort in qcom_rng driver
  1. Failed case: Kernel Crash — Synchronous External Abort in qcom_rng driver
  2. Root cause: Hardware random number generator (qcom_rng) driver triggered a synchronous external abort (bus error) at offset 0xc4 in qcom_rng_read() when the qcom_hwrng test attempted to read entropy from /dev/hwrng. The fault indicates the RNG hardware block is either not powered, not clocked, or the MMIO region is inaccessible on this shikra-iqs-evk board. The system panicked after the crash, preventing test completion and causing the 40-minute LAVA timeout.
  3. Possible fix: This is a pre-existing platform/hardware issue unrelated to PR Remove brownout risks from DPU #1083 (which modifies DRM display drivers). The qcom_rng driver requires proper power domain, clock, and IOMMU configuration for the RNG hardware block on shikra. Recommended actions: (1) verify RNG device tree node includes correct power-domains, clocks, and iommus properties for shikra SoC; (2) check if RNG hardware block is present and functional on this shikra-iqs-evk board revision; (3) if RNG is not supported on this board, disable CONFIG_HW_RANDOM_QCOM or mark the DT node as status="disabled" to prevent driver probe; (4) re-run the LAVA job after applying the fix to confirm test suite completes without crash.
  4. Detail analysis attachment: failed_case_job223502_10_detailed.md
  Case 11: Kernel Crash — Synchronous External Abort in qcom_rng driver
  1. Failed case: Kernel Crash — Synchronous External Abort in qcom_rng driver
  2. Root cause: Hardware register access fault in qcom_rng_read() at offset +0xc4 when reading from MMIO address 0x001036628887b000 during /dev/hwrng read operation; synchronous external abort (ESR 0x96000010) indicates the hardware block is not powered/clocked or the MMIO mapping is invalid for shikra-iqs-evk platform
  3. Possible fix: Verify qcom_rng device tree node for shikra (SM8750) includes correct reg address, required clocks, and power domain bindings; confirm RNG hardware block is enabled and accessible on this SoC variant; if RNG is not supported on shikra-iqs-evk, disable the qcom_hwrng test in the CI test suite for this platform
  4. Detail analysis attachment: failed_case_job223502_11_detailed.md
Job 223503 | SoC purwa-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/223503

Failed test cases in LAVA job 223503 (SoC: purwa-evk).

  Case 1: ** Probe_Failure_Check
  1. Failed case: ** Probe_Failure_Check
  2. Root cause: ** Pre-existing platform-specific probe failures on purwa-evk board for non-critical subsystems (qcom_qseecom_uefisecapp, qcom-pcie, qcom-spmi-lpg, regulatory.db firmware). These failures are unrelated to PR Remove brownout risks from DPU #1083, which only modifies DRM/MSM display driver power management and does not touch any of the failing drivers or subsystems.
  3. Possible fix: Add platform-specific suppression rules to Probe_Failure_Check test for purwa-evk to exclude these known-benign probe failures: qcom_qseecom_uefisecapp error -16 (UEFI secure app conflict), qcom-pcie 1bf8000/1bd0000 error -61 (unpopulated PCIe slots), qcom-spmi-lpg error -22 (DT multi-LED config mismatch), and regulatory.db error -2 (missing optional firmware). These represent board configuration issues, not kernel regressions introduced by the PR.
  4. Detail analysis attachment: failed_case_job223503_1_detailed.md
  Case 2: smmu
  1. Failed case: smmu
  2. Root cause: The SMMU test expects USB wrapper devices (a0f8800.usb, a2f8800.usb, a4f8800.usb, a6f8800.usb, a8f8800.usb) and video codec device (aa00000.video-codec) to be attached to IOMMU groups, but these devices are not probed/bound to drivers on the purwa-evk platform, causing the test's critical master validation to fail. This is a pre-existing platform/test configuration issue unrelated to the PR's DRM/MSM display driver changes.
  3. Possible fix: Update the SMMU test's critical master list for purwa-evk to exclude USB wrapper devices and video codec device that are not enabled/probed on this platform, or update the device tree/kernel config to enable these devices if they should be functional.
  4. Detail analysis attachment: failed_case_job223503_2_detailed.md
  Case 3: KVM_Driver — /dev/kvm device node not available
  1. Failed case: KVM_Driver — /dev/kvm device node not available
  2. Root cause: KVM driver initialization failed because the platform is running under the Gunyah hypervisor in a non-nested virtualization configuration where HYP mode (EL2) is not available to the Linux kernel. The kernel message "kvm [1]: HYP mode not available" at boot indicates KVM cannot initialize because EL2 is already claimed by the Gunyah hypervisor.
  3. Possible fix: This is not a regression introduced by PR Remove brownout risks from DPU #1083 (which only modifies DRM/MSM display power management). The failure is a platform configuration issue specific to purwa-evk running under Gunyah hypervisor. To enable KVM on this platform, either: (1) configure the hypervisor to support nested virtualization (if supported by Gunyah), or (2) exclude KVM tests from the purwa-evk test suite as this platform does not support KVM in its current hypervisor configuration.
  4. Detail analysis attachment: failed_case_job223503_3_detailed.md
  Case 4: ** KVM Driver Initialization Failure — HYP mode not available
  1. Failed case: ** KVM Driver Initialization Failure — HYP mode not available
  2. Root cause: ** KVM ARM64 driver initialization fails because the Purwa IoT EVK platform does not support or enable EL2 (hypervisor exception level). The hardware/firmware reports "HYP mode not available" during boot, preventing KVM from creating /dev/kvm and causing all KVM-dependent tests to fail at their availability gate.
  3. Possible fix: This is a platform configuration issue, not a kernel bug. To enable KVM on Purwa: (1) Verify the platform hardware supports ARMv8 virtualization extensions; (2) Update firmware (ABL/XBL) to enable EL2 for non-secure world; (3) Ensure TrustZone configuration allows EL2 access; (4) If the platform is not intended to support virtualization, mark KVM tests as "not applicable" for this target in the CI test matrix.
  4. Detail analysis attachment: failed_case_job223503_4_detailed.md
  Case 5: KVM_Infra — /dev/kvm device node unavailable
  1. Failed case: KVM_Infra — /dev/kvm device node unavailable
  2. Root cause: KVM driver initialization failed because the system is not running in HYP (EL2) mode; kernel log shows "kvm [1]: HYP mode not available" at boot, preventing /dev/kvm creation. This is a platform/firmware limitation on the Purwa IoT EVK board, not a kernel regression.
  3. Possible fix: This is not a PR-introduced failure (PR modifies only DRM/MSM display drivers). The Purwa IoT EVK platform does not boot Linux at EL2, which is a prerequisite for KVM. To enable KVM: (1) verify the bootloader/firmware boots Linux at EL2 (check UEFI/ABL configuration), (2) if the platform does not support EL2 boot, mark KVM tests as "not applicable" for this board in the LAVA test suite, or (3) use a different platform that supports virtualization (e.g., boards with hypervisor support).
  4. Detail analysis attachment: failed_case_job223503_5_detailed.md
  Case 6: Kernel Crash — Watchdog Hard LOCKUP (cpuidle hang)
  1. Failed case: Kernel Crash — Watchdog Hard LOCKUP (cpuidle hang)
  2. Root cause: CPU6 experienced a hard lockup while in cpuidle state (stuck at cpuidle_enter_state+0xf8/0x550 for >20 seconds), triggering the non-secure watchdog on CPU5; the system recovered and continued test execution, but LAVA marked the test suite as failed due to the watchdog event being logged during the qcom_hwrng test run.
  3. Possible fix: This is a transient cpuidle hang, not directly caused by the PR changes (which modify DRM/MSM display power management via dev_pm_opp_set_rate). The lockup occurred during an unrelated hardware RNG test. Recommended action: re-trigger the CI job to confirm this is not reproducible; if it recurs, investigate cpuidle driver state machine on purwa-evk (X5121) platform, specifically the interaction between deep idle states and hardware interrupt delivery during the qcom_hwrng test workload.
  4. Detail analysis attachment: failed_case_job223503_6_detailed.md
Job 223504 | SoC qcs9100-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/223504

Failed test cases in LAVA job 223504 (SoC: qcs9100-ride).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Test detected unresolved deferred probe for four PMIC temp-alarm devices (c440000.spmi:pmic@{0,2,4,6}:temp-alarm@a00) and two probe failures: regulatory.db firmware missing (benign cfg80211 regulatory database) and Aquantia AQR115C Ethernet PHY probe failure with error -22 (EINVAL, likely missing firmware-name DT property). These failures are pre-existing platform/configuration issues unrelated to the PR's DRM/MSM display power management changes.
  3. Possible fix: The deferred probe and Aquantia PHY failure are pre-existing board/DT configuration issues not introduced by this PR (which only modifies DRM display driver power management). The regulatory.db failure is benign (cfg80211 falls back to built-in regulatory rules). To resolve: (1) investigate why qpnp-temp-alarm driver dependency (likely IIO VADC channel) is not probing on qcs9100-ride; (2) add missing "firmware-name" DT property for Aquantia PHY or verify PHY firmware is present in rootfs. These are board-specific fixes outside PR scope.
  4. Detail analysis attachment: failed_case_job223504_1_detailed.md
  Case 2: ** smmu
  1. Failed case: ** smmu
  2. Root cause: ** Video codec device aa00000.video-codec is missing IOMMU group attachment on qcs9100-ride platform — this is a pre-existing device tree or platform configuration issue, not introduced by the PR (which only modifies DRM display power management).
  3. Possible fix: This is not a PR-blocking issue. The PR changes are unrelated to video codec IOMMU configuration. To resolve the underlying platform issue: verify the video codec DT node at aa00000.video-codec includes correct iommus property referencing the appropriate SMMU phandle and stream ID, then regenerate the DTB and retest.
  4. Detail analysis attachment: failed_case_job223504_2_detailed.md
  Case 3: USBHost
  1. Failed case: USBHost
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Verify the qcs9100-ride board's USB port connections in the LAVA lab. If the test requires a USB device to be permanently connected for CI validation, ensure one is physically attached (e.g., USB flash drive, keyboard, or hub with downstream devices). If this is a known lab limitation, add USBHost to the known-benign-failures suppression list or mark it as SKIP when no USB device is expected. The PR changes (DRM/MSM display driver fixes) are unrelated to USB functionality and did not cause this failure.
  4. Detail analysis attachment: failed_case_job223504_3_detailed.md
  Case 4: Ethernet_Basic_Validation
  1. Failed case: Ethernet_Basic_Validation
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Add the firmware-name property to the Aquantia AQR115C PHY device tree node under the stmmac-0 MDIO bus. The property should specify the path to the Aquantia PHY firmware file (typically Rhe-05.06-Candidate9-AQR_Mediatek_23B_P5_ID45824_LNXDRIVER.cld or similar). This is a device tree configuration issue on the qcs9100-ride platform, not a regression introduced by PR Remove brownout risks from DPU #1083 (which only modifies DRM/MSM display driver code).
  4. Detail analysis attachment: failed_case_job223504_4_detailed.md
  Case 5: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: KVM driver initialization fails because the qcs9100-ride platform runs Linux as a guest under Gunyah hypervisor (version gunyah-cdfb73831), which already claims EL2/HYP mode. KVM requires direct EL2 access to create /dev/kvm, but Linux is running at EL1 as a Gunyah guest VM, making "HYP mode not available" the expected behavior on this platform configuration.
  3. Possible fix: This is not a kernel regression or bug—it is the expected behavior for this platform configuration. The KVM_Driver test should be skipped on qcs9100-ride (and any platform running Gunyah hypervisor) because KVM host functionality is architecturally incompatible with running as a guest VM. Update the LAVA test suite to detect Gunyah presence (check for "Hypervisor cold boot, version: gunyah" in dmesg or /sys/hypervisor/type) and skip KVM host tests on such platforms.
  4. Detail analysis attachment: failed_case_job223504_5_detailed.md
  Case 6: KVM_EL2_DTB — Platform Configuration Limitation (HYP mode unavailable)
  1. Failed case: KVM_EL2_DTB — Platform Configuration Limitation (HYP mode unavailable)
  2. Root cause: The qcs9100-ride platform firmware/bootloader boots the kernel in EL1 (kernel mode) instead of EL2 (hypervisor mode), preventing KVM initialization. The kernel message kvm [1]: HYP mode not available at boot (timestamp 3.868720s) confirms the CPU is not running in EL2, which is a mandatory requirement for ARM64 KVM functionality. This is a pre-existing platform limitation, not a regression introduced by the PR (which only modifies DRM/MSM display driver power management code).
  3. Possible fix: Mark KVM tests as SKIP (not FAIL) on qcs9100-ride in the LAVA test definition, as this platform does not support virtualization. Add a platform check: if [ "$PLATFORM" = "qcs9100-ride" ]; then echo "[SKIP] KVM not supported (no EL2)"; exit 0; fi. If KVM support is required, the bootloader/firmware must be updated to boot the kernel in EL2, which may require TrustZone/secure firmware changes and platform security policy updates.
  4. Detail analysis attachment: failed_case_job223504_6_detailed.md
  Case 7: KVM_Infra — Platform Configuration Incompatibility (Not Applicable)
  1. Failed case: KVM_Infra — Platform Configuration Incompatibility (Not Applicable)
  2. Root cause: qcs9100-ride platform runs Gunyah Type-1 hypervisor which takes exclusive control of ARM EL2 (HYP mode). KVM requires EL2 access to initialize and create /dev/kvm. Kernel correctly detects "HYP mode not available" (log line 3126) and aborts KVM initialization. This is expected behavior on Gunyah-enabled platforms, not a kernel bug or PR-introduced regression.
  3. Possible fix: Update KVM test suite to detect Gunyah hypervisor presence (check /sys/hypervisor/type or device tree compatible = "gunyah-hypervisor") and report SKIP (not applicable) instead of FAIL. Alternatively, exclude KVM tests from qcs9100-ride CI test matrix, as this platform uses Gunyah for virtualization and does not support KVM by design.
  4. Detail analysis attachment: failed_case_job223504_7_detailed.md
  Case 8: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM virtualization is not supported on the qcs9100-ride (LeMans Ride Rev3) platform — kernel reports "HYP mode not available" during boot, preventing /dev/kvm device creation.
  3. Possible fix: This is a pre-existing platform limitation, not a PR-introduced regression. The PR changes only DRM/MSM display driver power management (dev_pm_opp_set_rate) and does not affect KVM/virtualization. Mark this test as expected-fail for qcs9100-ride, or exclude KVM tests from this platform's CI configuration.
  4. Detail analysis attachment: failed_case_job223504_8_detailed.md
Job 223505 | SoC qcs8300-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/223505

Failed test cases in LAVA job 223505 (SoC: qcs8300-ride).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Two pre-existing platform configuration issues detected: (1) Aquantia AQR115C Ethernet PHY probe fails with -EINVAL due to missing firmware-name property in qcs8300-ride device tree; (2) regulatory.db firmware load fails (benign — WiFi functional tests passed, indicating fallback to built-in regulatory data works correctly). Neither failure is related to the PR changes (DRM/MSM display driver power management fixes).
  3. Possible fix: (1) Add firmware-name = "Rhe-05.06-Candidate9-AQR_Mediatek_23B_P5_ID45824_LCLVER1.cld"; property to the Aquantia PHY node in arch/arm64/boot/dts/qcom/qcs8300-ride.dts to resolve the Ethernet PHY probe failure. (2) No action required for regulatory.db — this is expected behavior when the optional firmware file is absent; WiFi operates correctly using built-in regulatory data.
  4. Detail analysis attachment: failed_case_job223505_1_detailed.md
  Case 2: ** USBHost (Test Infrastructure Issue — No USB Devices Connected)
  1. Failed case: ** USBHost (Test Infrastructure Issue — No USB Devices Connected)
  2. Root cause: ** The USBHost test expects functional USB peripheral devices (keyboard, mouse, storage) to be physically connected to the qcs8300-ride board, but only the USB root hub is present. The USB host controller (xHCI) is working correctly — it successfully loaded, registered the USB bus, and enumerated the root hub. However, no external USB devices are connected to the board in the LAVA lab environment, causing the test to fail with "Only USB hubs detected, no functional USB devices."
  3. Possible fix: This is not a kernel bug or PR-introduced regression. The fix is a LAVA lab infrastructure action: physically connect a USB peripheral device (e.g., USB keyboard, mouse, or storage device) to the qcs8300-ride board's USB host port. If USB devices are already supposed to be connected, verify the physical USB cable connection and check if the USB device is powered and functional. Alternatively, if this test is not applicable to qcs8300-ride hardware configuration, mark the test as "skip" or "not applicable" for this board in the LAVA job definition.
  4. Detail analysis attachment: failed_case_job223505_2_detailed.md
  Case 3: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: KVM driver failed to initialize during boot — /dev/kvm device node was never created despite CONFIG_KVM being enabled. No KVM initialization messages appear in the kernel boot log, indicating KVM's init function either did not run or failed silently. On qcs8300-ride (Monaco), this typically occurs when the platform firmware does not enable EL2 (hypervisor mode) or when KVM detects incompatible CPU features during early initialization and aborts without creating the device node.
  3. Possible fix: Verify that the qcs8300-ride platform firmware enables EL2 (hypervisor exception level) and that the bootloader does not disable virtualization extensions. Check for silent KVM initialization failures by enabling KVM debug messages (add kvm-arm.mode= or loglevel=8 to kernel command line) and inspect dmesg for KVM init errors. If EL2 is unavailable on this platform, KVM cannot function and the test expectation should be adjusted to skip KVM tests on qcs8300-ride. The PR changes (DRM/MSM display driver OPP fixes) are unrelated to this failure — this is a pre-existing platform/configuration issue, not a regression introduced by the PR.
  4. Detail analysis attachment: failed_case_job223505_3_detailed.md
  Case 4: KVM Test Failure — Test Environment Configuration Mismatch
  1. Failed case: KVM Test Failure — Test Environment Configuration Mismatch
  2. Root cause: KVM cannot initialize on qcs8300-ride because the platform runs under the Gunyah hypervisor (nested virtualization). CONFIG_KVM is enabled in the kernel, but KVM requires EL2 (hypervisor mode) privileges which are already claimed by Gunyah. The kernel runs at EL1 as a guest, preventing /dev/kvm device node creation. This is expected behavior, not a kernel bug.
  3. Possible fix: Disable KVM tests for qcs8300-ride in the LAVA test suite, or add a pre-flight check to skip KVM tests when running under a hypervisor. The test should detect Hypervisor cold boot in dmesg and skip KVM validation. Alternatively, disable CONFIG_KVM in the kernel configuration for platforms that run under Gunyah.
  4. Detail analysis attachment: failed_case_job223505_4_detailed.md
  Case 5: KVM_Infra — KVM device node unavailable (nested virtualization not supported)
  1. Failed case: KVM_Infra — KVM device node unavailable (nested virtualization not supported)
  2. Root cause: The qcs8300-ride target is running under Gunyah hypervisor (guest VM mode), preventing KVM from initializing because the kernel runs at EL1 without access to EL2 (hypervisor exception level) required for KVM operation. CONFIG_KVM is enabled but /dev/kvm device node is never created because KVM initialization silently fails when nested virtualization is not available or not enabled.
  3. Possible fix: This is a platform/configuration limitation, not a kernel regression. To enable KVM testing on qcs8300-ride: (1) verify Gunyah hypervisor supports nested virtualization and enable it in the hypervisor configuration, OR (2) run the kernel directly on bare metal without Gunyah hypervisor for KVM testing, OR (3) exclude KVM tests from the CI test suite for qcs8300-ride targets running under Gunyah, as nested virtualization is not a standard configuration for this platform.
  4. Detail analysis attachment: failed_case_job223505_5_detailed.md
  Case 6: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: /dev/kvm device node is not created because the qcs8300-ride platform is running under the Gunyah Type-1 hypervisor (version gunyah-cdfb73831), which occupies EL2 and prevents KVM from initializing. KVM requires direct EL2 access to function, but on this platform EL2 is already claimed by Gunyah. CONFIG_KVM is enabled in the kernel configuration, but the KVM driver cannot create /dev/kvm because it cannot initialize without EL2 access.
  3. Possible fix: This is not a PR-introduced regression — it is a known platform limitation of qcs8300-ride when running under Gunyah hypervisor. The KVM tests should be skipped on qcs8300-ride (and other Gunyah-based platforms) in the LAVA test plan, as KVM cannot function on platforms where a Type-1 hypervisor already occupies EL2. Update the test plan to conditionally skip KVM_Driver, KVM_EL2_DTB, and KVM_Infra tests when the platform is qcs8300-ride or when Gunyah hypervisor is detected in the boot log.
  4. Detail analysis attachment: failed_case_job223505_6_detailed.md
Job 223506 | SoC hamoa-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/223506

Failed test cases in LAVA job 223506 (SoC: hamoa-evk).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Three pre-existing platform-specific probe failures unrelated to PR changes: (1) qcom_qseecom_uefisecapp probe fails with -EBUSY due to TrustZone/secure world resource conflict on hamoa-evk, (2) qcom-spmi-lpg probe fails with -EINVAL due to invalid multi-LED device tree configuration for PMIC PWM, (3) regulatory.db firmware file missing from rootfs (benign - WiFi regulatory database is optional).
  3. Possible fix: These are pre-existing platform/configuration issues not introduced by this PR (which only modifies DRM display driver power management). To resolve: (1) verify TZ firmware compatibility and qseecom device availability for hamoa-evk, (2) fix the PMIC PWM multi-LED device tree "reg" property in hamoa-evk DTS, (3) optionally add regulatory.db to rootfs if WiFi regulatory enforcement is required. The PR itself is not the cause and should not be blocked by this test failure.
  4. Detail analysis attachment: failed_case_job223506_1_detailed.md
  Case 2: smmu
  1. Failed case: smmu
  2. Root cause: Six critical USB DWC3 controller devices (a0f8800.usb, a2f8800.usb, a4f8800.usb, a6f8800.usb, a8f8800.usb) and one video codec device (aa00000.video-codec) on hamoa-evk are missing IOMMU group attachments, failing the SMMU test's critical master protection validation despite SMMU subsystem being functional.
  3. Possible fix: Add missing iommus property entries in the hamoa-evk device tree for USB DWC3 controllers at addresses 0xa0f8800, 0xa2f8800, 0xa4f8800, 0xa6f8800, 0xa8f8800 and video codec at 0xaa00000, referencing the appropriate SMMU phandle and stream IDs to enable IOMMU protection for these critical masters.
  4. Detail analysis attachment: failed_case_job223506_2_detailed.md
  Case 3: KVM_Driver — /dev/kvm device node not present
  1. Failed case: KVM_Driver — /dev/kvm device node not present
  2. Root cause: KVM driver initialization failed with "HYP mode not available" because the Gunyah hypervisor (gunyah-mobile-c487961e9) is already running at EL2 on the hamoa-evk platform. KVM requires exclusive EL2 access to create the /dev/kvm device node, but EL2 is occupied by Gunyah, preventing KVM from initializing.
  3. Possible fix: This is a platform configuration issue, not a kernel regression. The hamoa-evk board is configured to boot with Gunyah hypervisor at EL2, which is incompatible with KVM. To enable KVM: (1) reconfigure the board firmware to boot without Gunyah hypervisor, OR (2) mark KVM tests as "not applicable" for Gunyah-enabled platforms in the CI test matrix, OR (3) use a different test platform that boots Linux directly at EL2 without a hypervisor.
  4. Detail analysis attachment: failed_case_job223506_3_detailed.md
  Case 4: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: KVM driver initialization failed because Gunyah hypervisor is already running at EL2 (ARM Exception Level 2), preventing KVM from accessing HYP mode. On ARM64 architecture, only one entity can occupy EL2 at a time. This is a platform configuration issue, not a PR-introduced regression.
  3. Possible fix: This is not a bug to fix - it's an expected architectural limitation. To enable KVM on hamoa-evk, the platform must be reconfigured to boot without Gunyah hypervisor. This requires bootloader/firmware changes to not load Gunyah at EL2. Alternatively, accept that KVM tests are not applicable on Gunyah-enabled platforms and exclude them from the test suite for this SoC configuration.
  4. Detail analysis attachment: failed_case_job223506_4_detailed.md
  Case 5: KVM Infrastructure Test Failure — HYP Mode Unavailable
  1. Failed case: KVM Infrastructure Test Failure — HYP Mode Unavailable
  2. Root cause: KVM driver initialization failed because the Hamoa IoT EVK platform does not have EL2 (HYP mode) available at boot time, as evidenced by kernel message kvm [1]: HYP mode not available at 6.373s into boot. CONFIG_KVM is enabled but the platform firmware/bootloader is not entering the kernel at EL2, preventing KVM from initializing and creating the /dev/kvm device node.
  3. Possible fix: This is a pre-existing platform/firmware configuration issue unrelated to the PR (which only modifies DRM/MSM display drivers). To enable KVM on Hamoa: (1) verify the bootloader/firmware supports entering the kernel at EL2; (2) check if a hypervisor is already running at EL2 (preventing KVM); (3) update platform firmware/bootloader configuration to enable virtualization extensions and boot at EL2. If KVM support is not required for this platform, mark these tests as expected failures or skip them in the CI configuration for Hamoa.
  4. Detail analysis attachment: failed_case_job223506_5_detailed.md
  Case 6: ** KVM_Infra (Driver Initialization Failure)
  1. Failed case: ** KVM_Infra (Driver Initialization Failure)
  2. Root cause: ** KVM driver initialization failed because EL2 (Hypervisor mode) is not available on the Hamoa IoT EVK platform. The kernel message kvm [1]: HYP mode not available at boot time (6.373238s) indicates the CPU is not running at EL2 or EL2 is disabled by firmware/bootloader configuration. CONFIG_KVM is enabled but /dev/kvm device node cannot be created without EL2 support.
  3. Possible fix: This is a pre-existing platform configuration issue unrelated to the PR (which only modifies DRM/MSM display drivers). To enable KVM on Hamoa EVK: (1) verify the SoC supports virtualization extensions, (2) configure UEFI/bootloader to boot the kernel at EL2 instead of EL1, (3) ensure TrustZone firmware does not disable EL2, and (4) add device tree properties to enable virtualization if required by the platform. If Hamoa EVK does not support EL2, exclude KVM tests from the CI test suite for this platform.
  4. Detail analysis attachment: failed_case_job223506_6_detailed.md
Job 223507 | SoC qcs615-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/223507

Failed test cases in LAVA job 223507 (SoC: qcs615-ride).

  Case 1: login-action
  1. Failed case: login-action
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Verify and correct qcs615-ride device tree memory reservations and SWIOTLB buffer placement; ensure SWIOTLB buffers are allocated within valid DRAM range (typically below 4GB for 32-bit DMA devices); check bootloader memory map passed to kernel; this is a pre-existing platform issue, not a PR regression — PR Remove brownout risks from DPU #1083 changes only DRM/MSM display power management and does not touch USB/DMA/SWIOTLB code paths.
  4. Detail analysis attachment: failed_case_job223507_1_detailed.md
  Case 2: Kernel Crash — Synchronous External Abort during USB DMA
  1. Failed case: Kernel Crash — Synchronous External Abort during USB DMA
  2. Root cause: Synchronous external abort (ESR 0x96000010) in __pi_memcpy_generic during SWIOTLB bounce buffer operation for USB device enumeration. The CPU attempted to access a physical address that triggered a bus-level fault, indicating either an invalid IOMMU/SMMU mapping, a powered-down memory region, or a hardware interconnect issue on qcs615-ride during USB host controller DMA setup.
  3. Possible fix: This crash is likely a pre-existing platform issue unrelated to the PR's display power management changes. Recommended actions: (1) Verify USB host controller and SMMU power domain configuration in qcs615-ride device tree; (2) Check if USB DMA requires specific IOMMU reserved-memory regions that may be missing; (3) Enable SMMU fault debugging (CONFIG_IOMMU_TLBSYNC_DEBUG, CONFIG_ARM_SMMU_TESTBUS_DUMP) and collect SMMU register dumps to identify the faulting address and context; (4) Re-trigger the CI job to confirm reproducibility — if non-reproducible, this may be a transient hardware/timing issue in the LAVA lab.
  4. Detail analysis attachment: failed_case_job223507_2_detailed.md
  Case 3: Kernel Crash — Synchronous External Abort during USB DMA Operation
  1. Failed case: Kernel Crash — Synchronous External Abort during USB DMA Operation
  2. Root cause: Synchronous external abort (ESR 0x96000010) during swiotlb bounce buffer memory copy operation while USB hub worker attempted to enumerate a USB device. The crash occurred in __pi_memcpy_generic when copying from physical address 0xffff000091911020 to bounce buffer 0xffff00007febf000, indicating an invalid or unmapped physical memory access during DMA bounce buffer operation. This is a pre-existing kernel/hardware issue unrelated to the PR changes (which only modify DRM/MSM display driver power management).
  3. Possible fix: This is a pre-existing hardware/kernel issue on qcs615-ride, not introduced by PR1083. The PR modifies only DRM/MSM display subsystem (DPU/DP/DSI power management), while the crash occurs in USB/DMA/SWIOTLB subsystem during device enumeration. Recommended actions: (1) Re-trigger the CI job to confirm reproducibility; (2) If reproducible, investigate qcs615-ride USB controller DMA addressing constraints and SWIOTLB configuration; (3) Check if USB device connected to the board has DMA addressing issues requiring bounce buffers; (4) Verify IOMMU/SMMU configuration for USB controller on qcs615-ride; (5) Consider this a known qcs615-ride platform issue and track separately from PR validation.
  4. Detail analysis attachment: failed_case_job223507_3_detailed.md
  Case 4: Kernel Crash — Synchronous External Abort (USB/DMA)
  1. Failed case: Kernel Crash — Synchronous External Abort (USB/DMA)
  2. Root cause: Synchronous external abort at 7.781s during USB hub enumeration on qcs615-ride; kernel attempted to copy data from an invalid physical address (0x111911020) via swiotlb bounce buffer during USB device descriptor fetch, triggering a fatal bus-level memory access fault in __pi_memcpy_generic called from the DMA mapping path.
  3. Possible fix: This crash is unrelated to the PR patches (which modify DRM display power management). The fault occurs during USB enumeration when the xHCI controller attempts DMA to an invalid physical address. Immediate action: verify USB hardware connectivity and power state on qcs615-ride board 8602020515191D74. Check for known USB/DMA addressing issues on qcs615 platform in kernel 6.18.44-gd2f19f3a90fc. If reproducible, bisect to identify the commit that introduced the invalid DMA address mapping or swiotlb configuration regression.
  4. Detail analysis attachment: failed_case_job223507_4_detailed.md
Job 223508 | SoC monaco-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/223508

Failed test cases in LAVA job 223508 (SoC: monaco-evk).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Multiple probe failures detected: (1) ath11k_pci WiFi driver probe failed with -ETIMEDOUT (-110) due to missing firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin and MHI power-up timeout; (2) deferred probe loop for pinctrl device 3440000.pinctrl and sound device with error "snd-sc8280xp: HS0 MI2S Playback: error getting cpu dai name"; (3) Bluetooth firmware load failures (qca/wcnhpbtfw21.tlv, qca/hpbtfw21.tlv) are benign as BT_ON_OFF test passed. The PR patches modify DPU/DP/DSI power management (dev_pm_opp_set_rate) but are unrelated to WiFi/pinctrl/sound probe failures on monaco-evk.
  3. Possible fix: The probe failures are pre-existing platform/firmware issues not introduced by this PR. (1) For ath11k WiFi: add missing firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin to rootfs firmware directory, or update firmware path in device tree if using alternate firmware variant. (2) For pinctrl/sound deferred probe: verify device tree has correct pinctrl and sound codec DAI configuration for monaco-evk (sc8280xp-based) platform. The PR changes to DPU/DP/DSI dev_pm_opp_set_rate(0) removal do not affect these subsystems and should not block merge.
  4. Detail analysis attachment: failed_case_job223508_1_detailed.md
  Case 2: ** WiFi Driver Probe Failure — ath11k_pci probe timeout due to missing firmware
  1. Failed case: ** WiFi Driver Probe Failure — ath11k_pci probe timeout due to missing firmware
  2. Root cause: ** The ath11k_pci driver probe failed with error -110 (ETIMEDOUT) on monaco-evk because the required WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin is missing from the rootfs firmware directory. The MHI bus attempted to load the firmware during device power-on but received -ENOENT, causing the WCN6855 WiFi chip initialization to time out.
  3. Possible fix: Install the missing ath11k WCN6855 firmware package in the rootfs. For Yocto/meta-qcom builds, ensure linux-firmware-ath11k or the board-specific firmware package (e.g., linux-firmware-qcom-monaco-wifi) is included in the image recipe. For manual installation, copy the firmware files from linux-firmware.git/ath11k/WCN6855/hw2.1/ to /lib/firmware/ath11k/WCN6855/hw2.1/ on the target, ensuring the board-specific subdirectory (nfa765) and amss.bin are present.
  4. Detail analysis attachment: failed_case_job223508_2_detailed.md
  Case 3: ** WiFi Driver Probe Failure — ath11k_pci firmware dependency
  1. Failed case: ** WiFi Driver Probe Failure — ath11k_pci firmware dependency
  2. Root cause: ** ath11k_pci driver probe failed with -ETIMEDOUT (-110) because the required firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin is missing from the rootfs (-ENOENT, error -2). Without firmware, the MHI (Modem Host Interface) subsystem cannot power up the WCN6855 WiFi chip, causing the probe to time out. This is a pre-existing infrastructure/image issue, not introduced by the PR (which only modifies DRM display drivers).
  3. Possible fix: Add the missing firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin to the rootfs image under /lib/firmware/. Verify the firmware package (e.g., linux-firmware-ath11k or vendor-specific firmware package) is included in the Yocto/build recipe for Monaco EVK images.
  4. Detail analysis attachment: failed_case_job223508_3_detailed.md
  Case 4: ** 0_qcom-next-ci-premerge-tests
  1. Failed case: ** 0_qcom-next-ci-premerge-tests
  2. Root cause: ** WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin is missing from the test rootfs image, causing ath11k_pci probe to fail with -ETIMEDOUT (-110) during MHI power-up. This is a test infrastructure/firmware packaging issue unrelated to the PR patches, which only modify DRM/MSM display driver power management code.
  3. Possible fix: Add the missing WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin to the monaco-evk test rootfs image build recipe. Verify the linux-firmware package version includes WCN6855 hw2.1 nfa765 variant firmware, or add it from the appropriate Qualcomm firmware repository.
  4. Detail analysis attachment: failed_case_job223508_4_detailed.md
Job 223509 | SoC qcs6490-rb3gen2

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/223509

Failed test cases in LAVA job 223509 (SoC: qcs6490-rb3gen2).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Update the Probe_Failure_Check test to implement the suppression rules from lava-known-benign-failures.md Rule 2 — suppress WiFi firmware load failures when the WiFi_OnOff functional test passes.
  4. Detail analysis attachment: failed_case_job223509_1_detailed.md
  Case 2: USBHost
  1. Failed case: USBHost
  2. Root cause: USB controller on qcs6490-rb3gen2 is configured in gadget mode (systemd target "Hardware activated USB gadget" reached), but the USBHost test expects host mode with enumerated USB devices. No USB host controller driver (dwc3-host/xhci-hcd) initialization messages appear in the kernel log, and no USB devices are enumerated. The test script prints "Enumerated USB devices..." followed by an empty output, then immediately fails.
  3. Possible fix: Configure the USB controller in host mode for this test by either: (1) updating the device tree to set dr_mode = "host" for the USB controller node, or (2) modifying the test to dynamically switch USB role to host mode before enumeration, or (3) skip this test on platforms where USB is intentionally configured as gadget-only for the test scenario.
  4. Detail analysis attachment: failed_case_job223509_2_detailed.md
  Case 3: ** KVM Driver Initialization Failure — HYP mode not available
  1. Failed case: ** KVM Driver Initialization Failure — HYP mode not available
  2. Root cause: ** KVM driver cannot initialize on qcs6490-rb3gen2 because the platform is running Gunyah hypervisor which occupies EL2 (Exception Level 2). KVM and Gunyah are mutually exclusive — both require EL2 access, and Gunyah is already controlling it. The kernel correctly reports "HYP mode not available" and skips KVM initialization.
  3. Possible fix: This is not a kernel bug or PR regression. The test suite should either (1) skip KVM tests on Gunyah-enabled platforms by detecting Gunyah presence in device tree or boot logs, or (2) use a non-Gunyah firmware build for KVM validation. To enable KVM on this platform, the firmware must be rebuilt without Gunyah hypervisor support, which is a platform configuration decision outside the scope of this PR.
  4. Detail analysis attachment: failed_case_job223509_3_detailed.md
  Case 4: KVM_EL2_DTB — Platform Limitation (KVM Not Supported)
  1. Failed case: KVM_EL2_DTB — Platform Limitation (KVM Not Supported)
  2. Root cause: The qcs6490-rb3gen2 (Kodiak) platform does not support KVM because the CPU boots at EL1 without hypervisor (EL2) mode available. Kernel message at boot: "kvm [1]: HYP mode not available". CONFIG_KVM is enabled but /dev/kvm cannot be created without EL2 support. This is a platform hardware/firmware limitation, not a kernel bug.
  3. Possible fix: Exclude KVM tests from the CI test suite for qcs6490-rb3gen2 (Kodiak) platform, as this SoC does not support virtualization. The test should be skipped or marked as "not applicable" for platforms without EL2 support. Alternatively, update the test framework to detect EL2 availability before running KVM tests and report "SKIP" instead of "FAIL" when HYP mode is unavailable.
  4. Detail analysis attachment: failed_case_job223509_4_detailed.md
  Case 5: ** KVM Infrastructure Test Failure — HYP Mode Not Available
  1. Failed case: ** KVM Infrastructure Test Failure — HYP Mode Not Available
  2. Root cause: ** The qcs6490-rb3gen2 platform boots with Gunyah hypervisor enabled, running the Linux kernel as a guest VM at EL1. KVM requires EL2 (Hypervisor Exception Level) access to provide virtualization capabilities, which is not available when running as a guest under Gunyah. The kernel correctly detects this architectural limitation at boot (line 2962: "kvm [1]: HYP mode not available") and does not create the /dev/kvm device node, causing all KVM-dependent tests to fail.
  3. Possible fix: Disable the KVM_Infra, KVM_Driver, and KVM_EL2_DTB tests for qcs6490-rb3gen2 (and any other Gunyah-enabled platforms) in the LAVA test suite, as nested virtualization is not supported in this configuration. Alternatively, if KVM testing is required, boot the platform without Gunyah hypervisor (if supported by the bootloader/firmware) to allow the kernel to run at EL2.
  4. Detail analysis attachment: failed_case_job223509_5_detailed.md
  Case 6: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: Platform hardware limitation — qcs6490-rb3gen2 does not support EL2/HYP mode required for KVM virtualization. The KVM driver correctly detected this at boot (line 2506: "kvm [1]: HYP mode not available") and did not create the /dev/kvm device node. This is not a kernel regression; it is expected behavior on platforms without virtualization extensions enabled in firmware/hardware.
  3. Possible fix: Mark KVM_Infra test as expected-to-skip on qcs6490-rb3gen2 platform in the LAVA test suite configuration. This platform does not support KVM and the test should not be executed. Alternatively, update the test runner to check for /dev/kvm presence before marking as FAIL and report SKIP instead when the device node is absent due to platform limitations.
  4. Detail analysis attachment: failed_case_job223509_6_detailed.md

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants