Skip to content

FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended - #1088

Merged
Salendarsingh Gaud (sgaud-quic) merged 2 commits into
qualcomm-linux:qcom-6.18.yfrom
bibekpatro:priv_calls_runtime_handlers
Sep 12, 2026
Merged

Salendarsingh Gaud (sgaud-quic) merged 2 commits into
qualcomm-linux:qcom-6.18.yfrom
bibekpatro:priv_calls_runtime_handlers

Conversation

@bibekpatro

Copy link
Copy Markdown

qcom_adreno_smmu_get_fault_info() accesses SMMU registers without holding a runtime PM reference. A fault is raised while the SMMU is active, but the GPU may drop its power vote before the threaded fault handler reaches the callback, allowing the SMMU to runtime suspend.

Accessing the SMMU registers after suspend has started is unsafe and may cause subsequent register accesses during runtime resume to fail with a NoC error and an asynchronous SError.

Use pm_runtime_get_if_active() to keep the SMMU active while collecting the fault information, and skip the register reads if suspend has already started.

Link: https://lore.kernel.org/all/20260912-priv_call_runtime_handlers-v1-1-fc0c3a17523f@oss.qualcomm.com/

CRs-Fixed:4673688

qcom_adreno_smmu_get_fault_info() accesses SMMU registers without
holding a runtime PM reference. A fault is raised while the SMMU is
active, but the GPU may drop its power vote before the threaded fault
handler reaches the callback, allowing the SMMU to runtime suspend.

Accessing the SMMU registers after suspend has started is unsafe and
may cause subsequent register accesses during runtime resume to fail
with a NoC error and an asynchronous SError.

Use pm_runtime_get_if_active() to keep the SMMU active while collecting
the fault information, and skip the register reads if suspend has
already started.

Link: https://lore.kernel.org/all/20260912-priv_call_runtime_handlers-v1-1-fc0c3a17523f@oss.qualcomm.com/
Signed-off-by: Bibek Kumar Patro <bibek.patro@oss.qualcomm.com>
@qswat-orbit-external

Copy link
Copy Markdown

Merge Check Failed: No Change Task Found

No associated change tasks found for CR 4673688 on any of the following entities:

Entities:

  • kernel.qli.2.0

CR: 4673688

Please ensure the CR has a change task associated with at least one of the entities for this branch.

1 similar comment
@qswat-orbit-external

Copy link
Copy Markdown

Merge Check Failed: No Change Task Found

No associated change tasks found for CR 4673688 on any of the following entities:

Entities:

  • kernel.qli.2.0

CR: 4673688

Please ensure the CR has a change task associated with at least one of the entities for this branch.

@qcomlnxci

Copy link
Copy Markdown

Test Matrix

Test Case hamoa-iot-evk-multimedia lemans-evk-multimedia monaco-evk-multimedia purwa-iot-evk-multimedia qcs615-ride-multimedia qcs6490-rb3gen2-multimedia qcs8300-ride-multimedia qcs9100-ride-r3-multimedia shikra-iqs-evk-multimedia
Audio_Card_Registration ✅ Pass ◻️ ✅ Pass ◻️ ⚠️ skip ✅ Pass ⚠️ skip ⚠️ skip ⚠️ skip
BT_FW_KMD_Service ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
BT_ON_OFF ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
BT_SCAN ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ❌ Fail
CPUFreq_Validation ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
CPU_affinity ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
DSP_AudioPD ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
Ethernet_Basic_Validation ⚠️ skip ◻️ ✅ Pass ◻️ ⚠️ skip ⚠️ skip ❌ Fail ❌ Fail ⚠️ skip
Freq_Scaling ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
GIC ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ❌ Fail
IPA ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
Interrupts ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
KVM_Driver ❌ Fail ◻️ ✅ Pass ◻️ ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail
KVM_EL2_DTB ❌ Fail ◻️ ✅ Pass ◻️ ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail
KVM_Infra ❌ Fail ◻️ ✅ Pass ◻️ ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail
OpenCV ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
PCIe ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
Probe_Failure_Check ❌ Fail ◻️ ❌ Fail ◻️ ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail
RMNET ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
UFS_Validation ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
USBHost ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ❌ Fail ❌ Fail ❌ Fail ❌ Fail
WiFi_Firmware_Driver ✅ Pass ◻️ ❌ Fail ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
WiFi_OnOff ✅ Pass ◻️ ❌ Fail ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
adsp_remoteproc ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
cdsp_remoteproc ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
gpdsp_remoteproc ⚠️ skip ◻️ ✅ Pass ◻️ ⚠️ skip ⚠️ skip ✅ Pass ✅ Pass ⚠️ skip
hotplug ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
irq ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
kaslr ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
pinctrl ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
qcom_hwrng ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️
rngtest ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
shmbridge ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
smmu ❌ Fail ◻️ ✅ Pass ◻️ ❌ Fail ✅ Pass ✅ Pass ❌ Fail ✅ Pass
watchdog ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
wpss_remoteproc ✅ Pass ◻️ ✅ Pass ◻️ ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

iommu: arm-smmu-qcom: Ensure smmu is powered up in set_ttbr0_cfg

Apply right prefix, add Link: tag

@qswat-orbit-external

Copy link
Copy Markdown

Merge Check Failed: CR Not Eligible for Merge

CR 4673688 is not eligible for merge.

The parent software image for kernel.qli.2.0 is not development complete.

Entity: kernel.qli.2.0
CR: 4673688
Reason: CR_CANNOT_MERGE

Please ensure the CR passes both CCT (ComponentChangeTasks) and ICT (Integration Change Tasks) validations.

@qcomlnxci
qcomlnxci requested a review from a team September 12, 2026 05:52
…_cfg

arm_smmu_write_context_bank() assumes it is being called with RPM
active, but it turns out that is not guaranteed in the path from
qcom_adreno_smmu_set_ttbr0_cfg(), so it's possible for the register
writes to get lost when configuring the context bank while the GPU is
idle, leading to page faults later.
Add the RPM calls here to make sure the SMMU is active before we touch
it.

Link: https://lore.kernel.org/all/20260507-qcom_smmu_pmfix-v3-1-af8cd05831a2@gmail.com/
Signed-off-by: Anna Maniscalco <anna.maniscalco2000@gmail.com>
Reviewed-by: Rob Clark <rob.clark@oss.qualcomm.com>
Reviewed-by: Robin Murphy <robin.murphy@arm.com>
Tested-by: Xilin Wu <sophon@radxa.com> # sc8280xp-radxa-dragon-q8b
@qswat-orbit-external

Copy link
Copy Markdown

Merge Check Failed: CR Not Eligible for Merge

CR 4673688 is not eligible for merge.

The parent software image for kernel.qli.2.0 is not development complete.

Entity: kernel.qli.2.0
CR: 4673688
Reason: CR_CANNOT_MERGE

Please ensure the CR passes both CCT (ComponentChangeTasks) and ICT (Integration Change Tasks) validations.

@qcomlnxci

Copy link
Copy Markdown

Test Matrix

Test Case hamoa-iot-evk-multimedia lemans-evk-multimedia monaco-evk-multimedia purwa-iot-evk-multimedia qcs615-ride-multimedia qcs6490-rb3gen2-multimedia qcs8300-ride-multimedia qcs9100-ride-r3-multimedia shikra-iqs-evk-multimedia
Audio_Card_Registration ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip ✅ Pass ⚠️ skip ⚠️ skip ⚠️ skip
BT_FW_KMD_Service ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
BT_ON_OFF ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
BT_SCAN ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ❌ Fail ✅ Pass ✅ Pass ❌ Fail
CPUFreq_Validation ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
CPU_affinity ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
DSP_AudioPD ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
Ethernet_Basic_Validation ⚠️ skip ✅ Pass ✅ Pass ⚠️ skip ⚠️ skip ⚠️ skip ❌ Fail ❌ Fail ⚠️ skip
Freq_Scaling ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
GIC ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ❌ Fail
IPA ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
Interrupts ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
KVM_Driver ❌ Fail ✅ Pass ✅ Pass ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail
KVM_EL2_DTB ❌ Fail ✅ Pass ✅ Pass ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail
KVM_Infra ❌ Fail ✅ Pass ✅ Pass ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail
OpenCV ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
PCIe ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
Probe_Failure_Check ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail ❌ Fail
RMNET ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
UFS_Validation ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
USBHost ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ❌ Fail ❌ Fail ❌ Fail ❌ Fail
WiFi_Firmware_Driver ✅ Pass ✅ Pass ❌ Fail ✅ Pass ✅ Pass ✅ Pass ◻️ ✅ Pass ✅ Pass
WiFi_OnOff ✅ Pass ✅ Pass ❌ Fail ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
adsp_remoteproc ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ⚠️ skip
cdsp_remoteproc ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
gpdsp_remoteproc ⚠️ skip ✅ Pass ✅ Pass ⚠️ skip ⚠️ skip ⚠️ skip ✅ Pass ✅ Pass ⚠️ skip
hotplug ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
irq ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
kaslr ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
pinctrl ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
qcom_hwrng ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ◻️
rngtest ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
shmbridge ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
smmu ❌ Fail ❌ Fail ✅ Pass ❌ Fail ❌ Fail ✅ Pass ✅ Pass ❌ Fail ✅ Pass
watchdog ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass
wpss_remoteproc ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass ✅ Pass

@sgaud-quic

Copy link
Copy Markdown
Contributor

Merge Check Failed: CR Not Eligible for Merge

CR 4673688 is not eligible for merge.

The parent software image for kernel.qli.2.0 is not development complete.

Entity: kernel.qli.2.0 CR: 4673688 Reason: CR_CANNOT_MERGE

Please ensure the CR passes both CCT (ComponentChangeTasks) and ICT (Integration Change Tasks) validations.

Mainline PR is merged, only needs to moved to Dev-complete, merging the PR.

@sgaud-quic
Salendarsingh Gaud (sgaud-quic) merged commit b6e015a into qualcomm-linux:qcom-6.18.y Sep 12, 2026
6 of 9 checks passed
@qswat-orbit-external

Copy link
Copy Markdown

Dev Completion validation failed

CR: 4673688
Change Task: kernel.qli.2.0
Error: For CR '4673688' Change Task cannot be Development Complete until the Release Notes Status has transitioned beyond Update at least once.

The change task for this CR could not be moved to Dev Complete because of the error above. Please resolve the issue in Orbit and re-run the failed Orbit check.

@qlijarvis

Copy link
Copy Markdown

PR #1088 — validate-patch

PR: #1088

Verdict Issues Detailed Report
0 Full report

Final Summary

  1. Lore link present: Yes — both commits have correct lore.kernel.org links in Link: trailers

  2. Lore link matches PR commits: Yes — diff content is identical for both patches; line number shifts in commit 2/2 are expected due to commit 1/2 being applied first

  3. Upstream patch status:

    • Commit 1/2: ⏳ Decision Pending — posted Sep 12, 2026; under review
    • Commit 2/2: ✅ ACKed — applied to iommu maintainer tree (arm/smmu/updates)
  4. PR present in qcom-next/topics: Yes - all 2 commit(s) are present in qcom-next or topics

Verdict: ✅ — click to expand

🔍 Patch Validation

PR: #1088
Verdict: ✅ PASS


Commit 1/2: FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended

Upstream commit: https://lore.kernel.org/all/20260912-priv_call_runtime_handlers-v1-1-fc0c3a17523f@oss.qualcomm.com/

Commit Message

Check Status Note
Subject matches upstream Subject correctly adapted with FROMLIST: prefix
Body preserves rationale Commit body identical to lore patch
Fixes tag present/correct N/A No Fixes tag in upstream or PR
Authorship preserved From: matches lore author (Bibek Kumar Patro)
Backport note (if applicable) N/A FROMLIST: commit, not a backport

Diff

File Status Notes
drivers/iommu/arm/arm-smmu/arm-smmu-qcom.c Diff content identical to lore patch

Upstream Status

Decision Pending — Posted Sep 12, 2026; no maintainer decision yet. Thread shows no acceptance or rejection signals as of the fetched mbox.

Integration Presence

Present in topics — Found at d196108aa9efb3b713e8847dbc656556792423af in kernel-topics (per integration_presence_report.md)


Commit 2/2: FROMGIT: iommu: arm-smmu-qcom: Ensure smmu is powered up in set_ttbr0_cfg

Upstream commit: https://lore.kernel.org/all/20260507-qcom_smmu_pmfix-v3-1-af8cd05831a2@gmail.com/

Commit Message

Check Status Note
Subject matches upstream Subject correctly adapted with FROMGIT: prefix
Body preserves rationale Commit body identical to lore patch
Fixes tag present/correct N/A No Fixes tag in upstream or PR
Authorship preserved From: matches lore author (Anna Maniscalco)
Backport note (if applicable) N/A FROMGIT: commit, not a backport

Diff

File Status Notes
drivers/iommu/arm/arm-smmu/arm-smmu-qcom.c Diff content identical to lore patch; line number shift (231→275) due to first patch being applied

Additional Trailer

⚠️ Note: PR adds Tested-by: Xilin Wu <sophon@radxa.com> # sc8280xp-radxa-dragon-q8b which is not present in the lore v3 patch. This is acceptable — additional testing tags collected after upstream posting are legitimate additions.

Upstream Status

ACKed — Maintainer confirmed: "Applied to iommu (arm/smmu/updates), thanks!" (found in lore mbox). Patch has Reviewed-by: from Rob Clark and Robin Murphy.

Integration Presence

Present in qcom-next — All checked added lines are present in qcom-next (per integration_presence_report.md)


Issues

None. Both commits faithfully represent their upstream lore sources.


Verdict

Merge as-is — Both patches are faithful to their lore sources with correct prefixes (FROMLIST: for pending patch, FROMGIT: for accepted patch), proper authorship, preserved commit messages, and identical diff content. The additional Tested-by: tag in commit 2/2 is a legitimate enhancement.


Final Summary

  1. Lore link present: Yes — both commits have correct lore.kernel.org links in Link: trailers

  2. Lore link matches PR commits: Yes — diff content is identical for both patches; line number shifts in commit 2/2 are expected due to commit 1/2 being applied first

  3. Upstream patch status:

    • Commit 1/2: ⏳ Decision Pending — posted Sep 12, 2026; under review
    • Commit 2/2: ✅ ACKed — applied to iommu maintainer tree (arm/smmu/updates)
  4. PR present in qcom-next/topics: Yes — commit 1/2 found in topics; commit 2/2 present in qcom-next (per integration_presence_report.md)

Deterministic Integration Presence

Integration Presence Report

This report is generated by Jarvis before validate-patch runs.
It is the authoritative source for whether PR changes are already present
in qcom-next or in the kernel topic branches.

Kernel repo: /local/mnt/workspace/sgaud/Qgenie/image_pipeline/kernel
qcom-next ref: d49c33864d06e9672dce57738be8851384578fcf
topics remote: topics -> https://github.com/qualcomm-linux/kernel-topics
topics fetch: fetched

Commit Subject qcom-next topics Final
1/2 [PATCH 1/2] FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads partial - subject or partial tree evidence found, but full change was not verified present - exact patch-id match at d196108 present
2/2 [PATCH 2/2] FROMGIT: iommu: arm-smmu-qcom: Ensure smmu is powered up present - all checked added lines are present skipped - not checked because qcom-next already contains the change present

Final Status

overall_status: PASS
present_commits: 2/2
partial_commits: 0/2
missing_commits: 0/2
topics_checked_for_commits: 1/2
final_summary: PR present in qcom-next/topics: Yes - all 2 commit(s) are present in qcom-next or topics

@qlijarvis

Copy link
Copy Markdown

PR #1088 — checker-log-analyzer

PR: #1088
Checker run: https://github.com/qualcomm-linux/kernel-config/actions/runs/34654588046

Checker Result Summary
Checker Result Summary
checkpatch All commits pass style checks
dt-binding-check ⏭️ No DT binding changes
dtb-check ⏭️ No devicetree changes
sparse-check No new sparse warnings introduced
check-uapi-headers No UAPI changes
check-patch-compliance Commit 2/2 missing required prefix
tag-check Commit 2/2 missing required prefix (target: qcom-6.18.y)

Detailed report: Full report

Checker analysis — click to expand

🤖 CI Checker Analysis (checker-log-analyzer)

PR: #1088 - iommu: arm-smmu-qcom runtime PM fixes
Source: https://github.com/qualcomm-linux/kernel-config/actions/runs/34654588046

Checker Result Summary
checkpatch All commits pass style checks
dt-binding-check ⏭️ No DT binding changes
dtb-check ⏭️ No devicetree changes
sparse-check No new sparse warnings introduced
check-uapi-headers No UAPI changes
check-patch-compliance Commit 2/2 missing required prefix
tag-check Commit 2/2 missing required prefix (target: qcom-6.18.y)

❌ check-patch-compliance

Root cause: Commit 2d88b62 is missing the required subject-line prefix.

Failure details:

Checking commit: iommu: arm-smmu-qcom: Ensure smmu is powered up in set_ttbr0_cfg
Commit summary does not start with a required prefix

The second commit in the PR has subject:

iommu: arm-smmu-qcom: Ensure smmu is powered up in set_ttbr0_cfg

But it should start with one of: FROMLIST:, FROMGIT:, UPSTREAM:, BACKPORT:, QCLINUX:, PENDING:, or WORKAROUND:.

Based on the patch file metadata and the lore link in the commit body (https://lore.kernel.org/all/20260507-qcom_smmu_pmfix-v3-1-af8cd05831a2@gmail.com/), this patch was posted upstream and should use the FROMGIT: prefix (indicating it was taken from a maintainer git tree).

Fix:

git rebase -i 926f4ae52b245419be0322995f34beb1deabbd01
# Mark commit 2d88b6214b24 as 'edit'
git commit --amend -m "FROMGIT: iommu: arm-smmu-qcom: Ensure smmu is powered up in set_ttbr0_cfg

arm_smmu_write_context_bank() assumes it is being called with RPM
active, but it turns out that is not guaranteed in the path from
qcom_adreno_smmu_set_ttbr0_cfg(), so it's possible for the register
writes to get lost when configuring the context bank while the GPU is
idle, leading to page faults later.
Add the RPM calls here to make sure the SMMU is active before we touch
it.

Link: https://lore.kernel.org/all/20260507-qcom_smmu_pmfix-v3-1-af8cd05831a2@gmail.com/
Signed-off-by: Anna Maniscalco <anna.maniscalco2000@gmail.com>
Reviewed-by: Rob Clark <rob.clark@oss.qualcomm.com>
Reviewed-by: Robin Murphy <robin.murphy@arm.com>
Tested-by: Xilin Wu <sophon@radxa.com> # sc8280xp-radxa-dragon-q8b"

git rebase --continue
git push --force-with-lease

Reproduce locally:

cd kernel
git log 926f4ae52b24..2d88b6214b24 --format="%H %s" | while read sha subject; do
  if ! echo "$subject" | grep -qE '^(FROMLIST|FROMGIT|UPSTREAM|BACKPORT|QCLINUX|PENDING|WORKAROUND):'; then
    echo "MISSING prefix in commit $sha: $subject"
  fi
done

❌ tag-check

Root cause: Target branch qcom-6.18.y requires all commits to have a subject-line prefix; commit 2/2 is missing it.

Failure details:

The PR targets branch qcom-6.18.y, which is not qcom-next or qcom-next-staging. Therefore, every commit must start its subject line with a valid prefix tag.

Commit 2d88b62 has subject:

iommu: arm-smmu-qcom: Ensure smmu is powered up in set_ttbr0_cfg

This violates the mandatory prefix requirement for qcom-6.18.y.

Fix: Same as check-patch-compliance above — add FROMGIT: prefix to the commit subject.

Reproduce locally:

git log 926f4ae52b24..2d88b6214b24 --format="%H %s" | while read sha subject; do
  if ! echo "$subject" | grep -qE '^(FROMLIST|FROMGIT|UPSTREAM|BACKPORT|QCLINUX|PENDING|WORKAROUND):'; then
    echo "❌ tag-check FAIL: commit $sha missing prefix"
  fi
done

Verdict

1 blocker must be fixed before merge:

The second commit (2d88b62) is missing the required FROMGIT: prefix in its subject line. This is mandatory for the target branch qcom-6.18.y. Add the prefix and force-push the corrected commit.

All other checkers passed cleanly.

@qlijarvis

Copy link
Copy Markdown

LAVA Failed Case Triage Summary

PR: #1088

Job 225660 | SoC purwa-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225660

Failed test cases in LAVA job 225660 (SoC: purwa-evk).

  Case 1: Kernel Crash — NULL pointer dereference in DMA unmap path
  1. Failed case: Kernel Crash — NULL pointer dereference in DMA unmap path
  2. Root cause: The PR introduces runtime PM calls in arm-smmu-qcom that create a race condition: when cdsp crashes and fastrpc attempts cleanup via dma_unmap_sg_attrs, the SMMU device pointer (x0 = 0x0) is NULL because the device has been runtime-suspended or freed, causing a level-0 translation fault at address 0x00000009e6316259 when dereferencing dev->dma_ops in dma_unmap_sg_attrs+0x15c.
  3. Possible fix: Revert or fix the runtime PM handling in patches 1 and 2 of PR1088. The pm_runtime_get_if_active() in qcom_adreno_smmu_get_fault_info() and pm_runtime_resume_and_get() in qcom_adreno_smmu_set_ttbr0_cfg() must ensure the device reference remains valid for the entire lifecycle of DMA-mapped buffers. Add proper reference counting or defer cleanup until all DMA operations complete. Specifically, ensure system_heap_unmap_dma_buf() holds a device reference before calling dma_unmap_sg_attrs().
  4. Detail analysis attachment: failed_case_job225660_1_detailed.md
  Case 2: Kernel Crash — Use-After-Free in DMA unmapping (fastrpc)
  1. Failed case: Kernel Crash — Use-After-Free in DMA unmapping (fastrpc)
  2. Root cause: Kernel panic at 42.666s due to invalid pointer dereference (0x00000009e6316259) in dma_unmap_sg_attrs+0x15c/0x1b0 called from system_heap_unmap_dma_buf during fastrpc cleanup after CDSP remoteproc crash recovery. The cdsprpcd process (PID 2567) triggered a use-after-free when unmapping DMA buffers during CDSP fatal error handling ("sleep_statsi.c:537"). The crash occurred on purwa-evk (x1e80100 SoC) during remoteproc recovery flow.
  3. Possible fix: This is a pre-existing kernel bug in the fastrpc driver's DMA buffer lifecycle management during remoteproc crash recovery, not introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which modifies arm-smmu-qcom runtime PM handling). The PR changes are unrelated to fastrpc/remoteproc. To mitigate: (1) Apply upstream fastrpc fixes for DMA buffer cleanup races during remoteproc recovery; (2) Add proper synchronization between remoteproc stop/restart and fastrpc DMA unmapping; (3) Validate that fastrpc compute-cb devices are properly quiesced before DMA cleanup. Re-trigger the CI job to verify if this is a transient race or reproducible failure.
  4. Detail analysis attachment: failed_case_job225660_2_detailed.md
  Case 3: Kernel Crash — Null pointer dereference in fastrpc DMA unmap path
  1. Failed case: Kernel Crash — Null pointer dereference in fastrpc DMA unmap path
  2. Root cause: The fastrpc driver attempted to unmap a DMA buffer with a NULL or invalid device pointer in dma_unmap_sg_attrs(), causing a level 0 translation fault (null pointer dereference) at address 0x00000009e6316259 during a cdsprpcd ioctl call. The crash occurred in fastrpc_free_map()dma_buf_unmap_attachment_unlocked()system_heap_unmap_dma_buf()dma_unmap_sg_attrs(), where the device argument (x0 register = 0x0000000000000000) was NULL, leading to a page table walk failure when dereferencing device->dma_ops or related structures.
  3. Possible fix: This is a pre-existing kernel bug in the fastrpc driver, NOT introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which only modifies SMMU runtime PM handling). The PR changes are unrelated to fastrpc or DMA buffer management. To resolve: (1) Re-trigger the CI job to confirm reproducibility; (2) If reproducible, file a separate kernel bug report for the fastrpc driver maintainers with full crash log and reproduction steps; (3) The fastrpc driver needs to ensure the device pointer passed to DMA unmap operations is valid and non-NULL, likely by adding NULL checks in fastrpc_free_map() or ensuring proper device lifetime management during buffer cleanup.
  4. Detail analysis attachment: failed_case_job225660_3_detailed.md
  Case 4: Kernel Crash — NULL pointer dereference in DMA unmap path (fastrpc)
  1. Failed case: Kernel Crash — NULL pointer dereference in DMA unmap path (fastrpc)
  2. Root cause: NULL pointer dereference at dma_unmap_sg_attrs+0x15c triggered by fastrpc driver during DMA buffer unmapping; the crash occurred in cdsprpcd process (PID 2567) at 42.8 seconds into boot while the CDSP remoteproc was initializing, causing kernel panic and preventing login prompt from appearing; the fault address 0x00000009e6316259 shows a level 0 translation fault indicating a completely unmapped virtual address was dereferenced.
  3. Possible fix: This is a pre-existing kernel bug in the DMA/fastrpc subsystem interaction, not introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which only modifies SMMU runtime PM handling in arm-smmu-qcom.c); the crash occurs in core DMA mapping code (dma_unmap_sg_attrs) called from system_heap_unmap_dma_buffastrpc_free_map, suggesting a use-after-free or double-unmap scenario; recommend: (1) verify the fastrpc driver's DMA buffer lifecycle management, (2) check if recent upstream changes to DMA-buf or system heap introduced regressions, (3) re-trigger the CI job to confirm reproducibility — if the crash is intermittent, it may be a race condition in fastrpc's memory mapping cleanup path.
  4. Detail analysis attachment: failed_case_job225660_4_detailed.md
Job 225661 | SoC lemans-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225661

Failed test cases in LAVA job 225661 (SoC: lemans-evk).

  Case 1: Boot Hang — UFS initialization stall
  1. Failed case: Boot Hang — UFS initialization stall
  2. Root cause: System hung during UFS driver probe at 6.370 seconds with no further kernel output. UFS device (1d84000.ufshc) is assigned to IOMMU group 6 and uses arm-smmu-qcom. PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 adds pm_runtime_resume_and_get() call in qcom_adreno_smmu_set_ttbr0_cfg(), which is invoked during SMMU context bank configuration. On lemans-evk, this creates a runtime PM dependency deadlock: UFS probe requires SMMU context setup → SMMU context setup (via set_ttbr0_cfg) now requires SMMU device to be runtime-active → but SMMU may not be runtime-ready during early UFS probe, causing the pm_runtime_resume_and_get() call to block indefinitely or fail, stalling UFS initialization and preventing boot completion.
  3. Possible fix: Revert the second patch in PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (commit 17bce8d: "FROMGIT: iommu: arm-smmu-qcom: Ensure smmu is powered up in set_ttbr0_cfg") or modify qcom_adreno_smmu_set_ttbr0_cfg() to use pm_runtime_get_if_active() instead of pm_runtime_resume_and_get() to avoid forcing a resume during early device probe when runtime PM may not be fully initialized. The first patch (commit 41541f2) can remain as it uses get_if_active which is safe.
  4. Detail analysis attachment: failed_case_job225661_1_detailed.md
  Case 2: Complete System Hang — Boot Hang After UFS Driver Initialization
  1. Failed case: Complete System Hang — Boot Hang After UFS Driver Initialization
  2. Root cause: The kernel hung during boot at kernel timestamp 6.370s immediately after UFS driver initialization messages, with no further console output for over 9 minutes until LAVA timeout. The PR introduces runtime PM calls (pm_runtime_resume_and_get / pm_runtime_put_autosuspend) to the SMMU driver's qcom_adreno_smmu_set_ttbr0_cfg() and qcom_adreno_smmu_get_fault_info() functions. On lemans-evk, the UFS controller is behind the SMMU (added to iommu group 6 at kernel timestamp 3.557s). The hang occurs when the UFS driver attempts its first DMA transaction through the SMMU, triggering SMMU context bank configuration via qcom_adreno_smmu_set_ttbr0_cfg(), which now calls pm_runtime_resume_and_get(). If the SMMU device's runtime PM state is not correctly initialized or if there's a circular dependency (UFS needs SMMU, SMMU runtime PM may depend on storage/clocks that depend on UFS), the pm_runtime_resume_and_get() call can block indefinitely, causing a complete system hang with no forward progress.
  3. Possible fix: Revert the PR or modify the SMMU runtime PM implementation to handle early-boot scenarios where runtime PM infrastructure may not be fully initialized. Specifically, check if pm_runtime_enabled() is true before calling pm_runtime_resume_and_get() in qcom_adreno_smmu_set_ttbr0_cfg(), and fall back to direct register access if runtime PM is not yet enabled. Alternatively, ensure the SMMU device's runtime PM is initialized and enabled before any IOMMU-backed devices (like UFS) attempt their first DMA transaction. Test the fix on lemans-evk by verifying the system boots to login prompt without hanging.
  4. Detail analysis attachment: failed_case_job225661_2_detailed.md
  Case 3: minimal-boot — Login Timeout (System Hang During Boot)
  1. Failed case: minimal-boot — Login Timeout (System Hang During Boot)
  2. Root cause: System hangs completely during boot at 6.37 seconds after UFS driver initialization begins. The kernel boots successfully (Linux banner present at 0.0s, systemd-udevd starts at 5.3s), but all console output stops after the UFS regulator messages at 6.37s. LAVA login-action times out after 185 seconds waiting for a prompt that never appears. This is a complete system hang ("stopped dead") with no kernel panic, oops, or crash signature. The PR introduces runtime PM calls (pm_runtime_resume_and_get/pm_runtime_put_autosuspend) in the SMMU driver's set_ttbr0_cfg path, which may trigger a deadlock or circular dependency during early boot when UFS (a storage device requiring IOMMU/SMMU) attempts to initialize while the SMMU runtime PM state is not yet fully established.
  3. Possible fix: Revert or rework the runtime PM changes in the SMMU driver to avoid blocking during early boot. Specifically, investigate whether pm_runtime_resume_and_get() in qcom_adreno_smmu_set_ttbr0_cfg() can deadlock when called during UFS initialization (which itself may depend on SMMU being active). Consider using pm_runtime_get_if_active() instead of pm_runtime_resume_and_get() to avoid forcing a resume during early boot, or defer the runtime PM enforcement until after critical boot devices (UFS, eMMC) have initialized. Re-trigger the CI job after applying the fix to confirm boot completes and login prompt appears.
  4. Detail analysis attachment: failed_case_job225661_3_detailed.md
  Case 4: LAVA Job Timeout — Login Action Exhausted Retry Budget
  1. Failed case: LAVA Job Timeout — Login Action Exhausted Retry Budget
  2. Root cause: LAVA job timeout budget exhausted after 3 failed login-action attempts (each 200s). Kernel booted successfully (Linux version 6.18.44-g2d88b6214b24 at 0.0s, systemd-udevd started at 5.3s, UFS driver probing started at 6.3s), but system stopped producing console output during UFS initialization and never reached login prompt. LAVA dispatcher retried login 3 times over ~10 minutes before exhausting the block timeout budget. Error: "No time left for remaining 2 retries. You should either increase block timeout or decrease named action timeout."
  3. Possible fix: This is a LAVA job configuration issue, not a kernel regression. The PR changes (SMMU runtime PM fixes in arm-smmu-qcom.c) are unrelated to the boot hang. Recommended actions: (1) Re-trigger the LAVA job to rule out transient infrastructure issues; (2) If the hang reproduces, increase the LAVA job's block timeout from the current value to at least 15 minutes to allow more retry attempts; (3) Investigate why the lemans-evk board hung during UFS initialization — check for UFS driver issues, device tree configuration problems, or board-specific hardware/firmware issues on this specific lemans-evk instance.
  4. Detail analysis attachment: failed_case_job225661_4_detailed.md
Job 225662 | SoC qcs9100-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225662

Failed test cases in LAVA job 225662 (SoC: qcs9100-ride).

  Case 1: Probe_Failure_Check — Pre-existing Platform Configuration Issues
  1. Failed case: Probe_Failure_Check — Pre-existing Platform Configuration Issues
  2. Root cause: The test detected three categories of probe/firmware issues on qcs9100-ride, all pre-existing and unrelated to the PR: (1) Four PMIC temp-alarm devices stuck in deferred probe due to missing thermal zone or IIO ADC dependencies in DT; (2) Aquantia AQR115C Ethernet PHY probe failure (error -22) due to missing or malformed firmware-name DT property; (3) Benign regulatory.db firmware load failure (cfg80211 falls back to built-in rules). The PR only modifies SMMU runtime PM handling and does not touch PMIC, SPMI, thermal, PHY, or regulatory subsystems.
  3. Possible fix: These are not regressions introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088. The test failure reflects pre-existing board/DT configuration gaps on qcs9100-ride. To resolve: (1) Add thermal zone bindings for the four PMIC temp-alarm devices in qcs9100-ride DT; (2) Add firmware-name property to the Aquantia PHY node in qcs9100-ride DT (e.g., firmware-name = "Rhe-05.06-Candidate9-AQR_Mediatek_23B_P5_ID45824_LCLVER1.cld";); (3) Optionally add regulatory.db to rootfs firmware directory. For CI purposes, suppress these known board-specific failures in the Probe_Failure_Check test for qcs9100-ride, or mark them as expected/informational rather than hard failures.
  4. Detail analysis attachment: failed_case_job225662_1_detailed.md
  Case 2: smmu
  1. Failed case: smmu
  2. Root cause: Video codec device aa00000.video-codec exists in device tree but driver is not probing, preventing IOMMU group attachment; this is a pre-existing driver/platform issue unrelated to the PR's runtime PM changes in arm-smmu-qcom fault handling paths.
  3. Possible fix: Investigate why the video codec driver (qcom-iris or qcom-venus) is not probing on qcs9100-ride platform; check for missing dependencies (clocks, regulators, firmware, power domains), DT binding mismatches, or driver Kconfig/module loading issues; the PR changes to SMMU runtime PM handling do not affect device probe or IOMMU attachment flow.
  4. Detail analysis attachment: failed_case_job225662_2_detailed.md
  Case 3: USBHost (Test Infrastructure Issue — No CoT Applies)
  1. Failed case: USBHost (Test Infrastructure Issue — No CoT Applies)
  2. Root cause: Test infrastructure failure — no physical USB devices connected to the qcs9100-ride board's USB host ports; USB host controllers initialized successfully (3 buses enumerated: Bus 001/002/003 with root hubs), but test expects at least one non-hub USB device to be physically plugged in.
  3. Possible fix: Connect a physical USB device (e.g., USB flash drive, keyboard, or mouse) to one of the qcs9100-ride board's USB host ports before running the USBHost test; alternatively, update the test to SKIP (not FAIL) when no USB devices are present, as this is an expected hardware configuration state, not a kernel regression.
  4. Detail analysis attachment: failed_case_job225662_3_detailed.md
  Case 4: Ethernet_Basic_Validation
  1. Failed case: Ethernet_Basic_Validation
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Add the firmware-name property to the Aquantia AQR115C PHY device tree node under the MDIO bus for qcs9100-ride. The property should specify the PHY firmware file path (e.g., firmware-name = "Rhe-05.06-Candidate9-AQR_Mediatek_23B_P5_ID45824_LPDDR4X_VER_0.cld";). If the firmware file is not required for this PHY configuration, update the PHY driver to make the firmware-name property optional rather than mandatory.
  4. Detail analysis attachment: failed_case_job225662_4_detailed.md
  Case 5: KVM_Driver — Platform Configuration Limitation
  1. Failed case: KVM_Driver — Platform Configuration Limitation
  2. Root cause: KVM cannot initialize on qcs9100-ride because the platform boots with Gunyah hypervisor occupying EL2 (HYP mode). KVM requires exclusive EL2 access to create /dev/kvm, but Gunyah is already running at EL2, causing KVM to detect "HYP mode not available" and skip device node creation.
  3. Possible fix: This is not a kernel bug or PR regression. Either: (1) Exclude KVM tests from qcs9100-ride CI runs (platform does not support nested virtualization), or (2) Configure the platform to boot without Gunyah if KVM testing is required (mutually exclusive: Gunyah OR KVM, not both).
  4. Detail analysis attachment: failed_case_job225662_5_detailed.md
  Case 6: ** KVM Driver Initialization Failure — HYP mode not available
  1. Failed case: ** KVM Driver Initialization Failure — HYP mode not available
  2. Root cause: ** KVM initialization fails because EL2/HYP mode is not available to the Linux kernel on the qcs9100-ride platform; the Qualcomm hypervisor (hypvm.mbn) runs at EL2 and does not expose nested virtualization support to the guest Linux kernel running at EL1.
  3. Possible fix: This is a platform/firmware limitation, not a kernel regression. To enable KVM on this platform: (1) configure the hypervisor to expose nested virtualization (VHE/nVHE) to the guest kernel, or (2) boot Linux directly at EL2 without the Qualcomm hypervisor. For CI purposes, mark KVM tests as expected-to-skip on qcs9100-ride or exclude this platform from KVM test runs.
  4. Detail analysis attachment: failed_case_job225662_6_detailed.md
  Case 7: KVM Infrastructure Test Failure — HYP Mode Not Available
  1. Failed case: KVM Infrastructure Test Failure — HYP Mode Not Available
  2. Root cause: KVM cannot initialize because EL2 (HYP mode) is occupied by the Gunyah hypervisor on qcs9100-ride platform. The bootloader log shows "Gunyah based bootup" and kernel reports "kvm [1]: HYP mode not available" at boot time. KVM requires exclusive access to EL2 to function, which is incompatible with Gunyah hypervisor presence.
  3. Possible fix: This is a platform firmware configuration issue, not a kernel regression. The qcs9100-ride target is configured to run Gunyah hypervisor at EL2, which is mutually exclusive with KVM. To enable KVM on this platform: (1) reconfigure the bootloader/firmware to boot without Gunyah hypervisor, OR (2) exclude KVM tests from the test suite for Gunyah-enabled targets, OR (3) use nested virtualization if supported by the platform (requires Gunyah to expose EL2 to guest).
  4. Detail analysis attachment: failed_case_job225662_7_detailed.md
  Case 8: 0_qcom-next-ci-premerge-tests
  1. Failed case: 0_qcom-next-ci-premerge-tests
  2. Root cause: KVM test infrastructure failure due to platform limitation — qcs9100-ride (LeMans) does not support HYP mode (EL2 virtualization extensions are not available on this SoC/board configuration).
  3. Possible fix: Mark KVM tests as "not applicable" for qcs9100-ride in the LAVA test suite configuration, or skip KVM tests when /dev/kvm is not present and kernel reports "HYP mode not available" — this is a known platform limitation, not a kernel regression.
  4. Detail analysis attachment: failed_case_job225662_8_detailed.md
Job 225663 | SoC shikra-iqs-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225663

Failed test cases in LAVA job 225663 (SoC: shikra-iqs-evk).

  Case 1: GIC Test Infrastructure Bug — Test Script Assumes 8 CPUs
  1. Failed case: GIC Test Infrastructure Bug — Test Script Assumes 8 CPUs
  2. Root cause: The GIC test script at /lava-225663/0/tests/0_qcom-next-ci-premerge-tests/Runner/suites/Kernel/Baseport/GIC/run.sh line 75 attempts to validate timer interrupt increments for CPUs 0-7, but the shikra-iqs-evk platform only has 4 CPUs (0-3). When the script parses /proc/interrupts for non-existent CPUs 4-7, it encounters malformed data (text fields "GICv3", "Level", "arch_timer" instead of integer counts), triggering bash integer comparison errors and false test failures.
  3. Possible fix: Update the GIC test script to dynamically detect the number of online CPUs from /sys/devices/system/cpu/online or /proc/cpuinfo and only validate timer interrupts for CPUs that actually exist on the target platform. The script should not hardcode an assumption of 8 CPUs.
  4. Detail analysis attachment: failed_case_job225663_1_detailed.md
  Case 2: ** Probe_Failure_Check — Pre-existing Platform Probe Failures (Not PR-Introduced)
  1. Failed case: ** Probe_Failure_Check — Pre-existing Platform Probe Failures (Not PR-Introduced)
  2. Root cause: ** The test detected 7 probe failures and 3 deferred probe warnings on shikra-iqs-evk that are pre-existing platform-specific issues unrelated to PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088's SMMU runtime PM changes. Failures include: CoreSight ETM devices failing with -EINVAL (platform debug configuration issue), regulatory.db firmware missing (expected on minimal rootfs), cpufreq-dt duplicate registration (-EEXIST, benign), lt9611c display bridge I2C error (hardware/bus issue), and audio codec clock dependencies (transient deferrals).
  3. Possible fix: Refine the Probe_Failure_Check test to exclude known benign probe failures for shikra platform: (1) Add suppression rule for CoreSight ETM -EINVAL failures (known platform limitation), (2) Suppress regulatory.db -ENOENT (optional firmware), (3) Suppress cpufreq-dt -EEXIST (benign duplicate), (4) Classify deferred probes separately from hard failures. For PR validation, the test should PASS when no NEW probe failures are introduced relative to baseline.
  4. Detail analysis attachment: failed_case_job225663_2_detailed.md
  Case 3: USBHost
  1. Failed case: USBHost
  2. Root cause: Test infrastructure issue — no external USB devices physically connected to the shikra-iqs-evk board's USB host port during test execution. The kernel USB subsystem initialized correctly (usbcore, hub, usbhid drivers loaded), USB controller (4e00000.usb) was added to IOMMU group 5, but the test script found zero enumerated USB devices when querying the system.
  3. Possible fix: This is not a kernel regression introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which modifies arm-smmu-qcom runtime PM handling). The test requires a USB device (keyboard, mouse, or storage) to be physically connected to the board's USB host port before test execution. Either: (1) ensure the LAVA lab setup for shikra-iqs-evk includes a USB device connected to the host port, or (2) mark this test as SKIP for boards without permanent USB host peripherals attached in the lab environment.
  4. Detail analysis attachment: failed_case_job225663_3_detailed.md
  Case 4: ** BT_SCAN — Test Environment Issue (No Discoverable Devices)
  1. Failed case: ** BT_SCAN — Test Environment Issue (No Discoverable Devices)
  2. Root cause: ** Bluetooth hardware and driver are fully functional (BT_ON_OFF passed, hci0 powered and ready), but the BT_SCAN test failed because no nearby Bluetooth devices were available to discover in the LAVA lab environment during the 3 scan attempts (15 seconds each).
  3. Possible fix: This is not a kernel bug. To resolve: (1) Place a discoverable Bluetooth device (phone, speaker, beacon) within range of the shikra-iqs-evk board in the LAVA lab, or (2) Modify the BT_SCAN test to be optional/informational when no devices are found, or (3) Suppress BT_SCAN failures when BT_ON_OFF passes (similar to existing firmware suppression rules).
  4. Detail analysis attachment: failed_case_job225663_4_detailed.md
  Case 5: Kernel Crash — Synchronous External Abort in qcom_rng driver
  1. Failed case: Kernel Crash — Synchronous External Abort in qcom_rng driver
  2. Root cause: The qcom_rng driver accessed hardware MMIO registers (at offset +0xc4 in qcom_rng_read) while the RNG hardware block was powered down or clock-gated, triggering a synchronous external abort (bus error 0x96000010) on Shikra IQS EVK during the qcom_hwrng test.
  3. Possible fix: Add runtime PM protection to qcom_rng_read() using pm_runtime_get_if_active() before register access and pm_runtime_put_autosuspend() after, following the same pattern as the SMMU fix in this PR (patch 1/2). If the device is suspended, skip the read and return an appropriate error rather than crashing the kernel.
  4. Detail analysis attachment: failed_case_job225663_5_detailed.md
  Case 6: Kernel Crash — Synchronous External Abort (Hardware Access Fault)
  1. Failed case: Kernel Crash — Synchronous External Abort (Hardware Access Fault)
  2. Root cause: The qcom_rng driver accessed MMIO registers while the hardware was not powered or clocked, triggering a synchronous external abort at PC qcom_rng_read+0xc4. The crash occurred during the qcom_hwrng test when reading from /dev/hwrng, causing a kernel panic that prevented subsequent tests (including KVM_EL2_DTB) from running. The PR introduces runtime PM changes to SMMU but does not modify qcom_rng; this appears to be a pre-existing platform issue on shikra-iqs-evk where the RNG hardware is not properly powered/clocked at runtime.
  3. Possible fix: Add runtime PM support to the qcom_rng driver (drivers/char/hw_random/qcom-rng.c) to ensure the RNG hardware is powered and clocked before register access. Specifically, wrap register reads in qcom_rng_read() with pm_runtime_get_sync() / pm_runtime_put_autosuspend() calls, similar to the pattern used in the PR's SMMU fixes. As a short-term workaround, disable the qcom_hwrng test on shikra-iqs-evk until the driver fix is merged.
  4. Detail analysis attachment: failed_case_job225663_6_detailed.md
  Case 7: Kernel Crash — Synchronous External Abort in qcom_rng driver
  1. Failed case: Kernel Crash — Synchronous External Abort in qcom_rng driver
  2. Root cause: The qcom_rng driver accessed TRNG hardware registers without runtime PM protection while the hardware block was powered down, causing a bus timeout and synchronous external abort (error code 0x96000010) at qcom_rng_read+0xc4. The crash occurred during the qcom_hwrng test (not during KVM_Infra test itself). The KVM_Infra test failure (/dev/kvm not available) is benign and unrelated to the crash.
  3. Possible fix: Add runtime PM protection to qcom_rng_read() following the pattern in the PR under test: call pm_runtime_resume_and_get() before register access and pm_runtime_put_autosuspend() after. This is the same fix pattern used in patch 2/2 of PR#1088 for qcom_adreno_smmu_set_ttbr0_cfg(). Verify on Shikra IQS EVK that TRNG register access succeeds even when device was runtime suspended.
  4. Detail analysis attachment: failed_case_job225663_7_detailed.md
  Case 8: 0_qcom-next-ci-premerge-tests
  1. Failed case: 0_qcom-next-ci-premerge-tests
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Add runtime PM calls (pm_runtime_get_sync / pm_runtime_put) around hardware register access in the qcom_rng driver (drivers/crypto/qcom-rng.c) to ensure the RNG hardware is powered and clocked before register reads. Alternatively, verify that the RNG clock/power domain is correctly described in the shikra-iqs-evk device tree and that the driver probe sequence correctly enables them.
  4. Detail analysis attachment: failed_case_job225663_8_detailed.md
  Case 9: Kernel Crash — synchronous external abort in qcom_rng driver
  1. Failed case: Kernel Crash — synchronous external abort in qcom_rng driver
  2. Root cause: Hardware register access fault in qcom_rng_read() at offset +0xc4 while reading from the PRNG hardware block during the qcom_hwrng test. The crash triggered a kernel panic followed by EFI pstore write failures and system reset into EDL/ramdump mode, causing LAVA test-shell to timeout after 2400 seconds waiting for the board to recover.
  3. Possible fix: This is a pre-existing kernel bug in the qcom_rng driver unrelated to the PR changes (which only touch arm-smmu-qcom runtime PM handling). The crash indicates the PRNG hardware block became inaccessible, likely due to a power management or clock gating issue. Recommended actions: (1) Check if qcom_rng driver has proper runtime PM handling and clock management; (2) Verify PRNG hardware block power domain and clock dependencies in device tree; (3) Add error handling for hardware access failures in qcom_rng_read(); (4) Re-trigger the CI job to confirm if this is a transient hardware issue or reproducible regression.
  4. Detail analysis attachment: failed_case_job225663_9_detailed.md
  Case 10: **Kernel Crash — Synchronous External Abort in qcom_rng driver**
  1. Failed case: Kernel Crash — Synchronous External Abort in qcom_rng driver
  2. Root cause: The qcom_rng driver accessed SMMU-protected hardware registers while the SMMU was runtime-suspended, triggering a synchronous external abort (bus fault) at PC qcom_rng_read+0xc4/0x228 when reading from the RNG hardware during the qcom_hwrng test execution on shikra-iqs-evk.
  3. Possible fix: Apply runtime PM protection to qcom_rng driver register accesses: wrap hardware register reads in pm_runtime_resume_and_get() / pm_runtime_put_autosuspend() pairs in qcom_rng_read() to ensure the device (and its SMMU context) remains active during register access, following the same pattern introduced by this PR for arm-smmu-qcom fault handlers.
  4. Detail analysis attachment: failed_case_job225663_10_detailed.md
  Case 11: ** Kernel Crash — Synchronous External Abort in qcom_rng driver
  1. Failed case: ** Kernel Crash — Synchronous External Abort in qcom_rng driver
  2. Root cause: ** Hardware bus error (synchronous external abort 0x96000010) when qcom_rng driver attempted to read RNG hardware registers during the qcom_hwrng test. The crash occurred at qcom_rng_read+0xc4 when accessing MMIO registers, indicating the RNG hardware block was either not powered, not clocked, or inaccessible due to a runtime PM state mismatch on the shikra-iqs-evk platform.
  3. Possible fix: This is a pre-existing platform/driver issue unrelated to PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which modifies SMMU runtime PM handling). The qcom_rng driver needs runtime PM protection around register accesses similar to the SMMU fix in the PR. Short-term: disable the qcom_hwrng test on shikra-iqs-evk until the driver is fixed. Long-term: add pm_runtime_get_if_active() guards in qcom_rng_read() to prevent register access when the device is suspended, following the same pattern as the SMMU fix in this PR.
  4. Detail analysis attachment: failed_case_job225663_11_detailed.md
Job 225664 | SoC monaco-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225664

Failed test cases in LAVA job 225664 (SoC: monaco-evk).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: WiFi PCIe device (ath11k_pci WCN6855) probe failed with -ETIMEDOUT after firmware load failure; missing firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin in rootfs caused MHI power-up timeout during device initialization on monaco-evk.
  3. Possible fix: Add the missing WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin to the rootfs firmware directory (/lib/firmware/); verify firmware packaging in the build recipe includes the nfa765 variant for WCN6855 hw2.1.
  4. Detail analysis attachment: failed_case_job225664_1_detailed.md
  Case 2: Driver Probe Failure — WiFi ath11k_pci probe failed with error -110 (ETIMEDOUT)
  1. Failed case: Driver Probe Failure — WiFi ath11k_pci probe failed with error -110 (ETIMEDOUT)
  2. Root cause: WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin is missing from the rootfs firmware directory. The MHI bus controller attempted to load the firmware during ath11k_pci driver probe but received -ENOENT (-2), causing the MHI power-up sequence to timeout (-ETIMEDOUT, -110). This is a pre-existing infrastructure/build issue on the monaco-evk target, not introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which only modifies SMMU runtime PM handling in arm-smmu-qcom.c and does not touch WiFi, MHI, or firmware paths).
  3. Possible fix: Add the missing WCN6855 WiFi firmware files to the rootfs build for monaco-evk. Specifically, ensure ath11k/WCN6855/hw2.1/nfa765/amss.bin (and related board/regdb files) are included in the firmware package. This is a build/integration issue — the kernel driver is functioning correctly but cannot proceed without the required firmware blob. Re-trigger the CI job after updating the rootfs firmware manifest for monaco-evk.
  4. Detail analysis attachment: failed_case_job225664_2_detailed.md
  Case 3: WiFi_OnOff
  1. Failed case: WiFi_OnOff
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Add the missing WCN6855 firmware to the Yocto build by ensuring linux-firmware-ath11k (or the vendor-specific firmware package for Monaco/nfa765 variant) is included in IMAGE_INSTALL in the Yocto image recipe. Verify the firmware file exists at /lib/firmware/ath11k/WCN6855/hw2.1/nfa765/amss.bin in the built rootfs before flashing.
  4. Detail analysis attachment: failed_case_job225664_3_detailed.md
  Case 4: 0_qcom-next-ci-premerge-tests
  1. Failed case: 0_qcom-next-ci-premerge-tests
  2. Root cause: LAVA test runner marked the test definition as "unfinished" and failed it because WiFi tests (WiFi_Firmware_Driver and WiFi_OnOff) failed due to ath11k_pci probe failure (-110 ETIMEDOUT). The ath11k probe failed because the WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin was not found (-2 ENOENT), causing MHI power-up timeout. This is a pre-existing firmware packaging/availability issue on monaco-evk, not introduced by the PR (which modifies SMMU runtime PM handling).
  3. Possible fix: This is not a genuine PR-introduced regression. The WiFi firmware file is missing from the rootfs image used in this test run. To resolve: (1) verify the firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin is included in the rootfs build for monaco-evk, (2) re-trigger the CI job with a properly configured image. The PR changes (SMMU runtime PM fixes) are unrelated to this firmware availability issue and should not block merge.
  4. Detail analysis attachment: failed_case_job225664_4_detailed.md
Job 225665 | SoC qcs8300-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225665

Failed test cases in LAVA job 225665 (SoC: qcs8300-ride).

  Case 1: Probe_Failure_Check — WiFi regulatory.db firmware false positive (suppressed)
  1. Failed case: Probe_Failure_Check — WiFi regulatory.db firmware false positive (suppressed)
  2. Root cause: The Probe_Failure_Check test flagged a regulatory.db firmware load failure (error -2, ENOENT) during WiFi driver initialization. However, this is a known benign failure: the regulatory.db file is optional, WiFi uses compiled-in regulatory certificates as fallback, and both WiFi_Firmware_Driver and WiFi_OnOff functional tests passed, confirming WiFi is fully operational.
  3. Possible fix: No fix required. This failure should be suppressed per LAVA Known Benign Failure Rule 2 (WiFi firmware false positive when WiFi ON/OFF passes). Update the Probe_Failure_Check test to exclude regulatory.db firmware load failures from its error pattern matching, or add an exception when WiFi functional tests pass.
  4. Detail analysis attachment: failed_case_job225665_1_detailed.md
  Case 2: ** USBHost — Hardware Configuration Issue (No USB Devices Connected)
  1. Failed case: ** USBHost — Hardware Configuration Issue (No USB Devices Connected)
  2. Root cause: ** The qcs8300-ride board's USB host controller initialized successfully but reports "USB3 root hub has no ports" (only USB2 with 1 port available). The USBHost test detected only the root hub (Bus 001 Device 001: Linux Foundation 2.0 root hub) with no functional USB devices connected. This is a test infrastructure issue, not a kernel regression — the board requires a physical USB device to be connected to the USB port for the test to pass.
  3. Possible fix: Connect a functional USB device (e.g., USB flash drive, USB keyboard, or USB hub with devices) to the qcs8300-ride board's USB port before running the USBHost test. If the board is in a LAVA lab, verify the USB port is accessible and not blocked by the test fixture, and update the lab setup documentation to ensure a USB device is connected for this test case.
  4. Detail analysis attachment: failed_case_job225665_2_detailed.md
  Case 3: ** Ethernet_Basic_Validation — SGMII PHY SERDES Powerup Timeout
  1. Failed case: ** Ethernet_Basic_Validation — SGMII PHY SERDES Powerup Timeout
  2. Root cause: ** The qcom-dwmac-sgmii-phy driver fails to initialize the SERDES hardware during eth0 interface bring-up on qcs8300-ride, timing out while polling QSERDES_COM_C_READY_STATUS register. The failure is triggered by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088's changes to SMMU runtime PM handling, which alter the power state of the SMMU (IOMMU group 8) that the Ethernet device depends on, exposing a latent runtime PM dependency issue in the SGMII PHY driver.
  3. Possible fix: Revert PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 to confirm the regression, then investigate the SGMII PHY driver (drivers/net/ethernet/stmicro/stmmac/dwmac-qcom-ethqos.c and drivers/phy/qualcomm/phy-qcom-sgmii-eth.c) to add proper runtime PM handling for SERDES powerup operations, ensuring the SMMU and all dependent power domains are active before accessing QSERDES registers. Alternatively, ensure the SMMU runtime PM changes in the PR do not affect non-GPU devices by scoping the runtime PM calls more narrowly to Adreno-specific paths.
  4. Detail analysis attachment: failed_case_job225665_3_detailed.md
  Case 4: Pre-existing Platform Limitation — KVM Not Available Under Gunyah Hypervisor
  1. Failed case: Pre-existing Platform Limitation — KVM Not Available Under Gunyah Hypervisor
  2. Root cause: QCS8300 (Monaco) runs under Gunyah hypervisor (version gunyah-cdfb73831), which does not expose KVM functionality to the primary VM (HLOS). Hypervisor log shows "Failed to register KP" at boot, indicating KVM is not registered at the hypervisor level. CONFIG_KVM is enabled in the kernel, but /dev/kvm is never created because the underlying hypervisor does not provide the necessary EL2 virtualization support to the primary VM.
  3. Possible fix: This is not a bug. QCS8300 (Monaco) is designed to run under Gunyah hypervisor, which manages virtualization at EL2. KVM is not available in this configuration. To enable KVM testing, either: (1) use a platform that boots Linux directly at EL2 without a hypervisor, or (2) configure the Gunyah hypervisor to expose nested virtualization support (if supported by the platform). For CI purposes, mark KVM tests as "expected skip" on QCS8300/Monaco targets.
  4. Detail analysis attachment: failed_case_job225665_4_detailed.md
  Case 5: KVM Driver Initialization Failure — /dev/kvm device node not created
  1. Failed case: KVM Driver Initialization Failure — /dev/kvm device node not created
  2. Root cause: KVM driver (CONFIG_KVM=y) failed to initialize during early boot on qcs8300-ride platform; no KVM initialization messages in kernel log indicate silent initialization failure, likely due to missing EL2/virtualization hardware support or bootloader not enabling hypervisor mode for this SoC.
  3. Possible fix: Verify qcs8300 (Monaco) SoC supports ARM virtualization extensions (EL2/VHE) and that the bootloader/firmware enables EL2 mode. If hardware supports virtualization, add debug prints to arch/arm64/kvm/arm.c:kvm_arch_init() to identify why initialization is failing. If qcs8300 does not support virtualization, mark KVM tests as SKIP for this platform in CI configuration.
  4. Detail analysis attachment: failed_case_job225665_5_detailed.md
  Case 6: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM driver failed to initialize during kernel boot on qcs8300-ride (Monaco) — CONFIG_KVM is enabled but /dev/kvm device node was never created, indicating the KVM ARM driver probe failed silently without logging any error messages.
  3. Possible fix: Investigate why KVM ARM driver (arch/arm64/kvm/arm.c) failed to probe on qcs8300-ride: check if the platform boots at EL2 (required for KVM), verify CPU feature support for virtualization extensions (VHE/nVHE), and add debug logging to kvm_arch_init() and kvm_init() to identify the silent failure point; the PR changes to SMMU runtime PM are unrelated to this KVM initialization failure.
  4. Detail analysis attachment: failed_case_job225665_6_detailed.md
  Case 7: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM device node /dev/kvm is not present on qcs8300-ride platform despite CONFIG_KVM=y in kernel config — this is a pre-existing platform/kernel configuration issue unrelated to the PR's SMMU runtime PM changes.
  3. Possible fix: Investigate why KVM module is not creating /dev/kvm on qcs8300-ride: check if KVM ARM64 support is properly configured for this SoC, verify hypervisor mode (EL2) is available, check dmesg for KVM initialization errors, and ensure required KVM ARM dependencies are met.
  4. Detail analysis attachment: failed_case_job225665_7_detailed.md
Job 225666 | SoC hamoa-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225666

Failed test cases in LAVA job 225666 (SoC: hamoa-evk).

  Case 1: ** Probe_Failure_Check
  1. Failed case: ** Probe_Failure_Check
  2. Root cause: ** Three pre-existing platform configuration issues on hamoa-evk: (1) qcom_qseecom_uefisecapp driver fails probe with -EBUSY due to TrustZone secure app unavailability or resource conflict; (2) qcom-spmi-lpg driver fails probe with -EINVAL due to invalid device tree "reg" property for multi-LED configuration; (3) cfg80211 regulatory.db firmware file missing from rootfs firmware paths. None are related to the SMMU runtime PM changes in PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088.
  3. Possible fix: These failures are not PR-introduced regressions. To resolve: (1) verify TZ firmware includes UEFI secure app support for hamoa; (2) correct the SPMI LPG device tree "reg" property for multi-LED nodes in hamoa-evk DTS; (3) add wireless-regdb package to rootfs or include regulatory.db in /lib/firmware. The PR changes are safe to merge — they add runtime PM protection to SMMU register access and do not affect these unrelated drivers.
  4. Detail analysis attachment: failed_case_job225666_1_detailed.md
  Case 2: smmu (test validation failure — missing IOMMU group attachments for critical masters)
  1. Failed case: smmu (test validation failure — missing IOMMU group attachments for critical masters)
  2. Root cause: Six devices (five USB PHY controllers at addresses a0f8800.usb, a2f8800.usb, a4f8800.usb, a6f8800.usb, a8f8800.usb, and video codec aa00000.video-codec) are missing IOMMU group attachments on hamoa-evk. This is a pre-existing device tree configuration issue, not a regression introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088's runtime PM changes.
  3. Possible fix: Update the device tree for hamoa-evk to add iommus properties for the missing devices (if they require IOMMU protection), or update the test expectations to exclude these devices from the critical master check for this platform (if they do not perform DMA or do not require IOMMU protection). This issue is orthogonal to the PR and should be tracked separately.
  4. Detail analysis attachment: failed_case_job225666_2_detailed.md
  Case 3: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: If KVM support is required on Hamoa EVK, update the bootloader (ABL/UEFI) and firmware (TrustZone) to enable EL2 for the kernel, then verify /dev/kvm is created after reboot. If KVM is not supported on this IoT platform, update the LAVA test definition to skip KVM tests on Hamoa EVK by adding a platform check: if grep -q "Hamoa IoT EVK" /proc/device-tree/model 2>/dev/null; then echo "[SKIP] KVM not supported"; exit 0; fi.
  4. Detail analysis attachment: failed_case_job225666_3_detailed.md
  Case 4: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Update the LAVA job definition to skip (not fail) KVM tests on platforms without EL2 support. Add a pre-test check: if dmesg | grep -q "HYP mode not available", mark KVM tests as SKIP instead of FAIL. Alternatively, maintain a platform capability matrix in CI configuration and conditionally enable KVM tests only for EL2-capable platforms (e.g., rb3gen2, kodiak, not hamoa-evk).
  4. Detail analysis attachment: failed_case_job225666_4_detailed.md
  Case 5: KVM_Infra — Test Infrastructure Limitation (Not a Kernel Bug)
  1. Failed case: KVM_Infra — Test Infrastructure Limitation (Not a Kernel Bug)
  2. Root cause: KVM driver initialization failed because the Hamoa IoT EVK platform does not support ARM EL2 (Hypervisor mode). The kernel message kvm [1]: HYP mode not available indicates the CPU is not running in or does not support EL2, which is a hardware/firmware prerequisite for KVM. CONFIG_KVM is enabled in the kernel config, but the /dev/kvm device node is never created because the KVM driver exits early when it detects EL2 is unavailable.
  3. Possible fix: This is not a kernel regression introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which only touches SMMU runtime PM). The KVM test suite should be disabled or marked as expected-fail for the Hamoa IoT EVK platform in the LAVA job definition, as this platform does not support virtualization. If EL2 support is expected on this platform, verify the bootloader/firmware configuration and ensure the CPU is booted into EL2 mode (check UEFI/ABL settings and secure boot configuration).
  4. Detail analysis attachment: failed_case_job225666_5_detailed.md
  Case 6: 0_qcom-next-ci-premerge-tests
  1. Failed case: 0_qcom-next-ci-premerge-tests
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Re-trigger the CI job on hamoa-evk to verify if the lockups are reproducible. If lockups persist across multiple runs, investigate hamoa-evk platform-specific cpuidle or power management configuration. The SMMU PM patches in this PR are defensive and unlikely to be the root cause. Consider testing on a different hamoa-evk board or updating firmware/bootloader if the issue is board-specific.
  4. Detail analysis attachment: failed_case_job225666_6_detailed.md
Job 225667 | SoC qcs615-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225667

Failed test cases in LAVA job 225667 (SoC: qcs615-ride).

  Case 1: Probe_Failure_Check — cfg80211 regulatory.db firmware load failure (known benign)
  1. Failed case: Probe_Failure_Check — cfg80211 regulatory.db firmware load failure (known benign)
  2. Root cause: The cfg80211 wireless regulatory subsystem attempted to load the optional regulatory.db firmware file during boot, which is not present in the rootfs. Error -2 (ENOENT) indicates file not found. This is a known benign failure because: (1) regulatory.db is optional for WiFi operation, (2) WiFi functional tests (WiFi_OnOff, WiFi_Firmware_Driver) passed, confirming WiFi driver and firmware loaded correctly, and (3) the PR changes (SMMU runtime PM handling in arm-smmu-qcom.c) are unrelated to wireless regulatory database loading.
  3. Possible fix: Suppress this failure in the Probe_Failure_Check test by adding regulatory.db to the known-benign firmware load failure list, or include the regulatory.db file in the rootfs if strict compliance is required. This is not a PR-introduced regression — it is a pre-existing test environment configuration issue.
  4. Detail analysis attachment: failed_case_job225667_1_detailed.md
  Case 2: ** smmu (Test Infrastructure False Positive)
  1. Failed case: ** smmu (Test Infrastructure False Positive)
  2. Root cause: ** The LAVA smmu test incorrectly expects child devices aa00000.video-codec:video-decoder and aa00000.video-codec:video-encoder to appear as separate platform devices with IOMMU group attachments. These are V4L2 video device nodes created by the Venus video codec driver at runtime, not platform devices enumerated during boot. The parent device aa00000.video-codec is correctly attached to IOMMU group 7, and the kernel shows no SMMU errors.
  3. Possible fix: Update the LAVA smmu test script to exclude V4L2 video device nodes (video-decoder, video-encoder) from the critical master IOMMU attachment check, or verify IOMMU protection at the parent device level (aa00000.video-codec) only. The PR changes (SMMU runtime PM handling for Adreno GPU) are unrelated to this test failure and do not introduce any video codec regression.
  4. Detail analysis attachment: failed_case_job225667_2_detailed.md
  Case 3: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Update the LAVA test suite configuration for qcs615-ride to skip KVM tests when Gunyah hypervisor is present, or reconfigure the platform firmware to boot without Gunyah if KVM testing is required. The test expectation should be: if hypervisor detected at boot, mark KVM tests as SKIP (not FAIL).
  4. Detail analysis attachment: failed_case_job225667_3_detailed.md
  Case 4: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: KVM module initialization failed during boot because HYP (EL2) mode is not available on the qcs615-ride platform — the kernel detected at boot time that the CPU is not running in EL2 or that the hypervisor stub is not present, causing KVM to abort initialization and preventing /dev/kvm device node creation.
  3. Possible fix: This is a platform/firmware limitation, not a kernel regression. The qcs615-ride board does not support KVM virtualization (EL2/HYP mode is not enabled in the boot chain or is disabled by the secure firmware). Mark KVM tests as "skip" for this platform in the LAVA job definition, or enable EL2 support in the bootloader/firmware if the hardware supports it.
  4. Detail analysis attachment: failed_case_job225667_4_detailed.md
  Case 5: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: This is expected behavior for qcs615-ride with Gunyah hypervisor enabled. To enable KVM testing: (1) disable Gunyah hypervisor in the boot configuration, or (2) exclude KVM tests from the qcs615-ride test suite, or (3) use a different platform without a pre-loaded hypervisor for KVM validation. The PR patches (SMMU runtime PM fixes) are unrelated to this failure.
  4. Detail analysis attachment: failed_case_job225667_5_detailed.md
  Case 6: KVM_Infra — Platform Limitation (HYP mode not available)
  1. Failed case: KVM_Infra — Platform Limitation (HYP mode not available)
  2. Root cause: QCS615 platform does not support ARM Hypervisor (EL2/HYP) mode, as evidenced by kernel message "kvm [1]: HYP mode not available" at boot. CONFIG_KVM is enabled in the kernel config, but the hardware/firmware does not provide EL2 support, preventing /dev/kvm device creation.
  3. Possible fix: This is not a kernel bug or PR-introduced regression. The test should be skipped on platforms without EL2 support. Add platform capability detection to the LAVA test definition to skip KVM tests on QCS615 and other platforms that boot at EL1 without hypervisor support.
  4. Detail analysis attachment: failed_case_job225667_6_detailed.md
Job 225668 | SoC qcs6490-rb3gen2

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225668

Failed test cases in LAVA job 225668 (SoC: qcs6490-rb3gen2).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Two pre-existing boot-time firmware loading failures detected by the test: (1) cfg80211 regulatory database (regulatory.db) missing from rootfs at boot time (error -2 = -ENOENT), and (2) Renesas USB xHCI controller firmware (renesas_usb_fw.mem) missing from rootfs, causing xhci-pci-renesas driver probe failure. Both failures occurred during kernel boot (7.08s and 9.83s timestamps) and are unrelated to the PR changes (which only modify arm-smmu-qcom.c IOMMU driver runtime PM handling).
  3. Possible fix: Install missing firmware files in the rootfs: (1) Add linux-firmware package or manually copy regulatory.db to /lib/firmware/regulatory.db for cfg80211 wireless regulatory support, and (2) Add renesas_usb_fw.mem to /lib/firmware/ for Renesas USB controller support. These are rootfs/image packaging issues, not kernel regressions introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088. The Probe_Failure_Check test is correctly detecting pre-existing firmware gaps in the test image.
  4. Detail analysis attachment: failed_case_job225668_1_detailed.md
  Case 2: USBHost
  1. Failed case: USBHost
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Add the Renesas USB firmware package (typically linux-firmware or a vendor-specific firmware package containing renesas_usb_fw.mem) to the rootfs build recipe for qcs6490-rb3gen2. Verify the firmware file is present at /lib/firmware/renesas_usb_fw.mem in the deployed image. Re-run the LAVA job to confirm the xhci-pci-renesas driver probes successfully and USB devices are enumerated.
  4. Detail analysis attachment: failed_case_job225668_2_detailed.md
  Case 3: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: KVM is not available on qcs6490-rb3gen2 because the platform boots under a Gunyah hypervisor at EL2, preventing KVM from initializing. The kernel message "kvm [1]: HYP mode not available" confirms that KVM cannot run when another hypervisor already controls EL2. CONFIG_KVM is enabled in the kernel config, but /dev/kvm device node is not created because the KVM driver initialization fails early due to lack of HYP mode access.
  3. Possible fix: This is not a PR-introduced regression. The SMMU runtime PM fixes in PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 do not affect KVM availability. The KVM_Driver test should be skipped on platforms running Gunyah hypervisor, or the test should be updated to detect and report this as an expected configuration rather than a failure. To enable KVM on this platform, the system would need to boot without the Gunyah hypervisor, which is a platform/firmware configuration decision outside the scope of kernel changes.
  4. Detail analysis attachment: failed_case_job225668_3_detailed.md
  Case 4: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: KVM is unavailable on qcs6490-rb3gen2 (Kodiak) because the platform does not support EL2/HYP mode — kernel message at boot: "kvm [1]: HYP mode not available". CONFIG_KVM is enabled but /dev/kvm device node is never created because KVM initialization fails when HYP mode is not available.
  3. Possible fix: This is a pre-existing platform limitation, not a PR-introduced regression. The PR patches (SMMU runtime PM fixes) are unrelated to KVM. If KVM support is required on this platform, verify that: (1) the bootloader/firmware enables EL2 mode before handing control to the kernel, (2) the SoC/board supports virtualization extensions, and (3) secure firmware allows non-secure EL2 access. If KVM is not expected to work on this platform, mark these KVM test cases as expected failures for qcs6490-rb3gen2 in the CI test suite configuration.
  4. Detail analysis attachment: failed_case_job225668_4_detailed.md
  Case 5: KVM Infrastructure Test — Platform Limitation (EL1 boot, no hypervisor mode)
  1. Failed case: KVM Infrastructure Test — Platform Limitation (EL1 boot, no hypervisor mode)
  2. Root cause: The qcs6490-rb3gen2 platform boots the kernel at EL1 (Exception Level 1) instead of EL2 (hypervisor mode), as evidenced by the boot message "CPU: All CPU(s) started at EL1" and "kvm [1]: HYP mode not available". Without EL2 support, the KVM driver cannot initialize and /dev/kvm is never created, causing all KVM-dependent tests (KVM_Driver, KVM_EL2_DTB, KVM_Infra) to fail. This is a platform/firmware configuration limitation, not a kernel bug.
  3. Possible fix: This is not a PR-introduced regression. The PR patches (IOMMU/SMMU runtime PM fixes) are unrelated to KVM. To enable KVM on this platform: (1) verify the bootloader/firmware supports booting Linux at EL2 (requires PSCI firmware and bootloader configuration changes), (2) if EL2 boot is not supported by the SoC/firmware, mark KVM tests as "skip" or "not applicable" for qcs6490-rb3gen2 in the CI test matrix, or (3) run KVM tests only on platforms with confirmed EL2 boot support (e.g., platforms with hypervisor-enabled firmware).
  4. Detail analysis attachment: failed_case_job225668_5_detailed.md
  Case 6: KVM_Infra (also KVM_Driver, KVM_EL2_DTB — same root cause)
  1. Failed case: KVM_Infra (also KVM_Driver, KVM_EL2_DTB — same root cause)
  2. Root cause: KVM cannot initialize because the platform is not running in EL2 (hypervisor mode); kernel prints "kvm [1]: HYP mode not available" at boot, preventing /dev/kvm device creation.
  3. Possible fix: This is a pre-existing platform/firmware configuration issue unrelated to the PR (which only touches IOMMU/SMMU runtime PM). To enable KVM: (1) configure the bootloader/firmware to boot the kernel in EL2 mode, or (2) if the platform does not support EL2, mark these KVM tests as expected-fail for qcs6490-rb3gen2 in the LAVA test suite configuration.
  4. Detail analysis attachment: failed_case_job225668_6_detailed.md

@qlijarvis

Copy link
Copy Markdown

LAVA Failed Case Triage Summary

PR: #1088

Job 225912 | SoC lemans-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225912

Failed test cases in LAVA job 225912 (SoC: lemans-evk).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Four PMIC temp-alarm devices (c440000.spmi:pmic@{0,2,4,6}:temp-alarm@a00) remain in deferred probe state, and three firmware load failures (regulatory.db, qca/wcnhpbtfw21.tlv, qca/hpbtfw21.tlv) are detected. The Bluetooth firmware failures are benign (BT_ON_OFF test passed), but the deferred probe devices indicate a missing dependency preventing the qcom-spmi-temp-alarm driver from completing probe on lemans-evk.
  3. Possible fix: Investigate why the qcom-spmi-temp-alarm driver is deferring probe on lemans-evk. Check if the thermal zone framework dependency or IIO ADC channel dependency is missing in the device tree or kernel config. The Bluetooth firmware errors are false positives (suppressed per lava-known-benign-failures.md Rule 3) and the regulatory.db error is a known benign cfg80211 warning. Focus triage on the PMIC temp-alarm deferred probe issue.
  4. Detail analysis attachment: failed_case_job225912_1_detailed.md
  Case 2: smmu
  1. Failed case: smmu
  2. Root cause: Video codec device aa00000.video-codec is not attached to any IOMMU group on lemans-evk; the smmu test expects all critical hardware masters (GPU, USB, Display, Video) to be IOMMU-protected, but the video codec device is missing IOMMU group attachment despite the qcom_iris video driver module being loaded.
  3. Possible fix: This is a pre-existing platform/DT configuration issue unrelated to the PR (which only modifies SMMU runtime PM handling for GPU). Investigate why aa00000.video-codec platform device is not binding to an IOMMU group: check lemans-evk device tree for correct iommus property on the video-codec node, verify the video codec driver probe sequence, and confirm the SMMU stream-ID mapping is configured correctly for the video hardware block.
  4. Detail analysis attachment: failed_case_job225912_2_detailed.md
  Case 3: 0_qcom-next-ci-premerge-tests
  1. Failed case: 0_qcom-next-ci-premerge-tests
  2. Root cause: LAVA test suite marked as failed due to two individual test failures: (1) smmu test failed because Video codec device (aa00000.video-codec) is missing IOMMU group attachment, and (2) Probe_Failure_Check detected probe failures for temp-alarm devices and firmware load failures (benign). The PR introduces runtime PM handling in arm-smmu-qcom driver which may affect IOMMU group attachment timing for video codec on lemans-evk.
  3. Possible fix: Investigate why video codec device aa00000.video-codec is not being attached to an IOMMU group on lemans-evk after applying the PR patches that add runtime PM calls to arm-smmu-qcom driver. Check if the pm_runtime_resume_and_get() call in qcom_adreno_smmu_set_ttbr0_cfg() or the pm_runtime_get_if_active() call in qcom_adreno_smmu_get_fault_info() is preventing proper IOMMU group attachment during video codec probe. Verify device tree configuration for video codec IOMMU bindings on lemans platform.
  4. Detail analysis attachment: failed_case_job225912_3_detailed.md
Job 225913 | SoC qcs8300-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225913

Failed test cases in LAVA job 225913 (SoC: qcs8300-ride).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Test flagged two pre-existing platform configuration issues unrelated to PR changes: (1) missing regulatory.db firmware file (benign — cfg80211 regulatory subsystem operates correctly with built-in rules); (2) Aquantia AQR115C Ethernet PHY probe failure due to missing firmware-name DT property on qcs8300-ride board (error -22/EINVAL). Neither failure is introduced by the PR, which only modifies SMMU runtime PM handling in arm-smmu-qcom.c.
  3. Possible fix: Mark test as false positive for this PR. To resolve the underlying platform issues: (1) regulatory.db is optional and can be ignored; (2) add firmware-name = "..."; property to the Aquantia PHY node (stmmac-0:08) in arch/arm64/boot/dts/qcom/qcs8300-ride.dts if PHY firmware loading is required for this board, or mark the property as optional in the driver if the PHY can operate without firmware.
  4. Detail analysis attachment: failed_case_job225913_1_detailed.md
  Case 2: ** USBHost (Test Infrastructure Issue - No External USB Device Connected)
  1. Failed case: ** USBHost (Test Infrastructure Issue - No External USB Device Connected)
  2. Root cause: ** The USBHost test expects external USB devices (keyboard, mouse, flash drive, etc.) to be physically connected to the qcs8300-ride board's USB host port for enumeration verification. The kernel USB stack is functioning correctly (xHCI controller initialized, root hub detected with 1 port), but the LAVA lab hardware setup for this board has no external USB devices connected. Test output shows only Bus 001 Device 001: ID 1d6b:0002 Linux Foundation 2.0 root hub, triggering the test's failure condition: "Only USB hubs detected, no functional USB devices."
  3. Possible fix: This is not a kernel bug and requires no code fix. To resolve: (1) Connect a USB device (e.g., USB flash drive, keyboard) to the qcs8300-ride board's USB host port in the LAVA lab, OR (2) Mark this test as "skip" or "expected-fail" for qcs8300-ride in the LAVA job definition if USB host peripheral testing is not supported/required for this platform's CI validation, OR (3) Update the test to distinguish between "USB controller non-functional" (kernel bug) vs "no devices connected" (infrastructure limitation) and pass with a note when only the root hub is present.
  4. Detail analysis attachment: failed_case_job225913_2_detailed.md
  Case 3: ** Ethernet_Basic_Validation — PHY Driver Probe Failure
  1. Failed case: ** Ethernet_Basic_Validation — PHY Driver Probe Failure
  2. Root cause: ** The Aquantia AQR115C PHY driver probe fails at boot with error -22 (EINVAL) because the device tree is missing the required firmware-name property for the PHY node at stmmac-0:08. Without a functional PHY, the qcom-ethqos Ethernet driver cannot attach to the PHY during interface bring-up, causing phylink validation to fail with -EINVAL and preventing the eth0 interface from becoming operational.
  3. Possible fix: Add the missing firmware-name property to the Aquantia AQR115C PHY device tree node in the qcs8300-ride DTS file. The property should specify the correct firmware file path for the AQR115C PHY (typically "Rxx_AQR_Firmware.cld" or similar, depending on the PHY revision). This is a platform/board DT fix, not a kernel code change. The PR under test does not introduce this issue — it is a pre-existing platform configuration problem.
  4. Detail analysis attachment: failed_case_job225913_3_detailed.md
  Case 4: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: /dev/kvm device node is not present because the kernel is running as a guest under the Gunyah hypervisor at EL1, not at EL2 (hypervisor mode). KVM requires EL2 access to create the /dev/kvm device node, which is unavailable when Linux runs as a guest VM. This is a platform configuration issue, not a kernel bug.
  3. Possible fix: This is not a failure that can be "fixed" in the kernel. The test expectation is incorrect for this platform configuration. Either: (1) skip KVM tests when running under a hypervisor (detect via "Hypervisor cold boot" in boot log or check for paravirtualization features), or (2) run the test on bare-metal qcs8300-ride hardware where Linux boots at EL2, or (3) enable nested virtualization support in Gunyah if available.
  4. Detail analysis attachment: failed_case_job225913_4_detailed.md
  Case 5: KVM_EL2_DTB — KVM unavailable (architectural limitation)
  1. Failed case: KVM_EL2_DTB — KVM unavailable (architectural limitation)
  2. Root cause: The qcs8300-ride platform is running under the Gunyah hypervisor (detected at boot: "Hypervisor cold boot, version: gunyah-cdfb73831 perf"), which occupies EL2 (ARM hypervisor privilege level). KVM requires direct EL2 access to create the /dev/kvm device node and cannot function as a nested hypervisor under Gunyah. CONFIG_KVM is enabled, but the KVM driver does not initialize because EL2 is unavailable.
  3. Possible fix: This is not a bug or regression. KVM tests should be skipped on qcs8300-ride (and any Gunyah-based platform) via LAVA job definition or test suite configuration, as KVM cannot run under another hypervisor. If KVM functionality is required, the platform must boot without the Gunyah hypervisor (bare-metal or with KVM as the sole hypervisor).
  4. Detail analysis attachment: failed_case_job225913_5_detailed.md
  Case 6: ** KVM Infrastructure Test Failure — Missing /dev/kvm Device Node
  1. Failed case: ** KVM Infrastructure Test Failure — Missing /dev/kvm Device Node
  2. Root cause: ** QCS8300 (Monaco) platform does not support ARM virtualization extensions (EL2/KVM). CONFIG_KVM is enabled in the kernel configuration, but the KVM subsystem silently skips initialization when it detects the hardware does not provide virtualization support. The /dev/kvm device node is never created because the KVM driver never probes.
  3. Possible fix: Mark KVM tests as "not applicable" or "skip" for qcs8300-ride platform in the LAVA test suite configuration. Add platform capability detection to the test runner to automatically skip KVM tests on platforms without virtualization support. Alternatively, disable CONFIG_KVM in the kernel configuration for qcs8300 builds if virtualization is not a supported feature for this SoC.
  4. Detail analysis attachment: failed_case_job225913_6_detailed.md
  Case 7: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM kernel config (CONFIG_KVM) is enabled but /dev/kvm device node is not created, indicating KVM driver failed to initialize or the QCS8300 (Monaco) platform does not support ARM virtualization extensions (EL2/VHE) required for KVM operation.
  3. Possible fix: Verify QCS8300 hardware supports ARM virtualization extensions; check kernel boot log for KVM initialization messages or errors (none found in current log); if platform lacks VHE/EL2 support, disable CONFIG_KVM in kernel config for this SoC; if hardware supports virtualization, investigate why KVM driver probe is failing silently and add debug logging to arch/arm64/kvm/arm.c kvm_arch_init().
  4. Detail analysis attachment: failed_case_job225913_7_detailed.md
Job 225914 | SoC monaco-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225914

Failed test cases in LAVA job 225914 (SoC: monaco-evk).

  Case 1: Driver Probe Failure — ath11k_pci WiFi driver
  1. Failed case: Driver Probe Failure — ath11k_pci WiFi driver
  2. Root cause: WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin is missing from the test rootfs, causing MHI firmware load to fail with -ENOENT, which cascades to ath11k_pci probe timeout (-ETIMEDOUT). This is a pre-existing infrastructure issue unrelated to the PR's SMMU runtime PM changes.
  3. Possible fix: Add the missing WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin to the rootfs firmware directory (/lib/firmware/ath11k/WCN6855/hw2.1/nfa765/). Verify the firmware package (linux-firmware or vendor-specific) is installed in the Yocto/build recipe for monaco-evk.
  4. Detail analysis attachment: failed_case_job225914_1_detailed.md
  Case 2: WiFi_Firmware_Driver
  1. Failed case: WiFi_Firmware_Driver
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: Add the missing firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin to the rootfs firmware directory (/lib/firmware/ath11k/WCN6855/hw2.1/nfa765/). Verify the firmware package for WCN6855 hw2.1 is included in the Yocto build recipe or manually install the linux-firmware package containing this specific firmware variant for the Monaco EVK platform.
  4. Detail analysis attachment: failed_case_job225914_2_detailed.md
  Case 3: ** WiFi Driver Probe Failure — Missing Firmware
  1. Failed case: ** WiFi Driver Probe Failure — Missing Firmware
  2. Root cause: ** ath11k_pci driver probe failed with -110 (ETIMEDOUT) because the required WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin is missing from the rootfs. MHI firmware loader returned -2 (ENOENT), preventing the driver from powering up the WCN6855 hw2.1 WiFi chip on Monaco EVK.
  3. Possible fix: Add the missing WiFi firmware file ath11k/WCN6855/hw2.1/nfa765/amss.bin to the rootfs build. Verify the firmware package (linux-firmware-ath11k or equivalent) is included in the Yocto image recipe for monaco-evk, or manually install the firmware to /lib/firmware/ath11k/WCN6855/hw2.1/nfa765/ on the target.
  4. Detail analysis attachment: failed_case_job225914_3_detailed.md
  Case 4: WiFi Driver Probe Failure
  1. Failed case: WiFi Driver Probe Failure
  2. Root cause: The PR introduces runtime PM synchronization in qcom_adreno_smmu_set_ttbr0_cfg() that blocks during WiFi driver probe. The pm_runtime_resume_and_get() call added in patch 2 creates a timing issue where the SMMU device runtime PM state conflicts with the WiFi device probe sequence, causing the ath11k_pci driver to timeout (-110 ETIMEDOUT) waiting for IOMMU domain configuration to complete. This is a PR-introduced regression specific to the monaco-evk platform's WiFi PCIe device initialization flow.
  3. Possible fix: Revert patch 2 (FROMGIT: iommu: arm-smmu-qcom: Ensure smmu is powered up in set_ttbr0_cfg) or modify the runtime PM acquisition in qcom_adreno_smmu_set_ttbr0_cfg() to use pm_runtime_get_if_in_use() instead of pm_runtime_resume_and_get() to avoid forcing a synchronous resume during device probe. Alternatively, investigate whether the SMMU device's runtime PM initialization order needs adjustment relative to PCIe device enumeration on monaco-evk.
  4. Detail analysis attachment: failed_case_job225914_4_detailed.md
Job 225915 | SoC qcs615-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225915

Failed test cases in LAVA job 225915 (SoC: qcs615-ride).

  Case 1: Probe_Failure_Check — False Positive (Test Infrastructure Issue)
  1. Failed case: Probe_Failure_Check — False Positive (Test Infrastructure Issue)
  2. Root cause: The Probe_Failure_Check test flagged a benign cfg80211 regulatory.db firmware load failure (error -2 / -ENOENT) that does not impact WiFi functionality. WiFi loaded successfully and passed all functional tests (WiFi_Firmware_Driver: pass, WiFi_OnOff: pass). The regulatory.db file is an optional user-space regulatory database; cfg80211 falls back to compiled-in regulatory data when the file is absent. This error is unrelated to the PR changes (PR only modifies SMMU runtime PM in arm-smmu-qcom.c) and represents a pre-existing condition in the test environment.
  3. Possible fix: Suppress this specific regulatory.db firmware load failure in the Probe_Failure_Check test by adding it to the test's benign failure exclusion list, OR provide the regulatory.db firmware file in the rootfs at /lib/firmware/regulatory.db. The failure does not require kernel changes and should not block PR merge.
  4. Detail analysis attachment: failed_case_job225915_1_detailed.md
  Case 2: smmu (test validation failure — not a kernel crash or SMMU fault)
  1. Failed case: smmu (test validation failure — not a kernel crash or SMMU fault)
  2. Root cause: The LAVA smmu test expects video-decoder and video-encoder to exist as separate platform devices with individual IOMMU group attachments. However, the Venus video codec driver on qcs615-ride uses "non legacy binding" mode (confirmed by kernel log: qcom-venus aa00000.video-codec: non legacy binding), where video-decoder and video-encoder are internal components of the parent device aa00000.video-codec, not separate platform devices. The parent device IS correctly attached to IOMMU group 7. This is a test expectation mismatch for the qcs615 platform architecture, not a kernel regression.
  3. Possible fix: Update the LAVA smmu test script to recognize that on platforms using Venus "non legacy binding" mode (qcs615, and potentially other newer SoCs), video-decoder and video-encoder are not separate platform devices. The test should pass if the parent video-codec device is attached to an IOMMU group. Alternatively, add platform-specific test expectations that skip checking for video-decoder/video-encoder sub-devices on qcs615-ride and similar platforms that use non-legacy Venus binding.
  4. Detail analysis attachment: failed_case_job225915_2_detailed.md
  Case 3: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: KVM module initialization failed because the platform is not running in EL2 (hypervisor mode). The kernel message kvm [1]: HYP mode not available at boot indicates the CPU is running in EL1 (kernel mode) without hypervisor support, preventing KVM from creating the /dev/kvm device node required for virtualization.
  3. Possible fix: This is a platform/firmware limitation, not a kernel regression introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088. The PR changes only affect SMMU runtime PM handling in arm-smmu-qcom.c and do not impact KVM or EL2 mode availability. To enable KVM on qcs615-ride, the bootloader/firmware must be configured to boot Linux in EL2 mode (e.g., via ABL/XBL configuration or hypervisor stub). If KVM support is not required for this platform, mark the KVM test suite as expected-fail or skip for qcs615-ride in the LAVA job definition.
  4. Detail analysis attachment: failed_case_job225915_3_detailed.md
  Case 4: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: This is not a bug to fix but a known platform configuration limitation. The KVM_EL2_DTB test should be skipped on qcs615-ride when Gunyah hypervisor is enabled. Add a platform-specific test exclusion rule in the LAVA job definition or test suite configuration to skip all KVM tests on platforms running Gunyah. Alternatively, if KVM testing is required, boot the platform without the Gunyah hypervisor (native Linux mode), though this may disable other virtualization features that depend on Gunyah.
  4. Detail analysis attachment: failed_case_job225915_4_detailed.md
  Case 5: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: This is not a kernel bug. To enable KVM on this platform: (1) verify the QCS615 SoC supports virtualization extensions (check SoC datasheet), (2) if supported, enable EL2 in the bootloader/firmware configuration (ABL/XBL), (3) if the SoC does not support EL2, disable KVM tests for this platform in the CI test matrix as they will always fail.
  4. Detail analysis attachment: failed_case_job225915_5_detailed.md
  Case 6: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM infrastructure test failed because the qcs615-ride platform does not support EL2 (hypervisor mode) — kernel reports "HYP mode not available" at boot, preventing /dev/kvm device creation despite CONFIG_KVM being enabled.
  3. Possible fix: This is not a PR-introduced regression (PR changes SMMU runtime PM, unrelated to KVM). The qcs615-ride board lacks hardware virtualization support (EL2). Either: (1) skip KVM tests on qcs615-ride in CI configuration, or (2) run KVM tests only on platforms with EL2 support (e.g., sc8280xp, sm8450, sm8550).
  4. Detail analysis attachment: failed_case_job225915_6_detailed.md
Job 225916 | SoC qcs9100-ride

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225916

Failed test cases in LAVA job 225916 (SoC: qcs9100-ride).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Four PMIC temp-alarm devices remain in deferred probe state (c440000.spmi:pmic@{0,2,4,6}:temp-alarm@a00), and two genuine probe failures occurred: (1) regulatory.db firmware file missing (-ENOENT), and (2) Aquantia AQR115C Ethernet PHY probe failed with -EINVAL due to missing/invalid firmware-name DT property. These are pre-existing platform/configuration issues unrelated to the PR's SMMU runtime PM changes.
  3. Possible fix: (1) For temp-alarm deferred probes: investigate qcom-spmi-temp-alarm driver dependencies on qcs9100-ride — likely missing thermal zone bindings or PMIC thermal-sensor provider not probing; add required DT thermal zone references. (2) For regulatory.db: install wireless-regdb package in rootfs or add regulatory.db to /lib/firmware/. (3) For Aquantia PHY: add valid "firmware-name" property to stmmac-0:08 PHY node in qcs9100-ride DT, or remove the property if firmware is not required for this PHY configuration.
  4. Detail analysis attachment: failed_case_job225916_1_detailed.md
  Case 2: smmu
  1. Failed case: smmu
  2. Root cause: Video codec device aa00000.video-codec is not present in platform device enumeration on qcs9100-ride target; test expects this device to be attached to an IOMMU group but the device was never created during boot, indicating a device tree or platform configuration issue unrelated to the PR changes (which only modify SMMU runtime PM handling in fault paths).
  3. Possible fix: This is a pre-existing platform configuration issue, not a PR-introduced regression. The PR should not be blocked by this failure. Recommended actions: (1) Mark this test failure as a known platform issue for qcs9100-ride; (2) Investigate why aa00000.video-codec device is not being created from device tree on this target; (3) Verify the device tree node exists and has correct compatible string and status property; (4) Consider updating the smmu test to handle platforms where video codec is implemented via iris sub-devices only.
  4. Detail analysis attachment: failed_case_job225916_2_detailed.md
  Case 3: USBHost
  1. Failed case: USBHost
  2. Root cause: Test infrastructure issue — no physical USB devices connected to the qcs9100-ride board during test execution. USB host controllers initialized successfully (3 USB root hubs detected on buses 001, 002, 003), but the test expects at least one functional USB device beyond the root hubs to validate USB host functionality.
  3. Possible fix: Connect a functional USB device (e.g., USB flash drive, keyboard, or mouse) to one of the USB ports on the qcs9100-ride board before running the USBHost test. This is a lab/hardware setup requirement, not a kernel issue. The PR changes (IOMMU/SMMU runtime PM fixes) are unrelated to USB functionality and did not cause this failure.
  4. Detail analysis attachment: failed_case_job225916_3_detailed.md
  Case 4: Ethernet_Basic_Validation
  1. Failed case: Ethernet_Basic_Validation
  2. Root cause: Aquantia AQR115C PHY driver probe failure on stmmac-0:08 (MDIO address 8 on 23040000.ethernet) due to missing or invalid firmware-name device tree property (error -EINVAL), preventing the stmmac Ethernet driver from attaching to the PHY when bringing up interface end0.
  3. Possible fix: Add the required firmware-name property to the Aquantia PHY device tree node under the MDIO bus for 23040000.ethernet (end0), specifying the correct firmware file path for the AQR115C PHY (typically "Rhe-05.06-Candidate9-AQR_Mediatek_23B_P5_ID45824_LNXDRIVER.cld" or platform-specific variant), or verify that the firmware file exists in /lib/firmware/ if the property is already present.
  4. Detail analysis attachment: failed_case_job225916_4_detailed.md
  Case 5: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: KVM cannot initialize on qcs9100-ride because the platform runs under Gunyah hypervisor (Type-1 hypervisor at EL2), which prevents KVM from accessing HYP mode. Kernel log shows "kvm [1]: HYP mode not available" at boot. This is expected behavior when a Type-1 hypervisor occupies EL2.
  3. Possible fix: Update the KVM test suite to detect Gunyah hypervisor presence (check for "Gunyah based bootup" in dmesg or Gunyah DT nodes) and skip KVM tests with reason "KVM not available under Gunyah hypervisor". Alternatively, configure a separate qcs9100-ride test target without Gunyah if KVM testing is required.
  4. Detail analysis attachment: failed_case_job225916_5_detailed.md
  Case 6: KVM_EL2_DTB (Platform Configuration Issue — KVM Unavailable Under Gunyah Hypervisor)
  1. Failed case: KVM_EL2_DTB (Platform Configuration Issue — KVM Unavailable Under Gunyah Hypervisor)
  2. Root cause: The qcs9100-ride platform is running Linux as a guest under the Gunyah hypervisor (confirmed by boot log "Hypervisor cold boot, version: gunyah-cdfb73831"), which means Linux executes at EL1 without access to EL2 (HYP mode). KVM requires EL2 access to provide virtualization services, and the kernel correctly reports "kvm [1]: HYP mode not available" at boot. CONFIG_KVM is enabled in the kernel configuration, but /dev/kvm cannot be created because the KVM subsystem initialization fails when EL2 is unavailable.
  3. Possible fix: This is not a PR-introduced regression — the PR patches only modify SMMU runtime PM handling in arm-smmu-qcom.c and do not affect KVM or hypervisor configuration. The failure is a pre-existing platform limitation: KVM tests are not applicable on qcs9100-ride when booted under Gunyah hypervisor. Recommended action: exclude KVM test cases (KVM_Driver, KVM_EL2_DTB, KVM_Infra) from the LAVA test suite for qcs9100-ride, or configure the platform to boot Linux at EL2 (bare-metal or KVM-compatible hypervisor mode) if nested virtualization testing is required.
  4. Detail analysis attachment: failed_case_job225916_6_detailed.md
  Case 7: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM initialization failed because the qcs9100-ride platform is running under the Gunyah hypervisor in a non-nested virtualization configuration, preventing KVM from accessing EL2/HYP mode. The kernel message "kvm [1]: HYP mode not available" at boot indicates KVM detected it cannot initialize because EL2 is already claimed by Gunyah.
  3. Possible fix: This is a platform configuration issue, not a kernel regression. The qcs9100-ride board is configured to boot with Gunyah hypervisor (confirmed by bootloader message "Gunyah based bootup"), which prevents KVM from running. To enable KVM on this platform, either: (1) reconfigure the board firmware to boot without Gunyah hypervisor, or (2) skip KVM tests on qcs9100-ride in the CI test suite, as this platform does not support KVM in its current configuration.
  4. Detail analysis attachment: failed_case_job225916_7_detailed.md
  Case 8: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM is unavailable on qcs9100-ride because the platform does not support HYP (EL2) mode — kernel reports "kvm [1]: HYP mode not available" at boot, preventing /dev/kvm device creation.
  3. Possible fix: This is a platform/firmware limitation, not a kernel regression. Either: (1) enable EL2/hypervisor support in the bootloader/firmware for qcs9100-ride, or (2) mark KVM tests as expected-to-skip for this platform in the LAVA test definition, or (3) exclude qcs9100-ride from KVM CI validation until hypervisor support is enabled.
  4. Detail analysis attachment: failed_case_job225916_8_detailed.md
Job 225917 | SoC shikra-iqs-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225917

Failed test cases in LAVA job 225917 (SoC: shikra-iqs-evk).

  Case 1: GIC (Test Infrastructure Bug)
  1. Failed case: GIC (Test Infrastructure Bug)
  2. Root cause: The GIC test script (run.sh line 75) has a hardcoded assumption of 8 CPUs and incorrectly parses /proc/interrupts output. On shikra-iqs-evk (4-CPU platform), the script attempts to extract interrupt counts for non-existent CPUs 4-7, instead parsing the descriptor fields (GICv3, Level, arch_timer) as integer values, causing bash integer comparison errors and false test failures.
  3. Possible fix: Update the GIC test script to dynamically detect the number of online CPUs from /sys/devices/system/cpu/online or /proc/cpuinfo before parsing /proc/interrupts, and only validate interrupt counts for CPUs that actually exist on the target platform. The script should use awk field extraction based on the actual CPU count rather than hardcoded column positions.
  4. Detail analysis attachment: failed_case_job225917_1_detailed.md
  Case 2: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Pre-existing platform driver probe failures unrelated to PR changes — coresight-etm4x (-EINVAL, missing/misconfigured ETM resources in DT), cpufreq-dt (-EEXIST, driver already registered), lt9611c (-EIO, I2C communication failure), regulatory.db (-ENOENT, missing firmware file), and deferred probe for audio codec/sound card (missing clock dependencies).
  3. Possible fix: These failures are known platform issues on shikra-iqs-evk and do not represent a regression introduced by this PR. The PR adds runtime PM protection to SMMU fault handlers and context bank configuration — it does not touch coresight, cpufreq, display bridge, or audio subsystems. Mark this test case as a false positive for this PR and investigate the underlying platform issues separately (ETM DT bindings, cpufreq driver load order, lt9611c I2C/power sequencing, regulatory.db rootfs packaging, audio clock provider probe order).
  4. Detail analysis attachment: failed_case_job225917_2_detailed.md
  Case 3: USBHost — Test Infrastructure / Configuration Issue
  1. Failed case: USBHost — Test Infrastructure / Configuration Issue
  2. Root cause: USB host controller driver (DWC3/XHCI) not loaded or not configured in kernel — USB device at 4e00000.usb registered with IOMMU but no host controller driver probe occurred, resulting in no USB bus enumeration and zero USB devices detected by test.
  3. Possible fix: Verify USB host controller driver (CONFIG_USB_DWC3, CONFIG_USB_XHCI_HCD) is enabled in kernel config and built as module or built-in; if built as module, ensure it's loaded in initramfs or via modprobe before test execution; check device tree for USB controller node status and driver binding; this is a pre-existing platform/config issue unrelated to the IOMMU runtime PM fixes in PR#1088.
  4. Detail analysis attachment: failed_case_job225917_3_detailed.md
  Case 4: BT_SCAN
  1. Failed case: BT_SCAN
  2. Root cause: Test environment failure — no discoverable Bluetooth devices present in the LAVA lab environment during the scan window. Bluetooth hardware and firmware are functional (BT_ON_OFF and BT_FW_KMD_Service passed), hci0 adapter is operational, and scan operations execute successfully, but zero devices are discovered across 3 scan attempts totaling ~5 minutes of scan time.
  3. Possible fix: This is not a kernel regression. The PR changes SMMU/IOMMU runtime PM handling and do not affect Bluetooth functionality. To resolve: (1) ensure a discoverable Bluetooth device (phone, beacon, or test device) is powered on and in range of the shikra-iqs-evk board in the LAVA lab, OR (2) mark BT_SCAN as an optional/informational test that does not block PR merge when no external devices are available, OR (3) deploy a dedicated Bluetooth beacon device in the lab for consistent scan testing.
  4. Detail analysis attachment: failed_case_job225917_4_detailed.md
  Case 5: KVM_Driver — /dev/kvm device node missing
  1. Failed case: KVM_Driver — /dev/kvm device node missing
  2. Root cause: KVM driver initialization failed during boot with "kvm [1]: HYP mode not available" — the Shikra IQS EVK platform does not support ARM Hypervisor (EL2) mode, either due to firmware/bootloader configuration not enabling EL2, or hardware/SoC limitations preventing virtualization support.
  3. Possible fix: Verify that the bootloader (ABL/UEFI) is configured to boot the kernel at EL2 (not EL1); check that the SoC/platform supports virtualization extensions; if the platform fundamentally lacks EL2 support, mark KVM tests as "not applicable" for this board in the CI test matrix rather than treating them as failures.
  4. Detail analysis attachment: failed_case_job225917_5_detailed.md
  Case 6: Kernel Crash — Synchronous External Abort in qcom_rng Hardware Access
  1. Failed case: Kernel Crash — Synchronous External Abort in qcom_rng Hardware Access
  2. Root cause: The qcom_rng driver attempts to read hardware registers without holding a runtime PM reference, causing a synchronous external abort (ESR 0x96000010) when the RNG hardware block is powered down or clock-gated. This is a pre-existing kernel bug in the qcom_rng driver, not introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which only modifies GPU SMMU runtime PM handling).
  3. Possible fix: Add runtime PM protection to qcom_rng_read() in drivers/char/hw_random/qcom-rng.c: wrap hardware register access with pm_runtime_get_sync() / pm_runtime_put() to ensure the RNG hardware is powered and clocked before access. Reference similar fix pattern from PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 patch 1/2 (qcom_adreno_smmu_get_fault_info).
  4. Detail analysis attachment: failed_case_job225917_6_detailed.md
  Case 7: KVM_Infra — /dev/kvm device node not created
  1. Failed case: KVM_Infra — /dev/kvm device node not created
  2. Root cause: KVM driver initialization failed because the system is not booted in HYP (Hypervisor/EL2) mode. Kernel message at boot: kvm [1]: HYP mode not available. On ARM64, KVM requires the CPU to run in EL2 (Hypervisor exception level), but the Shikra IQS EVK bootloader/firmware is booting the kernel in EL1 (non-secure supervisor mode), preventing KVM from initializing and creating /dev/kvm.
  3. Possible fix: This is a pre-existing platform configuration issue, not introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which only touches SMMU runtime PM). To enable KVM on Shikra IQS EVK: (1) Update the bootloader/firmware to boot the kernel in EL2 mode instead of EL1, or (2) If the platform does not support EL2 boot, mark KVM tests as expected-to-skip for this board in the CI configuration, as KVM cannot function without hypervisor mode.
  4. Detail analysis attachment: failed_case_job225917_7_detailed.md
  Case 8: Kernel Crash — Synchronous External Abort (Hardware Bus Error)
  1. Failed case: Kernel Crash — Synchronous External Abort (Hardware Bus Error)
  2. Root cause: The qcom_rng driver accessed hardware registers while the RNG hardware block was powered down or clock-gated, causing a synchronous external abort (bus error 0x96000010) at PC qcom_rng_read+0xc4. The crash occurred during a qcom_hwrng test reading 20MB from /dev/hwrng. This is a pre-existing kernel bug unrelated to PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which modifies SMMU runtime PM, not RNG).
  3. Possible fix: Add runtime PM calls (pm_runtime_get_sync/pm_runtime_put) around register accesses in drivers/char/hw_random/qcom-rng.c:qcom_rng_read() to ensure the RNG hardware is powered and clocked before accessing its MMIO registers. This is the same pattern used in the PR's SMMU fixes.
  4. Detail analysis attachment: failed_case_job225917_8_detailed.md
  Case 9: Kernel Crash — Synchronous External Abort (Hardware Bus Error)
  1. Failed case: Kernel Crash — Synchronous External Abort (Hardware Bus Error)
  2. Root cause: qcom_rng driver attempted to read hardware registers (offset 0xc4 in qcom_rng_read) while the RNG hardware block was runtime-suspended or its interconnect path was powered down, causing a synchronous external abort. The PR adds runtime PM to arm-smmu-qcom, which exposes a pre-existing bug: qcom_rng lacks runtime PM support and assumes hardware is always accessible.
  3. Possible fix: Add runtime PM support to the qcom_rng driver (drivers/char/hw_random/qcom-rng.c): implement pm_runtime_get_sync() before register access in qcom_rng_read() and pm_runtime_put_autosuspend() after, following the pattern used in other Qualcomm drivers. Alternatively, mark the RNG device as always-on in the device tree or disable runtime PM for the RNG hardware block on Shikra until driver support is added.
  4. Detail analysis attachment: failed_case_job225917_9_detailed.md
  Case 10: ** Kernel Crash — Synchronous External Abort in qcom_rng Driver
  1. Failed case: ** Kernel Crash — Synchronous External Abort in qcom_rng Driver
  2. Root cause: ** The qcom_rng driver accessed hardware registers (at offset +0xc4 in qcom_rng_read) while the RNG hardware block was not powered or clocked, triggering a synchronous external abort (bus error ESR 0x96000010). The driver lacks runtime PM protection around register accesses, allowing reads from unpowered hardware. After the crash, the board entered EDL/ramdump mode, causing the LAVA test shell to time out after 2400 seconds waiting for test completion.
  3. Possible fix: Add runtime PM protection (pm_runtime_get_sync / pm_runtime_put) around hardware register accesses in drivers/char/hw_random/qcom-rng.c:qcom_rng_read(). Ensure the RNG hardware is powered and clocked before any MMIO access. Alternatively, verify the RNG device tree node has correct power-domain and clock bindings for the shikra-iqs-evk platform.
  4. Detail analysis attachment: failed_case_job225917_10_detailed.md
  Case 11: Test Timeout — qcom_hwrng test triggered kernel crash
  1. Failed case: Test Timeout — qcom_hwrng test triggered kernel crash
  2. Root cause: Synchronous external abort (ESR 0x96000010) in qcom_rng_read() at PC+0xc4 when accessing HWRNG registers; the qcom_rng driver attempted to read hardware registers while the device was runtime-suspended or powered down, causing a NoC (Network-on-Chip) timeout and system abort unrelated to the PR's SMMU runtime PM changes.
  3. Possible fix: Add runtime PM protection to qcom_rng_read() using pm_runtime_get_if_active() before register access and pm_runtime_put_autosuspend() after, following the same pattern as the PR's fix for arm-smmu-qcom; alternatively, ensure the HWRNG device remains powered during /dev/hwrng reads by holding a runtime PM reference in the rng_dev_read() path.
  4. Detail analysis attachment: failed_case_job225917_11_detailed.md
Job 225918 | SoC purwa-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225918

Failed test cases in LAVA job 225918 (SoC: purwa-evk).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: The test detected 5 probe/firmware errors during boot on purwa-evk (iq-x5121-evk): (1) qcom_qseecom_uefisecapp failed with -EBUSY (-16) due to device already in use, (2) two qcom-pcie instances (1bf8000, 1bd0000) failed with -ENODATA (-61) because PCIe PHY init sequences are missing for this SoC, (3) qcom-spmi-lpg failed with -EINVAL (-22) due to invalid multi-LED DT configuration, and (4) regulatory.db firmware file is missing from rootfs. None of these failures are introduced by the PR (which only modifies SMMU runtime PM handling); all are pre-existing platform/configuration issues on purwa-evk.
  3. Possible fix: These are known benign failures for purwa-evk and should be suppressed in the Probe_Failure_Check test for this platform: (1) qcom_qseecom_uefisecapp -EBUSY is expected when UEFI secure app is already loaded by firmware, (2) PCIe probe failures are expected on purwa-evk until PHY init sequences are added to the driver for this SoC variant, (3) qcom-spmi-lpg DT error requires fixing the device tree multi-LED node configuration, (4) regulatory.db is optional and its absence does not affect WiFi functionality (WiFi_OnOff test passed). Update the Probe_Failure_Check test to exclude these known platform-specific probe failures for purwa-evk, or add platform-specific suppression rules.
  4. Detail analysis attachment: failed_case_job225918_1_detailed.md
  Case 2: smmu
  1. Failed case: smmu
  2. Root cause: Six critical devices (five USB controllers at a0f8800, a2f8800, a4f8800, a6f8800, a8f8800 and one video codec at aa00000) are missing IOMMU group attachments in the purwa-evk device tree, failing the SMMU validation test's requirement that all critical masters be protected by IOMMU.
  3. Possible fix: Add missing iommus properties to the device tree nodes for USB controllers a0f8800.usb, a2f8800.usb, a4f8800.usb, a6f8800.usb, a8f8800.usb and video-codec aa00000.video-codec in arch/arm64/boot/dts/qcom/x1e80100.dtsi (or the purwa-specific overlay), referencing the appropriate SMMU phandle and stream IDs for purwa-evk platform.
  4. Detail analysis attachment: failed_case_job225918_2_detailed.md
  Case 3: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: KVM initialization failed because the purwa-evk platform is running under the Gunyah hypervisor (gunyah-mobile-c487961e9), which occupies EL2 (HYP mode). KVM requires exclusive access to EL2 to function, but when a hypervisor is already running at EL2, nested virtualization is not supported, resulting in "HYP mode not available" and no /dev/kvm device node creation.
  3. Possible fix: This is not a kernel regression introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which only modifies IOMMU/SMMU runtime PM handling). The failure is expected platform behavior on purwa-evk when Gunyah hypervisor is enabled in the firmware. To enable KVM testing: (1) disable Gunyah hypervisor in the firmware/bootloader configuration for purwa-evk, or (2) exclude KVM tests from the CI test suite for platforms that run with Gunyah enabled, or (3) add a test skip condition that detects hypervisor presence before attempting KVM tests.
  4. Detail analysis attachment: failed_case_job225918_3_detailed.md
  Case 4: KVM_EL2_DTB — /dev/kvm not available (HYP mode not available)
  1. Failed case: KVM_EL2_DTB — /dev/kvm not available (HYP mode not available)
  2. Root cause: KVM driver initialization failed at boot with "kvm [1]: HYP mode not available" because the purwa-evk platform does not support EL2 (hypervisor mode) — either the hardware does not implement EL2, the bootloader/firmware did not enter the kernel at EL2, or EL2 is disabled in the secure configuration.
  3. Possible fix: This is a platform limitation, not a kernel regression. The KVM_EL2_DTB test should be skipped on purwa-evk or the test suite should check for HYP mode availability before running KVM tests. If EL2 support is expected on this platform, verify the bootloader is configured to enter the kernel at EL2 and that TrustZone/secure firmware allows EL2 usage.
  4. Detail analysis attachment: failed_case_job225918_4_detailed.md
  Case 5: ** KVM Infrastructure Test Failure — HYP Mode Unavailable
  1. Failed case: ** KVM Infrastructure Test Failure — HYP Mode Unavailable
  2. Root cause: ** The Purwa IoT EVK platform boots with the Gunyah Type-1 hypervisor occupying EL2, preventing Linux from accessing HYP mode. KVM requires direct EL2 access to function and cannot initialize when Linux runs as a guest at EL1 under Gunyah. This is a platform configuration constraint, not a kernel regression. The PR (SMMU runtime PM fixes) does not touch KVM or virtualization code.
  3. Possible fix: Disable KVM tests for Purwa IoT EVK in the LAVA test suite, as this platform is configured for Gunyah-based virtualization and cannot support KVM. Alternatively, if KVM support is required, the platform must be reconfigured to boot Linux directly at EL2 without Gunyah (requires bootloader/firmware changes and is outside kernel scope).
  4. Detail analysis attachment: failed_case_job225918_5_detailed.md
  Case 6: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM infrastructure test failed because the Purwa IoT EVK platform does not support ARM virtualization extensions (EL2/HYP mode) — kernel reports "HYP mode not available" during boot, preventing /dev/kvm device creation despite CONFIG_KVM being enabled.
  3. Possible fix: This is a pre-existing platform limitation, not a PR-introduced regression. The PR modifies SMMU runtime PM handling and is unrelated to KVM. Recommended action: exclude KVM tests from the Purwa IoT EVK CI test suite, as this platform does not support virtualization extensions. If KVM support is required, verify that the bootloader/firmware enables EL2 and that the SoC variant includes virtualization extensions.
  4. Detail analysis attachment: failed_case_job225918_6_detailed.md
Job 225919 | SoC qcs6490-rb3gen2

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225919

Failed test cases in LAVA job 225919 (SoC: qcs6490-rb3gen2).

  Case 1: ** Probe_Failure_Check
  1. Failed case: ** Probe_Failure_Check
  2. Root cause: ** The Probe_Failure_Check test flags two benign missing-firmware errors as failures: (1) regulatory.db firmware missing for cfg80211 wireless regulatory database (non-critical; cfg80211 falls back to built-in regulatory rules), and (2) renesas_usb_fw.mem firmware missing for an optional Renesas xHCI PCIe controller at 0001:04:00.0 (non-critical; device is not required for qcs6490-rb3gen2 platform functionality). Neither error prevents system boot, causes device malfunction, or is introduced by the PR under test (PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 modifies only arm-smmu-qcom runtime PM handling in qcom_adreno_smmu_get_fault_info() and qcom_adreno_smmu_set_ttbr0_cfg(), unrelated to firmware loading or device probe paths). The test's grep-based probe failure detection does not distinguish between critical and benign firmware load failures.
  3. Possible fix: Refine the Probe_Failure_Check test to suppress known-benign firmware load failures: add exclusion patterns for regulatory.db (cfg80211 regulatory database — always optional, built-in rules used as fallback) and renesas_usb_fw.mem (optional PCIe USB controller firmware for non-essential hardware not required on qcs6490-rb3gen2). Alternatively, update the test to cross-check firmware load failures against functional device tests (e.g., if WiFi_OnOff passes, suppress WiFi-related firmware warnings; if USB enumeration succeeds, suppress USB controller firmware warnings). For this specific PR validation, mark Probe_Failure_Check as a false positive and approve PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 — the failures are pre-existing infrastructure noise unrelated to the SMMU runtime PM changes, and all critical platform tests (SMMU, WiFi_OnOff, IPA, UFS_Validation, CPU_affinity, Freq_Scaling) passed successfully.
  4. Detail analysis attachment: failed_case_job225919_1_detailed.md
  Case 2: USBHost
  1. Failed case: USBHost
  2. Root cause: The xhci-pci-renesas USB host controller driver failed to probe with error -2 (ENOENT) because the required firmware file renesas_usb_fw.mem is missing from the filesystem, preventing USB host functionality from being available on the PCIe-attached Renesas USB controller.
  3. Possible fix: Add the missing Renesas USB firmware package (linux-firmware or renesas-usb-firmware) to the root filesystem image used by this LAVA job, or if the Renesas USB controller is not required for this platform, disable CONFIG_USB_XHCI_PCI_RENESAS in the kernel configuration to prevent the probe attempt.
  4. Detail analysis attachment: failed_case_job225919_2_detailed.md
  Case 3: BT_SCAN — Bluetooth Device Discovery Test
  1. Failed case: BT_SCAN — Bluetooth Device Discovery Test
  2. Root cause: Test infrastructure limitation — no discoverable Bluetooth devices present in the LAVA lab environment during the 6-minute scan window (3 attempts with multiple fallback strategies). The Bluetooth stack is fully functional (hci0 adapter operational, BT_ON_OFF test passed, firmware loaded, discovery mode successfully enabled/disabled), but the test requires at least one external Bluetooth device within RF range to pass, which was not available during this CI run.
  3. Possible fix: This is not a kernel regression introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (SMMU runtime PM fixes). The failure is environmental. Recommended actions: (1) Deploy a dedicated Bluetooth beacon/test device in the LAVA lab rack for qcs6490-rb3gen2 to ensure consistent BT_SCAN test coverage, or (2) Mark BT_SCAN as an optional/informational test that does not block PR merges when the only failure mode is "no devices found" with all Bluetooth functional tests (BT_ON_OFF, BT_FW_KMD_Service) passing.
  4. Detail analysis attachment: failed_case_job225919_3_detailed.md
  Case 4: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: KVM cannot initialize because HYP (EL2 hypervisor) mode is not available on qcs6490-rb3gen2 — the kernel message "kvm [1]: HYP mode not available" at boot indicates the platform firmware/bootloader did not boot the kernel at EL2 or EL2 is disabled in the SoC configuration.
  3. Possible fix: This is a platform/firmware limitation, not a kernel regression introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which only modifies SMMU runtime PM handling). To enable KVM on qcs6490-rb3gen2: (1) verify the bootloader (ABL/UEFI) is configured to boot Linux at EL2, (2) confirm the SoC supports virtualization extensions and they are not fused off, (3) check if a hypervisor (e.g., Gunyah) is already running at EL2 preventing KVM from taking control. If the platform fundamentally does not support EL2 access for Linux, mark the KVM tests as expected-fail for this target.
  4. Detail analysis attachment: failed_case_job225919_4_detailed.md
  Case 5: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: KVM driver initialization failed at boot with "HYP mode not available" — the qcs6490-rb3gen2 platform is not running in EL2 hypervisor mode, preventing KVM from creating the /dev/kvm device node required by the test.
  3. Possible fix: This is a platform/firmware configuration issue, not a kernel regression introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which only modifies SMMU runtime PM handling). The board must boot with EL2 enabled in the bootloader/firmware to support KVM. If KVM support is required for this platform, update the bootloader configuration to enable EL2; otherwise, skip KVM tests on qcs6490-rb3gen2 in the CI test matrix.
  4. Detail analysis attachment: failed_case_job225919_5_detailed.md
  Case 6: KVM Infrastructure Test — KVM driver initialization failure
  1. Failed case: KVM Infrastructure Test — KVM driver initialization failure
  2. Root cause: KVM driver failed to initialize because HYP mode (EL2) is not available on the qcs6490-rb3gen2 platform. CONFIG_KVM is enabled in the kernel, but the hardware/firmware does not provide EL2 access required for ARM64 virtualization. This is a platform limitation, not a kernel regression.
  3. Possible fix: This is not a bug introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which only modifies SMMU runtime PM handling). The KVM tests should be disabled or marked as expected-fail for qcs6490-rb3gen2 in the LAVA test suite, as this platform does not support KVM/virtualization. If virtualization support is required, verify firmware configuration and ensure the bootloader enables EL2, or use a different SoC with virtualization support.
  4. Detail analysis attachment: failed_case_job225919_6_detailed.md
  Case 7: KVM_Infra
  1. Failed case: KVM_Infra
  2. Root cause: KVM/HYP mode not available on qcs6490-rb3gen2 platform — kernel message "kvm [1]: HYP mode not available" indicates the platform does not support virtualization extensions or they are not enabled in the bootloader/firmware configuration.
  3. Possible fix: This is a pre-existing platform limitation, not a PR-introduced regression. The PR changes (SMMU runtime PM fixes) are unrelated to KVM functionality. If KVM support is required on this platform, verify: (1) the SoC supports virtualization extensions (EL2), (2) the bootloader/firmware enables HYP mode, and (3) the device tree includes the necessary KVM/hypervisor configuration. Otherwise, mark KVM tests as "not applicable" for this platform in the CI test suite.
  4. Detail analysis attachment: failed_case_job225919_7_detailed.md
Job 225920 | SoC hamoa-evk

LAVA job: https://lava-oss.qualcomm.com/scheduler/job/225920

Failed test cases in LAVA job 225920 (SoC: hamoa-evk).

  Case 1: Probe_Failure_Check
  1. Failed case: Probe_Failure_Check
  2. Root cause: Could not be determined confidently from available logs.
  3. Possible fix: These failures are NOT introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (SMMU runtime PM changes). They are pre-existing hamoa-evk platform issues. To resolve: (1) Update TrustZone firmware or fix QSEECOM resource allocation, (2) Correct the device tree multi-LED "reg" property for c42d000.spmi:pmic@1:pwm node, (3) Add regulatory.db firmware file to the rootfs image. The PR itself is not the cause and should not be blocked by this test failure.
  4. Detail analysis attachment: failed_case_job225920_1_detailed.md
  Case 2: smmu (IOMMU Group Attachment Validation Failure)
  1. Failed case: smmu (IOMMU Group Attachment Validation Failure)
  2. Root cause: The smmu test validates that critical platform devices (USB PHYs and video codec) are attached to IOMMU groups for memory protection. On hamoa-evk (X1E80100), five USB PHY devices (a0f8800.usb, a2f8800.usb, a4f8800.usb, a6f8800.usb, a8f8800.usb) and one video codec device (aa00000.video-codec) are missing IOMMU group attachments. This is a pre-existing device tree or platform configuration issue, not introduced by PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 which only adds runtime PM protection to SMMU register access paths.
  3. Possible fix: This is not a PR-introduced regression. The PR changes (runtime PM protection in qcom_adreno_smmu_get_fault_info and qcom_adreno_smmu_set_ttbr0_cfg) do not affect IOMMU group attachment. The missing IOMMU groups indicate that the device tree for hamoa-evk either: (1) does not declare iommus properties for these USB PHY and video codec devices, or (2) the devices are not probed/registered. To fix: add iommus properties to the missing devices in arch/arm64/boot/dts/qcom/x1e80100*.dtsi, or mark the test as expected-fail for hamoa-evk if these devices are intentionally not IOMMU-protected on this platform.
  4. Detail analysis attachment: failed_case_job225920_2_detailed.md
  Case 3: KVM_Driver
  1. Failed case: KVM_Driver
  2. Root cause: KVM driver initialization failed because the hypervisor (Gunyah) is running at EL2, preventing KVM from accessing HYP mode — kernel log shows "kvm [1]: HYP mode not available" at boot time.
  3. Possible fix: This is a platform configuration issue, not a PR-introduced regression. The hamoa-evk board boots with Gunyah hypervisor at EL2, which prevents KVM from initializing. To enable KVM on this platform, either: (1) boot without the Gunyah hypervisor, or (2) use nested virtualization support if available. The PR changes (SMMU runtime PM fixes) are unrelated to KVM initialization and do not cause this failure.
  4. Detail analysis attachment: failed_case_job225920_3_detailed.md
  Case 4: KVM_EL2_DTB
  1. Failed case: KVM_EL2_DTB
  2. Root cause: Platform hardware limitation - Hamoa IoT EVK does not support ARM Virtualization Extensions (EL2/HYP mode), which is required for KVM to function; kernel correctly reports "HYP mode not available" and does not create /dev/kvm device node.
  3. Possible fix: Exclude KVM test suite from Hamoa IoT EVK CI runs by adding platform capability checks to the LAVA job definition; KVM tests should only run on platforms with virtualization hardware support (e.g., platforms where /proc/cpuinfo shows "virt" in CPU features or where EL2 is available).
  4. Detail analysis attachment: failed_case_job225920_4_detailed.md
  Case 5: ** KVM Infrastructure Unavailable — HYP Mode Not Available
  1. Failed case: ** KVM Infrastructure Unavailable — HYP Mode Not Available
  2. Root cause: ** ARM CPU not running in a mode that allows EL2 (hypervisor exception level) access; kernel message kvm [1]: HYP mode not available at boot time (line 3975) indicates KVM cannot initialize because ARM virtualization extensions are unavailable to Linux, likely because the bootloader/firmware on Hamoa EVK does not grant Linux EL2 access or because Gunyah hypervisor is occupying EL2 and running Linux as a guest VM.
  3. Possible fix: This is a pre-existing platform/firmware configuration issue unrelated to PR FROMLIST: iommu: arm-smmu-qcom: Skip fault-info reads when suspended #1088 (which only modifies SMMU runtime PM handling in arm-smmu-qcom.c). To enable KVM on Hamoa EVK: (1) configure the bootloader (ABL/UEFI) to boot Linux with EL2 or VHE (Virtualization Host Extensions) access, (2) verify that Gunyah hypervisor is not occupying EL2 (if Gunyah is present, KVM cannot coexist — choose one virtualization model), or (3) accept that this platform is configured for Gunyah-based virtualization and mark KVM tests as not applicable for Hamoa EVK in the CI test matrix.
  4. Detail analysis attachment: failed_case_job225920_5_detailed.md
  Case 6: ** KVM_Infra (driver initialization failure — platform does not support EL2/HYP mode)
  1. Failed case: ** KVM_Infra (driver initialization failure — platform does not support EL2/HYP mode)
  2. Root cause: ** The Hamoa IoT EVK platform does not support KVM virtualization because the CPU is not running in EL2 (Hypervisor Exception Level) mode. The kernel message "kvm [1]: HYP mode not available" at boot time (6.373348s) indicates that the bootloader/firmware boots the kernel in EL1 mode and does not enable EL2 access. This is a platform hardware/firmware configuration limitation, not a kernel software bug. The PR changes (SMMU runtime PM fixes in arm-smmu-qcom.c) are completely unrelated to KVM and did not cause this failure.
  3. Possible fix: Exclude KVM tests from the Hamoa IoT EVK test matrix in CI — the platform does not support virtualization, so these tests will always fail. Add a platform capability check in the LAVA job definition to skip KVM tests when /sys/module/kvm/parameters/ is absent or when the kernel log contains "HYP mode not available". If KVM support is required on this platform, contact the platform vendor (Qualcomm) to obtain firmware that enables EL2 mode, which may require bootloader (ABL/XBL) and TrustZone configuration changes.
  4. Detail analysis attachment: failed_case_job225920_6_detailed.md

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants