Skip to content

Fix #1: Overview stuck on "Reading system state" (combine starvation + 4 secondary defects) - #4

Merged
Dreamucxe merged 9 commits into
mainfrom
fix/issue-1-reading-system-state
Oct 9, 2026
Merged

Dreamucxe merged 9 commits into
mainfrom
fix/issue-1-reading-system-state

Conversation

@Dreamucxe

Copy link
Copy Markdown
Owner

Fixes #1 — Overview stuck on "Reading system state"

TL;DR

The dashboard hung on every launch, on every device, root or not. The SELinux
denials in the report are real and are fixed here too, but they are a second
problem — not the cause of the hang. The hang is a combine() that never emits.

Root cause

OverviewViewModel.state is a kotlinx.coroutines.flow.combine of five upstream
flows. combine is all-or-nothing: it emits nothing until every source has
emitted at least once. One source — observeCapabilities() — is a
MutableStateFlow(null) exposed through filterNotNull(), so it stays silent
until something calls refreshCapabilities(). Nothing on the cold-start path
ever does.
The state therefore never leaves its initialValue = State()
(system == null), and OverviewScreen renders LoadingBlock("Reading system state") forever.

Opening Settings → Access constructs AccessViewModel, whose
init { refresh() } writes the capabilities field. Because the repository is a
@Singleton, the flow then stays non-null for the rest of the process and every
screen unblocks at once. That is why "Probe root" looked like the fix — merely
navigating to that screen was enough; the root grant was incidental. Relaunching
resets the in-memory singleton to null, which is why it reproduced on every
launch.

The intended initial state already existed in the code —
SystemCapabilities.unknown(apiLevel), KDoc "Used only as an initial UI state
before the first real evaluation"
— with zero production callers. The fix
wires it up.

A note on the earlier read of this issue: the reporter's hypothesis (a non-root
collector treating denied reads as "not ready" and retrying forever) was a sharp,
reasonable guess from the avc: denied logs, and the thread understandably
converged on it. It turns out not to be what hangs the screen — every /proc
and /sys read already funnels through ProcFsReader, which catches the denial
and returns a terminal Observed.Restricted; a complete SystemState was being
produced the whole time and then silently dropped by the starved combine. The
denials were a real but separate defect, and they are addressed below. Full
evidence is in docs/ISSUE1_ROOT_CAUSE.md.

Changes

Five fixes, each in disjoint files, one small commit apiece:

# What Key files
F1 Seed SystemCapabilities.unknown() + an onStart kick so the capabilities flow emits before detection finishes — the dashboard resolves on frame one. data/repository/SystemRepositoryImpl.kt
F2 Per-metric negative cache: a terminal denial (platform-restricted, not-present, requires-elevated, unsupported-API) is attempted once per session, not re-read on every 2 s tick. Permission- and user-toggle refusals are deliberately not cached. At most one log line per metric per session; no raw denial spam or su output in release. new core/system/RestrictionCache.kt, core/system/ProcFsReader.kt
F3 Real shell deadline. Probes run on Dispatchers.IO under a wall-clock timeout; stdout/stderr are drained concurrently (the old drain-stdout-then-stderr order could deadlock on a full 64 KiB stderr pipe); on timeout/denial/exception the process is destroyed, streams closed, and collection continues non-root. new core/system/ProcessRunner.kt, core/system/RootShell.kt, core/system/ShizukuShell.kt
F4 Preserve a proven root grant across a route invalidation (invalidateAccess() used to wipe a just-obtained grant before routing re-evaluated, so root never actually engaged). core/system/CompositeSystemObserver.kt
F5 Dashboard always resolves: a spinner before the deadline, then an honest explanation (never withheld data); per-metric "N/A (restricted by Android)" with a hint that Shizuku/root unlocks it; a non-blocking no-root banner; a reachable manual re-probe. feature/overview/OverviewScreen.kt, feature/overview/OverviewViewModel.kt, core/designsystem/ScreenScaffold.kt

Infra: .github/workflows/build.yml (debug build + APK artifact on push); R8 + resource shrinking enabled for release by default; host-specific Gradle config moved out of the repo so one checkout builds on both CI and an ARM device.

Tests

New JVM unit tests, no device or Compose harness required — the decisions were
extracted into pure functions so the guarantees are directly assertable:

  • OverviewLoadStateTest — all-succeed / some-denied / all-denied all reach
    CONTENT; spinner before the deadline; UNRESOLVED (explanation, not a
    frozen spinner) after it; immediate resolve on a thrown pipeline.
  • RestrictionCacheTest — terminal denial memoised once; permission &
    sampling-disabled refusals re-read every time; failures and successes never
    cached; per-metric independence; one-shot announce budget.
  • ProcessRunnerTest — exit-zero, non-zero, su missing, su thrown, empty
    argv, and su hangs → times out, destroys the child, returns near the
    deadline; both pipes drained and kept separate; partial output before a hang
    survives.

Full suite: 435 passing, 0 failures. assembleDebug green.

Not included

No merge to main is requested here. Signing config, keystore.properties, and
the release keystore are intentionally untracked and absent from every commit.

Closes #1.

root added 9 commits October 7, 2026 08:26
Records the investigation before any code changes, from four independent
reads of the source.

The headline finding is that the SELinux denials in the report are not the
cause. Overview hangs because OverviewViewModel.state is a combine() of five
flows and one of them, observeCapabilities(), never emits on a cold start:
the backing MutableStateFlow is seeded null and exposed through
filterNotNull(), and nothing on the launch path writes it. combine publishes
nothing until every source has emitted, so the state never leaves its initial
value and the screen renders its loading block forever.

That also explains why "Probe root" appears to fix it. Navigating to
Settings > Access constructs AccessViewModel, whose init refreshes
capabilities; the repository is a singleton, so the flow stays non-null for
the rest of the process. Merely opening the screen is enough, and relaunching
resets it, which is why the hang reproduces every launch.

Four secondary defects are recorded with their consequences, including the
denial re-read on every poll tick that produced the repeating avc lines, and
a genuine pipe deadlock in both shells that makes their timeouts unreachable.
This is the fix for the hang in issue #1.

observeCapabilities() was capabilities.asStateFlow().filterNotNull(), over a
MutableStateFlow seeded with null. The filter swallowed the seed, so the flow
stayed silent until refreshCapabilities() happened to run. Twelve view models
combine() on it, and combine emits nothing until every source has produced a
value, so all twelve sat on their initial state. On Overview that is
system == null, which renders LoadingBlock("Reading system state") behind an
unconditional early return.

The repository now seeds the flow with SystemCapabilities.unknown(SDK_INT),
which already existed for exactly this purpose and whose KDoc says so: its
statuses map is empty, so every lookup falls back to "Not evaluated on this
device" and it renders as honest absence rather than invented capability.
With a real first value on the wire, combine resolves immediately and the
screen draws its content while detection is still in flight.

The first collector kicks that detection once per process via the application
scope, so a screen leaving the composition mid-detection cannot cancel work
the next screen needs, and a detection that throws leaves the placeholder in
place instead of restoring the spinner. refreshCapabilities() also marks the
kick as spent, so a caller that beats the first collector does not cause a
second evaluation.

Detection cannot prompt for root: RootShell.probe() is the only su spawn and
its sole caller is the explicit "Probe root" action on the Access screen.
gradle.properties carried the authoring device's own setup: an absolute
org.gradle.java.home, an aapt2FromMavenOverride pointing at /usr/bin/aapt2 for
an aarch64 host, and heap and daemon caps sized for roughly 1.8 GB of free RAM.
Those settings are correct for that machine and wrong everywhere else — a CI
runner reading them fails on the first path that only the authoring device has.

They are also redundant there. Gradle resolves GRADLE_USER_HOME/gradle.properties
ahead of the project file, and that is where this host already defines all of
them, so removing them from the repository changes nothing about how it builds
locally while making the checkout portable.

What stays is the portable half: AndroidX and non-transitive R classes, plus a
heap large enough for KSP and Compose in one invocation. A constrained host
overrides that downward in its own GRADLE_USER_HOME rather than lowering it in
the repository for everyone.
Release builds had isMinifyEnabled and isShrinkResources both false. The comment
explained it as a build-host concession: R8 is memory-hungry and the on-device
ARM host cannot run it. That is true, but it meant the shipped artifact was
unobfuscated and unshrunk because of where it happened to be compiled, which is
a property of the machine leaking into the product.

Both now default to on and are driven by a processlens.minify property. The
constrained host sets processlens.minify=false in its own GRADLE_USER_HOME, so
the opt-out stays on that one machine and every other build, CI included, gets a
minified and shrunk release. The keep-rules in proguard-rules.pro were already
written for Hilt, Room, Shizuku and Compose, so they now actually apply.

Shrinking resources requires shrinking code, so the two are driven by one flag
rather than being independently settable into an invalid combination.
There was no .github directory at all, so nothing verified that a change
compiled or that the tests passed outside the author's device.

The workflow runs the unit tests first and then assembles the debug APK,
uploading the APK as an artifact and the test report unconditionally — the
report is wanted precisely when the tests fail. Pinned to JDK 17 to match
AGP 8.5 and the Kotlin 2.0 toolchain. Concurrency is grouped by ref so a newer
push cancels the run it supersedes.

It deliberately does not patch any Gradle setting before building; that is only
possible because the host-specific configuration no longer lives in the
repository.
The note said twelve view models combine() on observeCapabilities() and then
listed eleven. Eleven is right.

Also records a second piece of evidence that the placeholder value was the
intended design rather than a convenient fix: AppShellViewModel.summarise()
already opens with a total == 0 branch returning "Checking what this device
allows...", and that total is the sum of the three statuses counts, so the
branch is reachable only from a capabilities object with an empty statuses map
— precisely what SystemCapabilities.unknown() produces and what nothing was
emitting. The handler, the value and the gate that made both unreachable were
all written by the same hand.

Notes that the empty statuses map is also the signal a consumer needs to tell
"not yet evaluated" from "evaluated, nothing allowed", since the placeholder
reports NORMAL access and UNAVAILABLE root before any detection has run.
A refusal is an answer. When SELinux denies /proc/loadavg or the thermal
nodes, the kernel returns the same denial to every subsequent identical
read, but the sampler asked again every one to ten seconds — one syscall,
one audit record and one `avc: denied` line per attempt. A field bug report
carried fifty minutes of that spam; nothing was learned after the first.

RestrictionCache remembers the first terminal refusal per logical metric and
returns the same Observed.Restricted to later callers without touching the
filesystem. Only platform-shaped reasons are terminal: PERMISSION_REQUIRED
and SAMPLING_DISABLED stay revocable, and Observed.Failed is never memoised
so a single bad sample cannot blank a screen for the session. The cache is
clearable so a root/Shizuku grant re-opens every read.
Both elevated shells drained stdout to EOF, then stderr, then consulted the
timeout — so the timeout was never reached while an `su` sat on an unanswered
superuser prompt, and a child that filled its ~64 KiB stderr pipe deadlocked
because it never closed stdout. ProcessRunner now drains both pipes
concurrently in a scope it owns, enforces the deadline by destroying the
process (the only thing that releases a reader parked in a blocking read()),
and tears down every stream on every exit path exactly once. Probes run on
Dispatchers.IO under a 3s budget; a timeout, denial or absent binary all come
back as a non-zero result the caller carries on without.

Routing no longer wipes a grant it was just handed: RootShell models its probe
result as Observed, invalidate() keeps a proven Value grant while clearing
stale negatives, and forgetGrant() drops even a proven grant when the user
turns root support off. CompositeSystemObserver and CapabilityDetector drop
the restriction cache on an access-level change so a grant re-opens reads.
The permanent "Reading system state" of issue #1 was a starved combine()
upstream (fixed in the repository), but it stayed invisible because nothing
on this screen questioned a silent source: a thrown collector took the whole
flow down without a trace, and the only reaction to "no state yet" was a
spinner with no deadline.

Three things now hold. A thrown pipeline is folded by .catch into an
Observed.Failed the screen renders through the same Section 42 components as
any other unreadable value, keeping whatever content was already shown.
Collection is restartable via an attempts counter behind flatMapLatest, so
the retry the user is offered genuinely re-runs the sources. And the screen
bounds its own wait: a spinner gives way to an UnresolvedBlock after a
deadline, with a real re-probe. The capability placeholder is kept
distinguishable from a probed "no elevated access" so the dashboard never
accuses a device of restrictions it has not yet been asked about.
@Dreamucxe
Dreamucxe merged commit 5d33103 into main Oct 9, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Not loading untill probed SU: 'Reading system state'

1 participant