Skip to content

fix(daemon)!: harden normal mode, settle loopback, remove Fast Allow - #56

Open
MotherSphere wants to merge 12 commits into
mainfrom
normal-mode-hardening
Open

MotherSphere wants to merge 12 commits into
mainfrom
normal-mode-hardening

Conversation

@MotherSphere

Copy link
Copy Markdown
Member

Hardens normal mode, removes the disabled Fast Allow path, settles the loopback policy, and adds the first CI job that proves the firewall really blocks.

Normal-mode hardening (8bbbb30)

Findings from the second Codex Security scan of 0.7.0: attribution now uses the real PID and process generation (a recycled PID no longer inherits a verdict), TCP flows are no longer attributed to listening sockets, and new loopback flows follow explicit rules instead of being accepted wholesale.

Loopback policy

8bbbb30 dropped oifname "lo" accept, which made every new loopback flow fail closed when no daemon runs (the systemd-resolved stub, CUPS and local dev servers all stopped). The outbound snippet now has

oifname "lo" ct state new queue num 0 bypass
ct state new queue num 0

While the daemon runs it judges loopback flows like any other (explicit rules apply, unmatched local IPC is allowed without prompting). While nothing listens on the queue, new loopback flows are accepted; every non-loopback new flow stays fail-closed. The limits are documented in HARDENING.md and TROUBLESHOOTING.md: in that window explicit loopback Deny rules are not enforced and nothing is logged, and the stub resolver only answers from its cache because its upstream queries are not loopback. scripts/vm-bench/plan.sh gains a loopback UDP latency measurement (lo-floor / lo-queue).

Fast Allow removed (userspace side)

Fast Allow was disabled during the Codex Security remediation because it opened bypasses, but about 300 references were still compiled in. The daemon, CLI, proto and xtask side is removed (-3247 lines). The legacy cleanup stays: the empty nft set is still flushed, and pinned eBPF maps left armed by a 0.4 to 0.6 daemon are disarmed and their sendmsg links unpinned at start. The kernel program keeps its ABI for now. The status field and proto entries are gone (field numbers reserved), hence the breaking marker.

Armed e2e job

.github/workflows/e2e.yml and scripts/armed-e2e.sh run the real daemon (eBPF off) against the shipped snippet in a network namespace on the runner and assert that a denied flow is dropped, an allowed flow gets through, a new flow is dropped after the daemon is killed, and a new loopback flow still works at that point. This PR is its first run.

Docs and CI hygiene

SECURITY.md now says beta and states plainly that there has been no external audit; the confinement mode is marked experimental; the README crate table matches the workspace; release.yml no longer claims to be AUR-submittable and passes the tag name through env:; dependabot groups minor and patch updates.

Checks run locally

  • cargo fmt --all -- --check, cargo clippy --workspace --all-targets --locked -- -D warnings, cargo clippy -p cfc-daemon --no-default-features --all-targets --locked -- -D warnings: clean.
  • cargo test --workspace --locked: 881 passed, 0 failed, 10 ignored (root-only). cargo test -p cfc-daemon --no-default-features --locked: 428 passed.
  • MSRV 1.89 cargo check, check-versions.sh, check-release-assets.sh, check-startup-protection.py: pass.
  • Both nft rulesets parse (nft -c -f in an unprivileged namespace).

Not run here: the root-only tests (cargo test --workspace -- --ignored with sudo), the eBPF kernel matrix, and the e2e job itself, which needs the CI runner.

Fast Allow was disabled during the security remediation because a socket
mark cannot prove which process sends a packet. Remove what stayed
compiled in: the grant writers, the eligibility ladder, mark drawing, the
heartbeat, the ALLOW_EVENTS consumer, the sendmsg attach and the status
reporting.

What remains is legacy cleanup for hosts upgrading from 0.4-0.6: one nft
flush of the fast_allow set at start (outside --dry-run), a disarm of the
pinned kernel maps on every load, and removal of the old sendmsg pins and
cookie-variants marker. The [ebpf] fast_allow keys still parse and only
log a warning. StatusResponse field 16 is reserved, and cfc --json status
no longer has a fast_allow key.
The loader no longer looks up cfc_sendmsg4/6, so the object check stops
demanding them. The FAST_ALLOW* maps and ALLOW_EVENTS stay required: they
are still pinned by name so startup can disarm what an older release
armed. The kernel matrix summary drops its "fast path" row.
Fast Allow opened bypasses because a socket mark cannot prove which
process sends, and its userspace side is gone. Say so in the README,
architecture, roadmap, TODO, troubleshooting and hardening notes, and
reduce the sample config to a note that the legacy keys are ignored.

Packaging comments now describe nft as a one-shot flush at start plus
the table probe. The VM bench drops its fast-N states and no longer
reads the removed fast_allow status key.
The root test arms the pinned FAST_ALLOW_MARK to prove the next load
disarms it, while the inherited connect hooks are still attached. Use a
deadline already in the past and a grant for a pid no process can hold,
so nothing on the test host is marked in that window.
The smoke job drives a --dry-run daemon, so it never proves a packet is
dropped. The new e2e workflow loads the shipped nftables snippet inside a
throwaway network namespace, runs colony-firewalld there and checks with
curl that an allow rule passes, a deny rule and an unmatched flow are
silently dropped, the audit log names the deciding rule, rules persist
across a restart, and new flows stay dropped after SIGTERM and SIGKILL.
Nothing outside the namespace is filtered.
SECURITY.md now says beta, states plainly that there has been no external
security audit, and marks explicit application confinement as experimental,
as does its README section. The README crate table lists all ten crates,
including cfc-ebpf-common, cfc-ebpf and xtask.

release.yml passes github.ref_name to run scripts through env instead of
interpolating it, and drops the AUR-submittable wording: the project is not
published on the AUR. Dependabot groups cargo minor+patch and GitHub Actions
updates, and covers transitive cargo dependencies.
Add `oifname "lo" ct state new queue num 0 bypass` before the final queue
rule. The daemon still judges loopback flows while it runs; when nothing
listens on the queue, local IPC such as the systemd-resolved stub keeps
working, and every other new flow stays fail-closed.

Docs, packaging notes and doc comments now describe the ruleset as
fail-closed for everything except new loopback flows. The VM bench gains a
loopback UDP echo measurement (floor vs armed), and the armed e2e asserts
that a new loopback TCP flow succeeds after the daemon is killed.
With the daemon down, the systemd-resolved stub still answers from its
cache, but its upstream queries are new non-loopback flows and stay
blocked. While no daemon listens, explicit loopback Deny rules are not
enforced, nothing is logged, and a loopback connection opened then keeps
its authorization afterwards. Say so in HARDENING, TROUBLESHOOTING, the
README and TODO, and fix the inbound snippet comment that still described
an unconditional outbound loopback accept.
Clippy 1.99, now the stable the CI resolves, flags the #[async_trait]
that tonic emits for the server trait. The code is generated, so the
lint is allowed on the generated module only; unknown_lints keeps older
clippy versions from rejecting the name.
The real-machine test pinned curl 8.21.0-1 and failed as soon as the host updated curl, although provenance still verified. Assert the owning package instead.
The previous commit left a second pinned curl version in the foreign-digest assertion. Compare with the package the same test just resolved.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant