Context
#135 / #136 fixed step attribution for the runner container itself on ARC: kernel-global vs namespace-local pids are now translated in the kernel, so the job's own processes attribute like on a hosted VM. What #136 does not cover is everything that lives outside the runner container's PID and cgroup namespaces, which on ARC means the Docker sidecar (containerMode: dind) and every container it launches.
Constraint for all of this: cargowall also runs as a sidecar in CodeCargo's own SaaS pods (non-GitHub mode, own pid/cgroup namespace, shared netns). Every item below must be opt-in or auto-detected in a way that leaves today's defaults untouched, and that deployment needs its own CI coverage (item F).
What still does not work on ARC, and why
- dind-launched containers are invisible to the cgroup half. The connect/sendmsg hooks (
cmd/start.go:554), cg_sock_create (pkg/steps/steps.go:302) and cg_origin_egress (pkg/origin/origin.go:335, :499) attach at /sys/fs/cgroup, which under a cgroup namespace is the runner container's cgroup. The dind container is a sibling; its subtree is never hooked. Container flows therefore reach TC post-NAT with pid=0, step_ordinal=0, there is no pre-NAT origin record for the enricher to join, and --container-egress observe|enforce is blind to them. A cgroup namespace cannot see or attach to ancestors, so there is no attach point inside the runner container that fixes this.
- dockerd's pids belong to another namespace.
State.Pid from inspect is numbered in the dind container's pid namespace; verifyContainerTask correctly rejects it and TagContainerProcess becomes a no-op. setns into an ancestor pid namespace is refused by the kernel, so this cannot be worked around from userspace either.
- Container DNS bypasses the proxy (inferred, not yet observed on a live pod).
network.ConfigureDockerDNS rewrites /etc/docker/daemon.json in the runner container's mount namespace; the sidecar dockerd never reads it. The DNS redirect rules (pkg/network/dns_redirect.go) are OUTPUT-only, so queries forwarded from docker0 are not DNAT'd. dind containers keep the pod resolver, resolve outside the proxy, and in enforce mode their connections hit IPs the proxy never allowed. Needs confirming on codecargo-dev-runner-* with a job that runs a container.
What is shared across the pod is the network namespace: eth0, docker0 and any br-* bridges, iptables, conntrack. That is what the proposal leans on.
Proposal
A. Pod-level cgroup attach (opt-in). Add a --cgroup-root flag (or auto-detect a host cgroup2 bind mount, e.g. hostPath /sys/fs/cgroup mounted at a known path). Locate our own container's cgroup in that tree by inode, take its parent as the pod cgroup, and attach every cgroup program there; also feed it to containers.Options.CgroupRoot (pkg/containers/tracker.go:221) so cgroup-id classification works for sibling subtrees. A bind mount of an existing cgroup2 mount should expose the full hierarchy despite the cgroup namespace (the namespace only constrains new mounts); this is the first thing to prove on our nodes. Default stays /sys/fs/cgroup.
B. Shared PID namespace. With shareProcessNamespace: true on the runner pod, the #136 iterator already covers every pod container's tasks, Runner.Worker ancestry discovery works from any container, and dockerd's State.Pid values become valid in our /proc. Check that verifyContainerTask accepts the /..-prefixed cgroup path a sibling-subtree process shows (the Contains(containerID) test should), and that the v2 cgroup id resolves through the root from A. Document it; no code change expected beyond A.
C. Bridge-side fallback when A/B are not available. A TC observer on docker0 and docker network bridges (shared netns, so it works even against a sidecar dockerd) sees container traffic pre-NAT and can emit the same origin Record shape the enricher already joins on via Observer.LookupV4(dst, dport, proto, srcPort). Container → step already comes from docker events (handleStart with timeNano, exec tracking), no pids required. Optional: enforce at bridge ingress. The cheaper alternative is a conntrack lookup for pid=0 TC events to recover the pre-NAT tuple (attribution only). Pick one; the bridge observer is the more complete option.
D. Container DNS on ARC. Add PREROUTING DNAT for udp/tcp 53 arriving on docker bridge interfaces to the proxy's container listen address (dockerBridgeIP:53 already exists), so dind containers are gated without needing daemon.json. Detect that dockerd is not local (its pid is not in our /proc) and say so instead of silently writing a daemon.json nobody reads.
E. Detection and operator UX. At startup, log clearly when: running in a non-init pid namespace (done in #136); dockerd runs in another pid namespace (with the shareProcessNamespace remedy); the cgroup root is a namespace root and no host cgroup mount was given. Ship a documented ARC runner template snippet: privileged, shareProcessNamespace: true, hostPath cgroup mount, --cgroup-root.
F. Guardrails. A CI job that runs cargowall in the sidecar shape — non-GitHub mode, own pid and cgroup namespace, shared netns with a client container — asserting enforcement and DNS redirect still work. Nothing in scripts/ci covers that today; it is the production SaaS surface. Ideally also an ARC-shaped e2e (kind: runner + dind + a job that starts a container) once A–D land.
Acceptance
- On an ARC runner with
containerMode: dind and the documented template: a docker run in a step attributes to that step in the summary and on the dashboard, --container-egress enforce blocks a disallowed destination from inside the container with a correctly attributed event, and container DNS goes through the proxy.
- Without the template changes: same behavior as today plus explicit warnings naming what is unavailable and why; no change for hosted runners.
- SaaS sidecar CI job green before and after.
Not in scope
A node-level DaemonSet agent (host pid/cgroup namespaces, per-pod policy) would remove the pod template requirements entirely and is the long-term shape for ARC at scale; separate discussion.
Related: #135, #136, code-cargo/cargowall-action#86.
Context
#135 / #136 fixed step attribution for the runner container itself on ARC: kernel-global vs namespace-local pids are now translated in the kernel, so the job's own processes attribute like on a hosted VM. What #136 does not cover is everything that lives outside the runner container's PID and cgroup namespaces, which on ARC means the Docker sidecar (
containerMode: dind) and every container it launches.Constraint for all of this: cargowall also runs as a sidecar in CodeCargo's own SaaS pods (non-GitHub mode, own pid/cgroup namespace, shared netns). Every item below must be opt-in or auto-detected in a way that leaves today's defaults untouched, and that deployment needs its own CI coverage (item F).
What still does not work on ARC, and why
cmd/start.go:554),cg_sock_create(pkg/steps/steps.go:302) andcg_origin_egress(pkg/origin/origin.go:335,:499) attach at/sys/fs/cgroup, which under a cgroup namespace is the runner container's cgroup. The dind container is a sibling; its subtree is never hooked. Container flows therefore reach TC post-NAT withpid=0, step_ordinal=0, there is no pre-NAT origin record for the enricher to join, and--container-egress observe|enforceis blind to them. A cgroup namespace cannot see or attach to ancestors, so there is no attach point inside the runner container that fixes this.State.Pidfrom inspect is numbered in the dind container's pid namespace;verifyContainerTaskcorrectly rejects it andTagContainerProcessbecomes a no-op.setnsinto an ancestor pid namespace is refused by the kernel, so this cannot be worked around from userspace either.network.ConfigureDockerDNSrewrites/etc/docker/daemon.jsonin the runner container's mount namespace; the sidecar dockerd never reads it. The DNS redirect rules (pkg/network/dns_redirect.go) are OUTPUT-only, so queries forwarded from docker0 are not DNAT'd. dind containers keep the pod resolver, resolve outside the proxy, and in enforce mode their connections hit IPs the proxy never allowed. Needs confirming oncodecargo-dev-runner-*with a job that runs a container.What is shared across the pod is the network namespace: eth0, docker0 and any
br-*bridges, iptables, conntrack. That is what the proposal leans on.Proposal
A. Pod-level cgroup attach (opt-in). Add a
--cgroup-rootflag (or auto-detect a host cgroup2 bind mount, e.g. hostPath/sys/fs/cgroupmounted at a known path). Locate our own container's cgroup in that tree by inode, take its parent as the pod cgroup, and attach every cgroup program there; also feed it tocontainers.Options.CgroupRoot(pkg/containers/tracker.go:221) so cgroup-id classification works for sibling subtrees. A bind mount of an existing cgroup2 mount should expose the full hierarchy despite the cgroup namespace (the namespace only constrains new mounts); this is the first thing to prove on our nodes. Default stays/sys/fs/cgroup.B. Shared PID namespace. With
shareProcessNamespace: trueon the runner pod, the #136 iterator already covers every pod container's tasks,Runner.Workerancestry discovery works from any container, and dockerd'sState.Pidvalues become valid in our/proc. Check thatverifyContainerTaskaccepts the/..-prefixed cgroup path a sibling-subtree process shows (theContains(containerID)test should), and that the v2 cgroup id resolves through the root from A. Document it; no code change expected beyond A.C. Bridge-side fallback when A/B are not available. A TC observer on docker0 and docker network bridges (shared netns, so it works even against a sidecar dockerd) sees container traffic pre-NAT and can emit the same origin
Recordshape the enricher already joins on viaObserver.LookupV4(dst, dport, proto, srcPort). Container → step already comes from docker events (handleStartwithtimeNano, exec tracking), no pids required. Optional: enforce at bridge ingress. The cheaper alternative is a conntrack lookup forpid=0TC events to recover the pre-NAT tuple (attribution only). Pick one; the bridge observer is the more complete option.D. Container DNS on ARC. Add PREROUTING DNAT for udp/tcp 53 arriving on docker bridge interfaces to the proxy's container listen address (
dockerBridgeIP:53already exists), so dind containers are gated without needingdaemon.json. Detect that dockerd is not local (its pid is not in our/proc) and say so instead of silently writing adaemon.jsonnobody reads.E. Detection and operator UX. At startup, log clearly when: running in a non-init pid namespace (done in #136); dockerd runs in another pid namespace (with the
shareProcessNamespaceremedy); the cgroup root is a namespace root and no host cgroup mount was given. Ship a documented ARC runner template snippet:privileged,shareProcessNamespace: true, hostPath cgroup mount,--cgroup-root.F. Guardrails. A CI job that runs cargowall in the sidecar shape — non-GitHub mode, own pid and cgroup namespace, shared netns with a client container — asserting enforcement and DNS redirect still work. Nothing in
scripts/cicovers that today; it is the production SaaS surface. Ideally also an ARC-shaped e2e (kind: runner + dind + a job that starts a container) once A–D land.Acceptance
containerMode: dindand the documented template: adocker runin a step attributes to that step in the summary and on the dashboard,--container-egress enforceblocks a disallowed destination from inside the container with a correctly attributed event, and container DNS goes through the proxy.Not in scope
A node-level DaemonSet agent (host pid/cgroup namespaces, per-pod policy) would remove the pod template requirements entirely and is the long-term shape for ARC at scale; separate discussion.
Related: #135, #136, code-cargo/cargowall-action#86.