Symptom
On Blacksmith runners (Firecracker VM, kernel 6.6.141, BTF present, TCX attach works) the cgroup egress program does not load, so --container-egress enforce and --tls-sni enforce-pinned are both silently off for the job. From a v2.0.0-rc.9 run on blacksmith-2vcpu-ubuntu-2404 (madhokie-io/dockcmd, run 35283540609):
::warning::Container egress hook disabled — enforcement falls back to TC egress only target_mode=enforce error=failed to load origin BPF objects: assign values: field CgOriginEgress: program cg_origin_egress: load program: argument list too long: BPF program is too large. Processed 1000001 insn (1315 line(s) omitted)
::warning::--tls-sni requested but the cgroup egress hook is not running — L7 SNI enforcement is OFF and shared-edge IPs stay L4-only
Everything else in the run is healthy: step attribution works, TC enforcement works, DNS gating works. The job passes its CargoWall check with the operator believing container egress and SNI pinning are enforced.
Cause
cg_origin_egress sits at ~705k processed instructions on the 6.17 Azure kernel CI verifies against (see the notes in [[design.md]] / the L7 SNI work). The 6.6 verifier explores the same program past the 1,000,000-instruction limit — older verifiers prune fewer states in the SNI/QUIC byte loops. The Lima gate is pinned to 6.17 precisely because 6.8 was more forgiving than CI; 6.6 turns out to be less forgiving, and nothing in CI measures it.
6.6 is a current LTS and the kernel on Blacksmith today; GitLab SaaS shared runners are on 5.15, which will be worse.
Ask
- Add a verifier-budget gate for the oldest kernels we claim to support (at minimum 6.6, ideally 5.15 for the GitLab preset) — a matrix job with a 6.6 kernel, or a
vmtest-style run — and record the processed-instruction count per program so regressions are visible before release.
- Bring
cg_origin_egress under budget on 6.6: split the SNI/QUIC parsing into a tail call or a separate program, bound the byte loops more tightly, or gate the L7 parse behind a rodata flag so an L4-only build of the program loads where the L7 one cannot (the L4 container-egress enforcement would then survive on 6.6 even if SNI does not).
- Make the degradation louder than a
::warning::. When the operator asked for container-egress: enforce or tls-sni: enforce-pinned and the hook cannot load, the job summary and the dashboard should say those postures were not applied; today it is only visible in the step log.
Related
Symptom
On Blacksmith runners (Firecracker VM, kernel 6.6.141, BTF present, TCX attach works) the cgroup egress program does not load, so
--container-egress enforceand--tls-sni enforce-pinnedare both silently off for the job. From a v2.0.0-rc.9 run onblacksmith-2vcpu-ubuntu-2404(madhokie-io/dockcmd, run 35283540609):Everything else in the run is healthy: step attribution works, TC enforcement works, DNS gating works. The job passes its CargoWall check with the operator believing container egress and SNI pinning are enforced.
Cause
cg_origin_egresssits at ~705k processed instructions on the 6.17 Azure kernel CI verifies against (see the notes in [[design.md]] / the L7 SNI work). The 6.6 verifier explores the same program past the 1,000,000-instruction limit — older verifiers prune fewer states in the SNI/QUIC byte loops. The Lima gate is pinned to 6.17 precisely because 6.8 was more forgiving than CI; 6.6 turns out to be less forgiving, and nothing in CI measures it.6.6 is a current LTS and the kernel on Blacksmith today; GitLab SaaS shared runners are on 5.15, which will be worse.
Ask
vmtest-style run — and record the processed-instruction count per program so regressions are visible before release.cg_origin_egressunder budget on 6.6: split the SNI/QUIC parsing into a tail call or a separate program, bound the byte loops more tightly, or gate the L7 parse behind a rodata flag so an L4-only build of the program loads where the L7 one cannot (the L4 container-egress enforcement would then survive on 6.6 even if SNI does not).::warning::. When the operator asked forcontainer-egress: enforceortls-sni: enforce-pinnedand the hook cannot load, the job summary and the dashboard should say those postures were not applied; today it is only visible in the step log.Related