From 9259b9e0c86b85a2b476a350a87b48d685db3355 Mon Sep 17 00:00:00 2001 From: Matthew DeVenny Date: Tue, 15 Sep 2026 13:44:44 -0700 Subject: [PATCH 1/8] #129 enforce the relayed synthetic answer like any upstream answer #128 relayed systemd-resolved's synthetic names to the stub as resolution only: the answer fed neither hostname tracking nor the firewall, so a runner whose host agent sits at the default gateway showed a block on a bare IP and could only be allowed by a provider CIDR. The relayed answer now goes through enforceDNSResponse exactly as an upstream one, still uncached: an allow rule naming _gateway opens the gateway's address on its ports, a deny closes it, and with no rule the connection is attributed to _gateway rather than a bare IP. The agent already accepted underscore rule values; the companion change in app admits them at the control plane. Verified live in Lima under enforce: without a rule, curl to the gateway logs "Connection blocked dst=_gateway"; with the rule, "DNS resolution: added to firewall hostname=_gateway action=allow" and "Connection allowed dst=_gateway". Signed-off-by: Matthew DeVenny --- README.md | 2 ++ design.md | 2 +- pkg/dns/synthetic.go | 9 ++++-- pkg/dns/synthetic_test.go | 66 ++++++++++++++++++++++++++++----------- 4 files changed, 58 insertions(+), 21 deletions(-) diff --git a/README.md b/README.md index 81bb28e..3d27826 100644 --- a/README.md +++ b/README.md @@ -104,6 +104,8 @@ allowed-hosts: | **.storage.azure.com ``` +Underscore labels are valid rule values. `_gateway` names whatever systemd-resolved resolves it to — the runner's default gateway — so a host-side agent reached at the gateway (Firecracker/gvproxy runners such as Blacksmith put theirs at `192.168.127.1`) can be allowed by name rather than by a provider-specific CIDR. Such a rule trusts the routing table the way any hostname rule trusts DNS: scope it to the repository or workflow that needs it, not the organization. + **Search domains** whitelist whole DNS suffixes for resolution without per-hostname tracking. Typical use case: you've allowed a VPC CIDR for internal traffic and want DNS resolution to work for any name under that VPC's internal suffix. ```yaml diff --git a/design.md b/design.md index cca40fb..00bdef7 100644 --- a/design.md +++ b/design.md @@ -698,7 +698,7 @@ documented residuals. - Scoped to dport 53: a full table flush would churn unrelated NAT state (Docker MASQUERADE bindings, established flows). Collateral within scope is deliberate: an in-flight UDP transaction costs one retry, while an established DNAT'd DNS-over-TCP stream is reset — its later packets miss the NAT verdict — which is the point at both call sites, since such a stream is either bypassing the proxy or aimed at a dead one. Loopback-destination entries (the DNAT'd `127.0.0.53` stub flows, direct `127.0.0.1` proxy flows) cost at most the same one-retry/reset when deleted: an in-flight reply recreates the entry — possibly reversed, which on `lo` matters not at all, since a reversed 127.x entry has loopback addresses on both sides and can never match a later external-resolver query. They stay in scope to keep the predicate simple, and the live test targets loopback for exactly that harmlessness. The proxy's own marked upstream flows re-match the mark RETURN rules on recreation - Reverse-direction guard: those costs only hold if the client speaks first. If the first packet after a delete comes from upstream (a reply in flight across the flush, a server-side segment on a DNS-over-TCP stream), conntrack re-creates the flow with upstream as ORIGINAL, stamps a null NAT binding (no nat rules face inbound), and every later client packet rides the reply direction — which never traverses nat OUTPUT — with an original tuple (dport = client's ephemeral port) no dport-53 flush can select. For a socket-reusing resolver (c-ares/Node holds one UDP socket per server) that is a persistent, invisible bypass. The redirect therefore installs `INPUT -p udp/tcp ! -i lo --sport 53 -m conntrack --ctstate NEW -j DROP` alongside the DNAT: dropping the packet destroys its unconfirmed entry, so the client's next packet is NEW forward and takes the DNAT — the already-budgeted retry/reset, now guaranteed regardless of which side speaks first. Legitimate replies ride ESTABLISHED entries and never match. Loopback is exempt (`! -i lo`): a reversed 127.x entry has loopback addresses on both sides, so it can only ever match loopback tuples — it can never capture a later external query, even though stub-destined `lo` traffic is now DNAT'd — and without the exemption a stub or proxy lookup in flight across the install flush would lose its reply and sit out the resolver timeout (5s for glibc) - Best-effort with a Warn at both call sites: idle entries age out in 30–120s, but an actively-used flow (an open DNS-over-TCP stream) refreshes its entry indefinitely — exactly the flow the flush exists to kill, which is why a failed flush is warned loudly rather than ignored - - Synthetic names (#126): systemd-resolved answers `_gateway`, `_outbound`, `_localdnsstub`, `_localdnsproxy` and the RFC 6761 localhost family (`localhost`, `localhost.localdomain`, anything beneath either) from local state and never sends them upstream. With the stub DNAT'd and the proxy's upstream deliberately the resolver *behind* resolved, they would NXDOMAIN for the whole run — REFUSED under enforce, silently under audit, where no connection is ever attempted and nothing is logged (`/etc/hosts` covers only the exact name `localhost`). The proxy relays them to the stub on its marked client socket (`pkg/dns/synthetic.go`; the mark RETURN rules pass it through the stub DNAT, as they do the startup cache peek) and writes resolved's own answer back — uncached, unenforced, never fed to hostname tracking or the firewall, so reaching the gateway stays an explicit CIDR rule. Ahead of the filter gate; host listeners and IN class only (a container's native path never resolved them). A stub that does not answer — resolved stopped, or never installed — sends the query down the ordinary path; `/run/systemd/resolve` outlives a stopped resolved (`RuntimeDirectoryPreserve=yes`), so directory presence is no probe, and the connection-refused round trip on loopback is the cheapest true one. The search-expanded form a stub resolver tries first (`_gateway.`, for a suffix on the host's own `resolv.conf` search list) is answered NXDOMAIN locally — what the upstream says natively — rather than REFUSED, which c-ares treats as terminal on a multi-label attempt and which therefore left `_gateway` unresolvable for c-ares clients under enforce; it is never relayed to the stub, which would send it upstream unmarked and DNAT it back into the proxy. This is specific to the synthetic set, whose expanded form never exists; for a real host the expanded form is the name that resolves (#127) + - Synthetic names (#126): systemd-resolved answers `_gateway`, `_outbound`, `_localdnsstub`, `_localdnsproxy` and the RFC 6761 localhost family (`localhost`, `localhost.localdomain`, anything beneath either) from local state and never sends them upstream. With the stub DNAT'd and the proxy's upstream deliberately the resolver *behind* resolved, they would NXDOMAIN for the whole run — REFUSED under enforce, silently under audit, where no connection is ever attempted and nothing is logged (`/etc/hosts` covers only the exact name `localhost`). The proxy relays them to the stub on its marked client socket (`pkg/dns/synthetic.go`; the mark RETURN rules pass it through the stub DNAT, as they do the startup cache peek) and writes resolved's own answer back — uncached, and enforced like any upstream answer (#129): a hostname rule naming `_gateway` opens the gateway's address on its ports, a deny rule closes it, and with no rule the connection is attributed to `_gateway` rather than a bare IP. Reaching the gateway stays an explicit rule; it is a hostname rule, portable across runner providers, rather than a provider CIDR. Such a rule trusts the routing table the way a hostname rule trusts DNS. Ahead of the filter gate; host listeners and IN class only (a container's native path never resolved them). A stub that does not answer — resolved stopped, or never installed — sends the query down the ordinary path; `/run/systemd/resolve` outlives a stopped resolved (`RuntimeDirectoryPreserve=yes`), so directory presence is no probe, and the connection-refused round trip on loopback is the cheapest true one. The search-expanded form a stub resolver tries first (`_gateway.`, for a suffix on the host's own `resolv.conf` search list) is answered NXDOMAIN locally — what the upstream says natively — rather than REFUSED, which c-ares treats as terminal on a multi-label attempt and which therefore left `_gateway` unresolvable for c-ares clients under enforce; it is never relayed to the stub, which would send it upstream unmarked and DNAT it back into the proxy. This is specific to the synthetic set, whose expanded form never exists; for a real host the expanded form is the name that resolves (#127) - **Sudo lockdown:** writes `/etc/sudoers.d/zz-cargowall-lockdown` with a NOPASSWD allowlist; removes the runner user from sudo-granting groups (`sudo`, `admin`, `wheel`) and the `docker` group; disables competing sudoers.d files by renaming them to `*.cargowall-disabled` - **Auto-infrastructure:** `EnsureInfraAllowed()` and `EnsureHostnameAllowed()` add rules for platform services (Azure IMDS, GitHub API, etc.) - **Logging:** `slog.Handler` that formats messages as GitHub workflow commands (`::error::`, `::warning::`, `::debug::`) diff --git a/pkg/dns/synthetic.go b/pkg/dns/synthetic.go index 7b76320..5a14c74 100644 --- a/pkg/dns/synthetic.go +++ b/pkg/dns/synthetic.go @@ -26,8 +26,10 @@ import ( // Names systemd-resolved synthesizes locally (#126): the underscore set, and // the RFC 6761 localhost family (localhost, localhost.localdomain, anything // beneath either). The proxy relays them to the stub on its marked client -// and writes resolved's answer back uncached and unenforced: it never feeds -// hostnameIPs or the firewall. +// and writes resolved's answer back, uncached but enforced like any upstream +// answer (#129): a rule naming "_gateway" opens the gateway's address on +// its ports, a deny closes it, and with no rule the connection is +// attributed to the name rather than a bare IP. var ( syntheticNames = []string{"_gateway", "_outbound", "_localdnsstub", "_localdnsproxy"} localhostRoots = []string{"localhost", "localhost.localdomain"} @@ -68,6 +70,9 @@ func (s *Server) serveSynthetic(w dns.ResponseWriter, r *dns.Msg) bool { return false } resp.Id = r.Id + if resp.Rcode == dns.RcodeSuccess { + s.enforceDNSResponse(name, resp, 0) + } w.WriteMsg(resp) return true } diff --git a/pkg/dns/synthetic_test.go b/pkg/dns/synthetic_test.go index 02c53f2..e28833e 100644 --- a/pkg/dns/synthetic_test.go +++ b/pkg/dns/synthetic_test.go @@ -106,17 +106,24 @@ func withFakeStub(t *testing.T) func() []dns.Question { // syntheticServer is a filtering, default-deny server — the configuration // under which these names would otherwise be REFUSED, so every answer below -// also proves precedence over the gate — with a fake stub and no host -// search list. The mock firewall carries no expectations: any enforcement -// side effect fails the test. -func syntheticServer(t *testing.T) (*Server, func() []dns.Question) { +// also proves precedence over the gate — with a fake stub, no host search +// list, and the given rules. The mock firewall carries no expectations +// unless a test adds them: with no rule naming a synthetic name, any +// firewall call fails the test. +func syntheticServer(t *testing.T, rules ...config.Rule) (*Server, func() []dns.Question, *firewall.MockFirewall) { + t.Helper() + return syntheticServerWithDefault(t, config.ActionDeny, rules...) +} + +func syntheticServerWithDefault(t *testing.T, defaultAction config.Action, rules ...config.Rule) (*Server, func() []dns.Question, *firewall.MockFirewall) { t.Helper() seen := withFakeStub(t) cfg := config.NewConfigManager() - require.NoError(t, cfg.LoadConfigFromRules(nil, config.ActionDeny)) - s := newTestServer(t, cfg, firewall.NewMockFirewall(t)) + require.NoError(t, cfg.LoadConfigFromRules(rules, defaultAction)) + mockFw := firewall.NewMockFirewall(t) + s := newTestServer(t, cfg, mockFw) s.filterQueries = true - return s, seen + return s, seen, mockFw } func ask(t *testing.T, s *Server, w *MockResponseWriter, q *dns.Msg) *dns.Msg { @@ -175,10 +182,11 @@ func TestSyntheticQuery(t *testing.T) { } // The bare name is relayed to the stub and resolved's answer written back -// as-is, ahead of a gate that would otherwise REFUSE it — and nothing is -// cached, tracked or enforced. +// as-is, ahead of a gate that would otherwise REFUSE it. Nothing is cached; +// the answer IS tracked, so a later connection to the address is attributed +// to the name — and with no rule naming it, the firewall is not touched. func TestServeSynthetic_RelaysBareNameToStub(t *testing.T) { - s, seen := syntheticServer(t) + s, seen, _ := syntheticServer(t) m := askName(t, s, "_gateway.", dns.TypeA) assert.Equal(t, dns.RcodeSuccess, m.Rcode) @@ -194,16 +202,38 @@ func TestServeSynthetic_RelaysBareNameToStub(t *testing.T) { _, cached := s.dnsCache.Get(s.generateCacheKey(&dns.Msg{Question: []dns.Question{{Name: "_gateway.", Qtype: dns.TypeA, Qclass: dns.ClassINET}}})) assert.False(t, cached, "stub answers are not cached") + assert.Equal(t, "_gateway", s.config.LookupHostnameByIP("192.168.5.2"), "the address is attributed to the name") s.hostnameIPsMutex.RLock() defer s.hostnameIPsMutex.RUnlock() - assert.Empty(t, s.hostnameIPs, "stub answers are not tracked") + assert.Contains(t, s.hostnameIPs, "_gateway") +} + +// A rule can name a synthetic name (#129): allow opens the gateway's address +// on the rule's ports, deny closes it — the same enforcement any upstream +// answer gets, so "_gateway" is a portable rule where a provider CIDR was. +func TestServeSynthetic_RuleNamesTheGateway(t *testing.T) { + ports := []config.Port{{Port: 1041, Protocol: config.ProtocolTCP}} + s, _, mockFw := syntheticServer(t, config.Rule{Type: config.RuleTypeHostname, Value: "_gateway", Ports: ports, Action: config.ActionAllow}) + mockFw.On("AddIP", mock.MatchedBy(func(ip net.IP) bool { return ip.Equal(net.ParseIP("192.168.5.2")) }), config.ActionAllow, ports). + Return(true, nil).Once() + m := askName(t, s, "_gateway.", dns.TypeA) + assert.Equal(t, dns.RcodeSuccess, m.Rcode) + mockFw.AssertExpectations(t) + + // A deny only needs a BPF entry when the default would otherwise allow. + s, _, mockFw = syntheticServerWithDefault(t, config.ActionAllow, config.Rule{Type: config.RuleTypeHostname, Value: "_gateway", Action: config.ActionDeny}) + mockFw.On("AddIP", mock.MatchedBy(func(ip net.IP) bool { return ip.Equal(net.ParseIP("192.168.5.2")) }), config.ActionDeny, []config.Port(nil)). + Return(true, nil).Once() + m = askName(t, s, "_gateway.", dns.TypeA) + assert.Equal(t, dns.RcodeSuccess, m.Rcode, "a deny rule still resolves the name; it closes the address") + mockFw.AssertExpectations(t) } // Whatever resolved says is what the client gets: NODATA for a family it // has no address for, NXDOMAIN for a name it does not synthesize. The proxy // shapes nothing itself. func TestServeSynthetic_RelaysStubVerdictVerbatim(t *testing.T) { - s, _ := syntheticServer(t) + s, _, _ := syntheticServer(t) m := askName(t, s, "_gateway.", dns.TypeAAAA) assert.Equal(t, dns.RcodeSuccess, m.Rcode) @@ -217,7 +247,7 @@ func TestServeSynthetic_RelaysStubVerdictVerbatim(t *testing.T) { // the name down the ordinary path, REFUSED here, rather than SERVFAIL: the // runtime directory outlives a stopped resolved, so this is the only probe. func TestServeSynthetic_StubUnreachableIsOrdinary(t *testing.T) { - s, _ := syntheticServer(t) + s, _, _ := syntheticServer(t) resolvedStubAddr = "127.0.0.1:1" // nothing listens; UDP gets ECONNREFUSED at once m := askName(t, s, "_gateway.", dns.TypeA) assert.Equal(t, dns.RcodeRefused, m.Rcode) @@ -225,7 +255,7 @@ func TestServeSynthetic_StubUnreachableIsOrdinary(t *testing.T) { // The localhost family resolved synthesizes beyond what /etc/hosts carries. func TestServeSynthetic_RelaysLocalhostFamily(t *testing.T) { - s, seen := syntheticServer(t) + s, seen, _ := syntheticServer(t) for _, q := range []string{"api.localhost.", "localhost.localdomain.", "foo.localhost.localdomain."} { m := askName(t, s, q, dns.TypeA) assert.Equal(t, dns.RcodeSuccess, m.Rcode, q) @@ -239,7 +269,7 @@ func TestServeSynthetic_RelaysLocalhostFamily(t *testing.T) { // bare name: NXDOMAIN, not REFUSED — the rcode every client, c-ares // included, treats as "try the next form" — and never relayed to the stub. func TestServeSynthetic_ExpandedIsNXDomain(t *testing.T) { - s, seen := syntheticServer(t) + s, seen, _ := syntheticServer(t) s.config.SetHostSearchDomains([]string{"lan", "vm.blacksmith.sh"}, s.logger) for _, q := range []string{"_gateway.lan.", "_gateway.vm.blacksmith.sh.", "_outbound.lan."} { @@ -255,7 +285,7 @@ func TestServeSynthetic_ExpandedIsNXDomain(t *testing.T) { // search list, or no search list at all, leaves the query on the ordinary // path — REFUSED here — rather than the proxy inventing a negative answer. func TestServeSynthetic_ExpandedOtherwiseOrdinary(t *testing.T) { - s, _ := syntheticServer(t) + s, _, _ := syntheticServer(t) m := askName(t, s, "_gateway.lan.", dns.TypeA) assert.Equal(t, dns.RcodeRefused, m.Rcode, "no search list") @@ -266,7 +296,7 @@ func TestServeSynthetic_ExpandedOtherwiseOrdinary(t *testing.T) { } func TestServeSynthetic_NonINClassIsOrdinary(t *testing.T) { - s, seen := syntheticServer(t) + s, seen, _ := syntheticServer(t) q := new(dns.Msg) q.SetQuestion("_gateway.", dns.TypeA) q.Question[0].Qclass = dns.ClassCHAOS @@ -286,7 +316,7 @@ func (c *containerListenerWriter) LocalAddr() net.Addr { // gateway is the wrong answer in a container netns: container listeners // fall through to the ordinary path. func TestServeSynthetic_ContainerListenerIsOrdinary(t *testing.T) { - s, seen := syntheticServer(t) + s, seen, _ := syntheticServer(t) s.listenerModes = map[string]listenerAttribution{"172.17.0.1": attributeContainerIP} q := new(dns.Msg) From 1604e8b9a2842ae8886731e77536aa960fb10b06 Mon Sep 17 00:00:00 2001 From: Matthew DeVenny Date: Tue, 15 Sep 2026 13:53:12 -0700 Subject: [PATCH 2/8] #129 enforce only the underscore set; the localhost family stays untracked MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Review on #131: syntheticQuery accepts any label under .localhost, and hostnameIPs has no eviction, so routing every relayed answer through enforcement let a job grow tracking without bound by issuing unique localhost aliases — no allow rule needed. Enforcement now applies to the four underscore names only, a fixed set. The localhost family is relayed as in #128 and never tracked: loopback is auto-allowed, and attributing 127.0.0.1 to an alias buys nothing. Signed-off-by: Matthew DeVenny --- design.md | 2 +- pkg/dns/synthetic.go | 12 +++++++----- pkg/dns/synthetic_test.go | 10 ++++++++-- 3 files changed, 16 insertions(+), 8 deletions(-) diff --git a/design.md b/design.md index 00bdef7..379c622 100644 --- a/design.md +++ b/design.md @@ -698,7 +698,7 @@ documented residuals. - Scoped to dport 53: a full table flush would churn unrelated NAT state (Docker MASQUERADE bindings, established flows). Collateral within scope is deliberate: an in-flight UDP transaction costs one retry, while an established DNAT'd DNS-over-TCP stream is reset — its later packets miss the NAT verdict — which is the point at both call sites, since such a stream is either bypassing the proxy or aimed at a dead one. Loopback-destination entries (the DNAT'd `127.0.0.53` stub flows, direct `127.0.0.1` proxy flows) cost at most the same one-retry/reset when deleted: an in-flight reply recreates the entry — possibly reversed, which on `lo` matters not at all, since a reversed 127.x entry has loopback addresses on both sides and can never match a later external-resolver query. They stay in scope to keep the predicate simple, and the live test targets loopback for exactly that harmlessness. The proxy's own marked upstream flows re-match the mark RETURN rules on recreation - Reverse-direction guard: those costs only hold if the client speaks first. If the first packet after a delete comes from upstream (a reply in flight across the flush, a server-side segment on a DNS-over-TCP stream), conntrack re-creates the flow with upstream as ORIGINAL, stamps a null NAT binding (no nat rules face inbound), and every later client packet rides the reply direction — which never traverses nat OUTPUT — with an original tuple (dport = client's ephemeral port) no dport-53 flush can select. For a socket-reusing resolver (c-ares/Node holds one UDP socket per server) that is a persistent, invisible bypass. The redirect therefore installs `INPUT -p udp/tcp ! -i lo --sport 53 -m conntrack --ctstate NEW -j DROP` alongside the DNAT: dropping the packet destroys its unconfirmed entry, so the client's next packet is NEW forward and takes the DNAT — the already-budgeted retry/reset, now guaranteed regardless of which side speaks first. Legitimate replies ride ESTABLISHED entries and never match. Loopback is exempt (`! -i lo`): a reversed 127.x entry has loopback addresses on both sides, so it can only ever match loopback tuples — it can never capture a later external query, even though stub-destined `lo` traffic is now DNAT'd — and without the exemption a stub or proxy lookup in flight across the install flush would lose its reply and sit out the resolver timeout (5s for glibc) - Best-effort with a Warn at both call sites: idle entries age out in 30–120s, but an actively-used flow (an open DNS-over-TCP stream) refreshes its entry indefinitely — exactly the flow the flush exists to kill, which is why a failed flush is warned loudly rather than ignored - - Synthetic names (#126): systemd-resolved answers `_gateway`, `_outbound`, `_localdnsstub`, `_localdnsproxy` and the RFC 6761 localhost family (`localhost`, `localhost.localdomain`, anything beneath either) from local state and never sends them upstream. With the stub DNAT'd and the proxy's upstream deliberately the resolver *behind* resolved, they would NXDOMAIN for the whole run — REFUSED under enforce, silently under audit, where no connection is ever attempted and nothing is logged (`/etc/hosts` covers only the exact name `localhost`). The proxy relays them to the stub on its marked client socket (`pkg/dns/synthetic.go`; the mark RETURN rules pass it through the stub DNAT, as they do the startup cache peek) and writes resolved's own answer back — uncached, and enforced like any upstream answer (#129): a hostname rule naming `_gateway` opens the gateway's address on its ports, a deny rule closes it, and with no rule the connection is attributed to `_gateway` rather than a bare IP. Reaching the gateway stays an explicit rule; it is a hostname rule, portable across runner providers, rather than a provider CIDR. Such a rule trusts the routing table the way a hostname rule trusts DNS. Ahead of the filter gate; host listeners and IN class only (a container's native path never resolved them). A stub that does not answer — resolved stopped, or never installed — sends the query down the ordinary path; `/run/systemd/resolve` outlives a stopped resolved (`RuntimeDirectoryPreserve=yes`), so directory presence is no probe, and the connection-refused round trip on loopback is the cheapest true one. The search-expanded form a stub resolver tries first (`_gateway.`, for a suffix on the host's own `resolv.conf` search list) is answered NXDOMAIN locally — what the upstream says natively — rather than REFUSED, which c-ares treats as terminal on a multi-label attempt and which therefore left `_gateway` unresolvable for c-ares clients under enforce; it is never relayed to the stub, which would send it upstream unmarked and DNAT it back into the proxy. This is specific to the synthetic set, whose expanded form never exists; for a real host the expanded form is the name that resolves (#127) + - Synthetic names (#126): systemd-resolved answers `_gateway`, `_outbound`, `_localdnsstub`, `_localdnsproxy` and the RFC 6761 localhost family (`localhost`, `localhost.localdomain`, anything beneath either) from local state and never sends them upstream. With the stub DNAT'd and the proxy's upstream deliberately the resolver *behind* resolved, they would NXDOMAIN for the whole run — REFUSED under enforce, silently under audit, where no connection is ever attempted and nothing is logged (`/etc/hosts` covers only the exact name `localhost`). The proxy relays them to the stub on its marked client socket (`pkg/dns/synthetic.go`; the mark RETURN rules pass it through the stub DNAT, as they do the startup cache peek) and writes resolved's own answer back — uncached. The underscore set is enforced like any upstream answer (#129): a hostname rule naming `_gateway` opens the gateway's address on its ports, a deny rule closes it, and with no rule the connection is attributed to `_gateway` rather than a bare IP. The localhost family is relayed only: loopback is auto-allowed, and tracking arbitrary `*.localhost` aliases would grow the per-host maps without bound. Reaching the gateway stays an explicit rule; it is a hostname rule, portable across runner providers, rather than a provider CIDR. Such a rule trusts the routing table the way a hostname rule trusts DNS. Ahead of the filter gate; host listeners and IN class only (a container's native path never resolved them). A stub that does not answer — resolved stopped, or never installed — sends the query down the ordinary path; `/run/systemd/resolve` outlives a stopped resolved (`RuntimeDirectoryPreserve=yes`), so directory presence is no probe, and the connection-refused round trip on loopback is the cheapest true one. The search-expanded form a stub resolver tries first (`_gateway.`, for a suffix on the host's own `resolv.conf` search list) is answered NXDOMAIN locally — what the upstream says natively — rather than REFUSED, which c-ares treats as terminal on a multi-label attempt and which therefore left `_gateway` unresolvable for c-ares clients under enforce; it is never relayed to the stub, which would send it upstream unmarked and DNAT it back into the proxy. This is specific to the synthetic set, whose expanded form never exists; for a real host the expanded form is the name that resolves (#127) - **Sudo lockdown:** writes `/etc/sudoers.d/zz-cargowall-lockdown` with a NOPASSWD allowlist; removes the runner user from sudo-granting groups (`sudo`, `admin`, `wheel`) and the `docker` group; disables competing sudoers.d files by renaming them to `*.cargowall-disabled` - **Auto-infrastructure:** `EnsureInfraAllowed()` and `EnsureHostnameAllowed()` add rules for platform services (Azure IMDS, GitHub API, etc.) - **Logging:** `slog.Handler` that formats messages as GitHub workflow commands (`::error::`, `::warning::`, `::debug::`) diff --git a/pkg/dns/synthetic.go b/pkg/dns/synthetic.go index 5a14c74..6ac0c39 100644 --- a/pkg/dns/synthetic.go +++ b/pkg/dns/synthetic.go @@ -26,10 +26,12 @@ import ( // Names systemd-resolved synthesizes locally (#126): the underscore set, and // the RFC 6761 localhost family (localhost, localhost.localdomain, anything // beneath either). The proxy relays them to the stub on its marked client -// and writes resolved's answer back, uncached but enforced like any upstream -// answer (#129): a rule naming "_gateway" opens the gateway's address on -// its ports, a deny closes it, and with no rule the connection is -// attributed to the name rather than a bare IP. +// and writes resolved's answer back, uncached. The underscore set is +// enforced like any upstream answer (#129): a rule naming "_gateway" opens +// the gateway's address on its ports, a deny closes it, and with no rule +// the connection is attributed to the name rather than a bare IP. The +// localhost family is relayed only: loopback is auto-allowed, and tracking +// arbitrary "*.localhost" aliases would grow hostnameIPs without bound. var ( syntheticNames = []string{"_gateway", "_outbound", "_localdnsstub", "_localdnsproxy"} localhostRoots = []string{"localhost", "localhost.localdomain"} @@ -70,7 +72,7 @@ func (s *Server) serveSynthetic(w dns.ResponseWriter, r *dns.Msg) bool { return false } resp.Id = r.Id - if resp.Rcode == dns.RcodeSuccess { + if resp.Rcode == dns.RcodeSuccess && slices.Contains(syntheticNames, name) { s.enforceDNSResponse(name, resp, 0) } w.WriteMsg(resp) diff --git a/pkg/dns/synthetic_test.go b/pkg/dns/synthetic_test.go index e28833e..fd6c715 100644 --- a/pkg/dns/synthetic_test.go +++ b/pkg/dns/synthetic_test.go @@ -253,8 +253,10 @@ func TestServeSynthetic_StubUnreachableIsOrdinary(t *testing.T) { assert.Equal(t, dns.RcodeRefused, m.Rcode) } -// The localhost family resolved synthesizes beyond what /etc/hosts carries. -func TestServeSynthetic_RelaysLocalhostFamily(t *testing.T) { +// The localhost family resolved synthesizes beyond what /etc/hosts carries — +// relayed, but not tracked: any label under .localhost is accepted, and +// retaining each alias would grow hostnameIPs without bound. +func TestServeSynthetic_RelaysLocalhostFamilyUntracked(t *testing.T) { s, seen, _ := syntheticServer(t) for _, q := range []string{"api.localhost.", "localhost.localdomain.", "foo.localhost.localdomain."} { m := askName(t, s, q, dns.TypeA) @@ -263,6 +265,10 @@ func TestServeSynthetic_RelaysLocalhostFamily(t *testing.T) { assert.Equal(t, "127.0.0.1", m.Answer[0].(*dns.A).A.String(), q) } assert.Len(t, seen(), 3) + assert.Empty(t, s.config.LookupHostnameByIP("127.0.0.1")) + s.hostnameIPsMutex.RLock() + defer s.hostnameIPsMutex.RUnlock() + assert.Empty(t, s.hostnameIPs) } // The search-expanded form is the first query a resolver sends for the From 1c2d21711a73198a71530b9266d72007027b0748 Mon Sep 17 00:00:00 2001 From: Matthew DeVenny Date: Tue, 15 Sep 2026 14:11:45 -0700 Subject: [PATCH 3/8] #129 relay the machine hostname; mint no L7 evidence for local aliases MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Review on #131, two findings. The machine's own hostname is a synthetic record too — its configured name, its first label, and the mDNS