diff --git a/README.md b/README.md index 81bb28e..c9f54c7 100644 --- a/README.md +++ b/README.md @@ -104,6 +104,8 @@ allowed-hosts: | **.storage.azure.com ``` +Underscore labels are valid rule values. `_gateway` names whatever systemd-resolved resolves it to — the runner's default gateway — so a host-side agent reached at the gateway (Firecracker/gvproxy runners such as Blacksmith put theirs at `192.168.127.1`) can be allowed by name rather than by a provider-specific CIDR. Two things to know: such a rule trusts the routing table the way any hostname rule trusts DNS, so scope it to the repository or workflow that needs it, not the organization; and `_gateway` is an address alias, never a name the peer presents, so the rule is L4-only — it opens the address on its ports like a CIDR allow and pins no L7 identity, so with `--tls-sni` on, HTTP to the agent carrying `Host: 192.168.127.1` is not an L7 miss. It needs systemd-resolved on the runner (the Ubuntu images have it): the name is answered by resolved's stub, so without one it does not resolve and the rule opens nothing — keep a CIDR on such a host. + **Search domains** whitelist whole DNS suffixes for resolution without per-hostname tracking. Typical use case: you've allowed a VPC CIDR for internal traffic and want DNS resolution to work for any name under that VPC's internal suffix. ```yaml diff --git a/cmd/start.go b/cmd/start.go index 6eaf2c7..5a62231 100644 --- a/cmd/start.go +++ b/cmd/start.go @@ -1617,16 +1617,7 @@ func prePopulateDNSCache(ctx context.Context, configMgr *config.Manager, dnsServ for hostname := range configMgr.GetTrackedHostnames() { lookupCtx, lookupCancel := context.WithTimeout(cacheCtx, 2*time.Second) if ips, err := cacheResolver.LookupHost(lookupCtx, hostname); err == nil { - for _, ip := range ips { - configMgr.UpdateDNSMapping(hostname, ip) - // A FORWARD lookup of a rule hostname — the same evidence - // class as the proxy's own answers, not a PTR — and these are - // the IPs live processes are already using. Recording it is - // what lets those IPs be L7-scoped at all (RegisterL7Identity - // scopes iff bound); without it the first flight to a - // systemd-resolved cached IP is name_not_at_ip under pin-ip. - configMgr.RecordForwardResolution(hostname, ip) - } + dnsServer.RecordSystemCacheAnswer(hostname, ips) } else { logger.Debug("System DNS cache miss", "hostname", hostname, "error", err) } diff --git a/design-l7.md b/design-l7.md index ef8bd07..a7739c8 100644 --- a/design-l7.md +++ b/design-l7.md @@ -288,7 +288,13 @@ destination we cannot bind stays L4-governed.** Startup pre-population therefore records its Phase-1 answers as forward resolutions — they are forward lookups of rule names through the system resolver, the same evidence class as the proxy's own answers — rather than leaving IPs live processes already use -un-scopeable. +un-scopeable. The one exception is a name the proxy itself answers from +resolved's local state: pre-population records through +`dns.Server.RecordSystemCacheAnswer`, so the proxy's own policy applies — a tracked +alias (`_gateway`, the machine's own name) is an address no peer presents by that +name, mapped for attribution but minting no evidence and staying L4-governed; a +loopback listener (`_localdnsstub`) is recorded nowhere, so the replay can never +write the stub's own address. **Lifecycle.** The store is count-bounded LRU with refresh-on-use, deliberately not TTL-swept: the L4 and `map_l7_scope` entries it gates never expire, so an diff --git a/design.md b/design.md index cca40fb..c5911a7 100644 --- a/design.md +++ b/design.md @@ -698,7 +698,7 @@ documented residuals. - Scoped to dport 53: a full table flush would churn unrelated NAT state (Docker MASQUERADE bindings, established flows). Collateral within scope is deliberate: an in-flight UDP transaction costs one retry, while an established DNAT'd DNS-over-TCP stream is reset — its later packets miss the NAT verdict — which is the point at both call sites, since such a stream is either bypassing the proxy or aimed at a dead one. Loopback-destination entries (the DNAT'd `127.0.0.53` stub flows, direct `127.0.0.1` proxy flows) cost at most the same one-retry/reset when deleted: an in-flight reply recreates the entry — possibly reversed, which on `lo` matters not at all, since a reversed 127.x entry has loopback addresses on both sides and can never match a later external-resolver query. They stay in scope to keep the predicate simple, and the live test targets loopback for exactly that harmlessness. The proxy's own marked upstream flows re-match the mark RETURN rules on recreation - Reverse-direction guard: those costs only hold if the client speaks first. If the first packet after a delete comes from upstream (a reply in flight across the flush, a server-side segment on a DNS-over-TCP stream), conntrack re-creates the flow with upstream as ORIGINAL, stamps a null NAT binding (no nat rules face inbound), and every later client packet rides the reply direction — which never traverses nat OUTPUT — with an original tuple (dport = client's ephemeral port) no dport-53 flush can select. For a socket-reusing resolver (c-ares/Node holds one UDP socket per server) that is a persistent, invisible bypass. The redirect therefore installs `INPUT -p udp/tcp ! -i lo --sport 53 -m conntrack --ctstate NEW -j DROP` alongside the DNAT: dropping the packet destroys its unconfirmed entry, so the client's next packet is NEW forward and takes the DNAT — the already-budgeted retry/reset, now guaranteed regardless of which side speaks first. Legitimate replies ride ESTABLISHED entries and never match. Loopback is exempt (`! -i lo`): a reversed 127.x entry has loopback addresses on both sides, so it can only ever match loopback tuples — it can never capture a later external query, even though stub-destined `lo` traffic is now DNAT'd — and without the exemption a stub or proxy lookup in flight across the install flush would lose its reply and sit out the resolver timeout (5s for glibc) - Best-effort with a Warn at both call sites: idle entries age out in 30–120s, but an actively-used flow (an open DNS-over-TCP stream) refreshes its entry indefinitely — exactly the flow the flush exists to kill, which is why a failed flush is warned loudly rather than ignored - - Synthetic names (#126): systemd-resolved answers `_gateway`, `_outbound`, `_localdnsstub`, `_localdnsproxy` and the RFC 6761 localhost family (`localhost`, `localhost.localdomain`, anything beneath either) from local state and never sends them upstream. With the stub DNAT'd and the proxy's upstream deliberately the resolver *behind* resolved, they would NXDOMAIN for the whole run — REFUSED under enforce, silently under audit, where no connection is ever attempted and nothing is logged (`/etc/hosts` covers only the exact name `localhost`). The proxy relays them to the stub on its marked client socket (`pkg/dns/synthetic.go`; the mark RETURN rules pass it through the stub DNAT, as they do the startup cache peek) and writes resolved's own answer back — uncached, unenforced, never fed to hostname tracking or the firewall, so reaching the gateway stays an explicit CIDR rule. Ahead of the filter gate; host listeners and IN class only (a container's native path never resolved them). A stub that does not answer — resolved stopped, or never installed — sends the query down the ordinary path; `/run/systemd/resolve` outlives a stopped resolved (`RuntimeDirectoryPreserve=yes`), so directory presence is no probe, and the connection-refused round trip on loopback is the cheapest true one. The search-expanded form a stub resolver tries first (`_gateway.`, for a suffix on the host's own `resolv.conf` search list) is answered NXDOMAIN locally — what the upstream says natively — rather than REFUSED, which c-ares treats as terminal on a multi-label attempt and which therefore left `_gateway` unresolvable for c-ares clients under enforce; it is never relayed to the stub, which would send it upstream unmarked and DNAT it back into the proxy. This is specific to the synthetic set, whose expanded form never exists; for a real host the expanded form is the name that resolves (#127) + - Synthetic names (#126): systemd-resolved answers `_gateway`, `_outbound`, `_localdnsstub`, `_localdnsproxy`, the machine's own hostname (its configured name and, when it carries a domain, its first label — read per lookup, so a rename mid-run is honoured; `/etc/hosts` does not always pin it, and `sudo` warns on every invocation when it fails) and the RFC 6761 localhost family (`localhost`, `localhost.localdomain`, anything beneath either) from local state and never sends them upstream. With the stub DNAT'd and the proxy's upstream deliberately the resolver *behind* resolved, they would NXDOMAIN for the whole run — REFUSED under enforce, silently under audit, where no connection is ever attempted and nothing is logged (`/etc/hosts` covers only the exact name `localhost`). The proxy relays them to the stub on its marked client socket (`pkg/dns/synthetic.go`; the mark RETURN rules pass it through the stub DNAT, as they do the startup cache peek) and writes resolved's own answer back — uncached. `_gateway`, `_outbound` and the machine names get L4 enforcement and attribution like any upstream answer (#129): a hostname rule naming `_gateway` opens the gateway's address on its ports, a deny rule closes it, and with no rule the connection is attributed to `_gateway` rather than a bare IP. They mint no L7 forward-resolution evidence — `enforceDNSResponse(…, localAnswer)` at the synthetic call site, and startup pre-population records its answers through `dns.Server.RecordSystemCacheAnswer`, the same policy: a tracked alias mapped without evidence, a loopback listener recorded nowhere at all — because an alias is never a name the peer presents, and binding it would L7-scope the address against a name no flow can show: an all-ports allow would then drop HTTP to the host agent for carrying `Host: `. SCOPE IFF BOUND makes the omission sufficient; nothing in the L7 matcher knows about synthetic names. `_localdnsstub`, `_localdnsproxy` and the localhost family are relayed only: loopback, which the proxy's own relay and cache peek depend on, and any label under `.localhost` is accepted, so tracking them would grow the per-host maps without bound. The mDNS `.local` form is not relayed: resolved renames it on conflict, and a `.local` query it does not synthesize is multicast to the LAN, whose answer is not local state. Ahead of the filter gate; host listeners and IN class only (a container's native path never resolved them). A stub that does not answer — resolved stopped, or never installed — sends the query down the ordinary path; `/run/systemd/resolve` outlives a stopped resolved (`RuntimeDirectoryPreserve=yes`), so directory presence is no probe, and the connection-refused round trip on loopback is the cheapest true one. The search-expanded form a stub resolver tries first (`.`, for a suffix on the host's own `resolv.conf` search list) is answered NXDOMAIN locally for every one of these names — what the upstream says natively — rather than REFUSED, which c-ares treats as terminal on a multi-label attempt and which therefore left `_gateway` unresolvable for c-ares clients under enforce; it is never relayed to the stub, which would send it upstream unmarked and DNAT it back into the proxy. This is specific to names resolved synthesizes, whose expanded form never exists; for a real host the expanded form is the name that resolves (#127). Residual: a machine hostname that is also a real DNS name resolves to resolved's synthesized local address rather than the upstream record — inherent to making the machine hostname resolve at all, and what the host does natively, since resolved synthesizes its own name before any upstream lookup - **Sudo lockdown:** writes `/etc/sudoers.d/zz-cargowall-lockdown` with a NOPASSWD allowlist; removes the runner user from sudo-granting groups (`sudo`, `admin`, `wheel`) and the `docker` group; disables competing sudoers.d files by renaming them to `*.cargowall-disabled` - **Auto-infrastructure:** `EnsureInfraAllowed()` and `EnsureHostnameAllowed()` add rules for platform services (Azure IMDS, GitHub API, etc.) - **Logging:** `slog.Handler` that formats messages as GitHub workflow commands (`::error::`, `::warning::`, `::debug::`) diff --git a/pkg/dns/hostsearch_test.go b/pkg/dns/hostsearch_test.go index 3ecdfd6..e51ff0c 100644 --- a/pkg/dns/hostsearch_test.go +++ b/pkg/dns/hostsearch_test.go @@ -37,6 +37,7 @@ import ( // opts in through withResolvConf. func TestMain(m *testing.M) { resolvConfPath = filepath.Join(os.TempDir(), "cargowall-dns-test-absent-resolv.conf") + machineHostname = func() (string, error) { return "", os.ErrNotExist } os.Exit(m.Run()) } @@ -121,6 +122,8 @@ func TestHandleDNSQuery_HostSearchExpandedFormResolvesAndEnforces(t *testing.T) s := newTestServer(t, cfg, mockFw) s.upstream = pc.LocalAddr().String() s.filterQueries = true + rec := &recordingRegistrar{} + s.SetL7Registrar(rec) q := new(dns.Msg) q.SetQuestion("myservice.corp.lan.", dns.TypeA) @@ -134,4 +137,10 @@ func TestHandleDNSQuery_HostSearchExpandedFormResolvesAndEnforces(t *testing.T) assert.Equal(t, dns.RcodeSuccess, w.msg.Rcode) require.Len(t, w.msg.Answer, 1) assert.Equal(t, "192.0.2.10", w.msg.Answer[0].(*dns.A).A.String()) + + // A real answer for an allowed name is a wire identity: evidence minted, + // address L7-scoped for the rule's (all) ports — the contrast to the + // synthetic path, which mints none. + assert.True(t, s.config.NameResolvedToIP("myservice.corp.lan", "192.0.2.10")) + assert.Equal(t, "all-ports", rec.scopes["192.0.2.10"]) } diff --git a/pkg/dns/server.go b/pkg/dns/server.go index 1f36e76..b5300c0 100644 --- a/pkg/dns/server.go +++ b/pkg/dns/server.go @@ -645,7 +645,7 @@ func (s *Server) handleDNSQuery(w dns.ResponseWriter, r *dns.Msg) { // DNS-path output consistent with the connection-event path, which logs // the lowercase hostname from the IP->hostname mapping (#65). if len(r.Question) > 0 && resp.Rcode == dns.RcodeSuccess { - s.enforceDNSResponse(strings.ToLower(strings.TrimSuffix(r.Question[0].Name, ".")), resp, 0) + s.enforceDNSResponse(strings.ToLower(strings.TrimSuffix(r.Question[0].Name, ".")), resp, 0, wireAnswer) } // Return response to client @@ -654,14 +654,32 @@ func (s *Server) handleDNSQuery(w dns.ResponseWriter, r *dns.Msg) { } } +// answerSource says where a DNS answer came from, for enforceDNSResponse's +// L7 evidence: only an answer that came off the wire, for a name a peer can +// present, may mint forward-resolution evidence. +type answerSource int + +const ( + // wireAnswer: an upstream answer, or a CNAME pre-resolve of one. + wireAnswer answerSource = iota + // localAnswer: relayed from resolved's own state — an address alias, not + // an identity. Scoping its address against it would make every flow + // there an L7 miss. + localAnswer +) + // enforceDNSResponse applies one successful DNS response to enforcement // state: rule/derived verdict matching, CNAME-chain learning, BPF map updates // for resolved IPs, late-allow reconciliation of previously blocked // connections, and pre-resolution of allowed CNAME-only responses. // canonicalHostname is the queried name, lowercased with the trailing dot // trimmed. depth bounds pre-resolve recursion (see preResolveCNAMETarget); -// handleDNSQuery passes 0. -func (s *Server) enforceDNSResponse(canonicalHostname string, resp *dns.Msg, depth int) { +// handleDNSQuery passes 0. source says whether the answer came off the wire +// (wireAnswer: upstream or pre-resolve, for a name a peer can present) and +// so may mint L7 forward-resolution evidence, or from resolved's local state +// (localAnswer: an address alias no peer presents) and mints none, so SCOPE +// IFF BOUND scopes nothing. +func (s *Server) enforceDNSResponse(canonicalHostname string, resp *dns.Msg, depth int, source answerSource) { // Extract IPs and TTLs from response ips, ttl := s.extractIPsFromResponse(resp) @@ -833,8 +851,11 @@ func (s *Server) enforceDNSResponse(canonicalHostname string, resp *dns.Msg, dep // This is THE forward-resolution path — a real DNS answer // traversing the proxy — so it (and RecordCNAMEChain below) // are the only seeds of the L7 per-IP binding evidence. - // Reverse-DNS paths must never record it (PTR forgery). - s.config.RecordForwardResolution(canonicalHostname, ip.String()) + // Reverse-DNS paths must never record it (PTR forgery), and + // an answer that did not come off the wire mints none. + if source == wireAnswer { + s.config.RecordForwardResolution(canonicalHostname, ip.String()) + } } // Track the IPs we've seen for this hostname. Accumulate @@ -1294,7 +1315,7 @@ func (s *Server) preResolveCNAMETarget(target string, depth int) { if resp.Rcode != dns.RcodeSuccess { continue } - s.enforceDNSResponse(target, resp, depth) + s.enforceDNSResponse(target, resp, depth, wireAnswer) } }() } diff --git a/pkg/dns/synthetic.go b/pkg/dns/synthetic.go index 7b76320..908c652 100644 --- a/pkg/dns/synthetic.go +++ b/pkg/dns/synthetic.go @@ -17,24 +17,122 @@ package dns import ( + "os" "slices" "strings" "github.com/miekg/dns" ) -// Names systemd-resolved synthesizes locally (#126): the underscore set, and -// the RFC 6761 localhost family (localhost, localhost.localdomain, anything -// beneath either). The proxy relays them to the stub on its marked client -// and writes resolved's answer back uncached and unenforced: it never feeds -// hostnameIPs or the firewall. +// aliasClass is what kind of name the proxy has in hand: an ordinary one, or +// one of the kinds it answers from resolved's local state (#126, #129). +type aliasClass int + +const ( + // notAlias: the ordinary path. + notAlias aliasClass = iota + // trackedAlias: _gateway, _outbound and the machine's own names — + // relayed, L4-enforced and attributed like an upstream answer, minting + // no L7 evidence. A real routed address a rule may name. + trackedAlias + // untrackedAlias: the stub's loopback listeners and the localhost + // family — relayed, recorded nowhere. Loopback the relay itself depends + // on, and any label under .localhost is accepted, so tracking would + // grow hostnameIPs without bound. + untrackedAlias + // expandedAlias: ., the form a stub resolver tries before + // the name itself — NXDOMAIN, recorded nowhere. Never relayed: the stub + // would query upstream unmarked and the DNAT would bring it back here. + expandedAlias +) + var ( - syntheticNames = []string{"_gateway", "_outbound", "_localdnsstub", "_localdnsproxy"} + enforcedNames = []string{"_gateway", "_outbound"} + relayOnlyNames = []string{"_localdnsstub", "_localdnsproxy"} localhostRoots = []string{"localhost", "localhost.localdomain"} ) -// resolvedStubAddr is a var so tests can point the relay at a fake stub. -var resolvedStubAddr = "127.0.0.53:53" +// Injection points for tests: the stub address and the machine hostname. +var ( + resolvedStubAddr = "127.0.0.53:53" + machineHostname = os.Hostname +) + +// machineHostnames returns the forms resolved synthesizes for the machine's +// own hostname: the configured name and, when it carries a domain, its +// first label (else ""). Read per lookup — gethostname is a trivial syscall +// — so a hostnamectl set-hostname mid-run is honoured. The mDNS +// "