Skip to content
Merged
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -104,6 +104,8 @@ allowed-hosts: |
**.storage.azure.com
```

Underscore labels are valid rule values. `_gateway` names whatever systemd-resolved resolves it to — the runner's default gateway — so a host-side agent reached at the gateway (Firecracker/gvproxy runners such as Blacksmith put theirs at `192.168.127.1`) can be allowed by name rather than by a provider-specific CIDR. Two things to know: such a rule trusts the routing table the way any hostname rule trusts DNS, so scope it to the repository or workflow that needs it, not the organization; and `_gateway` is an address alias, never a name the peer presents, so the rule is L4-only — it opens the address on its ports like a CIDR allow and pins no L7 identity, so with `--tls-sni` on, HTTP to the agent carrying `Host: 192.168.127.1` is not an L7 miss. It needs systemd-resolved on the runner (the Ubuntu images have it): the name is answered by resolved's stub, so without one it does not resolve and the rule opens nothing — keep a CIDR on such a host.

**Search domains** whitelist whole DNS suffixes for resolution without per-hostname tracking. Typical use case: you've allowed a VPC CIDR for internal traffic and want DNS resolution to work for any name under that VPC's internal suffix.

```yaml
Expand Down
11 changes: 1 addition & 10 deletions cmd/start.go
Original file line number Diff line number Diff line change
Expand Up @@ -1617,16 +1617,7 @@ func prePopulateDNSCache(ctx context.Context, configMgr *config.Manager, dnsServ
for hostname := range configMgr.GetTrackedHostnames() {
lookupCtx, lookupCancel := context.WithTimeout(cacheCtx, 2*time.Second)
if ips, err := cacheResolver.LookupHost(lookupCtx, hostname); err == nil {
for _, ip := range ips {
configMgr.UpdateDNSMapping(hostname, ip)
// A FORWARD lookup of a rule hostname — the same evidence
// class as the proxy's own answers, not a PTR — and these are
// the IPs live processes are already using. Recording it is
// what lets those IPs be L7-scoped at all (RegisterL7Identity
// scopes iff bound); without it the first flight to a
// systemd-resolved cached IP is name_not_at_ip under pin-ip.
configMgr.RecordForwardResolution(hostname, ip)
}
dnsServer.RecordSystemCacheAnswer(hostname, ips)
} else {
logger.Debug("System DNS cache miss", "hostname", hostname, "error", err)
}
Expand Down
8 changes: 7 additions & 1 deletion design-l7.md
Original file line number Diff line number Diff line change
Expand Up @@ -288,7 +288,13 @@ destination we cannot bind stays L4-governed.** Startup pre-population
therefore records its Phase-1 answers as forward resolutions — they are forward
lookups of rule names through the system resolver, the same evidence class as
the proxy's own answers — rather than leaving IPs live processes already use
un-scopeable.
un-scopeable. The one exception is a name the proxy itself answers from
resolved's local state: pre-population records through
`dns.Server.RecordSystemCacheAnswer`, so the proxy's own policy applies — a tracked
alias (`_gateway`, the machine's own name) is an address no peer presents by that
name, mapped for attribution but minting no evidence and staying L4-governed; a
loopback listener (`_localdnsstub`) is recorded nowhere, so the replay can never
write the stub's own address.

**Lifecycle.** The store is count-bounded LRU with refresh-on-use, deliberately
not TTL-swept: the L4 and `map_l7_scope` entries it gates never expire, so an
Expand Down
2 changes: 1 addition & 1 deletion design.md
Original file line number Diff line number Diff line change
Expand Up @@ -698,7 +698,7 @@ documented residuals.
- Scoped to dport 53: a full table flush would churn unrelated NAT state (Docker MASQUERADE bindings, established flows). Collateral within scope is deliberate: an in-flight UDP transaction costs one retry, while an established DNAT'd DNS-over-TCP stream is reset — its later packets miss the NAT verdict — which is the point at both call sites, since such a stream is either bypassing the proxy or aimed at a dead one. Loopback-destination entries (the DNAT'd `127.0.0.53` stub flows, direct `127.0.0.1` proxy flows) cost at most the same one-retry/reset when deleted: an in-flight reply recreates the entry — possibly reversed, which on `lo` matters not at all, since a reversed 127.x entry has loopback addresses on both sides and can never match a later external-resolver query. They stay in scope to keep the predicate simple, and the live test targets loopback for exactly that harmlessness. The proxy's own marked upstream flows re-match the mark RETURN rules on recreation
- Reverse-direction guard: those costs only hold if the client speaks first. If the first packet after a delete comes from upstream (a reply in flight across the flush, a server-side segment on a DNS-over-TCP stream), conntrack re-creates the flow with upstream as ORIGINAL, stamps a null NAT binding (no nat rules face inbound), and every later client packet rides the reply direction — which never traverses nat OUTPUT — with an original tuple (dport = client's ephemeral port) no dport-53 flush can select. For a socket-reusing resolver (c-ares/Node holds one UDP socket per server) that is a persistent, invisible bypass. The redirect therefore installs `INPUT -p udp/tcp ! -i lo --sport 53 -m conntrack --ctstate NEW -j DROP` alongside the DNAT: dropping the packet destroys its unconfirmed entry, so the client's next packet is NEW forward and takes the DNAT — the already-budgeted retry/reset, now guaranteed regardless of which side speaks first. Legitimate replies ride ESTABLISHED entries and never match. Loopback is exempt (`! -i lo`): a reversed 127.x entry has loopback addresses on both sides, so it can only ever match loopback tuples — it can never capture a later external query, even though stub-destined `lo` traffic is now DNAT'd — and without the exemption a stub or proxy lookup in flight across the install flush would lose its reply and sit out the resolver timeout (5s for glibc)
- Best-effort with a Warn at both call sites: idle entries age out in 30–120s, but an actively-used flow (an open DNS-over-TCP stream) refreshes its entry indefinitely — exactly the flow the flush exists to kill, which is why a failed flush is warned loudly rather than ignored
- Synthetic names (#126): systemd-resolved answers `_gateway`, `_outbound`, `_localdnsstub`, `_localdnsproxy` and the RFC 6761 localhost family (`localhost`, `localhost.localdomain`, anything beneath either) from local state and never sends them upstream. With the stub DNAT'd and the proxy's upstream deliberately the resolver *behind* resolved, they would NXDOMAIN for the whole run — REFUSED under enforce, silently under audit, where no connection is ever attempted and nothing is logged (`/etc/hosts` covers only the exact name `localhost`). The proxy relays them to the stub on its marked client socket (`pkg/dns/synthetic.go`; the mark RETURN rules pass it through the stub DNAT, as they do the startup cache peek) and writes resolved's own answer back — uncached, unenforced, never fed to hostname tracking or the firewall, so reaching the gateway stays an explicit CIDR rule. Ahead of the filter gate; host listeners and IN class only (a container's native path never resolved them). A stub that does not answer — resolved stopped, or never installed — sends the query down the ordinary path; `/run/systemd/resolve` outlives a stopped resolved (`RuntimeDirectoryPreserve=yes`), so directory presence is no probe, and the connection-refused round trip on loopback is the cheapest true one. The search-expanded form a stub resolver tries first (`_gateway.<search>`, for a suffix on the host's own `resolv.conf` search list) is answered NXDOMAIN locally — what the upstream says natively — rather than REFUSED, which c-ares treats as terminal on a multi-label attempt and which therefore left `_gateway` unresolvable for c-ares clients under enforce; it is never relayed to the stub, which would send it upstream unmarked and DNAT it back into the proxy. This is specific to the synthetic set, whose expanded form never exists; for a real host the expanded form is the name that resolves (#127)
- Synthetic names (#126): systemd-resolved answers `_gateway`, `_outbound`, `_localdnsstub`, `_localdnsproxy`, the machine's own hostname (its configured name and, when it carries a domain, its first label — read per lookup, so a rename mid-run is honoured; `/etc/hosts` does not always pin it, and `sudo` warns on every invocation when it fails) and the RFC 6761 localhost family (`localhost`, `localhost.localdomain`, anything beneath either) from local state and never sends them upstream. With the stub DNAT'd and the proxy's upstream deliberately the resolver *behind* resolved, they would NXDOMAIN for the whole run — REFUSED under enforce, silently under audit, where no connection is ever attempted and nothing is logged (`/etc/hosts` covers only the exact name `localhost`). The proxy relays them to the stub on its marked client socket (`pkg/dns/synthetic.go`; the mark RETURN rules pass it through the stub DNAT, as they do the startup cache peek) and writes resolved's own answer back — uncached. `_gateway`, `_outbound` and the machine names get L4 enforcement and attribution like any upstream answer (#129): a hostname rule naming `_gateway` opens the gateway's address on its ports, a deny rule closes it, and with no rule the connection is attributed to `_gateway` rather than a bare IP. They mint no L7 forward-resolution evidence — `enforceDNSResponse(…, localAnswer)` at the synthetic call site, and startup pre-population records its answers through `dns.Server.RecordSystemCacheAnswer`, the same policy: a tracked alias mapped without evidence, a loopback listener recorded nowhere at all — because an alias is never a name the peer presents, and binding it would L7-scope the address against a name no flow can show: an all-ports allow would then drop HTTP to the host agent for carrying `Host: <ip>`. SCOPE IFF BOUND makes the omission sufficient; nothing in the L7 matcher knows about synthetic names. `_localdnsstub`, `_localdnsproxy` and the localhost family are relayed only: loopback, which the proxy's own relay and cache peek depend on, and any label under `.localhost` is accepted, so tracking them would grow the per-host maps without bound. The mDNS `<hostname>.local` form is not relayed: resolved renames it on conflict, and a `.local` query it does not synthesize is multicast to the LAN, whose answer is not local state. Ahead of the filter gate; host listeners and IN class only (a container's native path never resolved them). A stub that does not answer — resolved stopped, or never installed — sends the query down the ordinary path; `/run/systemd/resolve` outlives a stopped resolved (`RuntimeDirectoryPreserve=yes`), so directory presence is no probe, and the connection-refused round trip on loopback is the cheapest true one. The search-expanded form a stub resolver tries first (`<name>.<search>`, for a suffix on the host's own `resolv.conf` search list) is answered NXDOMAIN locally for every one of these names — what the upstream says natively — rather than REFUSED, which c-ares treats as terminal on a multi-label attempt and which therefore left `_gateway` unresolvable for c-ares clients under enforce; it is never relayed to the stub, which would send it upstream unmarked and DNAT it back into the proxy. This is specific to names resolved synthesizes, whose expanded form never exists; for a real host the expanded form is the name that resolves (#127). Residual: a machine hostname that is also a real DNS name resolves to resolved's synthesized local address rather than the upstream record — inherent to making the machine hostname resolve at all, and what the host does natively, since resolved synthesizes its own name before any upstream lookup
- **Sudo lockdown:** writes `/etc/sudoers.d/zz-cargowall-lockdown` with a NOPASSWD allowlist; removes the runner user from sudo-granting groups (`sudo`, `admin`, `wheel`) and the `docker` group; disables competing sudoers.d files by renaming them to `*.cargowall-disabled`
- **Auto-infrastructure:** `EnsureInfraAllowed()` and `EnsureHostnameAllowed()` add rules for platform services (Azure IMDS, GitHub API, etc.)
- **Logging:** `slog.Handler` that formats messages as GitHub workflow commands (`::error::`, `::warning::`, `::debug::`)
9 changes: 9 additions & 0 deletions pkg/dns/hostsearch_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,7 @@ import (
// opts in through withResolvConf.
func TestMain(m *testing.M) {
resolvConfPath = filepath.Join(os.TempDir(), "cargowall-dns-test-absent-resolv.conf")
machineHostname = func() (string, error) { return "", os.ErrNotExist }
os.Exit(m.Run())
}

Expand Down Expand Up @@ -121,6 +122,8 @@ func TestHandleDNSQuery_HostSearchExpandedFormResolvesAndEnforces(t *testing.T)
s := newTestServer(t, cfg, mockFw)
s.upstream = pc.LocalAddr().String()
s.filterQueries = true
rec := &recordingRegistrar{}
s.SetL7Registrar(rec)

q := new(dns.Msg)
q.SetQuestion("myservice.corp.lan.", dns.TypeA)
Expand All @@ -134,4 +137,10 @@ func TestHandleDNSQuery_HostSearchExpandedFormResolvesAndEnforces(t *testing.T)
assert.Equal(t, dns.RcodeSuccess, w.msg.Rcode)
require.Len(t, w.msg.Answer, 1)
assert.Equal(t, "192.0.2.10", w.msg.Answer[0].(*dns.A).A.String())

// A real answer for an allowed name is a wire identity: evidence minted,
// address L7-scoped for the rule's (all) ports — the contrast to the
// synthetic path, which mints none.
assert.True(t, s.config.NameResolvedToIP("myservice.corp.lan", "192.0.2.10"))
assert.Equal(t, "all-ports", rec.scopes["192.0.2.10"])
}
33 changes: 27 additions & 6 deletions pkg/dns/server.go
Original file line number Diff line number Diff line change
Expand Up @@ -645,7 +645,7 @@ func (s *Server) handleDNSQuery(w dns.ResponseWriter, r *dns.Msg) {
// DNS-path output consistent with the connection-event path, which logs
// the lowercase hostname from the IP->hostname mapping (#65).
if len(r.Question) > 0 && resp.Rcode == dns.RcodeSuccess {
s.enforceDNSResponse(strings.ToLower(strings.TrimSuffix(r.Question[0].Name, ".")), resp, 0)
s.enforceDNSResponse(strings.ToLower(strings.TrimSuffix(r.Question[0].Name, ".")), resp, 0, wireAnswer)
}

// Return response to client
Expand All @@ -654,14 +654,32 @@ func (s *Server) handleDNSQuery(w dns.ResponseWriter, r *dns.Msg) {
}
}

// answerSource says where a DNS answer came from, for enforceDNSResponse's
// L7 evidence: only an answer that came off the wire, for a name a peer can
// present, may mint forward-resolution evidence.
type answerSource int

const (
// wireAnswer: an upstream answer, or a CNAME pre-resolve of one.
wireAnswer answerSource = iota
// localAnswer: relayed from resolved's own state — an address alias, not
// an identity. Scoping its address against it would make every flow
// there an L7 miss.
localAnswer
)

// enforceDNSResponse applies one successful DNS response to enforcement
// state: rule/derived verdict matching, CNAME-chain learning, BPF map updates
// for resolved IPs, late-allow reconciliation of previously blocked
// connections, and pre-resolution of allowed CNAME-only responses.
// canonicalHostname is the queried name, lowercased with the trailing dot
// trimmed. depth bounds pre-resolve recursion (see preResolveCNAMETarget);
// handleDNSQuery passes 0.
func (s *Server) enforceDNSResponse(canonicalHostname string, resp *dns.Msg, depth int) {
// handleDNSQuery passes 0. source says whether the answer came off the wire
// (wireAnswer: upstream or pre-resolve, for a name a peer can present) and
// so may mint L7 forward-resolution evidence, or from resolved's local state
// (localAnswer: an address alias no peer presents) and mints none, so SCOPE
// IFF BOUND scopes nothing.
func (s *Server) enforceDNSResponse(canonicalHostname string, resp *dns.Msg, depth int, source answerSource) {
// Extract IPs and TTLs from response
ips, ttl := s.extractIPsFromResponse(resp)

Expand Down Expand Up @@ -833,8 +851,11 @@ func (s *Server) enforceDNSResponse(canonicalHostname string, resp *dns.Msg, dep
// This is THE forward-resolution path — a real DNS answer
// traversing the proxy — so it (and RecordCNAMEChain below)
// are the only seeds of the L7 per-IP binding evidence.
// Reverse-DNS paths must never record it (PTR forgery).
s.config.RecordForwardResolution(canonicalHostname, ip.String())
// Reverse-DNS paths must never record it (PTR forgery), and
// an answer that did not come off the wire mints none.
if source == wireAnswer {
s.config.RecordForwardResolution(canonicalHostname, ip.String())
}
}

// Track the IPs we've seen for this hostname. Accumulate
Expand Down Expand Up @@ -1294,7 +1315,7 @@ func (s *Server) preResolveCNAMETarget(target string, depth int) {
if resp.Rcode != dns.RcodeSuccess {
continue
}
s.enforceDNSResponse(target, resp, depth)
s.enforceDNSResponse(target, resp, depth, wireAnswer)
}
}()
}
Expand Down
Loading
Loading