Skip to content

Resolver synthetic records die at the stub DNAT: _gateway stops resolving, silently in audit mode #126

Description

@matthewdevenny

Symptom

systemd-resolved synthesizes a handful of names that exist in no zone and are never sent upstream. The DNS redirect DNATs the resolved stub to our proxy (pkg/network/dns_redirect.go:75-76), and the proxy's upstream is deliberately not the stub — detectDnsUpstream classifies 127.0.0.53 as loopback (cargowall-action src/dns.ts:47) and falls back to /run/systemd/resolve/resolv.conf, the real upstream, mirroring detectSystemdResolvedUpstreams (cmd/start.go:1785). Net effect: every resolved synthetic record stops resolving for the length of the run.

Verified on the Lima box (Ubuntu, systemd-resolved, search lan):

$ resolvectl query _gateway
_gateway: 192.168.5.2        -- link: eth0
-- Data from: synthetic

$ resolvectl query _outbound        →  192.168.5.15   -- Data from: synthetic
$ resolvectl query _localdnsstub    →  127.0.0.53     -- Data from: synthetic

$ grep ^hosts /etc/nsswitch.conf
hosts:          files dns                     ← no libnss-resolve, so this is a real UDP/53 lookup

$ dig @1.1.1.1 _gateway
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN

localhost is synthetic too but survives, because files comes first in nsswitch and /etc/hosts carries it. _gateway, _outbound and _localdnsstub exist nowhere else, so with the redirect installed they NXDOMAIN.

Both postures break, differently:

  • enforce — no hostname rule can match (isQueryAllowed, pkg/dns/server.go:535-575), so the query is REFUSED and logged as dns_blocked.
  • audit — the query is forwarded, the real upstream NXDOMAINs it, and no event is recorded at all. The name fails and the run report is clean.

The audit case is the bad one: the failure is invisible in exactly the mode operators use to decide whether enforce is safe.

Where it bites

Found while answering a customer question about Blacksmith GitHub Actions runners. Their microVMs reach a host-side agent — sticky disks, job manager, log ingestion, metrics, cache manager, Bazel RE cache, VM metadata — and inject its address as BLACKSMITH_AGENT_ADDR. Confirmed from useblacksmith's own public CI logs (useblacksmith/checkout container job):

-e "BLACKSMITH_AGENT_ADDR=192.168.127.1"
-e "BLACKSMITH_BAZELRE_CACHE_GRPC=grpc://192.168.127.1:1055"
-e "BLACKSMITH_ACTIONS_RESULTS_URL=http://192.168.127.96:49193/"

192.168.127.1 is fixed across every host we sampled (10 distinct machines, two regions) — it's the gvisor-tap-vsock gateway inside each microVM, not a per-tenant allocation. Ports vary per VM (1026-1080, plus 49193 and 64001). Some runner generations name that host _gateway rather than the literal IP, and that is the configuration this issue is about.

Failure signature for the operator: their build gets slower and loses cache hits, while the Blacksmith actions warn and silently degrade —

BLACKSMITH_AGENT_ADDR or BLACKSMITH_STICKY_DISK_GRPC_PORT is not set; the Blacksmith
agent is not reachable from this runner, falling back to actions/checkout behavior

— and CargoWall reports nothing, because no connection was ever attempted.

Why the design is what it is

Neither half is wrong on its own. The stub DNAT exists for a real reason (dns_redirect.go:70-74): a stub-following client otherwise reaches resolved, which re-queries upstream from its own socket, laundering the querying process away from the sockdiag join and serving warm cache hits the proxy never sees. And skipping 127.0.0.53 as an upstream is what stops the proxy from looping back into itself.

The consequence is that we take over the stub's role without taking over its full behaviour. resolved answers synthetics before it answers anything else; our proxy answers none of them.

Proposal

1. Answer resolved's synthetic names in the proxy.

Handle them in handleDNSQuery before the filter gate and before the cache (pkg/dns/server.go:535), the same way the REFUSED path constructs a message locally (:569):

  • _gateway → the default-route gateway addresses, ordered by metric
  • _outbound → the source address chosen for outbound traffic
  • _localdnsstub → 127.0.0.53

Gate the whole thing on resolved actually running — probe /run/systemd/resolve, the same predicate FlushResolvedCache already uses (dns_redirect.go:122) — so we never invent names on a host whose resolver wouldn't have answered them either. A/AAAA and the matching NODATA for other qtypes; short TTL, since the route can change mid-run.

2. Don't let the filter gate refuse them.

Answering before the gate covers this, and is better than an allow-rule carve-out: these names are locally synthesized, never leave the box, and carry no tunneling channel.

3. Log the NXDOMAIN case until (1) lands.

An upstream NXDOMAIN for a name we know is synthetic is a CargoWall-caused failure and should not be silent in audit mode.

Out of scope

  • Auto-allowing the synthesized address. Resolution and egress are separate decisions. On Blacksmith the gateway fronts a metadata service that vends stickyDiskToken; granting reachability because we answered a DNS query would be a policy change smuggled in through a resolver fix. Operators allow it explicitly — 192.168.127.0/24, all ports, in Blacksmith's case.
  • Making _gateway writable as a hostname rule. The control plane's BeValidHostname regex has no underscore in its DNS label class, so _gateway is refused on every write path. That is correct — a leading underscore is not a legal hostname label, which is exactly why systemd chose the name — and CIDR is the right allow path regardless. Separate concern, no change requested here.
  • Undoing the stub DNAT. The attribution reason for it stands.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions