Skip to content

Re-attribute container-origin events from the start→tag window once the container is tagged #116

Description

@matthewdevenny

Symptom

A container's first DNS queries race the docker-event → tag window and land in the "Container (unattributed)" tier, while identical retries seconds later attribute to the step — the same domain shows up in both buckets.

Observed (Docker DNS Integration, https://github.com/code-cargo/cargowall-action/actions/runs/32531947802/attempts/1#summary-96925548617):

22:12:13.967  DNS query blocked domain=example.com from=172.17.0.2:56265 container_origin=true            ← no ordinal
22:12:13.968  Container attributed container=cd34651b6ef5 kind=start step_ordinal=2 tag_latency_ms=7.452  ← tag lands
22:12:16.470  DNS query blocked domain=example.com from=172.17.0.2:56265 step_ordinal=2 container_id=cd34651b6ef5

The container process resolves before dockerd's start event reaches the watcher, so at query time the container genuinely has no tag. This is the accepted start→tag residual from #106 phase 3a (design.md: "traffic from the start→tag window lands in the container tier, not a step") — correct-by-design, but it is now the last visibly unattributed traffic in an otherwise fully attributed run.

Proposal: late re-attribution on kind=start

Faster tagging can't win this race (the event-driven path cannot beat the container's first syscall). Instead, reconcile backwards: when a container is tagged with kind=start, re-attribute recent container-origin events from that container's IP that lack an ordinal — stamping the freshly learned step_ordinal/container_id.

  • Direct precedent: the Late-allowed IPs shipped to SaaS as blocks when A records bypass the DNS proxy #83 late-allow reconciliation (events.RecentBlocks) already buffers recently blocked connections and re-logs them when the firewall opens for their destination IP. Same shape, keyed on client IP instead.
  • Bounded window: tag latency is single-digit ms (7.45ms above); even a few seconds is safe causally — a veth IP cannot be recycled to a different container inside that window.
  • Emit as a reconciliation event (mirroring connection_late_allowed) rather than mutating history, so the audit log stays append-only; the summary/SaaS dedup then collapses the pair the way late-allow already does.
  • Keeps the "never guess" rule: only events whose source IP matches the tagged container, only within the window, only when the container resolved to a real ordinal.

Out of scope

  • exec-window equivalents (kind=exec) — same idea, separate consideration.
  • Anything policy-facing: DNS-path attribution stays audit-only (design.md, "DNS and per-step policy").

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions