Skip to content

feat: synthesize HTTP responses for HTTPS dependencies (opt-in TLS interception) - #4

Merged
achoimet merged 8 commits into
mainfrom
feat/https-response-injection
Sep 4, 2026
Merged

achoimet merged 8 commits into
mainfrom
feat/https-response-injection

Conversation

@achoimet

@achoimet achoimet commented Sep 2, 2026

Copy link
Copy Markdown
Member

What

An httpStatus fault could only ever be applied to cleartext HTTP. TLS connections were matched by SNI but always spliced through untouched — so an HTTPS dependency could be delayed or reset, but never made to return a status.

This adds an opt-in interception path. When a CA is configured and a matching rule carries httpStatus, the connection is terminated with a short-lived certificate minted for its SNI, and the response is synthesized inside TLS.

transparent-proxy \ --tls-ca-cert /etc/steadybit/intercept-ca.crt \ --tls-ca-key  /etc/steadybit/intercept-ca.key \ --fault-hosts api.stripe.com --fault-http-status 503

Design notes

HTTP/1.1 and HTTP/2, no new dependency. The forged response is served through net/http over the single already-handshaken connection, so ALPN picks the protocol and h2 works for free. The repo stays stdlib-only (no x/net, no go.sum) — deliberate, given this binary ships into customer network namespaces.

The CA is the customer's. They generate it, choose its validity, and install it in their workloads' truststores. We only sign per-SNI leaves with it — never create, rotate or renew. Leaves are clamped to never outlive the CA, and an already-expired CA is rejected at startup rather than failing every handshake with an opaque error.

One-sided by design. The real dependency is never dialed. So the proxy makes no trust decision about the origin's certificate, and a dependency behind mutual TLS is unaffected. The trade-off is that the response is fabricated, not a modified real one.

Failure is diagnosable, not silent. If the workload doesn't trust the CA (or pins certificates) the handshake fails. That is counted as tls_handshake_failures and deliberately not as an applied fault — so a non-zero value is the signal that the CA is missing from the truststore. Faults are still counted only at the point the response is actually delivered.

Backwards compatibility

Nothing changes without a CA. TLS is spliced through untouched when there is no CA, no SNI, no httpStatus rule, or the connection lost the probability roll. TestServer_TLSInject_DisabledPassesThrough pins this.

Tests

104 pass under -race. New coverage:

  • internal/tlsinject (16): CA validation (non-CA, mismatched key, garbage, expiry), per-SNI minting incl. IP SANs, caching, leaf validity clamped to the CA, SNI-required, hop-by-hop stripping, and real h1 and h2 round-trips against a live listener plus the untrusted-client and context-cancellation paths.
  • internal/proxy (3): end-to-end wiring — forged response instead of the real upstream (with per-host metrics), pass-through when disabled, and untrusted client counted as a handshake failure and never as a fault.

Not in this PR

Plumbing the CA through action-kit/proxyfault and the extensions (config, 443 guard, Helm/truststore docs) — needed for end-to-end testing via the platform.

An httpStatus fault could previously only be applied to cleartext HTTP: TLS
connections were matched by SNI but always spliced through untouched, so an
HTTPS dependency could be delayed or reset but never made to return a status.

Add an opt-in interception path. When a customer-supplied CA is configured and
a matching rule carries httpStatus, the connection is terminated with a
short-lived certificate minted for its SNI and the response is synthesized
inside TLS. The response is served through net/http, which supports HTTP/1.1
and HTTP/2 by ALPN without adding a dependency (the repo stays stdlib-only).

The CA belongs to the customer: they generate it, choose its validity, and
install it in their workloads' truststores. The proxy only signs per-SNI
leaves, clamped to never outlive the CA, and rejects an already-expired CA at
startup. Interception is one-sided — the dependency is never dialed — so no
upstream trust decision is made and mutual-TLS dependencies are unaffected.

Without a CA, without an SNI, or without an httpStatus rule, TLS is spliced
through exactly as before. A client that rejects the minted certificate is
recorded as tls_handshake_failures and never as an applied fault, so a missing
CA in the truststore is diagnosable instead of a silent no-op.
Smoke-testing against a real curl/OpenSSL client showed a rejected connection
being reported as a successful injection: connections_faulted and
http_responses_injected both incremented while the handshake-failure counter
stayed at zero, so the CA-not-trusted diagnostic never fired.

Two causes. Under TLS 1.3 the server's handshake completes before the client's
verdict on the certificate arrives — a client that rejects it does not fail the
handshake, it abandons the connection without sending a request. And success
was inferred from ServeForged returning, which happens whenever the connection
closes, even if net/http never invoked the handler.

Delivery is now tracked in the handler itself and is the only thing counted as
a fault. HandshakeError becomes RejectedError, carrying the stage so both the
TLS 1.2 (handshake) and TLS 1.3 (post-handshake) refusals are reported as what
they are. Teardown via a cancelled context is excluded, since that is not the
client's doing. The metric is renamed tls_intercept_rejected to match.

Verified end to end under real iptables: an untrusted client now yields
faulted=0, injected=0, rejected=1, and a trusted one faulted=1, injected=1.
achoimet and others added 6 commits September 2, 2026 15:13
… end

Seven fixes from review of this branch. The two that matter most:

Over HTTP/2 the fault counters stayed at zero for the whole attack. Delivery
was inferred from ServeForged returning, but that only happens once the client
goes away — and an h2 client (gRPC, pooled SDK clients: exactly the
dependencies this targets) holds the connection open for the duration. A
working attack therefore reported 'matched but never faulted', which is the
proxy's documented silent-no-op signature. Delivery is now reported from the
handler via Request.OnDelivered, the moment the response is written.

Conversely, teardown was counted as a delivered fault: a cancelled context
returned nil, and the caller booked an injected 503 for a connection that got
nothing. OnDelivered simply never fires there.

Also: a cancelled or timed-out handshake is no longer blamed on the truststore
(it was bucketed as 'client rejected the certificate'); cached leaves are
re-minted before they expire, and minting fails loudly once the CA has expired
mid-run, instead of serving certificates every client rejects; 1xx is treated
as out of range, since net/http does not commit an informational status and the
body would silently commit 200; 204/304 no longer carry a body or
Content-Length; a failed injection increments Dropped so matched still
reconciles with the outcome counters; and the CA is loaded after the --revert
branch, so teardown can never be blocked by a missing or expired CA.
Passing the CA by file path does not work for the runc backend. That backend
runs the proxy in a bundle whose rootfs is an overlay of the orchestrator's
"/", and an overlay does not carry the orchestrator's submounts — so a CA
mounted from a Kubernetes Secret is invisible by path inside the sidecar.
Verified against the real mount options: image-layer files are visible, while
the Secret mount point appears as an empty directory.

Left as-is this would have failed only at runtime, and failed misleadingly:
every handshake would abort and be reported as "the client rejected our
certificate", pointing the operator at their truststore rather than at a
proxy that never had a usable CA.

--tls-ca-stdin reads one PEM stream carrying both halves, in any order. It
works identically for both backends, keeps the key off the command line, and
never writes it to a filesystem the target could reach. The file flags stay for
standalone and manual use, and the two forms are mutually exclusive.

Verified under real iptables: the CA loads from stdin, a trusted client gets
the forged 503 over HTTP/2, an untrusted one is counted as rejected and not as
faulted, and untargeted traffic still passes through.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019wB5XrsrNJjU9MH6yegmTA
Six review findings, most of them variations on one theme: an error that is
ours getting reported as the client's.

The mid-run CA-expiry check added last commit exists precisely to avoid
misdiagnosis, and was then undone by the classification above it — every
non-context handshake error became a RejectedError, so an expired CA logged
"is the CA trusted by the target?" and sent the operator to inspect a truststore
that was fine. Certificate production failures now carry a CertError and are
reported as themselves. The same applied to a ClientHello with no SNI.

ServeForged returned nil for two non-delivery cases, so those connections were
counted in matched and in nothing else; at teardown with many connections in
flight, a whole batch could vanish from the outcome counters. Teardown now
returns ErrNotDelivered and the caller records it.

Delivery was read rather than claimed. An HTTP/2 handler can still be finishing
as ServeConn returns, so a plain read could see "not delivered", have the caller
count a rejection, and then have the straggler fire OnDelivered — one connection
booked into two mutually exclusive buckets. A CAS closes that window.

The stdin read was unbounded and untimed, and runs before the listener is bound,
before preflight, and before the deadman is armed: a writer that never closed
the pipe would hang the proxy forever with nothing installed. Now capped at
1 MiB with a 30s deadline.

Also: cached leaves were compared against a fixed renew window while their
validity is clamped to the CA's, so a CA with under an hour left made every
cached leaf permanently stale and re-signed on every handshake; the window is
now clamped to half the remaining validity. A passphrase-protected key is
refused with a pointed message instead of an opaque parse failure later. The
cleartext path normalises status the same way as the HTTPS path, so one rule no
longer produces different bytes depending on the dependency's protocol. And
--tls-ca-stdin is documented.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019wB5XrsrNJjU9MH6yegmTA
--tls-leaf-validity sets how long the per-SNI certificates the proxy mints stay
valid, so an operator can narrow the window in which a leaf that escaped the
proxy would be usable. Unset keeps the previous 24h.

Values below twice the renew window are raised to it: a leaf shorter than the
window is stale the moment it is issued, which would re-sign on every single
handshake. The value is still clamped to the CA's own expiry, so a leaf never
outlives its issuer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019wB5XrsrNJjU9MH6yegmTA
…stnames

The capture filter is deliberately broad — typically 0.0.0.0/0 on ports 80 and
443 — because the proxy decides what to fault by hostname once it has seen the
request. The flush cannot do that: it is a stateless iptables REJECT that knows
only addresses. Scoping it to the capture filter therefore reset every
established HTTP/HTTPS connection in the target, including ones to dependencies
the attack never named. Starting a slow-dependency attack on one hostname
briefly severed everything else the workload was talking to.

The targeted hostnames are now resolved up front, inside the target's network
namespace so the answers match what the workload sees, and only those addresses
are flushed. Capture stays broad, since hostname matching still needs it.

This is a snapshot: a dependency behind rotating IPs may hold a connection to an
address that no longer resolves, and that one is not reset — it is still faulted
when it next reconnects. Under-flushing is the right side to err on. A hostname
that fails to resolve warns rather than failing the attack, and an attack that
targets CIDRs rather than hostnames keeps the previous filter-wide flush, which
is what it actually wants.

Verified under real iptables: the flush chain names only the two addresses
example.com resolved to, while the REDIRECT chain still captures 0.0.0.0/0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019wB5XrsrNJjU9MH6yegmTA
… CA works

Only the signing certificate was presented alongside the minted leaf, so a
certificate authority that is itself issued by something else left the client
unable to build a path. That forced the operator towards handing over a root
key, which is the last thing anyone should be asked for.

The full supplied bundle is now presented, signing certificate first. An
operator can issue a short-lived intermediate from their own PKI, keep the root
offline, and hand the proxy only the intermediate and its key — the workloads
already trust the root, and the chain resolves. Constraining that intermediate
with nameConstraints then bounds what it can impersonate at all.

Covered by a test that verifies a client trusting only the root against a leaf
minted by an intermediate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019wB5XrsrNJjU9MH6yegmTA
@achoimet
achoimet merged commit 45e9866 into main Sep 4, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant