From 1ae261f1cf171fc8693b62bca7c15b153773bbfd Mon Sep 17 00:00:00 2001 From: vthwang Date: Sun, 16 Aug 2026 23:17:01 -0700 Subject: [PATCH] feat: move the ingress layer from ingress-nginx to Traefik MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The dids daemon warn-logs every request that claims a forwarded host from a peer outside its trusted-CIDR set, and an ingress puts X-Forwarded-Host on every request — so it logged one warning per request, forever. Trusting the ingress CIDR would have silenced it by granting host control to every pod on a flat network, which is exactly what the daemon's gate exists to prevent. Stripping the header at the ingress is impossible under ingress-nginx: snippet annotations are disabled cluster-wide (the admission webhook refuses them), and even enabled, the snippet renders after the controller's own `proxy_set_header X-Forwarded-Host $best_http_host` in the same location, where nginx sends both values rather than letting the later one win. Traefik writes the X-Forwarded-* set at the entrypoint, before middlewares run, so a headers middleware with an empty value actually removes it — verified against v3.3, and the dids Ingress now carries one per user namespace. The controller swap subsumes the rest: - ingressClassName is Traefik's, hardcoded in k8s.IngressClass. A different controller needs different entrypoint settings, default certificate and ACME solver class too, so one env var could never carry the switch. - The Ingresses lose every annotation. HTTP→HTTPS is now an entrypoint redirect and the wildcard certificate comes from the `default` TLSStore, so there is nothing left for them to say. The wildcard moves to kube-system because a TLSStore can only reference a Secret in its own namespace. - The HTTP-01 solver names the traefik class; Let's Encrypt follows the entrypoint redirect and does not validate the certificate it lands on. scripts/migrate-ingress-to-traefik.sh repoints Ingresses that already exist — this API only ever creates them, so without it they keep the nginx class and Traefik never adopts them. Dry-run by default, and it reports the chart-managed Ingresses it deliberately does not touch rather than editing a chart's output. Two things the dev migration taught, both recorded where they will be read: Helm accepts a values path that is one level off without complaining, so the entrypoint keys must be verified against the rendered args and not this file; and managed hostnames are proxied through Cloudflare, so verification has to pin the node IP or it reports Cloudflare's certificate instead of ours. Signed-off-by: vthwang --- .env.example | 2 +- CLAUDE.md | 45 ++++- README.md | 66 +++++-- docs/custom-domain-design.md | 14 +- docs/full-stack-setup-design.md | 4 +- docs/vta-setup-design.md | 12 +- .../templates/vtafarm-api/clusterrole.yaml | 6 + .../templates/vtafarm-api/ingress.yaml | 2 +- internal/k8s/component_resources.go | 33 ++-- internal/k8s/traefik.go | 96 ++++++++++ internal/k8s/vta_resources.go | 13 +- internal/setup/orchestrator_fullstack.go | 34 ++-- k8s/secret.yaml.example | 2 +- k8s/tls/certificate.yaml | 17 +- k8s/tls/clusterissuer-http01.yaml | 9 +- k8s/tls/rke2-ingress-nginx-config.yaml | 10 -- k8s/tls/rke2-traefik-config.yaml | 84 +++++++++ k8s/tls/tlsstore-default.yaml | 29 +++ scripts/migrate-ingress-to-traefik.sh | 165 ++++++++++++++++++ 19 files changed, 572 insertions(+), 71 deletions(-) create mode 100644 internal/k8s/traefik.go delete mode 100644 k8s/tls/rke2-ingress-nginx-config.yaml create mode 100644 k8s/tls/rke2-traefik-config.yaml create mode 100644 k8s/tls/tlsstore-default.yaml create mode 100755 scripts/migrate-ingress-to-traefik.sh diff --git a/.env.example b/.env.example index b540404..27ac9b1 100644 --- a/.env.example +++ b/.env.example @@ -24,7 +24,7 @@ KUBECONFIG= K8S_NAMESPACE_PREFIX=fpp-user # Cluster (required for setup wizard) -# External IP of the cluster's Ingress-NGINX LoadBalancer +# External IP of the cluster's Traefik LoadBalancer CLUSTER_INGRESS_IP= # Root domain managed by Cloudflare (subdomains will be created under this) CLUSTER_DOMAIN= diff --git a/CLAUDE.md b/CLAUDE.md index da010af..7ef0178 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -49,7 +49,7 @@ See `.env.example` for all options. Key ones: | `DID_HOSTING_DID` / `DID_HOSTING_PRIVATE_KEY` | — | vtafarm-api's **own** keypair (`make gen-keypair`) for the DID-hosting control API, enrolled in a daemon's ACL with `role=admin`. Not anything a daemon issued, so one keypair serves every daemon it is enrolled in. There are deliberately no DID-hosting **URLs** here — see "Shared infrastructure comes from the platform stack" below | | `CLOUDFLARE_API_TOKEN` | — | Cloudflare API token (`Zone:DNS:Edit` permission) | | `CLOUDFLARE_ZONE_ID` | — | Cloudflare Zone ID for the user's root domain | -| `CLUSTER_INGRESS_IP` | — | External IP of the cluster's Ingress-NGINX LoadBalancer | +| `CLUSTER_INGRESS_IP` | — | External IP of the Traefik LoadBalancer | | `ACME_CLUSTER_ISSUER` | `letsencrypt-http01` | The same issuer in every environment — there is no staging variant. A staging certificate passes `tls_provision` and then crash-loops the mediator, because the components resolve each other's `did:webvh` over HTTPS and reject an untrusted chain. Every environment therefore shares Let's Encrypt's unraisable allowances (5 certs per identical name set per week), so keep iteration on domains we own | ## Project Structure @@ -504,6 +504,46 @@ Assign the correct tag so it appears in the right group in Scalar: ## Kubernetes Design +### The ingress controller is Traefik + +`internal/k8s/traefik.go` holds everything this API knows about it. Two things +that were per-Ingress annotations under ingress-nginx are now settings on the +controller itself (`k8s/tls/rke2-traefik-config.yaml`), which is why the +Ingresses we create carry **no annotations at all** beyond the one below: + +- **HTTP→HTTPS** is an entrypoint redirect (`ports.web.redirections`), not + `ssl-redirect` per Ingress. +- **The wildcard certificate** comes from the `default` TLSStore + (`k8s/tls/tlsstore-default.yaml`), not `--default-ssl-certificate`. The + TLSStore and the Secret it names must both live in Traefik's own namespace — + which is why `k8s/tls/certificate.yaml` issues into `kube-system` rather than + `cert-manager`. Get this wrong and every managed/platform hostname silently + serves Traefik's self-signed certificate, which doesn't fail until + `step_vta_register_dids` — the components resolve each other's `did:webvh` + over HTTPS and reject an untrusted chain. +- Custom domains still name their own Secret in the Ingress `tls:` block; a + router-level certificate wins over the store. + +**The dids Ingress carries a `strip-forwarded-host` Middleware**, one per user +namespace. The daemon warn-logs every request that claims a forwarded host from +a peer outside its trusted-CIDR set, and an ingress puts `X-Forwarded-Host` on +every request — so without this it logs a warning per request, forever. +Stripping rather than trusting the ingress CIDR is deliberate: pod networking is +flat, so a CIDR wide enough to cover the ingress also lets any other pod dictate +the request host. Traefik can actually remove the header because it writes the +X-Forwarded-* set at the entrypoint, before middlewares run; ingress-nginx +emits its own `proxy_set_header` ahead of any snippet and sends **both** values, +so the same fix was impossible there. + +`ingressClassName` is hardcoded (`k8s.IngressClass`), not configurable — an +environment on a different controller would also need different entrypoint +settings, a different default certificate and a different ACME solver class, so +one env var could never carry the switch. + +Sessions created before the switch keep `ingressClassName: nginx` forever (we +only ever create Ingresses, never update them): +`scripts/migrate-ingress-to-traefik.sh` repoints them, dry-run by default. + ### Per-User Namespace Isolation Every user gets their own namespace: `vtafarm-user-{userID}`. @@ -587,4 +627,7 @@ rules: - apiGroups: ["cert-manager.io"] resources: ["certificates"] verbs: ["get", "list", "watch", "create", "delete"] +- apiGroups: ["traefik.io"] + resources: ["middlewares"] + verbs: ["get", "list", "create", "delete"] ``` diff --git a/README.md b/README.md index 7579263..ddae96b 100644 --- a/README.md +++ b/README.md @@ -104,7 +104,7 @@ Copy `.env.example` and adjust as needed: | `DB_NAME` | `vtafarm` | | | `JWT_SECRET` | _(required)_ | HS256 signing secret — must match the team, see below | | `ORCHESTRATOR_RESUME` | `true` | Re-attach interrupted sessions at startup. Set `false` locally — see [`docs/shared-dev-database.md`](docs/shared-dev-database.md) | -| `CLUSTER_INGRESS_IP` | _(required)_ | External IP of the cluster's Ingress-NGINX LoadBalancer | +| `CLUSTER_INGRESS_IP` | _(required)_ | External IP of the cluster's Traefik LoadBalancer | | `CLOUDFLARE_API_TOKEN` | _(optional)_ | Required for VTA setup wizard | | `CLOUDFLARE_ZONE_ID` | _(optional)_ | Required for VTA setup wizard | | `KUBECONFIG` | _(empty)_ | Auto-detects `~/.kube/config` when empty | @@ -133,8 +133,8 @@ The production stack is deployed to a Kubernetes cluster (RKE2) via Helm. ### 1. TLS — Wildcard Certificate via cert-manager (one-time cluster setup) All VTA sessions share a single `*.firstperson.dev` wildcard certificate managed by -cert-manager. nginx-ingress serves it as the default SSL certificate, so no -per-Ingress TLS configuration is needed. +cert-manager. Traefik serves it as its default certificate, so no per-Ingress TLS +configuration is needed. #### Step 1 — Create the Cloudflare API token Secret @@ -169,20 +169,66 @@ TXT record to Cloudflare, then removes it) and store the issued certificate. Check status with: ```bash -kubectl get certificate -n cert-manager firstperson-dev-wildcard +kubectl get certificate -n kube-system firstperson-dev-wildcard ``` -#### Step 4 — Configure nginx-ingress to use the wildcard cert by default +The Certificate issues into `kube-system` — Traefik's own namespace — because a +TLSStore can only reference a Secret alongside it. Adjust both files if Traefik +runs elsewhere in your cluster. -RKE2 manages its built-in nginx ingress controller via the `HelmChartConfig` CRD. +#### Step 4 — Configure Traefik + +RKE2 manages its bundled ingress controller through the `HelmChartConfig` CRD. +Two objects: the controller's entrypoints, and the default certificate. ```bash -kubectl apply -f k8s/tls/rke2-ingress-nginx-config.yaml +kubectl apply -f k8s/tls/rke2-traefik-config.yaml +kubectl apply -f k8s/tls/tlsstore-default.yaml ``` -RKE2 will reconcile the change and restart the ingress controller automatically. -After this, every VTA Ingress gets HTTPS automatically — no `tls:` block or -cert-manager annotation required on individual Ingress resources. +RKE2 reconciles the change and restarts Traefik automatically. After this every +Ingress gets HTTPS and an HTTP→HTTPS redirect with no annotation, `tls:` block +or cert-manager annotation of its own. + +Verify before going further — this is the step whose failure shows up several +minutes later as a mediator crash loop rather than as a TLS error. + +**Test against the origin, not the hostname.** Managed and platform records are +*proxied* through Cloudflare, so plain `curl https://` reports +Cloudflare's edge certificate (issuer: Google Trust Services) and Cloudflare's +status code — it tells you nothing about the cluster. Pin the node IP: + +```bash +IP=$(kubectl get nodes -o jsonpath='{.items[0].status.addresses[?(@.type=="InternalIP")].address}') + +# 1. The certificate the cluster itself serves — Let's Encrypt, not TRAEFIK DEFAULT CERT +openssl s_client -connect "$IP":443 -servername /dev/null \ + | openssl x509 -noout -issuer + +# 2. Routing. Pick a path the component actually serves: / on dids, /health on a +# VTA. A 404 on / from a VTA is correct and means nothing is wrong. +curl -skI --resolve :443:"$IP" https:// | head -1 + +# 3. The redirect really applied (see the note in rke2-traefik-config.yaml — +# a mistyped values path fails silently) +kubectl -n kube-system get ds rke2-traefik \ + -o jsonpath='{.spec.template.spec.containers[0].args}' | tr ',' '\n' | grep redirections +``` + +Note the ADDRESS column of `kubectl get ingress` stays **empty** under Traefik in +this layout, and that is not a fault — see the `publishedService` note in +`k8s/tls/rke2-traefik-config.yaml`. Read the CLASS column instead. + +#### Step 5 — Migrating an existing cluster off ingress-nginx + +Only when the cluster already ran sessions. vtafarm-api never updates an Ingress +after creating it, so pre-existing ones keep `ingressClassName: nginx` and +Traefik ignores them: + +```bash +KUBE_CONTEXT=k8s-fpp-dev ./scripts/migrate-ingress-to-traefik.sh # dry run +KUBE_CONTEXT=k8s-fpp-dev ./scripts/migrate-ingress-to-traefik.sh --apply +``` ### 2. HashiCorp Vault (one-time, before the API) diff --git a/docs/custom-domain-design.md b/docs/custom-domain-design.md index b0a3423..ce9b76f 100644 --- a/docs/custom-domain-design.md +++ b/docs/custom-domain-design.md @@ -107,9 +107,9 @@ what makes the platform stack nearly free: `vta-.firstperson.dev` and friends. vtafarm-api creates four **proxied** Cloudflare A records at session-create time. TLS comes from the -cluster-wide `*.firstperson.dev` wildcard that ingress-nginx serves as its -`default-ssl-certificate`. No Domains record involved, nothing for the user to -do. +cluster-wide `*.firstperson.dev` wildcard that Traefik serves as its default +certificate (the `default` TLSStore). No Domains record involved, nothing for +the user to do. ### 3.2 `custom` — new @@ -771,7 +771,7 @@ spec: solvers: - http01: ingress: - class: nginx + ingressClassName: traefik ``` **One issuer, every environment — there is deliberately no staging twin.** @@ -845,11 +845,11 @@ aaa.com allows letsencrypt.org. | Option | Verdict | | --- | --- | -| **Cloudflare for SaaS** — users CNAME to a proxied `lb`, Cloudflare issues and renews the edge certificate | **Deferred.** Free at this scale (100 custom hostnames included on every plan, then $0.10/mo each) and it would hide the origin IP. But origin-side TLS is unresolved: with a fallback origin, Cloudflare sends SNI = the custom hostname, our nginx answers with the wildcard, and Full (strict) rejects it. Fixing that means either downgrading the whole zone to Full, or running cert-manager **as well** — i.e. this option plus all of §8. Parked in §16.3; §4.1 keeps the migration path free. | +| **Cloudflare for SaaS** — users CNAME to a proxied `lb`, Cloudflare issues and renews the edge certificate | **Deferred.** Free at this scale (100 custom hostnames included on every plan, then $0.10/mo each) and it would hide the origin IP. But origin-side TLS is unresolved: with a fallback origin, Cloudflare sends SNI = the custom hostname, our ingress answers with the wildcard, and Full (strict) rejects it. Fixing that means either downgrading the whole zone to Full, or running cert-manager **as well** — i.e. this option plus all of §8. Parked in §16.3; §4.1 keeps the migration path free. | | **LE DNS-01 with `_acme-challenge` delegation** | Rejected — 4 more records for the user, and its benefits (no port-80 dependency, wildcard support) are ones we don't need. | | **LE DNS-01 with the user's DNS API credentials** | Rejected — a large trust ask, and provider-specific. | | **User-supplied certificates** | Rejected — manual renewal every 90 days. | -| **Caddy on-demand TLS** | Rejected — the slickest answer for arbitrary hostnames, but it means replacing RKE2's bundled ingress-nginx. | +| **Caddy on-demand TLS** | Rejected — the slickest answer for arbitrary hostnames, but it means replacing the bundled ingress controller. (Written when that was ingress-nginx; the cluster moved to Traefik afterwards, for unrelated reasons, and the verdict is unchanged — Caddy would still replace it.) | ### 8.6 On the origin IP @@ -862,7 +862,7 @@ openssl s_client 157.180.68.139:443 → CN=*.firstperson.dev curl --resolve dids-…:443:157.180.68.139 → HTTP 200 (identical to via-Cloudflare) ``` -ingress-nginx serves the wildcard certificate to *any* direct connection on +The ingress serves the wildcard certificate to *any* direct connection on :443, so internet-wide scanners (Censys, Shodan) already index the IP ↔ domain link. Choosing direct-to-origin therefore gives up nothing that is currently held. The real control would be firewalling the origin to diff --git a/docs/full-stack-setup-design.md b/docs/full-stack-setup-design.md index 9eb5abe..ed80940 100644 --- a/docs/full-stack-setup-design.md +++ b/docs/full-stack-setup-design.md @@ -178,8 +178,8 @@ A vtc-{vtc_name}.{domain} → {CLUSTER_INGRESS_IP} DNS must exist first because the rendered recipes embed the final `https://…` URLs (`public_url`, `webvh_url`, `[identity].public_url`, the VTC's `base_url`) into the DID -documents that get published. TLS is the cluster-wide wildcard default-ssl-certificate on -nginx-ingress (same as today's VTA Ingress — no per-host `tls:` block needed). +documents that get published. TLS is the cluster-wide wildcard, served by Traefik as its +default certificate (same as today's VTA Ingress — no per-host `tls:` block needed). `CreateARecord` already returns a record ID; store all four for teardown. diff --git a/docs/vta-setup-design.md b/docs/vta-setup-design.md index 5a30f71..091d3f9 100644 --- a/docs/vta-setup-design.md +++ b/docs/vta-setup-design.md @@ -165,7 +165,7 @@ The backend calls the Cloudflare API to create DNS records **before** running an | --- | --- | | `CLOUDFLARE_API_TOKEN` | Cloudflare API token with `Zone:DNS:Edit` permission | | `CLOUDFLARE_ZONE_ID` | Zone ID for the user's root domain (from Cloudflare dashboard) | -| `CLUSTER_INGRESS_IP` | External IP of the cluster's Nginx/Ingress-NGINX LoadBalancer | +| `CLUSTER_INGRESS_IP` | External IP of the cluster's Traefik LoadBalancer | | `CLUSTER_DOMAIN` | Root domain the generated subdomain is appended to (e.g. `example.com`) | ### DNS Records Created @@ -209,8 +209,8 @@ The returned record ID is stored on the session (`CFRecordID`) and used for tear ## Kubernetes Resource Provisioning All resources are created in the user's isolated namespace (`vtafarm-user-{userID}`). TLS is -served by the cluster-wide **wildcard default-ssl-certificate** on nginx-ingress — there is no -cert-manager and no per-host `tls:` block. +served by the cluster-wide wildcard, which Traefik serves as **its default certificate** (the +`default` TLSStore) — there is no per-host `tls:` block and nothing per-Ingress to configure. ### Resources (VTA Only) @@ -221,7 +221,7 @@ cert-manager and no per-host `tls:` block. | `Job` (setup) | `vta-setup-{sessionID}` | runs `vta setup --from …` | | `Job` (provision) | `vta-provision-{sessionID}` | runs `vta import-did` (+ `did-mgmt servers add`) | | `Service` | `vta-{sessionID}` | ClusterIP on 8100 | -| `Ingress` | `vta-{sessionID}` | nginx ingress (ssl-redirect annotation; wildcard TLS) | +| `Ingress` | `vta-{sessionID}` | Traefik ingress (no annotations; wildcard TLS + redirect both come from the controller) | | `Deployment` (server) | `vta-{sessionID}` | long-running VTA, starts after the provision Job | ### Ingress template @@ -232,10 +232,8 @@ kind: Ingress metadata: name: vta-{sessionID} namespace: vtafarm-user-{userID} - annotations: - nginx.ingress.kubernetes.io/ssl-redirect: "true" spec: - ingressClassName: nginx + ingressClassName: traefik rules: - host: {subdomain}.{CLUSTER_DOMAIN} http: diff --git a/helm/vtafarm-api/templates/vtafarm-api/clusterrole.yaml b/helm/vtafarm-api/templates/vtafarm-api/clusterrole.yaml index c6ba6db..05a5560 100644 --- a/helm/vtafarm-api/templates/vtafarm-api/clusterrole.yaml +++ b/helm/vtafarm-api/templates/vtafarm-api/clusterrole.yaml @@ -60,3 +60,9 @@ rules: - apiGroups: ["cert-manager.io"] resources: ["certificates"] verbs: ["get", "list", "watch", "create", "delete"] + # One Middleware per user namespace, stripping X-Forwarded-Host on the way to + # the dids daemon. Namespaced and not cross-namespace referenceable, so it + # cannot be a single shared object. + - apiGroups: ["traefik.io"] + resources: ["middlewares"] + verbs: ["get", "list", "create", "delete"] diff --git a/helm/vtafarm-api/templates/vtafarm-api/ingress.yaml b/helm/vtafarm-api/templates/vtafarm-api/ingress.yaml index 57aae60..13d7f8c 100644 --- a/helm/vtafarm-api/templates/vtafarm-api/ingress.yaml +++ b/helm/vtafarm-api/templates/vtafarm-api/ingress.yaml @@ -3,7 +3,7 @@ kind: Ingress metadata: name: {{ .Values.name }} spec: - ingressClassName: nginx + ingressClassName: traefik rules: - host: {{ .Values.ingress.host }} http: diff --git a/internal/k8s/component_resources.go b/internal/k8s/component_resources.go index bc03255..a48cfb2 100644 --- a/internal/k8s/component_resources.go +++ b/internal/k8s/component_resources.go @@ -223,28 +223,27 @@ type ComponentIngressSpec struct { Namespace, Name, ServiceName, Host string Port int32 // TLSSecret names the Secret serving this host's certificate. Empty selects - // the cluster-wide wildcard that ingress-nginx serves as its - // default-ssl-certificate — which covers every managed and platform - // hostname, so only custom domains ever set this. + // the cluster-wide wildcard Traefik serves as its default certificate — + // which covers every managed and platform hostname, so only custom domains + // ever set this. TLSSecret string + // StripForwardedHost attaches the namespace's strip-forwarded-host + // Middleware. Only the dids daemon needs it (see + // StripForwardedHostMiddleware); the caller must have created the + // Middleware first. + StripForwardedHost bool } -// CreateComponentIngress creates an nginx Ingress routing Host to ServiceName -// on Port. Idempotent. +// CreateComponentIngress creates an Ingress routing Host to ServiceName on +// Port. Idempotent. func (c *Client) CreateComponentIngress(ctx context.Context, spec ComponentIngressSpec) error { pathType := networkingv1.PathTypePrefix - ingressClass := "nginx" + ingressClass := IngressClass ingress := &networkingv1.Ingress{ ObjectMeta: metav1.ObjectMeta{ Name: spec.Name, Namespace: spec.Namespace, - Annotations: map[string]string{ - // cert-manager's HTTP-01 solver claims the more specific - // /.well-known/acme-challenge/ path, so this doesn't interfere - // with issuance. - "nginx.ingress.kubernetes.io/ssl-redirect": "true", - }, // Deliberately NO cert-manager.io/cluster-issuer annotation. We // create the Certificate ourselves (one covering all four hosts); // leaving the annotation here as well would have ingress-shim make @@ -280,6 +279,16 @@ func (c *Client) CreateComponentIngress(ctx context.Context, spec ComponentIngre SecretName: spec.TLSSecret, }} } + if spec.StripForwardedHost { + // Traefik fails the router outright when this names a Middleware that + // does not exist, so the object has to be in place before the Ingress — + // which is why the caller creates it, rather than this function doing it + // on the way past. + ingress.Annotations = map[string]string{ + "traefik.ingress.kubernetes.io/router.middlewares": MiddlewareRef( + spec.Namespace, StripForwardedHostMiddleware), + } + } _, err := c.kube.NetworkingV1().Ingresses(spec.Namespace).Create(ctx, ingress, metav1.CreateOptions{}) if err != nil && !k8serrors.IsAlreadyExists(err) { diff --git a/internal/k8s/traefik.go b/internal/k8s/traefik.go new file mode 100644 index 0000000..2c308b6 --- /dev/null +++ b/internal/k8s/traefik.go @@ -0,0 +1,96 @@ +package k8s + +import ( + "context" + "fmt" + + k8serrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime/schema" +) + +// Everything this API knows about the ingress controller lives in this file. +// +// Two things that were per-Ingress annotations under ingress-nginx are now +// entrypoint-level settings on the controller itself +// (k8s/tls/rke2-traefik-config.yaml): the HTTP→HTTPS redirect, and the wildcard +// certificate every managed and platform hostname is served with. That is why +// the Ingresses this package creates carry no annotations at all beyond the +// middleware reference below — there is nothing left for them to say. + +// IngressClass is the ingressClassName on every Ingress this API creates. +// Hardcoded rather than configurable: an environment that ran a different +// controller would also need different entrypoint settings, a different default +// certificate and a different ACME solver class, so one env var could never +// carry the switch on its own. +const IngressClass = "traefik" + +// StripForwardedHostMiddleware names the per-namespace Middleware that removes +// X-Forwarded-Host from requests on their way to the dids daemon. +// +// The daemon resolves a request's intended host from Host unless the direct TCP +// peer is inside its trusted-CIDR set, and warn-logs every request that claims a +// forwarded host from outside it. Traefik sets X-Forwarded-Host on every proxied +// request, so without this the daemon logs one warning per request, forever. +// +// Removing the header is the right half of that trade rather than trusting the +// ingress: pod networking is flat, so a CIDR wide enough to cover the ingress +// also lets any other pod in the cluster dictate the request host, and the whole +// point of the daemon's gate is that a spoofed X-Forwarded-Host cannot make a +// resolution answer for somebody else's domain. Stripping keeps Host-based +// routing — which is what everything already runs on — and grants nobody +// anything. +// +// Note this is only possible because Traefik writes the X-Forwarded-* headers at +// the entrypoint, before middlewares run. Under ingress-nginx the equivalent +// directive is emitted by the controller's own template ahead of any snippet, +// and nginx sends both values rather than letting the later one win, so the +// header could not be removed there at all. +const StripForwardedHostMiddleware = "strip-forwarded-host" + +// middlewareGVR reaches Traefik's CRD through the dynamic client, so we never +// import its Go module for one object. Same reasoning as certificateGVR. +var middlewareGVR = schema.GroupVersionResource{ + Group: "traefik.io", Version: "v1alpha1", Resource: "middlewares", +} + +// MiddlewareRef renders the value Traefik's Kubernetes Ingress provider expects +// in the router.middlewares annotation: -@kubernetescrd. +// +// The namespace is part of the reference because a Middleware is namespaced and +// cross-namespace references are refused unless the controller was started with +// allowCrossNamespace — which it is not. Hence one Middleware per user +// namespace rather than one shared object. +func MiddlewareRef(ns, name string) string { + return fmt.Sprintf("%s-%s@kubernetescrd", ns, name) +} + +// EnsureStripForwardedHostMiddleware creates the Middleware in ns. Idempotent — +// AlreadyExists is ignored, matching the other Create* helpers. +// +// An empty header value is Traefik's spelling of "delete this header", not "send +// it empty". +func (c *Client) EnsureStripForwardedHostMiddleware(ctx context.Context, ns string) error { + mw := &unstructured.Unstructured{Object: map[string]any{ + "apiVersion": "traefik.io/v1alpha1", + "kind": "Middleware", + "metadata": map[string]any{ + "name": StripForwardedHostMiddleware, + "namespace": ns, + }, + "spec": map[string]any{ + "headers": map[string]any{ + "customRequestHeaders": map[string]any{ + "X-Forwarded-Host": "", + }, + }, + }, + }} + + _, err := c.dyn.Resource(middlewareGVR).Namespace(ns).Create(ctx, mw, metav1.CreateOptions{}) + if err != nil && !k8serrors.IsAlreadyExists(err) { + return fmt.Errorf("create middleware %s: %w", StripForwardedHostMiddleware, err) + } + return nil +} diff --git a/internal/k8s/vta_resources.go b/internal/k8s/vta_resources.go index 0430396..750480d 100644 --- a/internal/k8s/vta_resources.go +++ b/internal/k8s/vta_resources.go @@ -126,23 +126,22 @@ func (c *Client) CreateVtaService(ctx context.Context, ns string, sessionID uint return nil } -// CreateVtaIngress creates an nginx Ingress routing the session FQDN to the VTA service. -// TLS is handled by the cluster-wide wildcard default-ssl-certificate on nginx-ingress; -// no tls: block is needed. Idempotent — AlreadyExists is ignored. +// CreateVtaIngress creates an Ingress routing the session FQDN to the VTA service. +// TLS is the cluster-wide wildcard, which Traefik serves as its default +// certificate, so no tls: block is needed — and the HTTP→HTTPS redirect is an +// entrypoint setting on the controller, so no annotation is either. +// Idempotent — AlreadyExists is ignored. func (c *Client) CreateVtaIngress(ctx context.Context, ns string, sessionID uint, fqdn string) error { name := vtaDeploymentName(sessionID) svcName := vtaServiceName(sessionID) port := intstr.FromInt32(8100) pathType := networkingv1.PathTypePrefix - ingressClass := "nginx" + ingressClass := IngressClass _, err := c.kube.NetworkingV1().Ingresses(ns).Create(ctx, &networkingv1.Ingress{ ObjectMeta: metav1.ObjectMeta{ Name: name, Namespace: ns, - Annotations: map[string]string{ - "nginx.ingress.kubernetes.io/ssl-redirect": "true", - }, }, Spec: networkingv1.IngressSpec{ IngressClassName: &ingressClass, diff --git a/internal/setup/orchestrator_fullstack.go b/internal/setup/orchestrator_fullstack.go index 20b8729..8b0265d 100644 --- a/internal/setup/orchestrator_fullstack.go +++ b/internal/setup/orchestrator_fullstack.go @@ -368,12 +368,18 @@ func (o *Orchestrator) fsK8sProvision(ctx context.Context, ns string, s *model.S name, fqdn string port int32 labels map[string]string + // stripXFH removes X-Forwarded-Host on the way to this component. Only + // the dids daemon reads that header — and warn-logs every request + // carrying one it is not configured to trust, which is every request + // once it is behind an ingress at all. See + // k8s.StripForwardedHostMiddleware. + stripXFH bool } svcs := []svc{ - {k8s.FSVtaName(s.ID), s.FQDN(), 8100, fsLabels("vta", s.ID)}, - {k8s.FSMediatorName(s.ID), s.MediatorFQDN(), 7037, fsLabels("mediator", s.ID)}, - {k8s.FSDidsName(s.ID), s.DidsFQDN(), 8534, fsLabels("dids", s.ID)}, - {k8s.FSVtcName(s.ID), s.VtcFQDN(), 8200, fsLabels("vtc", s.ID)}, + {k8s.FSVtaName(s.ID), s.FQDN(), 8100, fsLabels("vta", s.ID), false}, + {k8s.FSMediatorName(s.ID), s.MediatorFQDN(), 7037, fsLabels("mediator", s.ID), false}, + {k8s.FSDidsName(s.ID), s.DidsFQDN(), 8534, fsLabels("dids", s.ID), true}, + {k8s.FSVtcName(s.ID), s.VtcFQDN(), 8200, fsLabels("vtc", s.ID), false}, } // Only a custom domain needs a certificate of its own. Managed and platform @@ -391,17 +397,25 @@ func (o *Orchestrator) fsK8sProvision(ctx context.Context, ns string, s *model.S } } + // Before the Ingresses, not with them: the dids Ingress names this object, + // and Traefik fails a router that references a Middleware which isn't there + // yet. One per namespace, shared by every session the user owns. + if err := o.k8s.EnsureStripForwardedHostMiddleware(ctx, ns); err != nil { + return err + } + for _, sv := range svcs { if err := o.k8s.CreateComponentService(ctx, ns, sv.name, sv.labels, sv.port); err != nil { return err } if err := o.k8s.CreateComponentIngress(ctx, k8s.ComponentIngressSpec{ - Namespace: ns, - Name: sv.name, - ServiceName: sv.name, - Host: sv.fqdn, - Port: sv.port, - TLSSecret: tlsSecret, + Namespace: ns, + Name: sv.name, + ServiceName: sv.name, + Host: sv.fqdn, + Port: sv.port, + TLSSecret: tlsSecret, + StripForwardedHost: sv.stripXFH, }); err != nil { return err } diff --git a/k8s/secret.yaml.example b/k8s/secret.yaml.example index f058a18..c2be4c8 100644 --- a/k8s/secret.yaml.example +++ b/k8s/secret.yaml.example @@ -9,7 +9,7 @@ stringData: JWT_SECRET: "replace-with-strong-secret" # ── Cluster ────────────────────────────────────────────────────────────────── - # External IP of the cluster's Ingress-NGINX LoadBalancer + # External IP of the cluster's Traefik LoadBalancer CLUSTER_INGRESS_IP: "" # ── Cloudflare ─────────────────────────────────────────────────────────────── diff --git a/k8s/tls/certificate.yaml b/k8s/tls/certificate.yaml index 4f4f00a..c564bb4 100644 --- a/k8s/tls/certificate.yaml +++ b/k8s/tls/certificate.yaml @@ -1,8 +1,23 @@ +# The wildcard every managed and platform hostname is served with. +# +# kubectl apply -f k8s/tls/certificate.yaml +# +# ── Why kube-system ────────────────────────────────────────────────────────── +# +# Traefik serves this as its default certificate through the TLSStore named +# `default` (k8s/tls/tlsstore-default.yaml), and a TLSStore can only reference a +# Secret in its own namespace — which must be Traefik's. Under ingress-nginx the +# certificate lived in `cert-manager` and the controller was pointed at it by +# flag across namespaces; that indirection is gone. +# +# Moving it is one issuance against the "5 certificates per identical name set +# per week" limit. If you ever move it again, remember the old Secret keeps +# renewing until its Certificate object is deleted. apiVersion: cert-manager.io/v1 kind: Certificate metadata: name: firstperson-dev-wildcard - namespace: cert-manager + namespace: kube-system spec: secretName: firstperson-dev-wildcard-tls issuerRef: diff --git a/k8s/tls/clusterissuer-http01.yaml b/k8s/tls/clusterissuer-http01.yaml index cdc61e4..65401dd 100644 --- a/k8s/tls/clusterissuer-http01.yaml +++ b/k8s/tls/clusterissuer-http01.yaml @@ -59,4 +59,11 @@ spec: solvers: - http01: ingress: - class: nginx + # cert-manager creates a solver Ingress per challenge; this is the + # class it stamps on it. Wrong value = the challenge is served by + # nobody and tls_provision times out with no useful error. + # + # The entrypoint-level HTTP→HTTPS redirect does not break this: Let's + # Encrypt follows the redirect and does not validate the certificate + # it lands on. + ingressClassName: traefik diff --git a/k8s/tls/rke2-ingress-nginx-config.yaml b/k8s/tls/rke2-ingress-nginx-config.yaml deleted file mode 100644 index 577c4b3..0000000 --- a/k8s/tls/rke2-ingress-nginx-config.yaml +++ /dev/null @@ -1,10 +0,0 @@ -apiVersion: helm.cattle.io/v1 -kind: HelmChartConfig -metadata: - name: rke2-ingress-nginx - namespace: kube-system -spec: - valuesContent: |- - controller: - extraArgs: - default-ssl-certificate: cert-manager/firstperson-dev-wildcard-tls diff --git a/k8s/tls/rke2-traefik-config.yaml b/k8s/tls/rke2-traefik-config.yaml new file mode 100644 index 0000000..c8b9182 --- /dev/null +++ b/k8s/tls/rke2-traefik-config.yaml @@ -0,0 +1,84 @@ +# Ingress controller configuration — the settings that used to be per-Ingress +# annotations under ingress-nginx. +# +# kubectl apply -f k8s/tls/rke2-traefik-config.yaml +# +# RKE2 reconciles the change and restarts Traefik on its own. If Rancher manages +# this cluster it also writes to this object (it injects global.cattle.clusterId) +# — the values merge, but check the object after applying rather than assuming. +# +# Installing Traefik from the upstream chart instead? The same keys go straight +# into values.yaml; only the HelmChartConfig wrapper is RKE2-specific. +# +# **Mind the `http:` level.** Helm merges keys it does not recognise without +# complaining, so a values path that is one level off is accepted, produces no +# argument, and looks exactly like a working config. Both entrypoint settings +# below live under `http:` — verify against the rendered args after applying, +# never against this file: +# +# kubectl -n kube-system get ds rke2-traefik \ +# -o jsonpath='{.spec.template.spec.containers[0].args}' | tr ',' '\n' +# +# You want to see --entryPoints.web.http.redirections.entryPoint.to=websecure +# and --entryPoints.websecure.http.tls=true. (The TLS one is also the chart's +# default, so its presence proves nothing about whether this file applied — the +# redirect is the one to look for.) +# +# ── What each block is load-bearing for ────────────────────────────────────── +# +# websecure.tls.enabled Routers built from an Ingress with no tls: block get +# TLS from the entrypoint, and their certificate from the +# `default` TLSStore. This is what replaces +# --default-ssl-certificate. **Without it every managed +# and platform hostname serves plain HTTP on :443 only — +# i.e. no HTTPS at all — and the full_stack pipeline dies +# at step_vta_register_dids.** +# +# web.redirections Replaces nginx.ingress.kubernetes.io/ssl-redirect on +# every Ingress, which is why the ones this API creates +# now carry no annotations. cert-manager's HTTP-01 +# challenge is redirected too; Let's Encrypt follows the +# redirect and does not validate the certificate it lands +# on, so custom-domain issuance still works. +# +# kubernetesCRD Middleware is a CRD. allowCrossNamespace stays false — +# vtafarm-api creates one Middleware per user namespace +# precisely so it never needs the cluster-wide grant. +# +# publishedService Writes the published Service's external address into +# each Ingress's status. Note this leaves ADDRESS **empty** +# in the RKE2 layout, where Traefik takes traffic on +# hostPort 80/443 and its Service is ClusterIP — there is +# no external address to publish. Cosmetic either way, but +# it means an empty ADDRESS is NOT evidence that something +# is broken; check the CLASS column and curl instead. +apiVersion: helm.cattle.io/v1 +kind: HelmChartConfig +metadata: + name: rke2-traefik + namespace: kube-system +spec: + valuesContent: |- + ingressClass: + enabled: true + isDefaultClass: true + ports: + web: + http: + redirections: + entryPoint: + to: websecure + scheme: https + permanent: true + websecure: + http: + tls: + enabled: true + providers: + kubernetesCRD: + enabled: true + allowCrossNamespace: false + kubernetesIngress: + enabled: true + publishedService: + enabled: true diff --git a/k8s/tls/tlsstore-default.yaml b/k8s/tls/tlsstore-default.yaml new file mode 100644 index 0000000..e6d974d --- /dev/null +++ b/k8s/tls/tlsstore-default.yaml @@ -0,0 +1,29 @@ +# Traefik's default certificate — what every hostname under our own zone is +# served with. +# +# kubectl apply -f k8s/tls/tlsstore-default.yaml +# +# This is the direct replacement for ingress-nginx's --default-ssl-certificate +# flag, and it is what lets every Ingress this API creates carry no tls: block: +# the router has no certificate of its own, so Traefik falls back to this one. +# Custom domains are the exception — their Ingress names a Secret, and a +# router-level certificate wins over the store. +# +# Two constraints, both easy to get wrong: +# +# - The name MUST be `default`. Traefik uses no other store for this. +# - The namespace MUST be Traefik's own (kube-system for the RKE2 bundle), and +# the Secret has to live in that same namespace — hence +# k8s/tls/certificate.yaml issuing there. A store in the wrong namespace +# silently does nothing and every host falls back to Traefik's self-signed +# certificate, which fails the pipeline several steps later: from +# step_vta_register_dids on, the components resolve each other's did:webvh +# over HTTPS and reject an untrusted chain. +apiVersion: traefik.io/v1alpha1 +kind: TLSStore +metadata: + name: default + namespace: kube-system +spec: + defaultCertificate: + secretName: firstperson-dev-wildcard-tls diff --git a/scripts/migrate-ingress-to-traefik.sh b/scripts/migrate-ingress-to-traefik.sh new file mode 100755 index 0000000..cb757fc --- /dev/null +++ b/scripts/migrate-ingress-to-traefik.sh @@ -0,0 +1,165 @@ +#!/bin/bash +set -euo pipefail + +# Move every Ingress this API created off ingress-nginx and onto Traefik. +# +# ./scripts/migrate-ingress-to-traefik.sh # dry run — prints the plan +# ./scripts/migrate-ingress-to-traefik.sh --apply # do it +# +# KUBE_CONTEXT kubectl context to use (default: the current context) +# +# ── Why this has to exist ──────────────────────────────────────────────────── +# +# vtafarm-api only ever *creates* Ingresses — AlreadyExists is ignored and +# nothing updates them afterwards. So every session provisioned before the +# switch keeps `ingressClassName: nginx` forever, Traefik never adopts it, and +# the hostname 404s. Changing the Go constant fixes new sessions only; this +# fixes the ones already out there. +# +# It does three things per user namespace (label managed-by=vtafarm): +# +# 1. repoints every Ingress at the traefik class, +# 2. drops the ingress-nginx annotations, now that the HTTP→HTTPS redirect is +# an entrypoint setting, +# 3. for namespaces running a dids daemon: creates the strip-forwarded-host +# Middleware and references it from that daemon's Ingress. +# +# What it deliberately leaves alone: vtafarm-api's own Ingress (Helm-managed — +# `helm upgrade` carries it), and anything not labelled as ours. +# +# Safe to re-run. + +CTX="${KUBE_CONTEXT:-$(kubectl config current-context)}" +APPLY=false +[ "${1:-}" = "--apply" ] && APPLY=true + +# Keep in step with internal/k8s/traefik.go — the Go side creates this same +# object for every new session, under this same name. +MIDDLEWARE="strip-forwarded-host" +CLASS="traefik" + +k() { kubectl --context "$CTX" "$@"; } + +# run prints the command, and only executes it with --apply. +run() { + if $APPLY; then + printf ' → %s\n' "$*" + "$@" + else + printf ' would run: %s\n' "$*" + fi +} + +echo "context: $CTX" +$APPLY || echo "DRY RUN — nothing will be changed. Re-run with --apply." +echo + +# ── Preflight ──────────────────────────────────────────────────────────────── +# Both of these are cluster-side prerequisites (k8s/tls/rke2-traefik-config.yaml). +# Failing here is much cheaper than half-migrating and finding out at step 3. +fail=0 +if ! k get crd middlewares.traefik.io >/dev/null 2>&1; then + echo "ERROR: CRD middlewares.traefik.io not found — is Traefik installed with providers.kubernetesCRD enabled?" >&2 + fail=1 +fi +if ! k get ingressclass "$CLASS" >/dev/null 2>&1; then + echo "ERROR: IngressClass '$CLASS' not found — check ingressClass.enabled in the Traefik values." >&2 + fail=1 +fi +[ "$fail" -eq 0 ] || exit 1 + +namespaces=$(k get ns -l managed-by=vtafarm -o jsonpath='{range .items[*]}{.metadata.name}{"\n"}{end}') +if [ -z "$namespaces" ]; then + echo "No namespaces labelled managed-by=vtafarm. Nothing to do." + exit 0 +fi + +total=0 +for ns in $namespaces; do + ingresses=$(k get ingress -n "$ns" -o jsonpath='{range .items[*]}{.metadata.name}{"\n"}{end}' 2>/dev/null || true) + [ -n "$ingresses" ] || continue + + echo "namespace $ns" + + # The Middleware first: Traefik fails a router that references one which + # isn't there yet, so creating it after the annotation would break the daemon + # for as long as the gap lasts. + if echo "$ingresses" | grep -q -- '-dids$'; then + if $APPLY; then + echo " → create Middleware $MIDDLEWARE" + k apply -f - >/dev/null <:443: https:// + kubectl logs deploy/ -n # the forwarded-host warnings should have stopped +EOF +exit 0