Skip to content

Promote Keycloak theme to production (first deploy from main) - #20

Merged
saqibmanan merged 28 commits into
mainfrom
dev
Sep 8, 2026
Merged

saqibmanan merged 28 commits into
mainfrom
dev

Conversation

@saqibmanan

Copy link
Copy Markdown
Contributor

First production promotion under the new rule from #19. Merging this deploys to auth.civicdatalab.in.

Risk: low, and worth being precise about why

The box is deployed from dev today. Every one of these 26 commits is already running in production. Promoting them to main ships the same code, so this is a labelling and governance change rather than a code change.

What actually changes is where deploys come from: after this, pushing to dev no longer deploys, and main becomes the deploying branch. Verified already — merging #19 into dev fired no deploy.

What lands

26 commits, including the work from the migration:

Also fixes workflow_dispatch

deploy-keycloak-staging.yml has never been manually runnable, because GitHub only surfaces workflow_dispatch for workflows on the default branch and this file lived only on dev. Landing it on main makes the manual trigger usable, including the ref input.

Same defect exists in two other repos — tracked in CivicDataSpace-test#30.

Expected side effect

This is the first genuinely new run of the fixed keycloak-tests.yml (CivicDataSpace-test#29), so it should report 13 passed / 0 skipped, up from 12/1. A job re-run does not pick up an updated reusable workflow — confirmed today, a re-run still showed 12/1.

If it reports anything else, #29's diagnosis was wrong and should be reverted rather than patched.

Names deliberately left as "staging"

Container names, the kc_postgres_data volume, DEPLOY_PATH, and the keycloak-staging GitHub environment all keep their names. Each carries a comment saying why — renaming any of them is an auth outage for every product, and the environment specifically holds the deploy's SSH secrets, which GitHub cannot carry across a rename.

UdayRajSahai2 and others added 28 commits August 7, 2026 14:54
The lockfile was left inconsistent when @tabler/icons-react was added (it
looks to have been installed with yarn), so `npm ci` fails outright:
"Missing: @tabler/icons-react from lock file" plus a dozen version
mismatches. Regenerated with npm so reproducible installs work in CI.
Keeps node_modules, build output and .git out of the build context.
Multi-stage: node+JDK+Maven builds the Keycloakify jar, then it is copied
into quay.io/keycloak/keycloak:24.0.0 and registered with `kc.sh build`.

Three things this settles:

- The jar is copied *by name*. Keycloakify emits one jar per Keycloak
  version range; a glob would put several themes into providers/ and let
  Keycloak pick one.
- `kc.sh build` is required. Providers register at build time, so a jar
  dropped into providers/ on a running server does nothing. Staging has no
  build step today, which is why it re-augments for ~16s on every start.
- KC_DB / KC_HEALTH_ENABLED / KC_METRICS_ENABLED / KC_HTTP_RELATIVE_PATH
  are build-time options in KC 24. Setting them only in compose (as staging
  does) triggers that re-augmentation regardless. Baked in here to match.

The image carries the source commit and the theme jar's sha256 as labels,
so `docker inspect` answers which theme a box is running.
Builds the image in CI, pushes it to GHCR, and has the staging box pull it
-- replacing the manual build-jar/scp/rebuild-on-the-box process that on
2026-08-27 left staging running a seven-month-stale theme copied from a
retired server, with nothing logging an error.

Mirrors deploy-backend.yml in DataSpaceBackend: digest-pinned image,
health gate, automatic rollback to the previously-running image.

workflow_dispatch only for now, taking a ref input. push-to-main is left
disabled because main does not yet carry the theme work on
fix/login-ui-alignment, so auto-deploying it would regress staging.
The deploy job's own GITHUB_TOKEN can authenticate the box's `docker
compose pull`, and it expires with the run. That drops the GHCR_TOKEN
secret entirely and leaves no standing registry credential on the auth
server; a trap logs out on every exit path.
github.repository is CivicDataLab/DataSpaceKeycloakTheme; Docker rejects
uppercase in image references, so the digest-pinned image_ref would have
been invalid on the box.
dev becomes the integration branch for staging -- what is merged there is
what staging runs. A push trigger also means the workflow does not need to
sit on the default branch first, which workflow_dispatch would require.

workflow_dispatch is kept for deploying an arbitrary ref on demand.
Custom login theme, social provider icons, password wrapper, field error
icons, and the alert/social-divider alignment fixes.
ci: automate Keycloak theme deploys to staging
`keycloakify sync-extensions` shells out to `git ls-files`, and the
@keycloakify/email-native files under src/email/ are generated rather than
tracked, so it has real work to do on a clean checkout and dies with
"fatal: not a git repository".

This only passed locally because a host-side `npm install` had already
run the postinstall and `COPY . .` picked up the generated files.
Reproduced against a fresh clone of dev before fixing.
The first staging deploy still paid a 15.4s Quarkus re-augmentation on
start. Cause: KC_PROXY=edge and the `--http-relative-path=/auth` CLI flag
are build-time options in KC 24, and the box was passing both at runtime,
which overrides the baked build.

KC_PROXY is baked here; the box's compose now drops the build-time options
from `environment:` and starts with `--optimized`, which refuses to
re-augment silently rather than doing it at a cost of 15s per start.
Temporary. Builds a valid image that serves on /auth-rollback-test, so the
deploy's health probe against /auth/health/ready must fail and roll staging
back to the previous image. Reverted in the next commit.
…h gate"

Gate proven: the broken image built and pushed fine, the deploy's health
probe against /auth/health/ready failed for 3m, and staging was restored to
the previous image (commit 2d0b26c) with health back at 200. The run was
marked failed, as it should be.

This reverts commit 4998907.
Until now the box's ~/keycloak/docker-compose.yml was the one load-bearing
artifact with no home in git -- hand-edited over SSH, unreviewable, and
recoverable only from a backup directory sitting on the same box.

That is the same failure shape as the 2026-08-27 incident: a running server
silently diverging from what anyone believes is deployed. Adding `build: .`
back here would be invisible in exactly the same way.

Captured as-is; no behaviour change. The KC_BOOTSTRAP_ADMIN_* trap (a
Keycloak 26 name, silently ignored on 24) is now at least documented in
place rather than lurking on a box nobody reads.
scp the version-controlled compose file up, validate it with `docker
compose config` before swapping it in, and keep the live one as .prev so a
bad compose change can be rolled back as well as a bad image -- rollback
runs `docker compose up`, so it needs a compose file it can parse.

The box now holds nothing hand-written: only .env, which is never
committed.
Themes built with Keycloakify releases predating Keycloak 26 are
incompatible with Keycloak 26, so this is a prerequisite for the server
upgrade. package.json already allowed it (^11.8.42) -- the lockfile was
the pin, and npm install will not move an entry that already satisfies
the range, so this needed npm update.

Deployed on Keycloak 24 first, deliberately: if the login page breaks it
is this change, not the server upgrade. The image rollback still fully
works at this point because the database has not moved.
24.0.0 shipped February 2024. Keycloak has no LTS -- only the newest
release gets security fixes -- so staging has been unpatched for ~2 years.

Config changes forced by the jump:
- KC_PROXY=edge -> KC_PROXY_HEADERS=xforwarded (proxy removed in 26).
  nginx already sends the X-Forwarded-* headers this needs.
- hostname v1 removed: KC_HOSTNAME becomes the full URL
  https://auth.civicdatalab.in/auth, which is what makes HOSTNAME_STRICT
  safe behind a TLS-terminating proxy. KC_HOSTNAME_STRICT_HTTPS dropped.
- Health/metrics moved to management port 9000 in KC 25. Published on
  127.0.0.1:9004 and the deploy gate now probes it. Without publishing it
  a healthy deploy would look like a failed one and trigger a rollback.
- KC_HTTP_MANAGEMENT_RELATIVE_PATH pinned to / so the probe URL is
  explicit; unset, it silently inherits http-relative-path.
- KC_BOOTSTRAP_ADMIN_* are the correct names from 26 onward. They were
  silently ignored on 24, which cost an admin recovery on 2026-08-26.

Also fixes, incidentally: the jar we already ship
(keycloak-theme-for-kc-all-other-versions.jar) is the KC 26+ one. On 24
the correct jar was keycloak-theme-for-kc-22-to-25.jar, which bundles the
password-policy extension 22-25 lack natively -- so staging and prod have
both been running the wrong jar. 26 makes the shipped one correct.

Rehearsed end to end before pushing: staging's real database dumped,
restored locally, migrated by this exact image. Liquibase completed, 87
users and 7 clients intact, both realm themes preserved, issuer resolves
to https://auth.civicdatalab.in/auth/realms/DataSpace, login page renders
the theme, zero Quarkus augmentation.
URLs become https://auth.civicdatalab.in/realms/... instead of
https://auth.civicdatalab.in/auth/realms/...

Doing it now, deliberately: this changes the OIDC issuer, so every client
has to move at the same time. Nothing consumes staging yet -- dev
still authenticates against production Keycloak -- so the blast radius is
zero today and would be every client the moment dev is pointed here.

Dropping KC_HTTP_RELATIVE_PATH leaves it at the default /.
KC_HTTP_MANAGEMENT_RELATIVE_PATH stays pinned so the deploy's health probe
on :9000 cannot drift with it.

No database change needed: the account, account-console and
security-admin-console clients store root_url as ${authBaseUrl} with
paths already relative to the root, so they follow automatically.

nginx needs its 'location = / -> 301 /auth/' redirect removed; it does no
other path rewriting.
Fix Register Template headerNode type error
The theme ships independently of the app, so a theme change can break
sign-in while every product pipeline stays green. The deploy's existing
gate only proves Keycloak booted (/health/ready) — it says nothing about
whether a user can still reach the login page, whether "Continue With
Google" renders, or whether the privacy links this theme builds resolve.

Calls CivicDataSpace-test's keycloak-tests workflow, which runs only the
Keycloak-relevant tests rather than the full ~15 minute smoke suite.

Uses workflow_call rather than repository_dispatch: dispatching across
repositories needs a PAT with `repo` scope stored here, and this needs no
new secret. The listener accepts repository_dispatch too if a token is
ever preferred.

Deliberately does NOT gate the deploy. By the time this runs the theme is
already live, and Keycloak cannot be safely auto-rolled-back once a major
version has migrated the schema — the deploy's own rollback notes this. A
red run here is a signal to look, not an automatic revert.

KEYCLOAK_CLIENT_SECRET is passed but optional: present, the registration
tests also run; absent, they skip, because they create real accounts and
must be able to delete them.

Requires CivicDataSpace-test#28 to be merged to CI first.
…deploy

ci: run the Keycloak auth tests after a staging deploy
auth.civicdatalab.in is no longer a staging server. It is the auth
server for CivicDataSpace (dev and prod), ParakhAI (dev and prod),
Analytics (dev and prod) and the DRR-hosted DataSpace. An outage here
signs every user out of every product, so it should not redeploy on
every push to an integration branch.

Deploys now trigger on `main`. Work lands on `dev`, is reviewed, and
reaches users only when merged to `main`. Pushing to `dev` no longer
deploys.

This also fixes workflow_dispatch, which has never worked here: GitHub
only surfaces it for workflows present on the default branch, and this
file existed only on `dev`. Once it lands on `main` the manual trigger
becomes usable, including the `ref` input. The same defect exists in two
other repos - see CivicDataSpace-test#30.

Deliberately NOT renamed, despite the promotion:

- CONTAINER_NAME (keycloak-staging) and the compose file's container
  names. Changing them makes compose create a second container and
  orphan the running one.
- The kc_postgres_data volume. A renamed volume gives Keycloak an empty
  database - every realm, client and user gone.
- DEPLOY_PATH and the directory on the box.
- The `keycloak-staging` GitHub environment. EC2_HOST, EC2_USERNAME and
  EC2_PRIVATE_KEY are scoped to it, and GitHub cannot rename an
  environment while preserving secrets, so renaming would strip the
  deploy's SSH key.

Those names read as "staging" for historical reasons. That is cosmetic;
renaming any of them is an auth outage for every product. Each now
carries a comment saying so, so the next person does not tidy them.

Promoting dev to main deploys the same code that is already running -
the box is deployed from dev today - so this is a labelling and
governance change, not a code change.
…ction

ci: treat auth.civicdatalab.in as production and deploy from main
@saqibmanan
saqibmanan merged commit c9be852 into main Sep 8, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants