Skip to content

Add finkctl get balance, reporting what entered and left the broker - #6

Merged
fjammes merged 8 commits into
mainfrom
1216-run-fink-broker-on-ztf-data-for-one-week-at-cc-in2p3
Aug 26, 2026
Merged

fjammes merged 8 commits into
mainfrom
1216-run-fink-broker-on-ztf-data-for-one-week-at-cc-in2p3

Conversation

@fjammes

@fjammes fjammes commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

What

Adds finkctl get balance, which reports what entered and what left the broker
for each observing night, plus the packaging needed to run it as a container.

Balance per observing night -- prefix /user/185

     NIGHT  IN(kafka)  RAW(f)       RAW  SCI(f)       SCI  DISTRIB
  20260824       9984      45  575.9MiB      11  138.8MiB     3015
     TOTAL       9984                                         3015

Distribution per filter topic, per run window

  20260824  [2026-08-25T09:00Z -> 2026-08-26T09:00Z]
    fink_sso_ztf_candidates_ztf                               130
    fink_unknowns_ztf                                         602
    fink_vra_ztf                                               24
    ...

Why

The permanent daily ZTF run at CC-IN2P3 (fink-broker#1216) had no way to answer
"did last night actually go through the broker". Counts are read from the
sources of truth, so no Spark job is involved: consumed alerts come from the
last committed batch of the stream2raw checkpoint, distributed alerts from the
Kafka end offsets of the output topics.

Commits

  • 36b2851 make the pod exec logic reusable and cluster-aware
  • 9145459 add get balance
  • 0f6d5f4 publish finkctl as a container image
  • 25f4357 publish the finkctl image from CI
  • 4f989a3 gate the image push on the unit tests
  • 09442ce document the balance columns in the command help

Validation

Run against the live CC-IN2P3 cluster on 2026-08-25, against the night
20260824 while the three streaming jobs were running — output above.
go build ./..., go test ./cmd/... and gofmt are clean.

After the merge

fink-broker pins finkctl in two places, and both need the new tag before
report.enabled can be turned on — until then the report CronJob calls a
subcommand its image does not have:

  • chart/values.yaml, report.image.tag (currently v3.1.3-rc4)
  • .ciux, the finkctl/v3@... package pin

Note for whoever cuts the tag: v3.2.0-rc0 already exists and dates from
2025-01-16, so it is older than v3.1.3-rc4 (2026-07-06) despite the higher
number.

Context: astrolabsoftware/fink-broker#1216

The SPDY exec logic was inlined in getKafkaTopics, and setKubeClient only
ever read a kubeconfig file. Both stand in the way of running finkctl as a
pod, and of execing anywhere but the Kafka broker.

Extract execInPod, resolve pod names from the operator labels with a
fallback to the well-known ones, and fall back to the in-cluster service
account when no kubeconfig is around. Kafka and HDFS access details move
to constants.
Nothing reported what the broker ingested and redistributed: merge.py and
archive_statistics.py are not part of the Kubernetes deployment, and no
metrics are exposed.

Report it per observing night, from the sources of truth, without running
a Spark job:
  - alerts consumed come from the offsets of the last committed batch of
    the stream2raw checkpoint. The offsets a batch is planned to reach are
    written before it runs, so an uncommitted batch must not be counted;
  - raw and science hold the datasets written on both sides of the quality
    cuts;
  - distributed alerts come from the per-night counting topic, while the
    per-filter topics carry no night suffix and are bracketed by
    offset-by-timestamp lookups over the run window.

Bounds that no message reaches fall back to the high watermark, and a
window starting before the oldest retained message is flagged as an
undercount rather than reported as exact.
finkctl had no image, so it could not be run in the cluster. The report
CronJob of the fink-broker chart needs one.

finkctl only talks to the Kubernetes API, so a static binary on distroless
is enough. build.sh and push-image.sh mirror the fink-broker ones, and the
source pathes let ciux tag the image from the code it contains.
The image the fink-broker report CronJob runs was never published: the
Dockerfile and the build/push scripts existed, but no workflow ever ran
them, so the CronJob would land in ImagePullBackOff.

Add a 'push' job mirroring fink-broker's: ciux ignition, then build.sh
and push-image.sh, gated on the e2e job. Registry login comes first so
that ciux and build.sh see the same registry state when deciding whether
the image must be rebuilt.

The workflow could not run at all anyway: ubuntu-20.04 runners are
retired, and the last runs sat queued until they were cancelled. Move to
ubuntu-latest, refresh the actions, and align CIUX_VERSION with the one
fink-broker pins.
_e2e/e2e.sh drives 'finkctl run raw2science|stream2raw', subcommands
that no longer exist: they were dropped when fink-broker moved to
Helm-deployed SparkApplications, and 'finkctl run' now swallows them as
positional arguments, returning 0 without output. The suite has not run
since 2025 (ubuntu-20.04 runners), so nobody noticed.

Gate the image on 'go test ./...' instead, which covers get balance and
the pod exec logic the report CronJob relies on. The e2e job is left
untouched and failing: rewriting or removing it is a separate decision.
The table header is terse enough to be read wrong. Two confusions are worth
preventing: IN(kafka) looks like it comes from the offsets directory when it
deliberately does not, and the RAW(f)/SCI(f) pair invites a comparison that
means nothing -- they count micro-batch files, not alerts, so fewer science
files than raw ones is not a loss.

Spell out what each column measures and which pair actually balances.
Every assertion in _e2e/e2e.sh drove 'finkctl run stream2raw|raw2science|
distribution', subcommands dropped when fink-broker moved to Helm-deployed
SparkApplications. 'run' now takes no subcommand at all: cobra ignores the
argument, logs the configuration and exits 0, so the suite's first check --
which asserts that a malformed -N is rejected -- can never pass.

The suite was already broken beyond that: 2a7d068 deleted three of the four
.out.expected references without touching the assertions comparing against
them, so the run has been red since then.

Remove the script, its orphaned reference file and the job running it. The
image push was already gated on the unit tests alone, so nothing that was
actually protecting a release is lost; get_balance_test.go and the rest of
cmd/ keep that role. _e2e/finkctl.yaml.orig and finkctl.secret.yaml stay:
the README points at them as the configuration example.
Four mechanisms that repository has and this one lacked:

- fetch-tags plus an explicit git fetch --tags --force. The image tag comes
  from git describe, so a missing tag does not fail, it silently names the
  image after an older version.
- the build step now exports image-url and new-build, read back after
  build.sh since it runs its own ciux check in a subshell. Reading them
  before the push matters: afterwards the registry would answer that the
  image exists and new-build would flip.
- a job summary carrying the image URL and its docker pull line, whether the
  image was built or reused.
- a latest tag on main, for consumers that do not pin a version, and a Trivy
  scan of newly built images uploaded as SARIF.

Not ported: the workflow_call plumbing, which serves the seven image variants
built over there and has nothing to parameterise here, and the disk-space
reclaim steps, which a small Go image does not need. The login stays
unconditional because ciux probes the registry before anything else runs.
@fjammes
fjammes merged commit 37cb20d into main Aug 26, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant