Skip to content

Read HDFS natively in get balance, drop the namenode pod exec - #8

Merged
fjammes merged 3 commits into
mainfrom
1241-use-the-cc-in2p3-production-hdfs-outside-kubernetes-as-fink-broker-storage-at-cc
Sep 30, 2026
Merged

fjammes merged 3 commits into
mainfrom
1241-use-the-cc-in2p3-production-hdfs-outside-kubernetes-as-fink-broker-storage-at-cc

Conversation

@fjammes

@fjammes fjammes commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Summary

finkctl get balance now reads HDFS with a native Go client (github.com/colinmarc/hdfs/v2) instead of exec-ing hdfs dfs into the Stackable namenode pod. Part of astrolabsoftware/fink-broker#1241, where the broker moves to an HDFS outside Kubernetes and there is no namenode pod to exec into.

  • --namenode host:port[,host:port] is required. List every namenode of an HA pair: the client follows the active one. --hdfs-user defaults to $HADOOP_USER_NAME, else 185.
  • DataNodes are reached by IP, since their host names often do not resolve outside the HDFS cluster. Every dial is bounded (10 s), so an unreachable HDFS fails the report instead of hanging it.
  • Before the first run, <prefix>/raw does not exist: the command now prints No observing night under <prefix>/raw yet and succeeds. Before, the fink-broker report CronJob failed on every fresh deployment.
  • The pod exec reader, the parsing of the hdfs dfs text output and the HDFS pod constants are removed. The image stays a static binary on distroless.
  • The report logic is unit tested against an in-memory HDFS.

Kafka is still read by exec into the broker pod; see #7.

Validation

  • go vet, go test ./...
  • On the CC cluster, in-cluster Stackable HDFS (HA), from a pod with the fink-report ServiceAccount:
    • native mode vs exec mode on the existing data (~30 nights): identical reports (diff empty), 281 s vs 512 s;
    • empty HDFS: No observing night under /user/185/raw yet, rc=0 (the exec version failed there);
    • unreachable NameNode: fails with an error.

Follow-up in fink-broker

Pin the released image in chart/values.yaml (report.image.tag) and in .ciux. The chart now passes --namenode and no longer grants the report any right in the hdfs namespace.

🤖 Generated with Claude Code

get balance read HDFS by exec-ing hdfs dfs into the Stackable namenode pod.
With an HDFS outside the cluster (fink-broker#1241) there is no such pod.

The HDFS reads (list, count, cat) now go through an hdfsReader interface:
- podReader keeps the exec into the namenode pod, still the default, used
  by CI and when finkctl runs outside the cluster;
- nativeReader, selected by --namenode host:port[,host:port], talks to the
  NameNode RPC port and the DataNodes with github.com/colinmarc/hdfs, a pure
  Go client: the image stays a static binary. DataNodes are reached by IP,
  as their host names often do not resolve outside the HDFS cluster, and
  every dial is bounded so an unreachable HDFS fails instead of hanging.

--hdfs-user defaults to $HADOOP_USER_NAME, else 185 like the Spark
containers. The reporter logic is unchanged and now unit tested against an
in-memory reader.
On a fresh deployment, or any HDFS no run has written to yet, <prefix>/raw
does not exist: hdfs dfs -ls failed and get balance exited in error. The
daily report CronJob of fink-broker then failed every time, and kept
restarting until its backoffLimit.

A missing path now reads as os.ErrNotExist in both readers (the pod reader
recognises the "No such file or directory" message of hdfs dfs, the native
client already returns it), and a missing raw dataset lists no night: the
command says so and succeeds. Any other failure to list it, such as a denied
exec or an unreachable NameNode, still fails the report.
get balance read HDFS either by exec-ing hdfs dfs into the Stackable
namenode pod or, with --namenode, through the native client. The exec path
needed pods/exec rights in the hdfs namespace, tied the command to the
Stackable pod names and could not work with an HDFS outside the cluster.

--namenode is now required and the native client is the only reader. The
exec reader, the parsing of the hdfs dfs text output and the HDFS pod
constants are removed. The command must run where the NameNode and the
DataNodes are reachable, typically in a pod of the cluster.
@fjammes
fjammes merged commit 4c89062 into main Sep 30, 2026
7 checks passed
fjammes added a commit to astrolabsoftware/fink-broker that referenced this pull request Sep 30, 2026
v3.1.3-rc6 reads HDFS natively (--namenode) and reports an empty balance
before the first run instead of failing. It replaces the development image
the report was pinned on while astrolabsoftware/finkctl#8 was under review.
fjammes added a commit to astrolabsoftware/fink-broker that referenced this pull request Oct 5, 2026
v3.1.3-rc6 reads HDFS natively (--namenode) and reports an empty balance
before the first run instead of failing. It replaces the development image
the report was pinned on while astrolabsoftware/finkctl#8 was under review.
fjammes added a commit to astrolabsoftware/fink-broker that referenced this pull request Oct 6, 2026
v3.1.3-rc6 reads HDFS natively (--namenode) and reports an empty balance
before the first run instead of failing. It replaces the development image
the report was pinned on while astrolabsoftware/finkctl#8 was under review.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant