The single source of truth for the words we use — in the product UI, docs, website, decks, CLI, and support. One concept, one word.
This file is authoritative. All copy defers here — including the
communication,persona, andfounder-voiceskills, which reference this file and must not redefine terms. If a term isn't here, add it here first.Status: working draft — started 2026-07-11, being aligned over the coming days. Rows marked ✅ are decided; rows marked 🔵 OPEN await a team decision (see Open questions at the bottom).
We align the vocabulary here first. The "words to retire" section lists the drift terms currently live in the code/UI, recorded from a full cross-repo sweep for awareness — not as a signal to start renaming. The load-bearing question #1 (what we call the on-prem environment) is now decided — secure environment (2026-07-12); the remaining open questions still gate the broader rename. This file leads; the codebase follows later.
- Use the Preferred term. Avoid everything in Don't use.
- One concept = one word. If two words mean the same thing, pick one and retire the other.
- Applies everywhere customer- or team-facing: docs, app UI, website, decks, CLI copy, support replies, and code-facing copy (comments, log lines).
"Secure environment" — the quotable definition (restates RFC-CLI-0003 §1 in the terms below — house wording, not a verbatim quote): the tracebloc software running on infrastructure the owner controls (a laptop, a cloud cluster, or bare metal), with exactly three ingress channels, one egress channel, and nothing sideways — in: the data ingestor (raw datasets), the platform (models, weights, experiment instructions), and the signed release artifacts (digest-pinned images, chart, CLI); out: the platform (trained weights, metrics, status). Raw data is never an egress channel — it stays inside your infrastructure, and no inbound ports are required (the environment only dials out).
Wording discipline (external copy): say "defined, auditable ingress and egress — and raw data is never an egress channel." Never call it "air-gapped" — a secure environment requires outbound connectivity (it authenticates to the platform to obtain its messaging credentials; no platform egress ⇒ experiments sit Pending), so a literal air-gap claim fails the first serious security review. "Almost air-gapped" is internal shorthand only.
| Concept | Preferred | Don't use | Status | Definition |
|---|---|---|---|---|
| The software you run on your infra (+ what it gives you) | secure environment | workspace, edge device, agent, box, node, cluster, instance, deployment, site | ✅ | DECIDED 2026-07-12 (Lukas): the on-prem environment is a secure environment everywhere user-facing (matches the marketing hero "Your own secure environment" + the installer/home-screen work). Supersedes the 2026-06-05 "workspace" pick and lifts the old "environment" soft-ban. client survives as the deep-tech / federated-learning synonym collaborators know (and in Client ID), but the default word is secure environment. The load-bearing rename. |
| The command-line tool | the tracebloc CLI (the tracebloc / tb command) |
the client, the binary, the agent | ✅ | What you run to manage your secure environment — ingest data, check status, diagnose. tb is the short alias. |
| The credential connecting it to the platform | Client ID | client key, token, API key | ✅ | Created on the clients page; identifies your secure environment. "client" as a bare noun survives only here. |
| The hosted tracebloc service | the platform | the cloud, the server, the backend, SaaS | ✅ | The hosted side (ai.tracebloc.io) collaborators connect through. |
| The web app you log into | the dashboard (at ai.tracebloc.io) | portal, console, "the platform" (for the UI), Hub | 🔵 OPEN | The browser UI. Brand called this the Hub — keep "dashboard" or adopt "Hub"? Pair once: dashboard = ai.tracebloc.io, docs = docs.tracebloc.io. |
| The user's own servers / laptop | your infrastructure (the specific host: this machine) | your box, on-prem (as a noun) | ✅ | The hardware the secure environment runs on. Distinct from the secure environment above (the tracebloc software/runtime it gives you) — don't conflate the two. In CLI/installer copy the specific host is "this machine". |
| The user's data | dataset | data source, "data set" (two words), "client dataset" | ✅ | Training & test data ingested and staged locally. (It's "your dataset" — never "client data set".) |
| Bringing data in | ingest | upload, import, load, push, send, transfer | ✅ | Copying a dataset into your secure environment's storage. Raw data never leaves your infrastructure. |
| The component that does the ingesting | data ingestor (the ingestor on second mention) | ingester (the -er spelling), importer, loader, connector, ETL job, data pipeline | ✅ | The containerized service that validates a dataset and stages it into your secure environment's storage — the first of the three ingress channels in the definition above. ingestor, never ingester (the distribution is tracebloc-ingestor, the import tracebloc_ingestor — see §9). It is a component; ingest above is the verb. |
| Removing a dataset | delete | rm, drop, teardown | ✅ | Say "removes the dataset from your secure environment (the record is kept)"; keep table/PVC detail behind --verbose. |
| Connection status | Online / Offline | connected/disconnected, up/down | ✅ | Whether your secure environment has an active secure connection. |
| Concept | Preferred | Don't use | Status | Definition |
|---|---|---|---|---|
| The people who build & train models on your data | collaborators | vendor, contributor, participant, expert, "the user" | ✅ | Invited, whitelisted data scientists who train models and never see the raw data. (Decided 2026-07-11 — supersedes the seed's "contributor".) |
| The person who deploys & owns the secure environment | owner (label TBD) | admin, host, customer, "the user" | 🔵 OPEN | Deploys the secure environment, ingests data, creates use cases, controls sharing. Was "workspace owner" — needs a new label now that workspace is retired (data owner / admin panel — decide). |
Say task on every user surface. The wire/spec field is category — internal only. Don't use (user-facing): category, task category, "task categories", task_type, use_case (as the type).
| task_id (internal) | Display name (user-facing) | Gloss |
|---|---|---|
| image_classification | Image classification | sort images into classes |
| object_detection | Object detection | draw boxes around objects |
| keypoint_detection | Keypoint detection | locate landmark points (e.g. pose) |
| semantic_segmentation | Semantic segmentation | label every pixel |
| text_classification | Text classification | sort text into classes |
| masked_language_modeling | Masked language modeling (fill-mask) | predict masked-out words (no labels) |
| token_classification | Token classification | label each word in a sequence |
| sentence_pair_classification | Sentence-pair classification | label how two texts relate |
| causal_language_modeling | Causal language modeling | predict the next word |
| seq2seq | Sequence-to-sequence (translation / summarization) | map input sequence → output |
| embeddings | Embeddings | learn vector representations from text pairs |
| tabular_classification | Tabular classification | predict a class from table columns |
| tabular_regression | Tabular regression | predict a number from table columns |
| time_series_forecasting | Time-series forecasting | predict future values from past |
| time_series_classification | Time-series classification | predict a class per whole sequence |
| time_to_event_prediction | Survival analysis | predict how long until an event |
How the user lays out each task's data, and what the label/target column is actually called (from the di#347 layout contract + the ingestor modality specs + shipped template CSVs). ⇥ = tab.
| Task | Input layout | Sidecars | Label / target column | Supervision |
|---|---|---|---|---|
| Image classification | folder images/ + CSV (filename,label) |
— | label |
labeled |
| Object detection | images/ + CSV (filename,image_label) + annotations/*.xml |
Pascal-VOC XML, paired by filename stem | boxes+class in the XML; CSV image_label is coarse |
labeled |
| Keypoint detection | images/ + CSV (filename,Annotation,Visibility,image_label) |
— | class image_label; coords in Annotation (JSON) |
labeled |
| Semantic segmentation | images/ + CSV (filename,mask_id,image_label) + masks/*.png |
masks/*.png linked by mask_id |
mask PNG via mask_id; class image_label |
labeled |
| Text classification | folder texts/ + CSV (filename,extension,label) |
— | label |
labeled |
| Masked language modeling | folder sequences/ + CSV (filename,extension) |
— | none | self-supervised |
| Token classification | texts/ + CSV (filename,extension,label) |
— | label (BIO/IOB2 tag sequence) |
labeled |
| Sentence-pair classification | texts/ (text_a⇥text_b) + CSV (…,label) |
— | label |
labeled |
| Causal language modeling | texts/ (raw or prompt⇥completion) + CSV (filename,extension) |
— | none | self-supervised |
| Seq2seq | texts/ (source⇥target) + CSV (filename,extension) |
— | none | self-supervised |
| Embeddings | texts/ (anchor⇥positive[⇥negative]) + CSV (filename,extension) |
— | none (contrastive) | self-supervised |
| Tabular classification | single CSV: id + features + label |
— | label |
labeled |
| Tabular regression | single CSV: id + features + numeric target | — | configurable numeric target — no fixed word (sample price); label.policy=bucket |
labeled (regression) |
| Time-series forecasting | single CSV: timestamp + features + numeric target |
— | configurable numeric target (sample value); time col timestamp |
labeled (regression) |
| Time-series classification | single grouped CSV: sequence_id,timestamp,…,label |
— | label (one per sequence); group sequence_id, time timestamp |
labeled |
| Survival analysis (time-to-event) | single CSV: features + time + event indicator |
— | duration col time (fixed) + event indicator = the label column (sample DEATH_EVENT) |
labeled (survival) |
Consistent, keep: filename (every file-bearing task), optional extension (text tasks), mask_id (semseg link), Annotation (keypoint coords), sequence_id/timestamp (time-series grouping).
🔵 Column-naming inconsistencies to decide (folded into Open Q#7):
labelvsimage_label— classification/text/tabular/TSC uselabel; the three vision tasks (object detection, keypoint, semseg) useimage_label. Canonicalize on one, or documentlabel(classification) vsimage_label(vision image-level)?timestampvstime— time-series usestimestamp; survival usestime. Same concept, two words.- Survival event indicator has no platform name — it's just the configurable label column;
DEATH_EVENTis only the shipped Heart-Failure sample. Docs must say "event indicator (the label column)" + the fixedtimecolumn — never presentDEATH_EVENTas a tracebloc column. masked_language_modelingstages fromsequences/, the other three self-supervised text tasks fromtexts/— spell this folder difference out.
| Concept | Preferred | Don't use | Status | Definition |
|---|---|---|---|---|
| A collaborator's attempt within a use case | experiment | training run, training job, training plan | ✅ | The unit a collaborator runs against a use case. |
| The act of training a model | training (run a training) | — | ✅ | A training happens inside an experiment. |
| The Kubernetes/infra work object | job | — | ✅ | Infra-internal only — never user-facing. "job" = the K8s Job; don't use it in product/docs/CLI copy. |
| Concept | Preferred | Don't use | Status | Definition |
|---|---|---|---|---|
| The ML model a collaborator trains | model | architecture | ✅ | — |
| A ready-made starter model/project | template | sample, demo | ✅ | model-zoo starters only. |
| The act of a collaborator adding their model | submit a model or upload a model | — | 🔵 OPEN | SDK method is upload_model. Is the verb "upload", "submit", or "contribute" a model? (Data is never "uploaded" — but a model legitimately goes to the platform.) |
| Concept | Preferred | Don't use | Status | Definition |
|---|---|---|---|---|
| Training across secure environments + combining weights | federated training ("build together") | distributed learning, FedAvg (in user copy) | ✅ | User-facing. |
| The infra that merges the weights | averaging service / federated averaging | — | ✅ | Internal component name. |
| Concept | Preferred | Don't use | Status |
|---|---|---|---|
| Our category | Collaborative AI / "build AI together" | "AI collaboration platform", "secure AI" | ✅ |
| The results view | Leaderboard | dashboard, scoreboard | ✅ |
| Pricing / compute unit | PetaFLOPs (PF) | credits, tokens | ✅ |
| Our security story | Compliance by architecture | enterprise-grade security, compliance by design | ✅ |
| Surface | Preferred | Don't use |
|---|---|---|
| load data | tracebloc data ingest <path> |
dataset push, upload |
| list datasets | tracebloc data list |
— |
| remove a dataset | tracebloc data delete <name> |
dataset rm |
| the task flag | --task |
--category (hidden alias only) |
| the name flag | --name (rename delete's positional <table> → <name>) |
--table (hidden alias only) |
| train/test flag | --split (train|test) (RFC-0002 — not yet implemented; today --intent) |
--intent (after migration) |
Command group is data (alias dataset one cycle). Keep cluster/namespace/PVC/"stage pod" jargon behind --verbose.
"client" today names four different things (a fifth — the training-images repo — is already resolved by the tracebloc-engine rename) — disambiguate 🔵:
| Thing | Today | Proposed |
|---|---|---|
| Helm chart that deploys the secure environment | client repo |
environment-chart |
| Training-execution container images | tracebloc-engine repo |
✅ renamed from tracebloc-client; describe as "the engine" / "training images" |
| In-cluster pod orchestration | client-runtime repo |
keep; "runtime" |
| The credential | Client ID | unchanged (the one legit "client") |
| CLI command group | tracebloc client |
installer-internal, hidden |
- Ingestor: distribution
tracebloc-ingestor(PyPI) / importtracebloc_ingestor— state both (fix data-ingestors/CLAUDE.md which claims the PyPI name is the underscore form). - SDK methods: snake_case (
upload_model,link_model_dataset); camelCase (uploadModel) forms are deprecated aliases — fix the model-zoo README quick-start. - Backend model:
EdgeDevice/edge= the secure environment (clientin FL terms);Competition/PrivateCompetition= a use case. Biggest reality-vs-canon gap (rename = migration + API; own epic).
- tracebloc is always lowercase — including sentence-start and mid-sentence. Not "Tracebloc"/"TraceBloc" (except unavoidable title-case contexts).
- Avoid hype/corporate filler (revolutionary, cutting-edge, seamless, leverage, end-to-end, enterprise-grade, AI-powered, democratize, …) — full banned list in the
communicationskill, which references this file.
- competition / PrivateCompetition → use case (frontend ×782, backend
Competitionmodel + route) - workspace → secure environment (docs, website hero, decks, and the
communicationskill's Terminology Bible) - edge / edge device / EdgeDevice → secure environment — client only in FL-technical contexts, never in UI copy (frontend "edge" ×449, backend
EdgeDevice) - vendor → collaborators (client README, SDK README, frontend "Vendor Testing Platform" ×54, several docs pages)
- push /
dataset push→ ingest /data ingest(client README quick-install, docs cli.mdx) dataset rm→data delete(cli README)- ingester → data ingestor (exactly one instance org-wide: the
Data ingesterchannel row incli/docs/rfcs/0003-storage-and-offboard-hygiene.md§1 — the RFC the definition above restates. Left in place: that RFC is Accepted with a dated decision log, so it is a review target here, not a silent edit there.) - "data set" (two words) / "client dataset" → dataset / "your dataset"
- category → task (docs cli.mdx, "9 task categories")
- uploadModel() → upload_model() (model-zoo README)
- Tracebloc (capitalized) → tracebloc (docs key-terms.mdx, frontend alt text)
- admin / "workspace owner" → owner (label TBD) (pending Q4 — the owner label is undecided now that workspace is retired)
- "compliance by design" → Compliance by architecture (the website hero currently says "compliance by design" — fix it)
- ✅ What we call the on-prem thing — DECIDED 2026-07-12 (Lukas): secure environment. Supersedes the 2026-06-05 "workspace" pick + lifts the old "environment" ban. workspace → retired user-facing; client survives as the deep-tech/FL synonym + in Client ID. Propagation (do next): the
communicationskill's embedded "Terminology Bible" still says workspace + "don't say environment" — replace it with a reference to this file; sweep the website + decks; then the code rename (edge / EdgeDevice → secure environment) is unblocked (its own epic). - Hub vs the dashboard/platform — one name for the web UI.
- The model verb — submit / upload / contribute a model?
- workspace owner vs admin — esp. the UI "admin panel".
- The
clientcomponent overload — rename theclientrepo, or only fix user-facing copy? - Internal ML-type field — keep
categoryon the wire, or move totask_type? (User-facing is "task" either way.) - Per-task column names — (a)
labelvsimage_label(classification/text/tabular/TSC vs the 3 vision tasks); (b)timestampvstime(time-series vs survival). Unify each, or document the split? (The survival event indicator stays "the label column" — no fixed name;timeandmask_idstay fixed.)