Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 23 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,29 @@ All notable changes to this project are documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased]

### Added

- Apple `container` is now supported as an alternative container backend on
Apple silicon (macOS 26+). Select it with the new global `container_backend`
setting (`docker`, `apple`, or `auto`). Docker remains the default; Apple and
automatic selection are opt-in. Container services run unchanged on either
backend. The TUI and `bench status` show the active backend. See
`docs/apple-container.md`.
- Container images with no usable variant for the host architecture (e.g. an
amd64-only image on Apple silicon) now fail terminally with a clear message
instead of looping through the restart policy.
- New readiness kind `container_exec` runs a command inside the service's own
container. Workbench supplies the container name and the backend's CLI, so
`command: pg_isready -U bench` works unchanged on either container backend —
unlike `kind: exec` with a hand-written `docker exec <name> …`, which breaks
when the backend resolves to Apple `container` or `container_prefix` changes.
See `docs/configuration.md`.
- Setup hooks now support `kind: exec` and `kind: container_exec`, using the
same configuration as readiness hooks. Existing command-only setup hooks
remain host-side `exec` hooks and emit a deprecation warning.

## [0.6.10] - 2026-07-07

### Fixed
Expand Down
107 changes: 107 additions & 0 deletions docs/apple-container.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,107 @@
# Apple Container backend

On Apple silicon, workbench can run container services on
[Apple's `container`](https://github.com/apple/container) tool instead of
Docker. Apple `container` runs each Linux container in its own lightweight
virtual machine and integrates directly with Virtualization.framework, avoiding
the Docker Desktop dependency.

The `bench.yaml` service schema is unchanged — the same `container:` block runs
on either backend. Only the global `container_backend` setting differs.

## Requirements

- Apple silicon Mac (arm64).
- macOS 26 (Tahoe) or later.
- The [`container`](https://github.com/apple/container) CLI installed and its
system service running:
```bash
container system start
```

## Selecting the backend

```yaml
global:
container_backend: apple # docker (default) | apple | auto
apple:
gateway_ip: 192.168.64.1
```

- **`docker`** (default) — always use Docker.
- **`apple`** — always use Apple `container`. Startup fails with a clear message
if the host doesn't meet the requirements above.
- **`auto`** — use Apple `container` when running on Apple silicon with the
`container` binary installed; otherwise use Docker.

The active backend is shown per container service in the TUI detail pane and in
`bench status` (the `TYPE` column reads `container/apple` or `container/docker`,
and `--json` includes a `backend` field).

## Host connectivity (tracing)

Docker exposes `host.docker.internal` so a container can reach services on the
host (workbench uses this for the OTLP trace collector). Apple `container` has
no such alias. Instead, containers reach the host at the vmnet **gateway IP**,
which defaults to `192.168.64.1`.

Workbench injects that gateway IP as the OTLP endpoint host for Apple-backend
container services automatically. If you've changed the `container` default
subnet (in `~/.config/container/config.toml`), set `apple.gateway_ip` to the
matching gateway address.

## Commands that exec into a container

Use `kind: container_exec` rather than `kind: exec` with a hand-written `docker
exec`. Workbench supplies the container and the backend's CLI, so readiness
probes and setup hooks are portable across backends:

```yaml
readiness:
kind: container_exec
command: pg_isready -U bench -d bench
```

A probe written as `kind: exec` with `command: docker exec my-postgres
pg_isready …` keeps working on Docker but fails here — the container isn't in
Docker's namespace, so every attempt reports `No such container: my-postgres`
until `max_attempts` is exhausted. The service itself is usually healthy the
whole time; only the probe is broken. If a readiness failure appears right after
switching backends, check the log buffer's `probe` lines for `No such
container` — that's this, and `container_exec` is the fix.

## Differences from Docker

- **Isolation** — one lightweight VM per container, rather than shared-kernel
namespaces. Startup is slightly slower but isolation is stronger.
- **Exit codes** — `container` has no `wait` subcommand and does not reliably
expose a container's process exit code via `inspect`. Workbench detects that a
container has stopped by polling its status; the reported exit code is
best-effort. This mainly affects `restart.policy: on-failure`, which may not
distinguish a clean exit from a crash as precisely as it does on Docker.
- **Anonymous volumes** — Docker's `-v` removal flag drops anonymous volumes on
cleanup. Apple `container` has no equivalent; anonymous volumes are not
auto-removed.
- **Host networking / `--add-host`** — not used by workbench on this backend;
host access goes through the gateway IP instead.

## Images without an arm64 variant

Apple `container` runs images for the host architecture (arm64), using Rosetta
to run amd64 images where possible. An image with **no usable variant** for the
host — e.g. an amd64-only image whose binaries Rosetta can't run — fails at
start with `does not support required platforms` or `exec format error`.

Workbench treats this as a **terminal failure**: the service goes straight to
`failed` with a clear message (`image does not support this platform: …`) and is
*not* retried, even under `restart.policy: always`. Retrying can never succeed,
so looping would only bury the real reason in noise. Rebuild or source a
multi-arch (arm64) image to fix it.

## Out of scope

- Building images (`container build`) — workbench only runs pre-built images.
- `container system dns` domains — workbench uses the gateway IP so it never
needs `sudo` or to disable iCloud Private Relay.
- Starting the `container` system service — workbench reports if it's not
running but does not run `container system start` for you.
68 changes: 55 additions & 13 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,7 +65,7 @@ Unknown YAML fields are rejected at parse time. A typo such as `expect_status: 2

```
error: parsing config bench.yml: parsing config: yaml: unmarshal errors:
line 9: field expect_status not found in type config.ReadinessConfig
line 9: field expect_status not found in type config.ServiceHookConfig
```

Run `bench validate` to surface these errors without starting any services.
Expand All @@ -90,9 +90,25 @@ Run `bench validate` to surface these errors without starting any services.
| `watch_debounce` | duration | `300ms` | Default debounce for file watchers |
| `env` | map | | Global environment variables applied to all services |
| `env_file` | path | | Global .env file loaded for all services |
| `container_prefix` | string | dirname | Prefix for Docker container names (e.g. `{prefix}-{service}`) |
| `container_prefix` | string | dirname | Prefix for container names (e.g. `{prefix}-{service}`) |
| `container_backend`| string | `docker`| Container runtime: `docker`, `apple`, or `auto` |
| `apple` | object | | Apple `container` backend settings |
| `tracing` | object | | Tracing configuration |

#### Container backend

Container services run on Docker by default. On Apple silicon you can run them
on [Apple's `container`](apple-container.md) tool instead.

| Field | Type | Default | Description |
| ------------------- | ------ | ------- | -------------------------------------------------------- |
| `container_backend` | string | `docker`| `docker`, `apple`, or `auto` (prefer Apple when present) |
| `apple.gateway_ip` | string | `192.168.64.1` | Host IP an Apple container uses to reach the host |

`auto` is opt-in. It selects the Apple backend when running on Apple silicon
with the `container` binary installed, otherwise Docker. See
[apple-container.md](apple-container.md) for requirements and caveats.

#### Tracing

| Field | Type | Default | Description |
Expand Down Expand Up @@ -216,15 +232,15 @@ Common noisy directories (`.git`, `node_modules`, `__pycache__`) are always excl

| Field | Type | Description |
| --------------- | -------------- | --------------------------------------------------------------------------------- |
| `kind` | string | `none`, `log_pattern`, `tcp`, `http`, `exec`, or `grpc` |
| `kind` | string | `none`, `log_pattern`, `tcp`, `http`, `exec`, `container_exec`, or `grpc` |
| `pattern` | string | Go regular expression matched against log lines (for `log_pattern`) |
| `address` | string | TCP address to dial, `host:port` (for `tcp` and `grpc`) |
| `url` | string | HTTP URL to GET; any 2xx response means ready (for `http`) |
| `command` | string or list | Shell command or argv to run (for `exec`); exit 0 = ready |
| `command` | string or list | Shell command or argv to run (for `exec` and `container_exec`); exit 0 = ready |
| `service` | string | gRPC service name (for `grpc`); empty = overall server health |
| `timeout` | duration | Per-attempt probe timeout (default `2s`) |
| `initial_delay` | duration | Delay before the first probe attempt |
| `interval` | duration | Sleep between failed attempts (default `500ms`); applies to `tcp`, `http`, `exec` |
| `interval` | duration | Sleep between failed attempts (default `500ms`); applies to `tcp`, `http`, `exec`, `container_exec` |
| `max_attempts` | integer | Cap on probe attempts before giving up (default `0` = unlimited) |
| `settle` | duration | Delay between probe-success and the Ready transition |

Expand All @@ -245,9 +261,29 @@ dependents parked in Pending.
successful connect wins.
- **`http`** issues `GET url` using an `http.Client` with `timeout`. Any 2xx
response marks the service Ready.
- **`exec`** runs `command` with a `timeout` deadline per attempt. Exit 0 = ready.
stdout/stderr from the probe is appended to the service's log buffer tagged
with stream `probe`, so you can see what the probe is observing.
- **`exec`** runs `command` on the **host** with a `timeout` deadline per attempt.
Exit 0 = ready. stdout/stderr from the probe is appended to the service's log
buffer tagged with stream `probe`, so you can see what the probe is observing.
- **`container_exec`** runs `command` **inside the service's own container**,
otherwise behaving exactly like `exec`. Only valid on a service with a
`container:` block; `bench validate` rejects it elsewhere. Workbench supplies
both the container and the runtime CLI, so the probe works unchanged on either
[container backend](apple-container.md) and does not depend on
`container_prefix` or the service key:

```yaml
services:
postgres:
container:
image: postgres:16-alpine
readiness:
kind: container_exec
command: pg_isready -U bench -d bench
```

Prefer this over `exec` with a hand-written `docker exec <name> …`, which
hardcodes both the runtime and the container name and so breaks when the
backend resolves to Apple `container` or the prefix changes.
- **`grpc`** issues a `grpc.health.v1.Health/Check` call against `address`. Ready
when the server responds with status `SERVING`. Set `service` to probe a
specific gRPC service registered for health reporting; leave it empty to
Expand Down Expand Up @@ -304,11 +340,12 @@ passes (and after any `settle` delay), before the service transitions to
bootstrap that's logically part of bringing this service up — creating a dev
environment in a flag service, seeding a default DB user, applying migrations.

| Field | Type | Description |
| --------- | -------------- | ------------------------------------------------------ |
| `command` | string or list | Shell command or argv to run; exit 0 = setup succeeded |
| `timeout` | duration | Cap on setup runtime (default `60s`) |
| `env` | map | Extra env applied on top of the service's env |
| Field | Type | Description |
| --------- | -------------- | ------------------------------------------------------------- |
| `kind` | string | `exec` (host) or `container_exec` (the service's container) |
| `command` | string or list | Shell command or argv to run; exit 0 = setup succeeded |
| `timeout` | duration | Cap on setup runtime (default `60s`) |
| `env` | map | Extra environment for `exec`; unsupported by `container_exec` |

```yaml
services:
Expand All @@ -318,10 +355,15 @@ services:
kind: http
url: http://localhost:4242/health
setup:
kind: exec
command: ./bin/flagman create-env development
timeout: 30s
```

`container_exec` is only valid for container services and uses the configured
container backend. A legacy setup block containing `command` without `kind`
still runs as `exec`, with a deprecation warning.

The status flow is `Running → Setup → Ready`. On non-zero exit or timeout the
supervisor stops the service and marks it **Failed** with the setup error in
`last_error`, so dependents cascade just as they would for any other failure.
Expand Down
13 changes: 10 additions & 3 deletions docs/troubleshooting.md
Original file line number Diff line number Diff line change
Expand Up @@ -119,17 +119,24 @@ a **container** service produces nothing, check the endpoint the service is
actually exporting to. A container's `localhost` is its own loopback, not the
host — so `http://localhost:<port>` silently fails with connection-refused.

Workbench injects `http://host.docker.internal:<port>` for container services
and adds a `host.docker.internal:host-gateway` alias to each container run. If
spans still don't arrive:
Workbench injects the OTLP endpoint using a backend-specific host: on Docker it
uses `http://host.docker.internal:<port>` and adds a
`host.docker.internal:host-gateway` alias to each container run; on the
[Apple backend](apple-container.md) it uses the vmnet gateway IP
(`http://192.168.64.1:<port>` by default). If spans still don't arrive:

- Confirm the service didn't override `OTEL_EXPORTER_OTLP_ENDPOINT` itself (any
env layer outranks the injected default — see `docs/configuration.md`). A
hardcoded `localhost` in the service's own config is the usual culprit.
- Verify the container can reach the host collector:
```bash
# Docker
docker exec <container> getent hosts host.docker.internal
# Apple container
container exec <container> sh -c 'nc -z -v 192.168.64.1 <port>'
```
On the Apple backend, if you've changed the `container` default subnet, set
`global.apple.gateway_ip` to the matching gateway address.
- Confirm the collector is listening on the host: `lsof -nP -iTCP:<port> -sTCP:LISTEN`.

## Getting debug output
Expand Down
2 changes: 2 additions & 0 deletions internal/api/handlers.go
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@ type ServiceStatus struct {
DisplayName string `json:"display_name"`
Status string `json:"status"`
Type string `json:"type"`
Backend string `json:"backend,omitempty"`
PID int `json:"pid,omitempty"`
ContainerID string `json:"container_id,omitempty"`
Image string `json:"image,omitempty"`
Expand Down Expand Up @@ -92,6 +93,7 @@ func (s *Server) buildServiceStatus(key string) ServiceStatus {
DisplayName: snap.Name(),
Status: snap.Status.String(),
Type: snap.ServiceType,
Backend: snap.Backend,
PID: snap.PID,
ContainerID: snap.ContainerID,
Image: snap.Image,
Expand Down
Loading