Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -101,6 +101,9 @@ instance/
# Sphinx documentation
docs/_build/

# Local agent planning artifacts
docs/superpowers/

# Jupyter Notebook
.ipynb_checkpoints

Expand Down
4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Added

- **Spur (Crusoe) scheduler support** (#157): Spur ships SLURM-compatible CLI shims, but `srun` cannot fan out across nodes (`SLURM_PROCID` is empty), `scontrol show hostname[s]` is unsupported, and the Raft-based control plane makes `squeue`/`sacct` eventually consistent. A new `SpurDeployment` backend reuses the SLURM template and presets but drives multi-node runs with a job **array** of single-node tasks that self-form the cluster through a shared-filesystem rendezvous, and detects completion from per-rank marker files instead of `sacct`. Select it with `"slurm": {"scheduler": "spur"}` — the `slurm` block is otherwise unchanged, so spur is reachable from both `madengine build` and `madengine run --additional-context`. Peers that do not see rank 0's `MASTER_ADDR` within `slurm.rendezvous_timeout` (default 900s) fail with a diagnostic instead of starting with an empty address. The `srun`-based node health preflight is skipped on spur (it cannot target a node there, and the `--nodelist` it pins conflicts with the job array), an explicit `"deploy": "spur"` selects the backend at run time as well as at build time, and monitoring gives up with a diagnostic if the array never appears in `squeue` instead of polling forever. Stock SLURM behaviour is unchanged.

### Docs

- **README rewritten as a concise landing page** (#161): Trimmed the root README from 707 to 258 lines by moving deep reference material (profiling tables, extended config/usage recipes, tips) into `docs/` and linking out. Replaced the stale ASCII architecture block and unreferenced `docs/img` PNGs with accurate inline Mermaid diagrams for the layered architecture, build→run→report pipeline, and deployment-target inference; added matching diagrams to `docs/deployment.md` and `docs/README.md`. Also corrects numerous stale references across docs: `--csv-file` → `--csv-file-path`/`--file`, missing `database` command flags (`--unique-key`/`-k`, `--batch-size`, `--no-upsert`, `--no-index`, `--dry-run`, `MONGO_AUTH_SOURCE`/`MONGO_TIMEOUT_MS`), wrong `run --output`/`--tools-config` defaults, `megatron` → `megatron-lm` launcher name, fabricated `timeout_multiplier`/`service_account` config keys, missing Kubernetes/SLURM `additional_context` keys, `DOCKER_CONFIG`/`MAD_SKIP_DOCKER_LOGIN` documentation, and corrected SGLang Disaggregated minimum node counts/split formula for SLURM vs. Kubernetes.
Expand Down
37 changes: 37 additions & 0 deletions docs/deployment.md
Original file line number Diff line number Diff line change
Expand Up @@ -249,9 +249,46 @@ The deployment target is automatically detected from the `slurm` key in the conf
- `reservation`: SLURM reservation name; forwarded to srun health/cleanup commands
- `time`: Wall time limit (HH:MM:SS)
- `exclusive`: Exclusive node access (default: `true`)
- `scheduler`: SLURM flavor - `slurm` (default) or `spur` (see below)
- `rendezvous_timeout`: spur only; seconds a node waits for rank 0 to publish `MASTER_ADDR` (default: 900)

See [examples/slurm-configs/](../examples/slurm-configs/) for complete examples.

### Spur (Crusoe) Scheduler

Spur exposes SLURM-compatible CLI shims but `srun` cannot fan tasks out across
nodes, so multi-node runs use a job **array** of single-node tasks instead: each
array task runs on one node, `SLURM_ARRAY_TASK_ID` is the node rank, and the
tasks self-form the cluster through a shared-filesystem rendezvous (rank 0
publishes its transport IP; the other ranks read it as `MASTER_ADDR`).

Select it with `slurm.scheduler`; everything else in the `slurm` block is
unchanged:

```json
{
"slurm": {
"scheduler": "spur",
"partition": "gpu",
"nodes": 4,
"gpus_per_node": 8,
"time": "02:00:00"
}
}
```

A job array carries no gang-scheduling guarantee, so tasks may start minutes
apart. A node that does not see rank 0's address within `rendezvous_timeout`
fails the run with a diagnostic rather than starting with an empty
`MASTER_ADDR`; raise the timeout if your queue wait is longer than that.

`slurm.output_dir` must be on a filesystem shared by every node - it holds the
rendezvous files.

The `srun`-based node health preflight (`enable_node_check`) is skipped on spur:
`srun -w <node>` does not run on the requested node there, and the `--nodelist`
it would pin conflicts with the job array, whose tasks each request one node.

### Multi-Node Training

For distributed training across SLURM nodes:
Expand Down
Loading