Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 14 additions & 10 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -112,29 +112,33 @@ When `run_cluster = BashSLURM` or `run_cluster = BashLSF`, config-driven runs wa
Dry-run a config to inspect the resolved phase calls:

```bash
python run.py -c run_configs/uci_binary_hcc.cfg --dry_run
python run.py -c run_configs/local/uci_binary_hcc.cfg --dry_run
```

Run the configured pipeline:

```bash
python run.py -c run_configs/uci_binary_hcc.cfg
python run.py -c run_configs/local/uci_binary_hcc.cfg
```

Useful controls:

```bash
python run.py -c run_configs/uci_binary_hcc.cfg --start_at p4
python run.py -c run_configs/uci_binary_hcc.cfg --stop_after p8
python run.py -c run_configs/uci_binary_hcc.cfg --only p6,p8,p11
python run.py -c run_configs/uci_binary_hcc.cfg --skip p3,p4
python run.py -c run_configs/local/uci_binary_hcc.cfg --start_at p4
python run.py -c run_configs/local/uci_binary_hcc.cfg --stop_after p8
python run.py -c run_configs/local/uci_binary_hcc.cfg --only p6,p8,p11
python run.py -c run_configs/local/uci_binary_hcc.cfg --skip p3,p4
```

Example configs are included for the three UCI demos:
Example configs are included for the three UCI demos and are organized by run environment:

- `run_configs/uci_binary_hcc.cfg`
- `run_configs/uci_multiclass_student.cfg`
- `run_configs/uci_regression_auto_mpg.cfg`
- `run_configs/local/uci_binary_hcc.cfg`
- `run_configs/local/uci_multiclass_student.cfg`
- `run_configs/local/uci_regression_auto_mpg.cfg`
- `run_configs/hpc/cedars_slurm_hcc.cfg`
- `run_configs/hpc/upenn_lsf_hcc.cfg`

The original top-level demo configs are still kept for backward compatibility.

Phase 10 runs only when replication paths are configured, unless it is explicitly enabled. Phase 7 is automatically skipped for continuous/regression runs because the current ensemble registry is classification-only.

Expand Down
30 changes: 30 additions & 0 deletions docs/source/changelog.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,36 @@ Older public release notes are based on the
[GitHub Releases](https://github.com/UrbsLab/STREAMLINE/releases) entries, with
minor wording cleanup for readability.

## v1.0.1 - Bug Fix Release

STREAMLINE v1.0.1 is a focused bug-fix release for regression reporting,
regression EDA robustness, and multiclass weighted feature-importance
visualization.

### Fixed

* Fixed Phase 1 regression EDA so skewness and kurtosis calculations do not
crash on mixed-type, missing, infinite, or nonnumeric continuous outcome
values.
* Fixed Phase 11 report task detection so explicit saved or CLI `outcome_type`
values are respected before falling back to dataset inference.
* Fixed regression reports for low-cardinality continuous outcomes that could
otherwise be misidentified as multiclass based on `ClassCounts.csv`.
* Fixed regression report performance tables so both display-style metric names
and snake_case metric keys are recognized.
* Fixed multiclass weighted composite feature-importance plots so the
balanced-accuracy no-skill baseline is `1 / number_of_classes` rather than
always `0.5`.

### Added

* Added an HPC and cluster-running documentation page with Conda setup notes,
`tmux` workflow basics, SLURM/LSF monitoring commands, scheduler config
explanations, and recovery/rerun guidance.
* Added a UPenn/LSF HCC demo config template alongside the existing
Cedars/SLURM template and updated README, installation, running, and
parameter documentation to point users to both HPC examples.

## v1.0.0 - Main Release

STREAMLINE v1.0.0 is a major reorganization and expansion of STREAMLINE into a
Expand Down
2 changes: 1 addition & 1 deletion docs/source/conf.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
project = "STREAMLINE"
copyright = "2026, Ryan Urbanowicz, Harsh Bandhey"
author = "Ryan Urbanowicz, Harsh Bandhey"
release = "1.0.0"
release = "1.0.1"

extensions = [
"sphinx.ext.autodoc",
Expand Down
2 changes: 1 addition & 1 deletion docs/source/development.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ Prefer the existing registry patterns when adding new components:
## Tests

The default pytest configuration collects only the current main end-to-end
tests. Legacy and phase-level subtests were removed from the maintained v1.0.0
tests. Legacy and phase-level subtests were removed from the maintained v1.0.1
test path so routine testing stays focused on the binary, multiclass, and
regression demo pipelines.

Expand Down
215 changes: 215 additions & 0 deletions docs/source/hpc.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,215 @@
# HPC and Cluster Runs

STREAMLINE can run small examples on a laptop, but paper-scale runs often need a
cluster. The cluster path is still the same pipeline: edit a `.cfg`, dry-run it,
then launch the config runner. The difference is that selected phases submit
many scheduler jobs through SLURM or LSF and the config runner waits for those
jobs to finish before moving to the next phase.

## When To Use Each Execution Mode

| Mode | Best use |
| --- | --- |
| `Serial` | Debugging, small demos, and first config checks. |
| `Parallel` | A single machine or one allocated compute node using joblib multiprocessing. |
| `Local` | A local Dask cluster on one machine. |
| `BashSLURM` | HPC systems that submit jobs with `sbatch`. |
| `BashLSF` | HPC systems that submit jobs with `bsub`. |
| Named Dask cluster | Site-specific Dask jobqueue execution when configured by the user/site. |

Use a scheduler mode for long P4/P6/P8/P10/P11-style workloads or any analysis
that would be inappropriate to run directly on a login node. Use `Parallel` only
inside an interactive allocation or on a machine where it is acceptable to use
multiple local cores.

## Included HPC Config Templates

HPC configs live in `run_configs/hpc/`.

| Config | Scheduler | Intended starting point |
| --- | --- | --- |
| `run_configs/hpc/cedars_slurm_hcc.cfg` | SLURM | Cedars/Sinai-style SLURM clusters using `run_cluster = BashSLURM`. |
| `run_configs/hpc/upenn_lsf_hcc.cfg` | LSF | UPenn/I2C2-style LSF clusters using `run_cluster = BashLSF`. |

Both templates run the HCC binary demo by default. Copy one of them before using
it for a real project and edit at least `output_path`, `experiment_name`,
`data_path`, `queue`, `reserved_memory`, model list, and modeling budget.

## Basic Cluster Setup

From a login node:

```bash
ssh <user>@<cluster-host>
git clone --single-branch https://github.com/UrbsLab/STREAMLINE.git
cd STREAMLINE
conda create -n streamline python=3.11 pip
conda activate streamline
pip install -r requirements.txt
python run.py --help
```

Many clusters require modules before Conda, Python, or compiled libraries are
available. If your site uses modules, load the same modules before installation
and before running STREAMLINE jobs. Also make sure the repository, data, and
`output_path` are on a filesystem visible to compute nodes.

## Conda Installation Quickstart

If Conda is already available on the cluster, either directly or through a
module, create a dedicated STREAMLINE environment from the repository root:

```bash
module load anaconda # omit or change this if your cluster uses a different module name
conda create -n streamline python=3.11 pip
conda activate streamline
pip install -r requirements.txt
```

If Conda is not available, install Miniconda in your home or project space using
your cluster's approved download method:

```bash
mkdir -p ~/miniconda3
curl -L https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh -o /tmp/miniconda.sh
bash /tmp/miniconda.sh -b -p ~/miniconda3
source ~/miniconda3/etc/profile.d/conda.sh
conda create -n streamline python=3.11 pip
conda activate streamline
pip install -r requirements.txt
```

Some HPC systems block outbound internet from compute nodes. In that case,
install packages from the login node, a site Conda mirror, or an administrator
provided module/wheelhouse, then run STREAMLINE from the same environment.

## Use tmux For Long Runs

The config runner is the phase orchestrator. Scheduler jobs can keep running if
your SSH connection drops, but the runner may stop waiting and the next phases
may not launch. Use `tmux` or `screen` for long runs.

```bash
tmux new -s streamline
conda activate streamline
python run.py -c run_configs/hpc/cedars_slurm_hcc.cfg --dry_run
python run.py -c run_configs/hpc/cedars_slurm_hcc.cfg
```

Useful `tmux` commands:

```bash
# Detach from the session without stopping STREAMLINE:
Ctrl-b, then d

# List sessions:
tmux ls

# Reattach later:
tmux attach -t streamline

# Kill the session after the run is done:
tmux kill-session -t streamline
```

The same pattern works for the UPenn LSF template:

```bash
tmux new -s streamline
conda activate streamline
python run.py -c run_configs/hpc/upenn_lsf_hcc.cfg --dry_run
python run.py -c run_configs/hpc/upenn_lsf_hcc.cfg
```

## Scheduler Settings In Configs

The core cluster settings live in the `[run]` section:

```ini
run_cluster = BashSLURM
wait_for_cluster_completion = True
cluster_phase_timeout = 86400
cluster_phase_poll_interval = 30
queue = defq
reserved_memory = 4
```

For UPenn/LSF, the same fields look like:

```ini
run_cluster = BashLSF
queue = i2c2_normal
reserved_memory = 4
```

`queue` maps to the scheduler queue or partition. `reserved_memory` is the memory
request in GB used when STREAMLINE writes scheduler scripts. The exact queue
names and memory limits are site-specific, so treat the included values as
starting points.

`wait_for_cluster_completion = True` tells the config runner to wait for
STREAMLINE completion markers in `jobsCompleted/` before it starts the next
phase. This is important because later phases depend on files written by earlier
scheduler jobs.

## Monitoring Jobs

STREAMLINE writes scheduler scripts to the experiment `jobs/` folder and
stdout/stderr files to `logs/`.

Common SLURM commands:

```bash
squeue -u $USER
sacct -j <job_id>
scancel <job_id>
```

Common LSF commands:

```bash
bjobs
bjobs -l <job_id>
bkill <job_id>
```

If a phase appears stuck, check the scheduler first, then inspect
`<output_path>/<experiment_name>/logs/` and the `jobsCompleted/` markers.

## Recovery And Reruns

Use a dry run before every large launch:

```bash
python run.py -c run_configs/hpc/cedars_slurm_hcc.cfg --dry_run
```

If one phase fails, restart from that phase instead of repeating the full run:

```bash
python run.py -c run_configs/hpc/cedars_slurm_hcc.cfg --start_at p6
python run.py -c run_configs/hpc/cedars_slurm_hcc.cfg --only p8,p11
```

Phase 6 reruns and overwrites requested model jobs by default. For recovery,
set `skip_completed_models = True` in `[p6]` or pass
`--skip_completed_models 1` to the P6 CLI. That runs missing or failed model/CV
jobs while leaving completed model jobs in place.

If the config runner times out while scheduler jobs are still queued or running,
increase `cluster_phase_timeout` and rerun from the interrupted phase after
checking the logs.

## Practical HPC Checklist

Before a paper-scale cluster run:

* Confirm the config with `--dry_run`.
* Use absolute paths for project data and outputs when running outside the repo.
* Keep `output_path` on shared storage visible to login and compute nodes.
* Start from small `models`, `n_trials`, `timeout`, and `n_splits` values.
* Use `tmux` or `screen` for any run that may outlive an SSH session.
* Confirm the Conda environment is available on compute nodes.
* Check `logs/` and `jobsCompleted/` before restarting a failed phase.
* Use `skip_completed_models = True` only for Phase 6 recovery runs where you do
not want to overwrite completed model artifacts.
14 changes: 8 additions & 6 deletions docs/source/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -7,16 +7,16 @@ Overview
--------------------------------------

STREAMLINE is an end-to-end automated machine learning pipeline for
supervised tabular data. The v1.0.0 main release supports binary classification,
supervised tabular data. The v1.0.1 release supports binary classification,
multiclass classification, and regression, with integrated
data processing, imputation, scaling, feature learning, feature importance,
feature selection, model training, classification ensembles, summary
statistics, dataset comparison, replication, and PDF reporting.

The schematic below summarizes the STREAMLINE v1.0.0 workflow.
The schematic below summarizes the STREAMLINE v1.0.1 workflow.

.. image:: pictures/STREAMLINE_v3_paper_new_lightcolor.png
:alt: STREAMLINE v1.0.0 automated machine learning pipeline overview
:alt: STREAMLINE v1.0.1 automated machine learning pipeline overview
:width: 100%

The repository is organized around eleven explicit phases:
Expand Down Expand Up @@ -80,8 +80,8 @@ For most users, the easiest local route is:
conda create -n streamline python=3.11 pip
conda activate streamline
pip install -r requirements.txt
python run.py -c run_configs/uci_binary_hcc.cfg --dry_run
python run.py -c run_configs/uci_binary_hcc.cfg
python run.py -c run_configs/local/uci_binary_hcc.cfg --dry_run
python run.py -c run_configs/local/uci_binary_hcc.cfg

The notebooks expose the same major settings as the config files and are a
better starting point for interactive tutorials, Colab demos, and custom data
Expand All @@ -93,6 +93,7 @@ How This Documentation Is Organized
* Use :doc:`install` to prepare a local environment.
* Use :doc:`data` to format custom datasets and understand the included UCI demos.
* Use :doc:`running` for notebooks, config-driven runs, and phase-by-phase CLI commands.
* Use :doc:`hpc` for SLURM/LSF configs, tmux basics, scheduler monitoring, and cluster recovery.
* Use :doc:`parameters` when editing ``.cfg`` files or command-line calls.
* Use :doc:`output` to navigate experiment folders and reports.
* Use :doc:`pipeline` for a phase-by-phase explanation of what STREAMLINE does.
Expand All @@ -101,7 +102,7 @@ How This Documentation Is Organized
Version History
--------------------------------------

This site documents the STREAMLINE v1.0.0 main release. See
This site documents the STREAMLINE v1.0.1 release. See
:doc:`changelog` for dated release entries and notable changes.

Current Scope
Expand Down Expand Up @@ -143,6 +144,7 @@ questions, contact Harsh Bandhey at ``harsh.bandhey@cshs.org``.
install
tabpfn_token
running
hpc
parameters
model_params_json
output
Expand Down
10 changes: 10 additions & 0 deletions docs/source/install.md
Original file line number Diff line number Diff line change
Expand Up @@ -92,6 +92,16 @@ depending on phase support. For long runs, use a persistent terminal session
such as `tmux` or `screen` so orchestration is not interrupted if your SSH
connection drops.

Use the scheduler templates in `run_configs/hpc/` as starting points:

```bash
python run.py -c run_configs/hpc/cedars_slurm_hcc.cfg --dry_run
python run.py -c run_configs/hpc/upenn_lsf_hcc.cfg --dry_run
```

See [HPC and Cluster Runs](hpc.md) for tmux basics, SLURM/LSF monitoring
commands, scheduler config fields, and recovery notes.

## Known Installation Issues

Some modeling and reporting packages include compiled dependencies. On macOS,
Expand Down
Loading
Loading