Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 15 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ Package information: ![Python 3.9](https://img.shields.io/badge/python-3.9-blue.

scikit-FIBERS can be used directly as a **modeling strategy**, by training a bin population and using the predict() function to apply the discovered bin with the highest fitness as a predictive model of risk group assigment. It can also be used as a **feature learning algorithm**, by training a bin population and using the transform() function to convert each discovered bin in the population into corresponding dataset features for additional downstream machine learning modeling.

The scikit-FIBERS algorithm seeks to automatically identify and optimize a population of 'candidate bins' that maximize time-to-event differences between high and low risk groups. A 'bin' is a subset of features and an associated 'burden threshold' that together differentiate instances into high vs. low risk instance groups. Instances that have a bin sum (of feature values) greater than the threshold are assigend to the high-risk group, and all others to the low-risk group. The fitness (i.e. quality) of bins in the candidate bin population drives evolutionary algorithm learning.
The scikit-FIBERS algorithm seeks to automatically identify and optimize a population of 'candidate bins' that maximize time-to-event differences between high and low risk groups. A 'bin' is a subset of features and an associated 'burden threshold' that together differentiate instances into two risk groups. Instances with a bin sum greater than the threshold are assigned to the above-threshold group, and all others to the below-threshold group. The fitness (i.e. quality) of bins in the candidate bin population drives evolutionary algorithm learning.

A schematic detailing how the scikit-FIBERS algorithm works is given below:

Expand Down Expand Up @@ -145,6 +145,20 @@ While scikit_FIBERS has a number of available hyperparameters only a few are con
| *manual_bin_init* | Dataframe of FIBERS-formatted bin population used to initialize the bin population | dataframe, None | None |
| *pop_clean* | Optional bin population cleanup strategy | 'group_strata', None | None |

### Directional bin-effect modes

For every candidate bin and burden threshold, FIBERS separates the training samples into a below-threshold group (bin sum <= threshold) and an above-threshold group (bin sum > threshold). In the directional modes, it fits Kaplan-Meier survival estimates for both groups and compares their restricted mean survival time (RMST). RMST is evaluated through the latest follow-up time shared by both groups, which makes the direction check censoring-aware.

* `desired_bin_effect="default"` preserves the original, non-directional behavior. Bins are ranked using the selected fitness metric without requiring the above-threshold group to have better or worse survival.
* `desired_bin_effect="protective"` requires the above-threshold group to have a strictly higher RMST than the below-threshold group. A larger burden of the bin's features is therefore associated with better survival.
* `desired_bin_effect="high_risk"` requires the above-threshold group to have a strictly lower RMST than the below-threshold group. A larger feature burden is therefore associated with worse survival.

If a bin points in the wrong direction, its applicable log-rank and/or residual fitness score is set to zero. With adaptive thresholding (`group_thresh=None`), FIBERS first looks for the best threshold that satisfies both the requested direction and *group_strata_min*. If none does, it prefers a directionally valid threshold and applies the group-balance penalty. If no threshold has the requested direction, the fallback remains directionally invalid and receives zero directional fitness.

The binary output convention does not change between modes: `predict()` and `transform(..., full_sums=False)` encode the above-threshold group as `1` and the below-threshold group as `0`. Consequently, `1` identifies the protective group in `protective` mode and the high-risk group in `high_risk` mode.

RMST acts as a direction gate rather than the optimization score: candidates that pass are still ranked by the selected log-rank, residual, or product fitness metric. The RMST comparison is based on the observed survival and censoring columns; covariates can affect residual-based fitness, but they do not adjust the direction gate itself. The constraint is applied during initialization and offspring evaluation throughout training, not only as a post-hoc filter. See the [protective and high-risk mode guide](docs/source/directional_modes.md) for the complete evaluation sequence, fallback rules, a worked example, and interpretation cautions.

* The remaining hyperparameters in the table below can largely be left to their default values by most users.

| Hyperparameter | Description | Type/Options | Default Value |
Expand Down
2 changes: 1 addition & 1 deletion docs/source/conf.py
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@
author = 'Ryan Urbanowicz, Harsh Bandhey'

# The full version, including alpha/beta/rc tags
release = '1.0.0'
release = '2.2.1'


# -- General configuration ---------------------------------------------------
Expand Down
143 changes: 143 additions & 0 deletions docs/source/directional_modes.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,143 @@
# Protective and high-risk modes

The *desired_bin_effect* parameter lets FIBERS restrict its evolutionary search to bins whose above-threshold group has a requested survival direction. The direction check is applied while candidate bins are evaluated; it is not merely a label assigned after training.

The supported values are:

| Mode | Required survival direction for the above-threshold group | Meaning of binary output `1` |
| ---- | --------------------------------------------------------- | ---------------------------- |
| `default` | No direction is required | Above-threshold membership only |
| `protective` | Better survival than the below-threshold group | Protective-burden membership |
| `high_risk` | Worse survival than the below-threshold group | High-risk-burden membership |

## Forming the two threshold groups

For a candidate bin, FIBERS first sums the values of the features included in that bin for each sample. Given a candidate threshold, the samples are separated as follows:

* **Below-threshold group:** bin sum <= threshold.
* **Above-threshold group:** bin sum > threshold.

The terms *above threshold* and *below threshold* describe burden membership only. They should not automatically be interpreted as high and low risk because the above-threshold group may be protective.

## Censoring-aware direction test

For `protective` and `high_risk`, FIBERS fits a Kaplan-Meier survival estimate for each threshold group. It then chooses a shared evaluation horizon:

```
shared_time = min(maximum below-threshold duration,
maximum above-threshold duration)
```

The restricted mean survival time (RMST) for each group is calculated from time zero through `shared_time`. Restricting the comparison to this shared horizon avoids comparing one group beyond the follow-up supported by the other group and allows censored samples to contribute through the Kaplan-Meier estimate.

The direction rules use strict inequalities:

* `protective` is valid when `above_RMST > below_RMST`.
* `high_risk` is valid when `above_RMST < below_RMST`.

Equal RMST values do not satisfy either directional mode. A threshold is also treated as directionally invalid if a group is empty or its RMST cannot be calculated.

## Direction is a gate, not the fitness score

RMST determines whether a candidate has the requested direction, but the size of the RMST difference is not itself the optimization score. Candidates that pass the direction gate continue to be ranked using the configured *fitness_metric*:

| Fitness metric | Score for a directionally valid candidate | Score for a wrong-direction candidate |
| -------------- | ----------------------------------------- | --------------------------------------- |
| `log_rank` | Log-rank test statistic | `0` |
| `residuals` | Absolute rank-sum statistic for deviance residuals | `0` |
| `log_rank_residuals` | Log-rank score multiplied by the residual score | `0` |

The associated p-value is unavailable for a score that was set to zero by the directional gate. A candidate can satisfy the RMST direction with a small difference, so directionally valid does not by itself mean statistically significant. Statistical separation is still represented by the selected fitness statistic and its associated output.

The RMST direction check uses the observed survival and censoring columns and is therefore an unadjusted direction check. If covariates are supplied, `residuals` or `log_rank_residuals` can incorporate covariate-adjusted deviance residuals into candidate ranking, but the protective/high-risk gate itself remains based on the two Kaplan-Meier RMST estimates.

## Fixed and adaptive thresholds

When *group_thresh* is a fixed integer, FIBERS evaluates that threshold directly. If the resulting groups do not have the requested direction, the applicable fitness score is zero.

When `group_thresh=None`, FIBERS evaluates the candidate thresholds from *min_thresh* through *max_thresh*. Directional modes use the following selection order:

1. Select the highest-scoring threshold that satisfies both the requested RMST direction and *group_strata_min*.
2. If none satisfies both requirements, prefer the highest-scoring directionally valid threshold even when its groups are too imbalanced. The regular group-balance penalty is applied, followed by an additional fallback penalty.
3. If no threshold has the requested direction, retain the best available fallback for algorithm continuity. Its directional fitness remains zero.

With the default penalty calculation, a directionally valid fallback that misses *group_strata_min* can receive two multiplicative `(1 - penalty)` reductions: one for group imbalance and one for requiring the fallback.

This threshold logic keeps the requested biological direction as the priority while still allowing the evolutionary process to continue when a candidate cannot simultaneously satisfy the direction and balance constraints.

## Where the direction constraint is applied

The selected mode is used throughout training:

* When random bins are initialized.
* When a manually supplied starting population is evaluated.
* When crossover, mutation, or merge operations create offspring.
* When adaptive thresholds are re-evaluated, including the final iteration.

As a result, `protective` and `high_risk` influence which candidates receive useful fitness throughout evolution rather than filtering only the final population.

## Prediction and transformation semantics

FIBERS does not reverse its binary encoding in protective mode:

```
0 = bin sum <= threshold
1 = bin sum > threshold
```

Therefore:

* In `protective` mode, a prediction of `1` means membership in the group with better censoring-aware survival during training.
* In `high_risk` mode, a prediction of `1` means membership in the group with worse censoring-aware survival during training.
* In `default` mode, a prediction of `1` means above-threshold membership without a guaranteed survival direction.

The same convention is used by `transform(..., full_sums=False)`. When `full_sums=True`, the transformed feature contains the raw bin sum instead of a binary group assignment.

## Worked example

Suppose a candidate bin contains three features and uses a threshold of `1`. Samples with zero or one included feature value are below threshold, while samples with a bin sum of two or more are above threshold.

If the Kaplan-Meier estimates produce the following RMST values over their shared follow-up horizon:

| Group | RMST |
| ----- | ---- |
| Below threshold | 5.7 years |
| Above threshold | 8.4 years |

the candidate passes `protective` because `8.4 > 5.7`. The same candidate fails `high_risk` and receives zero directional fitness in that mode. If the RMST values were reversed, it would pass `high_risk` and fail `protective`.

After passing the direction gate, the candidate still competes using its configured log-rank, residual, or product fitness score. RMST establishes the direction; the fitness metric determines its evolutionary ranking.

## Configuration example

```python
common_options = {
"outcome_label": "Duration",
"censor_label": "Censoring",
"fitness_metric": "log_rank",
"group_thresh": None,
"min_thresh": 0,
"max_thresh": 5,
"group_strata_min": 0.2,
}

protective_model = FIBERS(
**common_options,
desired_bin_effect="protective",
).fit(train_data)

high_risk_model = FIBERS(
**common_options,
desired_bin_effect="high_risk",
).fit(train_data)
```

Use separate models when both directions are scientifically relevant. Comparing the discovered populations can reveal whether different feature combinations characterize protective and high-risk burdens.

## Interpretation cautions

* A directional result is an association in the training data and does not establish a causal protective or harmful effect.
* The direction gate does not replace statistical assessment, validation on held-out data, or replication in an independent cohort.
* The RMST gate is unadjusted even when residual-based fitness incorporates covariates.
* `default` remains the appropriate choice when either direction is meaningful or backward-compatible behavior is required.

5 changes: 2 additions & 3 deletions docs/source/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ FIBERS (Feature Inclusion Bin Evolver for Risk Stratification) is an evolutionar

scikit-FIBERS can be used directly as a 'modeling strategy', by training a bin population and using the predict() function to apply the discovered bin with the highest fitness as a predictive model of risk group assigment. It can also be used as a 'feature learning algorithm', by training a bin population and using the transform() function to convert each discovered bin in the population into corresponding dataset features for additional downstream machine learning modeling.

The scikit-FIBERS algorithm seeks to automatically identify and optimize a population of 'candidate bins' that maximize time-to-event differences between high and low risk groups. A 'bin' is a subset of features and an associated 'burden threshold' that together differentiate instances into high vs. low risk instance groups. Instances that have a bin sum (of feature values) greater than the threshold are assigend to the high-risk group, and all others to the low-risk group. The fitness (i.e. quality) of bins in the candidate bin population drives evolutionary algorithm learning.
The scikit-FIBERS algorithm seeks to automatically identify and optimize a population of 'candidate bins' that maximize time-to-event differences between high and low risk groups. A 'bin' is a subset of features and an associated 'burden threshold' that together differentiate instances into two risk groups. Instances with a bin sum greater than the threshold are assigned to the above-threshold group, and all others to the below-threshold group. The fitness (i.e. quality) of bins in the candidate bin population drives evolutionary algorithm learning.

A schematic detailing how the scikit-FIBERS algorithm works is given below:

Expand Down Expand Up @@ -60,8 +60,7 @@ inquiries related to scikit-FIBERS.
data
running
parameters
directional_modes
history
citation
modules


Loading
Loading