Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions docs/changes.rst
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,18 @@ v0.17
with an error while it is read, before anything in it is audited or
constructed, instead of failing with an unrelated error, or being accepted,
during construction. :pr:`547` by `Adrin Jalali`_.
- Add support for pandas objects: :class:`~pandas.DataFrame`,
:class:`~pandas.Series`, every kind of :class:`~pandas.Index`, extension
arrays and extension dtypes can now be saved and loaded. They are stored as
the numpy arrays and scalars they are made of and rebuilt through the public
pandas constructors, so no pandas internals end up in the file, and they are
trusted by default. Estimators from other libraries that keep pandas objects
in their fitted attributes, such as ``category_encoders``, can now be
persisted. A file written with one pandas version loads with any other from
2.0 on, keeping the dtypes of the version that wrote it. Not preserved are
the ``freq`` of datetime-like indexes and arrays, the ``attrs`` and ``flags``
of a Series or DataFrame, and the storage, python or pyarrow, of a string
dtype. :pr:`552` by `Adrin Jalali`_.
- Fix a regression since v0.12.0 where saving an object whose ``__reduce__``
raises failed at dump time. ``__reduce__`` is called on every object to
detect a plain constructor call, but Cython extension types with a
Expand Down
10 changes: 9 additions & 1 deletion docs/persistence.rst
Original file line number Diff line number Diff line change
Expand Up @@ -250,7 +250,15 @@ Supported libraries
Skops intends to support all of **scikit-learn**, that is, not only its
estimators, but also other classes like cross validation splitters. Furthermore,
most types from **numpy** and **scipy** should be supported, such as (sparse)
arrays, dtypes, random generators, and ufuncs.
arrays, dtypes, random generators, and ufuncs. **pandas** objects, that is
``DataFrame``, ``Series``, every kind of ``Index``, extension arrays and
extension dtypes, are supported as well with pandas 2.0 or later: they are
stored as the arrays they are made of and rebuilt through the public pandas
constructors, so that no pandas internals end up in the file, and a file
written with one pandas version loads with any other. Not preserved are the
``freq`` of datetime-like indexes and arrays, the ``attrs`` and ``flags`` of a
``Series`` or ``DataFrame``, and the storage, python or pyarrow, of a string
dtype, which is an environment choice over the same values.

Apart from this core, we plan to support machine learning libraries commonly
used be the community. So far, we have tested the following libraries:
Expand Down
2 changes: 1 addition & 1 deletion docs/requirements.txt
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# to be synced with the versions in pyproject.toml
matplotlib>=3.3
pandas>=1
pandas>=2
fairlearn>=0.7.0
sphinx>=3.2.0
sphinx-gallery>=0.7.0
Expand Down
1,360 changes: 1,086 additions & 274 deletions pixi.lock

Large diffs are not rendered by default.

17 changes: 13 additions & 4 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -131,7 +131,6 @@ skops = { path = ".", editable = true }
[tool.pixi.feature.docs.dependencies]
# To be synced with the versions in docs/requirements.txt
matplotlib = ">=3.3"
pandas = ">=1"
sphinx = ">=3.2.0"
sphinx-gallery = ">=0.7.0"
sphinx-rtd-theme = ">=1"
Expand All @@ -141,8 +140,10 @@ sphinx-issues = ">=1.2.0"

[tool.pixi.feature.docs.pypi-dependencies]
# everything that depends on scikit-learn needs to be a pypi dependency so that this
# spec is compatible with the nightly build environment.
# spec is compatible with the nightly build environment. The same holds for pandas,
# whose dev version the nightly build environment installs from pypi.
fairlearn = ">=0.7.0"
pandas = ">=2"

[tool.pixi.feature.tests.dependencies]
pytest = ">=7"
Expand All @@ -151,14 +152,18 @@ flaky = ">=3.7.0"
pandoc = ">=3.6.4"
rich = ">=12"
matplotlib = ">=3.3"
pandas = ">=1"

[tool.pixi.feature.tests.pypi-dependencies]
# these are packages that require scikit-learn. They need to be as a pypi dependency
# because otherwise there will be a package resolution conflict between pypi and conda
# when installing pre-release nightly release.
lightgbm = ">=3"
xgboost = ">=1.6"
# skops.io supports pandas 2.0 and later; each CI environment pins one minor
# version so that the whole range is tested, see the sklearn* features below.
# A pypi dependency for the same reason as above: the nightly environment
# installs the dev version of pandas from pypi.
pandas = ">=2"

[tool.pixi.feature.lint.dependencies]
pre-commit = "*"
Expand Down Expand Up @@ -242,6 +247,8 @@ numpy = "~=2.5.0"
scipy = "~=1.18.0"
catboost = ">=1.0"
quantile-forest = "~=1.4.0"
# keeps pandas objects in fitted attributes, see the test for issue #450
category_encoders = ">=2.6"
python = "~=3.14.0"

# [tool.pixi.feature.sklearn17]
Expand All @@ -260,7 +267,9 @@ extra-index-urls = ["https://pypi.anaconda.org/scientific-python-nightly-wheels/
# The version value here needs to be exact, hence == instead of ~=
scikit-learn = "==1.10.dev0"
fairlearn = "*"
pandas = "*"
# The dev version of pandas from the nightly index; "*" would pick the latest
# release, since pre-releases are only considered when named explicitly.
pandas = "==3.1.0.dev0"
numpy = "*"
scipy = "*"

Expand Down
Loading
Loading