This is development documentation. Some features are not available in the stable release. View stable documentation.

Pipeline

Pipeline composes a small, reproducible sequence of SpectroChemPy preprocessing transformers and a final transformer or supervised estimator. It is useful when you want to fit preprocessing and a model together, then reuse exactly the preprocessing learned from calibration spectra on new spectra.

The first public version is intentionally modest. Its goal is not to replace a workflow engine or scikit-learn’s full Pipeline API, but to make common SpectroChemPy estimator workflows easier to repeat without data leakage: preprocessing steps are fitted only when Pipeline.fit() is called, so test spectra do not accidentally influence centering or scaling statistics.

The synthetic concentrations, spectral profiles, and noise below remain NDDataset objects. Reshaping concentration as a column lets the public arithmetic operations broadcast it over the spectral dimension while preserving the sample and spectral coordinates.

[1]:
import spectrochempy as scp

samples = scp.Coord.arange(12, title="sample")
concentration = scp.linspace(
    0.1,
    1.2,
    samples.size,
    coordset=[samples],
    dims=["y"],
)

wavenumbers = scp.linspace(
    1000.0,
    1200.0,
    80,
    dims=["x"],
    units="cm^-1",
)
wavenumbers.set_coordset(x=scp.Coord(wavenumbers, title="wavenumber"))
band_a = scp.exp(
    -0.5 * ((wavenumbers - 1060.0 * scp.ur("cm^-1")) / (12.0 * scp.ur("cm^-1"))) ** 2
)
band_b = scp.exp(
    -0.5 * ((wavenumbers - 1140.0 * scp.ur("cm^-1")) / (18.0 * scp.ur("cm^-1"))) ** 2
)
baseline = (wavenumbers - wavenumbers.mean()) * (0.0005 * scp.ur.cm)
baseline += 0.03

concentration_column = concentration.reshape((samples.size, 1), dims=("y", "x"))
contribution_a = concentration_column * band_a
contribution_b = 0.35 * concentration_column * band_b

noise = scp.normal(
    scale=0.015,
    size=(samples.size, wavenumbers.size),
    seed=42,
    dims=["y", "x"],
)
dataset = contribution_a + baseline + contribution_b + noise
dataset.units = "absorbance"
dataset.title = "calibration spectra"

target = concentration.copy()
target.units = "mol/L"
target.title = "concentration"
_ = dataset.plot(show=False)
../../_images/userguide_analysis_pipeline_1_1.png

Broadcasting is positional and right-aligned: dimension names do not align or reorder operands, and coordinates are never interpolated. A singleton axis inherits the name and coordinate of the operand providing the non-singleton axis. Duplicate result dimension names are rejected explicitly.

A complete calibration/test example

A typical supervised use is to fit preprocessing and regression together on a calibration subset, then apply the fitted pipeline to separate test spectra. The scaler below learns its statistics from X_cal only; those fitted statistics are reused when predicting X_test.

[2]:
test_indices = [2, 5, 8, 11]
cal_indices = [i for i in range(concentration.size) if i not in test_indices]

X_cal = dataset[cal_indices]
y_cal = target[cal_indices]
X_test = dataset[test_indices]
y_test = target[test_indices]

regression_pipeline = scp.Pipeline(
    [
        ("scale", scp.AutoscaleTransformer(dim="y")),
        ("pls", scp.PLSRegression(n_components=2, scale=False)),
    ]
)

regression_pipeline.fit(X_cal, y_cal)
y_pred = regression_pipeline.predict(X_test)
residuals = y_test - y_pred
rmse = float(((residuals**2).mean() ** 0.5).magnitude)

summary = scp.concatenate(y_test, y_pred, residuals, axis=1)
summary.x = scp.Coord.arange(
    3,
    labels=["observed", "predicted", "residual"],
    title="quantity",
)
summary.title = f"test predictions, RMSE = {rmse:.3f}"
summary
[2]:
NDDataset — float64, shape: (y:4, x:3), mol⋅l⁻¹
name
:
NDDataset_7a52be68
author
:
runner@runnervmtr4k5
created
:
2026-09-29 19:17:57+00:00
description
:
Concatenation of 3 datasets:
( , NDDataset_79cf210a_PLSRegression.prediction, )
history
:
2026-09-29 19:17:57+00:00> Created by concatenate from 3 datasets: , NDDataset_79cf210a_PLSRegression.prediction,
Data
title
:
test predictions, RMSE = 0.019
values
:
[[ 0.3 0.3079 -0.007874]
[ 0.6 0.5957 0.004253]
[ 0.9 0.8976 0.002387]
[ 1.2 1.163 0.03665]] mol⋅l⁻¹
shape
:
(y:4, x:3)
Dimension `x`
size
:
3
title
:
quantity
coordinates
:
[ 0 1 2]
labels
:
[ observed predicted residual]
Dimension `y`
size
:
4
title
:
sample
coordinates
:
[ 2 5 8 11]

The fitted final estimator can still be inspected directly. Here we reuse its parity-plot helper and add the independent test predictions in red.

[3]:
fitted_pls = regression_pipeline.fitted_named_steps_["pls"]
ax = fitted_pls.plot_parity(label="calibration", s=120, show=False)
_ = fitted_pls.plot_parity(
    y_test,
    y_pred,
    ax=ax,
    s=120,
    c="red",
    label="test",
    clear=False,
    show=False,
)
_ = ax.legend(loc="lower right")
../../_images/userguide_analysis_pipeline_6_0.png

Transformer-final pipelines

A transformer-final pipeline ends with a preprocessing transformer or with PCA. fit_transform(X) is equivalent to fit(X).transform(X).

[4]:
pca_pipeline = scp.Pipeline(
    [
        ("center", scp.CenterTransformer(dim="y")),
        ("pca", scp.PCA(n_components=3)),
    ]
)
scores = pca_pipeline.fit_transform(dataset)
scores
[4]:
NDDataset [NDDataset_7a52bf25_PCA.scores] — float64, shape: (y:12, k:3)
name
:
NDDataset_7a52bf25_PCA.scores
author
:
runner@runnervmtr4k5
created
:
2026-09-29 19:17:57+00:00
description
:
scores from PCA fit of NDDataset_7a52bf25.
history
:
2026-09-29 19:17:57+00:00> Created analysis output scores with PCA from NDDataset_7a52bf25.
Data
title
:
scores
values
:
[[ -1.722 -0.0005223 -0.05279]
[ -1.429 -0.02479 -0.07234]
...
[ 1.403 0.03685 -0.004356]
[ 1.754 -0.06336 -0.05766]]
shape
:
(y:12, k:3)
Dimension `k`
size
:
3
title
:
components
labels
:
[ PC1 PC2 PC3]
Dimension `y`
size
:
12
title
:
sample
coordinates
:
[ 0 1 ... 10 11]

Template and fitted steps

The steps passed to Pipeline are templates. Calling fit() clones those templates, fits the clones, and leaves the original step objects unchanged. Template steps remain available through steps and named_steps. Learned runtime state is available after fitting through fitted_steps_ and fitted_named_steps_.

[5]:
print(
    regression_pipeline.named_steps["scale"]
    is regression_pipeline.fitted_named_steps_["scale"]
)
print(regression_pipeline.steps[0][1] is regression_pipeline.fitted_steps_[0][1])
False
False

Nested parameters

Template parameters are visible and editable with step__parameter names, using the step name followed by a double underscore and the parameter name. Any effective change invalidates the fitted state, so you must call fit() again before using predict() or transform().

[6]:
regression_pipeline.get_params(deep=True)["pls__n_components"]
regression_pipeline.set_params(pls__n_components=1)
regression_pipeline.fit(X_cal, y_cal)
regression_pipeline.predict(X_test)
[6]:
NDDataset [NDDataset_7a52bfcf_PLSRegression.prediction] — float64, size: 4, mol⋅l⁻¹
name
:
NDDataset_7a52bfcf_PLSRegression.prediction
author
:
runner@runnervmtr4k5
created
:
2026-09-29 19:17:57+00:00
description
:
Prediction from PLSRegression fit of NDDataset_7a52bf8d + NDDataset_79cf2087 applied to NDDataset_7a52bfcf.
history
:
2026-09-29 19:17:57+00:00> Created analysis output prediction with PLSRegression from NDDataset_7a52bf8d + NDDataset_79cf2087; applied to NDDataset_7a52bfcf.
Data
title
:
prediction
values
:
[ 0.3072 0.5908 0.8955 1.159] mol⋅l⁻¹
size
:
4
Dimension `y`
size
:
4
title
:
sample
coordinates
:
[ 2 5 8 11]

What Pipeline v1 does and does not do

Pipeline is deliberately linear. Steps are ordered (name, step) pairs, intermediate steps must transform an NDDataset into another NDDataset, and the final step is either a transformer or a supervised estimator. A final transformer supports transform() and fit_transform(). A final supervised estimator supports predict() and, when available on that estimator, score().

Pipeline does not choose train/test splits, run cross-validation, search hyperparameters, cache intermediate results, branch into multiple paths, or route arbitrary fit parameters. For cross-validation, the split loop must fit a fresh pipeline inside each fold so that every preprocessing statistic stays fold-local.

Supported v1 classes

Intermediate positions accept CenterTransformer, AutoscaleTransformer, ParetoScaleTransformer, RangeScaleTransformer, RobustScaleTransformer, SNVTransformer, NormalizeTransformer, MSCTransformer and LogTransformer.

Final positions accept those transformers, PCA, PLSRegression, LSTSQ and NNLS.

Version 1 intentionally excludes Baseline, SVD, PSD, Filter, functional processing wrappers, MCRALS, Optimize, optional steps, branching, caching, persistence guarantees and arbitrary fit-parameter routing.