Skip to content

OPLS-DA

Bases: ClassifierMixin, BaseEstimator

Binary OPLS Discriminant Analysis.

The two class labels are encoded as a -1/+1 dummy response and fitted with OPLS. decision_function returns the raw signed OPLS regression output (positive favours classes_[1]) and predict returns class labels from its sign. For class probabilities, wrap in CalibratedClassifierCV (cross-fitted, robust) when each class has enough samples for the chosen calibration CV split.

Parameters:

Name Type Description Default
n_components int

Number of predictive PLS components fitted on the orthogonally filtered X block by the inner OPLS.

1
n_orthogonal int

Number of X-orthogonal components removed before fitting the predictive PLS model. To choose this by cross-validated score, wrap OPLSDA in GridSearchCV over n_orthogonal.

1
scale ('none', 'center', 'pareto', 'standard')

Column preprocessing applied to X. Note: unlike the boolean scale parameter of PLSRegression, this is a string mode; passing True/False raises an error.

"none"
copy bool

Whether the input arrays are copied during validation. Note that copy=False is passed to sklearn input validation; OPLS filtering still allocates working arrays.

True

Attributes:

Name Type Description
classes_ ndarray

The two class labels seen during fit.

opls_ OPLS

The fitted underlying OPLS regressor against a -1/+1 dummy response.

n_orthogonal_ int

Number of orthogonal components actually used by the inner OPLS.

n_features_in_ int

Number of features seen during fit.

feature_names_in_ ndarray of shape (n_features_in_,)

Names of features seen during fit. Defined only when X has feature names that are all strings.

vip_, ortho_vip_ ndarray of shape (n_features,)

Predictive / orthogonal Variable Importance in Projection scores computed by the inner opls_ estimator. Use with SelectFromModel via importance_getter="vip_".

See Also

OPLS : Underlying OPLS regressor fitted against the -1/+1 dummy response. O2PLS : Two-block variant that also models Y-specific orthogonal structure. sklearn.calibration.CalibratedClassifierCV : Wrapper providing calibrated class probabilities from decision_function.

References

.. [1] Bylesjo, M., Rantalainen, M., Cloarec, O., Nicholson, J. K., Holmes, E. & Trygg, J. (2006). OPLS discriminant analysis: combining the strengths of PLS-DA and SIMCA classification. Journal of Chemometrics, 20(8-10), 341-351. https://doi.org/10.1002/cem.1006 .. [2] Trygg, J. & Wold, S. (2002). Orthogonal projections to latent structures (O-PLS). Journal of Chemometrics, 16(3), 119-128. https://doi.org/10.1002/cem.695

Examples:

>>> import numpy as np
>>> from scikit_opls import OPLSDA
>>> rng = np.random.default_rng(0)
>>> X = rng.normal(size=(20, 5))
>>> y = np.where(X[:, 0] > 0, "case", "control")
>>> clf = OPLSDA(n_orthogonal=1).fit(X, y)
>>> clf.classes_.tolist()
['case', 'control']
>>> clf.predict(X[:2]).shape
(2,)

vip_ property

vip_: NDArray[float64]

Predictive VIP per feature, delegated to the inner OPLS.

ortho_vip_ property

ortho_vip_: NDArray[float64]

Orthogonal VIP per feature, delegated to the inner OPLS.

fit

fit(X: ArrayLike, y: ArrayLike) -> OPLSDA

Fit the binary OPLS-DA classifier.

Parameters:

Name Type Description Default
X array-like of shape (n_samples, n_features)

Training predictors.

required
y array-like of shape (n_samples,)

Binary class labels.

required

Returns:

Name Type Description
self OPLSDA

The fitted estimator.

decision_function

decision_function(X: ArrayLike) -> NDArray[np.float64]

Raw signed OPLS regression output; positive favours classes_[1].

Parameters:

Name Type Description Default
X array-like of shape (n_samples, n_features)

Samples to score.

required

Returns:

Name Type Description
scores ndarray of shape (n_samples,)

Signed confidence; > 0 predicts classes_[1]. Scores equal to zero are assigned to classes_[0] by predict.

predict

predict(X: ArrayLike) -> NDArray

Predict class labels.

Parameters:

Name Type Description Default
X array-like of shape (n_samples, n_features)

Samples to classify.

required

Returns:

Name Type Description
y_pred ndarray of shape (n_samples,)

Predicted labels drawn from classes_.

score_distance

score_distance(
    X: ArrayLike, *, kind: str = "predictive"
) -> NDArray[np.float64]

Return Hotelling-like score distances from the inner OPLS model.

Parameters:

Name Type Description Default
X array-like of shape (n_samples, n_features)

Samples in raw feature space. Pass raw X. Do not manually center or scale before calling diagnostics. The outer classifier validates feature names and feature counts before delegating to the fitted inner OPLS model.

required
kind ('predictive', 'orthogonal', 'all')

Which coordinate space to use.

"predictive"

Returns:

Name Type Description
score_dist ndarray of shape (n_samples,)

Computed Hotelling-like distance per sample.

q_residuals

q_residuals(
    X: ArrayLike, *, space: str = "full"
) -> NDArray[np.float64]

Return Q residuals from the inner OPLS model.

Parameters:

Name Type Description Default
X array-like of shape (n_samples, n_features)

Samples in raw feature space. Pass raw X. Do not manually center or scale before calling diagnostics. The outer classifier validates feature names and feature counts before delegating to the fitted inner OPLS model.

required
space ('full', 'predictive')

Which model reconstruction space to use.

"full"

Returns:

Name Type Description
q ndarray of shape (n_samples,)

Squared residual norm per sample.