skfolio.seriation.HierarchicalSeriation#

class skfolio.seriation.HierarchicalSeriation(*, hierarchical_clustering_estimator=None, optimal_ordering=True)[source]#

Order assets by the leaves of a hierarchical clustering tree.

Parameters:
hierarchical_clustering_estimatorHierarchicalClustering, optional

Hierarchical clustering estimator, including compatible subclasses. Each fit trains a fresh clone when at least two assets are investable. The default (None) uses HierarchicalClustering with Ward linkage.

optimal_orderingbool, default=True

If True, minimize distances between adjacent leaves without changing the tree [1]. If False, use the leaf order produced by the clustering tree.

Attributes:
ordering_ndarray of shape (n_investable_assets,)

Positions in the original input matrix, listed in the computed order. Each investable asset appears exactly once. Empty when no assets are investable. A single investable asset produces its original position.

investable_mask_ndarray of shape (n_assets,)

Boolean mask selecting investable assets for the current ordering.

hierarchical_clustering_estimator_HierarchicalClustering or None

Fitted clustering estimator, or None with fewer than two investable assets.

ordered_linkage_matrix_ndarray of shape (max(n_investable_assets - 1, 0), 4)

Linkage matrix with the selected leaf ordering. Leaf indices refer to the compact input order, given by np.flatnonzero(investable_mask_).

n_features_in_int

Number of assets in the full schema.

feature_names_in_ndarray of shape (n_features_in_,)

Asset names, defined when the input names are all strings.

Methods

fit(X[, y])

Start a new hierarchical ordering from a distance snapshot.

get_metadata_routing()

Route fitting metadata to the clustering estimator.

get_params([deep])

Get parameters for this estimator.

set_params(**params)

Set the parameters of this estimator.

References

[1]

“Fast optimal leaf ordering for hierarchical clustering”. Ziv Bar-Joseph, David K. Gifford and Tommi S. Jaakkola, Bioinformatics (2001).

fit(X, y=None, **fit_params)[source]#

Start a new hierarchical ordering from a distance snapshot.

Parameters:
Xarray-like of shape (n_assets, n_assets)

Distance snapshot. NaN diagonal entries mark non-investable assets.

yIgnored

Not used, present for API consistency by convention.

**fit_paramsdict

Parameters to pass to the clustering estimator. Only available if enable_metadata_routing=True, which can be set by using sklearn.set_config(enable_metadata_routing=True). See Metadata Routing User Guide for more details.

Returns:
selfHierarchicalSeriation

Fitted estimator.

get_metadata_routing()[source]#

Route fitting metadata to the clustering estimator.

get_params(deep=True)#

Get parameters for this estimator.

Parameters:
deepbool, default=True

If True, will return the parameters for this estimator and contained subobjects that are estimators.

Returns:
paramsdict

Parameter names mapped to their values.

set_params(**params)#

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:
**paramsdict

Estimator parameters.

Returns:
selfestimator instance

Estimator instance.