<a id="skfolio-moments-graphicallassocv"></a>

# skfolio.moments.GraphicalLassoCV

<a id="skfolio.moments.GraphicalLassoCV"></a>

### *class* skfolio.moments.GraphicalLassoCV(alphas=4, n_refinements=4, cv=None, tol=0.0001, enet_tol=0.0001, max_iter=100, mode='cd', n_jobs=None, verbose=False, assume_centered=False, nearest=True, higham=False, higham_max_iteration=100)

Sparse inverse covariance with cross-validated choice of the l1 penalty.

Read more in [scikit-learn](https://scikit-learn.org/stable/auto_examples/covariance/plot_sparse_cov.html).

* **Parameters:**
  **alphas** *int or array-like of shape (n_alphas,), dtype=float, default=4*
  : If an integer is given, it fixes the number of points on the
    grids of alpha to be used. If a list is given, it gives the
    grid to be used. See the notes in the class docstring for
    more details. Range is [1, inf) for an integer.
    Range is (0, inf] for an array-like of floats.

  **n_refinements** *int, default=4*
  : The number of times the grid is refined. Not used if explicit
    values of alphas are passed. Range is [1, inf).

  **cv** *int, cross-validation generator or iterable, default=None*
  : Determines the cross-validation splitting strategy.
    Possible inputs for cv are:
    - None, to use the default 5-fold cross-validation,
    - integer, to specify the number of folds.
    - `CV splitter`,
    - An iterable yielding (train, test) splits as arrays of indices.
    <br/>
    For integer/None inputs `KFold` is used.

  **tol** *float, default=1e-4*
  : The tolerance to declare convergence: if the dual gap goes below
    this value, iterations are stopped. Range is (0, inf].

  **enet_tol** *float, default=1e-4*
  : The tolerance for the elastic net solver used to calculate the descent
    direction. This parameter controls the accuracy of the search direction
    for a given column update, not of the overall parameter estimate. Only
    used for mode=’cd’. Range is (0, inf].

  **max_iter** *int, default=100*
  : Maximum number of iterations.

  **mode** *{‘cd’, ‘lars’}, default=’cd’*
  : The Lasso solver to use: coordinate descent or LARS. Use LARS for
    very sparse underlying graphs, where number of features is greater
    than number of samples. Elsewhere prefer cd which is more numerically
    stable.

  **n_jobs** *int, default=None*
  : Number of jobs to run in parallel.
    `None` means 1 unless in a `joblib.parallel_backend` context.
    `-1` means using all processors.

  **verbose** *bool, default=False*
  : If verbose is True, the objective function and duality gap are
    printed at each iteration.

  **assume_centered** *bool, default=False*
  : If True, data are not centered before computation.
    Useful when working with data whose mean is almost, but not exactly
    zero.
    If False, data are centered before computation.
* **Attributes:**
  **covariance_** *ndarray of shape (n_assets, n_assets)*
  : Estimated covariance.

  **location_** *ndarray of shape (n_assets,)*
  : Estimated location, i.e. the estimated mean.

  **precision_** *ndarray of shape (n_assets, n_assets)*
  : Estimated pseudo inverse matrix.
    (stored only if store_precision is True)

  **alpha_** *float*
  : Penalization parameter selected.

  **cv_results_** *dict of ndarrays*
  : A dict with keys:
    <br/>
    alphas *ndarray of shape (n_alphas,)*
    : All penalization parameters explored.
    <br/>
    split(k)_test_score *ndarray of shape (n_alphas,)*
    : Log-likelihood score on left-out data across (k)th fold.
      <br/>
      #### Versionadded
      Added in version 1.0.
    <br/>
    mean_test_score *ndarray of shape (n_alphas,)*
    : Mean of scores over the folds.
      <br/>
      #### Versionadded
      Added in version 1.0.
    <br/>
    std_test_score *ndarray of shape (n_alphas,)*
    : Standard deviation of scores over the folds.
      <br/>
      #### Versionadded
      Added in version 1.0.

  **n_iter_** *int*
  : Number of iterations run for the optimal alpha.

  **n_features_in_** *int*
  : Number of assets seen during `fit`.

  **feature_names_in_** *ndarray of shape (`n_features_in_`,)*
  : Names of features seen during `fit`. Defined only when `X`
    has feature names that are all strings.

### Methods

| [`error_norm`](#skfolio.moments.GraphicalLassoCV.error_norm)(comp_cov[, norm, scaling, squared])   | Compute the Mean Squared Error between two covariance estimators.          |
|---------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------|
| [`fit`](#skfolio.moments.GraphicalLassoCV.fit)(X[, y])                                      | Fit the GraphicalLasso covariance model to X.                              |
| [`get_metadata_routing`](#skfolio.moments.GraphicalLassoCV.get_metadata_routing)()                           | Get metadata routing of this object.                                       |
| [`get_params`](#skfolio.moments.GraphicalLassoCV.get_params)([deep])                               | Get parameters for this estimator.                                         |
| [`get_precision`](#skfolio.moments.GraphicalLassoCV.get_precision)()                                  | Getter for the precision matrix.                                           |
| [`mahalanobis`](#skfolio.moments.GraphicalLassoCV.mahalanobis)(X_test)                              | Compute the squared Mahalanobis distance of observations.                  |
| [`score`](#skfolio.moments.GraphicalLassoCV.score)(X_test[, y])                               | Compute the mean log-likelihood of observations under the estimated model. |
| [`set_params`](#skfolio.moments.GraphicalLassoCV.set_params)(\*\*params)                           | Set the parameters of this estimator.                                      |
| [`set_score_request`](#skfolio.moments.GraphicalLassoCV.set_score_request)()                              | No-op.                                                                     |

### Notes

The search for the optimal penalization parameter (`alpha`) is done on an
iteratively refined grid: first the cross-validated scores on a grid are
computed, then a new refined grid is centered around the maximum, and so
on.

One of the challenges which is faced here is that the solvers can
fail to converge to a well-conditioned estimate. The corresponding
values of `alpha` then come out as missing values, but the optimum may
be close to these missing values.

In `fit`, once the best parameter `alpha` is found through
cross-validation, the model is fit again using the entire training set.

<a id="skfolio.moments.GraphicalLassoCV.error_norm"></a>

#### error_norm(comp_cov, norm='frobenius', scaling=True, squared=True)

Compute the Mean Squared Error between two covariance estimators.

* **Parameters:**
  **comp_cov** *array-like of shape (n_features, n_features)*
  : The covariance to compare with.

  **norm** *{“frobenius”, “spectral”}, default=”frobenius”*
  : The type of norm used to compute the error. Available error types:
    - ‘frobenius’ (default): sqrt(tr(A^t.A))
    - ‘spectral’: sqrt(max(eigenvalues(A^t.A))
    where A is the error `(comp_cov - self.covariance_)`.

  **scaling** *bool, default=True*
  : If True (default), the squared error norm is divided by n_features.
    If False, the squared error norm is not rescaled.

  **squared** *bool, default=True*
  : Whether to compute the squared error norm or the error norm.
    If True (default), the squared error norm is returned.
    If False, the error norm is returned.
* **Returns:**
  **result** *float*
  : The Mean Squared Error (in the sense of the Frobenius norm) between
    `self` and `comp_cov` covariance estimators.

<a id="skfolio.moments.GraphicalLassoCV.fit"></a>

#### fit(X, y=None, \*\*fit_params)

Fit the GraphicalLasso covariance model to X.

* **Parameters:**
  **X** *array-like of shape (n_observations, n_assets)*
  : Price returns of the assets.

  **y** *Ignored*
  : Not used, present for API consistency by convention.
* **Returns:**
  **self** *GraphicalLassoCV*
  : Fitted estimator.

<a id="skfolio.moments.GraphicalLassoCV.get_metadata_routing"></a>

#### get_metadata_routing()

Get metadata routing of this object.

Please check [User Guide](https://skfolio.org/user_guide/metadata_routing.html.md#metadata-routing) on how the routing
mechanism works.

#### Versionadded
Added in version 1.5.

* **Returns:**
  **routing** *MetadataRouter*
  : A `MetadataRouter` encapsulating
    routing information.

<a id="skfolio.moments.GraphicalLassoCV.get_params"></a>

#### get_params(deep=True)

Get parameters for this estimator.

* **Parameters:**
  **deep** *bool, default=True*
  : If True, will return the parameters for this estimator and
    contained subobjects that are estimators.
* **Returns:**
  **params** *dict*
  : Parameter names mapped to their values.

<a id="skfolio.moments.GraphicalLassoCV.get_precision"></a>

#### get_precision()

Getter for the precision matrix.

* **Returns:**
  **precision_** *array-like of shape (n_features, n_features)*
  : The precision matrix associated to the current covariance object.

<a id="skfolio.moments.GraphicalLassoCV.mahalanobis"></a>

#### mahalanobis(X_test)

Compute the squared Mahalanobis distance of observations.

The squared Mahalanobis distance of an observation $r$ is defined as:

$$
d^2 = (r - \mu)^T \Sigma^{-1} (r - \mu)

$$

where $\Sigma$ is the estimated covariance matrix (`self.covariance_`)
and $\mu$ is the estimated mean (`self.location_` if available, otherwise
zero).

This distance measure accounts for correlations between assets and is useful
for:

* Outlier detection in portfolio returns
* Risk-adjusted distance calculations
* Identifying unusual market regimes

* **Parameters:**
  **X_test** *array-like of shape (n_observations, n_assets) or (n_assets,)*
  : Observations for which to compute the squared Mahalanobis distance.
    Each row represents one observation. If 1D, treated as a single
    observation. Assets with non-finite fitted variance are excluded from
    inference. After this asset-level filtering, each row is evaluated
    using the remaining available values only, covering row-level missing
    values such as market holidays or pre/post-listing. When rows have
    different observation patterns, the returned distances follow
    $\chi^2$ distributions with different degrees of freedom.
    Rows with no finite retained observation return NaN.
* **Returns:**
  **distances** *ndarray of shape (n_observations,) or float*
  : Squared Mahalanobis distance for each observation. Returns a scalar
    if input is 1D.

### Examples

```pycon
>>> import numpy as np
>>> from skfolio.moments import EmpiricalCovariance
>>> rng = np.random.default_rng(0)
>>> X = rng.standard_normal((100, 3))
>>> model = EmpiricalCovariance()
>>> model.fit(X)
EmpiricalCovariance()
>>> distances = model.mahalanobis(X)
>>> # The mean squared distance should be close to the number of assets (3).
>>> print(distances.mean())
2.9...
```

<a id="skfolio.moments.GraphicalLassoCV.score"></a>

#### score(X_test, y=None)

Compute the mean log-likelihood of observations under the estimated model.

Evaluates how well the fitted covariance matrix explains new observations,
assuming a multivariate Gaussian distribution. This is useful for:

* Model selection (comparing different covariance estimators)
* Cross-validation of covariance estimation methods
* Assessing goodness-of-fit

The log-likelihood for a single observation $r$ is:

$$
\log p(r | \mu, \Sigma) = -\frac{1}{2} \left[
    n \log(2\pi) + \log|\Sigma| + (r - \mu)^T \Sigma^{-1} (r - \mu)
\right]

$$

where $n$ is the number of assets, $\Sigma$ is the estimated
covariance matrix (`self.covariance_`), and $\mu$ is the estimated
mean (`self.location_` if available, otherwise zero).

* **Parameters:**
  **X_test** *array-like of shape (n_observations, n_assets)*
  : Observations for which to compute the log-likelihood.
    Typically held-out test data not used during fitting.
    Assets with non-finite fitted variance are excluded from inference. This
    typically happens when the fitted covariance cannot be estimated for an
    asset, for example before listing, after delisting, or during a warmup
    period. After this asset-level filtering, each row of `X_test` is scored
    using the remaining available values only. This covers row-level missing
    values in `X_test`, such as market holidays or pre/post-listing.

  **y** *Ignored*
  : Not used, present for scikit-learn API consistency.
* **Returns:**
  **score** *float*
  : Mean log-likelihood of the observations. Higher values indicate better fit.
    The score is averaged over all observations.

### Examples

```pycon
>>> import numpy as np
>>> from skfolio.moments import EmpiricalCovariance, LedoitWolf
>>> rng = np.random.default_rng(0)
>>> X_train = rng.standard_normal((100, 5))
>>> X_test = rng.standard_normal((50, 5))
>>> emp = EmpiricalCovariance().fit(X_train)
>>> lw = LedoitWolf().fit(X_train)
>>> # Compare models on held-out data
>>> print("Empirical:", emp.score(X_test))
Empirical: -6.97...
>>> print("LedoitWolf:", lw.score(X_test))
LedoitWolf: -6.88...
```

<a id="skfolio.moments.GraphicalLassoCV.set_params"></a>

#### set_params(\*\*params)

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects
(such as `Pipeline`). The latter have
parameters of the form `<component>__<parameter>` so that it’s
possible to update each component of a nested object.

* **Parameters:**
  **\*\*params** *dict*
  : Estimator parameters.
* **Returns:**
  **self** *estimator instance*
  : Estimator instance.

<a id="skfolio.moments.GraphicalLassoCV.set_score_request"></a>

#### set_score_request()

No-op.

Calling this method has no effect.

* **Returns:**
  **self** *object*
  : The updated object.

