#### NOTE
[Go to the end](#sphx-glr-download-auto-examples-metadata-routing-plot-1-implied-volatility-py)
to download the full example code or to run this example in your browser via JupyterLite.

<a id="sphx-glr-auto-examples-metadata-routing-plot-1-implied-volatility-py"></a>

<a id="using-implied-volatility-with-metadata-routing"></a>

# Using Implied Volatility with Metadata Routing

This tutorial shows how to use [metadata routing](https://skfolio.org/user_guide/metadata_routing.html.md#metadata-routing).

We will use the [`ImpliedCovariance`](https://skfolio.org/generated/skfolio.moments.ImpliedCovariance.html.md#skfolio.moments.ImpliedCovariance) estimator inside
optimization models and grid search procedures to show how the implied volatility
time series can be routed.

<a id="load-datasets"></a>

## Load Datasets

We load the S&P 500 [dataset](https://skfolio.org/user_guide/datasets.html.md#datasets) composed of the daily prices of 20
assets from the S&P 500 Index composition and the implied volatility time series
of these 20 assets starting from 2010-01-04 up to 2022-12-28.

```Python
import numpy as np
import pandas as pd
import plotly.express as px
from plotly.io import show
from sklearn import set_config
from sklearn.model_selection import GridSearchCV, train_test_split

from skfolio import Population, RatioMeasure
from skfolio.datasets import load_sp500_dataset, load_sp500_implied_vol_dataset
from skfolio.metrics import make_scorer
from skfolio.model_selection import WalkForward, cross_val_predict
from skfolio.moments import (
    EmpiricalCovariance,
    GerberCovariance,
    ImpliedCovariance,
    LedoitWolf,
)
from skfolio.optimization import InverseVolatility, MeanRisk
from skfolio.preprocessing import prices_to_returns
from skfolio.prior import EmpiricalPrior

prices = load_sp500_dataset()
implied_vol = load_sp500_implied_vol_dataset()

X = prices_to_returns(prices)
X = X.loc["2010":]

implied_vol.head()
```

[plotly figure stripped from llms output]<style>html[data-theme="dark"] div.output_subarea:has(.plotly-graph-div){background:#fff;border-radius:0.25rem;padding:0.5rem}@media (prefers-color-scheme: dark){html:not([data-theme="light"]) div.output_subarea:has(.plotly-graph-div){background:#fff;border-radius:0.25rem;padding:0.5rem}}</style><script>if (!window.plotlySphinxGalleryResize) {window.plotlySphinxGalleryResize = true;window.addEventListener("load", function () {document.querySelectorAll(".plotly-graph-div").forEach(function (gd) { Plotly.Plots.resize(gd); });});}</script>

<br/>

Let’s print the average R2 per window size:

```Python
print({k: f"{np.mean(v):0.1%}" for k, v in coefs.items()})
```

```none
{10: '35.9%', 20: '36.8%', 60: '28.1%', 100: '31.4%'}
```

The highest R2 is achieved for a window size of 20 observations.

<a id="inverse-volatility"></a>

## Inverse Volatility

To use the `ImpliedCovariance` estimator inside a meta-estimator such as the
`InverseVolatility`, you must enable metadata routing with `set_config` and
specify where to route the implied vol using `set_fit_request` as shown below:

```Python
set_config(enable_metadata_routing=True)

model = InverseVolatility(
    prior_estimator=EmpiricalPrior(
        covariance_estimator=ImpliedCovariance().set_fit_request(implied_vol=True)
    )
)
```

Then you can pass the implied volatility to the `fit` method of the meta-estimator:

```Python
model.fit(X, implied_vol=implied_vol)
print(model.weights_)
```

```none
[0.03629576 0.02554287 0.04135536 0.02979163 0.04207543 0.03703667
 0.04592872 0.08520002 0.04462864 0.0739193  0.04745674 0.06701373
 0.03752756 0.07489027 0.05358558 0.07275256 0.02274707 0.05618856
 0.06368034 0.04238321]
```

<a id="cross-validation"></a>

## Cross Validation

In this section, we show how to use metadata routing with `cross_val_predict`.
First, we create a `WalkForward` cross-validator to rebalance our portfolio every 20
business days by re-fitting the model on the previous 400 business days (~ 1.5 years):

```Python
cv = WalkForward(train_size=400, test_size=20)
```

We use the model created above and pass the implied volatility in `params`:

```Python
pred_model = cross_val_predict(model, X, cv=cv, params={"implied_vol": implied_vol})
pred_model.name = "Implied Vol"
```

Let’s compare the model with a benchmark using `InverseVolatility` with the default
`EmpiricalCovariance` estimator:

```Python
benchmark = InverseVolatility()
pred_bench = cross_val_predict(benchmark, X, cv=cv)
pred_bench.name = "Benchmark"
```

For easier analysis, we add both predicted portfolios into a `Population`:

```Python
population = Population([pred_bench, pred_model])
summary = population.summary()
print(summary.loc[["Annualized Standard Deviation", "Annualized Sharpe Ratio"]])
```

```none
                              Benchmark Implied Vol
Annualized Standard Deviation    16.24%      16.09%
Annualized Sharpe Ratio            1.00        1.04
```

Let’s plot the Composition and Cumulative returns:

```Python
population.plot_composition(display_sub_ptf_name=False)
```

<style>html[data-theme="dark"] div.output_subarea:has(.plotly-graph-div){background:#fff;border-radius:0.25rem;padding:0.5rem}@media (prefers-color-scheme: dark){html:not([data-theme="light"]) div.output_subarea:has(.plotly-graph-div){background:#fff;border-radius:0.25rem;padding:0.5rem}}</style><script>if (!window.plotlySphinxGalleryResize) {window.plotlySphinxGalleryResize = true;window.addEventListener("load", function () {document.querySelectorAll(".plotly-graph-div").forEach(function (gd) { Plotly.Plots.resize(gd); });});}</script>[plotly figure stripped from llms output]
<br />
<br />

```Python
population.plot_cumulative_returns()
```

<style>html[data-theme="dark"] div.output_subarea:has(.plotly-graph-div){background:#fff;border-radius:0.25rem;padding:0.5rem}@media (prefers-color-scheme: dark){html:not([data-theme="light"]) div.output_subarea:has(.plotly-graph-div){background:#fff;border-radius:0.25rem;padding:0.5rem}}</style><script>if (!window.plotlySphinxGalleryResize) {window.plotlySphinxGalleryResize = true;window.addEventListener("load", function () {document.querySelectorAll(".plotly-graph-div").forEach(function (gd) { Plotly.Plots.resize(gd); });});}</script>[plotly figure stripped from llms output]
<br />
<br />

<a id="hyper-parameters-tuning"></a>

## Hyper-Parameters Tuning

In this section, we show how to use metadata routing with `GridSearchCV`.
First, we split the data into a train and a test set:

```Python
X_train, X_test, implied_vol_train, implied_vol_test = train_test_split(
    X, implied_vol, test_size=1 / 2, shuffle=False
)
```

We create a Minimum Variance that uses the `ImpliedCovariance` estimator:

```Python
model = MeanRisk(
    prior_estimator=EmpiricalPrior(
        covariance_estimator=ImpliedCovariance().set_fit_request(implied_vol=True)
    )
)
```

Then, we find the hyper-parameters of the `ImpliedCovariance` estimator that
maximizes the out-of-sample Sharpe Ratio of the Minimum Variance model:

```Python
grid_search = GridSearchCV(
    estimator=model,
    param_grid={
        "prior_estimator__covariance_estimator__window_size": np.arange(5, 50, 3),
        "prior_estimator__covariance_estimator__prior_covariance_estimator": [
            LedoitWolf(),
            GerberCovariance(),
            EmpiricalCovariance(),
        ],
    },
    return_train_score=True,
    scoring=make_scorer(RatioMeasure.ANNUALIZED_SHARPE_RATIO),
    n_jobs=-1,
    cv=cv,
)
grid_search.fit(X_train, implied_vol=implied_vol_train)
gs_model = grid_search.best_estimator_
print(gs_model)
```

```none
MeanRisk(prior_estimator=EmpiricalPrior(covariance_estimator=ImpliedCovariance(annualization_factor=252.0,
                                                                               prior_covariance_estimator=GerberCovariance(),
                                                                               window_size=np.int64(17))))
```

Let’s plot the out-of-sample Sharpe Ratio as a function of the window size and
the prior covariance estimator used to compute the correlation matrix:

```Python
cv_results = grid_search.cv_results_

df = pd.DataFrame(
    {
        "Prior Cov Estimator": [
            str(x)
            for x in cv_results[
                "param_prior_estimator__covariance_estimator__prior_covariance_estimator"
            ]
        ],
        "Window Size": cv_results[
            "param_prior_estimator__covariance_estimator__window_size"
        ],
        "Test Sharpe Ratio": cv_results["mean_test_score"],
        "error": cv_results["std_test_score"] / 10,  # one tenth of std for readability
    }
)
px.line(
    df,
    x="Window Size",
    y="Test Sharpe Ratio",
    color="Prior Cov Estimator",
    error_y="error",
    title="Out-of-Sample Sharpe Ratio",
)
```

<style>html[data-theme="dark"] div.output_subarea:has(.plotly-graph-div){background:#fff;border-radius:0.25rem;padding:0.5rem}@media (prefers-color-scheme: dark){html:not([data-theme="light"]) div.output_subarea:has(.plotly-graph-div){background:#fff;border-radius:0.25rem;padding:0.5rem}}</style><script>if (!window.plotlySphinxGalleryResize) {window.plotlySphinxGalleryResize = true;window.addEventListener("load", function () {document.querySelectorAll(".plotly-graph-div").forEach(function (gd) { Plotly.Plots.resize(gd); });});}</script>[plotly figure stripped from llms output]
<br />
<br />

Finally, we compare the optimal Grid Search model with a naive Minimum Variance
benchmark on the **test set**:

```Python
pred_gs_model = cross_val_predict(
    gs_model, X_test, params={"implied_vol": implied_vol_test}, cv=cv, n_jobs=-1
)
pred_gs_model.name = "GS Model"

benchmark = MeanRisk()
pred_bench = cross_val_predict(benchmark, X_test, cv=cv)
pred_bench.name = "Benchmark"

population = Population([pred_bench, pred_gs_model])
summary = population.summary()
print(summary.loc[["Annualized Standard Deviation", "Annualized Sharpe Ratio"]])
```

```none
                              Benchmark GS Model
Annualized Standard Deviation    18.12%   18.28%
Annualized Sharpe Ratio            0.59     0.67
```

```Python
population.plot_cumulative_returns()
```

<style>html[data-theme="dark"] div.output_subarea:has(.plotly-graph-div){background:#fff;border-radius:0.25rem;padding:0.5rem}@media (prefers-color-scheme: dark){html:not([data-theme="light"]) div.output_subarea:has(.plotly-graph-div){background:#fff;border-radius:0.25rem;padding:0.5rem}}</style><script>if (!window.plotlySphinxGalleryResize) {window.plotlySphinxGalleryResize = true;window.addEventListener("load", function () {document.querySelectorAll(".plotly-graph-div").forEach(function (gd) { Plotly.Plots.resize(gd); });});}</script>[plotly figure stripped from llms output]
<br />
<br />

<a id="conclusion"></a>

## Conclusion

This was a toy example to introduce the metadata routing API.
For more information, see [Metadata Routing User Guide](https://skfolio.org/user_guide/metadata_routing.html.md#metadata-routing).

**Total running time of the script:** (1 minutes 7.118 seconds)

<a id="sphx-glr-download-auto-examples-metadata-routing-plot-1-implied-volatility-py"></a>
[![Launch JupyterLite](auto_examples/metadata_routing/images/jupyterlite_badge_logo.svg)](../../lite/lab/index.html?path=auto_examples/metadata_routing/plot_1_implied_volatility.ipynb)

[`Download Jupyter notebook: plot_1_implied_volatility.ipynb`](https://skfolio.org/auto_examples/metadata_routing/_downloads/e48165c5b00d9d5f038444a4e041dd91/plot_1_implied_volatility.ipynb)

[`Download Python source code: plot_1_implied_volatility.py`](https://skfolio.org/auto_examples/metadata_routing/_downloads/f621598e19b9fa853c8ed59240063635/plot_1_implied_volatility.py)

[`Download zipped: plot_1_implied_volatility.zip`](https://skfolio.org/auto_examples/metadata_routing/_downloads/2cacef42211ad9bfc45fb395723eff2d/plot_1_implied_volatility.zip)

[Gallery generated by Sphinx-Gallery](https://sphinx-gallery.github.io)
