Skip to content
angelo-saporito24Public

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

qvol

v0.6.0 — run qvol.self_check() after installing; it verifies the build is current rather than letting a stale install surface as an ImportError mid-notebook.

Quantile methods for volatility and spread tail risk, with the backtest validity machinery written first.

Grew out of an MSc capstone applying walk-forward CAViaR models to ERCOT day-ahead/real-time electricity spreads. The estimation approach — direct two-sided conditional quantiles driven by exogenous state variables, matched-baseline attribution, thresholds frozen ex ante, joint statistical and economic evaluation — transfers to the volatility complex. The market does not: this repo tests that transfer rather than assuming it.


Colab quickstart

Open In Colab

Click the badge to open 00_setup.ipynb directly, or paste this into any notebook:

!pip install -q git+https://github.com/angelo-saporito24/qvol.git

from qvol.colab import setup, print_env
paths = setup(project="qvol", persist=True)   # mounts Drive
print_env()

Keep persist=True. Colab wipes its filesystem on disconnect. The trial registry is append-only across sessions — that is what it is for — so on ephemeral disk n_trials silently resets, the deflated Sharpe correction shrinks, and every result looks better than it is. Persistence is a correctness requirement here, not a convenience.

Check numba is present. The CAViaR recursion is sequential, so JIT is the difference between a 6-second test suite and a 60-second one, and between a 10-second notebook and one that will not finish. Colab normally ships with it; notebooks/00_setup.ipynb reports whether yours does.

Notebooks

Notebook Purpose
00_setup.ipynb Install, mount Drive, run tests, check CBOE endpoints, confirm persistence
01_validate_demo.ipynb The harness refusing a 30-candidate search with zero signal
02_vix_term_structure.ipynb Constant-maturity curve, spread target, exogenous state variables, VRP
03_caviar_vix.ipynb CAViaR fitted to the spread: matched baseline, DM comparison, calibration, crossing
10_st498_audit.ipynb Template: point the harness at the real capstone results

Install

git clone https://github.com/angelo-saporito24/qvol.git
cd qvol
make install          # pip install -e ".[dev]"
make test             # 99 tests

Dependencies are deliberately minimal — numpy, pandas, scipy — all preinstalled on Colab, so a fresh runtime installs in seconds with no resolution step.


Layout

src/qvol/
├── validate/         backtest validity: costs, purged CV, DSR, PBO, trials
├── caviar/           CAViaR estimation kernel, ported from ST498
├── data/             CBOE ingest, term-structure features, synthetic generator
└── colab.py          Drive persistence and environment reporting
notebooks/            generated by tools/make_notebooks.py
tests/                99 tests, all offline

qvol.validate

Written before the strategies it judges, and never modified to accommodate a result.

Module Fills
costs.py Transaction costs and slippage, charged on position changes. Plus breakeven_half_spread, the fastest viability check available.
cv.py Purged + embargoed k-fold and combinatorial purged CV.
metrics.py Probabilistic and deflated Sharpe, expected-maximum-under-null, minimum track record length.
pbo.py Probability of backtest overfitting via combinatorially symmetric CV, plus IS→OOS degradation slope.
trials.py Content-addressed, append-only trial registry. Makes n_trials mechanical rather than remembered.

qvol.data

Module Fills
cboe.py VIX-family index and VX futures settlement loaders, with disk caching and verify_endpoints().
features.py Constant-maturity construction, term-structure features, realised variance, VRP, and the point-in-time information guard.
synthetic.py Plausible VIX futures panels so everything is testable offline.

qvol.caviar

The ST498 estimation kernel, ported. Direct conditional-quantile modelling (Engle & Manganelli 2004) in three functional forms — SAV, Asymmetric Slope, Indirect GARCH — each accepting exogenous regressors.

Module Fills
models.py The three recursions. Numba optional: ~100x speedup where available, correct pure-Python fallback where not.
objective.py Exact (unsmoothed) pinball objective, tail-side semantics, upper-tail reflection.
starting.py Stationarity-derived starting values and training-window standardisation.
fit.py Ranked multi-start optimisation with admissibility rules; quantile sets and crossing diagnostics.
calibration.py Kupiec, Christoffersen, Dynamic Quantile, and Diebold-Mariano on pinball differentials.
from qvol.caviar import fit_caviar, calibration_report, standardise

Xs, mu, sd = standardise(X, n_train)
fit = fit_caviar(y, theta=0.05, X=Xs, model="AS", lag=1, n_train=n_train)
print(calibration_report(fit, y)["verdict"])

Two things that did NOT transfer from ERCOT. The lag-48 information constraint encoded a day-ahead bid deadline with no analogue in the volatility complex — the default here is 1, and carrying 48 across would model a deadline that does not exist. And the regressors: net load and thermal outages measured physical scarcity in a non-storable good. The method transfers; the model does not.

Expect Kupiec to pass while Christoffersen and DQ reject. Getting the long-run violation rate right is much easier than getting the timing right. The capstone rejected conditional coverage and DQ in all 816 combinations while Kupiec passed in two thirds — and said so. Report all three.


Gates

Gate Threshold Function
Cost breakeven half-spread > what you would actually pay breakeven_half_spread
Multiple testing DSR > 0.95 deflated_sharpe_ratio
Overfitting PBO < 0.5, ideally < 0.1 cscv_pbo
Selection stability one variant chosen in most combinations cscv_pbo → selected
Degradation IS→OOS slope > 0 performance_degradation

01_validate_demo.ipynb runs a 30-candidate search on data with no signal by construction:

annualised Sharpe        :  2.61     looks like a finding
PSR (ignores selection)  :  0.979    passes the naive test
DSR (corrects for it)    :  0.938    fails at 30 trials
PBO                      :  0.758    worse than random selection
IS→OOS slope             : -0.905    in-sample ranking is misleading
VERDICT: REJECT

That gap is the reason the layer exists.


Two design decisions worth knowing

Constant maturity, not raw M1−M2. The raw front-minus-second spread carries a sawtooth: as the front contract nears expiry it converges on spot and the spread narrows for purely calendar reasons. Modelling that means modelling an artefact. Same class of problem as back-adjusting a futures continuous contract — two analysts on identical raw data can reach different conclusions through this choice alone. front_two is retained for execution realism; constant_maturity_curve is the modelling target.

The information guard. apply_information_lag is the analogue of the lag-48 constraint in the ERCOT work, where the day-ahead bid deadline made same-hour realised data unavailable. The volatility analogue is milder — settlements publish after the close, so lag=1 for a daily book — but it is enforced in code rather than by intention, and lag=0 warns.


Caveats

  • CBOE endpoints are unverified. The loaders were written without live network access and CBOE has moved its data paths before. Run verify_endpoints() first. Everything downstream works from a DataFrame, so a moved URL costs one line, not a rewrite.
  • deflated_sharpe_ratio needs per-period returns, never annualised ones: the √(T−1) term assumes the Sharpe and the observation count are on the same frequency.
  • Kurtosis is raw (normal = 3), not excess. Passing excess kurtosis silently understates the deflation.
  • synthetic.py is a plausible generator, not a calibrated model. Use it to verify that code does what it claims, never to make claims about the market.

References

  • Engle & Manganelli (2004), JBES 22(4):367–381 — CAViaR.
  • Bailey & López de Prado (2014), JPM 40(5):94–107 — deflated Sharpe.
  • Bailey, Borwein, López de Prado & Zhu (2014), Notices of the AMS 61(5):458–471 — backtest overfitting, PBO.
  • López de Prado (2018), Advances in Financial Machine Learning, ch. 7, 12 — purged CV, CPCV.
  • Rockafellar & Uryasev (2000), Journal of Risk 2:21–41 — CVaR optimisation.

MIT licensed.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages