Skip to content
Research

Machine Learning

Questions about evaluation, assumptions and the limits of learned models.

Overview

Financial time series are non-stationary, serially dependent and have a low signal-to-noise ratio — three properties that break the independence and stationarity assumptions behind most standard machine-learning evaluation.

Core questions

  • How should training and validation splits respect the temporal and label-overlap structure of financial data?
  • When does an improvement in validation performance reflect a real pattern rather than leakage from overlapping labels?
  • How does a model's performance degrade as the underlying data-generating process drifts?

Mathematical formulation

Purged cross-validation constraint

If a label at time tᵢ is determined by information up to tᵢ + hᵢ (its label window), a valid split removes any training observation whose label window overlaps a test observation's label window. Training on an overlapping window leaks the test outcome into the training set.

Methods we use

  • Purged and embargoed cross-validation

    López de Prado, M. (2018). Advances in Financial Machine Learning. Wiley.

  • Deflated performance measures applied to model selection

    Bailey, D. H., & López de Prado, M. (2014). The deflated Sharpe ratio. Journal of Portfolio Management, 40(5), 94–107.

Open problems

  • What embargo length is sufficient when label windows are themselves of uncertain or variable duration?
  • How should feature importance be attributed when features are highly collinear and their relationships are regime-dependent?

This page describes the field's established methods, not DaraHoosh's own results, parameters or current use of them.