Cookie & Analytics Notice
With your consent, Q72 records public-page interactions and technical error signals for analytics. Interaction replays are kept for 48 hours; analytics summaries remain. Known Q72 links or submitted contact details may associate your visit with an existing contact. No keystrokes, form values or IP addresses are recorded in replays.
TL;DR
In-sample Sharpe ratios systematically overstate future performance because the optimizer finds the allocation that looks best on training data by construction. Q72 uses a rolling walk-forward framework: data is held out entirely, the optimizer runs on the training window, and the allocation is evaluated on unseen data. All summary statistics come from this out-of-sample distribution.
In-sample Sharpe ratios look good in every backtest. We explain why Q72 only shows out-of-sample validation, how it is implemented, and what it means for the reliability of the recommendations you act on.
There is a number that appears in almost every portfolio optimization output, and it is almost always misleading: the in-sample Sharpe ratio. It is the ratio of return to volatility calculated on the same historical period that was used to build the portfolio. It is the number that optimization software puts in the summary box, the number that gets copied into client presentations, and the number that, in the vast majority of cases, significantly overstates the strategy's likely future performance.
The reason is straightforward. An optimizer that is given a set of historical return data and asked to find the best allocation will find the allocation that looks best on that data. This is what optimization means. The resulting portfolio has, by construction, a high Sharpe ratio over the period it was trained on. What it does not have — and what the in-sample number tells you nothing about — is a demonstrated ability to perform on data the optimizer never saw.
Portfolio overfitting is not a subtle or academic concern. It is the reason that strategies which look compelling in backtests so frequently disappoint in live trading. The optimizer has found patterns in the historical data — some structural, some spurious — and it has no mechanism to distinguish between them. The more freedom the optimizer has, the more aggressively it will exploit spurious patterns, and the larger the gap between in-sample and out-of-sample performance.
This gap has a name in the quantitative finance literature: the backtest overfitting premium. It describes the systematic bias toward inflated in-sample Sharpe ratios that results from optimization against historical data. Research consistently shows that the gap between in-sample and out-of-sample Sharpe ratios is larger for more complex optimization methods, larger for shorter historical windows, and larger when the optimizer has more degrees of freedom.
Q72 uses a rolling window validation framework. For each optimization run, a portion of the available historical data is held out entirely — it is not used in the optimization, not used in the Confidence Alpha scoring, and not used in any parameter calibration. The optimizer runs on the training window. The resulting allocation is then evaluated on the held-out window as if it had been deployed at the beginning of that period.
This process is repeated across multiple rolling windows, producing a distribution of out-of-sample results rather than a single number. The summary statistics shown in every Q72 output — Sharpe ratio, maximum drawdown, annualized volatility — are calculated from this distribution, not from the training period.
All four classical methodologies are held to the same walk-forward out-of-sample reporting framework. Within Q72 Confidence Alpha, the acceptance rule comparing a quantum candidate with Q72 Classic is separate from that OOS evidence: a quantum candidate is retained only when it improves the same Q72 objective without increasing volatility. Any historical figure shown for the final quantum-refined weights is a replay of those final live weights, not a walk-forward OOS track record.
Out-of-sample framing also changes the quality of an investment discussion. A more conservative metric drawn from data the model did not use for construction is not automatically a weaker result; it is a different and more defensible type of evidence. The figures remain historical tests, not forecasts or guarantees.
Continue with Q72
Start Free Trial