All work
04Walk-Forward Quant Research2026

StockCraft

The honest backtest was the one that showed the edge was noise, and that result is the actual contribution.

PythonFlaskscikit-learnPyTorchSQLite

Context & problem

Most backtests leak future information into features without realizing it, producing an apparent edge that vanishes in live trading. The harder, more useful question is whether a signal survives a validation process that never lets the model see tomorrow.

Architecture

18,825 bars of raw price and volume history.

The interesting engineering: The honest backtest was the failed one

The model reached 54.34% directional accuracy. The baseline reached 54.29%. Calling that an edge would have made a better headline and a worse project.

Walk-forward validation forced every prediction to use only information available at that point in time. After eight folds and 7,740 out-of-sample predictions, the apparent gain was effectively noise. Probability ties revealed another uncomfortable truth: part of the ensemble was adding complexity without useful separation.

Result

54.34%
directional accuracy

vs. 54.29% baseline, statistically indistinguishable

7,740
out-of-sample predictions

8 walk-forward folds, 18,825 daily bars

0
lookahead bias

21 indicators, causal by construction

What I learned

  • The pipeline is the work: causal indicators, calibration, leakage controls, diagnostics, and reproducibility are more valuable than a backtest that quietly benefits from tomorrow’s data.
  • Reporting a negative result honestly is a stronger signal of engineering judgment than reporting a marginal, likely-noise positive one would have been.

Source & demo

No hosted demo. This is a research pipeline, not a product, so full methodology, folds, and diagnostics in the repo are the evidence.

View on GitHub

Next project

InferGate