Skip to content

feat(model): add StockMixer (AAAI 2024) to model zoo and benchmarks - #2363

Open
Sourish-07 wants to merge 3 commits into
microsoft:mainfrom
Sourish-07:feat/stockmixer-model
Open

Sourish-07 wants to merge 3 commits into
microsoft:mainfrom
Sourish-07:feat/stockmixer-model

Conversation

@Sourish-07

@Sourish-07 Sourish-07 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Description

Adds a Qlib-native implementation of StockMixer (AAAI 2024), an
MLP-based architecture for stock price forecasting, to the contrib model
zoo.

qlib/contrib/model/pytorch_stockmixer.py follows the same skeleton as
pytorch_tcn.py (Model interface, fit/predict, early stopping,
seed/GPU handling, count_parameters logging). The network implements
the paper's three mixers: indicator mixing, multi-scale time mixing
(upper-triangular TriU mixing of the raw series and a conv-down-sampled
view), and market-aware stock mixing (NoGraphMixer, no prior stock
graph required) - plus the paper's MSE + α·pairwise-ranking loss.

Key design decisions:

  • Uses DatasetH (not TSDatasetH), like TCN. Since the stock-mixing
    block needs the full daily cross-section, the wrapper batches
    day by day (grouped by the datetime index level); each day is fed
    as one (n_stock, time_steps, d_feat) tensor.
  • Each day's cross-section is zero-padded to a fixed n_stock. A boolean
    mask is threaded through the model so padded rows are excluded from the
    NoGraphMixer LayerNorm statistics and zeroed before/after its dense
    layers - the only layer in the network that mixes across stocks; every
    other layer operates per-stock, where padding is harmless by
    construction. Padded rows are also masked out of the loss and dropped
    from predictions. A day with more real instruments than n_stock
    raises a clear error.
  • Alpha360: d_feat=6, time_steps=60 (paper-faithful). Alpha158:
    d_feat=157, time_steps=1; the multi-scale time-mixing branch
    degrades to a single-step mapping in that case (documented in the
    model's docstring).

New files:

  • qlib/contrib/model/pytorch_stockmixer.py
  • examples/benchmarks/StockMixer/workflow_config_stockmixer_Alpha360.yaml
  • examples/benchmarks/StockMixer/workflow_config_stockmixer_Alpha158.yaml
  • examples/benchmarks/StockMixer/requirements.txt
  • tests/model/test_stockmixer.py

README / model-zoo documentation and the official 20-seed benchmark-table
numbers are intentionally left out of this PR and can be added in a
follow-up once more thorough runs and hyperparameter tuning are done.

Motivation and Context

StockMixer is a recent MLP-based architecture that avoids relying on a
pre-defined stock graph while still modeling cross-sectional relationships.
Adding it expands Qlib's Quant Model Zoo with an architecture distinct
from the existing RNN/GNN/Transformer baselines.

Paper: https://ojs.aaai.org/index.php/AAAI/article/view/28681
Official code: https://github.com/SJTU-DMTai/StockMixer

How Has This Been Tested?

  • Pass the test by running: pytest qlib/tests/test_all_pipeline.py under upper directory of qlib.
  • If you are adding a new feature, test on your own test scripts.

Automated tests (tests/model/test_stockmixer.py, 5 tests, all passing):
instantiation from both real workflow configs; forward/backward for both
time_steps=60 and time_steps=1; a correctness test proving padded
stock rows cannot influence real stocks' outputs through the masked
NoGraphMixer (torch.allclose under extreme injected noise in the
padding); the n_stock overflow error path; a full fit/predict cycle
on synthetic data.

Full regression: pytest tests/test_all_pipeline.py - 3 passed, no
failures (3 pre-existing unrelated warnings).

Real-data runs on CSI300 (via qrun, full unmodified configs,
n_epochs: 200, early_stop: 20, no crashes, no NaN/Inf anywhere):

  • Alpha158, 5 seeds (mean ± std - 5 seeds, not the standard 20):
    IC 0.0264 ± 0.0028, ICIR 0.207 ± 0.044, Rank IC 0.0326 ± 0.0041,
    Rank ICIR 0.246 ± 0.020, annualized return (with cost) 5.45% ± 1.17%,
    information ratio (with cost) 0.686 ± 0.164. All 5 runs early-stopped
    within epochs 20–22, best epoch between 0–2.
  • Alpha360, 1 seed: IC 0.0054, Rank IC 0.0229, early-stopped at
    epoch 21 (best @ epoch 1), with-cost annualized return −8.8%,
    information ratio −0.946.

Known limitation: both configs' validation score
peaks very early (epoch 0–2) and then degrades, triggering early stop
well before 200 epochs. Alpha158 ends up modestly positive; Alpha360's
backtest is currently negative after cost. This looks like a
hyperparameter-tuning issue (learning rate / rank-loss weight / possibly
batch composition) rather than a correctness bug - the architecture
itself is verified correct via the masking test above, and training
losses decrease smoothly and finitely throughout. I'd value input from
maintainers familiar with similar architectures on reasonable defaults,
and plan to tune further before/alongside the full 20-seed benchmark run.

Screenshots of Test Results (if appropriate):

  1. Pipeline test: tests/test_all_pipeline.py - 3 passed, 3 pre-existing warnings, exit 0.
  2. Your own tests: tests/model/test_stockmixer.py - 5 passed, exit 0.

Types of changes

  • Fix bugs
  • Add new feature
  • Update documentation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant