Temporal Validation and Incremental News-Text Analysis of a Lightweight Cross-Sectional Equity Ranking Model

Model Construction, Matched Comparisons, and Cost Sensitivity for M8 N0

Authors

  • Jiexiao Zhang discoveryopen.org Author

Keywords:

cross-sectional equity ranking, event metadata, lightweight neural network, temporal validation, pairwise ranking, reproducibility

Abstract

Cross-sectional equity prediction requires explicit information-timing rules, consistent treatment of differences in samples across trading dates, and careful separation of model selection from final evaluation. This study examines whether structured market features and news-event metadata support same-day open-to-close return ranking when inputs are frozen before the market opens. It also assesses the incremental contribution of news text in a matched comparison. M8 N0 combines 23 numerical features, their missingness indicators, and two log-transformed event counts into a 48-dimensional input. A multilayer perceptron with 5,345 trainable parameters produces cross-sectional ranking scores. Training uses within-date rank-normalized targets, a joint SmoothL1 and pairwise ranking objective, and optimization batches organized around complete trading dates. Three expanding-window folds use preprocessing statistics fitted only on their respective training samples, temporal gaps, and controls for duplicate events across partitions.

Across three development-validation folds with seed 42, the model covers 298,769 symbol-day observations and 1,350 trading dates. Mean daily Spearman IC is 0.02088, with positive IC on 55.48% of dates. Fold-level mean IC values are 0.01633, 0.02948, and 0.01668. In the matched comparison on fold 0, the numerical baseline, additive text-residual model, and text branch with relation-weighted pooling achieve IC values of 0.01633, 0.01587, and 0.00396, respectively. The zero-text control matches the numerical baseline. These findings document an observable historical ranking relationship in the structured inputs within the completed configurations and sample scope, while the tested text addition does not improve mean IC on this fold. Post-hoc moving-block bootstrap analyses characterize uncertainty in the saved daily sequences under temporal dependence, and yearly summaries describe temporal variation. An exploratory portfolio analysis demonstrates that ranking metrics and returns after costs require separate interpretation: an equal-weight long-short portfolio selecting 20% of each tail has approximately -3.11% arithmetic annualized return under a simplified cost of 2 basis points per one-way execution. Reload and release-consistency checks support reproducible inference. Because validation data were used for checkpoint selection, the results are developmental rather than an untouched out-of-sample test. The study provides an auditable research workflow connecting information timing, a numerical baseline, incremental-text controls, and released artifacts, establishing an explicit starting point for independent temporal evaluation and multi-seed research.

References

[1] Zihan Dong, Xinyu Fan, Zhiyuan Peng. FNSPID: A Comprehensive Financial News Dataset in Time Series. arXiv preprint arXiv:2402.06698, 2024. DOI 10.48550/arXiv.2402.06698. https://arxiv.org/abs/2402.06698

[2] Shihao Gu, Bryan Kelly, Dacheng Xiu. Empirical Asset Pricing via Machine Learning. The Review of Financial Studies, 33(5): 2223–2273, 2020. DOI 10.1093/rfs/hhaa009. https://doi.org/10.1093/rfs/hhaa009

[3] Dogu Araci. FinBERT: Financial Sentiment Analysis with Pre-trained Language Models. arXiv preprint arXiv:1908.10063; University of Amsterdam master's thesis, 2019. DOI 10.48550/arXiv.1908.10063. https://arxiv.org/abs/1908.10063

[4] Yi Yang, Mark Christopher Siy Uy, Allen Huang. FinBERT: A Pretrained Language Model for Financial Communications. arXiv preprint arXiv:2006.08097, 2020. DOI 10.48550/arXiv.2006.08097. https://arxiv.org/abs/2006.08097

[5] Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, Greg Hullender. Learning to Rank Using Gradient Descent. Proceedings of the 22nd International Conference on Machine Learning (ICML): 89–96, 2005. DOI 10.1145/1102351.1102363. https://doi.org/10.1145/1102351.1102363

[6] Peter J. Huber. Robust Estimation of a Location Parameter. The Annals of Mathematical Statistics, 35(1): 73–101, 1964. DOI 10.1214/aoms/1177703732. https://doi.org/10.1214/aoms/1177703732

[7] Ilya Loshchilov, Frank Hutter. Decoupled Weight Decay Regularization. International Conference on Learning Representations (ICLR), 2019. DOI 10.48550/arXiv.1711.05101. https://arxiv.org/abs/1711.05101

[8] David H. Bailey, Jonathan M. Borwein, Marcos López de Prado, Qiji Jim Zhu. The Probability of Backtest Overfitting. Journal of Computational Finance, 20(4): 39–69, 2017. DOI 10.21314/JCF.2016.322. https://doi.org/10.21314/JCF.2016.322

[9] Campbell R. Harvey, Yan Liu, Heqing Zhu. … and the Cross-Section of Expected Returns. The Review of Financial Studies, 29(1): 5–68, 2016. DOI 10.1093/rfs/hhv059. https://doi.org/10.1093/rfs/hhv059

[10] Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, Tie-Yan Liu. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. Advances in Neural Information Processing Systems, 30, 2017. https://papers.nips.cc/paper/2017/hash/6449f44a102fde848669bdd9eb6b76fa-Abstract.html

[11] Yury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem Babenko. Revisiting Deep Learning Models for Tabular Data. Advances in Neural Information Processing Systems, 34, 2021. https://proceedings.neurips.cc/paper/2021/hash/9d86d83f925f2149e9edb0ac3b49229c-Abstract.html

[12] Léo Grinsztajn, Edouard Oyallon, Gaël Varoquaux. Why Do Tree-Based Models Still Outperform Deep Learning on Typical Tabular Data?. Advances in Neural Information Processing Systems, Datasets and Benchmarks Track, 35, 2022. DOI 10.52202/068431-0037. https://papers.neurips.cc/paper_files/paper/2022/hash/0378c7692da36807bdec87ab043cdadc-Abstract-Datasets_and_Benchmarks.html

[13] Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of NAACL-HLT 2019, Volume 1 (Long and Short Papers): 4171–4186, 2019. DOI 10.18653/v1/N19-1423. https://aclanthology.org/N19-1423/

[14] Halbert White. A Reality Check for Data Snooping. Econometrica, 68(5): 1097–1126, 2000. DOI 10.1111/1468-0262.00152. https://onlinelibrary.wiley.com/doi/10.1111/1468-0262.00152

Downloads

Published

2026-09-26

Issue

Section

Research Articles