A Chemically Interpretable and Leakage‐Safe Machine Learning Framework for Predicting Molecular Refractive Index (nD)
CHEMISTRYSELECT, cilt.11, sa.29, ss.20-40, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 11 Sayı: 29
- Basım Tarihi: 2026
- Doi Numarası: 10.1002/slct.73822
- Dergi Adı: CHEMISTRYSELECT
- Derginin Tarandığı İndeksler: Academic Search Ultimate (EBSCO), Scopus, Science Citation Index Expanded (SCI-EXPANDED), Chemical Abstracts Core
- Sayfa Sayıları: ss.20-40
- Akdeniz Üniversitesi Adresli: Evet
Özet
The molecular refractive index (
) governs light–matter interactions and underpins optical biosensors, bioimaging agents, lab-on-a-chip devices, and functional biomaterials. Existing machine-learning models for
emphasize predictive accuracy but provide limited chemical interpretability and safeguards against information leakage. Here, we present an interpretable, leakage-safe framework for predicting
from structure-derived descriptors and fingerprints. A curated dataset is processed via multi-toolkit feature generation, removal of low-information channels, diversity-aware filtering, and ANOVA F-test feature selection within a pipeline that performs median imputation and standardization. Sparse and generalized linear models enable attribution of
variation to chemically meaningful descriptors. The final Lasso model shows strong generalization, with a mean fold-specific test performance of
(MAE
), the highest fold-specific test performance of
, and a definitive held-out unseen test-set performance of
(MAE
) after retraining on the complete training set. Ablation studies identify feature selection, pipeline-based preprocessing, and targeted tuning as the primary contributors to predictive performance in the selected framework.