A systematic reproducibility framework for machine learning in additive manufacturing: comparative validation, cross-dataset generalization


Rizwan M., Shafi I., Çağlar A., Martinez Espinosa J. C., Uc Rios C. E., Ashraf I.

International Journal of Advanced Manufacturing Technology, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1007/s00170-026-18994-7
  • Dergi Adı: International Journal of Advanced Manufacturing Technology
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, IBZ Online, Compendex, INSPEC, DIALNET, Academic Search Ultimate (EBSCO), Engineering Source (EBSCO), Materials Science & Engineering Collection (ProQuest), Technology Collection (ProQuest)
  • Anahtar Kelimeler: 3D printing, Additive manufacturing, Machine learning, ManufacturingNet, Reproducibility, Uncertainty quantification
  • Akdeniz Üniversitesi Adresli: Evet

Özet

Machine learning (ML) is widely applied to quality prediction in additive manufacturing (AM), but few published models have been independently verified, and reproducibility assessments that do exist report whether a reproduction succeeded rather than why it failed. This paper proposes a five-pillar framework that quantifies the contribution of each undocumented methodological choice to the observed reproduction error, and that separates implementation paradigm from algorithm choice through a controlled factorial design. The framework is applied to two contrasting case studies: ManufacturingNet, an open-source ML toolbox, and a custom Random Forest (RF) model for surface roughness prediction in laser powder bed fusion (LPBF). Reproducing the published RF models gives a test of 0.375 and 0.396 against claimed values of 0.865 and 0.907. Decomposing that shortfall shows that the unreported random seed explains almost none of it and that the train–test partitioning rule explains a substantial share, while the undocumented descriptor formulation dominates, with physically plausible interpretations spanning from to 0.49. This ordering inverts the priority that current reproducibility checklists assign to seed reporting. The factorial experiment shows that the toolbox and a custom script produce numerically identical linear regression results on matched partitions, because the toolbox wraps the same estimator; differences between the two case studies are therefore attributable to algorithm, feature representation and evaluation protocol rather than to implementation paradigm. A feature-ablation test on the classification case shows that a reported accuracy near unity rests largely on near-disjoint nozzle temperature ranges between the two material classes. Bootstrap confidence intervals for all reproduced regression models overlap on the available test set, so none of the observed differences is statistically distinguishable. These are observations about two case studies and are not offered as field-level conclusions.