Training neural networks on mechanistic simulations improves scientific inference

Nature作者:Carson Dudley2026年8月12日正文已收录本站
  • Article
  • Open access
  • Published:

Scientific Reports (2026) Cite this article

We are providing an unedited version of this manuscript to give early access to its findings. Before final publication, the manuscript will undergo further editing. Please note there may be errors present which affect the content, and all legal disclaimers apply.

Abstract

Scientific modeling faces a tradeoff between the interpretability of mechanistic theory and the predictive power of machine learning. While existing hybrid approaches have made progress by incorporating domain knowledge into machine learning methods as functional constraints, they can be limited by a reliance on precise mathematical specifications. When the underlying equations are partially unknown or misspecified, enforcing rigid constraints can introduce bias and hinder a model’s ability to learn from data. We introduce Simulation-Grounded Neural Networks (SGNNs), a framework that incorporates scientific theory by using mechanistic simulations as training data for neural networks. By pretraining on diverse synthetic corpora that span multiple model structures and realistic observational noise, SGNNs internalize the underlying dynamics of a system as a structural prior. We evaluated SGNNs across multiple disciplines, including epidemiology, ecology, social science, and chemistry. In forecasting tasks, SGNNs outperformed both standard data-driven baselines and physics-constrained hybrid models. They nearly tripled the forecasting skill of the average CDC models in COVID-19 mortality forecasts and accurately forecasted high-dimensional ecological systems. SGNNs demonstrated robustness to model misspecification, performing well even when trained on data with incorrect assumptions. Our framework also introduces back-to-simulation attribution, a method for mechanistic interpretability that explains real-world dynamics by identifying their most similar counterparts within the simulated corpus. By unifying these techniques into a single framework, we demonstrate that diverse mechanistic simulations can serve as effective training data for robust scientific inference.

Subjects

Funding

This project was made possible by the Insight Net cooperative agreement with University of Michigan (5 NU38FT000002-02-00) from the CDC’s Center for Forecasting and Outbreak Analytics (CDC-RFA-FT-23-0069). Its contents are solely the responsibility of the authors and do not necessarily represent the official views of the Centers for Disease Control and Prevention.

Author information

Authors and Affiliations

  1. Department of Mathematics, University of Michigan, Ann Arbor, MI 48109, USA

    Carson Dudley & Marisa Eisenberg

  2. School of Public Health, University of Michigan, Ann Arbor, MI 48109, USA

    Carson Dudley, Reiden Magdaleno & Marisa Eisenberg

  3. Department of Physics, University of Michigan, Ann Arbor, MI 48109, USA

    Christopher Harding

  4. Center for the Study of Complex Systems, University of Michigan, Ann Arbor, MI 48109, USA

    Marisa Eisenberg

Authors

  1. Carson Dudley
  2. Reiden Magdaleno
  3. Christopher Harding
  4. Marisa Eisenberg

Corresponding author

Correspondence to Carson Dudley.

Ethics declarations

Competing interests

The authors declare no competing interests.

Additional information

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary Information

About this article

Check for updates. Verify currency and authenticity via CrossMark

Cite this article

Dudley, C., Magdaleno, R., Harding, C. et al. Training neural networks on mechanistic simulations improves scientific inference. Sci Rep (2026). https://doi.org/10.1038/s41598-026-64106-6

Download citation

  • Received:

  • Accepted:

  • Published:

  • DOI: https://doi.org/10.1038/s41598-026-64106-6