Search NASA⌕ Search

SEARCH · Search NASA

Results for “data modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Moving beyond post hoc explainable artificial intelligence: a perspective paper on lessons learned from dynamical climate modeling

AI models are criticized as being black boxes, potentially subjecting climate science to greater uncertainty. Explainable artificial intelligence (XAI) has been proposed to probe AI models and increase trust. In this review and perspective paper, we suggest that, in addition to using XAI methods, AI researchers in climate science can learn from past successes in the development of physics-based dynamical climate models. Dynamical models are complex but have gained trust because their successes and failures can sometimes be attributed to specific components or sub-models, such as when model bias is explained by pointing to a particular parameterization. We propose three types of understanding as a basis to evaluate trust in dynamical and AI models alike: (1) instrumental understanding, which is obtained when a model has passed a functional test; (2) statistical understanding, obtained when researchers can make sense of the modeling results using statistical techniques to identify input–output relationships; and (3) component-level understanding, which refers to modelers' ability to point to specific model components or parts in the model architecture as the culprit for erratic model behaviors or as the crucial reason why the model functions well. We demonstrate how component-level understanding has been sought and achieved via climate model intercomparison projects over the past several decades. Such component-level understanding routinely leads to model improvements and may also serve as a template for thinking about AI-driven climate science. Currently, XAI methods can help explain the behaviors of AI models by focusing on the mapping between input and output, thereby increasing the statistical understanding of AI models. Yet, to further increase our understanding of AI models, we will have to build AI models that have interpretable components amenable to component-level understanding. We give recent examples from the AI climate science literature to highlight some recent, albeit limited, successes in achieving component-level understanding and thereby explaining model behavior. The merit of such interpretable AI models is that they serve as a stronger basis for trust in climate modeling and, by extension, downstream uses of climate model data.

54 ENVIRONMENTAL SCIENCES↗

Horne et al. (2026) supporting files - WRF-LES model outputs for a summer heatwave event on June 2025 in Baltimore, MD

Brief Description Shown is the supporting information for Horne et al. (2026). These files include all outputs from the WRF model simulations and the observational datasets used for comparison in the study. Scripts are provided so users can recreate the manuscript's figures using the provided observational and modeling data. For more information regarding the study, please contact the primary author of the associated manuscript, Jason Horne. Horne, J. P., Pan, Y., Davis, K. J., Waugh, D., Ahlswede, B. J., Prince, N. E. (2026). Simulating near-surface environments in urban neighborhoods using WRF-LES: A case study of classic atmospheric boundary layer (ABL) during a heatwave event JAMES. (to be submitted)

atmosphere↗

Horne et al. (2026) supporting files - WRF-LES model outputs for a summer heatwave event on June 2025 in Baltimore, MD

Brief Description Shown is the supporting information for Horne et al. (2026). These files include all outputs from the WRF model simulations and the observational datasets used for comparison in the study. Scripts are provided so users can recreate the manuscript's figures using the provided observational and modeling data. For more information regarding the study, please contact the primary author of the associated manuscript, Jason Horne. Horne, J. P., Pan, Y., Davis, K. J., Waugh, D., Ahlswede, B. J., Prince, N. E. (2026). Simulating near-surface environments in urban neighborhoods using WRF-LES: A case study of classic atmospheric boundary layer (ABL) during a heatwave event JAMES.

atmosphere↗

Examination of simulated behavior of ND-LAr and TMS detectors using CAFAna for DUNE analysis framework

Simulations are run using the new DUNE CAFAna framework to generate pseudo-data modeling the interactions of neutrinos in the DUNE near detector at truth-level and detector-level. Truth-level analysis of neutrino kinematics reveals strong agreement with expected behavior, validating the kinematic portion of the simulation. Examination of the detector-level reconstructions of coordinates of interaction vertex appear consistent with an interaction density independent of detector position. Track lengths of particles resultant from neutrino interactions are aligned with varied particle identities, but are misaligned with prediction of uniform position density.

Fein, Jarrett [Fermilab]↗

Mapping Incidence and Prevalence Peak Data for SIR Modeling Applications

Infectious disease modeling and forecasting have played a key role in helping assess and respond to epidemics and pandemics. Recent work has leveraged data on disease peak infection and peak hospital incidence to fit compartmental models for the purpose of forecasting and describing the dynamics of a disease outbreak. Incorporating these data can greatly stabilize a compartmental model fit on early observations, where slight perturbations in the data may lead to model fits that forecast wildly unrealistic peak infection. We introduce a new method for incorporating historic data on the value and time of peak incidence of hospitalization into the fit for a Susceptible-Infectious-Recovered (SIR) model by formulating the relationship between an SIR model’s starting parameters and peak incidence as a system of two equations that can be solved computationally. We demonstrate how to calculate SIR parameter estimates – which describe disease dynamics such as transmission and recovery rates – using this method, and determine that there is a noticeable loss in accuracy whenever prevalence data is misspecified as incidence data. To exhibit the modeling potential, we update the Dirichlet-Beta State Space modeling framework to use hospital incidence data, as this framework was previously formulated to incorporate only data on total infections. This approach is assessed for practicality in terms of accuracy and speed of computation via simulation.

97 MATHEMATICS AND COMPUTING↗

Leveraging ARM Data to Improve Models for Predictive Understanding of Energy and Security Challenges

Extreme weather and natural hazards can disrupt the energy sector, affecting demand, generation, transmission, distribution, consumption and operational planning at regional and national scales. These disruptions stem from a broad range of atmospheric phenomena, including winter storms, freezing rain, wet snow loading, severe convection, flooding and landslides, wildfires, prolonged heat, and drought. Many of these same phenomena can also affect national security through impacts to transportation and infrastructure. To support the U.S. Department of Energy (DOE) focus on energy resilience and national security, the Atmospheric Radiation Measurement (ARM) User Facility is uniquely positioned to contribute measurement data, analyses, and modeling frameworks that can significantly improve predictive understanding of these hazards to mitigate their effects. To explore this opportunity, ARM convened a two-part virtual workshop in November 2025. The workshop engaged interdisciplinary experts in atmospheric science, energy systems, modeling, and operations. The goal of the meeting was to engage with these interdisciplinary experts to address three questions: • What are examples of atmospheric processes that represent significant risks to energy security or national security and where are those risks greatest? • What measurements or measurement strategies would improve ARM’s capacity to address these issues? • How can ARM and users of the ARM facility better work with the Energy Exascale Earth System Model (E3SM) and multi-sector modeling communities to apply ARM data to improving E3SM simulations of these phenomena? Participants were asked to submit white papers ahead of the meeting to initiate thinking on these themes and to help organize discussions. Workshop sessions were then organized around themes identified in the white papers. First from the white papers and then through subsequent discussions, workshop participants identified many examples that address the three questions listed above. Participants called out energy system vulnerabilities to weather phenomena such as the impact of freezing rain, strong winds, and excessive heat on power grids. They also noted the effects that weather phenomena could have on energy demand or supply (e.g., through effects of extreme temperatures). They called out security vulnerabilities such as impacts to crops from aerosol-borne pathogens and risks to industry due to melting permafrost in the Arctic. In all, over a dozen meteorological phenomena were linked to energy or security vulnerabilities. For many of the identified phenomena, participants pointed out where ARM was well poised to address issues (e.g., through measurements of cloud microphysics to inform studies of freezing rain) but also noted needs for additional measurements or modified measurement strategies. For example, adaptive scanning of severe weather would be valuable for probing winter storms or severe convection. Participants pointed out the value in integrating external observations with ARM measurements and with applying artificial intelligence (AI) to ARM observation analysis and they advocated for using model simulations to help optimize measurement strategies through Observing System Simulation Experiments (OSSEs). It was clear from the workshop that there are many ways that ARM observations can be used to mitigate energy and security concerns, but meeting participants were also asked to identify what they considered to be the greatest opportunities by ranking issues pertaining to the three workshop questions. This was accomplished through a survey administered to participants between the two virtual sessions. The highest-priority phenomena identified were winter storms, severe convection, and arctic processes. Discussion in the second session, therefore, focused primarily on these three areas, which were most fully developed in exploring ARM opportunities. Nevertheless, it was also clear that ARM has opportunities to contribute to all the identified topics. This report describes the workshop, including input from discussion and white papers (Sections 2 and 3) and a list of priority recommendations (section 4). Many other ideas for ARM contributions are discussed in individual white papers (Appendix D).

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Cluster-Graph Fingerprinting: A Framework for Quantitative Analysis of Machine-Learned Interatomic Model Training and Simulation Data

Machine-learned interatomic models represent a significant advancement in simulation methods, extending the predictive ability of first-principles methods to previously inaccessible length and time scales. However, the data-driven nature of these models can lead to difficult-to-detect errors that can compromise prediction accuracy. To address this challenge, we introduce a novel fingerprinting approach based on the Chebyshev Interaction Model for Efficient Simulation (ChIMES) ML-IAM graph-based descriptor. Our strategy enables efficient and statistically rigorous analysis of system configurations used in ML-IAM training and those generated by their application, e.g., in molecular dynamics simulations. We demonstrate that these fingerprints can effectively assess novelty of a configuration relative to an existing data set and determine dissimilarity among individual configurations, which are two key tasks in workflows for active learning-based ML-IAM training, data set curation, and on-the-fly uncertainty quantification.

36 MATERIALS SCIENCE↗

mphys-surrogate-model

This repository contains python scripts for building and studying reduced-order-modeling representations of droplet coalescence for eventual use in atmospheric models. The included data are generated from high-fidelity superdroplet methods and are utilized by machine learning pipelines to build data-driven models of droplet size distributions that evolve under coalescence. This repository further includes scripts to determine prediction (uncertainty) intervals on the data-driven model products based on conformal prediction.

Katona, JonasE [Lawrence Livermore National Labora↗

Nuclear Data Impact on Key Metrics for a Representative Molten Chloride Fast Reactor Model

Nuclear data are an essential component of the foundation on which all modeling and simulation methods and tools are relying upon, from the front end to the back end of the nuclear fuel cycle. In this study, the impact of uncertainties in nuclear data is investigated for a representative molten chloride fast reactor, for several important metrics, including eigenvalue, reactivity differences, and nuclide inventories in fuel at 5-yr irradiation. Uncertainty of keff for a full core model was found to be similar between the fresh fuel and the irradiated fuel states (1.7-1.8%), with its primary driver being the uncertainty in the 235U (n,γ) cross section. The results obtained for the reactivity differences show large uncertainties, of over 100%, in elastic scattering sensitivities of several nuclides, which led to large uncertainties of temperature reactivity differences for cladding and reflector. These results provide evidence that the currently applied methods may not be sufficiently adequate for ensuring the reliable determination of such metrics.

Procop, Germina [ORNL] (ORCID:0000000342226393)↗

Data-scarce surrogate modeling of shock-induced pore collapse process

Understanding the mechanisms of shock-induced pore collapse is of great interest in various disciplines in sciences and engineering, including materials science, biological sciences, and geophysics. However, numerical modeling of the complex pore collapse processes can be costly. To this end, a strong need exists to develop surrogate models for generating economic predictions of pore collapse processes. Here, in this work, we study the use of a data-driven reduced-order model, namely dynamic mode decomposition, and a deep generative model, namely conditional generative adversarial networks, to resemble the numerical simulations of the pore collapse process at representative training shock pressures. Since the simulations are expensive, the training data are scarce, which makes training an accurate surrogate model challenging. To overcome the difficulties posed by the complex physics phenomena, we make several crucial treatments to the plain original form of the methods to increase the capability of approximating and predicting the dynamics. In particular, physics information is used as indicators or conditional inputs to guide the prediction. In realizing these methods, the training of each dynamic mode composition model takes only around 30 s on CPU. In contrast, training a generative adversarial network model takes 8 h on GPU. Moreover, using dynamic mode decomposition, the final-time relative error is around 0.3% in the reproductive cases. We also demonstrate the predictive power of the methods at unseen testing shock pressures, where the error ranges from 1.3 to 5% in the interpolatory cases and 8 to 9% in extrapolatory cases.

97 MATHEMATICS AND COMPUTING↗

Modeling performance of data collection systems for high-energy physics

Exponential increases in scientific experimental data are outpacing silicon technology progress, necessitating heterogeneous computing systems—particularly those utilizing machine learning (ML)—to meet future scientific computing demands. The growing importance and complexity of heterogeneous computing systems require systematic modeling to understand and predict the effective roles for ML. We present a model that addresses this need by framing the key aspects of data collection pipelines and constraints and combining them with the important vectors of technology that shape alternatives, computing metrics that allow complex alternatives to be compared. For instance, a data collection pipeline may be characterized by parameters such as sensor sampling rates and the overall relevancy of retrieved samples. Alternatives to this pipeline are enabled by development vectors including ML, parallelization, advancing CMOS, and neuromorphic computing. By calculating metrics for each alternative such as overall F1 score, power, hardware cost, and energy expended per relevant sample, our model allows alternative data collection systems to be rigorously compared. We apply this model to the Compact Muon Solenoid experiment and its planned high luminosity-large hadron collider upgrade, evaluating novel technologies for the data acquisition system (DAQ), including ML-based filtering and parallelized software. The results demonstrate that improvements to early DAQ stages significantly reduce resources required later, with a power reduction of 60% and increased relevant data retrieval per unit power (from 0.065 to 0.31 samples/kJ). However, we predict that further advances will be required in order to meet overall power and cost constraints for the DAQ.

Olin-Ammentorp, Wilkie (ORCID:0000000224729862)↗

Exploring the Whole Set of Accurate Sparse Interpretable Models

In data science applications, there are often many models that fit the data well. This phenomenon was called the Rashomon Effect by Leo Breiman. The set of good models is called the Rashomon Set, and the goal of this project is to locate, store, and study the Rashomon sets for classes of interpretable models, including decision trees and generalized additive models.

97 MATHEMATICS AND COMPUTING↗

Data-driven background model for the CUORE experiment

Here, we present the model we developed to reconstruct the CUORE radioactive background based on the analysis of an experimental exposure of 1038.4 kg yr. The data reconstruction relies on a simultaneous Bayesian fit applied to energy spectra over a broad energy range. The high granularity of the CUORE detector, together with the large exposure and extended stable operations, allow for an in-depth exploration of both spatial and time dependence of backgrounds. We achieve high sensitivity to both bulk and surface activities of the materials of the setup, detecting levels as low as 10 nBq kg −1 and 0.1 nBq cm −2 , respectively. We compare the contamination levels we extract from the background model with prior radio-assay data, which informs future background risk mitigation strategies. The results of this background model play a crucial role in constructing the background budget for the CUPID experiment as it will exploit the same CUORE infrastructure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Augmenting a Simulation Campaign for Hybrid Computer Model and Field Data Experiments

The Kennedy and O’Hagan (KOH) calibration framework uses coupled Gaussian processes (GPs) to meta-model an expensive simulator (first GP), tune its “knobs” (calibration inputs) to best match observations from a real physical/field experiment and correct for any modeling bias (second GP) when predicting under new field conditions (design inputs). There are well-established methods for placement of design inputs for data-efficient planning of a simulation campaign in isolation, that is, without field data: space-filling, or via criterion like minimum integrated mean-squared prediction error (IMSPE). Analogues within the coupled GP KOH framework are mostly absent from the literature. Here, in this study, we derive a closed form IMSPE criterion for sequentially acquiring new simulator data for KOH. We illustrate how acquisitions space-fill in design space, but concentrate in calibration space. Closed form IMSPE precipitates a closed-form gradient for efficient numerical optimization. We demonstrate that our KOH-IMSPE strategy leads to a more efficient simulation campaign on benchmark problems, and conclude with a showcase on an application to equilibrium concentrations of rare earth elements for a liquid–liquid extraction reaction.

97 MATHEMATICS AND COMPUTING↗

Wind Plant Flow Physics and Power Performance in Complex Environments: Cooperative Research and Development (Final Report)

Cornell University will partner with NLR on the topic of wind farm wake effects to improve understanding of interactions between complex atmospheric flows, terrain, and wind turbine wakes and plant efficiency. Wind plant flow simulation tools will also be validated. The work performed will help improve wind farm modeling by analyzing data, applying models, designing and performing experiments to acquire additional wind farm data, and develop better models.

17 WIND ENERGY↗

A DATA EFFICIENT SPARSE MODELING FRAMEWORK FOR POWER ESTIMATION IN WATER TREATMENT SENSING OPERATIONS

With increasing freshwater scarcity, advanced process design mechanisms such as Closed-Circuit Reverse Osmosis (CCRO) and Digital/Physical Twin systems are gaining traction in water treatment and reuse operations. While digital and physical twin models enable improved system insight and control, their development is often expensive and computationally intensive, requiring large volumes of synthetic or experimental data to characterize underlying process dynamics. This work introduces a sparse surrogate modeling framework to estimate power consumption from measured flow and pressure variables, along with their nonlinear polynomial and interaction expansions. To ensure model reliability and reduce overfitting, a two-stage pipeline is proposed. First, a dynamic data filtering algorithm is employed to remove uninformative observations and transient operational states. Second, a sparse penalized regression technique is applied to select a minimal set of parsimonious features. The proposed model achieves high sparsity, retaining only 7 out of 34 candidate features (≈79.41% sparsity) while delivering a root mean square error (RMSE) of 0.072 on the test dataset.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Toward equitable environmental exposure modeling through convergence of data, open, and citizen sciences: an example of air pollution exposure modeling amidst increasing wildfire smoke

Exposure modeling is critical in environmental epidemiology and human health but may face challenges (e.g., skewed data, unequal error, context-insensitive validation, and computational demands). Modeling decisions reflect the intended use of the models and the values that modelers prioritize. We aimed to provide a conceptual framework and machine learning (ML) modeling protocols that address these issues. With 500m-gridded hourly PM 2.5 and O 3 levels in Illinois before, during, and after the 2023 Canadian wildfire season as a motivating example, we conducted modeling experiments to evaluate modeling methods, guided by three domains we propose based on theories of science: 1) Data Diversity, leveraging open and citizen science data to enhance inclusivity, parsimony, and representativeness; 2) Equitable Accuracy, ensuring fairly distributed uncertainties across subpopulations; and 3) Sustainable Modeling, balancing accuracy with reducing computational demands to promote accessibility for under-resourced researchers. Here, we found that ML with publicly available data can achieve high accuracy. Depending on methods, performance may vary substantially, even with identical input data. Large but skewed data may reduce performance. Misuse of cross-validation protocols can underestimate prediction error; although we observed R 2 s of ∼98 %, the modeled estimates varied significantly, indicating the need for careful model validation. By using new modeling protocols including representativeness-considered training and validation data and a new loss function, we achieved high agreement between estimates and ground-based measurements (e.g., R 2 = ∼90 % for PM 2.5 ; ∼80 % for O 3 ), equally distributed errors across sociodemographic strata and urban–rural divides, and reduction in computation time—from several weeks or months to a few days.

Exposure assessment↗