Search NASA⌕ Search

SEARCH · Search NASA

Results for “Models, Statistical”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Moving beyond post hoc explainable artificial intelligence: a perspective paper on lessons learned from dynamical climate modeling

AI models are criticized as being black boxes, potentially subjecting climate science to greater uncertainty. Explainable artificial intelligence (XAI) has been proposed to probe AI models and increase trust. In this review and perspective paper, we suggest that, in addition to using XAI methods, AI researchers in climate science can learn from past successes in the development of physics-based dynamical climate models. Dynamical models are complex but have gained trust because their successes and failures can sometimes be attributed to specific components or sub-models, such as when model bias is explained by pointing to a particular parameterization. We propose three types of understanding as a basis to evaluate trust in dynamical and AI models alike: (1) instrumental understanding, which is obtained when a model has passed a functional test; (2) statistical understanding, obtained when researchers can make sense of the modeling results using statistical techniques to identify input–output relationships; and (3) component-level understanding, which refers to modelers' ability to point to specific model components or parts in the model architecture as the culprit for erratic model behaviors or as the crucial reason why the model functions well. We demonstrate how component-level understanding has been sought and achieved via climate model intercomparison projects over the past several decades. Such component-level understanding routinely leads to model improvements and may also serve as a template for thinking about AI-driven climate science. Currently, XAI methods can help explain the behaviors of AI models by focusing on the mapping between input and output, thereby increasing the statistical understanding of AI models. Yet, to further increase our understanding of AI models, we will have to build AI models that have interpretable components amenable to component-level understanding. We give recent examples from the AI climate science literature to highlight some recent, albeit limited, successes in achieving component-level understanding and thereby explaining model behavior. The merit of such interpretable AI models is that they serve as a stronger basis for trust in climate modeling and, by extension, downstream uses of climate model data.

54 ENVIRONMENTAL SCIENCES↗

DRDMannTurb: A Python package for scalable, data-driven synthetic turbulence

Synthetic turbulence models (STMs) are used in wind engineering to generate realistic flow fields and are employed as inputs to industrial wind simulations. Examples include prescribing inlet conditions in large eddy simulations that model loads on wind turbines and tall buildings. We are interested in STMs capable of generating fluctuations based on prescribed second-moment statistics since such models can simulate environmental conditions that closely resemble on-site observations. To this end, the widely used Mann model (see Mann, 1994, 1998) is the inspiration for DRDMannTurb. The Mann model is described by three physical parameters: a magnitude parameter influencing the global variance of the wind field and corresponding to the Kolmogorov constant multiplied by the rate of viscous dissipation of the turbulent kinetic energy to the two-thirds, αϵ 2/3 , a turbulence length scale parameter L, and a nondimensional parameter Γ related to the lifetime of the eddies. A number of studies, as well as international standards (e.g., those by the International Electrotechnical Commission (IEC)), include recommended values for these three parameters with the goal of standardizing wind simulations according to observed energy spectra. Yet, having only three parameters, the Mann model faces limitations in accurately representing the diversity of observable spectra. This Python package enables users to extend the Mann model and more accurately fit field measurements through flexible neural network models of the eddy lifetime function. Following Keith et al. (2021), we refer to this class of models as Deep Rapid Distortion (DRD) models. DRDMannTurb also includes a general module implementing an efficient method for synthetic turbulence generation based on a domain decomposition technique. This technique is also described in Keith et al. (2021).

17 WIND ENERGY↗

Testing the thermal Sunyaev-Zel’dovich power spectrum of a halo model using hydrodynamical simulations

Statistical properties of large-scale cosmological structures serve as powerful tools for constraining the cosmological properties of our Universe. Tracing the gas pressure, the thermal Sunyaev-Zel’dovich (tSZ) effect is a biased probe of mass distribution and, hence, can be used to test the physics of feedback or cosmological models. Therefore, it is crucial to develop robust modelling of hot gas pressure for applications to tSZ surveys. Since gas collapses into bound structures, it is expected that most of the tSZ signal is within halos produced by cosmic accretion shocks. Hence, simple empirical halo models can be used to predict the tSZ power spectra. In this study, we employed the HMx halo model to compare the tSZ power spectra with those of several hydrodynamical simulations: the Horizon suite and the Magneticum simulation. We examine various contributions to the tSZ power spectrum across different redshifts, including the one- and two-halo term decomposition, the amount of bound gas, the importance of different masses, and the electron pressure profiles. Our comparison of the tSZ power spectrum reveals discrepancies between the halo model and cosmological simulations that increase with redshift. We find a 20% to 50% difference between the measured and predicted tSZ angular power spectrum over the multipole range ℓ = 10 3 − 10 4 . Our analysis reveals that these differences are driven by the excess of power in the predicted two-halo term at low k and in the one-halo term at high k . At higher redshifts ( z ∼ 3), simulations indicate that more power comes from outside the virial radius than from inside, suggesting a limitation in the applicability of the halo model. We also observe differences in the pressure profiles, despite the fair level of agreement on the tSZ power spectrum at low redshift with the default calibration of the halo model. In conclusion, our study suggests that the properties of the halo model need to be carefully controlled against real or mock data to be proven useful for cosmological purposes.

Ayçoberry, Emma (ORCID:0000000292351195)↗

Low responsiveness of machine learning models to critical or deteriorating health conditions

Machine learning (ML) based mortality prediction models can be immensely useful in intensive care units. Such a model should generate warnings to alert physicians when a patient’s condition rapidly deteriorates, or their vitals are in highly abnormal ranges. Before clinical deployment, it is important to comprehensively assess a model’s ability to recognize critical patient conditions. We develop multiple medical ML testing approaches, including a gradient ascent method and neural activation map. We systematically assess these machine learning models’ ability to respond to serious medical conditions using additional test cases, some of which are time series. Guided by medical doctors, our evaluation involves multiple machine learning models, resampling techniques, and four datasets for two clinical prediction tasks. We identify serious deficiencies in the models’ responsiveness, with the models being unable to recognize severely impaired medical conditions or rapidly deteriorating health. For in-hospital mortality prediction, the models tested using our synthesized cases fail to recognize 66% of the injuries. In some instances, the models fail to generate adequate mortality risk scores for all test cases. Our study identifies similar kinds of deficiencies in the responsiveness of 5-year breast and lung cancer prediction models. Using generated test cases, we find that statistical machine-learning models trained solely from patient data are grossly insufficient and have many dangerous blind spots. Most of the ML models tested fail to respond adequately to critically ill patients. How to incorporate medical knowledge into clinical machine learning models is an important future research direction.

60 APPLIED LIFE SCIENCES↗

MPACT Safeguards Modeling: FY25 Update

Sandia National Laboratories develops and maintains several open-source software packages to support material accountancy analyses. This includes the Material Accountancy Performance Indicator Toolkit (MAPIT), the Fissile Facility Flow Modeler (F3M) and the Separation and Safeguards Performance Model Library (SSPM-L). MAPIT is responsible for performing statistical safeguards analyses on bulk and itemized data from nuclear fuel cycle facilities and can operate on real or synthetic data. MAPIT is the only open-source software for such analyses. F3M is a library of modules, built in MATLAB Simulink, that contain pre made blocks to represent different generic fuel cycle processes. These blocks can be used together in a modular fashion to represent and simulate nuclear fuel cycle processes with the goal of improving facility-level accountancy during the design phase. F3M is also an open-source library. Finally, the SSPM-L library is a series of completed models built from F3M. The library includes facility models such as a generic PUREX facility and a fuel fabrication facility. The SSPM-L library is not open source, but is available to collaborators with a relevant use case. These tools include modeling and simulation pipelines to simulate nuclear fuel cycle facilities and the underlying software needed to simulate measurement uncertainty and perform statistical analyses. Together, these tools can perform end-to-end nuclear material accountancy analyses. This report documents the various improvements made to these tools in FY25. Specifically, we added new statistical test, new statistical modeling capabilities, new fuel cycle facility models, and launched a new open-source model component library.

97 MATHEMATICS AND COMPUTING↗

Performance evaluation of CMIP6 models on the Arctic-Siberian Plain teleconnection affecting the East Asian heat waves

The frequency and intensity of summer heat waves in East Asia have increased sharply in recent decades, significantly impacting public health and the economy. The Arctic-Siberian Plain (ASP) teleconnection pattern has been identified as a key driver, with ASP warming amplifying atmospheric circulation patterns conducive to extreme temperatures. This study evaluates the ability of Coupled Model Inter-comparison Project phase 6 models to simulate the ASP pattern across interannual variability (IAV) and intra-seasonal variability (ISV) timescales using the Common Basis Function method. The multi-model mean shows statistically significant pattern correlations with ERA5 reanalysis, with correlation coefficients of 0.90 and 0.99 for IAV and ISV, respectively. While the ASP pattern is generally well captured, models exhibit substantial inter-model diversity in the intensity and position of anticyclonic anomalies over the ASP and East Asia. Models with ASP pattern variability similar to reanalysis better reproduce extreme East Asian temperatures, whereas those over- or underestimating ASP variability exhibit lower skill. These performance differences are related to differences in simulating key variables associated with the development of the ASP pattern. Our findings highlight the role of the ASP pattern in modulating extreme heat events, as models with improved ASP simulations align more closely with observed temperature extremes. Refining ASP representations in models could enhance seasonal heat wave predictions, improving climate adaptation strategies.

Arctic-Siberian Plain (ASP)↗

Statistical Estimation of EV Driver Charging Behavior and Influential Factors

INL received data collected via telematics from battery electric vehicles (BEVs), and these vehicles were owned by retail customers who had entered into a telematics user agreement. The goal of analyzing these data was to develop mathematical models to characterize how different sets of BEV drivers use charging infrastructure at home and away from home (i.e., public charging) and quantify how various factors influence BEV drivers’ decision to charge and use available infrastructure. The data used in this analysis are unique because they provide real world BEV driving and charging behavior at the individual driving and parking event level. In this study we seek to leverage this data to quantify BEV charging and driving metrics to help inform models that predict quantities like the specific times when loads are imposed on the electrical grid due to BEV charging. Most models that have been developed to predict electrical grid load due to BEV charging, use simulations of BEV driving events and rely on assumptions such as every vehicle charges every night. Using a statistical modelling framework, we seek to investigate BEV charging behavior and quantitatively assess these common assumptions of BEV charging behavior.

33 - ADVANCED PROPULSION SYSTEMS↗

Thermal Gradient Effects on Local Hotspot Ignition in 1,3,5,7‐Tetranitro‐1,3,5,7‐tetrazocane (HMX)

Understanding hotspot ignition, growth, and criticality, as well as the timescales of each, is crucial for parameterizing mesoscale and continuum‐level models that rely on a statistical understanding of hotspots. However, these models often consider hotspots to have uniform temperatures or for hotspots of a given temperature and size to always behave the same. Therefore, using molecular dynamics simulations, we assess the influence of thermal distribution effects on hotspot local ignition and time to ignition in 1,3,5,7‐tetranitro‐1,3,5,7‐tetrazocane (HMX) nanoscale hotspots. Finally, by assessing hotspots with a gradient driven from an initial two‐temperature core–shell setup, we show that small increases in the core temperature under a constant average temperature can lead to order‐of‐magnitude effects on reaction and local ignition timescales.

36 MATERIALS SCIENCE↗

Model orthogonalization and Bayesian forecast mixing via principal component analysis

One can improve predictability in the unknown domain by combining forecasts of imperfect complex computational models using a Bayesian statistical machine learning framework. In many cases, however, the models used in the mixing process are similar. In addition to contaminating the model space, the existence of such similar, or even redundant, models during the multimodeling process can result in misinterpretation of results and deterioration of predictive performance. In this paper we describe a method based on the principal component analysis that eliminates model redundancy. We show that by adding model orthogonalization to the proposed Bayesian model combination framework, one can arrive at better prediction accuracy and reach excellent uncertainty quantification performance.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Improving North American Wildfire Prediction by Integrating a Machine-Learning Fire Model in a Land Surface Model

Wildfires have shown increasing trends in both frequency and severity across the Contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth System Models (ESMs). Alternatively, machine learning (ML) based fire models, which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ML4Fire-XGB) that integrates a pretrained eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM) version 2.1. A Fortran-C-Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001-2020, the ML4Fire-XGB model outperforms process-based fire models in terms of spatial distribution and seasonal variations. Sensitivity analysis confirms that the ML4Fire-XGB well captures the responses of the burned area to rising temperatures. The ML4Fire-XGB model has proved to be a new tool for studying vegetation-fire interactions, and more importantly, enables seamless exploration of climate-fire feedback, working as an active component in E3SM.

54 ENVIRONMENTAL SCIENCES↗

Simulated wildfire burned area over the CONUS during 2001-2020

Wildfires have shown increasing trends in both frequency and severity across the Contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth System Models (ESMs). Alternatively, machine learning (ML) based fire models, which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ML4Fire-XGB) that integrates a pretrained eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM). A Fortran-C-Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001-2020, the ML4Fire-XGB model outperforms process-based fire models in terms of spatial distribution and seasonal variations. Sensitivity analysis confirms that the ML4Fire-XGB well captures the responses of the burned area to rising temperatures. The ML4Fire-XGB model has proved to be a new tool for studying vegetation-fire interactions, and more importantly, enables seamless exploration of climate-fire feedback, working as an active component in E3SM.

Liu, Ye↗

Inverse design of a pyrochlore lattice of DNA origami through model-driven experiments

Sophisticated statistical mechanics approaches and human intuition have demonstrated the possibility of self-assembling complex lattices or finite-size constructs. However, attempts so far have mostly only been successful in silico and often fail in experiment because of unpredicted traps associated with kinetic slowing down (gelation, glass transition) and competing ordered structures. Theoretical predictions also face the difficulty of encoding the desired interparticle interaction potential with the experimentally available nano- and micrometer-sized particles. To overcome these issues, we combine SAT assembly (a patchy-particle interaction design algorithm based on constrained optimization) with coarse-grained simulations of DNA nanotechnology to experimentally realize trap-free self-assembly pathways. In this paper, we use this approach to assemble a pyrochlore three-dimensional lattice, coveted for its promise in the construction of optical metamaterials, and characterize it with small-angle x-ray scattering and scanning electron microscopy visualization.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection↗

PaleoSTeHM v1.0: a modern, scalable spatiotemporal hierarchical modeling framework for paleo-environmental data

Abstract. Geological records of past environmental change provide crucial insights into long-term climate variability, trends, non-stationarity, and nonlinear feedback mechanisms. However, reconstructing spatiotemporal fields from these records is statistically challenging due to their sparse, indirect, and noisy nature. Here, we present PaleoSTeHM, a scalable and modern framework for spatiotemporal hierarchical modeling of paleo-environmental data. This framework enables the implementation of flexible statistical models that rigorously quantify spatial and temporal variability from geological data while clearly distinguishing measurement and inferential uncertainty from process variability. We illustrate its application by reconstructing temporal and spatiotemporal paleo-sea-level changes across multiple locations. Using various modeling and analysis choices, PaleoSTeHM demonstrates the impact of different methods on inference results and computational efficiency. Our results highlight the critical role of model selection in addressing specific paleo-environmental questions, showcasing the PaleoSTeHM framework's potential to enhance the robustness and transparency of paleo-environmental reconstructions.

58 GEOSCIENCES↗

Mercury’s Chaotic Secular Evolution as a Subdiffusive Process

Abstract Mercury’s orbit can destabilize, generally resulting in a collision with either Venus or the Sun. Chaotic evolution can causeg 1 to decrease to the approximately constant value ofg 5 and create a resonance. Previous work has approximated the variation ing 1 as stochastic diffusion, which leads to a phenomological model that can reproduce the Mercury instability statistics of secular andN-body models on timescales longer than 10 Gyr. Here we show that the diffusive model significantly underpredicts the Mercury instability probability on timescales less than 5 Gyr, the remaining lifespan of the solar system. This is becauseg 1 exhibits larger variations on short timescales than the diffusive model would suggest. To better model the variations on short timescales, we build a new subdiffusive phenomological model forg 1 . Subdiffusion is similar to diffusion but exhibits larger displacements on short timescales and smaller displacements on long timescales. We choose model parameters based on the behavior of theg 1 trajectories in theN-body simulations, leading to a tuned model that can reproduce Mercury instability statistics from 1–40 Gyr. This work motivates fundamental questions in solar system dynamics: why does subdiffusion better approximate the variation ing 1 than standard diffusion? Why is there an upper bound ong 1 , but not a lower bound that would prevent it from reachingg 5 ?

Astronomy & Astrophysics↗

Stochastic room temperature creep of 316 L stainless steel

The creep behavior of 316 L stainless steel at room temperature was evaluated as a function of time and applied stress using a new high-throughput approach. Several common creep models were evaluated against the observations, leading to deeper analysis of a stress-dependent modified logarithmic creep model. Within this model, multiple sources of uncertainty were compared. Aleatoric stochastic variation between samples under nominally identical conditions was identified as the primary contributor to uncertainty in creep response. Under any particular set of conditions, the sample-to-sample variability in creep strain was as high as a factor of two, highlighting the engineering importance of characterizing large statistical datasets. The model's extrapolation capabilities were assessed by comparing predictions derived from calibration on partial, shorter-duration subsets of the data. In conclusion, these findings underscore the importance of accounting for stochastic effects in predictive modeling of aging phenomena.

High-throughput↗

Uncovering the truth about M101, NGC 3938, and their significant others through radiative transfer

ABSTRACT Solving the inverse problem in spiral galaxies, that allows the derivation of the spatial distribution of dust, gas, and stars, together with their associated physical properties, directly from panchromatic imaging observations, is one of the main goals of this work. To this end, we used radiative transfer models to decode the spatial and spectral distributions of the nearby face-on galaxies M101 and NGC 3938. In both cases, we provide excellent fits to the surface-brightness distributions derived from GALEX, SDSS, 2MASS, Spitzer, and Herschel imaging observations. Together with previous results from M33, NGC 628, M51, and the Milky Way, we obtain a small statistical sample of modelled nearby galaxies that we analyse in this work. We find that in all cases Milky Way-type dust with Draine-like optical properties provide consistent and successful solutions. We do not find any ‘submm excess’, and no need for modified dust-grain properties. Intrinsic fundamental quantities like star-formation rates (SFR), specific SFR (sSFR), dust opacities, and attenuations are derived as a function of position in the galaxy and overall trends are discussed. In the SFR surface density versus stellar mass surface density space, we find a structurally resolved relation (SRR) for the morphological components of our galaxies, that is steeper than the main sequence (MS). Exception to this is for NGC 628, where the SRR is parallel to the MS.

Pricopi, D.↗