Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection↗

PaleoSTeHM v1.0: a modern, scalable spatiotemporal hierarchical modeling framework for paleo-environmental data

Abstract. Geological records of past environmental change provide crucial insights into long-term climate variability, trends, non-stationarity, and nonlinear feedback mechanisms. However, reconstructing spatiotemporal fields from these records is statistically challenging due to their sparse, indirect, and noisy nature. Here, we present PaleoSTeHM, a scalable and modern framework for spatiotemporal hierarchical modeling of paleo-environmental data. This framework enables the implementation of flexible statistical models that rigorously quantify spatial and temporal variability from geological data while clearly distinguishing measurement and inferential uncertainty from process variability. We illustrate its application by reconstructing temporal and spatiotemporal paleo-sea-level changes across multiple locations. Using various modeling and analysis choices, PaleoSTeHM demonstrates the impact of different methods on inference results and computational efficiency. Our results highlight the critical role of model selection in addressing specific paleo-environmental questions, showcasing the PaleoSTeHM framework's potential to enhance the robustness and transparency of paleo-environmental reconstructions.

58 GEOSCIENCES↗

Mercury’s Chaotic Secular Evolution as a Subdiffusive Process

Abstract Mercury’s orbit can destabilize, generally resulting in a collision with either Venus or the Sun. Chaotic evolution can causeg 1 to decrease to the approximately constant value ofg 5 and create a resonance. Previous work has approximated the variation ing 1 as stochastic diffusion, which leads to a phenomological model that can reproduce the Mercury instability statistics of secular andN-body models on timescales longer than 10 Gyr. Here we show that the diffusive model significantly underpredicts the Mercury instability probability on timescales less than 5 Gyr, the remaining lifespan of the solar system. This is becauseg 1 exhibits larger variations on short timescales than the diffusive model would suggest. To better model the variations on short timescales, we build a new subdiffusive phenomological model forg 1 . Subdiffusion is similar to diffusion but exhibits larger displacements on short timescales and smaller displacements on long timescales. We choose model parameters based on the behavior of theg 1 trajectories in theN-body simulations, leading to a tuned model that can reproduce Mercury instability statistics from 1–40 Gyr. This work motivates fundamental questions in solar system dynamics: why does subdiffusion better approximate the variation ing 1 than standard diffusion? Why is there an upper bound ong 1 , but not a lower bound that would prevent it from reachingg 5 ?

Astronomy & Astrophysics↗

Stochastic room temperature creep of 316 L stainless steel

The creep behavior of 316 L stainless steel at room temperature was evaluated as a function of time and applied stress using a new high-throughput approach. Several common creep models were evaluated against the observations, leading to deeper analysis of a stress-dependent modified logarithmic creep model. Within this model, multiple sources of uncertainty were compared. Aleatoric stochastic variation between samples under nominally identical conditions was identified as the primary contributor to uncertainty in creep response. Under any particular set of conditions, the sample-to-sample variability in creep strain was as high as a factor of two, highlighting the engineering importance of characterizing large statistical datasets. The model's extrapolation capabilities were assessed by comparing predictions derived from calibration on partial, shorter-duration subsets of the data. In conclusion, these findings underscore the importance of accounting for stochastic effects in predictive modeling of aging phenomena.

High-throughput↗

Uncovering the truth about M101, NGC 3938, and their significant others through radiative transfer

ABSTRACT Solving the inverse problem in spiral galaxies, that allows the derivation of the spatial distribution of dust, gas, and stars, together with their associated physical properties, directly from panchromatic imaging observations, is one of the main goals of this work. To this end, we used radiative transfer models to decode the spatial and spectral distributions of the nearby face-on galaxies M101 and NGC 3938. In both cases, we provide excellent fits to the surface-brightness distributions derived from GALEX, SDSS, 2MASS, Spitzer, and Herschel imaging observations. Together with previous results from M33, NGC 628, M51, and the Milky Way, we obtain a small statistical sample of modelled nearby galaxies that we analyse in this work. We find that in all cases Milky Way-type dust with Draine-like optical properties provide consistent and successful solutions. We do not find any ‘submm excess’, and no need for modified dust-grain properties. Intrinsic fundamental quantities like star-formation rates (SFR), specific SFR (sSFR), dust opacities, and attenuations are derived as a function of position in the galaxy and overall trends are discussed. In the SFR surface density versus stellar mass surface density space, we find a structurally resolved relation (SRR) for the morphological components of our galaxies, that is steeper than the main sequence (MS). Exception to this is for NGC 628, where the SRR is parallel to the MS.

Pricopi, D.↗

Mass of 101 Sn and Bayesian extrapolations to the proton drip line

The favorable energy configurations of nuclei at magic numbers of 𝑁 neutrons and 𝑍 protons are fundamental for understanding the evolution of nuclear structure. The 𝑍 = 50 (tin) isotopic chain is a frontier for such studies, with particular interest at and around the doubly magic 100 Sn isotope, for which the mass is a topic of debate. Precise mass values for neutron-deficient isotopes provide necessary anchor points for mass models to test extrapolations near the proton drip line, where experimental studies remain out of reach. In this work, we report a Penning trap mass measurement of 101 Sn . The determined mass excess of −59889.89⁢(96) keV for 101 Sn represents a factor-of-300 improvement over the current precision and indicates that 101 Sn is less bound than previously thought. Mass predictions from a recently developed Bayesian model combination framework employing statistical machine learning and nuclear masses computed within seven global models based on nuclear density functional theory agree within 1⁢𝜎 with experimental masses from the 48 ≤ 𝑍 ≤ 52 isotopic chains. The framework's resilience to new mass data gave confidence in the extrapolation of tin masses down to 𝑁 = 46. Our calculations suggest that 96 Sn is a two-proton drip line nucleus and predict a mass excess of −58090⁢(800) keV for 100 Sn , showing a preference within 1⁢𝜎 for the mass of 100 Sn derived from the 𝛽-delayed 𝑄 value measured at GSI.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Recent evolution of risk analyses in atomic bomb survivor studies: new methods and applications

Abstract Several decades ago a dramatic leap forward occurred in the development and application of statistical methods for modeling radiation risk at the Radiation Effects Research Foundation (RERF). Poisson regression analysis for grouped person-year cohort data and the linear excess relative risk model were introduced, and subsequently a devoted software system, Epicure® (https://www.hirosoft.com), was developed by researchers at RERF and at the U.S. National Cancer Institute. Numerous advancements in understanding radiation effects on humans were made possible with these methods, which are still the state-of-the-art for risk assessment at RERF and have remained part of the standard toolbox for radiation—and other environmental—epidemiological studies worldwide. Nevertheless, as our understanding of radiation risk has increased, so have the breadth and depth of questions that require answers based on emerging data that are not amenable to these conventional methods. This overview briefly recounts the conventional methods and then describes our recent diversification into the use or development of new statistical approaches to meet the challenges of burgeoning biological data and emerging mechanistic information. We briefly discuss the development and application of new methods, current and planned, that are part of the RERF Statistics Department’s role in supporting institution-wide research, especially in our collaborations involving the Life Span Study, Adult Health Study, and First-generation Offspring Clinical Study. Some approaches to modeling and assessing radiation risk with newer methods mentioned herein have already been published, while some are still in development or are only beginning at the proposal stage.

Oncology↗

Collisional Excitation of HCN by CO to Refine the Modeling of Cometary Comae

Here, we present the first dataset of collisional (de)-excitation rate coefficients of HCN induced by CO, one of the main perturbing gases in cometary atmospheres. The dataset spans the temperature range of 5–50 K. It includes both state-to-state rate coefficients involving the lowest ten and nine rotational levels of HCN and CO, respectively, and the so-called “thermalized” rate coefficients over the rotational population of CO at each kinetic temperature. The derivation of these coefficients exploited the good performance of the statistical adiabatic channel model (SACM) on top of an accurate interaction potential computed at the CCSD(T)-F12b/CBS level of theory. The reliability of the SACM approach was validated by comparison with full quantum calculations restricted at the lowest total angular momentum of the system. These results provide essential input to accurately model the distribution among the rotational energy levels and the abundance of HCN in cometary atmospheres, accounting for deviations from local thermodynamic equilibrium that typically occurs in such environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Measuring quasiparticle dynamics for particle impact reconstruction in a superconducting qubit chip

Quasiparticle poisoning following particle impacts poses a significant challenge to the development of fault-tolerant superconducting quantum computers, as a sudden excess of quasiparticles can simultaneously degrade the coherence of multiple qubits across large device arrays. In this work, we present a statistical analysis that models the time evolution of radiation-induced qubit energy relaxation through quasiparticle density dynamics. This study provides insight into quasiparticle loss processes by distinguishing between recombination and trapping decay channels and assessing their respective impact on qubit performance. We precisely measure quasiparticle recombination in multiple transmon qubits and uncover an unexpected dependence of qubit relaxation dynamics on deposited energy. By linking correlated relaxation events across qubits to ballistic phonon propagation, we introduce a statistical localization approach to extract the energy deposited in the substrate, which is in good agreement with Monte Carlo simulation. This work establishes the quantitative framework for using an arbitrary subset of superconducting transmon qubits in a QPU as energy-resolving witness particle detectors.

Celi, E. [Northwestern U.]↗

Yet Another Discriminant Analysis (YADA): A Probabilistic Model for Machine Learning Applications

This paper presents a probabilistic model for various machine learning (ML) applications. While deep learning (DL) has produced state-of-the-art results in many domains, DL models are complex and over-parameterized, which leads to high uncertainty about what the model has learned, as well as its decision process. Further, DL models are not probabilistic, making reasoning about their output challenging. In contrast, the proposed model, referred to as Yet Another Discriminate Analysis(YADA), is less complex than other methods, is based on a mathematically rigorous foundation, and can be utilized for a wide variety of ML tasks including classification, explainability, and uncertainty quantification. YADA is thus competitive in most cases with many state-of-the-art DL models. Ideally, a probabilistic model would represent the full joint probability distribution of its features, but doing so is often computationally expensive and intractable. Hence, many probabilistic models assume that the features are either normally distributed, mutually independent, or both, which can severely limit their performance. YADA is an intermediate model that (1) captures the marginal distributions of each variable and the pairwise correlations between variables and (2) explicitly maps features to the space of multivariate Gaussian variables. Numerous mathematical properties of the YADA model can be derived, thereby improving the theoretic underpinnings of ML. Validation of the model can be statistically verified on new or held-out data using native properties of YADA. However, there are some engineering and practical challenges that we enumerate to make YADA more useful.

97 MATHEMATICS AND COMPUTING↗

Forest aboveground biomass estimation through integration of sentinel-2 and PALSAR-2 time series: assessing models trained on GEDI and field inventory benchmarks

Accurate and spatially explicit forest Aboveground Biomass (AGB) mapping through remote sensing is critical for quantifying terrestrial carbon stocks and informing effective forest management strategies. However, AGB estimation in dense forests with complex terrain remains challenging due to satellite sensor signal saturation problem (saturation issue occurs in high biomass forests), structural complexity, and limited ground truth for calibration. This study presents a novel framework that integrates multi-temporal Sentinel-2 optical imagery, ALOS PALSAR-2 Synthetic Aperture Radar (SAR) data, and topographic variables with explainable Machine Learning to map AGB across mountainous forests within subtropical and temperate oceanic climate zones of Mexico. We evaluate the effects of temporal granularity and sensor synergy by comparing multiple temporal inputs and sensor configurations (Sentinel-2, PALSAR-2, and their fusion), and assess model performance using two reference datasets: NASA GEDI LiDAR-derived biomass and Mexico’s National Forest and Soil Inventory (INFyS). Our results showed that models trained on INFyS consistently outperformed those trained on GEDI, highlighting limitations in GEDI’s reliability in biomass estimates within this study region. Furthermore, the integration of Sentinel-2 and PALSAR-2 provided improved predictions compared to single-sensor models, particularly when combined with temporally explicit yearly statistics. The best-performing model, which was trained on INFyS data, and considered both Sentinel-2 and PALSAR-2 yearly statistics, as well as topographic variables, achieved an R2 of 0.64, RMSE of 51.10 Mg/ha, and relative RMSE (rRMSE) of 58.69%. Explainable ML analysis identified Sentinel-2 spectral indices and topographic features as key predictors, while PALSAR-2 metrics provided complementary information, partially mitigating saturation effects in high-biomass areas. Specifically, integrating both sensors substantially improved AGB estimation in high biomass forest (≥200 Mg/ha), yielding 98% gains over optical-only model, with resulting estimates exceeding GEDI L4B by 29% and ESA-CCI-BIOMASS by 174%. Terrain-stratified analysis indicated close agreement with GEDI in low-slope areas, with increasing divergence as slope steepness increased, while estimates remained consistently higher than ESA-CCI-BIOMASS across all slope classes. The proposed approach advances multi-sensor fusion and temporal feature engineering for AGB mapping using open-access satellite datasets, providing a scalable and reproducible framework for annual biomass monitoring in topographically complex mountainous forests. The resulting 25 m resolution biomass product has the potential to provide spatially detailed information for forest monitoring and may support applications in carbon accounting and forest management.

54 ENVIRONMENTAL SCIENCES↗

Mesoscale Convective Systems Tracking Method Intercomparison (MCSMIP): Application to DYAMOND Global km‐Scale Simulations

Abstract Global kilometer‐scale models represent the future of Earth system modeling, enabling explicit simulation of organized convective storms and their associated extreme weather. Here, we comprehensively evaluate tropical mesoscale convective system (MCS) characteristics in the DYAMOND (DYnamics of the atmospheric general circulation modeled on non‐hydrostatic domains) simulations for both summer and winter phases. Using 10 different feature trackers applied to simulations and satellite observations, we assess MCS frequency, precipitation, and other key characteristics. Substantial differences (a factor of 2–3) arise among trackers in observed MCS frequency and their precipitation contribution, but model‐observation differences in MCS statistics are more consistent across trackers. DYAMOND models are generally skillful in simulating tropical mean MCS frequency, with multi‐model mean biases ranging from −2%–8% over land and −8%–8% over ocean (summer vs. winter). However, most DYAMOND models underestimate MCS precipitation amount (23%) and their contribution to total precipitation (17%). Biases in precipitation contributions are generally smaller over land (13%) than over ocean (21%), with moderate inter‐model variability. While models better simulate MCS diurnal cycles and cloud shield characteristics, they overestimate MCS precipitation intensity and underestimate stratiform rain contributions (up to a factor of 2), particularly over land, albeit observational uncertainties exist. Additionally, models exhibit a wide range of precipitable water in the tropics compared to reanalysis and satellite observations, with many models showing exaggerated sensitivity of MCS precipitation intensity to precipitable water. The MCS metrics developed here provide process‐oriented diagnostics to guide future model development.

54 ENVIRONMENTAL SCIENCES↗

Micromechanical Surrogate Machine Learning Model for Creep Deformation Modeling

Process variability during the manufacture of gas turbine engine hot section components can significantly affect the material’s resulting microstructure. In casting, for instance, geometric variation within a component (thin sections versus thick sections, radial location) influences cooling rates and the resulting grain size. The high temperature creep response is known to be sensitive to grain size owing to a diffusional creep mechanism which occurs more readily along grain boundaries. Microstructural variation correspondingly drives mechanical behavior which propagates into component scale performance uncertainty. These factors are essential when planning inspection, maintenance, and repair strategies within a reliability framework. These benefits provide opportunities to increase overall energy efficiency through refined margins. Critically, there is an opportunity to bolster existing data-driven reliability models using physics-driven process-structure-property relations. Here we present recent work establishing a framework for evaluating the probabilistic creep performance of high-temperature materials. A novel microstructure-sensitive crystal plasticity finite element model is established that captures both grain boundary and crystallographic deformation effects. The computationally expensive physics model is calibrated using a statistical approach and this high-fidelity model is subsequently used to train a computationally efficient machine learning surrogate model. The surrogate model is essential for sampling a large ensemble of simulated structure-property pair results. The ensemble data are then mined to extract salient trends to be incorporated into a microstructure-sensitive reliability model. The proposed approach represents a novel way to capture microstructure-sensitive trends from physics-based models within a modern reliability framework.

Fernandez-Zelaia, Patxi [ORNL]↗

Data driven investigation to understand the influence of total solids on biological biogas upgrading

In situ biogas upgrading achieves CO 2 conversion to CH 4 via hydrogenotrophic methanogenesis; however, gas-liquid mass transfer constraints limit the upgrading performance. Recognizing that optimization studies often underrepresent the effects of total solids (TS) and organic loading rate (OLR), this study undertook a holistic, statistics driven assessment of operating conditions for in situ H 2 assisted biogas upgrading, centering the analysis on TS and OLR. A dataset of 31 studies was compiled and comprised 99 observations. A rigorous analytical framework was employed, combining data standardization, fixed- and random-effects (REML) weighted regressions with cluster-robust errors, stratified analyses, and machine learning. Mixed-effects meta regression indicated that TS was the main factor explaining differences of methane fraction (CH 4 %) when considering the between studies heterogeneity. Focusing on a near-stoichiometric subset (H 2 /CO 2 ≈ 4:1), TS remained significant. Stratified results showed a stronger negative relationship between TS and CH 4 % in UASB reactors than in CSTRs, with a negative effect under mesophilic conditions and no significant effect under thermophilic conditions. A Random Forest model corroborated the statistical findings, consistently ranking H 2 /CO 2 ratio, OLR, TS, and hydrogen injection rate (HIR) as the most influential predictors. These findings delineate trends across increasing TS levels, particularly between 1% and 10%, and provide preliminary insights for TS above 15% in in situ biogas upgrading. They further provide insights for the influence of TS by reactor type and temperature, thereby advancing the evidence base for implementing biological CO 2 conversion to CH 4 in practice.

In situ biogas upgrading↗

Analysis on Evaluations of Monterey Bay Aquarium Research Institute’s Wave Energy Converter’s Field Data Using WEC-Sim and Gazebo: A Simulation Tool Comparison

Although many studies have validated wave energy converter (WEC) numerical models against scaled prototype experimental data, there remains a notable lack of validation using data from full-scale deployed WECs. This paper compares two numerical models of Monterey Bay Aquarium Research Institute’s Wave Energy Converter (MBARI-WEC), a two-body point absorber with an electro-hydraulic power take-off system (PTO). The models are implemented in WEC-Sim/Simscape and Gazebo Simulator. A statistical analysis of the models was performed, and field results were obtained to compare the models’ accuracy in predicting the RMS piston velocity, RMS motor speed, and mean electric power compared to field data for 56 observations across varying sea states. The Gazebo model demonstrated a closer agreement across all three parameters for a majority of the observations. When compared to the field data, the Gazebo and WEC-Sim models exhibited average mean electric power overestimations of 13% and 22%, respectively.

16 TIDAL AND WAVE POWER↗

Analysis of the Challenges in Developing Sample-Based Multi-fidelity Estimators for Non-deterministic Models

Multifidelity (MF) uncertainty quantification (UQ) seeks to leverage and fuse information from a collection of models to achieve greater statistical accuracy with respect to a single-fidelity counterpart, while maintaining an efficient use of computational resources. Despite many recent advancements in MF UQ, several challenges remain and these often limit its practical impact in certain application areas. In this manuscript, we focus on the challenges introduced by nondeterministic models to sampling MF UQ estimators. Nondeterministic models produce different responses for the same inputs, which means their outputs are effectively noisy. MF UQ is complicated by this noise since many state-of-the-art approaches rely on statistics, e.g., the correlation among models, to optimally fuse information and allocate computational resources. Here, we demonstrate how the statistics of the quantities of interest, which impact the design, effectiveness, and use of existing MF UQ techniques, change as functions of the noise. With this in hand, we extend the unifying approximate control variate framework to account for nondeterminism, providing for the first time a rigorous means of comparing the effect of nondeterminism on different multifidelity estimators and analyzing their performance with respect to one another. Numerical examples are presented throughout the manuscript to illustrate and discuss the consequences of the presented theoretical results.

97 MATHEMATICS AND COMPUTING↗

Precise Modeling of a Complex Solenoidal Magnetic Field Using a Combination of Analytic Functions and a PINN

We demonstrate an iterative approach to modeling a sparsely measured magnetic field in a large-bore solenoid. This approach uses a hybrid of traditional and machine learning techniques. The traditional technique is a linear least-squares fit using a series solution to Laplace's equation, while the machine learning technique involves the training of a physics-informed neural network (PINN) on the least-squares fit residuals. We use a newly defined activation function "DELTAsnake," a modification to the snake activation function proposed by Ziyin et al. that allows for stronger curvature and non-monotonicity. The combined model approximately obeys Maxwell's equations to a level sufficient for producing high quality physics simulations and analysis. Our approach is applied to a highly realistic calculation of the expected magnetic field in the Mu2e experiment's Detector Solenoid which includes a simple model for the expected statistical measurement uncertainties. Using ten toy measurement simulations, we demonstrate the capabilities of our model in comparison to the least-squares method alone; the least-squares method alone results in a reduced chi-squared statistic of ${2.15 \pm 0.01}$, while our approach improves the reduced chi-square to ${1.034 \pm 0.005}$. Furthermore, for an average toy simulation, we show that the range of the RMS of the three field component residuals reduces from ${0.07-0.37}$ Gauss to ${0.05-0.07}$ Gauss. We find that this novel method is robust against a realistic systematic uncertainty deriving from Hall probe calibration bias and can be used to significantly reduce the number of measurements required to achieve an accurate model.

Kampa, Cole [Caltech] (ORCID:0000000192972920)↗

Performance Prediction of High‐Entropy Perovskites La 0.8 Sr 0.2 Mn x Co y Fe z O 3 with Automated High‐Throughput Characterization of Combinatorial Libraries and Machine Learning

Perovskite oxides form a large family of materials with applications across various fields, owing to their structural and chemical flexibility. Efficient exploration of this extensive compositional space is now achievable through automated high-throughput experimentation combined with machine learning. In this study, we investigate the composition–structure–performance relationships of high-entropy La 0.8 Sr 0.2 Mn x Co y Fe z O 3±𝞭 perovskite oxides (0 < x, y, z <1; x+y+z≈1) for application as oxygen electrodes in Solid Oxide Cells. Following the deposition of a continuous compositional map using thin-film combinatorial pulsed laser deposition, compositional, structural, and performance properties are characterized using six different techniques with mapping capabilities. Random forests effectively model electrochemical performance, consistently identifying Fe-rich oxides as optimal compounds with the lowest area-specific resistance values for oxygen electrodes at 700 °C. Additionally, the models identify a statistical correlation between oxygen sublattice distortion—derived from spectral analysis of Raman-active modes—and enhanced performance.

high entropy oxides↗