Search NASASearch

SEARCH · Search NASA

Results for “Statistical methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP)

Anthropogenic climate change is unfolding rapidly, yet its regional manifestation can be obscured by internal variability. A primary goal of climate science is to identify the externally forced climate response from among the noise of internal variability. Separating the forced response from internal variability can be addressed in climate models by using a large ensemble to average over different possible realizations of internal variability. However, with only one realization of the real world, it is a major challenge to isolate the forced response directly in observations. In the Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP), contributors used existing and newly developed statistical and machine learning methods to estimate the forced response over 1950–2022 within individual realizations of the climate system. Participants used neural networks, linear inverse models, fingerprinting methods, and low-frequency component analysis, among other approaches. These methods were trained using large ensembles from multiple climate models and then applied to observations. Here, we evaluate method performance within large ensembles and investigate the estimates of the forced response in observations. Our results show that many different types of methods are skillful for estimating the forced response in climate models, though the relative skill of individual methods varies depending on the variable and evaluation metric. Methods with comparable skill in models can give a wide range of estimates of the forced response pattern in observations, illustrating the epistemic uncertainty in forced response estimates. ForceSMIP gives new insights into the forced response in observations, its uncertainty, and methods for its estimation.

Climate attribution

Testing convolutional neural network based deep learning systems: a statistical metamorphic approach

Machine learning technology spans many areas and today plays a significant role in addressing a wide range of problems in critical domains,i.e., healthcare, autonomous driving, finance, manufacturing, cybersecurity,etc. Metamorphic testing (MT) is considered a simple but very powerful approach in testing such computationally complex systems for which either an oracle is not available or is available but difficult to apply. Conventional metamorphic testing techniques have certain limitations in verifying deep learning-based models (i.e., convolutional neural networks (CNNs)) that have a stochastic nature (because of randomly initializing the network weights) in their training. In this article, we attempt to address this problem by using a statistical metamorphic testing (SMT) technique that does not require software testers to worry about fixing the random seeds (to get deterministic results) to verify the metamorphic relations (MRs). We propose seven MRs combined with different statistical methods to statistically verify whether the program under test adheres to the relation(s) specified in the MR(s). We further use mutation testing techniques to show the usefulness of the proposed approach in the healthcare space and test two CNN-based deep learning models (used for pneumonia detection among patients). The empirical results show that our proposed approach uncovers 85.71% of the implementation faults in the classifiers under test (CUT). Furthermore, we also propose an MRs minimization algorithm for the CUT, thus saving computational costs and organizational testing resources.

Computer Science

Causal relationships of vegetation productivity with root zone water availability and atmospheric dryness at the catchment scale

Abstract. This study explores the causal relationships between catchment water availability, vapor pressure deficit, and gross primary productivity (GPP) across 341 catchments in the contiguous US. Seasonal climatic, hydrological, and vegetation characteristics were represented using the Horton index, ecological aridity index, evaporative fraction index, and carbon uptake efficiency. Statistical methods, including circularity statistics, correlation analysis, and causality tests, were employed to determine the complex interactions between catchment wetness, atmospheric dryness, and vegetation carbon uptake. The results revealed a maximum lag of 2 months in the intra-annual variability of catchment water supply–productivity and atmospheric water demand–productivity relationships, with hysteresis patterns varying with the catchment's hydrological characteristics. In catchments not permanently under water-limited or energy-limited conditions, vegetation experiences hydrological stress during the peak growing period, coinciding with the highest gross primary productivity and carbon uptake efficiency being out of phase with the Horton index and in phase with the evaporative fraction index. Causality analysis highlights strong temporal continuity in GPP seasonal characteristics, with a cause–effect relationship between catchment water supply, atmospheric demand, and vegetation productivity spanning a maximum of 2 months. These findings underscore the need for a comprehensive functional framework that integrates catchment water supply, atmospheric demand, and vegetation productivity to enhance our understanding and predictive capabilities with regard to ecosystem responses to climate change.

54 ENVIRONMENTAL SCIENCES

New Probe of Cosmic Birefringence Using Galaxy Polarization and Shapes

We propose a novel statistical method to measure cosmic birefringence and demonstrate its power in probing parity violation due to axions. Exploiting an empirical correlation between the integrated radio polarization direction of a spiral galaxy and its apparent shape, we devise an unbiased minimum-variance estimator for the rotation angle, which should achieve an uncertainty of 5°–15° per galaxy. In conclusion, large galaxy samples from the forthcoming SKA continuum surveys, together with optical shape catalogs, promise a comparable or even lower noise power spectrum for the rotation angle than in the CMB Stage-IV (CMB-S4) experiment, with different systematics.

Axion-like particles

Unique & challenging aspects of plutonium metal standards exchange program for actinide measurements

The Los Alamos National Laboratory exchange program is the only program of its kind for the distribution of plutonium (Pu) standards materials with a range of impurity contents to multiple laboratories for destructive measurements of elemental concentration. This paper discusses statistical methods used to address challenges in Pu metal exchange data by way of two case studies. Challenges include how to evaluate a data set when a large fraction of the values are minimum detection limits (MDLs), and how to determine potential outliers with limited in-formation on the true spread of the data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Recent evolution of risk analyses in atomic bomb survivor studies: new methods and applications

Abstract Several decades ago a dramatic leap forward occurred in the development and application of statistical methods for modeling radiation risk at the Radiation Effects Research Foundation (RERF). Poisson regression analysis for grouped person-year cohort data and the linear excess relative risk model were introduced, and subsequently a devoted software system, Epicure® (https://www.hirosoft.com), was developed by researchers at RERF and at the U.S. National Cancer Institute. Numerous advancements in understanding radiation effects on humans were made possible with these methods, which are still the state-of-the-art for risk assessment at RERF and have remained part of the standard toolbox for radiation—and other environmental—epidemiological studies worldwide. Nevertheless, as our understanding of radiation risk has increased, so have the breadth and depth of questions that require answers based on emerging data that are not amenable to these conventional methods. This overview briefly recounts the conventional methods and then describes our recent diversification into the use or development of new statistical approaches to meet the challenges of burgeoning biological data and emerging mechanistic information. We briefly discuss the development and application of new methods, current and planned, that are part of the RERF Statistics Department’s role in supporting institution-wide research, especially in our collaborations involving the Life Span Study, Adult Health Study, and First-generation Offspring Clinical Study. Some approaches to modeling and assessing radiation risk with newer methods mentioned herein have already been published, while some are still in development or are only beginning at the proposal stage.

Oncology

Quantification and prediction of solidification textures under additive manufacturing conditions

Crystallographic textures are a major determinant of the macroscale anisotropic properties of polycrystalline metallic alloys produced in a wide range of additive manufacturing (AM) processes. Here, we introduce a statistical method that can accurately quantify the degree of orientational order of textures despite the large random fluctuations in the orientation of individual grains inherent in AM processes. The method, demonstrated for laser and resolidification of AlSi thin films, extends Z-scoring to a dynamical regime to assess the statistical significance of observed textures compared to randomly generated ones at different stages of solidification. We further show that, combined with phase-field modeling, this method can be used to infer fundamental anisotropic properties of the solid-liquid interface that are essential for texture prediction, and are compared here to the results of atomistic simulations. In addition, phase-field modeling reveals that, even at rapid AM solidification rates, the observed 〈110〉-dominated textures in the AlSi thin films are controlled predominantly by the anisotropy of the interface free-energy and sheds light on the physical mechanism of grain competition. These results significantly enhance both the existing tools for the quantification and prediction of AM crystallographic textures and our basic understanding of their formation.

36 MATERIALS SCIENCE

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection

Measurement of the W-boson mass and width with the ATLAS detector using proton–proton collisions at $\sqrt{s}=7$ TeV

Proton-proton collision data recorded by the ATLAS detector in 2011, at a centre-of-mass energy of 7 TeV, have been used for an improved determination of the W-boson mass and a first measurement of the W-boson width at the LHC. Recent fits to the proton parton distribution functions are incorporated in the measurement procedure and an improved statistical method is used to increase the measurement precision. The measurement of the W-boson mass yields a value of m w = 80,366.5 ± 9.8 (stat.) ± 12.5 (syst.) MeV = 80,366.5 ± 15.9 MeV, and the width is measured as Γ w = 2202 ± 32 (stat.) ± 34 (syst.) MeV = 2202 ± 47 MeV. The first uncertainty components are statistical and the second correspond to the experimental and physics-modelling systematic uncertainties. Both results are consistent with the expectation from fits to electroweak precision data. The present measurement of m w is compatible with and supersedes the previous measurement performed using the same data.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Destructive Analysis of TRISO Particles: Crush/Burn/Leach Followed by Davies-Gray Titration and IDMS

The accurate accounting of nuclear materials is a cornerstone of international nuclear safeguards. One emerging challenge in this domain is the fabrication of TRIstructural ISOtropic (TRISO) particle fuels. Although these innovative fuel forms are critical for advanced reactor applications, their robust refractory ceramics and coating compositions present significant obstacles to destructive analysis (DA) methods. Ensuring full and quantitative recovery from these particles is essential for accurate mass accountancy. The current study was initiated to address these challenges, first by validating a previously established destructive method developed by Oak Ridge National Laboratory (ORNL) for the quantitative recovery of uranium from TRISO particles and then following that process with uranium content determination through isotope dilution mass spectrometry (IDMS) and Davies-Gray titration. This study expands on the scope of a digestive method that was developed under the Advanced Gas Reactor Fuel Development and Qualification program and is currently implemented in both the Coated Particle Fuel Development Laboratory and Irradiated Fuels Examination Laboratory at ORNL. The success of the previous Advanced Gas Reactor work relied on developing a DA method to evaluate the fabrication process and reactor experiments. The methodology described in this report was designed to rigorously investigate the efficacy of the crush/burn/leach sample preparation of TRISO particles; it aims to quantify uranium recovery while also assessing the effects of TRISO constituents (e.g., silicon and zirconium) on analytical precision and accuracy. By comparing the results from the titration method and IDMS, we sought to determine whether existing analytical procedures accepted by the International Atomic Energy Agency (IAEA) could be effectively translated to TRISO fuel forms. The team employed an approach that involved processing replicate TRISO samples, optimizing the milling (i.e., crushing) step and performing serial leaches. The elemental composition of the analytical samples was examined to prepare for interference studies in the second year of this project. The integration of gamma spectrometry to verify residual uranium activity further strengthened the validation. Statistical methods were applied to the collected data to evaluate the uncertainties arising from sampling, sample preparation and uranium quantification. These uncertainties were then compared to the IAEA’s international target values (ITVs). Additional data collected in upcoming project work will strengthen the uncertainty estimates. Ultimately, it is hoped that this project will contribute materially to the body of work related to characterization of TRISO based fuels for the purpose of material accountancy and its applications to international safeguards.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Data-driven upper bounds and event attribution for unprecedented heatwaves

The last decade has seen numerous record-shattering heatwaves in all corners of the globe. In the aftermath of these devastating events, there is interest in identifying worst-case thresholds or upper bounds that quantify just how hot temperatures can become. Generalized Extreme Value theory provides a data-driven estimate of extreme thresholds; however, upper bounds may be exceeded by future events, which undermines attribution and planning for heatwave impacts. Here, we show how the occurrence and relative probability of observed yet unprecedented events that exceed a priori upper bound estimates, so-called “impossible” temperatures, has changed over time. We find that many unprecedented events are actually within data-driven upper bounds, but only when using modern spatial statistical methods. Furthermore, there are clear connections between anthropogenic forcing and the “impossibility” of the most extreme temperatures. Robust understanding of heatwave thresholds provides critical information about future record-breaking events and how their extremity relates to historical measurements.

54 ENVIRONMENTAL SCIENCES

Quantifying biases in stellar masses of JWST high- z quasar host galaxies caused by quasar subtraction

The James Webb Space Telescope (JWST) has enabled dozens of high-z quasar host galaxy detections. Many of these observations imply galaxies with black holes that are overmassive compared to their low-z counterparts. However, the bright quasar point source removal can cause significant biases in recovered host magnitudes and stellar mass measurements due to the degeneracy in host galaxy and quasar light. We develop a statistical method to disentangle the quasar host galaxy stellar mass measurements from observational biases during the point source removal assuming the PSF is modelled perfectly. We use the BlueTides simulation to generate mock images and perform point source removal on thousands of simulated high-z quasar host galaxies, constructing corrected host magnitude posteriors. We find that removing a bright quasar in JWST photometry tends to either correctly recover or modestly misestimate host magnitudes, with a maximum magnitude underestimate of 0.2 mag. With our corrected magnitude posteriors, we perform SED fitting on each quasar host galaxy and compare the stellar mass measurement before and after the correction. We find that stellar mass estimates are generally robust, or misestimated by $<$ 0.3 dex. We also find that the stellar masses of a subset of hosts (J0844−0132, J0911+0152, and J1146−0005) remain unconstrained, as key photometric bands provide only flux upper limits. Accounting for observational biases does not resolve the apparent mismatch between black hole and host galaxy growth at high-z, where some quasars appear to host overmassive black holes while others reside in relatively massive galaxies.

79 ASTRONOMY AND ASTROPHYSICS

Predicting RNA structure and dynamics with deep learning and solution scattering

Advanced deep learning and statistical methods can predict structural models for RNA molecules. However, RNAs are flexible, and it remains difficult to describe their macromolecular conformations in solutions where varying conditions can induce conformational changes. Small-angle x-ray scattering (SAXS) in solution is an efficient technique to validate structural predictions by comparing the experimental SAXS profile with those calculated from predicted structures. There are two main challenges in comparing SAXS profiles to RNA structures: the absence of cations essential for stability and charge neutralization in predicted structures and the inadequacy of a single structure to represent RNA’s conformational plasticity. We introduce a solution conformation predictor for RNA (SCOPER) to address these challenges. This pipeline integrates kinematics-based conformational sampling with the innovative deep learning model, IonNet, designed for predicting Mg 2+ ion binding sites. Validated through benchmarking against 14 experimental data sets, SCOPER significantly improved the quality of SAXS profile fits by including Mg 2+ ions and sampling of conformational plasticity. We observe that an increased content of monovalent and bivalent ions leads to decreased RNA plasticity. Therefore, carefully adjusting the plasticity and ion density is crucial to avoid overfitting experimental SAXS data. SCOPER is an efficient tool for accurately validating the solution state of RNAs given an initial, sufficiently accurate structure and provides the corrected atomistic model, including ions.

59 BASIC BIOLOGICAL SCIENCES

Leveraging interpolation models and error bounds for verifiable scientific machine learning

Effective verification and validation techniques for modern scientific machine learning workflows are challenging to devise. Statistical methods are abundant and easily deployed, but often rely on speculative assumptions about the data and methods involved. Error bounds for classical interpolation techniques can provide mathematically rigorous estimates of accuracy, but often are difficult or impractical to determine computationally. Here, in this work, we present a best-of-both-worlds approach to verifiable scientific machine learning by demonstrating that (1) multiple standard interpolation techniques have informative error bounds that can be computed or estimated efficiently; (2) comparative performance among distinct interpolants can aid in validation goals; (3) deploying interpolation methods on latent spaces generated by deep learning techniques enables some interpretability for black-box models. We present a detailed case study of our approach for predicting lift-drag ratios from airfoil images. Code developed for this work is available in a public Github repository.

97 MATHEMATICS AND COMPUTING

Formation of 1 H -Phenalene (C 13 H 10 ) in the Taurus Molecular Cloud via Methylidyne Addition-Cyclization-Aromatization (MACA)

The formation of 1H-phenalene (C 13 H 10 ) in cold molecular clouds, such as the Taurus Molecular Cloud-1 (TMC-1), presents a significant challenge to traditional astrochemical models, which predominantly suggest high-temperature pathways for polycyclic aromatic hydrocarbon (PAH) formation. In this study, we explore computationally the Methylidyne Addition-Cyclization-Aromatization (MACA) mechanism as a viable, barrierless pathway for phenalene synthesis under low-temperature conditions. Through electronic structure calculations and Rice–Ramsperger–Kassel–Marcus (RRKM) statistical methods, we demonstrate that the reaction of 1-vinylnaphthalene (C 10 H 7 C 2 H 3 ) with the methylidyne radical (CH) leads to the formation of 1H-phenalene via a bimolecular reaction, a process that is exoergic and without entrance barrier. The MACA mechanism facilitates the growth of the aromatic carbon backbone via a [5 + 1] ring annulation, providing a new insight into PAH formation in cold molecular clouds. Notably, the MACA mechanism has previously been shown to form indene (C9H8), which was detected in TMC-1 as well, via a [4 + 1] annulation, demonstrating its potential to produce a variety of complex PAHs by addition of a five- and six-membered ring to a benzene moiety via [4 + 1] and [5 + 1] annulation, respectively. As a result, this work highlights the importance of barrierless, exoergic reactions involving MACA in the synthesis of complex aromatic molecules in space, expanding our physicochemical understanding of carbon-rich chemistry in cold molecular clouds.

Aromatic compounds

Autonomous platform for solution processing of electronic polymers

The manipulation of electronic polymers’ solid-state properties through processing is crucial in electronics and energy research. Yet, efficiently processing electronic polymer solutions into thin films with specific properties remains a formidable challenge. We introduce Polybot, an artificial intelligence (AI) driven automated material laboratory designed to autonomously explore processing pathways for achieving high-conductivity, low-defect electronic polymers films. Leveraging importance-guided Bayesian optimization, Polybot efficiently navigates a complex 7-dimensional processing space. In particular, the automated workflow and algorithms effectively explore the search space, mitigate biases, employ statistical methods to ensure data repeatability, and concurrently optimize multiple objectives with precision. The experimental campaign yields scale-up fabrication recipes, producing transparent conductive thin films with averaged conductivity exceeding 4500 S/cm. Feature importance analysis and morphological characterizations reveal key design factors. This work signifies a significant step towards transforming the manufacturing of electronic polymers, highlighting the potential of AI-driven automation in material science.

Wang, Chengshi [Argonne National Laboratory (ANL),

Anticipating decoherence in quantum systems

Large-scale quantum technologies require coherence across distant nodes, necessitating indistinguishable quantum states. However, environmental disorder, including dephasing, spectral diffusion, and spin-bath interactions, undermines coherence. Using statistical methods, we uncover correlations in decoherence channels induced by slowly varying environments. Spectral diffusion serves as a representative demonstration case that can be extended to other remote, disordered systems such as spins in nitrogen-vacancy centers and quantum-dot spin qubits, as well as flux noise in superconducting qubits. In this work, we employ replica-theory-inspired trajectory analysis to reveal predictable temporal structures in decoherence dynamics, and validate these through an anticipatory systems framework with internal prediction of unseen spectral dynamics in multiple quantum systems, showing that this framework could, if implemented, reduce spectral shift by average factors of approximately 2 to 19, depending on emitter stability, thereby enabling enhanced coherence and multi-node synchronization for scalable quantum communication, computation, imaging, and sensing.

Maan, Pranshu [Purdue University]