Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Reliable statistics-based detection and investigation of anomalies in a SMART valve system

Reliable anomaly detection and diagnosis are critical for the safe operation of complex engineered systems. This study presents a unified framework that integrates statistical, model-based, and data-driven techniques for anomaly detection and investigation, demonstrated on SMART valve systems in hybrid energy applications. Four detection methods—mean deviation, seasonal extreme studentized deviate, ARIMA forecasting, and matrix profiling—were implemented and compared. Matrix profiling was particularly effective in revealing subtle deviations and hidden relationships among variables. Anomaly investigation was performed by analyzing variable-level and grouped signal profiles, with system topology incorporated to distinguish primary faults from propagated effects. Grouping signals by type enhanced interpretability, enabling accurate localization of anomalies across multi-dimensional datasets. Experimental results confirmed the framework's capability to consistently detect and isolate anomalies while providing actionable insights into system interdependencies. The proposed methodology offers a robust, interpretable, and scalable solution for condition monitoring, with potential applications in safety-critical domains such as nuclear energy, aerospace, and process industries.

ARIMA models↗

Thermal shock resistance of additively manufactured alumina

The mechanical behavior and cracking patterns of thermally-shocked additively manufactured alumina were investigated. The flexural strength of test specimens that had been heated to temperatures ranging from 200°C to 1000°C and then rapidly quenched in water was determined at ambient temperature by four-point bending. Results indicated that the surface cracking patterns had a multifractal structure and that an increase in the thermal shock temperature led to an increase in the density and uniformity of the crack network. Further, the flexural strength results were analyzed with Weibull statistics, where the Weibull moduli for most of the thermal shock conditions tested were found to be statistically indistinguishable. It was also found that a significant decrease (~50%) in flexural strength occurred for heating temperatures ≥300°C. The effect of the manufacturing method on cracking patterns is discussed, as well as the implication of the material behavior for practical applications of these materials.

36 MATERIALS SCIENCE↗

elm-diagnostics

elm-diagnostics is a Python package for computing diagnostic analyses and visualizations for the E3SM Land Model (ELM) component and is meant to support new feature development in ELM. The tool reads model history files and performs quantitative analyses including budget-closure checking, variable transformations, temporal aggregations, and statistical summaries to support model evaluation, validation, and scientific interpretation. The framework is designed for extensibility, with modular architecture enabling straightforward addition of new diagnostic methods, derived variables, analysis types, visualization approaches, and model-specific adaptations

Hoffman, Matt [Los Alamos National Laboratory]↗

Modal Field Reconstruction in Resonant Cavities in the Fundamental and Undermoded Frequency Regimes

Theory, simulations, and experiments are presented that demonstrate reconstruction of electromagnetic fields in a cavity from sparse probe measurements. Such techniques are often referred to as virtual sensing, allowing fields at unobserved locations to be predicted. These methods are appropriate for the fundamental and undermoded regimes, providing the ability to estimate fields (and shielding effectiveness) throughout an arbitrarily shaped cavity from a few judiciously spaced probes. A modal simulation method is implemented that allows the response of arbitrarily shaped cavities to be rapidly computed with respect to varying probe locations and slot parameters, enabling statistical analysis of probe placement on reconstruction performance. A cylindrical vessel with numerous probe holes is developed for experiments, referred to as Perforated Vessel 2 (PV2). Experiments are performed on the vessel with and without a steel box inside, where transmit power is delivered into the vessel either through probes (probe injection) or through slots using an external antenna (slot excitation). Simulations and experiments illustrate that when the number of probes is minimal (equal to the number of mode coefficients to be estimated at each frequency), probe placement is critical to avoid missed peaks and to have acceptable reconstruction error. Probe placement becomes less important as the number of probes is increased, but care is still required to avoid probe locations giving poor performance.

42 ENGINEERING↗

Mining Product Reviews for Important Product Features of Refurbished iPhones

Problem: Remanufacturers want to increase consumer interest in refurbished products, which motivates the need to understand which product features are important to buyers of refurbished products such as mobile phones. Research Questions: This study addresses two questions. First, which product features are most important for buyers of refurbished iPhones? Second, how do those preferences differ from the preferences of buyers of new iPhones? Methods: Online reviews of iPhones are obtained and converted into a document–term matrix. Using this text model, three subsets of features are identified using statistical analysis of frequency of mention: most frequent, average, and least frequent. A logistic regression (LR) model is then used to identify which features are most predictive of whether a review is for a new or refurbished phone. Results: Buyers of refurbished phones mention battery health, screen/display, shell condition, and brand significantly more often than other features. Directly contrasting reviews of refurbished versus new phones shows that shell condition, brand, speaker, and charger are found to be the most predictive product features indicated in reviews for refurbished phones. Of those, the shell condition is significantly more predictive than the others. Implications: The results identify product features that remanufacturers of iPhones can emphasize to increase customer demand.

Anisi, Atefeh↗

Integrated machine learning-molecular dynamics framework for electrolyte property prediction

Electrochemical stability windows determine the operating range of battery electrolytes, yet accurate prediction remains challenging because stability emerges from statistical ensembles of local solvation environments rather than single ground-state molecular structures. Traditional density functional theory calculations on energy-minimized clusters cannot capture the thermal variations in local coordination environments and geometries that govern decomposition, while SMILES-based machine learning methods lack explicit representation of three-dimensional solvation structure and ion pairing. Here, we introduce a structure-aware machine learning framework that predicts frontier orbital energies (HOMO and LUMO) directly from molecular dynamics-sampled solvation configurations, achieving sub-0.6 eV accuracy at computational costs 3–4 orders of magnitude lower than first-principles methods. Across twelve representative battery electrolytes, we demonstrate that solvent-separated and contact ion pairs exhibit strong size- and local chemistry dependent electronic stability, with variations in coordination shifts of HOMO or LUMO level by 2–3 eV, and that extended solvation structure and partially desolvated environment further modulate stability by up to 3 eV. By encoding the statistical nature of electrochemical failure through ensemble sampling of explicit solvation geometries, our approach enables high-throughput screening and rational design of next-generation battery electrolytes with mechanistic understanding of structure–property relationships.

Energy - Storage↗

A Computational Review of Privacy-Preserving Mechanisms for the Smart Grid

Smart grid technologies have rapidly become one of the largest and most comprehensive sources of data for the modern utility. For the most part, data streams are seen as an essential tool that enable utilities to carry their day-to-day business operations, but they also create the need for efficient and secure data management strategies. In the context of the smart grid, ensuring data privacy is becoming an increasing concern due to a combination of factors that range from shifts in operational paradigms and rapid technology evolution to changes in legislation. Furthermore, researchers have highlighted the risks associated with improperly protected energy records. For example, energy consumption data from homes could be used to infer the behaviors and habits of home occupants through activity recognition or user profiling (Fan, 2017), which may lead to unfair service pricing, targeted advertising, or other personal security violations. Similarly, Electric Vehicles’ (EVs) charging metadata could be used to reveal private information about the owner such as their payment methods, preferred charging stations, and other locational and timing information that could be used to reconstruct the vehicle owner’s behaviors. The privacy of user data, even when used for statistical analysis or machine learning training processes, also needs to be carefully considered, as an individual’s private traits may still be vulnerable if their inclusion/exclusion greatly impacts the result or could be linked to a public dataset through cross-reference. The breach of user privacy also has severe impacts for organizations that store, transmit, or work on the data in the form of diminishing the public’s trust in them while potentially incurring legal consequences (e.g., fines and suspensions under the European Union General Data Protection Regulation, Health Insurance Portability and Accountability Act, etc.). Because of these risks, several privacy-preserving mechanisms are available to help organizations comply with privacy legislations and prevent the unauthorized and malicious use of user data. In light of these concerns, this report focuses on performing a computational review of privacy-preserving mechanisms that have received a significant amount of interest in literature. It specifically focuses on 1) homomorphic encryption, 2) zero-knowledge proofs, 3) differential privacy, and 4) federated learning. It is worth noting that although many of the methods presented in this document rely on cryptographic primitives, their intent is not to provide perfect secrecy, but rather to enable users to maintain privacy, and thus they shall not be compared or equated to other constructs that are aimed to address cybersecurity constructs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

High-resolution hypernuclear decay pion spectroscopy at MAMI and future

Hypernuclear decay pion spectroscopy was established in 2012 at MAMI as a mass spectroscopy method for light hypernuclei. A monochromatic pion peak from $^4_Λ$H was successfully observed, and the Λ binding energy was determined to be B Λ = 2.157±0.005(stat.)±0.077(syst.) MeV in the 2014 run. In 2022, an upgrade experiment for $^3_Λ$H spectroscopy was conducted using a newly developed Li target. The absolute electron beam energy will be measured by the synchrotron radiation interferometry, which will be applied with the spectrometer calibration to improve the systematic error. The decay pion spectroscopy is planned to be performed at JLab Hall-C, which would significantly increase the statistics thanks to the higher energy beam, the better K + identification, and the faster data acquisition system. This experiment has been submitted as a Letter of Intent in JLab PAC51. The upgraded decay pion spectroscopy method is expected to provide new, accurate hypernuclear data, which will contribute to the advancement of our understanding of hypernuclear physics

47 OTHER INSTRUMENTATION↗

Mitigation of DESI fiber assignment incompleteness effect on two-point clustering with small angular scale truncated estimators

We present a method to mitigate the effects of fiber assignment incompleteness in two-point power spectrum and correlation function measurements from galaxy spectroscopic surveys, by truncating small angular scales from estimators. We derive the corresponding modified correlation function and power spectrum windows to account for the small angular scale truncation in the theory prediction. We validate this approach on simulations reproducing the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) with and without fiber assignment. We show that we recover unbiased cosmological constraints using small angular scale truncated estimators from simulations with fiber assignment incompleteness, with respect to standard estimators from complete simulations. Additionally, we present an approach to remove the sensitivity of the fits to high k modes in the theoretical power spectrum, by applying a transformation to the data vector and window matrix. We find that our method efficiently mitigates the effect of fiber assignment incompleteness in two-point correlation function and power spectrum measurements, at low computational cost and with little statistical loss.

79 ASTRONOMY AND ASTROPHYSICS↗

Deep inference of simulated strong lenses in ground-based surveys

The large number of strong lenses discoverable in future astronomical surveys will likely enhance the value of strong gravitational lensing as a cosmic probe of dark energy and dark matter. However, leveraging the increased statistical power of such large samples will require further development of automated lens modeling techniques. We show that deep learning and simulation-based inference (SBI) methods produce informative and reliable estimates of parameter posteriors for strong lensing systems in ground-based surveys. We present the examination and comparison of two approaches to lens parameter estimation for strong galaxy-galaxy lenses — Neural Posterior Estimation (NPE) and Bayesian Neural Networks (BNNs). We perform inference on 1-, 5-, and 12-parameter lens models for ground-based imaging data that mimics the Dark Energy Survey (DES). We find that NPE outperforms BNNs, producing posterior distributions that are more accurate, precise, and well-calibrated for most parameters. For the 12-parameter NPE model, the calibration is consistently within <10% of optimal calibration for all parameters, while the BNN is rarely within 20% of optimal calibration for any of the parameters. Similarly, residuals for most of the parameters are smaller (by up to an order of magnitude) with the NPE model than the BNN model. This work takes important steps in the systematic comparison of methods for different levels of model complexity.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Direct statistical simulation of the Lorenz96 system in model reduction approaches

Direct statistical simulation (DSS) of nonlinear dynamical systems bypasses the traditional route of accumulating statistics by lengthy direct numerical simulations by solving the equations that govern the statistics themselves. DSS suffers, however, from the curse of dimensionality as the statistics (such as correlations) generally have higher dimensions than the underlying dynamical variables. Here we investigate two approaches to reduce the dimensionality of DSS, illustrating each method with numerical experiments with the Lorenz96 dynamical system. The forms of DSS chosen here involve approximate closures at second and third order in the equal-time cumulants. We demonstrate significant reduction in computational effort that can be achieved without sacrificing the accuracy of DSS. The methods developed here can be applied to turbulent fluid and magnetohydrodynamical systems. Published by the American Physical Society 2025

Li, Kuan↗

Monte Carlo Simulations of Crystal Defects in Open Ensembles

Zero- and two-dimensional crystal defects form in open statistical ensembles, such as the grand canonical, that are usually inaccessible with conventional simulation techniques. This longstanding challenge is overcome with a new Hamiltonian Monte Carlo method that samples energy-biased gradual transformations. In conclusion, the method enables free energy calculations for nonideal point defects and the direct prediction of finite-temperature interface structures.

Grain boundaries↗

Characterizing the Uncertainty of Measurement of Traceable Isotope Ratios with Bayesian Statistical Techniques

Analytical techniques such as multicollector—inductively coupled plasma—mass spectrometry (MC-ICP-MS) are routinely employed at SRNL, other National Laboratories, and in academia to determine the precise isotopic composition of diverse natural and anthropogenic samples (e.g., rocks and nuclear materials). Quantifying and reporting uncertainty in such analyses, while regularly performed, have a rigorous statistical foundation. The Guide to the Expression of Uncertainty in Measurement 4 (GUM) outlines conventional techniques used to assess such uncertainty. As the accessibility and speed of statistical computing increase, there is a need to modernize conventional techniques. For example, Supplement 1 to the 3rd to the GUM suggests the use of approximation methods as an updated approach to the GUM.

McLarty, Ellis C.↗

Nonlinear Topological Photonics: Capturing Nonlinear Dynamics and Optical Thermodynamics

Combining multiple optical resonators or engineering dispersion of complex media has provided an effective method for demonstrating topological physics controlling photons in unprecedented ways such as unidirectional light propagation and spatially localized modes between an interface or on a corner. Further, adding nonlinear responses to those topological photonic systems has enabled achieving diverse phases of photons in both space and time, allowing for more functionalities in photonic devices that provide a new playground for studying dynamic features of nonlinear topological systems. However, most methods for describing nonlinear topological photonic systems rely on linear topological theories, making it challenging to accurately characterize the topology of nonlinear systems. Thus, substantial efforts have focused on rigorously describing nonlinear topological phases and developing effective tools to analyze nonlinear topological effects. Meanwhile, coupled multimode optical waveguides with nonlinear dynamic responses provide an excellent platform for the statistical description of photons, opening a new paradigm called “optical thermodynamics”. This review will introduce the basic concepts of nonlinear topological photonics and the recent development of theoretical approaches focusing on data-driven approaches for creating phase diagrams as well as the spectral localizer framework and the pseudospectrum method for understanding optical nonlinearities in topological systems. In addition, the new concept of optical thermodynamics will be introduced with some recent theoretical works.

Topological photonics↗

A chain stretch-based gradient-enhanced model for damage and fracture in elastomers

Similar to quasi-brittle materials, it has been recently shown that elastomers can exhibit a macroscopically diffuse damage zone that accompanies the fracture process. In this study, we introduce a stretch-based gradient-enhanced damage (GED) model that allows the fracture to localize and also captures the development of a physically diffuse damage zone. This capability contrasts with the paradigm of the phase field method for fracture, where a sharp crack is numerically approximated in a diffuse manner. Capturing fracture localization and diffuse damage in our approach is achieved by considering nonlocal effects that encompass network topology, heterogeneity, and imperfections. These considerations motivate the use of a statistical damage function dependent upon the nonlocal deformation state. From this model, fracture toughness is realized as an output. While GED models have been classically utilized for damage modeling of structural engineering materials (e.g., concrete), they face challenges when trying to capture the cascade from damage to fracture, often leading to damage zone broadening (de Borst and Verhoosel, 2016). This deficiency contributed to the popularity of the phase-field method over the GED model for elastomers and other quasi-brittle materials. Other groups have proceeded with damage-based GED formulations that prove identical to the phase-field method (Lorentz et al., 2012), but these inherit the aforementioned limitations. To address this issue in a thermodynamically consistent framework, we implement two modeling features (a nonlocal driving force bound and a simple relaxation function) specifically designed to capture the evolution of a physically meaningful damage field and the simultaneous localization of fracture, thereby overcoming a longstanding obstacle in the development of these nonlocal strain- or stretch-based approaches. Here, we discuss several numerical examples to understand the features of the approach at the limit of incompressibility, and compare them to the phase-field method as a benchmark for the macroscopic response and fracture energy predictions.

Elastomers↗

Dose-efficient automatic differentiation for ptychographic reconstruction

Ptychography, as a powerful lensless imaging method, has become a popular member of the coherent diffractive imaging family over decades of development. The ability to utilize low-dose X-rays and/or fast scans offers a big advantage in a ptychographic measurement (for example, when measuring radiation-sensitive samples), but results in low-photon statistics, making the subsequent phase retrieval challenging. Here, we demonstrate a dose-efficient automatic differentiation framework for ptychographic reconstruction (DAP) at low-photon statistics and low overlap ratio. As no reciprocal space constraint is required in this DAP framework, the framework, based on various forward models, shows superior performance under these conditions. It effectively suppresses potential artifacts in the reconstructed images, especially for the inherent periodic artifact in a raster scan. We validate the effectiveness and robustness of this method using both simulated and measured datasets.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗