Search NASASearch

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Operations on Graphical Models with Plates

This paper explains how graphical models, for instance Bayesian or Markov networks, can be extended to model problems in data analysis and learning. This provides a unified framework that combines lessons learned from the artificial intelligence, statistical and connectionist communities. This also offers a set of principles for developing a software generator for data analysis, whereby a learning or discovery system can be compiled from specifications. Many of the popular learning algorithms can be compiled in this way from graphical specifications. While in a sense this paper is a multidisciplinary review of learning, the main contribution here is the presentation of the material within the unifying framework of graphical models, and the observation that, as a result, the process of developing learning algorithms can be partly automated.

Buntine, Wray L.

Learning classification trees

Algorithms for learning classification trees have had successes in artificial intelligence and statistics over many years. How a tree learning algorithm can be derived from Bayesian decision theory is outlined. This introduces Bayesian techniques for splitting, smoothing, and tree averaging. The splitting rule turns out to be similar to Quinlan's information gain splitting rule, while smoothing and averaging replace pruning. Comparative experiments with reimplementations of a minimum encoding approach, Quinlan's C4 and Breiman et al. Cart show the full Bayesian algorithm is consistently as good, or more accurate than these other approaches though at a computational price.

Buntine, Wray

Statistical Engineering

This webinar provides an overview of the International Statistical Engineering Association (ISEA), and it illustrates the practice of statistical engineering at NASA. ISEA was formed to promote the study of how successful data-based problem-solving methods are leveraged to realize innovative opportunities and solve problems sustainably. ISEA is comprised of statisticians, engineers, scientists, and other professionals that exchange ideas and experiences in the development and application of statistical engineering theories. ISEA is building the body of knowledge of the statistical engineering discipline with a particular focus on improving academic preparation for tackling complex problems. Over the past 15 years, the practice of statistical engineering has gained recognition within NASA by spurring innovation and efficiency, and it has demonstrated significant impact. Aerospace research and development benefits from an application-focused statistical engineering perspective to accelerate learning, maximize knowledge, ensure strategic resource investment, and inform data-driven decisions. The second portion of this presentation provides an overview of infusing statistical engineering at NASA through pioneering case studies in aeronautics, space exploration, and atmospheric science.

Peter A Parker

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection

Graph Convolutional Network-Strengthened Topic Modeling for Scientific Papers

Machine learning has been woven into statistics to modernize topic modeling over textual documents written in natural language, and scientific paper search and recommendation can consequently offer higher accuracy instead of counting on traditional keyword-based search. However, topic distribution of a paper resulted from existing topic modeling techniques only relies on the statistics of words contained in the paper itself. We argue that community users’ views of a paper may also provide insights at the time of recommendation. For example, if a paper on fake image detection has been cited heavily by machine learning papers, such a feature should be absorbed in the embedding of this paper, so that it can be recommended for future query on machine learning. In this paper, we present a Graph Convolutional Network-strengthened Topic Modeling (GCN-TM) method, which employs GCN technique to refine topic modeling of scientific papers. A citation-oriented knowledge graph is constructed, and topic modeling is mapped to feature embedding of the comprising papers. On top of its own topics carried in its content, each paper learns topics from its neighbors and revise its embedding accordingly. Our empirical studies over real-life scientific literature has proved the necessity and effectiveness of our proposed approach.

Jia Zhang

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE

Diagnostics, Prognostics, and Optimization for Lithium-Ion Battery Systems

Health management of lithium-ion battery systems presents a host of challenges due to their complex physics, large numbers of components, and a wide variety of degradation behaviors across different battery types. Dr. Paul Gasper will present on research from the Electrochemical Energy Storage Group on Lithium-ion battery diagnostics, prognostics, and optimization. Diagnostics research, including state-estimation via machine-learning from electrochemical impedance spectroscopy and DC pulses as well as continuous state-estimation via Kalman filters, will highlight the ongoing challenges for accurately measuring the state of batteries without performing time-consuming characterization tests. NLR's industry-recognized battery prognostics work, which predicts real-world battery degradation by identifying degradation rate models from accelerated aging data using statistical modeling and machine-learning, will be used to demonstrate the critical impact of battery controls, thermal management, and operating strategy on durability and lifetime. Finally, the use of prognostic models for financial or lifetime optimization will be discussed.

25 ENERGY STORAGE

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry

Spatiotemporal Learning in Power Modules: Wavelet-Enhanced Forecasting of Thermomechanical Degradation

Detecting internal defects in power electronics packages is critical for their performance and reliability, especially under extreme operating conditions, as these defects can lead to catastrophic failure if not properly addressed. Confocal scanning acoustic microscopy (C-SAM) plays a key role in the nondestructive evaluation of bond layer degradation within a power electronics package by detecting defects such as delamination, voids, and cracks. However, accurately quantifying and predicting these defects from C-SAM images remains a significant challenge due to the low noise-to-signal ratio, which typically arises from both imaging process and bond patterns itself. In this paper, we explore machine learning strategies for processing C-SAM images and providing predictive models of defect growth. We use C-SAM images of sintered copper and sintered silver samples, which are obtained under accelerated thermal experiments, as the representative dataset for our study. We investigate the effect of Fourier transforms and wavelet transforms on these datasets to remove high-frequency noise and address noise across multiple scales with histogram equalization to enhance the contrast and improve the visibility of defects. As a result, defect boundaries can be clearly distinguished, enabling more accurate tracking of their growth over time. We then employ different time-series forecasting algorithms on the denoised images to formulate an image-based lifetime prediction model. Statistical models and deep-learning techniques are trained on images obtained in the early stages of thermal shock, and defect growth in the later stages is predicted. Our work serves as a preliminary attempt to improve the accuracy of lifetime prediction models of power electronics packages, which is critical under extreme operating environments.

24 POWER TRANSMISSION AND DISTRIBUTION

The Statistical Emulators of GGCMI Phase 2: Responses of Year-to-Year Variation of Crop Yield to CO 2 , Temperature, Water and Nitrogen Perturbations

Understanding the impact of climate change on year-to-year variation of crop yield is critical to global food stability and security. While crop model emulators are believed to be lightweight tools to replace the models, few emulators have been developed to capture such interannual variation of crop yield in response to climate variability. In this study, we developed a statistical emulator with a machine learning algorithm to reproduce the response of year-to-year variation of four crop yields to CO 2 (C), temperature (T), water (W), and nitrogen (N) perturbations defined in the Global Gridded Crop Model Intercomparison Project (GGCMI) phase 2. The emulators were able to explain more than 52 % of the variance of simulated yield and performed well in capturing the year-to-year variation of global average and gridded crop yield over current croplands in the baseline. With the changes in CO 2 –temperature–water–nitrogen (CTWN) perturbations, the emulators could reproduce the year-to-year variation of crop yield well over most current cropland. The variation of R and the mean absolute error was small under the single CTWN perturbations and dual-factor perturbations. These emulators thus provide statistical response surfaces of yield, including both its mean and interannual variability, to climate factors. They could facilitate spatiotemporal downscaling of crop model simulation, projecting the changes in crop yield variability in the future and serving as a lightweight tool for multi-model ensemble simulation. The emulators enhanced the flexibility of crop yield estimates and expanded the application of large-ensemble simulations of crop yield under climate change.

Climate change

Personalized Tucker Decomposition: Modeling Commonality and Peculiarity on Tensor Data

In this paper, we propose a personalized Tucker decomposition (perTucker) to address the limitations of traditional tensor decomposition methods in capturing heterogeneity across different datasets. perTucker decomposes tensor data into shared global components and personalized local components. We introduce an order orthogonality assumption and develop a proximal gradient regularized block coordinate descent algorithm guaranteed to converge to a stationary point. The unique and common representations learned by perTucker reveal intrinsic statistical patterns in data and provide valuable information for a wide range of downstream analytics, including anomaly detection, source classification, and clustering. We demonstrate perTucker’s effectiveness through a simulation study and two case studies on solar flare detection and tonnage signal classification.

14 SOLAR ENERGY

Curating a Standardized Dataset for Statistical Biosignature Classification

In recent years, machine learning has been explored as a toolkit for planetary science and operations [Helbert, Azari]. Machine learning has been used to improve our understanding of possible biosignatures and mineral signatures to improve science return on future missions [Warren-Rhodes, Cleaves].

Biosignatures

Code for the manuscript "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Mode

We disclose a python/pytorch implementation of the physics-informed machine learning algorithm described in "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Modeling", LA-UR-24-30678. Direct numerical simulation (DNS) of ubiquitous turbulence phenomena is computationally infeasible for realistic flows. As a result, reduced modeling for turbulent flows aim to reduce the number of resolved scales while retaining accurate representations of the small-scale physics. The dynamics of the velocity gradient tensor (VGT) is a key ingredient in reduced or subgrid turbulence models. The evolution equation for the VGT involves nonlocal terms, requiring closure modeling. This implementation of the novel methodology of Lagrangian Attention Tensor Networks (LATN), utilizes a structured representation of the history of the VGT to inform a physics-informed machine learning algorithm. This addition of structured memory terms is shown to outperform previous models when trained and evaluated on DNS data.

Livescu, Daniel [LANL]

Lack of clear standards and usable comparisons of downscaled climate projections pose a roadblock for US climate discovery and adaptation

Abstract The release of global climate projections coupled with the demand for local-resolution climate-forced meteorology has prompted many research groups to downscale these projections using various statistical, dynamical, and current machine learning techniques. Such downscaled datasets are being used to plan infrastructure and other community needs over the coming decades. Faced with roughly a dozen available US downscaled datasets, many practitioners ask, ‘What are the relevant differences between datasets?’ This work highlights the difficulty of comparing downscaled datasets and illustrates ways in which datasets differ even when using identical climate model input data. We show that substantial variability in precipitation projections arises from downscaling alone and that the downscaled dataset agreement varies depending on global climate projection. This analysis emphasizes the need for greater coordination and movement toward rigorous benchmarking of downscaling strategies within the downscaling research community, à la the land-modeling community, to better quantify downscaling dataset differences, strengths, and weaknesses for practitioners.

Hartke, Samantha H. (ORCID:0000000202394723)

Predicting fusion ignition at the National Ignition Facility with physics-informed deep learning

Here, an inertial confinement fusion experiment, carried out at the National Ignition Facility, has achieved ignition by generating fusion energy exceeding the laser energy that drove the experiment. Prior to the experiment, a generative machine learning model that combines radiation hydrodynamics simulations, deep learning, experimental data, and Bayesian statistics was used to predict, with a probability greater than 70%, that ignition was the most likely outcome for this shot.

Spears, Brian K. [Lawrence Livermore National Labo

Laser Powder Bed Fusion Microstructure Surrogate Model

SAND2025-11467O The Laser Powder Bed Fusion (LPBF) Microstructure Surrogate Model is a machine-learning-based tool. It predicts statistics of microstructures that are produced by the LPBF additive manufacturing process. It includes a series of codes for training, testing, and analyzing the model as well as utility scripts for handling data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Moser, Daniel [Sandia National Lab. (SNL-CA), Live

NASA DEVELOP’s Communication Strategy for Sharing Applied Science Outcomes

NASA’s DEVELOP Program conducts rapid feasibility projects that enable the future workforce and current decision makers to collaborate and build capacity to use Earth science data to enhance environmental management. The program communicates its results and applications to a broad audience through multiple methods: social media, website blog posts, print materials, short project videos, an “Ambassador Corps,” scientific and policy conference presentations, community engagement activities, end-of-project presentations, project “hand-offs,” traditional publications (both gray literature and peer-reviewed), and a project archive gallery. This presentation will describe the various methods employed by DEVELOP to communicate the program’s scientific outputs, target audiences, general statistics, community response, and lessons learned.

Capacity Building

CLAIRE: Enabling Heterogeneous Communication Network Optimization for Robust and Resilient Operations

In this paper, we present the capabilities of the CLAIRE System which provides resilient communications for NASA in presence of interference and congestion for a heterogeneous multi-vendor network. CLAIRE increases mission science data return to improve resource efficiencies and ensures resilience in the unpredictable space environment for NASA missions and communication networks. CLAIRE provides technology / waveform agnostic cognitive control plane that is instantiated at the Application Layer (APP) so that it can ride on NASA’s HDTN bundle protocol or any other protocol stack that is used by the network. The cognitive control plane is instantiated using Heartbeats (HTBTs). CLAIRE is assisted by Wideband UHF-Ka Band RF Sensing that leverages advances in the Direct Digital Transceiver (DDTRX) technology. The Wideband RF Sensing is driven by statistical signal processing and machine learning algorithms. Interference is mitigated using Dynamic Spectrum Access (DSA). Finally, CLAIRE addresses congestion using spectrum aware packet forwarding algorithm. CLAIRE provides an extensible protocol that allows passing of RF spectrum situational awareness, cross-layer sensing, delay tolerant networking and dynamic spectrum access information that can help with network optimization. Cross-Layer Sensing (CLS) and CLAIRE Decision Engine (CDE) enable spectrum and delay aware packet forwarding and Dynamic Spectrum Access during cases of severe interference.

cognitive communications